Risk control model automatic variable screening method and system based on marginal contribution
Through the automated variable screening method of risk control model based on marginal contribution, the problems of complex variable screening, insufficient model stability and low computational efficiency in the existing technology are solved, and a more efficient and stable credit risk management model is achieved.
Patent Information
- Application Number
- CN202411858210.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-17
- Publication Date
- 2025-05-06
AI Technical Summary
In the credit risk management of the prior art, variable screening is complex, model stability is insufficient, and computational efficiency is low, making it difficult to effectively process large-scale data and high-dimensional features.
A risk control model based on marginal contribution is used to automatically filter variables, and the WoE value, IV value and PSI value are initially screened, low variance characteristics and high correlation characteristics are eliminated, and the feature set is optimized using the marginal contribution algorithm until the variable screening is stable.
It improves the accuracy and stability of model prediction, simplifies model complexity, improves computing efficiency, enhances feature utilization efficiency, and ensures the consistency and reliability of credit decisions.
Smart Images

Figure CN119939202A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of financial risk control technology, and in particular to a method and system for automatic variable screening of a risk control model based on marginal contribution. Background Art
[0002] In credit risk management, the scorecard model is a commonly used tool to assess the borrower's default risk. Traditional scorecard models often rely on manual experience and statistical methods, such as univariate analysis and stepwise regression, in the process of variable screening and model construction. Although these methods are effective, they face many challenges when dealing with large-scale data and high-dimensional features:
[0003] High complexity in feature screening: With the diversification of data sources and the increase in data volume, the variable screening process has become complex and time-consuming. Traditional methods find it difficult to accurately characterize the contribution of variables to the model, and therefore it is difficult to screen effective variables. Insufficient model stability: In the actual application of the scorecard model, model stability between different time periods and data sets is an important consideration. Traditional methods often fail to fully consider the long-term stability of variables when selecting features, resulting in large fluctuations in model performance. Low computational efficiency: Traditional variable screening methods have low computational efficiency when processing large-scale data, especially when the feature dimension is high, and the time cost of model training and optimization is high.
[0004] Therefore, there is an urgent need for an automated variable screening method and system for risk control models based on marginal contribution. Summary of the invention
[0005] The present invention provides a method and system for automatic variable screening of a risk control model based on marginal contribution to solve the above-mentioned problems existing in the prior art.
[0006] In order to achieve the above object, the present invention provides the following technical solutions:
[0007] A method for automatically screening variables in a risk control model based on marginal contribution, characterized by comprising:
[0008] S101: acquiring raw data based on several data sources, extracting continuous variables after cleaning the raw data, binning the continuous variables, and calculating the variable weight value WoE of each bin;
[0009] S102: Based on the WoE value, identify and delete low variance features and high correlation features, and then preliminarily screen variable features by calculating the information content IV and stability index PSI of the remaining features to obtain a preliminarily screened feature set;
[0010] S103: using the initially screened feature set, selecting a certain loss function, training an initial scorecard model, and obtaining a trained scorecard model;
[0011] S104: Based on the marginal contributions in model MC of the variables in the trained scorecard model, remove the variables whose marginal contributions are lower than the set threshold, add the variables whose marginal contributions Out Model MC of the external variables are higher than the preset threshold, obtain the updated feature set, use the updated feature set to train the risk control model and calculate the marginal contribution;
[0012] S105: Repeat steps S101 to S104 to screen continuous variables until the continuous variable screening is stable.
[0013] Wherein, step S101 includes:
[0014] S1011: obtaining raw data through a number of data sources, the number of data sources including credit reports, customer application forms, and transaction records;
[0015] S1012: Perform outlier and missing value operations on the original data, obtain preprocessed data, extract continuous variables in the preprocessed data, and perform binning operations on the continuous variables;
[0016] S1013: Calculate the WoE value of each bin, and quantify the impact of the bin on the probability of default through the WoE value.
[0017] Wherein, step S102 includes:
[0018] S1021: Perform variance analysis on the WoE value, identify low variance features, calculate the correlation matrix between features based on the WoE value, identify high correlation features, delete low variance features and high correlation features, and retain the remaining features;
[0019] S1022: Calculate the IV value for each remaining feature, and calculate the PSI value of each remaining feature in different time windows or different sample sets;
[0020] S1023: Based on the IV value and the PSI value, remove the features whose IV value is lower than the preset threshold, remove the features whose PSI value is higher than the preset threshold, use the remaining feature set after removing the low IV value and the high PSI value as the feature set after preliminary screening, and output the feature set after preliminary screening.
[0021] Wherein, step S103 includes:
[0022] S1031: Obtain a feature set after preliminary screening, which includes income, age, and credit card usage rate, and select a corresponding loss function based on the model objective and data characteristics;
[0023] S1032: Construct an initial scoring card model, adopt a linear model structure, use the initially screened features as input to the initial scoring card model, apply the selected loss function to the model training process, optimize the model parameters, and adjust the model parameters through iterative training.
[0024] Wherein, step S104 includes:
[0025] S1041: Set the coefficient of a certain input variable to 0, calculate the model loss function value, and subtract the newly obtained model loss function value from the original model loss function value to obtain the in model MC of the input variable;
[0026] S1042: Use a variable not included in the model and the credit score predicted by the existing scorecard model as two-dimensional features to train the risk control model, calculate the new risk control model loss function value, use the difference between the newly obtained model loss function value and the original model loss function value as the Out Model MC of the variable not included in the model, and add variables whose Out Model MC is higher than the set threshold;
[0027] S1043: Use the updated feature set to train the risk control model, recalculate the updated In Model MC and OutModel MC, and repeat the process of removing low-contribution model variables and adding high-contribution non-model variables until no variables can be removed or added, or the preset maximum number of iterations is reached.
[0028] Wherein, after step S105, the following steps are included:
[0029] Perform K-fold cross-validation on the initial model, calculate the model performance index of each fold of cross-validation, and calculate the average value of the risk control model performance index of all folds as the cross-validation performance evaluation result of the risk control model. Based on the cross-validation performance evaluation result, determine whether the risk control model parameters need to be further optimized.
[0030] Among them, add variables whose Out Model MC is higher than the set threshold, including:
[0031] Obtain the target person’s credit score;
[0032] Identify the unmodeled variables and use them together with the target person’s credit score as two-dimensional features;
[0033] Train a new risk control model based on two-dimensional features;
[0034] Calculate the loss function value of the new risk control model;
[0035] Obtain the loss function value of the original risk control model;
[0036] Calculate the difference between the newly acquired model loss function value and the original model loss function value as the Out Model MC of the unmodeled variable;
[0037] Add variables whose Out Model MC is higher than the set threshold to the risk control model;
[0038] Credit scoring and risk assessment are performed based on the updated risk control model.
[0039] Among them, calculating the loss function value of the new risk control model includes:
[0040] Define the loss function, including the error term and the regularization term;
[0041] Calculate the error of the new risk control model on the training set;
[0042] Calculate regularization terms to prevent model overfitting;
[0043] Add the error term and the regularization term to get the total loss function value.
[0044] Among them, an automated variable screening system for risk control models based on marginal contribution includes:
[0045] A data processing unit is used to obtain raw data based on several data sources, extract continuous variables after cleaning the raw data, bin the continuous variables, and calculate the variable weight value WoE of each bin;
[0046] The preliminary feature screening unit is used to identify and delete low variance features and high correlation features based on the WoE value, and then preliminarily screen the variable features by calculating the information content IV and stability index PSI of the remaining features to obtain the feature set after preliminary screening;
[0047] The model training unit is used to use the feature set after preliminary screening, select a certain loss function, train the initial scorecard model, and obtain the trained scorecard model;
[0048] The feature screening and optimization unit is used to remove variables whose marginal contributions are lower than the set threshold based on the marginal contributions in modelMC of the variables in the trained scorecard model, add variables whose marginal contributions Out Model MC of the external variables are higher than the preset threshold, obtain the updated feature set, use the updated feature set to train the risk control model and calculate the marginal contributions, and repeat the continuous variable screening process until the continuous variable screening is stable;
[0049] The model optimization and verification unit is used to evaluate and optimize the performance of the risk control model through cross-validation and parameter tuning.
[0050] The data processing unit includes:
[0051] A data collection subunit, for obtaining raw data from a number of data sources, including credit reports, customer application forms, and transaction records;
[0052] The data cleaning subunit is used to process missing values and outliers in the original data and obtain preprocessed data;
[0053] The feature engineering subunit is used to extract continuous variables from the preprocessed data, perform binning operations on the continuous variables, calculate the WoE value of each bin, and quantify the impact of binning on the probability of default through the WoE value.
[0054] Compared with the prior art, the present invention has the following advantages:
[0055] An automated variable screening method for a risk control model based on marginal contribution includes: obtaining raw data based on several data sources, extracting continuous variables after cleaning the raw data, binning the continuous variables, and calculating the variable weight value WoE of each bin; based on the WoE value, identifying and deleting low variance features and high correlation features, and then preliminarily screening variable features by calculating the information volume IV and stability index PSI of the remaining features to obtain a feature set after preliminary screening; using the feature set after preliminary screening, selecting a certain loss function, training an initial scorecard model, and obtaining a trained scorecard model; based on the marginal contribution in model MC of the variables in the trained scorecard model, eliminating variables whose marginal contribution is lower than a set threshold, adding variables whose marginal contribution Out Model MC of the external variables is higher than a preset threshold, obtaining an updated feature set, and using the updated feature set to train the risk control model and calculate the marginal contribution. The out model MC algorithm provided by the method only needs to fit the scorecard of two-dimensional features instead of using all the input model variables to fit the scorecard, which greatly improves the calculation efficiency, thereby increasing the landing value of the method in the big data credit risk control scenario. Furthermore, by screening out high-gain features that have not been included in the model, the model's ability to identify credit risks can be enhanced and the efficiency of feature utilization can be improved.
[0056] Other features and advantages of the present invention will be set forth in the description which follows, and in part will be apparent from the description, or may be learned by practice of the present invention.
[0057] The technical solution of the present invention is further described in detail below through the accompanying drawings and embodiments. BRIEF DESCRIPTION OF THE DRAWINGS
[0058] The accompanying drawings are used to provide a further understanding of the present invention and constitute a part of the specification. Together with the embodiments of the present invention, they are used to explain the present invention and do not constitute a limitation of the present invention. In the accompanying drawings:
[0059] Figure 1 It is a flow chart of a method for automatic variable screening of a risk control model based on marginal contribution in an embodiment of the present invention;
[0060] Figure 2 This is a flow chart of calculating the variable weight value WoE of each bin in an embodiment of the present invention;
[0061] Figure 3 This is a flow chart of obtaining a feature set after preliminary screening in an embodiment of the present invention. DETAILED DESCRIPTION
[0062] The preferred embodiments of the present invention are described below in conjunction with the accompanying drawings. It should be understood that the preferred embodiments described herein are only used to illustrate and explain the present invention, and are not used to limit the present invention.
[0063] The embodiment of the present invention provides a method for automatically screening variables of a risk control model based on marginal contribution, comprising:
[0064] S101: acquiring raw data based on several data sources, extracting continuous variables after cleaning the raw data, binning the continuous variables, and calculating the variable weight value WoE of each bin;
[0065] S102: Based on the WoE value, identify and delete low variance features and high correlation features, and then preliminarily screen variable features by calculating the information content IV and stability index PSI of the remaining features to obtain a preliminarily screened feature set;
[0066] S103: using the initially screened feature set, selecting a certain loss function, training an initial scorecard model, and obtaining a trained scorecard model;
[0067] S104: Based on the marginal contributions in model MC of the variables in the trained scorecard model, remove the variables whose marginal contributions are lower than the set threshold, add the variables whose marginal contributions Out Model MC of the external variables are higher than the preset threshold, obtain the updated feature set, use the updated feature set to train the risk control model and calculate the marginal contribution;
[0068] S105: Repeat steps S101 to S104 to screen continuous variables until the continuous variable screening is stable.
[0069] The working principle of the above technical solution is as follows: Data preparation. Handle missing values and outliers, and bin continuous variables (such as WoE bins). Calculate the WoE value of each bin to ensure that the variable can be effectively used in the scorecard model.
[0070] Preliminary feature screening: Remove low variance features and high correlation features to avoid redundant information. Use IV and PSI to preliminarily screen features to ensure that the model only contains important and stable features.
[0071] Initial model training: Use the initially screened feature set, select a loss function, and train the initial scorecard model.
[0072] Feature screening and optimization. Eliminate variables with marginal contributions below the threshold based on in-model MC. Add variables with out-model MC above the threshold. Retrain the model with the updated feature set and calculate the marginal contribution. Repeat the above screening process until the variable screening is stable.
[0073] Model optimization and validation. Use cross-validation to evaluate model performance after each iteration to ensure model stability and generalization. Adjust model parameters through grid search or Bayesian optimization to further improve model performance.
[0074] The beneficial effects of the above technical solution are: Improve the prediction accuracy of the model: By removing low-contribution features and adding high-contribution features, ensure that the model only contains the most important variables, thereby improving the prediction accuracy of the model and enhancing the ability of the scorecard model to identify high-risk groups. Enhance model stability: Use PSI to screen variables to ensure the stability of the model in different time periods and data sets, reduce model performance fluctuations, and ensure the consistency and reliability of credit decisions. Simplify model complexity: By removing redundant and low-contribution features, the complexity of the model is simplified, the interpretability and ease of use of the model are improved, and it is easier for the credit risk management department to understand and use it. Improve model efficiency: Through iterative optimization and feature screening, the model training and prediction process is more efficient, reducing the consumption of computing resources and time, and speeding up the credit approval process. Ensure model compliance: Through a systematic variable screening process, the model construction process is transparent and traceable, meets regulatory requirements in the field of credit risk control, and helps meet regulatory compliance checks.
[0075] In another embodiment, step S101 includes:
[0076] S1011: obtaining raw data through a number of data sources, the number of data sources including credit reports, customer application forms, and transaction records;
[0077] S1012: Perform outlier and missing value operations on the original data, obtain preprocessed data, extract continuous variables in the preprocessed data, and perform binning operations on the continuous variables;
[0078] S1013: Calculate the WoE value of each bin, and quantify the impact of the bin on the probability of default through the WoE value.
[0079] The calculation of the WoE value for each bin includes:
[0080] Extract continuous variables from preprocessed data and define them as a set of candidate variables;
[0081] Perform a binning operation on each candidate variable in the candidate variable set, divide the value range of each candidate variable into multiple binning intervals, and obtain a binning set;
[0082] Count the number of default samples and non-default samples in each bin interval, and calculate the probability of default and non-default in each bin;
[0083] According to the default probability and non-default probability of each sub-box, the WoE value of each sub-box is calculated. The calculation formula of the WoE value is:
[0084]
[0085] Based on the calculated WoE value of each bin, the impact of each bin on the probability of default is quantified to form a sequence of WoE values of candidate variables;
[0086] The WoE value sequence is tested to determine whether there are abnormal fluctuations or monotonicity problems. If so, the bins are adjusted and the WoE value is recalculated;
[0087] Apply the optimized WoE value to the scorecard model to evaluate the predictive effect of the candidate variables in the model and ensure that the candidate variables are effectively used in the scorecard model.
[0088] The working principle of the above technical solution is as follows: data usually comes from multiple sources, such as credit reports, customer application forms, transaction records, etc. The data needs to be cleaned, normalized and binned for subsequent scorecard modeling. For example, collect customer credit reports from different credit agencies and unify the format. This ensures the consistency and accuracy of the data and improves the quality of model training.
[0089] Handle missing values, outliers, and bin continuous variables (such as WoE binning). For example, fill missing income values with the median and correct abnormally high transaction amounts. Enhance the stability of the model by reducing noise and anomalies in the data.
[0090] Calculate the WoE value for each bin to quantify the impact of binning on the probability of default and ensure that the variable can be effectively used in the scorecard model. For example, divide customer income into three intervals: less than 5,000 yuan, 5,000-10,000 yuan, and more than 10,000 yuan, and calculate the WoE value for each. Quantify the impact of binning on default risk through WoE values and improve the interpretability of the model.
[0091] The beneficial effects of the above technical solution are: through multi-source data fusion, strict data preprocessing, scientific variable binning and WoE conversion, a solid foundation is laid for building a high-quality credit scorecard model. It not only improves the predictive ability and stability of the model, but also enhances the interpretability and practicality of the model. This is of great significance for financial institutions to conduct risk assessment, formulate credit strategies and optimize customer management, and ultimately help improve the accuracy and efficiency of credit decisions and reduce financial risks.
[0092] In actual implementation, binning is a data preprocessing technique, including equal-width binning, equal-frequency binning, binning based on business logic, using decision trees, and clustering binning. In the implementation of the above schemes, converting continuous variables into discrete bins may lose detailed information in the data, especially when the bins are wide. This information loss may affect the predictive ability of the model.
[0093] Therefore, in the actual implementation process, this application provides a binning auxiliary method, including:
[0094] Get the starting and ending variables of the continuous variables in the preprocessed data to generate variable text data; in this process, it is necessary to collect raw data containing continuous variables, such as credit reports, customer application forms, transaction records, etc. Annotate the dataset, which contains continuous variables and corresponding labels (for example, default / non-default).
[0095] Then, the variable text data is cleaned to handle outliers and missing values;
[0096] Then, use the preset decision tree to determine the binning boundaries of the continuous variables; ensure that the binning process is based on the labeled data set to reflect the relationship between the continuous variables and the labels to prevent missing data. Analyze the importance of each continuous variable in the model to determine which variables are most critical for predicting the label.
[0097] Then, based on feature importance analysis, optimize the binning boundaries to ensure that the binning can better reflect the predictive power of the variable. Calculate the WOE value of each bin to quantify the impact of the binning on the label. Based on the binning results and WOE values, select the bins with the greatest impact on label prediction as key variables. The binning boundaries of these key variables are used as "keyword sets".
[0098] Finally, a response dialogue is generated based on the binning information of key variables.
[0099] In actual implementation, supervised binning uses labeled data sets to guide the binning process. Common supervised learning methods, such as decision trees or random forests, can determine the optimal binning boundaries based on the relationship between data features and target labels. The best split points are found by training models (such as decision trees), which are the boundaries of the bins. WOE is a statistic that measures the relationship between variable bins and target variables (such as default probability). The bins that have the greatest impact on model predictions are selected as keywords, which can represent important features of the data. Use keyword sets and starting sentences to construct dialogue replies so that the reply content is closely related to the user's questions or needs, and always issue alarms for variable anomalies.
[0100] In another embodiment, step S102 includes:
[0101] S1021: Perform variance analysis on the WoE value, identify low variance features, calculate the correlation matrix between features based on the WoE value, identify high correlation features, delete low variance features and high correlation features, and retain the remaining features;
[0102] S1022: Calculate the IV value for each remaining feature, and calculate the PSI value of each remaining feature in different time windows or different sample sets;
[0103] S1023: Based on the IV value and the PSI value, remove the features whose IV value is lower than the preset threshold, remove the features whose PSI value is higher than the preset threshold, use the remaining feature set after removing the low IV value and the high PSI value as the feature set after preliminary screening, and output the feature set after preliminary screening.
[0104] Among them, the output of the initially screened feature set includes:
[0105] Perform variance analysis on the WoE value of each feature to obtain the variance of the feature;
[0106] Based on the variance size, determine whether the variance of each feature is lower than the preset threshold;
[0107] If the variance of a feature is lower than a preset threshold, the feature is deleted to avoid redundant information;
[0108] For the remaining features, the correlation matrix between the features is calculated based on the WoE value to determine the correlation coefficient between the features;
[0109] Determine whether the correlation coefficient between the features is higher than a preset threshold;
[0110] If the correlation coefficient is higher than the preset threshold, one feature is deleted from the high-correlation feature pair to reduce multicollinearity;
[0111] By removing low variance features and high correlation features, redundant information and multicollinearity are avoided.
[0112] For each feature in the remaining feature set after removing low variance features and high correlation features, the IV value is calculated to determine the IV value of each feature. The IV value calculation formula is as follows:
[0113]
[0114] Among them, P i and Q i are the percentages of a certain feature in a certain category, and n is the number of feature categories;
[0115] Determine whether the IV value of each feature is lower than a preset IV value threshold (e.g., 0.1);
[0116] If the IV value is lower than the preset IV value threshold, the feature will be eliminated because the current feature contributes less to predicting default risk;
[0117] For the remaining features, calculate the PSI value of each feature in different time windows or different sample sets to determine the PSI value of each feature;
[0118] Determine whether the PSI value of each feature is higher than a preset PSI value threshold;
[0119] If the PSI value is higher than the preset PSI value threshold, the feature is removed because its data distribution varies greatly in different time periods, resulting in unstable model performance;
[0120] The remaining feature set after removing low IV values and high PSI values is used as the feature set after preliminary screening, and the feature set after preliminary screening is output.
[0121] The working principle of the above technical solution is to avoid redundant information and multicollinearity by removing low-variance features and high-correlation features. For example, remove the feature that all customers have the same marital status, and keep the features with low correlation between income and savings amount. This can reduce the complexity of the model and improve the stability and performance of the model.
[0122] Features with IV values lower than 0.1 are removed because they contribute less to predicting default risk; features with PSI values higher than 0.25 are removed because their data distribution varies greatly in different time periods, which may lead to unstable model performance.
[0123] The beneficial effects of the above technical solution are: through multi-dimensional analysis and screening, the feature set is effectively optimized. It not only improves the predictive ability and stability of the model, but also enhances the interpretability and practicality of the model. This is of great significance for financial institutions to conduct accurate risk assessment, formulate effective credit strategies and optimize resource allocation. Ultimately, this process helps to improve the accuracy and efficiency of credit decisions, reduce financial risks, and provide customers with more fair and reasonable credit services.
[0124] In another embodiment, step S103 includes:
[0125] S1031: Obtain a feature set after preliminary screening, which includes income, age, and credit card usage rate, and select a corresponding loss function based on the model objective and data characteristics;
[0126] S1032: Construct an initial scoring card model, adopt a linear model structure, use the initially screened features as input to the initial scoring card model, apply the selected loss function to the model training process, optimize the model parameters, and adjust the model parameters through iterative training.
[0127] Among them, the linear model structure includes multiple linear combination units, the features after preliminary screening include preprocessed feature variables, the selected loss function is used to evaluate the model prediction error, and the iterative training includes multiple rounds of parameter updating processes.
[0128] The working principle of the above technical solution is: the initially screened feature set should have a high correlation with the target variable of credit risk. For example, income: the level of income can reflect the borrower's repayment ability. The higher the income, the lower the user's probability of default may be; age: according to data analysis, the probability of default of users in certain age groups may be low, while the probability of default of other specific age groups is high. Age can be used as an important factor in predicting the probability of default; credit card utilization rate: the credit card utilization rate refers to the ratio of the amount used by the user on the credit card to the credit limit of the credit card. The higher the utilization rate, the more likely the user is to have a heavier debt and a higher risk of default.
[0129] Once the feature selection is completed, we build a scorecard model based on these feature sets. Scorecard models in the financial field usually use linear models because they have strong explanatory properties and can clearly clarify the contribution of each feature to the credit score.
[0130] Use linear models: We use linear models to establish the relationship between features and target variables (such as "whether to default"). In the linear model, the weights of each feature (i.e. the scores in the scorecard) are set, and the user's risk score is calculated based on the values and weights of these features. The higher the score, the higher the risk. Choice of loss function: During the model training process, choosing an appropriate loss function is an important step to improve the model's performance. For example: Logarithmic loss function: The logarithmic loss function is commonly used in binary classification scenarios and can quantify the model's error in predicting default and non-default classifications. The model minimizes the cross-entropy loss to make the predicted default probability closer to the true value.
[0131] For the binary classification problem, the loss function formula is as follows:
[0132] Loss=-[ylog(p)+(1-y)log(1-p)]
[0133] Where y is the actual label (0 or 1) and p is the default probability predicted by the model.
[0134] If our model predicts that the probability of default of a customer is 0.8, but the actual situation is that the customer has not defaulted (the label is 0), then the loss will be larger, and vice versa.
[0135] In scenarios where certain features (such as credit card usage rate) have a strong ability to identify defaults, the Divergence loss function can more effectively perform linear discrimination on the relationship between features and default probability, thereby optimizing the model. Using this loss function can more clearly distinguish between high and low distributions of default probability, making it easier for the model to identify high-risk users in the scorecard.
[0136] Iterative parameter optimization: During the training of the scorecard model, we will repeatedly adjust the model parameters (i.e., the weights of each feature) and continuously reduce the value of the loss function through iterative optimization. The final parameter set makes the model perform best on the validation set, thus forming the final scorecard model.
[0137] Suppose we find in the training set data that "credit card usage rate" has a significant positive correlation with default risk, while "income" is negatively correlated with default risk. In the final scorecard, users with high credit card usage rates will be given higher risk scores, while users with higher incomes will have their scores lowered accordingly.
[0138] By applying different loss functions, we can make the model pay more attention to the identification of high-risk users. For example, using the Divergence loss function can more significantly distinguish high-risk and low-risk users, ensuring that the scorecard has a higher ability to identify high-risk users in actual use.
[0139] The beneficial effects of the above technical solution are: through preliminary screening, irrelevant or redundant features are eliminated, the model training time is reduced, the model complexity is reduced, and the model prediction accuracy is improved; the number of features is reduced, making the model easier to understand and interpret, and facilitating the analysis of the impact of features on the target variable; the number of features is reduced, and the cost of data storage and processing is reduced; the linear model has a simple structure, is easy to train and understand, and can quickly build an initial scoring card model to provide a basis for subsequent model optimization; the initial model can provide preliminary evaluation indicators, such as AUC, KS, etc., to provide a reference for subsequent model optimization; by observing the model parameters, the degree of influence of the features on the target variable can be preliminarily judged, providing guidance for subsequent feature engineering; through feature screening and linear model construction, a basic scoring card model can be quickly established, and preliminary evaluation indicators and feature importance information can be provided; the initial model can be used as the starting point for subsequent model optimization, for example, you can try to add nonlinear features, use more complex model structures, etc., to further improve model performance.
[0140] In another embodiment, step S104 includes:
[0141] S1041: Set the coefficient of a certain input variable to 0, calculate the model loss function value, and subtract the newly obtained model loss function value from the original model loss function value to obtain the in model MC of the input variable;
[0142] S1042: Use a variable not included in the model and the credit score predicted by the existing scorecard model as two-dimensional features to train the risk control model, calculate the new risk control model loss function value, use the difference between the newly obtained model loss function value and the original model loss function value as the Out Model MC of the variable not included in the model, and add variables whose Out Model MC is higher than the set threshold;
[0143] S1043: Use the updated feature set to train the risk control model, recalculate the updated In Model MC and OutModel MC, and repeat the process of removing low-contribution model variables and adding high-contribution non-model variables until no variables can be removed or added, or the preset maximum number of iterations is reached.
[0144] The working principle of the above technical solution is: according to the in model MC, the variables whose marginal contribution is lower than the set threshold are eliminated. In the linear model, the specific algorithm of the In model MC is:
[0145] 1. Set the coefficient of a certain input variable to 0;
[0146] 2. Calculate the model loss function value according to the formula;
[0147] 3. Subtract the newly obtained model loss function value from the original model loss function value to obtain the inmodel MC of the input variable.
[0148] Add variables whose out model MC is higher than the set threshold. In the linear model, the specific algorithm of out model MC is:
[0149] 1. Use a certain unmodeled variable together with the credit score predicted by the existing scorecard model as two-dimensional features to train a scorecard model.
[0150] 2. For this new scorecard, recalculate the model loss function value.
[0151] 3. Subtract the newly obtained model loss function value from the original model loss function value to obtain the out model MC.
[0152] Retrain the scorecard model using the updated feature set, recalculate the updated in model MC and outmodel MC, and repeat the process of removing low-contribution in-model variables and adding high-contribution out-of-model variables until no variables can be removed or added, or the preset maximum number of iterations is reached. Repeat feature removal and addition until the model performance is stable.
[0153] The beneficial effects of the above technical solution are as follows: by setting the variable coefficient to 0 and observing the change of the model loss function, the contribution of the variable to the model prediction ability can be quantified; if the In Model MC of a variable is low, it means that the variable contributes less to the model and may be a redundant variable, which can be considered to be eliminated; by eliminating low-contribution variables, the model structure can be simplified and the model efficiency and interpretability can be improved. By training the risk control model with the unmodeled variables and the credit score predicted by the existing scorecard model, the potential contribution of the variable to the model prediction ability can be evaluated; if the Out Model MC of a unmodeled variable is high, it means that the variable may contain important information, and it can be considered to be added to the scorecard model; by adding high-contribution unmodeled variables, the prediction ability of the scorecard model can be expanded and the accuracy of the model can be improved. By continuously iteratively eliminating low-contribution modeled variables and adding high-contribution unmodeled variables, the optimal feature set can be automatically selected without manual intervention; by optimizing the feature set, the prediction ability of the scorecard model can be improved and the model risk can be reduced; through iterative optimization, the model can be prevented from over-relying on certain specific features and the stability and generalization ability of the model can be improved.
[0154] In another embodiment, the step S105 includes:
[0155] Perform K-fold cross-validation on the initial model, calculate the model performance index of each fold of cross-validation, and calculate the average value of the risk control model performance index of all folds as the cross-validation performance evaluation result of the risk control model. Based on the cross-validation performance evaluation result, determine whether the risk control model parameters need to be further optimized.
[0156] The working principle of the above technical solution is: use the cross-validation method to evaluate the model performance after each iteration to ensure the stability and generalization ability of the model. Adjust the model parameters through grid search or Bayesian optimization to further improve the model performance.
[0157] The model performance was evaluated through 10-fold cross validation to ensure that the model performed consistently on different datasets and had good generalization capabilities.
[0158] Adjust model parameters through grid search or Bayesian optimization to further improve model performance. For example, adjust the regularization parameter λ to optimize model performance.
[0159] Based on the cross-validation performance evaluation result, determining whether it is necessary to further optimize the risk control model parameters includes: comparing the first indicator to be evaluated with the second indicator to be evaluated based on the cross-validation technology to obtain a first performance score;
[0160] If the first performance score is greater than or equal to the preset performance threshold, the execution result of the target evaluation task is that the evaluation passes, otherwise the evaluation fails;
[0161] If the execution results of the preset number of target evaluation tasks in the target evaluation task set are all evaluated as passed, the parameter optimization phase will be entered, otherwise it will return to the model iteration phase;
[0162] In the parameter optimization stage, a preset parameter search space is constructed, including multiple parameters to be optimized and the corresponding value ranges;
[0163] Perform iterative search in the parameter search space based on grid search or Bayesian optimization algorithm, and evaluate model performance after each iteration;
[0164] Record the parameter combination and corresponding performance score of each iteration, and build a parameter-performance mapping relationship;
[0165] Based on the parameter-performance mapping relationship, the optimal parameter combination is selected using a preset optimization strategy;
[0166] Apply the optimal parameter combination to the model to obtain the optimized model;
[0167] Repeat the above evaluation task for the optimized model. If the evaluation result is better than that before optimization, the optimized model is adopted; otherwise, the original model remains unchanged.
[0168] The above steps are executed repeatedly until the preset number of iterations is reached or the performance improvement is less than the preset threshold.
[0169] The beneficial effects of the above technical solution are: performing K-fold cross-validation on the initial model can effectively evaluate the generalization ability of the model, determine whether the model parameters need to be further optimized, select the best model parameters, and improve the reliability and stability of the model.
[0170] In another embodiment, adding variables whose Out Model MC is higher than a set threshold includes:
[0171] Obtain the target person’s credit score;
[0172] Identify the unmodeled variables and use them together with the target person’s credit score as two-dimensional features;
[0173] Train a new risk control model based on two-dimensional features;
[0174] Calculate the loss function value of the new risk control model;
[0175] Obtain the loss function value of the original risk control model;
[0176] Calculate the difference between the newly acquired model loss function value and the original model loss function value as the Out Model MC of the unmodeled variable;
[0177] Add variables whose Out Model MC is higher than the set threshold to the risk control model;
[0178] Credit scoring and risk assessment are performed based on the updated risk control model.
[0179] The working principle of the above technical solution is as follows: first, obtain the credit score of the target person from the existing data set, and this score is calculated based on the existing risk control model; identify variables that have not yet been included in the current risk control model from the data set, which may be newly introduced data or features that have not been considered before; combine the unmodeled variables with the credit score of the target person to form a new two-dimensional feature, and this feature combination can be used to train a new risk control model; use these new two-dimensional features to train a new risk control model. The purpose of this step is to evaluate whether the unmodeled variables can improve the prediction performance of the model; calculate the loss function value of the new model on the test set, which reflects the prediction error or performance of the model; calculate the loss function value of the original risk control model on the same test set as a reference; compare the loss function values of the new model and the original model, and calculate the difference between the two. This difference is used to evaluate the contribution of the unmodeled variables to the model improvement, which is called Out Model MC (model contribution); set a threshold, and only when the Out Model MC of the unmodeled variable is higher than this threshold, the variable is included in the risk control model. If the contribution of the variable is high enough, it means that the variable helps to improve the performance of the model. Use updated risk control models for credit scoring and risk assessment to ensure that the new models can predict risks more accurately.
[0180] The beneficial effects of the above technical solution are: by introducing variables with high Out Model MC, the predictive performance of the risk control model can be significantly improved. This method helps to discover potentially useful variables, thereby improving the accuracy and reliability of the model. Continuously evaluate and add new variables to ensure that the model can adapt to the latest data and features, and enhance the adaptability and foresight of the model. The newly added variables may provide additional information, which helps to better identify and assess risks, thereby optimizing credit scoring and risk management strategies.
[0181] In another embodiment, calculating a new risk control model loss function value includes:
[0182] Define the loss function, including the error term and the regularization term;
[0183] Calculate the error of the new risk control model on the training set;
[0184] Calculate regularization terms to prevent model overfitting;
[0185] Add the error term and the regularization term to get the total loss function value.
[0186] The working principle of the above technical solution is as follows: the loss function of the risk control model includes an error term and a regularization term. The error term is used to measure the deviation between the predicted value of the model and the true value, and usually uses metrics such as mean square error (MSE) and absolute error (MAE). The regularization term is used to control the complexity of the model and prevent overfitting. Commonly used regularization methods include L1 regularization (Lasso) and L2 regularization (Ridge).
[0187] Apply the new risk control model to the training data set and calculate the model error. The error value reflects the model's ability to fit the training data and is usually calculated as the difference between the predicted value and the actual value.
[0188] The complexity of the model is evaluated by calculating the regularization term. Common regularization forms include L1 regularization and L2 regularization. L1 regularization (Lasso) penalizes the absolute value sum of the variables in the model, making some coefficients zero, thereby performing feature selection. L2 regularization (Ridge) penalizes the sum of the squares of the variables to limit the size of the coefficients, thereby preventing the model from being too complex.
[0189] The total loss function value is the sum of the error term and the regularization term, which comprehensively reflects the prediction performance and complexity of the model. This value is used to evaluate the overall performance of the model, taking into account both the accuracy of the model and the complexity of the model.
[0190] The beneficial effects of the above technical solution are: by defining a clear loss function and calculating the total loss, the performance of the model on the training set can be objectively evaluated, including prediction accuracy and complexity, which helps to select and optimize the most suitable model. By introducing regularization terms to limit the complexity of the model and effectively prevent overfitting, regularization can help the model generalize to unseen data, thereby improving the stability and reliability of the model in practical applications. The calculation of the error term reflects the predictive ability of the model. By optimizing the loss function value, the model's fitting effect on the training data can be improved, thereby improving the model's prediction accuracy. During the training process, by adjusting the weights of the error term and the regularization term, the model's parameters can be optimized, the optimal model configuration can be found, and the overall performance of the model can be enhanced.
[0191] In another embodiment, a risk control model automatic variable screening system based on marginal contribution includes:
[0192] A data processing unit is used to obtain raw data based on several data sources, extract continuous variables after cleaning the raw data, bin the continuous variables, and calculate the variable weight value WoE of each bin;
[0193] The preliminary feature screening unit is used to identify and delete low variance features and high correlation features based on the WoE value, and then preliminarily screen the variable features by calculating the information content IV and stability index PSI of the remaining features to obtain the feature set after preliminary screening;
[0194] The model training unit is used to use the feature set after preliminary screening, select a certain loss function, train the initial scorecard model, and obtain the trained scorecard model;
[0195] The feature screening and optimization unit is used to remove variables whose marginal contributions are lower than the set threshold based on the marginal contributions in modelMC of the variables in the trained scorecard model, add variables whose marginal contributions Out Model MC of the external variables are higher than the preset threshold, obtain the updated feature set, use the updated feature set to train the risk control model and calculate the marginal contributions, and repeat the continuous variable screening process until the continuous variable screening is stable;
[0196] The model optimization and verification unit is used to evaluate and optimize the performance of the risk control model through cross-validation and parameter tuning.
[0197] The working principle of the above technical solution is as follows: Data processing unit: responsible for collecting data from various data sources, cleaning data, and performing feature engineering processing to ensure data quality and model effectiveness. It includes data collection subunit, data cleaning subunit and feature engineering subunit.
[0198] Data collection subunit: Collect data from multiple data sources such as credit reports, customer application forms, transaction records, etc. to ensure the comprehensiveness of the data. For example, collect customer credit reports from different credit agencies and unify the format. This ensures the consistency and accuracy of the data and improves the quality of model training.
[0199] Data cleaning subunit: handles missing values and outliers, unifies data formats, and improves data quality. For example, missing income values are filled with medians, and abnormally high transaction amounts are corrected. By reducing noise and anomalies in the data, the stability of the model is enhanced.
[0200] Feature Engineering Subunit: Transform and extract features from data, perform WoE binning on continuous variables, and calculate the WoE value for each bin. For example, divide customer income into three intervals: less than 5,000 yuan, 5,000-10,000 yuan, and more than 10,000 yuan, and calculate their respective WoE values. The WoE value is used to quantify the impact of binning on default risk and improve the interpretability of the model.
[0201] Preliminary feature screening unit: By analyzing and screening features, redundant and invalid features are removed to ensure that the model contains useful and stable features. It includes variance analysis subunit, correlation analysis subunit, and IV and PSI analysis subunit.
[0202] Variance analysis subunit: By analyzing the variance of features, features with low variance are eliminated to avoid the impact of invalid information on the model. For example, the feature that all customers have the same marital status is deleted. This can reduce the complexity of the model and improve the stability and performance of the model.
[0203] Correlation analysis subunit: Use the Pearson correlation coefficient to filter features, retain features that are highly correlated with the target variable, and reduce redundant information. For example, calculate the correlation coefficient between each feature and the target variable, and filter out features with high correlation for model training. Ensure that the model only contains features that have a significant impact on the prediction results, and improve the prediction accuracy of the model.
[0204] IV and PSI analysis subunit: Calculate the IV and PSI values of each feature, remove features with low IV values and high PSI values, and ensure the accuracy and stability of the model. For example, calculate the IV value of each feature and remove features with low IV values. Calculate the PSI value of each feature and remove features with high PSI values. In the credit risk control business, using IV preliminary screening variables can ensure that the model contains features with strong predictive power and improve the accuracy of the model; using PSI preliminary screening variables can ensure that the model contains stable features and improve the stability of the model in different time periods and data sets.
[0205] Model training unit: used for initial model training and marginal contribution calculation, identifying and quantifying the contribution of important features. It includes initial model training subunit and marginal contribution calculation subunit.
[0206] Initial model training subunit: Use the initially screened feature set to train the scorecard. Use the initially screened features as the full set and combine them with the initial model input variables specified by the user to train the model. For example, use user-specified features such as income, age, and credit card usage rate to train the initial model. This provides a basis for subsequent marginal contribution calculations.
[0207] The loss function selects the subunit:
[0208] Provide users with different loss function settings, such as logarithmic loss function, Divergence, Hinge loss function, etc. Users can select a suitable loss function for scorecard training based on the distribution of good and bad samples and business experience. The choice of loss function will affect the calculation of marginal contribution.
[0209] Feature screening and optimization unit: By removing and adding features, the scoring card model is continuously optimized to ensure the stability and accuracy of the model performance. It includes feature removal subunit, feature addition subunit and iterative optimization subunit.
[0210] Feature elimination subunit: Calculate the in model MC and eliminate features whose marginal contribution is lower than the set threshold according to the in model MC. For example, it is found that among the initial features provided by the user, the "number of customer address changes" has a low model gain, so it is eliminated. The advantage of doing this is that removing features with small model gain helps to simplify the model and improve the credit score stability of the scorecard.
[0211] Feature adding subunit: Calculate out model MC, and add features with marginal contributions higher than the set threshold according to out model MC to enhance the model's predictive ability. For example, during the screening process, it may be found that the gain of "credit card usage years" is high, so this feature is added to the scorecard. By introducing new high-gain features, the scorecard's ability to characterize and identify credit risk is enhanced.
[0212] Iterative optimization subunit: Repeat the process of feature removal and addition, retrain the model with the updated feature set, and calculate the in model MC and out model MC until the model performance is stable. This subunit ensures that the model achieves optimal performance and improves the stability and prediction accuracy of the model through repeated iterative optimization.
[0213] Model optimization and verification unit: Evaluate and optimize model performance through cross-validation and parameter tuning to ensure the stability and generalization ability of the model.
[0214] Cross-validation subunit: Use cross-validation to evaluate model performance and ensure that the model has good stability and generalization ability. For example, evaluate model performance through 10-fold cross-validation to ensure that the model performs consistently on different data sets and has good generalization ability.
[0215] Parameter tuning subunit: Adjust model parameters through grid search or Bayesian optimization to further improve model performance. For example, adjust the regularization parameter λ to optimize model performance. Further improve the model's predictive ability and stability to ensure the best performance of the model.
[0216] The beneficial effects of the above technical solution are: by deleting low-contribution features and adding high-contribution features, it is ensured that the model only contains the most important variables, thereby improving the prediction accuracy of the model and enhancing the ability of the scorecard model to identify high-risk groups. The use of PSI to screen variables ensures the stability of the model in different time periods and data sets, reduces model performance fluctuations, and ensures the consistency and reliability of credit decisions. By eliminating redundant and low-contribution features, the complexity of the model is simplified, the interpretability and ease of use of the model are improved, and it is easier for the credit risk management department to understand and use it. Through iterative optimization and feature screening, the model training and prediction process is more efficient, reducing the consumption of computing resources and time, and speeding up the credit approval process. Through a systematic variable screening process, the model building process is transparent and traceable, meets regulatory requirements in the field of credit risk control, and helps meet regulatory compliance checks.
[0217] In another embodiment, the data processing unit comprises:
[0218] A data collection subunit, for obtaining raw data from a number of data sources, including credit reports, customer application forms, and transaction records;
[0219] The data cleaning subunit is used to process missing values and outliers in the original data and obtain preprocessed data;
[0220] The feature engineering subunit is used to extract continuous variables from the preprocessed data, perform binning operations on the continuous variables, calculate the WoE value of each bin, and quantify the impact of binning on the probability of default through the WoE value.
[0221] The working principle of the above technical solution is as follows: Data collection subunit: collects data from multiple data sources such as credit reports, customer application forms, transaction records, etc. to ensure the comprehensiveness of the data. For example, collect customer credit reports from different credit institutions and unify the format. This ensures the consistency and accuracy of the data and improves the quality of model training.
[0222] Data cleaning subunit: handles missing values and outliers, unifies data formats, and improves data quality. For example, missing income values are filled with medians, and abnormally high transaction amounts are corrected. By reducing noise and anomalies in the data, the stability of the model is enhanced.
[0223] Feature Engineering Subunit: Transform and extract features from data, perform WoE binning on continuous variables, and calculate the WoE value for each bin. For example, divide customer income into three intervals: less than 5,000 yuan, 5,000-10,000 yuan, and more than 10,000 yuan, and calculate their respective WoE values. The WoE value is used to quantify the impact of binning on default risk and improve the interpretability of the model.
[0224] The beneficial effect of the above technical solution is: by deleting low-contribution features and adding high-contribution features, it is ensured that the model only contains the most important variables, thereby improving the prediction accuracy of the model and enhancing the scoring card model's ability to identify high-risk groups.
[0225] Obviously, those skilled in the art can make various changes and modifications to the present invention without departing from the spirit and scope of the present invention. Thus, if these modifications and variations of the present invention fall within the scope of the claims of the present invention and their equivalents, the present invention is also intended to include these modifications and variations.
Claims
1. A method for automatically screening variables in a risk control model based on marginal contribution, characterized in that: include: S101: acquiring raw data based on several data sources, extracting continuous variables after cleaning the raw data, binning the continuous variables, and calculating the variable weight value WoE of each bin; S102: Based on the WoE value, identify and delete low variance features and high correlation features, and then preliminarily screen variable features by calculating the information content IV and stability index PSI of the remaining features to obtain a preliminarily screened feature set; S103: using the initially screened feature set, selecting a certain loss function, training an initial scorecard model, and obtaining a trained scorecard model; S104: Based on the marginal contributions in modelMC of the internal variables of the trained scorecard model, remove the variables whose marginal contributions are lower than the set threshold, add the variables whose marginal contributions Out ModelMC of the external variables are higher than the preset threshold, obtain the updated feature set, use the updated feature set to train the risk control model and calculate the marginal contribution; S105: Repeat steps S101 to S104 to screen continuous variables until the continuous variable screening is stable.
2. The method for automatically screening variables in a risk control model based on marginal contribution according to claim 1, characterized in that: Step S101 includes: S1011: obtaining raw data from a number of data sources, including credit reports, customer application forms, and transaction records; S1012: Perform outlier and missing value operations on the original data, obtain preprocessed data, extract continuous variables in the preprocessed data, and perform binning operations on the continuous variables; S1013: Calculate the WoE value of each bin, and quantify the impact of the bin on the probability of default through the WoE value.
3. The method for automatic variable screening of risk control models based on marginal contribution according to claim 1, characterized in that: Step S102 includes: S1021: Perform variance analysis on the WoE value, identify low variance features, calculate the correlation matrix between features based on the WoE value, identify high correlation features, delete low variance features and high correlation features, and retain the remaining features; S1022: Calculate the IV value for each remaining feature, and calculate the PSI value of each remaining feature in different time windows or different sample sets; S1023: Based on the IV value and the PSI value, remove the features whose IV value is lower than the preset threshold, remove the features whose PSI value is higher than the preset threshold, use the remaining feature set after removing the low IV value and the high PSI value as the feature set after preliminary screening, and output the feature set after preliminary screening.
4. The method for automatic variable screening of a risk control model based on marginal contribution according to claim 1, characterized in that: Step S103 includes: S1031: Obtain a feature set after preliminary screening, which includes income, age, and credit card usage rate, and select a corresponding loss function based on the model objective and data characteristics; S1032: Construct an initial scoring card model, adopt a linear model structure, use the initially screened features as input to the initial scoring card model, apply the selected loss function to the model training process, optimize the model parameters, and adjust the model parameters through iterative training.
5. The method for automatic variable screening of risk control models based on marginal contribution according to claim 1 is characterized in that: Step S104 includes: S1041: Set the coefficient of a certain input variable to 0, calculate the model loss function value, and subtract the newly obtained model loss function value from the original model loss function value to obtain the in model MC of the input variable; S1042: Use a variable not included in the model and the credit score predicted by the existing scorecard model as two-dimensional features to train the risk control model, calculate the new risk control model loss function value, use the difference between the newly obtained model loss function value and the original model loss function value as the Out Model MC of the variable not included in the model, and add variables whose Out Model MC is higher than the set threshold; S1043: Use the updated feature set to train the risk control model, recalculate the updated In Model MC and Out Model MC, and repeat the process of eliminating low-contribution variables and adding high-contribution non-model variables until no variables can be eliminated or added, or the preset maximum number of iterations is reached.
6. The method for automatic variable screening of a risk control model based on marginal contribution according to claim 1, characterized in that: The following steps are included after step S105: Perform K-fold cross-validation on the initial model, calculate the model performance index of each fold of cross-validation, and calculate the average value of the risk control model performance index of all folds as the cross-validation performance evaluation result of the risk control model. Based on the cross-validation performance evaluation result, determine whether the risk control model parameters need to be further optimized.
7. The method for automatic variable screening of a risk control model based on marginal contribution according to claim 5, characterized in that: Add Out Model MC variables above the set threshold, including: Obtain the target person’s credit score; Identify the unmodeled variables and use them together with the target person’s credit score as two-dimensional features; Train a new risk control model based on two-dimensional features; Calculate the loss function value of the new risk control model; Obtain the loss function value of the original risk control model; Calculate the difference between the newly acquired model loss function value and the original model loss function value as the OutModel MC of the unmodeled variable; Add variables whose Out Model MC is higher than the set threshold to the risk control model; Credit scoring and risk assessment are performed based on the updated risk control model.
8. The method for automatic variable screening of risk control models based on marginal contribution according to claim 5 is characterized in that: Calculate the new risk control model loss function value, including: Define the loss function, including the error term and the regularization term; Calculate the error of the new risk control model on the training set; Calculate regularization terms to prevent model overfitting; Add the error term and the regularization term to get the total loss function value.
9. An automated variable screening system for risk control models based on marginal contribution, characterized in that: include: A data processing unit is used to obtain raw data based on several data sources, extract continuous variables after cleaning the raw data, bin the continuous variables, and calculate the variable weight value WoE of each bin; The preliminary feature screening unit is used to identify and delete low variance features and high correlation features based on the WoE value, and then preliminarily screen the variable features by calculating the information content IV and stability index PSI of the remaining features to obtain the feature set after preliminary screening; The model training unit is used to use the feature set after preliminary screening, select a certain loss function, train the initial scorecard model, and obtain the trained scorecard model; The feature screening and optimization unit is used to remove variables whose marginal contributions are lower than the set threshold based on the marginal contributions in model MC of the variables in the trained scorecard model, add variables whose marginal contributions Out ModelMC of the external variables are higher than the preset threshold, obtain the updated feature set, use the updated feature set to train the risk control model and calculate the marginal contributions, and repeat the continuous variable screening process until the continuous variable screening is stable; The model optimization and verification unit is used to evaluate and optimize the performance of the risk control model through cross-validation and parameter tuning.
10. The automated variable screening system for risk control models based on marginal contribution according to claim 9, characterized in that: The data processing unit includes: A data collection subunit, for obtaining raw data from a number of data sources, including credit reports, customer application forms, and transaction records; The data cleaning subunit is used to process missing values and outliers in the original data and obtain pre-processed data; The feature engineering subunit is used to extract continuous variables from the preprocessed data, perform binning operations on the continuous variables, calculate the WoE value of each bin, and quantify the impact of binning on the probability of default through the WoE value.
Citation Information
Cited By
Pipe network leakage risk evaluation and identification method and system and storage medium
CN122310283A