A risk control variable intelligent mining and identification method based on causal inference and feature importance fusion

By combining causal structure learning and dual machine learning, a causal directed acyclic graph and a comprehensive scoring system are constructed, which solves the problems of pseudo-correlation redundancy and insufficient cross-time generalization in risk control variable mining. This enables automated screening and efficient interpretability of risk control variables, and improves the interpretability and generalization ability of the risk control model.

CN122434646APending Publication Date: 2026-07-21CHINA NAT BUILDING MATERIALS TECH CO LTD +1
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
CHINA NAT BUILDING MATERIALS TECH CO LTD
Filing Date
2026-04-23
Publication Date
2026-07-21

AI Technical Summary

Technical Problem

Existing risk control variable mining techniques suffer from problems such as prominent spurious correlation redundancy, lack of causal logic, insufficient cross-time generalization, and high dependence of screening thresholds on human experience, resulting in insufficient interpretability and generalization ability of risk control models.

Method used

A constrained causal structure learning algorithm is used to construct a directed acyclic graph. Combined with dual machine learning and propensity score matching methods, unbiased causal effect estimation and significance testing are performed. Through adaptive threshold screening and causal redundancy elimination, a comprehensive scoring system integrating causal effects and stability-weighted feature importance is constructed to achieve automated and unbiased variable screening.

Benefits of technology

Accurately identifying backdoor paths and confusing variables, eliminating spurious redundancy, improves the causal explanatory power of risk control variables and their generalization ability across business scenarios, reduces labor costs and technical barriers, and meets financial regulatory compliance requirements.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122434646A_ABST
    Figure CN122434646A_ABST
Patent Text Reader

Abstract

The present application relates to the field of financial risk control and artificial intelligence, in particular to a risk control variable intelligent mining and identification method based on causal inference and feature importance fusion, comprising the following steps: first, collecting credit risk control sample data for preprocessing, and dividing the training set and cross-time OOT validation set; constructing a causal DAG through business constraint-based causal structure learning, identifying a backdoor path confounder set, calculating the modified conditional average treatment effect of each candidate variable and completing significance test; calculating the stability weighted feature importance of the variable; constructing a fusion score model to calculate the comprehensive mining score, and outputting the final interpretable risk control variable set through adaptive threshold screening, causal redundancy elimination and multi-dimensional generalization verification. The present application takes into account the causal significance, predictive discrimination and cross-time stability of the variable, improves the interpretability and generalization ability of the risk control variable, adapts to the full business scenario of credit risk control, and meets the regulatory compliance requirements.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the fields of financial risk control and artificial intelligence technology, specifically to a method for intelligent mining and identification of risk control variables based on the fusion of causal inference and feature importance. Background Technology

[0002] With the deepening digital transformation of credit operations, the predictive accuracy, generalization stability, and business interpretability of risk control models have become core indicators for credit risk management in financial institutions. The discovery and selection of risk control variables is a crucial step in the entire risk control modeling process, directly determining the final performance and compliance of the model. Currently, traditional risk control variable selection methods are mostly based on association analysis frameworks, using indicators such as Pearson correlation coefficient, information value (IV), and variance inflation factor (VIF) for initial variable screening, and then combining these with the feature importance of tree models to complete the final variable selection. These methods can only identify the relationship between variables and risk control objectives, but cannot distinguish between true causal effects and spurious correlations. They easily select a large number of redundant variables that only have coincidental data connections with the target variable. This not only increases model complexity but also leads to insufficient model interpretability, making it difficult to meet the rigid requirements of financial regulators for the interpretability of risk control models.

[0003] To address the issue of insufficient variable interpretability, the industry has gradually introduced causal inference techniques for feature selection. Among them, CN120596876A discloses a multi-path feature selection method based on causal inference. The core technical solution of this method is as follows: First, target data is collected to form a dataset and preprocessed to identify feature variables in the dataset. A causal graph corresponding to the preprocessed dataset is constructed, and initial feature screening is performed based on the first correlation coefficient between the feature variables and the target variable in the causal graph to obtain candidate feature variables. Then, multiple feature selection paths, consisting of filtering paths and feature importance evaluation paths, are used to further filter the candidate feature variables, obtaining a subset of feature variables corresponding to each path. Finally, the frequency of each selected feature is calculated and normalized, and the final input feature variables are determined based on the normalization results. This scheme improves feature interpretability through causal graph construction and balances feature relevance and predictive contribution through multi-path filtering, providing a feasible approach for the application of causal inference in the field of feature selection.

[0004] However, the aforementioned comparative documents and existing similar technologies still have several technical shortcomings: First, the causal effect assessment in the comparative documents is based solely on correlation coefficients and partial correlation coefficients, without performing unbiased quantitative calculations of the causal effect between the variable and the target variable. Especially in risk control scenarios with high-dimensional confounding variables, it is impossible to eliminate the causal effect estimation bias caused by overfitting, making it difficult to accurately distinguish between the true causal effect and spurious correlation of variables, resulting in insufficient accuracy and reliability of causal screening. Second, the feature importance assessment in the comparative documents is based solely on the contribution of the tree model on a single training set, without considering the feature stability issues caused by cross-time sample distribution shifts in credit risk control scenarios. The selected variables are prone to performance degradation in cross-cycle business scenarios, and their generalization ability cannot meet the requirements of long-term, high-robust credit risk control. Third, the comparison documents use frequency voting to fuse multi-path routing results, failing to achieve deep quantitative fusion of causal effects and feature importance. Furthermore, the variable selection thresholds are highly dependent on manual experience, making adaptive unbiased selection impossible, thus limiting automation and selection accuracy. Fourth, the comparison documents do not design suitable causal effect calculation and significance testing processes for variables of different data types widely present in risk control scenarios, such as continuous, binary, and ordered multi-class variables. This results in poor adaptability for causal identification of discrete risk control variables. Additionally, the lack of redundancy elimination and multi-dimensional systematic verification based on causal logic leads to collinear redundancy in the selected variable set, failing to simultaneously meet the multiple requirements of risk control models for causal interpretability, predictive discrimination, and cross-time stability. Summary of the Invention

[0005] The purpose of this invention is to provide a method for intelligent mining and identification of risk control variables based on the fusion of causal inference and feature importance, so as to solve the problems mentioned in the background art of existing risk control variable mining technology, such as prominent spurious correlation redundancy, lack of causal logic, insufficient cross-time generalization, and high dependence of screening threshold on human experience.

[0006] To achieve the above objectives, the present invention provides the following technical solution: A method for intelligent mining and identification of risk control variables based on the fusion of causal inference and feature importance includes the following steps: S1: Collect full sample data from credit risk control scenarios, complete data preprocessing, divide the data into a training set and a cross-time OOT validation set, and determine the risk control target variable. With the candidate variable set; S2: Based on candidate variable set and target variable A constrained causal structure learning algorithm is used to construct a directed acyclic graph (DAG) to identify backdoor path confusion variable sets. Calculate the modified conditional average treatment effect for each candidate variable. And complete the significance test of causal effect; S3: Based on the training set and the OOT validation set, an ensemble learning model with monotonic constraints is used to calculate the stability-weighted feature importance of each candidate variable. ; S4: Based on standardization and A fusion scoring model was constructed to calculate the comprehensive mining score of each candidate variable. After sorting by score in descending order, an initial set of risk control variables is obtained through adaptive threshold filtering; S5: Perform causal redundancy removal and generalization verification on the initial set of risk control variables, and output the final interpretable set of risk control variables.

[0007] Preferably, step S2 is for continuous candidate variables. Corrected conditional average treatment effect The calculation formula is: ; in, The total number of samples in the training set. As the target variable for risk control, Candidate variables The preset unit marginal increment, To obfuscate the variable set for backdoor paths based on causal DAG identification, For the conditional expectation operator, For the first Candidate variables corresponding to each sample The value of , For the first The value vector of the set of confounding variables corresponding to each sample.

[0008] Preferably, in step S2, the constrained causal structure learning algorithm employs an improved PC algorithm, and the preset causal constraints include: risk control target variables. A variable can only be a child node of a causal DAG and cannot be the parent node of any node. There must be no directed edges between derived variables in the same dimension. In the time dimension, a variable that occurs earlier can only be the parent node of a variable that occurs later.

[0009] Preferably, in step S3, the importance of the stability-weighted feature... The calculation formula is: ; in, Candidate variables Feature importance calculated based on tree model split gain on the training set. To determine the importance of this variable as a feature of the same dimension on the OOT validation set, The preset stability penalty coefficient, To prevent positive minimum values ​​where the denominator is 0.

[0010] Preferably, the calculation formula for the fusion scoring model in step S4 is: ; in, , For all candidate variables The sample mean and standard deviation, , For all candidate variables The sample mean and standard deviation, for The corresponding t-test p-value, At the preset significance level, This is an indicator function that takes the value when the condition within the parentheses is met. Otherwise, the value is .

[0011] Preferably, step S2 employs a dual machine learning DML framework for estimation. The expected condition includes the following sub-steps: S21: Divide the training set samples into... Constructing a series of non-overlapping folds In folded cross-validation, samples are grouped so that each folded sample serves as a validation subset, and the remaining samples are... Fold as a training subset; S22: For each training subset, obfuscate the variable set. As input features, with risk control target variables To fit the labels, fit the first set of residual prediction models and output the target variable residuals for the corresponding validation subset. ; S23: For the same training subset, use a confounding variable set As input features, with candidate variables To fit the labels, a second set of residual prediction models is fitted, outputting the candidate variable residuals for the corresponding validation subset. ; S24: Target variable residuals based on the full sample Residuals of candidate variables By fitting a linear regression model, the unbiased conditional average treatment effect is obtained, thus eliminating the overfitting bias caused by high-dimensional confounding variables.

[0012] Preferably, in step S4, the method for determining the adaptive threshold specifically includes the following sub-steps: S41: Comprehensive Mining and Scoring of All Candidate Variables Sort the scores in descending order of numerical value to generate an ordered score sequence. ,in Let be the total number of candidate variables, and satisfy . ; S42: Calculate the first-order difference sequence of the ordered score sequence. ,in , The range of values ​​is arrive ; S43: Calculate the absolute value of each element in the difference sequence. Iterate through the sequence to find the position of the element with the largest absolute value difference. ; S44: In an ordered scoring sequence Location-related score As an adaptive screening threshold, candidate variables with scores greater than or equal to the threshold in the ordered scoring sequence are retained to form the initial risk control variable set.

[0013] Preferably, the causal redundancy removal in step S5 specifically includes the following sub-steps: S511: Based on the causal DAG constructed in step S2, identify the risk control target variable. A Markov blanket containing a target variable. The set of parent nodes, the set of child nodes, and the set of other parent nodes of the child nodes; S512: Iterate through all variables in the initial risk control variable set and remove those not in the target variable. Within the Markov blanket range of variables, a subset of variables with causal redundancy is obtained; S513: Perform a multicollinearity test on the subset of variables after removing causal redundancy, and calculate the variance inflation factor for each variable. value; S514: According to Values ​​are removed in descending order of importance. Variables whose values ​​are greater than a preset threshold, until all remaining variables... All values ​​are less than or equal to the preset threshold, thus completing the collinearity redundancy removal.

[0014] Preferably, step S5, the generalization verification specifically includes the following sub-steps: S521: Based on the cross-time OOT validation set, calculate the population stability index of each variable in the variable subset after redundancy removal. The The calculation interval is a preset equal-frequency binning interval, and the preset number of bins is [number missing]. to box; S522: Removal Variables whose values ​​are greater than a preset stability threshold are selected as a subset of variables that meet the stability criteria. S523: For the subset of variables that meet stability criteria, calculate the correlation between each variable and the risk control target variable based on the OOT validation set. Rank correlation coefficient, discrimination Value, remove Variables whose values ​​are less than the preset discrimination threshold are selected as a subset of variables that meet the discrimination criteria. S524: Input the subset of variables that meet the discrimination criteria into the pre-set logistic regression benchmark model, and calculate the default prediction of each univariate on the OOT validation set. Value, remove For variables whose values ​​are less than the preset prediction performance threshold, the generalization verification is completed.

[0015] Preferably, step S2 targets discrete candidate variables. This includes binary and ordinal categorical variables, and the adjusted conditional average treatment effect. The calculation process specifically includes the following sub-steps: S25 Discrete candidate variables Value encoding is performed, and binary classification variables are encoded using... - Dumb variable coding, for ordered multi-category variables, uses adjacent category coding to determine the values ​​for the treatment group. Values ​​compared to the control group ; S26 Backdoor Path Confusion Variable Set Based on Causal DAG Recognition The propensity score matching (PSM) method was used to match a predetermined number of control group samples to each treatment group sample. The matching caliper value was preset to be 1 / 3 of the propensity score standard deviation. times; S27 Based on the matched sample set, calculate candidate variables The modified conditional average treatment effect is calculated using the following formula: ; in, This is the set of samples for the matched processing group. To determine the number of samples in the processing group, This is the matched control group sample set. This represents the sample size of the control group. , These represent the values ​​of the risk control target variables for the corresponding samples; S28 For the calculated The bootstrap method is used for repeated sampling inspection, with the number of repeated samplings preset to [number]. Next, the standard error for calculating causal effects and the two-tailed test. The value is used to complete the significance test of causal effect.

[0016] Compared with the prior art, the beneficial effects of the present invention are: 1. This invention constructs a causal directed acyclic graph (DAG) through causal structure learning with business constraints, accurately identifying confounding variables related to backdoor paths. Combining dual machine learning and propensity score matching methods, it performs unbiased causal effect estimation and significance testing for different types of variables, clarifying the true path of variable impact on risk control objectives from a causal relationship perspective. It fundamentally eliminates redundant variables with only spurious correlations to the target variable. Furthermore, this invention performs causal redundancy removal based on a Markov blanket of the target variable, ensuring minimal information redundancy and maximum causal explanatory power in the selected variable set. This solves the core technical problem of insufficient interpretability of traditional risk control variables and aligns with regulatory compliance requirements in the financial risk control field.

[0017] 2. This invention constructs a comprehensive scoring system that integrates causal effects and stability-weighted feature importance. By introducing a stability penalty term based on the difference in feature importance across time-series OOT validation sets, it quantifies the stability of variables over time. Simultaneously, through standardized fusion and adaptive threshold screening, it achieves automated and unbiased variable selection, balancing causal significance and predictive discriminative ability. This approach effectively reduces the performance degradation problem after variable deployment caused by overfitting and sample distribution shifts in traditional methods, significantly improving the generalization ability and robustness of risk control variables in cross-cycle business scenarios.

[0018] 3. This invention designs appropriate causal effect calculation and significance testing processes for candidate variables of different data types, such as continuous, binary, and ordered multi-class classification, covering all business scenarios of credit risk control, including pre-loan access, mid-loan risk monitoring, and post-loan collection early warning. Simultaneously, this invention ensures the statistical validity and business applicability of the output variables through a multi-dimensional closed-loop verification system including multicollinearity testing, population stability testing, discrimination testing, and predictive performance testing. The method can be directly embedded into existing risk control modeling processes without requiring large-scale modifications to existing business systems, reducing the manual costs and technical barriers of risk control variable mining, and improving the automation level and execution efficiency of the entire risk control modeling process. Attached Figure Description

[0019] The accompanying drawings are provided to further illustrate the invention and form part of the specification. They are explained in detail together with the embodiments of the invention, but do not constitute a limitation thereof.

[0020] Figure 1 The main flowchart of the present invention, which is based on the fusion of causal inference and feature importance, is shown. Figure 2This is a flowchart of the data acquisition and causal effect calculation process of the present invention; Figure 3 This is a flowchart illustrating the feature importance and comprehensive scoring selection process for this invention. Figure 4 This is a flowchart of the redundancy removal and generalization verification process of the present invention. Detailed Implementation

[0021] The technical solutions of the present invention will be clearly and completely described below with reference to the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the scope of protection of the present invention.

[0022] Example 1: Variable Mining for Pre-Loan Access Risk Control in Personal Consumer Credit This embodiment is applied to the pre-loan access risk control scenario for personal consumer credit in commercial banks, with the target variable being... The system assigns a second-category default label to users to indicate whether they have defaulted for more than 30 days within 12 months of receiving the loan. A default label value of 1 indicates a default has occurred, while a value of 0 indicates no default has occurred.

[0023] In this embodiment, the total number of samples collected is 120,000, of which the number of samples in the training set is [number missing]. The initial set contains 100,000 candidate variables, spanning from January 2025 to December 2025. The OOT validation set contains 20,000 candidate variables, spanning from January 2026 to March 2026. After preprocessing, the total number of candidate variables in the candidate variable set is... There are 200 variables, including three main categories: user credit information derived variables, transaction behavior variables, and consumption behavior variables.

[0024] In step S2, the improved PC algorithm constructs a causal DAG and identifies the backdoor path confusion variable set. It includes four basic variables: user age, gender, education level, and average monthly income. The pre-set significance level is... It is 0.05. Number of folds in cross-validation The unit marginal increment of a continuous variable is 5. Take 0.1 times the standard deviation of the variable itself.

[0025] With continuous candidate variables Taking the average monthly spending over the past 6 months as an example, the calculation process is demonstrated. The standard deviation of this variable is 2800 yuan, therefore... The value is set to 280 yuan. Total number of training set samples. The value is 100,000. The conditional expectation is obtained by fitting the data using a DML framework, and then substituted into... Calculation formula: ; Calculated The value is -0.0012, corresponding to the t-test. The value is 0.002, which satisfies the condition. The significance requirement.

[0026] In step S3, a LightGBM model with monotonicity constraints is used, with a stability penalty coefficient. The value is 1. Values Regarding variables Calculations yielded It is 0.082. The value is 0.076. Substituting this into the stability-weighted feature importance formula: ; Calculated It is 0.0762.

[0027] In step S4, all 200 candidate variables... It is 0.0003. It is 0.0008. It is 0.005. The value is 0.012. Substituting this into the fusion scoring formula: ; Calculated It is 4.033.

[0028] All variables by After sorting in descending order, an ordered score sequence is generated. The first-order difference sequence is calculated to obtain the position corresponding to the largest absolute difference. The value is 32, corresponding to an adaptive threshold of 0.62. Variables with scores greater than or equal to 0.62 are retained, resulting in a preliminary risk control variable set of 32 variables.

[0029] In step S5, the causal redundancy elimination process identifies the target variable. The Markov blanket was used to identify 26 variables. Six variables not included in the Markov blanket were removed. The VIF test had a preset threshold of 10, and four variables with a VIF greater than 10 were removed, leaving 22 variables. For generalization validation, the PSI, KS, and AUC thresholds were preset to 0.2 and 0.55 respectively. Finally, 18 interpretable risk control variables were selected, including core variables such as average monthly spending over the past 6 months, number of credit inquiries over the past 3 months, and average monthly repayment-to-income ratio.

[0030] Example 2: Risk Monitoring Variable Mining in Micro and Small Enterprise Business Loans This embodiment is applied to a risk monitoring scenario in the lending process for small and micro enterprises in city commercial banks, with the target variable being... This is a continuous label for the number of overdue days per month for enterprises, used to monitor credit risk fluctuations caused by changes in the enterprise's operating conditions.

[0031] In this embodiment, the total number of samples collected is 36,000, of which the number of samples in the training set is [number missing]. The initial sample set contains 30,000 entries, spanning from January 2024 to June 2025. The OOT validation set contains 6,000 entries, spanning from July 2025 to December 2025. After preprocessing, the total number of candidate variables in the candidate variable set is... There are 150 variables, including four categories: business operating cash flow variables, tax declaration variables, upstream and downstream transaction stability variables, and business owner credit variables.

[0032] In step S2, the improved PC algorithm constructs a causal DAG and identifies the backdoor path confusion variable set. The variables include four basic variables: company establishment year, industry, owner's age, and registered capital. A pre-set significance level is used. It is 0.05. Number of folds in cross-validation The ratio was 5, the number of PSM-matched control group matches was 1:4, and the number of bootstrap repeated samplings was 1000.

[0033] Discrete candidate variables Taking the example of whether there is a break in social security contributions within the past 3 months, the calculation process is demonstrated. This variable is a binary variable, using 0-1 dummy variable coding, and the processing group takes values... A value of 1 indicates a history of missed payments; the control group uses the value... A value of 0 indicates no record of missed payments. After PSM matching, the number of samples in the treatment group... The number is 2100, the sample size of the control group. The value is 8400. Substitute this into the discrete variable... Calculation formula: ; Calculation of the treatment group samples The mean was 12.3 days, compared to the control group sample. The average is 3.6 days, therefore The value is 8.7. The bootstrap test yielded a two-tailed result. The value is 0.001, which satisfies the condition. The significance requirement.

[0034] In step S3, an XGBoost model with monotonicity constraints is used, and the stability penalty coefficient is... The value is 0.8. Values Regarding variables Calculations yielded It is 0.105. The value is 0.098. Substituting this into the stability-weighted feature importance formula: ; Calculated It is 0.0996.

[0035] In step S4, all 150 candidate variables... It is 1.2. It is 2.1. It is 0.0067. The value is 0.015. Substituting this into the fusion scoring formula: ; Calculated It is 9.78.

[0036] All variables by After sorting in descending order, an ordered score sequence is generated. The first-order difference sequence is calculated to obtain the position corresponding to the largest absolute difference. The value is 25, corresponding to an adaptive threshold of 0.58. Variables with scores greater than or equal to 0.58 are retained, resulting in a preliminary risk control variable set of 25 variables.

[0037] In step S5, the causal redundancy elimination process identifies the target variable. The Markov blanket was used, and after removing 4 variables not included in the Markov blanket, 21 variables remained. The VIF test had a preset threshold of 10, and after removing 3 variables with a VIF greater than 10, 18 variables remained. In the generalization validation phase, the PSI preset threshold was 0.25, the KS preset threshold was 0.18, and the AUC preset threshold was 0.54. Finally, a set of 15 explainable risk control variables was obtained, including core variables such as social security payment interruption records for the past 3 months, volatility of operating cash flow for the past 6 months, and year-on-year change rate of average monthly tax payment.

[0038] Example 3: Variable Mining for Credit Card Post-Loan Collection Early Warning Risk Control This embodiment is applied to a post-loan collection early warning scenario for credit cards in joint-stock banks, with the target variable being... This is a binary label indicating whether a user is on a high-risk debt collection list. A value of 1 indicates that the user is on the high-risk debt collection list, while a value of 0 indicates that the user is making normal repayments.

[0039] In this embodiment, the total number of samples collected is 150,000, of which the number of samples in the training set is [number missing]. The initial dataset contains 120,000 candidate variables, spanning from March 2024 to February 2025. The OOT validation set contains 30,000 candidate variables, spanning from March 2025 to April 2025. After preprocessing, the total number of candidate variables in the candidate variable set is... There are 180 variables, including four main categories: user card usage behavior variables, repayment behavior variables, bill installment variables, and credit limit usage variables.

[0040] In step S2, the improved PC algorithm constructs a causal DAG and identifies the backdoor path confusion variable set. It includes four basic variables: user age, cardholder duration, credit limit, and education level. A pre-set significance level is used. It is 0.05. Number of folds in cross-validation The unit marginal increment of a continuous variable is 5. Take 0.15 times the standard deviation of the variable itself.

[0041] With ordered multi-category candidate variables The calculation process is demonstrated using the number of minimum repayments in the past three months as an example. This variable is an ordered multi-category variable, taking values ​​of 0, 1, 2, 3, and above. Adjacent category codes are used to determine the values ​​for the treatment group. Values ​​of 3 and above are used for the control group. The value is 0. After PSM matching, the number of samples in the treatment group is... The number is 3600, the sample size of the control group. The value is 14400. Substitute this into the discrete variable... Calculation formula: ; Calculation of the treatment group samples The mean value was 0.28, and the control group sample... The mean is 0.03, therefore The value is 0.25. The bootstrap test yields a two-tailed result. The value is 0.0005, which satisfies the condition. The significance requirement.

[0042] In step S3, a LightGBM model with monotonicity constraints is used, with a stability penalty coefficient. The value is 1.2. Values Regarding variables Calculations yielded It is 0.126. Given values ​​of 0.076 and 0.118, substituting these values ​​into the stability-weighted feature importance formula: ; Calculated It is 0.1167.

[0043] In step S4, all 180 candidate variables... It is 0.02. It is 0.06. It is 0.0056. The value is 0.014. Substituting this into the fusion scoring formula: ; Calculated It is 11.79.

[0044] All variables by After sorting in descending order, an ordered score sequence is generated. The first-order difference sequence is calculated to obtain the position corresponding to the largest absolute difference. The initial set of risk control variables is 28, corresponding to an adaptive threshold of 0.6. Variables with scores greater than or equal to 0.6 are retained, resulting in an initial set of 28 variables.

[0045] In step S5, the causal redundancy elimination process identifies the target variable. The Markov blanket was used to identify 23 variables. Five variables not included in the Markov blanket were removed. The VIF test had a preset threshold of 10; three variables with a VIF greater than 10 were removed, leaving 20 variables. For generalization validation, the PSI, KS, and AUC thresholds were preset to 0.2 and 0.22 respectively. Finally, a set of 16 explainable risk control variables was obtained, including core variables such as the number of minimum repayments in the past 3 months, the installment payment ratio in the past 6 months, and the average monthly volatility of credit limit utilization.

[0046] This invention constructs a causal directed acyclic graph (DAG) through causal structure learning with business constraints, accurately identifying confounding variables related to backdoor paths. Combining dual machine learning and propensity score matching methods, it performs unbiased causal effect estimation and significance testing for different types of variables. This clarifies the true path of variable impact on risk control objectives at the causal relationship level, fundamentally eliminating redundant variables with only spurious correlations to the target variable. Furthermore, this invention performs causal redundancy removal based on a Markov blanket of the target variable, ensuring minimal information redundancy and maximum causal explanatory power in the selected variable set. This solves the core technical problem of insufficient interpretability of traditional risk control variables and aligns with regulatory compliance requirements in the financial risk control field.

[0047] This invention constructs a comprehensive scoring system that integrates causal effects and stability-weighted feature importance. By introducing a stability penalty term based on the difference in feature importance across time-series OOT validation sets, it quantifies the stability of variables over time. Simultaneously, through standardized fusion and adaptive threshold selection, it achieves automated and unbiased variable selection, balancing causal significance and predictive discriminative ability. This approach effectively reduces the performance degradation problem after variable deployment caused by overfitting and sample distribution shifts in traditional methods, significantly improving the generalization ability and robustness of risk control variables in cross-cycle business scenarios.

[0048] This invention designs appropriate causal effect calculation and significance testing processes for candidate variables of different data types, including continuous, binary, and ordered multi-class classifications, covering all business scenarios of credit risk control, from pre-loan access to mid-loan risk monitoring and post-loan collection early warning. Simultaneously, this invention ensures the statistical validity and business applicability of the output variables through a multi-dimensional closed-loop verification system including multicollinearity testing, population stability testing, discrimination testing, and predictive performance testing. The method can be directly embedded into existing risk control modeling processes without requiring large-scale modifications to existing business systems, reducing the manual costs and technical barriers of risk control variable mining, and improving the automation level and execution efficiency of the entire risk control modeling process.

[0049] The foregoing has shown and described the basic principles, main features, and advantages of the present invention. Those skilled in the art should understand that the present invention is not limited to the above embodiments. The embodiments and descriptions in the specification are merely preferred examples and are not intended to limit the invention. Various changes and modifications can be made to the invention without departing from its spirit and scope, and all such changes and modifications fall within the scope of the present invention as claimed. The scope of protection of the present invention is defined by the appended claims and their equivalents.

Claims

1. A method for intelligent mining and identification of risk control variables based on the fusion of causal inference and feature importance, comprising the following steps: S1: Collect full sample data from credit risk control scenarios, complete data preprocessing, divide the data into a training set and a cross-time OOT validation set, and determine the target variable for risk control. With the candidate variable set; S2: Based on candidate variable set and target variable A constrained causal structure learning algorithm is used to construct a directed acyclic graph (DAG) to identify backdoor path confusion variable sets. Calculate the modified conditional average treatment effect for each candidate variable. And complete the significance test of causal effect; S3: Based on the training set and the OOT validation set, an ensemble learning model with monotonic constraints is used to calculate the stability-weighted feature importance of each candidate variable. ; S4: Based on standardization and A fusion scoring model was constructed to calculate the comprehensive mining score of each candidate variable. After sorting by score in descending order, an initial set of risk control variables is obtained through adaptive threshold filtering; S5: Perform causal redundancy removal and generalization verification on the initial set of risk control variables, and output the final interpretable set of risk control variables.

2. The intelligent mining and identification method for risk control variables based on causal inference and feature importance fusion as described in claim 1, characterized in that, In step S2, for continuous candidate variables Corrected conditional average treatment effect The calculation formula is: ; in, The total number of samples in the training set. As the target variable for risk control, Candidate variables The preset unit marginal increment, To obfuscate the variable set for backdoor paths based on causal DAG identification, For the conditional expectation operator, For the first Candidate variables corresponding to each sample The value of , For the first The value vector of the set of confounding variables corresponding to each sample.

3. The intelligent mining and identification method for risk control variables based on the fusion of causal inference and feature importance as described in claim 1, characterized in that, In step S2, the constrained causal structure learning algorithm employs an improved PC algorithm, and the preset causal constraints include: risk control target variables. A variable can only be a child node of a causal DAG and cannot be the parent node of any node. There must be no directed edges between derived variables in the same dimension. In the time dimension, a variable that occurs earlier can only be the parent node of a variable that occurs later.

4. The intelligent mining and identification method for risk control variables based on causal inference and feature importance fusion as described in claim 1, characterized in that, Importance of stability-weighted features mentioned in step S3 The calculation formula is: ; in, Candidate variables Feature importance calculated based on tree model split gain on the training set. To determine the importance of this variable as a feature of the same dimension on the OOT validation set, The preset stability penalty coefficient, To prevent positive minimum values ​​where the denominator is 0.

5. The intelligent mining and identification method for risk control variables based on causal inference and feature importance fusion as described in claim 1, characterized in that, The calculation formula for the fusion scoring model mentioned in step S4 is as follows: ; in, , For all candidate variables The sample mean and standard deviation, , For all candidate variables The sample mean and standard deviation, for The corresponding t-test p-value, At the preset significance level, This is an indicator function that takes the value when the condition within the parentheses is met. Otherwise, the value is .

6. The intelligent mining and identification method for risk control variables based on causal inference and feature importance fusion as described in claim 1, characterized in that, Step S2 employs a dual machine learning DML framework for estimation. The expected condition includes the following sub-steps: S21: Divide the training set samples into... Constructing a series of non-overlapping folds In folded cross-validation, samples are grouped so that each folded sample serves as a validation subset, and the remaining samples are... Fold as a training subset; S22: For each training subset, obfuscate the variable set. As input features, with risk control target variables To fit the labels, fit the first set of residual prediction models and output the target variable residuals for the corresponding validation subset. ; S23: For the same training subset, use a confounding variable set As input features, with candidate variables To fit the labels, a second set of residual prediction models is fitted, outputting the candidate variable residuals for the corresponding validation subset. ; S24: Target variable residuals based on the full sample Residuals of candidate variables By fitting a linear regression model, the unbiased conditional average treatment effect is obtained, thus eliminating the overfitting bias caused by high-dimensional confounding variables.

7. The intelligent mining and identification method for risk control variables based on the fusion of causal inference and feature importance as described in claim 1, characterized in that, The method for determining the adaptive threshold in step S4 specifically includes the following sub-steps: S41: Comprehensive Mining and Scoring of All Candidate Variables Sort the scores in descending order of numerical value to generate an ordered score sequence. ,in Let be the total number of candidate variables, and satisfy . ; S42: Calculate the first-order difference sequence of the ordered score sequence. ,in , The range of values ​​is arrive ; S43: Calculate the absolute value of each element in the difference sequence. Iterate through the sequence to find the position of the element with the largest absolute value difference. ; S44: In an ordered scoring sequence Location-related rating As an adaptive screening threshold, candidate variables with scores greater than or equal to the threshold in the ordered scoring sequence are retained to form the initial risk control variable set.

8. The intelligent mining and identification method for risk control variables based on causal inference and feature importance fusion as described in claim 1, characterized in that, The causal redundancy removal described in step S5 specifically includes the following sub-steps: S511: Based on the causal DAG constructed in step S2, identify the risk control target variable. A Markov blanket containing a target variable. The set of parent nodes, the set of child nodes, and the set of other parent nodes of the child nodes; S512: Iterate through all variables in the initial risk control variable set and remove those not in the target variable. Within the Markov blanket range of variables, a subset of variables with causal redundancy is obtained; S513: Perform a multicollinearity test on the subset of variables after removing causal redundancy, and calculate the variance inflation factor for each variable. value; S514: According to Values ​​are removed in descending order of importance. Variables whose values ​​are greater than a preset threshold, until all remaining variables... All values ​​are less than or equal to the preset threshold, thus completing the collinearity redundancy removal.

9. The intelligent mining and identification method for risk control variables based on the fusion of causal inference and feature importance as described in claim 1, characterized in that, The generalization verification in step S5 specifically includes the following sub-steps: S521: Based on the cross-time OOT validation set, calculate the population stability index of each variable in the variable subset after redundancy removal. The The calculation interval is a preset equal-frequency binning interval, and the preset number of bins is [number missing]. to box; S522: Remove Variables whose values ​​are greater than a preset stability threshold are selected as a subset of variables that meet the stability criteria. S523: For the subset of variables that meet stability criteria, calculate the correlation between each variable and the risk control target variable based on the OOT validation set. Rank correlation coefficient, discrimination Value, remove Variables whose values ​​are less than the preset discrimination threshold are selected as a subset of variables that meet the discrimination criteria. S524: Input the subset of variables that meet the discrimination criteria into the pre-set logistic regression benchmark model, and calculate the default prediction of each univariate on the OOT validation set. Value, remove For variables whose values ​​are less than the preset prediction performance threshold, the generalization verification is completed.

10. The intelligent mining and identification method for risk control variables based on causal inference and feature importance fusion as described in claim 1, characterized in that, In step S2, for discrete candidate variables This includes binary and ordinal categorical variables, and the adjusted conditional average treatment effect. The calculation process specifically includes the following sub-steps: S25 Discrete candidate variables Value encoding is performed, and binary classification variables are encoded using... - Dumb variable coding, for ordered multi-category variables, uses adjacent category coding to determine the values ​​for the treatment group. Values ​​compared to the control group ; S26 Backdoor Path Confusion Variable Set Based on Causal DAG Recognition The propensity score matching (PSM) method was used to match a predetermined number of control group samples to each treatment group sample. The matching caliper value was preset to be 1 / 3 of the propensity score standard deviation. times; S27 Based on the matched sample set, calculate candidate variables The modified conditional average treatment effect is calculated using the following formula: ; in, This is the set of samples for the matched processing group. To determine the number of samples in the processing group, This is the matched control group sample set. This represents the sample size of the control group. , These represent the values ​​of the risk control target variables for the corresponding samples; S28 For the calculated The bootstrap method is used for repeated sampling inspection, with the number of repeated samplings preset to [number]. Next, the standard error for calculating causal effects and the two-tailed test. The value is used to complete the significance test of the causal effect.