Risk control assessment method based on consensus refusal inference

By introducing co-occurrence performance data and dynamic monitoring of rejected orders, the problem of bias in the risk estimation of all users in the credit risk control model was solved, the stability and accuracy of the model were improved, the vicious cycle was broken, and the control of lending rate and bad debt rate was enhanced.

CN120852039APending Publication Date: 2025-10-28重庆富民银行股份有限公司
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510995581.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-18
Publication Date
2025-10-28

AI Technical Summary

Technical Problem

Existing credit risk control models are based solely on approved loan samples, leading to a systematic bias in the risk estimation of all credit applicants. This results in an inability to accurately reflect the true risk level, which in turn leads to a vicious cycle of rising bad debt rates and declining lending rates.

Method used

By acquiring full credit application order data, including approved and rejected order data, an initial risk control model is trained. Then, the performance data of rejected orders are introduced for annotation and extended training to form an extended training sample set. The risk control model is updated, and the model performance is dynamically monitored by combining the lending rate threshold and performance evaluation to ensure that the model is adapted to different scenarios.

Benefits of technology

It improves the accuracy of risk estimation for all credit applicants, avoids the increase in bad debt rate and the decrease in lending rate due to model bias, ensures that the model maintains a stable risk identification capability in actual credit approval, and dynamically adapts to different lending rate scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120852039A_ABST
    Figure CN120852039A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of credit risk control, discloses a risk control evaluation method for refusing inference based on a syngenetic representation mode, and aims to solve the problem of model estimation deviation caused by modeling by only using approved lending samples. The method comprises the steps of firstly obtaining full-amount credit application data, training an initial risk control model based on lended data and evaluating the sorting performance of the initial risk control model, and triggering refusal inference according to whether a lending rate is in a preset interval or not; according to the method, internal multi-product after-line lending performance or external credit investigation data subjected to validity verification is used as syngenetic performance data, rejected orders are marked in a preset time window to deduce good or bad samples, a new air control model is updated after integration, and stable performance of the model is ensured through continuous monitoring. The method can enrich the diversity of modeling samples, reduces the risk estimation deviation, gives consideration to the lending rate while reducing the bad debt rate, breaks the risk control vicious circle, and improves the model reliability and scene adaptability.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of credit risk control technology, specifically to a risk control assessment method based on rejection inference of co-existing performance patterns. Background Technology

[0002] Throughout the entire lending process, the risk control system acts like a sieve, filtering out relatively high-quality customers layer by layer, and ultimately deciding whether to grant a loan. The loan approval process typically includes anti-fraud strategies, policy regulations, credit approval strategies, and manual credit checks. For all application orders, only a small percentage are actually loaned out; the actual amount depends on each lending institution's risk control strategy.

[0003] Risk assessment models are a core tool for financial institutions in making credit approval decisions, and their accuracy directly affects the institution's bad debt ratio control, lending efficiency, and business sustainability. Traditional risk control model construction usually relies on the post-loan performance data of approved loan orders (i.e., "good samples" or "bad sample" labels). Through feature analysis of these samples and model training, a model is formed to assess the risk of new loan applicants.

[0004] However, this approach of modeling solely based on approved samples has significant limitations. Approved loan orders are the result of prior risk control strategies, and their user characteristic distribution naturally differs from that of all loan applicants (including those whose orders were rejected). The model's learning of risk patterns solely from approved samples leads to a systematic bias in risk estimation for rejected users, thus failing to accurately reflect the true risk level of all applicants. For example, customers with severe multiple borrowing are typically rejected, resulting in a loan sample largely consisting of customers with relatively few multiple borrowings. When the model built using these loan samples is applied to the full application sample, the problem of "estimating the whole from a part" leads to inaccurate risk estimations for all applicants. Over time, the trained model deviates further from reality, even approving a large number of potentially rejected bad users, resulting in substantial bad debts. To reduce the bad debt rate, risk control strategies will be further tightened, making it difficult to improve the lending rate, thus trapping the risk control system in a vicious cycle. Summary of the Invention

[0005] The present invention aims to provide a risk control assessment method based on the rejection of inference based on co-occurrence behavior, in order to solve the model estimation bias caused by modeling based only on approved loan samples in the prior art, and improve the model quality for risk estimation of all credit applicants.

[0006] To achieve the above objectives, the present invention adopts the following technical solution: A risk control assessment method based on rejection of inference in co-occurrence behavior includes: S1, Obtain full credit application order data, which includes approved loan order data and rejected order data; S2, train an initial risk control model based on the approved loan order data, evaluate the ranking of the initial risk control model, and if the initial risk control model meets the preset ranking threshold, then proceed to step S3; S3, calculate the lending rate in the full volume of credit application order data. If the lending rate is within a preset effective threshold range, trigger the rejection inference process and execute step S4. S4, Acquire and validate co-existing performance data: S41, determine whether there is internal multi-product line data. If so, obtain the post-loan performance data of the customer corresponding to the rejected order in other internal product lines as the first co-occurrence performance data; if not, obtain external credit data as the second co-occurrence performance data. S42, verify the validity of the second co-occurrence performance data by selecting a batch of approved loan orders as verification samples, labeling the verification samples as good or bad based on the internal post-loan performance tags and the second co-occurrence performance data, and calculating the consistency of the two labeling results; if the consistency is greater than a preset consistency threshold, the second co-occurrence performance data is determined to be valid. S5, label the rejected orders as good or bad based on the effective co-occurrence performance data; the labeling is limited to a preset time window before and after the application of the rejected order, and the rejected orders are labeled as inferred good samples, inferred bad samples or uncertain samples by comprehensively considering the credit records of customers in the effective co-occurrence performance data; S6, integrate the inferred good samples and inferred bad samples with the approved loan order data to form an extended training sample set, update the risk control model based on the extended training sample set, and obtain the target risk control model; S7. The target risk control model is evaluated for performance and continuously monitored in actual credit approval scenarios. If the model performance drift exceeds a preset threshold, the process returns to step S6 to update the model again.

[0007] The principles and advantages of this solution are as follows: In practical applications, by introducing inference samples from rejected orders, the diversity of modeling samples is enriched, the difference between training samples and the full application sample is reduced, and the risk estimation bias of the model for the full user base is decreased, thus solving the problem of estimating the overall population with partial samples. It avoids the chain reaction of tightened risk control strategies and decreased lending rates due to a large number of bad debts caused by model bias, thus improving the lending rate while reducing the bad debt rate and breaking the vicious cycle of the risk control system. Rejected orders are labeled with co-existing performance data (internal multi-product line post-loan performance or external credit data with verified effectiveness), ensuring the reliability of inference samples and improving the accuracy of rejection inference. Rejection inference is triggered based on the lending rate threshold, avoiding the introduction of invalid inference samples that lead to model performance degradation in high / low lending rate scenarios, and dynamically adapting to different scenarios. Through comprehensive performance evaluation and dynamic monitoring, model performance drift is detected and updated in a timely manner, ensuring that the model maintains a stable risk identification capability in actual credit approval and guaranteeing the model's continued effectiveness.

[0008] Preferably, as an improvement, the ranking is evaluated by at least one of the following metrics: KS value, AUC value, and Lorentz curve fit.

[0009] Technical effect: KS value and AUC value are core indicators for quantifying the risk discrimination ability of a model. Using them to evaluate ranking can objectively and accurately determine whether the initial model has the basic performance to trigger rejection inference, avoiding the introduction of inference samples when the model itself has insufficient ranking ability, which would lead to further deterioration of model performance.

[0010] Preferably, as an improvement, the consistency is calculated by a confusion matrix, including the proportion of the number of matches between good samples and good samples, uncertain samples and uncertain samples, and bad samples and bad samples in the internal post-loan performance label and the second co-occurrence performance data label to the total number of verification samples.

[0011] Technical effect: The confusion matrix can comprehensively display the matching details of internal and external data on the three categories of good / uncertain / bad. The consistency calculation in this way can accurately verify the fit between external credit data and internal standards, ensuring the validity of the external data used for rejection inference and guaranteeing the accuracy of rejection order labeling from the source.

[0012] Preferably, as an improvement, the labeling rule for inferring bad samples is: within the preset time window, there are records in the valid co-existing performance data that are continuously overdue for more than a preset number of days.

[0013] Technical benefits: It helps avoid inference bias caused by differences in annotation rules, ensures that the inferred bad samples can truly reflect the user's risk characteristics, and improves the model's ability to identify high-risk users.

[0014] Preferably, as an improvement, the preset number of days is determined by: statistically analyzing the overdue days distribution of all bad samples in the approved loan order data, taking the median of the distribution as the initial value of the preset number of days, and then adjusting it to match the internal bad sample definition through iterative verification.

[0015] Technical effect: It makes it easier to ensure that the labeling criteria for bad samples in rejected orders are highly consistent with the existing internal bad sample definitions.

[0016] Preferably, as an improvement, the construction of the extended training sample set further includes a data cleaning step: handling missing values, removing outliers, and performing feature caliber logic verification on the integrated samples.

[0017] Technical benefits: Facilitates ensuring the consistency of sample features.

[0018] Preferably, as an improvement, the performance evaluation includes a discrimination index, a stability index, and a prediction accuracy. The discrimination index includes the KS value and the AUC value. The stability index includes the PSI value. The prediction accuracy includes the bad sample prediction accuracy and the good sample discrimination accuracy.

[0019] Technical performance: The discrimination index reflects the model's ability to distinguish risks, the stability index reflects the model's consistency over time, and the prediction accuracy directly reflects the precision of risk assessment.

[0020] Preferably, as an improvement, the criterion for determining the performance drift of the model is: any performance index exceeds a preset baseline by ±10%.

[0021] Technical benefits: The ±10% drift threshold can quantify the degree of model performance degradation, ensuring timely detection and correction of model biases and avoiding risk misjudgments caused by performance degradation.

[0022] Preferably, as an improvement, in step S3, if the lending rate is not within a preset effective threshold range, the initial risk control model is maintained, and the distribution characteristics of the current full volume of credit application order data are recorded.

[0023] Technical benefits: Facilitates subsequent model iteration analysis.

[0024] Preferably, as an improvement, in S6, the method of updating the risk control model includes: when the distribution difference between the expanded training sample set and the initial training sample set is less than or equal to a preset difference threshold, an incremental update method is adopted; when the distribution difference is greater than the preset difference threshold, a retraining method is adopted; the distribution difference is calculated through the feature PSI mean.

[0025] Technical benefits: By quantifying the sample distribution difference through the mean of the feature PSI, incremental updates can reduce computational resource consumption and quickly adapt to sample changes when the distribution difference is small; when the distribution difference is large, retraining can ensure that the model is fully adapted to the new sample distribution, balancing model update efficiency and performance stability, and extending the effective lifespan of the model. Attached Figure Description

[0026] Figure 1 This is a flowchart illustrating a risk control assessment method that rejects inference based on co-occurrence behavior. Detailed Implementation

[0027] The following detailed description illustrates the specific implementation method: The basic implementation examples are as follows: Figure 1 As shown: A risk control assessment method based on rejection of inference in co-occurrence behavior includes: S1. Obtain all loan application order data, which includes approved loan order data and rejected loan order data. The approved loan order data includes customer credit history, financial status, and post-loan performance tags, while the rejected loan order data includes customer basic information and traceable application records.

[0028] S2, train an initial risk control model based on the approved loan order data. The approved loan order data can predict the customer's risk level and determine whether to approve the loan application. Evaluate the ranking of the initial risk control model. If the initial risk control model meets a preset ranking threshold, proceed to step S3.

[0029] Rejecting an inference is merely an inference, not a fact, and its impact on the model is not necessarily positive. Therefore, priority is given to ensuring that the initial risk control model built based on the lending samples has good ranking properties. If the ranking property is satisfied, the rejection inference mechanism is triggered. The ranking property is evaluated using at least one of the following indicators: KS value, AUC value, and Lorenz curve fit. In this embodiment, the preset ranking property threshold is KS value ≥ 0.3 or AUC value ≥ 0.7.

[0030] S3, calculate the loan disbursement rate in the full volume of credit application order data. If the loan disbursement rate is within a preset effective threshold range, trigger the rejection inference process and execute step S4. When the loan disbursement rate is very high, the sample bias problem is not obvious, and rejection inference is unnecessary; when the loan disbursement rate is very low, the model performance deteriorates due to the large difference between the rejection inference and the actual post-loan performance, and rejection inference is unnecessary. In this embodiment, the preset effective threshold range is 5% ≤ loan disbursement rate ≤ 90%; if the loan disbursement rate < 5% or the loan disbursement rate > 90%, the rejection inference process is not executed, the initial risk control model is maintained, and the distribution characteristics of the current full volume of credit application order data are recorded for subsequent model iteration analysis.

[0031] S4, Acquire and validate co-existing performance data: S41, determine whether internal multi-product-line data exists. If it exists, obtain the customer's post-loan performance data across other internal product lines corresponding to the rejected order as the first co-occurrence performance data; if it does not exist, obtain external credit data as the second co-occurrence performance data. Post-loan performance data across other internal product lines includes, but is not limited to: the customer's overdue records, repayment behavior, credit limit usage, and risk level labels across other product lines. External credit data includes central bank credit reports, third-party credit scores, multiple borrowing records, and post-loan performance data from other lending institutions.

[0032] S42, validate the validity of the second co-occurrence performance data. Select a batch of approved loan orders as validation samples. Label the validation samples with good and bad tags based on the internal post-loan performance tags and the second co-occurrence performance data, respectively, and calculate the consistency between the two labeling results. If the consistency is greater than a preset consistency threshold, the second co-occurrence performance data is determined to be valid. The good and bad tags include good samples, uncertain samples, and bad samples. The consistency is calculated using a confusion matrix, specifically: the proportion of good sample-good sample, uncertain sample-uncertain sample, and bad sample-bad sample matches between the internal post-loan performance tags and the second co-occurrence performance data tags to the total number of validation samples. The preset consistency threshold is ≥90%.

[0033] S5, label the rejected orders as good or bad based on the effective co-occurrence performance data; the effective co-occurrence performance data is the first co-occurrence performance data obtained through S41 or the second co-occurrence performance data verified through S42; the labeling is limited to a preset time window before and after the application of the rejected order, and the rejected orders are labeled as inferred good samples, inferred bad samples or uncertain samples by comprehensively considering the customer's credit record in the effective co-occurrence performance data.

[0034] The labeling rule for uncertain samples is as follows: within the preset time window, the credit records of customers in the effective co-occurrence performance data do not meet the conditions for inferring good samples or bad samples.

[0035] In this embodiment, the preset time window is 90 days before and after the date of the rejected order application. The labeling rule for inferring bad samples is: within the preset time window, there are records in the valid peer-to-peer performance data that are continuously overdue for more than N days, where N is a threshold determined based on the statistical characteristics of overdue samples in the approved loan order data. The method for determining N is: statistically analyzing the distribution of overdue days of all bad samples in the approved loan order data, taking the median of the distribution as the initial value of N, and then adjusting it through iterative verification until the matching degree with the internal bad sample definition is ≥95%.

[0036] S6, integrate the inferred good samples and inferred bad samples with the approved loan order data to form an extended training sample set. Update the risk control model based on the extended training sample set to obtain the target risk control model. The construction of the extended training sample set also includes a data cleaning step: handling missing values, removing outliers, and performing logical verification of feature calibers on the integrated samples to ensure the consistency of sample features.

[0037] The risk control model is updated in two ways: when the distribution difference between the expanded training sample set and the initial training sample set is less than or equal to a preset difference threshold, an incremental update method is used; when the distribution difference is greater than the preset difference threshold, a retraining method is used; the distribution difference is calculated using the mean of the feature PSI. In this embodiment, the preset difference threshold is a feature PSI mean ≤ 0.2.

[0038] S7. The target risk control model is evaluated for performance and continuously monitored in actual credit approval scenarios. If the model performance drift exceeds a preset threshold, the process returns to step S6 to update the model again.

[0039] The discrimination index includes KS value and AUC value; the stability index includes PSI value; and the prediction accuracy includes bad sample prediction accuracy and good sample discrimination accuracy. The criterion for determining model performance drift is: any one of the indicators exceeds the preset baseline by ±10%.

[0040] The above descriptions are merely embodiments of the present invention, and common knowledge such as specific technical solutions and / or characteristics are not described in detail here. It should be noted that those skilled in the art can make various modifications and improvements without departing from the technical solutions of the present invention, and these should also be considered within the scope of protection of the present invention. These modifications and improvements will not affect the effectiveness of the implementation of the present invention or the practicality of the patent. The scope of protection claimed in this application should be determined by the content of its claims, and the specific embodiments described in the specification can be used to interpret the content of the claims.

Claims

1. A risk control assessment method based on rejection of inference based on co-occurrence behavior, characterized in that... ,include: S1, Obtain full credit application order data, which includes approved loan order data and rejected order data; S2, train an initial risk control model based on the approved loan order data, evaluate the ranking of the initial risk control model, and if the initial risk control model meets the preset ranking threshold, then proceed to step S3; S3, calculate the lending rate in the full volume of credit application order data. If the lending rate is within a preset effective threshold range, trigger the rejection inference process and execute step S4. S4, Acquire and validate co-existing performance data: S41, determine whether there is internal multi-product line data. If so, obtain the post-loan performance data of the customer corresponding to the rejected order in other internal product lines as the first co-occurrence performance data; if not, obtain external credit data as the second co-occurrence performance data. S42, verify the validity of the second co-occurrence performance data by selecting a batch of approved loan orders as verification samples, labeling the verification samples as good or bad based on the internal post-loan performance tags and the second co-occurrence performance data, and calculating the consistency of the two labeling results; if the consistency is greater than a preset consistency threshold, the second co-occurrence performance data is determined to be valid. S5, label the rejected orders as good or bad based on the effective co-occurrence performance data; the labeling is limited to a preset time window before and after the application of the rejected order, and the rejected orders are labeled as inferred good samples, inferred bad samples or uncertain samples by comprehensively considering the credit records of customers in the effective co-occurrence performance data; S6, integrate the inferred good samples and inferred bad samples with the approved loan order data to form an extended training sample set, update the risk control model based on the extended training sample set, and obtain the target risk control model; S7. The target risk control model is evaluated for performance and continuously monitored in actual credit approval scenarios. If the model performance drift exceeds a preset threshold, the process returns to step S6 to update the model again.

2. The risk control assessment method based on rejection inference of co-occurrence behavior as described in claim 1, characterized in that: The ranking is evaluated using at least one of the following metrics: KS value, AUC value, and Lorentz curve fit.

3. The risk control assessment method based on rejection inference of co-occurrence behavior as described in claim 1, characterized in that: The consistency is calculated using a confusion matrix, which includes the proportion of matches between good samples, uncertain samples, and bad samples in the internal post-loan performance label and the second co-occurrence performance data label, relative to the total number of verification samples.

4. The risk control assessment method based on rejection inference of co-occurrence behavior as described in claim 1, characterized in that, The labeling rule for inferring bad samples is: within the preset time window, there are records in the valid co-existing performance data that are continuously overdue for more than a preset number of days.

5. The risk control assessment method based on rejection inference of co-occurrence behavior as described in claim 4, characterized in that, The preset number of days is determined by: statistically analyzing the overdue days distribution of all bad samples in the approved loan order data, taking the median of the distribution as the initial value of the preset number of days, and then adjusting it to match the internal bad sample definition through iterative verification.

6. The risk control assessment method based on rejection inference of co-occurrence behavior as described in claim 1, characterized in that, The construction of the extended training sample set also includes a data cleaning step: handling missing values, removing outliers, and performing feature caliber logic verification on the integrated samples.

7. The risk control assessment method based on rejection inference of co-occurrence behavior as described in claim 1, characterized in that: The performance evaluation includes discrimination index, stability index and prediction accuracy. The discrimination index includes KS value and AUC value. The stability index includes PSI value. The prediction accuracy includes bad sample prediction accuracy and good sample discrimination accuracy.

8. The risk control assessment method based on rejection inference of co-occurrence behavior as described in claim 7, characterized in that: The criterion for determining model performance drift is: any performance indicator exceeds the preset baseline by ±10%.

9. The risk control assessment method based on rejection inference of co-occurrence behavior as described in claim 1, characterized in that: In step S3, if the lending rate is not within the preset effective threshold range, the initial risk control model is maintained, and the distribution characteristics of the current full volume of credit application order data are recorded.

10. The risk control assessment method based on rejection inference of co-occurrence behavior as described in claim 1, characterized in that, In step S6, the risk control model is updated in the following ways: when the distribution difference between the expanded training sample set and the initial training sample set is less than or equal to a preset difference threshold, an incremental update method is used; when the distribution difference is greater than the preset difference threshold, a retraining method is used; the distribution difference is calculated using the mean of the feature PSI.