Credit risk rating assessment method, device, equipment and storage medium

By using a credit risk rating assessment model based on the Adaboost algorithm, combined with feature selection and data preprocessing technology, the data quality problem of the credit risk assessment model is solved, achieving more accurate credit risk assessment and a more efficient loan approval process.

CN119693125BActive Publication Date: 2025-09-26INDUSTRIAL AND COMMERCIAL BANK OF CHINA
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411781674.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-12-05
Publication Date
2025-09-26
Estimated Expiration
2044-12-05

AI Technical Summary

Technical Problem

The existing credit risk rating assessment model is unable to accurately assess users' credit risks due to quality issues with the model training data, which increases the credit risk in the credit business of banks and other financial institutions.

Method used

A credit risk rating assessment model based on the Adaboost algorithm is adopted. By obtaining the credit data to be assessed of the target object and using the sample credit data set for training, a target assessment model is constructed by combining feature selection and data preprocessing techniques such as the Pearson correlation coefficient and the fuzzy C-means clustering algorithm to solve the problems of data imbalance and missingness and improve the quality of model training data.

Benefits of technology

It improves the accuracy of credit risk assessment, reduces the credit risk of banks and other financial institutions, improves loan approval efficiency, and shortens business processing time.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119693125B_ABST
    Figure CN119693125B_ABST
Patent Text Reader

Abstract

The present invention discloses a credit risk rating assessment method, apparatus, device, and storage medium, belonging to the field of artificial intelligence. The method comprises: obtaining credit data to be assessed of a target object; inputting the credit data to be assessed into a target assessment model to obtain a credit risk rating corresponding to the target object; wherein the target assessment model is obtained by training a credit risk rating assessment model based on an Adaboost algorithm based on a sample credit data set. The present invention improves the accuracy of user credit risk assessment and reduces credit risk in the credit business of financial institutions such as banks. At the same time, there is no need for credit approval personnel to review the borrower's loan application materials one by one, thereby improving the loan approval efficiency of financial institutions such as banks and shortening the business processing time of borrowers.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of artificial intelligence, and in particular to a credit risk rating assessment method, apparatus, device, and storage medium. Background Art

[0002] In the financial industry, when a user applies for a loan from a financial institution, banks and other financial institutions typically assess the user's credit risk based on credit data such as the user's credit history and financial status. Accurately assessing a user's credit risk is crucial for banks and other financial institutions to maintain stability and profitability.

[0003] However, existing credit risk rating assessment models are often unable to accurately assess users' credit risks due to quality issues with model training data, thereby increasing the credit risk in the credit business of banks and other financial institutions. Summary of the Invention

[0004] The present invention provides a credit risk rating assessment method, device, equipment and storage medium to improve the accuracy of user credit risk assessment and reduce credit risks in the credit business of financial institutions such as banks.

[0005] According to one aspect of the present invention, a credit risk rating assessment method is provided, the method comprising:

[0006] Obtain the credit data of the target object to be evaluated;

[0007] The credit data to be evaluated is input into the target evaluation model to obtain the credit risk level corresponding to the target object; wherein, the target evaluation model is obtained by training a credit risk level evaluation model based on the Adaboost algorithm based on a sample credit data set.

[0008] According to another aspect of the present invention, there is provided a credit risk level assessment device, the device comprising:

[0009] A credit data acquisition module for obtaining credit data to be evaluated, used to obtain the credit data to be evaluated of the target object;

[0010] The credit risk rating assessment module is used to input the credit data to be assessed into the target assessment model to obtain the credit risk rating corresponding to the target object; wherein, the target assessment model is obtained by training the credit risk rating assessment model based on the Adaboost algorithm based on the sample credit data set.

[0011] According to another aspect of the present invention, an electronic device is provided, comprising:

[0012] at least one processor; and

[0013] a memory communicatively connected to at least one processor; wherein,

[0014] The memory stores a computer program that can be executed by at least one processor. The computer program is executed by the at least one processor so that the at least one processor can execute the credit risk rating assessment method according to any embodiment of the present invention.

[0015] According to another aspect of the present invention, a computer-readable storage medium is provided, wherein the computer-readable storage medium stores computer instructions, which are used to enable a processor to implement the credit risk rating assessment method according to any embodiment of the present invention when executed.

[0016] According to another aspect of the present invention, a computer program product is provided, comprising a computer program, which, when executed by a processor, implements the credit risk rating assessment method according to any embodiment of the present invention.

[0017] The technical solution of the embodiment of the present invention obtains the credit data to be evaluated of the target object; inputs the credit data to be evaluated into the target evaluation model to obtain the credit risk level corresponding to the target object; wherein the target evaluation model is obtained by training a credit risk level evaluation model based on the Adaboost algorithm based on a sample credit data set. The above technical solution processes the credit data to be evaluated of the target object through the target evaluation model to obtain the credit risk level corresponding to the target object, thereby improving the accuracy of the credit risk assessment of the user and reducing the credit risk in the credit business of financial institutions such as banks. At the same time, there is no need for credit approval personnel to review the borrower's loan application materials one by one, thereby improving the loan approval efficiency of banks and other financial institutions and shortening the business processing time of the borrower.

[0018] It should be understood that the content described in this section is not intended to identify the key or important features of the embodiments of the present invention, nor is it intended to limit the scope of the present invention. Other features of the present invention will become readily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS

[0019] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.

[0020] Figure 1 This is a flow chart of a credit risk rating assessment method provided according to the first embodiment of the present invention;

[0021] Figure 2This is a flow chart of a credit risk rating assessment method provided according to the second embodiment of the present invention;

[0022] Figure 3 This is a flow chart of a credit risk rating assessment method provided in accordance with a third embodiment of the present invention;

[0023] Figure 4 This is a schematic diagram of the structure of a credit risk level assessment device provided according to a fourth embodiment of the present invention;

[0024] Figure 5 It is a structural diagram of an electronic device for implementing the credit risk level assessment method according to an embodiment of the present invention. DETAILED DESCRIPTION

[0025] In order to enable those skilled in the art to better understand the solutions of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the embodiments described are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts should fall within the scope of protection of the present invention.

[0026] It should be noted that the terms "target", "sample", "first" and "second" in the description and claims of the present invention and the above-mentioned drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that the numbers used in this way can be interchanged where appropriate, so that the embodiments of the present invention described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusions. For example, a process, method, system, product or device that includes a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.

[0027] It should be noted that the information collected in the present invention, such as the target object's credit data to be evaluated and the sample credit data set, is information and data authorized by the user or fully authorized by all parties, and the collection, storage, use, processing, transmission, provision, disclosure and application of the relevant data comply with the relevant laws, regulations and standards of the relevant countries and regions, take necessary confidentiality measures, do not violate public order and good morals, and provide corresponding operation entrances for users to choose to authorize or refuse.

[0028] Example 1

[0029] Figure 1This is a flowchart of a credit risk rating assessment method provided in the first embodiment of the present invention. This embodiment is applicable to the situation of assessing the credit risk rating of customers of financial institutions. The method can be executed by a credit risk rating assessment device, which can be implemented in the form of hardware and / or software and can be configured in an electronic device. Figure 1 As shown, the method includes:

[0030] S101: Obtain credit data to be evaluated of the target object.

[0031] The target subject refers to a financial institution's clients who borrow money from the financial institution, referred to as the borrower. The credit data to be assessed refers to the credit data required to assess the target subject's credit risk level; optionally, the credit data to be assessed includes, but is not limited to, the target subject's personal information, financial information, credit history, and social activity information. Credit history refers to the credit records and historical data generated by the target subject's past borrowing activities, including repayment records, loan limits, and credit card usage.

[0032] Specifically, the target subject's credit data to be assessed can be obtained based on a pre-defined data acquisition method. For example, the target subject's object identifier can be used as an index to obtain the target subject's credit data from a financial institution's database and a third-party data platform associated with the financial institution. The object identifier uniquely identifies the target subject; optionally, the object identifier can be in the form of numbers, letters, or a combination of numbers and letters. Third-party data platforms include, but are not limited to, credit reporting agencies, social networking platforms, and financial data service providers.

[0033] S102: Input the credit data to be evaluated into the target evaluation model to obtain the credit risk level corresponding to the target object.

[0034] The target assessment model is obtained by training a credit risk rating assessment model based on the Adaboost algorithm based on a sample credit dataset. The target assessment model refers to a trained credit risk rating assessment model, which is used to assess the credit risk rating corresponding to the target object. The credit risk rating is used to characterize the degree of credit risk of the target object. The sample credit dataset refers to the dataset required to train the credit risk rating assessment model; optionally, the sample credit dataset consists of sample credit data from a large number of sample objects. The sample object refers to an object randomly selected when training the credit risk rating assessment model. The sample credit data refers to the credit data of the sample object within a historical time period; optionally, the sample credit data includes but is not limited to the sample object's personal information, financial information, credit history, credit risk rating, and social activity information.

[0035] The Adaboost (Adaptive Boosting) algorithm is an ensemble learning method based on the Boosting strategy. Its core concept is to train different weak classifiers on the same training set and then combine these weak classifiers to form a stronger final classifier (i.e., a strong classifier). A weak classifier is one whose classification accuracy is only slightly better than random guessing. After each iteration, the Adaboost algorithm updates the weights of the samples in the training set and the weights of the weak classifiers. Finally, the classifiers obtained from each training cycle are combined to form the final decision classifier.

[0036] Specifically, the credit data to be evaluated is input into the target evaluation model, and after being processed by the target evaluation model, the credit risk level corresponding to the target object is obtained.

[0037] The technical solution of the embodiment of the present invention obtains the credit data to be evaluated of the target object; inputs the credit data to be evaluated into the target evaluation model to obtain the credit risk level corresponding to the target object; wherein the target evaluation model is obtained by training a credit risk level evaluation model based on the Adaboost algorithm based on a sample credit data set. The above technical solution processes the credit data to be evaluated of the target object through the target evaluation model to obtain the credit risk level corresponding to the target object, thereby improving the accuracy of the credit risk assessment of the user and reducing the credit risk in the credit business of financial institutions such as banks. At the same time, there is no need for credit approval personnel to review the borrower's loan application materials one by one, thereby improving the loan approval efficiency of banks and other financial institutions and shortening the business processing time of the borrower.

[0038] Example 2

[0039] Figure 2 This is a flowchart of a credit risk rating assessment method provided by Example 2 of the present invention. Based on the above examples, this example further optimizes the "determination of the target assessment model" and provides an optional implementation plan. It should be noted that for parts not described in detail in the examples of the present invention, reference can be made to the relevant descriptions of other examples. Figure 2 As shown, the method includes:

[0040] S201. Obtain a sample credit dataset.

[0041] S202: Perform feature selection on the sample credit dataset to obtain a key credit dataset.

[0042] Among them, the key credit data set refers to a data set consisting of data in the sample credit data set that has a greater impact on the credit risk level assessment.

[0043] Specifically, the first data feature and the second data feature in the sample credit data set can be determined; based on the Pearson correlation coefficient, the correlation between each sub-data feature in the second data feature and the first data feature is calculated; according to the correlation between each sub-data feature in the second data feature and the first data feature, feature selection is performed on the sample credit data set to obtain the key credit data set.

[0044] The Pearson correlation coefficient is a statistic that measures the strength of the linear relationship between variables. Its value range is between -1 and 1, reflecting the degree of correlation between two variables. If the correlation coefficient is close to 1, it indicates that there is a completely positive linear relationship between the two variables; if the correlation coefficient is close to -1, it indicates that there is a completely negative linear relationship between the two variables; if the correlation coefficient is close to 0, it indicates that there is no linear relationship between the two variables.

[0045] The first data feature refers to the credit risk level in the sample credit data set, and the second data feature refers to the data feature in the sample credit data set other than the first data feature.

[0046] More specifically, the first data feature and the second data feature in the sample credit data set can be determined according to actual business needs; then, for the i-th sub-data feature X in the second data feature, i , based on the Pearson correlation coefficient, calculate the i-th sub-data feature X in the second data feature by the following formula i Correlation with the first data feature:

[0047]

[0048] Wherein, M represents the total number of eigenvalues. It should be noted that the total number of eigenvalues ​​corresponding to each sub-data feature in the second data feature is the same, and is equal to the total number of eigenvalues ​​corresponding to the first data feature. j Represents X i The corresponding j-th eigenvalue is, Represents X i The average value of the corresponding M eigenvalues, y j represents the j-th eigenvalue corresponding to the first data feature, represents the average value of the M eigenvalues ​​corresponding to the first data feature, r i Represents the i-th sub-data feature X in the second data feature i The correlation between the first data feature and each sub-data feature in the second data feature can be obtained in the same way, and are recorded as r1, r2, ... and r n . Wherein, n represents the total number of sub-data features in the second data feature.

[0049] The correlation between each sub-data feature and the first data feature is then compared with a correlation threshold. Sub-data features with correlations greater than the correlation threshold, along with the first data feature, are saved to an empty dataset, thereby obtaining a key credit dataset. The correlation threshold can be pre-set based on actual business needs and is not specifically limited in this embodiment of the present invention.

[0050] It can be understood that by performing feature selection on the data features in the sample credit data set based on the Pearson correlation coefficient, only the data features in the sample credit data set that are highly correlated with the credit risk level are retained, thereby reducing the data volume of the model training data and improving the data quality of the model training data, thereby reducing the consumption of computing resources and improving the data processing efficiency during model training.

[0051] S203: Construct a training sample set based on the key credit data set, and train a credit risk rating assessment model based on the Adaboost algorithm based on the training sample set to obtain a target assessment model.

[0052] The training sample set refers to the data set actually used to train the credit risk rating assessment model.

[0053] Specifically, the data in the key credit data set can be processed based on the data processing model to obtain a training sample set output by the data processing model; wherein the data processing model is used to solve the data missing problem and data imbalance problem in the data set; then, the total number of training samples in the training sample set can be determined by a statistical algorithm; the total number of classifiers of the basic classifier in the Adaboost algorithm can be determined by a cross-validation strategy; wherein the cross-validation strategy can be a K-fold cross-validation strategy or a leave-one-out cross-validation strategy; the initial weight value of each training sample in the training sample set is determined based on the total number of samples, that is, the reciprocal of the total number of samples is used as the initial weight value of each training sample in the training sample set, thereby obtaining the initial weight value distribution of the training sample set:

[0054]

[0055] Where D1(i) represents the initial weight value distribution of the training sample set, and also represents the weight value distribution of the training sample set during the first training; i represents the i-th training sample in the training sample set, and i = 1, 2, ..., m; m represents the total number of training samples in the training sample set; w m It is important to note that the basic classifiers in the Adaboost algorithm are all decision tree stumps.

[0056] Afterwards, all the basic classifiers in the Adaboost algorithm are trained based on the training sample set that assigns a weight value to each training sample, and then the basic classifiers are combined to obtain the target evaluation model. More specifically, during the kth training, the training sample set T obtained after the k-1th training is k-1 Output to the kth basic classifier in the Adaboost algorithm, and get the kth basic classifier for the training sample set T k-1 The classification results of each training sample in ; According to the classification results obtained, and the training sample set T k-1 The actual classification results corresponding to each training sample in the Adaboost algorithm are used to determine the classification error rate of the k-th basic classifier:

[0057]

[0058] Among them, e k represents the classification error rate of the kth basic classifier in the Adaboost algorithm; x i Represents the training sample set T k-1 The i-th training sample in G k (x i ) represents the kth basic classifier for the training sample set T k-1 The classification result of the i-th training sample in y i Represents the training sample set T k-1 The actual classification result corresponding to the i-th training sample in P(G k (x i )≠y i ) indicates that the kth basic classifier in the Adaboost algorithm will train the sample set T k-1 The probability of misclassification of training samples in the training set T k-1 The total number of training samples misclassified by the kth basic classifier; w ki Represents the training sample set T during the kth training k-1 It should be noted that the number of training times for the credit risk rating assessment model based on the Adaboost algorithm is equal to the total number of classifiers in the basic classifier of the Adaboost algorithm. k (x i )≠y i ) is an indicator function, if G k )x i )≠y i If it is established, then I)G k )x i )≠y i )=1; if G k )x i )≠yi Not true, I(G k (x i )≠y i )=0.

[0059] Afterwards, according to the classification error rate of the kth basic classifier in the Adaboost algorithm, the weight of the kth basic classifier in the Adaboost algorithm is determined by the following formula:

[0060]

[0061] Among them, α k Represents the weight of the kth basic classifier in the Adaboost algorithm.

[0062] Afterwards, according to the weight of the kth basic classifier in the Adaboost algorithm, the training sample set T is updated by the following formula: k-1 The weight value distribution of the training sample set T k .

[0063]

[0064] Among them, D k (i) represents the training sample set T obtained after the kth training k The weight value distribution, z t represents the normalization constant,

[0065] After determining the weight of each basic classifier in the Adaboost algorithm, a weighted voting mechanism is used to linearly combine the basic classifiers in the Adaboost algorithm to obtain the target evaluation model, namely:

[0066]

[0067] Where G(x) represents the target evaluation model; K represents the total number of classifiers in the basic classifier of the Adaboost algorithm; G k (x) represents the kth basic classifier.

[0068] S204: Obtain the credit data of the target object to be evaluated.

[0069] S205: Input the credit data to be evaluated into the target evaluation model to obtain the credit risk level corresponding to the target object.

[0070] The technical solution of an embodiment of the present invention comprises obtaining a sample credit dataset; performing feature selection on the sample credit dataset to obtain a key credit dataset; constructing a training sample set based on the key credit dataset, and training a credit risk rating assessment model based on the Adaboost algorithm based on the training sample set to obtain a target assessment model; obtaining credit data of a target object to be assessed; and inputting the credit data to be assessed into the target assessment model to obtain the credit risk rating corresponding to the target object. After obtaining the sample credit dataset, the above technical solution performs a series of data preprocessing, such as feature selection, on the sample credit dataset to obtain a training sample set. This reduces the amount of training data required for training the credit risk rating assessment model and improves the data quality of the training data. Subsequently, the credit risk rating assessment model based on the Adaboost algorithm is trained based on the training sample set to obtain a target assessment model. This reduces the consumption of computing resources and model training time during training of the credit risk rating assessment model, improving data processing efficiency during training of the credit risk rating assessment model; and also improves the accuracy and reliability of the target assessment model. Afterwards, the target assessment model is used to process the target subject's credit data to be assessed, making the corresponding credit risk level of the target subject more accurate. This improves the accuracy of user credit risk assessments and reduces the credit risk in the credit operations of banks and other financial institutions. This allows financial institutions to adjust credit limits and terms in a targeted manner, avoiding excessive concentration on high-risk projects, thereby reducing the overall risk of the financial institutions' credit assets. At the same time, credit approval personnel no longer need to review each borrower's loan application materials one by one, thereby improving the loan approval efficiency of banks and other financial institutions and shortening the borrower's business processing time.

[0071] Example 3

[0072] Figure 3 This is a flowchart of a credit risk rating assessment method provided by Example 3 of the present invention. Based on the above examples, this example further optimizes the "construction of a training sample set based on a key credit data set" and provides an optional implementation plan. It should be noted that for the parts not described in detail in the examples of the present invention, reference can be made to the relevant descriptions of other examples. Figure 3 As shown, the method includes:

[0073] S301. Obtain a sample credit dataset.

[0074] S302: Perform feature selection on the sample credit data set to obtain a key credit data set.

[0075] S303. Based on the fuzzy C-means clustering algorithm, cluster the key credit data set to obtain a clustering result.

[0076] Among them, the Fuzzy C-means algorithm (FCM) is a soft clustering method. Its core idea is to use membership to represent the relationship between data in a data set, thereby determining the data cluster to which each data belongs.

[0077] The clustering result refers to the result obtained after clustering the key credit data set; optionally, the clustering result includes information of each data cluster.

[0078] Specifically, the number of clusters can be determined based on the specific number of credit risk levels. For example, if the credit risk level is divided into the following five levels: normal, special mention, substandard, doubtful, and loss, the number of clusters is 5. The number of clusters and the key credit data set are then input into the fuzzy C-means clustering algorithm, and the clustering results are obtained after processing by the fuzzy C-means clustering algorithm. The fuzzy C-means clustering algorithm processes the number of clusters and the key credit data set as follows:

[0079] Step a: Initialize the fuzzy factor, current number of iterations, membership matrix, iterative allowable error and maximum number of iterations.

[0080] Specifically, the fuzzy factor, the current number of iterations, the membership matrix, the iterative allowable error and the maximum number of iterations may be initialized according to actual business requirements and the experience of those skilled in the art.

[0081] Step b: Check whether the current number of iterations is less than the maximum number of iterations.

[0082] Specifically, if the current number of iterations is less than the maximum number of iterations, then go to step c; otherwise, then go to step e.

[0083] Step c: determining the cluster center according to the membership in the membership matrix, and updating the membership in the membership matrix according to the cluster center to obtain an updated membership matrix.

[0084] Specifically, according to the membership in the membership matrix, the cluster center is determined by the following cluster center determination formula:

[0085]

[0086] Among them, N j represents the jth cluster center; q represents the fuzzy factor, and the value of q is usually a real number greater than 1; u ij Indicates the membership degree of the i-th sample credit data in the key credit data set to the j-th cluster; x i represents the i-th sample credit data in the key credit dataset, and c represents the total number of sample credit data in the key credit dataset.

[0087] Afterwards, according to the cluster center, the membership in the membership matrix is ​​updated by the following membership update formula to obtain the updated membership matrix:

[0088]

[0089] Where n represents the number of clusters; N k represents the kth cluster center.

[0090] Step d: Detect whether the objective function of the fuzzy C-means clustering algorithm meets the iteration stopping condition.

[0091] Among them, the objective function of the fuzzy C-means clustering algorithm is as follows:

[0092]

[0093] Among them, d ij Represents the i-th sample credit data x in the key credit dataset i With the jth cluster center N j It should be noted that in order to minimize the objective function of the fuzzy C-means clustering algorithm, the objective function of the fuzzy C-means clustering algorithm needs to meet the following constraints:

[0094]

[0095] Where c represents the total number of sample credit data in the key credit dataset; n represents the number of clusters.

[0096] The iterative stopping condition refers to the condition that stops the iteration of the fuzzy C-means clustering algorithm. Optionally, the iterative stopping condition can be determined based on the iterative allowable error and the two adjacent objective functions. For example, the iterative stopping condition can be: || J t (U,V)-J t-1 (U,V)||<ε. Among them, J t (U, V) represents the objective function of the fuzzy C-means clustering algorithm at the tth iteration; J t-1 (U, V) represents the objective function of the fuzzy C-means clustering algorithm at the t-1th iteration.

[0097] Specifically, if the objective function of the fuzzy C-means clustering algorithm meets the iteration stopping condition, step e is executed; otherwise, the current iteration number is increased by 1, and step b is continued.

[0098] Step e: Output the clustering results.

[0099] It can be understood that clustering the key credit data set based on the fuzzy C-means clustering algorithm can bring together sample credit data of the same credit risk level in the key credit data set, making it easier to better mine the potential information between similar data and provide more potential information for subsequent multiple interpolation regression analysis.

[0100] Optionally, after obtaining the clustering results, the clustering labels of each data cluster in the clustering results can be uniquely encoded to obtain the clustering features corresponding to each data cluster; for each data cluster, the cluster corresponding to the data cluster is added to the feature information of each sample credit data in the data cluster to enrich the feature dimensions of the sample credit data in the key credit data set.

[0101] S304. Based on the clustering results, a chain equation multiple interpolation algorithm is used to perform data interpolation processing on the key credit data set to obtain a training sample set.

[0102] Among them, the Multiple Imputation by Chained Equations (MICE) algorithm fills the missing values ​​in the data set through an iterative prediction model. It should be noted that, in an embodiment of the present invention, the iterative prediction model used in the Multiple Imputation by Chained Equations algorithm is an iterative prediction model based on the Gradient Boosting Decision Tree (GBDT). Among them, GBDT is an integrated learning algorithm, and its core idea is to iteratively construct a series of decision trees, each decision tree trying to correct the prediction error of the previous decision tree. In each round of iteration, for each training sample, the error of the current model is calculated; using the error of the current model as the target, a new decision tree is trained to obtain a new decision tree; in order to minimize the loss, the gradient of the loss function is estimated and the learning rate is determined; the result of multiplying the new decision tree by the learning rate is added to the current model to obtain an updated current model; the above process is repeated until a predetermined number of iterations is reached, or until the performance of the model is no longer significantly improved, thereby obtaining the final iterative prediction model.

[0103] For example, for training sample 1, in the first iteration, the predicted value of the current model for training sample 1 is calculated, and the difference between the actual observed value and the predicted value of training sample 1 is taken as the error of the current model, which is recorded as e i ; Use the error e of the current model i As the goal, train a new decision tree and get the following new decision tree:

[0104] y=a i +(b i *X)+e i ;

[0105] Among them, y represents the new decision tree, X represents the training sample 1, a i and b i is the preset coefficient.

[0106] Afterwards, the final iterative prediction model is obtained at the end of the iteration:

[0107]

[0108] Where n represents the predetermined number of iterations; e n Represents the error of the current model in iteration n.

[0109] Specifically, based on the clustering results, the minority and majority clusters in the key credit data set can be determined; the minority clusters can be oversampled to obtain new data clusters corresponding to the minority clusters; and the chain equation multiple interpolation algorithm can be used to interpolate the new data clusters corresponding to the majority and minority clusters to obtain the training sample set. The minority cluster refers to the data cluster with a small number of sample credit data; correspondingly, the majority cluster refers to the data cluster with a large number of sample credit data.

[0110] More specifically, the key credit data set can be classified according to the total data volume of each data cluster in the clustering results to obtain minority clusters and majority clusters. The total data volume refers to the total number of sample credit data in the data cluster. For example, if the clustering results include three data clusters, namely data cluster 1, data cluster 2, and data cluster 3; wherein the total data volume of data cluster 1 is 50, the total data volume of data cluster 2 is 60, and the total data volume of data cluster 3 is 40, then the data cluster with the largest total data volume in the clustering results (i.e., data cluster 2) can be used as the majority cluster, and the remaining data clusters in the clustering results (i.e., data cluster 1 and data cluster 3) can be used as minority clusters.

[0111] Afterwards, an oversampling algorithm can be used to oversample the minority class cluster to obtain a new data cluster corresponding to the minority class cluster. The oversampling algorithm can be pre-set according to actual business needs. For example, the oversampling algorithm can be one of a random oversampling algorithm, a synthetic minority class oversampling algorithm (Synthetic Minority Oversampling Technique, SMOTE) and an adaptive synthetic sampling algorithm (Adaptive Synthetic Sampling, ADASYN). Based on the above example, the synthetic minority class oversampling algorithm can be used to oversample data cluster 1 and data cluster 3 respectively to obtain a new data cluster corresponding to data cluster 1 and a new data cluster corresponding to data cluster 3.

[0112] Afterwards, the chain equation multiple interpolation algorithm is used to perform data interpolation on the new data clusters corresponding to the majority cluster and the minority cluster respectively to obtain the training sample set.

[0113] It can be understood that, based on the clustering results, the minority clusters and majority clusters in the key credit data set are determined, which can quickly screen out the minority clusters in the key credit data set; then, the minority clusters are oversampled to obtain new data clusters corresponding to the minority clusters, which increases the amount of data in the minority clusters and solves the data category imbalance problem in the key credit data set; then, the chain equation multiple interpolation algorithm is used to perform data interpolation on the new data clusters corresponding to the majority clusters and the minority clusters, respectively, to obtain a training sample set, solving the data missing problem in the key credit data set, thereby improving the data quality of the training sample set. In the process of solving the data missing problem in the key credit data set, the data category imbalance problem in the key credit data set is also solved simultaneously, avoiding further amplification of the interpolation process error during the sampling process, thereby reducing the risk of overfitting during the model training process, improving data utilization, and thus improving the robustness of the target evaluation model in the face of different situations.

[0114] S305: Train the credit risk rating assessment model based on the Adaboost algorithm according to the training sample set to obtain a target assessment model.

[0115] S306: Obtain the credit data of the target object to be evaluated.

[0116] S307: Input the credit data to be evaluated into the target evaluation model to obtain the credit risk level corresponding to the target object.

[0117] The technical solution of the embodiment of the present invention is as follows: obtaining a sample credit data set; performing feature selection on the sample credit data set to obtain a key credit data set; clustering the key credit data set based on the fuzzy C-means clustering algorithm to obtain a clustering result; using the chain equation multiple interpolation algorithm according to the clustering result, performing data interpolation on the key credit data set to obtain a training sample set; training a credit risk rating assessment model based on the Adaboost algorithm according to the training sample set to obtain a target assessment model; obtaining the credit data to be assessed of the target object; inputting the credit data to be assessed into the target assessment model to obtain the credit risk rating corresponding to the target object. The above technical solution, after obtaining the sample credit data set, obtains the training sample set by performing a series of data preprocessing such as feature selection, fuzzy clustering, minority oversampling and multiple interpolation on the sample credit data set, thereby solving the data category imbalance and data missing problems in the training data when training the credit risk rating assessment model, reducing the data volume of the training data when training the credit risk rating assessment model, and thus improving the data quality of the training data when training the credit risk rating assessment model. While addressing the missing data problem, the data imbalance problem is simultaneously addressed, preventing further errors from the interpolation process during the sampling process. This reduces the risk of overfitting and improves data utilization. Subsequently, a credit risk assessment model based on the Adaboost algorithm is trained on the training sample set to obtain a target assessment model. This reduces computing resource consumption and training time during the training of the credit risk assessment model, improving data processing efficiency during training. Furthermore, it also improves the accuracy and reliability of the target assessment model. The target assessment model is then used to process the credit data of the target subject to be assessed, resulting in a more accurate credit risk rating for the target subject. This improves the accuracy of user credit risk assessments and reduces credit risk in the credit operations of banks and other financial institutions. This allows financial institutions to adjust credit limits and terms in a targeted manner, avoiding excessive concentration on high-risk projects and thus reducing the overall risk of their credit assets. Furthermore, it eliminates the need for credit approval personnel to individually review each borrower's loan application materials, thereby improving loan approval efficiency and shortening processing time for borrowers.

[0118] Example 4

[0119] Figure 4 This is a schematic diagram of the structure of a credit risk rating assessment device provided by the fourth embodiment of the present invention. This embodiment is applicable to the situation of assessing the credit risk rating of customers of financial institutions. The device can be implemented in the form of hardware and / or software and can be configured in an electronic device. Figure 4 As shown, the device includes:

[0120] The credit data acquisition module 401 is used to acquire the credit data of the target object.

[0121] The credit risk rating assessment module 402 is used to input the credit data to be assessed into the target assessment model to obtain the credit risk rating corresponding to the target object; wherein the target assessment model is obtained by training the credit risk rating assessment model based on the Adaboost algorithm based on the sample credit data set.

[0122] The technical solution of the embodiment of the present invention obtains the credit data to be evaluated of the target object; inputs the credit data to be evaluated into the target evaluation model to obtain the credit risk level corresponding to the target object; wherein the target evaluation model is obtained by training a credit risk level evaluation model based on the Adaboost algorithm based on a sample credit data set. The above technical solution processes the credit data to be evaluated of the target object through the target evaluation model to obtain the credit risk level corresponding to the target object, thereby improving the accuracy of the credit risk assessment of the user and reducing the credit risk in the credit business of financial institutions such as banks. At the same time, there is no need for credit approval personnel to review the borrower's loan application materials one by one, thereby improving the loan approval efficiency of banks and other financial institutions and shortening the business processing time of the borrower.

[0123] Optionally, the device further includes a target evaluation model determination module, wherein the target evaluation model determination module includes:

[0124] A sample credit data set acquisition unit, used to acquire a sample credit data set;

[0125] A key credit data set determination unit, used to perform feature selection on the sample credit data set to obtain the key credit data set;

[0126] The target assessment model determination unit is used to construct a training sample set based on the key credit data set, and train the credit risk rating assessment model based on the Adaboost algorithm based on the training sample set to obtain the target assessment model.

[0127] Optionally, a target assessment model determination unit includes:

[0128] A clustering result determination subunit is used to perform clustering processing on the key credit data set based on the fuzzy C-means clustering algorithm to obtain a clustering result;

[0129] The training sample set determines the subunits, which are used to perform data interpolation processing on the key credit data set based on the clustering results using the chain equation multiple interpolation algorithm to obtain the training sample set.

[0130] Optionally, the training sample set determines a subunit, specifically for:

[0131] Based on the clustering results, determine the minority and majority clusters in the key credit data set;

[0132] Oversampling is performed on the minority cluster to obtain a new data cluster corresponding to the minority cluster;

[0133] The chain equation multiple interpolation algorithm is used to perform data interpolation on the new data clusters corresponding to the majority cluster and the minority cluster respectively to obtain the training sample set.

[0134] Optionally, based on the clustering results, minority clusters and majority clusters in the key credit data set are determined. Specifically, the key credit data set is classified according to the total data volume of each data cluster in the clustering results to obtain minority clusters and majority clusters.

[0135] Optional, key credit data set determination unit, specifically used for:

[0136] determining a first data feature and a second data feature in a sample credit data set;

[0137] Calculating the correlation between each sub-data feature in the second data feature and the first data feature based on the Pearson correlation coefficient;

[0138] Based on the correlation between each sub-data feature in the second data feature and the first data feature, feature selection is performed on the sample credit data set to obtain a key credit data set.

[0139] The credit risk level assessment device provided in the embodiment of the present invention can execute the credit risk level assessment method provided in any embodiment of the present invention, and has the corresponding functional modules and beneficial effects for executing each credit risk level assessment method.

[0140] According to an embodiment of the present invention, the present invention further provides an electronic device, a readable storage medium and a computer program product.

[0141] Example 5

[0142] Figure 5 A schematic diagram of the structure of an electronic device 10 that can be used to implement an embodiment of the present invention is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital processing, cellular phones, smart phones, wearable devices (such as helmets, glasses, watches, etc.) and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present invention described and / or claimed herein.

[0143] like Figure 5 As shown, the electronic device 10 includes at least one processor 11, and a memory connected to the at least one processor 11, such as a read-only memory (ROM) 12, a random access memory (RAM) 13, etc., wherein the memory stores a computer program that can be executed by the at least one processor, and the processor 11 can perform various appropriate actions and processes according to the computer program stored in the read-only memory (ROM) 12 or the computer program loaded from the storage unit 18 to the random access memory (RAM) 13. Various programs and data required for the operation of the electronic device 10 can also be stored in the RAM 13. The processor 11, ROM 12 and RAM 13 are connected to each other via a bus 14. An input / output (I / O) interface 15 is also connected to the bus 14.

[0144] Multiple components in the electronic device 10 are connected to the I / O interface 15, including an input unit 16, such as a keyboard, a mouse, etc.; an output unit 17, such as various types of displays, speakers, etc.; a storage unit 18, such as a magnetic disk, an optical disk, etc.; and a communication unit 19, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 19 allows the electronic device 10 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.

[0145] Processor 11 can be any general-purpose and / or specialized processing component with processing and computing capabilities. Some examples of processor 11 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various specialized artificial intelligence (AI) computing chips, various processors running machine learning model algorithms, digital signal processors (DSPs), and any other suitable processor, controller, microcontroller, etc. Processor 11 executes the various methods and processes described above, such as the credit risk rating assessment method.

[0146] In some embodiments, the credit risk rating assessment method may be implemented as a computer program tangibly contained in a computer-readable storage medium, such as storage unit 18. In some embodiments, part or all of the computer program may be loaded and / or installed on electronic device 10 via ROM 12 and / or communication unit 19. When the computer program is loaded into RAM 13 and executed by processor 11, one or more steps of the credit risk rating assessment method described above may be performed. Alternatively, in other embodiments, processor 11 may be configured to execute the credit risk rating assessment method in any other appropriate manner (e.g., by means of firmware).

[0147] Various embodiments of the systems and techniques described herein can be implemented in digital electronic circuit systems, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), system-on-chip systems (SOCs), programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include being implemented in one or more computer programs that are executable and / or interpreted on a programmable system that includes at least one programmable processor, which can be a special purpose or general purpose programmable processor that can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit data and instructions to the storage system, the at least one input device, and the at least one output device.

[0148] Computer programs for implementing the methods of the present invention may be written in any combination of one or more programming languages. These computer programs may be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device, such that when the computer program is executed by the processor, the functions / operations specified in the flowcharts and / or block diagrams are implemented. The computer program may be executed entirely on the machine, partially on the machine, as a stand-alone software package, partially on the machine and partially on a remote machine, or entirely on a remote machine or server.

[0149] In the context of the present invention, computer-readable storage media can be tangible media that can contain or store a computer program for use with an instruction execution system, device or equipment or used in combination with an instruction execution system, device or equipment. Computer-readable storage media can include but are not limited to electronic, magnetic, optical, electromagnetic, infrared or semiconductor systems, devices or equipment, or any suitable combination of the foregoing. Alternatively, computer-readable storage media can be machine-readable signal media. More specific examples of machine-readable storage media can include electrical connections based on one or more lines, portable computer disks, hard disks, random access memories (RAM), read-only memories (ROM), erasable programmable read-only memories (EPROM or flash memory), optical fibers, portable compact disk read-only memories (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.

[0150] To provide interaction with a user, the systems and techniques described herein can be implemented on an electronic device having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user can provide input to the electronic device. Other types of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).

[0151] The systems and techniques described herein can be implemented in a computing system that includes back-end components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes front-end components (e.g., a user computer with a graphical user interface or web browser through which a user can interact with implementations of the systems and techniques described herein), or a computing system that includes any combination of such back-end components, middleware components, or front-end components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include: a local area network (LAN), a wide area network (WAN), a blockchain network, and the Internet.

[0152] A computing system may include clients and servers. The clients and servers are typically remote from each other and typically interact via a communication network. This client-server relationship arises through computer programs running on the respective computers, creating a client-server relationship. The server may be a cloud server, also known as a cloud computing server or cloud host. This server is a hosting product within the cloud computing service ecosystem that addresses the management difficulties and limited scalability of traditional physical hosting and VPS services.

[0153] It should be understood that the various forms of the processes shown above can be used to reorder, add, or delete steps. For example, the steps described in the present invention can be performed in parallel, sequentially, or in a different order, as long as the desired results of the technical solution of the present invention can be achieved. This is not limited herein.

[0154] The above specific embodiments do not limit the scope of protection of the present invention. Those skilled in the art will appreciate that various modifications, combinations, sub-combinations, and substitutions may be made based on design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of the present invention are intended to be included within the scope of protection of the present invention.

Claims

1. A credit risk rating assessment method, characterized in that: include: Obtain the credit data of the target object to be evaluated; Inputting the credit data to be evaluated into a target evaluation model to obtain a credit risk level corresponding to the target object; wherein the target evaluation model is obtained by training a credit risk level evaluation model based on an Adaboost algorithm based on a sample credit data set; The determination process of the target evaluation model is as follows: Get a sample credit dataset; determining a first data feature and a second data feature in the sample credit data set; Calculating the correlation between each sub-data feature in the second data feature and the first data feature based on the Pearson correlation coefficient; performing feature selection on the sample credit dataset based on the correlation between each sub-data feature in the second data feature and the first data feature to obtain a key credit dataset; Based on the fuzzy C-means clustering algorithm, clustering processing is performed on the key credit data set to obtain a clustering result; Based on the clustering results, a chain equation multiple interpolation algorithm is used to perform data interpolation processing on the key credit data set to obtain a training sample set; The credit risk rating assessment model based on the Adaboost algorithm is trained according to the training sample set to obtain a target assessment model.

2. The method according to claim 1, characterized in that According to the clustering results, a chain equation multiple interpolation algorithm is used to perform data interpolation processing on the key credit data set to obtain a training sample set, including: Determining a minority cluster and a majority cluster in the key credit data set based on the clustering results; Oversampling the minority cluster to obtain a new data cluster corresponding to the minority cluster; A chain equation multiple interpolation algorithm is used to perform data interpolation processing on the new data clusters corresponding to the majority cluster and the minority cluster respectively to obtain a training sample set.

3. The method according to claim 2, characterized in that Determining the minority cluster and the majority cluster in the key credit data set according to the clustering result includes: The key credit data set is classified according to the total data volume of each data cluster in the clustering result to obtain minority clusters and majority clusters.

4. A credit risk level assessment device, characterized in that: include: A target assessment model determination module is used to obtain a sample credit dataset; determining a first data feature and a second data feature in the sample credit data set; calculating, based on the Pearson correlation coefficient, the correlation between each sub-data feature in the second data feature and the first data feature; performing feature selection on the sample credit dataset based on the correlation between each sub-data feature in the second data feature and the first data feature to obtain a key credit dataset; performing clustering processing on the key credit dataset based on a fuzzy C-means clustering algorithm to obtain a clustering result; and performing data interpolation processing on the key credit dataset using a chain equation multiple interpolation algorithm based on the clustering result to obtain a training sample set; Training a credit risk rating assessment model based on the Adaboost algorithm according to the training sample set to obtain a target assessment model; A credit data acquisition module for obtaining credit data to be evaluated, used to obtain the credit data to be evaluated of the target object; The credit risk rating assessment module is used to input the credit data to be assessed into the target assessment model to obtain the credit risk rating corresponding to the target object.

5. An electronic device, characterized in that: The electronic device comprises: at least one processor; and a memory communicatively connected to the at least one processor; wherein, The memory stores a computer program executable by the at least one processor, and the computer program is executed by the at least one processor so that the at least one processor can execute the credit risk rating assessment method according to any one of claims 1 to 3.

6. A computer-readable storage medium, characterized in that The computer-readable storage medium stores computer instructions, and the computer instructions are used to enable a processor to implement the credit risk rating assessment method according to any one of claims 1 to 3 when executed.

7. A computer program product, comprising a computer program, wherein when the computer program is executed by a processor, the computer program implements the credit risk rating assessment method according to any one of claims 1 to 3.

Citation Information

Patent Citations

  • Enterprise credit risk security assessment method and device, equipment and storage medium

    CN116757834A

  • BP_adaboost model-based method and system for predicting credit card user default

    WO2018090657A1