Loan application processing method and apparatus, electronic device, and storage medium

By combining a credit rating model with first and second probability prediction models, the approval result of a loan application is determined, which solves the problem that the credit rating model cannot objectively process loan applications and achieves objectivity and accuracy in loan approval results.

CN117291707BActive Publication Date: 2026-07-21CHINA CONSTRUCTION BANK +1
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
CHINA CONSTRUCTION BANK
Filing Date
2023-09-13
Publication Date
2026-07-21

AI Technical Summary

Technical Problem

In existing technologies, credit rating models for loan applications cannot objectively handle applications that cannot be approved, resulting in subjective loan approval results. The application of rules based on expert experience is also highly subjective.

Method used

After obtaining a credit score through a credit rating model, the first probability prediction model is used to determine the predicted probability of a user belonging to multiple approval methods. Based on the predicted probabilities, a target group is determined. The second probability prediction model of the target group is used to determine the probability of default or non-default after the loan. Finally, the approval result of the loan application is determined based on the predicted probabilities and the approval threshold.

Benefits of technology

It improves the objectivity and accuracy of loan application approval results, enhances the accuracy of user credit assessment through multi-model prediction, and ensures the fairness of the approval results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117291707B_ABST
    Figure CN117291707B_ABST
Patent Text Reader

Abstract

The present disclosure provides a loan application processing method and device, electronic equipment and storage medium, wherein the method comprises: performing credit prediction on a user applying for a loan based on a credit evaluation model to obtain a credit score corresponding to the user; if the credit score is within a preset range, determining a first prediction probability that the user belongs to each first label by a first probability prediction model; determining a target cluster to which the user belongs according to each first prediction probability corresponding to the user; determining a second prediction probability that the user belongs to a second label by using a second probability prediction model corresponding to the target cluster; and determining an approval result of the loan application of the user according to the second prediction probability and an approval threshold value corresponding to the second probability prediction model. Thus, for users whose approval results cannot be determined by the credit evaluation model, the first probability prediction model and the second probability prediction model can be used to assist the credit evaluation model in processing the loan applications of these users, thereby ensuring the objectivity of the loan application approval result.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to the fields of artificial intelligence and risk control technology, and in particular to a loan application processing method, apparatus, electronic device and storage medium. Background Technology

[0002] When customers apply for loans, credit scoring models can be used to screen customers and assess default risk. However, some loan applications still cannot be approved through these models. Related technologies typically rely on rules derived from expert experience to approve these customers. However, the expert experience used in these rules is often subjective, making the loan approval outcome subjective as well. Summary of the Invention

[0003] This disclosure provides a loan application processing method, apparatus, electronic device, and storage medium. The technical solution of this disclosure is as follows:

[0004] According to a first aspect of the present disclosure, a loan application processing method is provided, comprising:

[0005] Credit prediction is performed on loan applicants based on a credit rating model to obtain their corresponding credit scores.

[0006] If the credit score is within a preset range, the first probability prediction model is used to determine the first predicted probability of the user belonging to each first label, where the first label is used to indicate the approval method and the corresponding approval result.

[0007] Based on the first predicted probability corresponding to each user, determine the target group to which the user belongs;

[0008] Using the second probability prediction model corresponding to the target group, the second predicted probability of a user belonging to the second label is determined, where the second label is used to indicate whether the user defaults after the loan or does not default after the loan.

[0009] The approval result of the user's loan application is determined based on the second predicted probability and the approval threshold corresponding to the second probability prediction model.

[0010] According to a second aspect of the present disclosure, a loan application processing apparatus is provided, comprising:

[0011] The first acquisition module is used to predict the creditworthiness of loan applicants based on a credit rating model in order to obtain the corresponding credit score of the user.

[0012] The first determining module is used to determine the first predicted probability of a user belonging to each first label by using a first probability prediction model when the credit score is within a preset range. The first label is used to indicate the approval method and the corresponding approval result.

[0013] The second determining module is used to determine the target group to which the user belongs based on the first predicted probabilities corresponding to the user.

[0014] The third determination module is used to determine the second predicted probability of a user belonging to the second label by using the second probability prediction model corresponding to the target group. The second label is used to indicate whether the user defaults after the loan or does not default after the loan.

[0015] The fourth determination module is used to determine the approval result of the user's loan application based on the second predicted probability and the approval threshold corresponding to the second probability prediction model.

[0016] According to a third aspect of the present disclosure, an electronic device is provided, comprising: a processor; and a memory for storing processor-executable instructions; wherein the processor is configured to execute the instructions to implement the method described in the above embodiments of the present disclosure.

[0017] According to a fourth aspect of the present disclosure, a computer-readable storage medium is provided that, when instructions in the computer-readable storage medium are executed by a processor of an electronic device, enables the electronic device to perform the methods described in the above embodiments of the present disclosure.

[0018] According to a fifth aspect of the present disclosure, a computer program product is provided, comprising: a computer program that, when executed by a processor, implements the method described in the above embodiments of the present disclosure.

[0019] The technical solutions provided by the embodiments of this disclosure offer at least the following beneficial effects: For users whose loan applications cannot be approved by the credit rating model, a first probability prediction model and a second probability prediction model can be used to assist the credit rating model in processing these users' loan applications, ensuring the objectivity of the loan application approval results. Furthermore, by using the prediction results of the first probability prediction model to determine the target group to which the user belongs, and using the second probability prediction model corresponding to the target group to determine the probability that the user belongs to the second label, the accuracy of the prediction results can be improved, thereby improving the accuracy of the approval results.

[0020] It should be understood that the above general description and the following detailed description are exemplary and explanatory only, and are not intended to limit this disclosure. Attached Figure Description

[0021] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this disclosure and, together with the description, serve to explain the principles of this disclosure, and are not intended to unduly limit this disclosure.

[0022] Figure 1 This is a schematic flowchart of the loan application processing method shown in the first embodiment of this disclosure.

[0023] Figure 2 This is a schematic flowchart of the loan application processing method shown in the second embodiment of this disclosure;

[0024] Figure 3 This is a schematic flowchart of the loan application processing method shown in the third embodiment of this disclosure;

[0025] Figure 4 This is a schematic diagram of the structure of the loan application processing device shown in the fourth embodiment of this disclosure;

[0026] Figure 5 This is a schematic diagram of the structure of an electronic device shown in an exemplary embodiment of the present disclosure. Detailed Implementation

[0027] To enable those skilled in the art to better understand the technical solutions of this disclosure, the technical solutions in the embodiments of this disclosure will be clearly and completely described below with reference to the accompanying drawings.

[0028] It should be noted that the terms "first," "second," etc., used in the specification, claims, and accompanying drawings of this disclosure are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of this disclosure described herein can be implemented in orders other than those illustrated or described herein. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with this disclosure. Rather, they are merely examples of apparatuses and methods consistent with some aspects of this disclosure as detailed in the appended claims.

[0029] It should be noted that the acquisition, storage, use, and processing of data in this disclosed technical solution comply with the relevant provisions of national laws and regulations and do not violate public order and good morals.

[0030] When customers apply for loans, credit scoring models can be used to screen customers and assess default risk. However, some loan applications still cannot be approved through these models. Related technologies typically rely on rules derived from expert experience to approve these customers. However, the expert experience used in these rules is often subjective, making the loan approval outcome subjective as well.

[0031] Therefore, in response to at least one of the aforementioned problems, this disclosure provides a loan application processing method, apparatus, electronic device, and storage medium.

[0032] The following description, with reference to the accompanying drawings, outlines a loan application processing method, apparatus, electronic device, and storage medium according to embodiments of the present disclosure. Figure 1This is a schematic flowchart of the loan application processing method shown in the first embodiment of this disclosure.

[0033] This disclosure illustrates the example of a loan application processing method configured in a loan application processing device, which can be applied to any electronic device to enable the electronic device to perform loan application processing functions.

[0034] Among them, electronic devices can be any device with computing capabilities, such as personal computers, mobile terminals, servers, etc. Mobile terminals can be hardware devices with various operating systems, touch screens and / or displays, such as in-vehicle devices, mobile phones, tablets, personal digital assistants, wearable devices, etc.

[0035] like Figure 1 As shown, the loan application processing method may include:

[0036] Step 101: Based on the credit rating model, perform credit prediction on the loan applicant to obtain the corresponding credit score.

[0037] In this disclosure, for users applying for loans, a credit scoring model can be used to predict their creditworthiness based on their loan application data, resulting in a credit score. This loan application data may include the loan amount requested, the user's age, income status, and any prior loan records.

[0038] Credit scores can be used to indicate a user's creditworthiness. For example, a relatively low credit score indicates that the user has low creditworthiness, meaning that the user may have a higher risk of default (such as low repayment ability).

[0039] Step 102: If the credit score is within a preset range, determine the first predicted probability of the user belonging to each first label using the first probability prediction model.

[0040] In this disclosure, if the credit score is within a preset range, it can be assumed that the credit evaluation model cannot provide an approval result. In this case, the first probability prediction model can be used to determine the first prediction probability of the user belonging to each first label.

[0041] The first label can be used to indicate the approval method and the corresponding approval result. For example, there are four first labels: automatic approval using the credit rating model (hereinafter referred to as automatic approval), automatic approval using the credit rating model (hereinafter referred to as automatic approval failure), manual approval without a credit rating model (hereinafter referred to as manual approval), and manual approval without a credit rating model (hereinafter referred to as manual approval failure).

[0042] In this disclosure, the first probability prediction model can be trained using sample users who have applied for loans and the first label to which the sample users belong.

[0043] In this disclosure, if a user's credit score is greater than the upper limit of the preset range, the user's credit is considered good and the user's loan application is approved; if the credit score is less than the lower limit of the preset range, the user's credit is considered relatively low and the user's loan application is not approved.

[0044] For example, if the preset range is [a1, a2], and user A1's credit score is greater than a1 and less than a2, it means that the credit rating model cannot automatically give the approval result. Instead, the first probability prediction model can be used to predict the first prediction probability of user A1 belonging to each first label. If user A2's credit score is greater than a2, user A2's loan application is approved. If user A3's credit score is less than a1, user A3's loan application is not approved.

[0045] Step 103: Determine the target group to which the user belongs based on the first prediction probabilities corresponding to each user.

[0046] In this disclosure, the clustering can be obtained by clustering sample users who apply for loans, and sample users in the same cluster are similar.

[0047] In this disclosure, a sample user can be randomly selected from each subgroup, and the absolute value of the difference between the probability of the sample user belonging to each first label and the first predicted probability of the user belonging to the same first label can be calculated. The subgroup to which the sample user with the smallest average absolute value of the difference belongs can be taken as the subgroup of the user, i.e., the target subgroup.

[0048] Step 104: Using the second probability prediction model corresponding to the target cluster, determine the second predicted probability that the user belongs to the second label.

[0049] The second label can be used to indicate whether a loan has been defaulted on or not.

[0050] In this disclosure, each cluster has a corresponding second probability prediction model, which can be trained using feature data of sample users in the cluster. The feature data can include multiple feature dimensions, and different feature dimensions can be used for different types of users or different types of loan applications; for example, the feature dimensions used for individual users and corporate users can be different.

[0051] In this disclosure, user feature data can be input into a second probability prediction model corresponding to the target group for prediction, thereby obtaining a second predicted probability that the user belongs to the second label.

[0052] The second probability prediction model can output the probability that a user will not default after taking out a loan and the probability that a user will default after taking out a loan, or the second probability prediction model can output the probability that a user will not default after taking out a loan and calculate the probability that a user will default after taking out a loan, or the second probability prediction model can output the probability that a user will default after taking out a loan and calculate the probability that a user will not default after taking out a loan. This disclosure does not limit the specific model.

[0053] Step 105: Determine the approval result of the user's loan application based on the second predicted probability and the approval threshold corresponding to the second probability prediction model.

[0054] In this disclosure, each subgroup has a corresponding second probability prediction model, and the second probability prediction model has a corresponding approval threshold.

[0055] In this disclosure, the second predicted probability can be compared with the approval threshold corresponding to the second probability prediction model of the target cluster, and the approval result of the user's loan application can be determined based on the comparison result.

[0056] As one implementation method, the second label is used to indicate that there is no default after the loan is issued. If the second predicted probability is greater than or equal to the approval threshold, the approval result of the user's loan application can be determined as approved. If the second predicted probability is less than the approval threshold, the approval result of the user's loan application can be determined as not approved.

[0057] As another implementation, the second label is used to indicate loan default. If the second predicted probability is less than or equal to the approval threshold, the approval result of the user's loan application can be determined as approved. If the second predicted probability is greater than the approval threshold, the approval result of the user's loan application can be determined as unapproved.

[0058] It should be noted that the approval result when the second predicted probability equals the approval threshold can be either approval or disapproval, and can be set according to actual needs. This disclosure does not limit this.

[0059] In this embodiment, for users whose loan application approval result cannot be determined by the credit rating model, a first probability prediction model can be used to determine the first predicted probability of the user belonging to each first label. Then, based on the first predicted probability, the target group to which the user belongs can be determined. Based on the second probability prediction model corresponding to the target group, a second predicted probability of the user belonging to the second label can be determined. Finally, based on the second predicted probability and the approval threshold corresponding to the second probability prediction model, the approval result of the user's loan application can be determined. Therefore, for users whose loan application approval result cannot be determined by the credit rating model, the first and second probability prediction models can be used to assist the credit rating model in processing these users' loan applications, ensuring the objectivity of the loan application approval result. Furthermore, by using the prediction result of the first predicted probability to determine the target group to which the user belongs, and using the second probability prediction model corresponding to the target group to determine the probability of the user belonging to the second label, the accuracy of the prediction result can be improved, thereby improving the accuracy of the approval result.

[0060] Figure 2 This is a flowchart illustrating the loan application processing method shown in the second embodiment of this disclosure.

[0061] like Figure 2 As shown, the loan application processing method includes:

[0062] Step 201: Based on the credit rating model, perform credit prediction on the loan applicant to obtain the corresponding credit score.

[0063] Step 202: If the credit score is within a preset range, determine the first predicted probability of the user belonging to each first label using the first probability prediction model.

[0064] In this disclosure, steps 201-202 can be implemented in any of the embodiments of this disclosure. This disclosure does not limit them and will not elaborate further.

[0065] Step 203: Vectorize the user according to each first predicted probability to obtain the first vector corresponding to the user.

[0066] In this disclosure, each first predicted probability can be arranged into a vector according to a certain label order to obtain the first vector corresponding to the user.

[0067] For example, if there are four first labels, namely label 1, label 2, label 3, and label 4, and the first predicted probabilities of a user belonging to these four labels are p1, p2, p3, and p4 respectively, then the first vector can be (p1, p2, p3, p4).

[0068] Step 204: Determine the target cluster from the multiple clusters based on the distances between the first vector and the cluster centers of the multiple clusters.

[0069] The cluster center can be determined based on the probability that each sample user in the cluster belongs to each first label. The cluster center can be represented by a vector, and the number of elements in the vector corresponding to the cluster center is the same as the number of first labels.

[0070] In this disclosure, the distance between the first vector and the cluster center of each cluster can be calculated, and the cluster with the smallest distance can be identified as the target cluster. For example, the Euclidean distance between the first vector and the cluster center of each cluster can be calculated, and the cluster corresponding to the minimum Euclidean distance can be identified as the target cluster.

[0071] Step 205: Using the second probability prediction model corresponding to the target cluster, determine the second predicted probability that the user belongs to the second label.

[0072] Step 206: Determine the approval result of the user's loan application based on the second predicted probability and the approval threshold corresponding to the second probability prediction model.

[0073] In this disclosure, steps 205-206 can be implemented in any of the embodiments of this disclosure. This disclosure does not limit them and will not elaborate further.

[0074] In this embodiment of the disclosure, the first predicted probability of a user belonging to each first label can be vectorized to obtain a first vector, and the target group to which the user belongs can be determined based on the distance between the first vector and the group center of each group. This not only improves accuracy but is also simple and convenient.

[0075] Figure 3 This is a schematic flowchart of the loan application processing method shown in the third embodiment of this disclosure.

[0076] like Figure 3 As shown, the loan application processing method includes:

[0077] Step 301: Based on the credit prediction results of multiple sample users applying for loans using the credit rating model, determine the first label to which each sample user belongs.

[0078] In this disclosure, a credit rating model can be used to predict the creditworthiness of each sample user to obtain a credit score. If the credit score is greater than the upper limit of a preset range, the sample user's loan application is approved. If the credit score is less than the lower limit of the preset range, the sample user's loan application is rejected. If the credit score is within the preset range, the sample user's loan application can be manually reviewed, with the result being either approval or rejection. Therefore, based on the loan approval method and result for each sample user, the user's primary tag can be determined.

[0079] For example, there are four primary labels: label 1, label 2, label 3, and label 4. Label 1 indicates that the automatic approval has been approved, label 2 indicates that the automatic approval has not been approved, label 3 indicates that the manual approval has been approved, and label 4 indicates that the manual approval has not been approved. If a sample user's loan application is approved by the credit rating model, the primary label of the sample user can be determined to be label 1.

[0080] Step 302: Train the initial first probability prediction model based on the first label of each sample user and the feature data corresponding to each sample user to obtain the first probability prediction model.

[0081] In this disclosure, the feature data of sample users can be input into an initial first probability prediction model to obtain the predicted labels of the sample users. The model loss is determined based on the difference between the predicted labels and the first label. The parameters of the initial first probability prediction model are adjusted based on the model loss. The model with adjusted parameters is then trained until the conditions for training termination are met, thus obtaining the first probability prediction model.

[0082] The training termination condition can be that the model loss is less than a preset threshold, or that the number of training iterations reaches a preset number, etc., which can be determined according to actual needs. This disclosure does not limit this.

[0083] For example, customers applying for loans can be categorized and tagged based on a credit rating model, as shown in Table 1 below.

[0084] Indicates customer category and tag

[0085]

[0086] Here, the set {AutoAccept} represents the group of customers who are automatically approved based on the credit rating model online_model. The judgment logic is that if the credit score calculated by the online_model is greater than the automatic approval score threshold baseScoreUp, the first label value of the customers in the set {AutoAccept} is assigned to 1.

[0087] The set {AutoRefuse} represents the group of customers whose automatic approval was rejected based on the credit rating model online_model. The judgment logic is that if the credit score calculated by the customer through online_model is less than the automatic rejection score threshold baseScoreDown, the first label value of the customer in the set {AutoRefuse} is assigned to 2.

[0088] The set {ManulAccept} represents customers who cannot be automatically approved or rejected by the credit rating model online_model and require manual intervention to determine approval. The judgment logic is that the credit score calculated by the online_model is greater than or equal to the automatic rejection score threshold baseScoreDown and less than or equal to the automatic pass score threshold baseScoreUp. If the customer is approved after manual review, the first label value of the customer in the set {ManulAccept} is assigned to 3.

[0089] The set {ManulRefuse} represents a customer group whose approval or rejection cannot be automatically determined by the credit rating model online_model and requires manual intervention. The judgment logic is that if the credit score calculated by the online_model is greater than or equal to the automatic rejection score threshold baseScoreDown and less than or equal to the automatic pass score threshold baseScoreUp, and the approval is determined to be unsuccessful after manual review, the first label value of the customer in the set {ManulRefuse} is assigned to 4.

[0090] Merging the aforementioned customer sets {AutoAccept}, {AutoRefuse}, {ManulAccept}, and {ManulRefuse} into a single customer set {MergeSample}, and combining this with customer characteristics and tags, we can generate the following Table 2:

[0091] Table 2. Example of a Customer Sample

[0092]

[0093] Among them, S i Let {MergeSample} represent the customers in the customer set, and let feature j represent some feature dimensions corresponding to the customer. Let feature j be the candidate independent variable and the label be the dependent variable. We will use common supervised classification algorithms (such as logistic regression, random forest, and the optimized distributed gradient boosting library XGBOOST) to build a first probability prediction model. This model can be used to determine the probability that a customer belongs to label 1, label 2, label 3, and label 4 respectively.

[0094] In this embodiment of the disclosure, a first label to which each sample user belongs can be determined based on a credit rating model. Then, based on the feature data of the sample users and the first label to which the sample users belong, a first probability prediction model can be trained, thereby enabling the prediction of the probability that a loan applicant belongs to each first label through the model.

[0095] In one embodiment of this disclosure, a first probability prediction model obtained through training can be used to predict each sample user to obtain a third prediction probability that each sample user belongs to each first label. Then, each sample user is vectorized according to the third prediction probability that the sample user belongs to each first label to obtain a second vector corresponding to each sample user. Then, the sample users are clustered according to the feature data of the sample users to obtain multiple clusters. And the cluster center of each cluster is determined according to the second vector corresponding to the sample users in each cluster.

[0096] It should be noted that the method for obtaining the second vector in this disclosure is similar to that for the first vector, so it will not be described again here.

[0097] When determining the cluster center of each cluster, the average of the third predicted probabilities of each sample user in the cluster belonging to the same first label can be calculated. The cluster center is obtained based on the average of each first label.

[0098] For example, a certain cluster includes N sample users, where N is an integer greater than 1, and the second vector of the i-th sample user is (p i1 p i2 p i3 p i4 (), where i is a positive integer less than or equal to N, and the cluster center of this cluster can be...

[0099] Taking the aforementioned customer set {MergeSample} as an example, the probability of each customer in the set {MergeSample} belonging to label 1, label 2, label 3, and label 4 is predicted using the first probability prediction model, accept_model. The results are shown in Table 3 below:

[0100] Table 3. Prediction Results of the First Probability Prediction Model

[0101] client Tag 1 Tag 2 Tag 3 Tag 4 <![CDATA[S1]]> <![CDATA[P 1,1 ]]> <![CDATA[P 1,2 ]]> <![CDATA[P 1,3 ]]> <![CDATA[P 1,4 ]]> <![CDATA[S2]]> <![CDATA[P 2,1 ]]> <![CDATA[P 2,2 ]]> <![CDATA[P 2,3 ]]> <![CDATA[P 2,4 ]]> <![CDATA[S3]]> <![CDATA[P 3,1 ]]> <![CDATA[P 3,2 ]]> <![CDATA[P 3,3 ]]> <![CDATA[P 3,4 ]]> … … … … … <![CDATA[S i ]]> <![CDATA[P i,1 ]]> <![CDATA[P i,2 ]]> <![CDATA[P i,3 ]]> <![CDATA[P i,4 ]]> … … … … … <![CDATA[S m ]]> <![CDATA[P m,1 ]]> <![CDATA[P m,2 ]]> <![CDATA[P m,3 ]]> <![CDATA[P m,4 ]]>

[0102] Among them, P i,1 P i,2 P i,3 and P i,4 They represent customer S respectively i The probabilities of a customer belonging to label 1, label 2, label 3, and label 4 are calculated, and the customer is vectorized using the predicted probabilities of belonging to label 1, label 2, label 3, and label 4 respectively, to obtain S. i = <P i,1 P i,2 P i,3 P i,4 >

[0103] All customers in the customer set {MergeSample} are vectorized. Based on the vectorization results, unsupervised clustering algorithms, such as KMEANS (K-means) and density-based clustering of applications with noise (DBSCAN), are used to construct a customer clustering model. The final clustering results are shown in Table 4 below.

[0104] Table 4 Clustering Results

[0105]

[0106] in, <AVG_P i,1 ,AVG_P i,2 ,AVG_P i,3 ,AVG_P i,4 > represents the center point of cluster {Cluster_i}, that is, the vector S corresponding to all customers in cluster {Cluster_i}. i = <P i,1 P i,2 P i,3 P i,4 The average value of >, i.e., AVG_P i,1 =AVG(P i,1 ), AVG_P i,2 =AVG(P i,2 ), AVG_P i,3 =AVG(P i,3 ), AVG_P i,4 =AVG(P i,4 ).

[0107] The union of the clusters is the set {MergeSample}, i.e., {MergeSample} = {Cluster_1}U{Cluster_2}U…U{Cluster_i}U…U{Cluster_n}, and the intersection of any two clusters is empty.

[0108] In this embodiment of the disclosure, a first probability prediction model can be used to determine the third predicted probability of each sample user belonging to each first label. The sample users are vectorized according to the third predicted probability of each sample user belonging to each first label in the cluster, and the sample users are clustered to obtain multiple clusters. The second vector of the sample users in the cluster is used to determine the cluster center of the cluster. Thus, the sample users are classified by clustering, and the cluster center is determined based on the prediction result of the first probability prediction model, which improves the accuracy.

[0109] In one embodiment of this disclosure, for each cluster, a third probability prediction model can be obtained by training an initial second probability prediction model based on the feature data of the first target sample users in the cluster and the second label to which the first target sample users belong. Here, the first target sample users can refer to sample users in the cluster whose first label is the first target label. The first target label can include both automatically approved and manually approved users; that is, the first target sample users refer to both automatically approved and manually approved sample users in the cluster.

[0110] After obtaining the third probability prediction model, the feature data of the second target sample users in the cluster can be input into the third probability prediction model to determine the fourth prediction probability that the second target sample users belong to the second label. Here, the second sample users can refer to sample users in the cluster whose first label is the second target label. The second target label can include both automatically approved and manually approved users; that is, the second target sample users can refer to sample users in the cluster whose automatic approval was rejected and whose manual approval was rejected.

[0111] For example, the second label is used to indicate that the loan was not defaulted. The third probability prediction model corresponding to the cluster can be used to predict the probability that the sample users whose loans were not approved automatically did not default, as well as the probability that the users whose loans were not approved manually did not default.

[0112] Next, the fourth predicted probability that the second target sample user belongs to the second label can be labeled as a weight, and the predicted probability that the second sample user belongs to another second label can also be determined and labeled as a weight. The second target sample users belonging to the second label, the second target sample users belonging to the other second label, and the remaining sample users in the cluster (excluding the second target sample users) are merged to obtain the training sample set corresponding to the cluster. The second label to which the sample belongs is trained using the feature dataset of the training samples in this training sample set, and the initial second probability prediction model is trained to obtain the second probability prediction model corresponding to the cluster. The training method of the second probability prediction model is similar to that of the first probability prediction model, and therefore will not be described in detail here.

[0113] Taking Table 4 as an example, the second label for automatically approved customers and the second label for manually approved customers can be determined based on whether the customers who were automatically approved and manually approved in each group defaulted on their loans. For example, for ease of understanding, the second label for customers who did not default after being instructed to take out a loan can be marked as "good," and the second label for customers who defaulted after being instructed to take out a loan can be marked as "bad," as shown in Table 5 below:

[0114] Table 5: Wide table of customers approved by automatic and manual approval.

[0115]

[0116] Among them, S i Let {Cluster_i} represent customers who have been automatically approved and customers who have been manually approved. Let feature j represent some feature dimensions corresponding to the customer. Let feature j be the candidate independent variable and the labels good and bad be the dependent variables. A third probabilistic prediction model good_model_A is constructed using supervised classification algorithms (such as logistic regression, random forest, XGBOOST, etc.).

[0117] The third probability prediction model good_model_A is used to predict the probability of customers who failed automatic approval and those who failed manual approval in cluster {Cluster_i}. For each customer, the probability of not defaulting after the loan P(Good) and the probability of defaulting after the loan P(Bad) can be calculated, where P(Good) = 1 - P(Bad).

[0118] Customers who failed automatic approval and those who failed manual approval in cluster {Cluster_i} can be copied into two copies. One copy is labeled as "Good" with a weight of P(Good), and the other copy is labeled as "Bad" with a weight of P(Bad).

[0119] Next, the weighted customers who failed automatic and manual approvals after copying and conversion are merged with the customers who passed automatic and manual approvals. Specifically, the weighted good customers from the customers who failed automatic and manual approvals are merged with the good customers from the customers who passed automatic and manual approvals. Then, the weighted bad customers from the customers who failed automatic and manual approvals are merged with the bad customers from the customers who passed automatic and manual approvals. A second probabilistic prediction model good_model_B is then constructed for the merged customers using a supervised classification algorithm (such as logistic regression, random forest, XGBoost, etc.). A second probabilistic prediction model good_model_B can be obtained for each customer group, as shown in Table 6 below:

[0120] Table 6 Prediction of Default Probability for Different Customer Groups

[0121] Group number Grouped Customers Second probability prediction model Group 1 {Cluster_1} good_model_B_1 Group 2 {Cluster_2} good_model_B_2 …. …. …. Grouping i {Cluster_i} good_model_B_i …. …. …. Group n {Cluster_n} good_model_B_n

[0122] In this embodiment of the disclosure, a third probability prediction model can be trained first based on the second label of the sample users whose automatic and manual approvals are approved in the group. The third probability prediction model is then used to predict the probability that the sample users whose automatic and manual approvals are not approved belong to the second label, so as to construct a training sample set. The accuracy of the second probability prediction model is improved by training the second probability prediction model through the training sample set.

[0123] It should be noted that, in this disclosure, the third probability prediction model can be directly determined as the second probability prediction model, or the weighted good customers among the customers who failed automatic approval and those who failed manual approval in the cluster can be merged with the users who passed automatic approval and those who passed manual approval to obtain a training sample set, and the second probability prediction model corresponding to the cluster can be trained using this training sample set. Alternatively, the weighted bad customers among the customers who failed automatic approval and those who failed manual approval in the cluster can be merged with the users who passed automatic approval and those who passed manual approval to obtain a training sample set, and the second probability prediction model corresponding to the cluster can be trained using this training sample set. This disclosure does not limit this approach.

[0124] In one embodiment of this disclosure, the approval threshold corresponding to the second probability prediction model can also be determined by utilizing the second probability prediction model corresponding to the cluster.

[0125] In this disclosure, for each cluster, the fifth predicted probability of each sample user in the cluster belonging to the second label can be determined by using the second probability prediction model corresponding to the cluster. Based on the order of the fifth predicted probability from high to low, the sample users in the cluster are divided into multiple bins with equal frequency.

[0126] As an example, sample users can be divided into groups according to a preset number of users, either in descending or ascending order of the fifth prediction probability. The lowest fifth prediction probability in each group can be used as the lower limit of the binning for that group, and the highest fifth prediction probability in each group can be used as the upper limit of the binning for that group.

[0127] For example, if a cluster contains 100 sample users, divided into 5 bins with equal frequency (i.e., 20 sample users per group), the bins can be sorted in descending order of the fifth prediction probability. The 1st to 20th sample users can be grouped together, the 21st to 40th together, the 41st to 60th together, the 61st to 80th together, and the 81st to 100th together. The bins for each group are determined based on the minimum and maximum fifth prediction probabilities within each group. For example, for the first group, positive infinity can be used as the upper bound probability, and the minimum fifth prediction probability of the first group as the lower bound probability. For the fifth group, negative infinity can be used as the lower bound probability, and the maximum fifth prediction probability of the first group as the upper bound probability. For other groups, the minimum fifth prediction probability within the group can be used as the lower bound probability, and the maximum fifth prediction probability within the group as the upper bound probability.

[0128] After obtaining multiple bins, for each bin, we can identify sample users whose fifth probability in the cluster is greater than the lower limit probability of the bin, and use them as the third target sample users. We can also determine the number of sample users who defaulted on their loans among the third target sample users in the cluster. The ratio of the number of sample users who defaulted on their loans to the total number of third target sample users is determined as the cumulative non-performing loan rate for the bin. The ratio of the number of sample users whose loan applications were approved to the total number of sample users in the cluster is determined as the cumulative approval rate for the bin. Based on the cumulative approval rate for the bin and the cumulative approval rate, we can determine the approval threshold.

[0129] When determining the approval threshold, the absolute value of the difference between the cumulative defect rate and the actual defect rate can be determined as the defect rate difference, the absolute value of the difference between the cumulative pass rate and the actual pass rate can be determined as the pass rate difference, and the lower limit probability of the box with the smallest sum of the defect rate difference and the pass rate difference can be determined as the approval threshold.

[0130] The actual approval rate and actual default rate can be obtained by statistically analyzing sample users who previously applied for loans based on a credit rating model. The actual approval rate can be the ratio of the sum of the number of sample users who were approved automatically and manually to the total number of sample users. The actual default rate can be the ratio of the number of sample users who defaulted on their loans after being approved automatically and manually to the sum of the number of sample users who were approved automatically and manually.

[0131] Taking the example of a user with the second tag indicating that they did not default on a loan, the process of determining the approval threshold is described below in conjunction with Table 5 above.

[0132] Based on the credit rating model online_model, the online customer acceptance rate (online_accept_rate) and the default rate (online_bad_rate) are statistically analyzed. The calculation formula is as follows:

[0133]

[0134]

[0135] Wherein, s1 represents the number of customers who were automatically approved, s2 represents the number of customers who were manually approved, s3 represents the number of customers who were not automatically approved, s4 represents the number of customers who were not manually approved, s11 represents the number of customers who defaulted on their loans after being automatically approved, and s21 represents the number of customers who defaulted on their loans after being manually approved.

[0136] The second probability prediction model good_model_B_i corresponding to cluster {Cluster_i} is used to predict the probability of good or bad customers for all customers in {Cluster_i}. The prediction results are sorted from high to low, and then divided into m bins according to equal frequency, where m is an adjustable parameter, as shown in Table 7 below:

[0137] Table 7. Statistics on Failure Rate and Pass Rate

[0138]

[0139] Among them, score i The sum_bad_rate represents the prediction result of the second probability prediction model good_model_B_i for customers in {Cluster_i}. i and sum_accept_rate i For binning [score] i score i-1 The corresponding cumulative defect rate and cumulative pass rate are calculated using the following formulas:

[0140]

[0141]

[0142] Among them, h i1 This indicates that the prediction result of model good_model_B_i in {Cluster_i} is greater than the score. i The number of customers who actually defaulted among h's customers i2 This indicates that the prediction result of model good_model_B_i in {Cluster_i} is greater than the score. i The total number of customers, h i3This indicates that the prediction result of model good_model_B_i in {Cluster_i} is greater than the score. i The number of customers whose applications were actually approved among the customers, h i This represents the number of customers in {Cluster_i}.

[0143] The absolute difference between the cumulative defect rate and online_bad_rate for each sub-box in Table 6, and the absolute difference between the cumulative pass rate and online_accept_rate, are shown in Table 8 below:

[0144] Table 8. Statistical Table of the Difference Between Failure Rate and Pass Rate

[0145] Where d_badi = abs(sum_bad_ratei - online_bad_rate)

[0146] d_accept i =abs(sum_accept_rate) i -online_accept_rate)

[0147] sum i =d_bad i +d_accept i

[0148] Then, sum the difference. i Sort the data in ascending order and take the minimum value min(sum). i The corresponding bin, let's assume the bin is [score]. i score i-1 If the score is not met, then take the score. i The second probabilistic prediction model good_model_B_i, corresponding to cluster {Cluster_i}, automatically passes the approval threshold.

[0149] Therefore, when the credit rating model fails to provide an approval result for a user's loan application, the probability that the user will not default after taking out the loan can be determined by using the good_model_B_i corresponding to the target segment to which the user belongs. If the probability is greater than the automatic approval threshold of good_model_B_i, the approval result of the user's loan application can be determined as approved. If the probability is less than or equal to the automatic approval threshold of good_model_B_i, the approval result of the user's loan application can be determined as not approved.

[0150] It should be noted that if the second label is used to indicate loan default, the above method can also be used to determine the second probability prediction model and the corresponding approval threshold for each subgroup.

[0151] Corresponding to the loan application processing method provided in the above embodiments, this disclosure also provides a loan application processing device. Since the loan application processing device provided in this disclosure corresponds to the loan application processing method provided in the above embodiments, the implementation method of the loan application processing method is also applicable to the loan application processing device provided in this disclosure, and will not be described in detail in this disclosure.

[0152] Figure 4 This is a schematic diagram of the structure of the loan application processing device shown in the fourth embodiment of this disclosure.

[0153] Reference Figure 4 The loan application processing device 100 may include:

[0154] The first acquisition module 410 is used to predict the credit of loan applicants based on a credit rating model in order to obtain the credit score of the user.

[0155] The first determining module 420 is used to determine the first predicted probability of a user belonging to each first label by means of a first probability prediction model when the credit score is within a preset range. The first label is used to indicate the approval method and the corresponding approval result.

[0156] The second determining module 430 is used to determine the target group to which the user belongs based on the first predicted probabilities corresponding to the user.

[0157] The third determining module 440 is used to determine the second predicted probability of a user belonging to the second label by using the second probability prediction model corresponding to the target group, wherein the second label is used to indicate whether the user defaults after the loan or does not default after the loan.

[0158] The fourth determining module 450 is used to determine the approval result of the user's loan application based on the second predicted probability and the approval threshold corresponding to the second probability prediction model.

[0159] Optionally, the second determining module 430 is used for:

[0160] The user is vectorized according to each of the first predicted probabilities to obtain the first vector corresponding to the user.

[0161] The target cluster is determined from the multiple clusters based on the distances between the first vector and the cluster centers of the multiple clusters.

[0162] Optionally, the second label is used to indicate that there has been no default after the loan was issued, and the fourth determination module 450 is used for:

[0163] If the second predicted probability is greater than or equal to the approval threshold, the approval result is determined as approval passed;

[0164] If the second predicted probability is less than the approval threshold, the approval result is determined to be "approval not approved".

[0165] Optionally, the second label is used to indicate post-loan default, and the fourth determination module 450 is used for:

[0166] If the second predicted probability is less than or equal to the approval threshold, the approval result is determined as approval passed;

[0167] If the second predicted probability is greater than the approval threshold, the approval result is determined to be "approval not approved".

[0168] Optionally, the device may further include:

[0169] The fifth determination module is used to determine the first label of each sample user based on the credit prediction results of multiple sample users applying for loans using the credit rating model.

[0170] The first training module is used to train the initial first probability prediction model based on the first label of each sample user and the feature data corresponding to each sample user, so as to obtain the first probability prediction model.

[0171] Optionally, the device may further include:

[0172] The second acquisition module is used to predict each sample user using the first probability prediction model to obtain the third prediction probability of each sample user belonging to each first label.

[0173] The third acquisition module is used to vectorize each sample user according to the third predicted probability of each sample user belonging to each first label, so as to obtain the second vector corresponding to each sample user.

[0174] The clustering module is used to cluster multiple sample users to obtain multiple groups;

[0175] The sixth determination module is used to determine the cluster center of each cluster based on the second vector corresponding to the sample users in each cluster.

[0176] Optionally, the device may further include:

[0177] The second training module is used to train the initial second probability prediction model to obtain the third probability prediction model for each group based on the feature data of the first sample user in the group and the second label to which the first target sample user belongs. The first target sample user refers to the sample user whose first label is the first target label.

[0178] The seventh determination module is used to determine the fourth prediction probability that the second target sample user in the cluster belongs to the second label using the third probability prediction model, wherein the second sample user refers to the sample user whose first label is the second target label;

[0179] The construction module is used to construct the training sample set corresponding to the cluster based on the second target sample user, the fourth predicted probability that the second target sample user belongs to the second label, the other sample users in the cluster except for the second target sample user, and the second label to which the other sample users belong.

[0180] The third training module is used to train the initial second probability prediction model using the feature data of the training samples in the training sample set and the second label to which the training samples belong, so as to obtain the second probability prediction model corresponding to the cluster.

[0181] Optionally, the device may further include:

[0182] The eighth determination module is used to determine the fifth prediction probability of each sample user in the cluster belonging to the second label by using the second probability prediction model corresponding to the cluster for each cluster.

[0183] The partitioning module is used to divide the sample users in the cluster into multiple bins based on the order of the fifth prediction probability.

[0184] The ninth determination module is used to determine the cumulative non-performing loan rate for each bin based on the ratio of the number of sample users who defaulted on their loans in the third target sample users to the number of third target sample users. Here, the third target sample users refer to sample users in the cluster whose fifth probability is greater than the lower limit probability of the bin.

[0185] The tenth determination module is used to determine the cumulative pass rate of a subgroup based on the ratio of the number of sample users whose loan applications have been approved in the third target sample users to the number of sample users in the subgroup.

[0186] The eleventh module is used to determine the approval threshold based on the cumulative pass rate and the cumulative pass rate corresponding to the sub-box.

[0187] In this embodiment, for users whose loan application approval result cannot be determined by the credit rating model, a first probability prediction model can be used to determine the first predicted probability of the user belonging to each first label. Then, based on the first predicted probability, the target group to which the user belongs can be determined. Based on the second probability prediction model corresponding to the target group, a second predicted probability of the user belonging to the second label can be determined. Finally, based on the second predicted probability and the approval threshold corresponding to the second probability prediction model, the approval result of the user's loan application can be determined. Therefore, for users whose loan application approval result cannot be determined by the credit rating model, the first and second probability prediction models can be used to assist the credit rating model in processing these users' loan applications, ensuring the objectivity of the loan application approval result. Furthermore, by using the prediction result of the first predicted probability to determine the target group to which the user belongs, and using the second probability prediction model corresponding to the target group to determine the probability of the user belonging to the second label, the accuracy of the prediction result can be improved, thereby improving the accuracy of the approval result.

[0188] In an exemplary embodiment, an electronic device is also proposed.

[0189] The electronic devices include:

[0190] processor;

[0191] Memory used to store processor-executable instructions;

[0192] The processor is configured to execute instructions to implement the loan application processing method as presented in any of the foregoing embodiments.

[0193] As an example, Figure 5 This is a schematic diagram of the structure of an electronic device 500 as shown in an exemplary embodiment of this disclosure, as follows: Figure 5 As shown, the above-mentioned electronic device 500 may further include:

[0194] The system includes a memory 510 and a processor 520, and a bus 530 connecting different components (including the memory 510 and the processor 520). The memory 510 stores a computer program, which, when executed by the processor 520, implements the loan application processing method described in this embodiment.

[0195] Bus 530 represents one or more of several bus architectures, including a memory bus or memory controller, a peripheral bus, a graphics acceleration port, a processor, or a local bus using any of the various bus architectures. Examples of these architectures include, but are not limited to, the Industry Standard Architecture (ISA) bus, the Micro Channel Architecture (MAC) bus, the Enhanced ISA bus, the Video Electronics Standards Association (VESA) local bus, and the Peripheral Component Interconnect (PCI) bus.

[0196] Electronic device 500 typically includes a variety of electronic device readable media. These media can be any available media that can be accessed by electronic device 500, including volatile and non-volatile media, removable and non-removable media.

[0197] Memory 510 may also include computer system readable media in the form of volatile memory, such as random access memory (RAM) 540 and / or cache memory 550. Electronic device 500 may further include other removable / non-removable, volatile / non-volatile computer system storage media. By way of example only, storage system 560 may be used to read and write non-removable, non-volatile magnetic media (… Figure 5 Not shown; usually referred to as a "hard drive"). Although Figure 5 Not shown, a disk drive for reading and writing to a removable non-volatile disk (e.g., a "floppy disk") and an optical disk drive for reading and writing to a removable non-volatile optical disk (e.g., a CD-ROM, DVD-ROM, or other optical media) may be provided. In these cases, each drive may be connected to bus 530 via one or more data media interfaces. Memory 510 may include at least one program product having a set (e.g., at least one) of program modules configured to perform the functions of the embodiments of this disclosure.

[0198] A program / utility 580 having a set (at least one) of program modules 570 may be stored in, for example, memory 510. Such program modules 570 include—but are not limited to—an operating system, one or more application programs, other program modules, and program data. Each or some combination of these examples may include an implementation of a network environment. Program modules 570 typically perform the functions and / or methods described in the embodiments of this disclosure.

[0199] Electronic device 500 can also communicate with one or more external devices 590 (e.g., keyboard, pointing device, display 591, etc.), and with one or more devices that enable a user to interact with electronic device 500, and / or with any device that enables electronic device 500 to communicate with one or more other computing devices (e.g., network card, modem, etc.). This communication can be performed via input / output (I / O) interface 592. Furthermore, electronic device 500 can also communicate with one or more networks (e.g., local area network (LAN), wide area network (WAN), and / or public networks, such as the Internet) via network adapter 593. As shown, network adapter 593 communicates with other modules of electronic device 500 via bus 530. It should be understood that, although not shown in the figures, other hardware and / or software modules can be used in conjunction with electronic device 500, including but not limited to: microcode, device drivers, redundant processing units, external disk drive arrays, RAID systems, tape drives, and data backup storage systems.

[0200] The processor 520 performs various functional applications and data processing by running programs stored in the memory 510.

[0201] It should be noted that the implementation process and technical principles of the electronic device in this embodiment are explained in the foregoing description of the loan application processing method of this disclosure embodiment, and will not be repeated here.

[0202] In an exemplary embodiment, a computer-readable storage medium including instructions is also provided, such as a memory including instructions that can be executed by a processor of an electronic device to perform the method proposed in any of the above embodiments. Optionally, the computer-readable storage medium may be a ROM, random access memory (RAM), CD-ROM, magnetic tape, floppy disk, and optical data storage device, etc.

[0203] In an exemplary embodiment, a computer program product is also provided, including a computer program / instructions, characterized in that the computer program / instructions, when executed by a processor, implement the method proposed in any of the above embodiments.

[0204] Other embodiments of this disclosure will readily occur to those skilled in the art upon consideration of the specification and practice of the invention disclosed herein. This disclosure is intended to cover any variations, uses, or adaptations of this disclosure that follow the general principles of this disclosure and include common knowledge or customary techniques in the art not disclosed herein. The specification and examples are to be considered exemplary only, and the true scope and spirit of this disclosure are indicated by the following claims.

[0205] It should be understood that this disclosure is not limited to the precise structures described above and shown in the accompanying drawings, and various modifications and changes can be made without departing from its scope. The scope of this disclosure is limited only by the appended claims.

Claims

1. A method for processing loan applications, characterized in that, include: Credit prediction is performed on loan applicants based on a credit rating model to obtain their corresponding credit scores. When the credit score is within a preset range, the first probability prediction model determines the first predicted probability of the user belonging to each first label. Here, the credit score being within the preset range means that the credit evaluation model cannot provide an approval result. The first label is used to indicate the approval method and the corresponding approval result. The first label includes automatic approval using the credit evaluation model, automatic approval using the credit evaluation model, manual approval without approval result provided by the credit evaluation model, and manual approval without approval result provided by the credit evaluation model. Based on the first predicted probabilities corresponding to each user, the target group to which the user belongs is determined; Using the second probability prediction model corresponding to the target group, the second predicted probability of the user belonging to the second label is determined, wherein the second label is used to indicate whether the user defaults on the loan or does not default on the loan. The approval result of the user's loan application is determined based on the second predicted probability and the approval threshold corresponding to the second probability prediction model.

2. The method as described in claim 1, characterized in that, The step of determining the target group to which the user belongs based on each first predicted probability corresponding to the user includes: The user is vectorized according to each first predicted probability corresponding to the user to obtain the first vector corresponding to the user; The target cluster is determined from the multiple clusters based on the distances between the first vector and the cluster centers of the multiple clusters.

3. The method as described in claim 1, characterized in that, The second label is used to indicate that there has been no default after the loan was issued. The step of determining the approval result of the user's loan application based on the second predicted probability and the approval threshold corresponding to the second probability prediction model includes: If the second predicted probability is greater than or equal to the approval threshold, the approval result is determined to be approval passed; If the second predicted probability is less than the approval threshold, the approval result is determined to be approval failure.

4. The method as described in claim 1, characterized in that, The second label is used to indicate loan default. The step of determining the approval result of the user's loan application based on the second predicted probability and the approval threshold corresponding to the second probability prediction model includes: If the second predicted probability is less than or equal to the approval threshold, the approval result is determined to be approval passed; If the second predicted probability is greater than the approval threshold, the approval result is determined to be that the approval was not approved.

5. The method according to any one of claims 1-4, characterized in that, Also includes: Based on the credit prediction results of multiple sample users applying for loans using the credit rating model, a first label is determined for each sample user; The first probability prediction model is trained based on the first label of each sample user and the feature data corresponding to each sample user to obtain the first probability prediction model.

6. The method as described in claim 5, characterized in that, Also includes: The first probability prediction model is used to predict each of the sample users to obtain a third predicted probability that each of the sample users belongs to each of the first tags; Each sample user is vectorized according to the third predicted probability of each sample user belonging to each of the first labels to obtain the second vector corresponding to each sample user; Cluster the multiple sample users to obtain multiple groups; The cluster center of each cluster is determined based on the second vector corresponding to the sample users in each cluster.

7. The method as described in claim 6, characterized in that, Also includes: For each subgroup, a third probability prediction model is trained on the initial second probability prediction model based on the feature data of the first target sample user in the subgroup and the second label to which the first target sample user belongs. Here, the first target sample user refers to the sample user whose first label is the first target label. Using the third probability prediction model, a fourth prediction probability is determined for the second target sample user in the cluster to belong to the second label, wherein the second target sample user refers to the sample user whose first label is the second target label; Based on the second target sample user, the fourth predicted probability that the second target sample user belongs to the second label, the other sample users in the cluster except for the second target sample user, and the second label to which the other sample users belong, construct the training sample set corresponding to the cluster; The initial second probability prediction model is trained using the feature data of the training samples in the training sample set and the second label to which the training samples belong, so as to obtain the second probability prediction model corresponding to the cluster.

8. The method as described in claim 7, characterized in that, Also includes: For each cluster, the fifth predicted probability of each sample user in the cluster belonging to the second label is determined using the second probability prediction model corresponding to the cluster. Based on the order of the fifth predicted probability, the sample users in the cluster are divided into equal-frequency groups to obtain multiple bins. For each bin, the cumulative default rate corresponding to the bin is determined based on the ratio of the number of sample users who defaulted on their loans in the third target sample users to the number of the third target sample users. The third target sample users refer to sample users in the cluster whose fifth probability is greater than the lower limit probability of the bin. The cumulative approval rate corresponding to the subgroup is determined by the ratio of the number of sample users whose loan applications have been approved in the third target sample users to the number of sample users in the subgroup. The approval threshold is determined based on the cumulative pass rate and cumulative pass rate corresponding to the sub-boxes.

9. A loan application processing device, characterized in that, include: The first acquisition module is used to predict the creditworthiness of loan applicants based on a credit rating model in order to obtain the credit score corresponding to the user. The first determining module is used to determine the first predicted probability of the user belonging to each first label by using a first probability prediction model when the credit score is within a preset range. The preset range means that the credit evaluation model cannot provide an approval result. The first label is used to indicate the approval method and the corresponding approval result. The first label includes automatic approval using the credit evaluation model, automatic approval using the credit evaluation model, manual approval when the credit evaluation model does not provide an approval result, and manual approval when the credit evaluation model does not provide an approval result. The second determining module is used to determine the target group to which the user belongs based on the first predicted probabilities corresponding to the user. The third determining module is used to determine the second predicted probability of the user belonging to the second label by using the second probability prediction model corresponding to the target group, wherein the second label is used to indicate whether the user defaults after the loan or does not default after the loan. The fourth determining module is used to determine the approval result of the user's loan application based on the second predicted probability and the approval threshold corresponding to the second probability prediction model.

10. The apparatus as claimed in claim 9, characterized in that, The second determining module is used for: The user is vectorized according to each first predicted probability corresponding to the user to obtain the first vector corresponding to the user; The target cluster is determined from the multiple clusters based on the distances between the first vector and the cluster centers of the multiple clusters.

11. The apparatus as claimed in claim 9, characterized in that, The second tag is used to indicate that there has been no default after the loan was issued, and the fourth determining module is used to: If the second predicted probability is greater than or equal to the approval threshold, the approval result is determined to be approval passed; If the second predicted probability is less than the approval threshold, the approval result is determined to be approval failure.

12. The apparatus as claimed in claim 9, characterized in that, The second label is used to indicate loan default, and the fourth determining module is used to: If the second predicted probability is less than or equal to the approval threshold, the approval result is determined to be approval passed; If the second predicted probability is greater than the approval threshold, the approval result is determined to be that the approval was not approved.

13. The apparatus according to any one of claims 9-12, characterized in that, Also includes: The fifth determining module is used to determine the first label to which each of the sample users belongs based on the credit prediction results of the credit rating model for multiple sample users applying for loans; The first training module is used to train an initial first probability prediction model based on the first label to which each sample user belongs and the feature data corresponding to each sample user, so as to obtain the first probability prediction model.

14. The apparatus as claimed in claim 13, characterized in that, Also includes: The second acquisition module is used to predict each of the sample users using the first probability prediction model to obtain a third prediction probability that each of the sample users belongs to each of the first tags. The third acquisition module is used to vectorize each sample user according to the third predicted probability of each sample user belonging to each of the first labels, so as to obtain the second vector corresponding to each sample user. The clustering module is used to cluster multiple sample users to obtain multiple groups; The sixth determination module is used to determine the cluster center of each cluster based on the second vector corresponding to the sample users in each cluster.

15. The apparatus as claimed in claim 14, characterized in that, Also includes: The second training module is used to train the initial second probability prediction model to obtain a third probability prediction model for each subgroup based on the feature data of the first target sample user in the subgroup and the second label to which the first target sample user belongs. The first target sample user refers to the sample user whose first label is the first target label. The seventh determining module is used to determine the fourth predicted probability that the second target sample user in the cluster belongs to the second label using the third probability prediction model, wherein the second target sample user refers to the sample user whose first label is the second target label; The construction module is used to construct the training sample set corresponding to the cluster based on the second target sample user, the fourth predicted probability that the second target sample user belongs to the second label, the other sample users in the cluster except for the second target sample user, and the second label to which the other sample users belong. The third training module is used to train the initial second probability prediction model using the feature data of the training samples in the training sample set and the second label to which the training samples belong, so as to obtain the second probability prediction model corresponding to the cluster.

16. The apparatus as claimed in claim 15, characterized in that, Also includes: The eighth determining module is used to determine, for each subgroup, the fifth predicted probability that each sample user in the subgroup belongs to the second label using the second probability prediction model corresponding to the subgroup; The partitioning module is used to partition the sample users in the cluster into multiple bins based on the order of the fifth predicted probability. The ninth determination module is used to determine the cumulative default rate of each sub-box based on the ratio of the number of sample users who defaulted on loans in the third target sample users to the number of the third target sample users. The third target sample users refer to sample users in the sub-group whose fifth probability is greater than the lower limit probability of the sub-box. The tenth determining module is used to determine the cumulative approval rate corresponding to the subgroup based on the ratio of the number of sample users whose loan applications have been approved among the third target sample users to the number of sample users in the subgroup; The eleventh determining module is used to determine the approval threshold based on the cumulative pass rate and cumulative pass rate corresponding to the sub-box.

17. An electronic device, characterized in that, include: processor; Memory used to store the processor's executable instructions; The processor is configured to execute the instructions to implement the loan application processing method as described in any one of claims 1 to 8.

18. A computer-readable storage medium, wherein instructions in the computer-readable storage medium, when executed by a processor of an electronic device, enable the electronic device to perform the loan application processing method as described in any one of claims 1 to 8.

19. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by a processor, it implements the loan application processing method as described in any one of claims 1-8.