A method for predicting default time and related devices
By executing the default time prediction method of multi-model collaborative work on the server, the problem of insufficient default time prediction in the prior art is solved, and more accurate default time prediction and more efficient credit risk control are achieved.
Patent Information
- Application Number
- CN202210460740.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-04-28
- Publication Date
- 2025-06-20
- Estimated Expiration
- 2042-04-28
AI Technical Summary
The existing credit risk assessment methods are mainly to predict the probability of default of customers, and the lack of prediction of default time has led to low efficiency in risk control in financial institutions.
Through the default time prediction method executed by the server, multiple models (first model, second model and third model) work together to predict the customer's default time. The first model is used to predict the default probability, the second model is used to predict the time interval to which the default time belongs, and the third model is used to predict the specific default time.
This method can not only predict the default probability of customers, but also specifically predict the default time, provide more comprehensive information, help financial institutions formulate more reasonable risk control measures and improve the efficiency of credit risk control.
Smart Images

Figure CN114676936B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of artificial intelligence, and in particular, to a method for predicting default time and related devices. Background Art
[0002] With the rapid development of China's economy and finance, consumer credit has gradually been accepted by people. In recent years, various credit businesses such as auto loans, education loans, small cash loans, and beauty loans have flourished. For credit businesses, the credit risk assessment of customers by financial institutions is crucial. In the big data era, although the data for credit risk assessment is becoming increasingly rich, it also brings many challenges to credit risk assessment.
[0003] Currently, the mainstream credit risk assessment method is to use statistical models to predict whether a customer will default or calculate the default probability of a customer. However, simply predicting the default probability of a customer is not comprehensive enough for the risk control of financial institutions, and the efficiency of risk control is not high. Summary of the Invention
[0004] This application provides a method for predicting default time and related devices. By predicting the default time of a customer, financial institutions can formulate more reasonable risk control measures for the default time, thereby improving the efficiency of credit risk control of financial institutions.
[0005] In a first aspect, this application provides a method for predicting default time. This method can be executed by a server, or by a component (such as a chip, a chip system, etc.) configured in the server, or by a logic module or software capable of implementing all or part of the server functions. This application makes no limitation thereto.
[0006] Among them, the above server is configured with a first model, a second model, and at least one third model. The first model is used to predict the default probability of a customer. The second model is used to predict the time interval to which the default time of the customer belongs. The at least one third model corresponds to at least one time interval, and each third model is used to predict the number of defaults and the default time of the customer within the corresponding time interval.
[0007] Exemplarily, the method includes: obtaining data of a first customer, where the data of the first customer includes parameters for reflecting the credit risk of the first customer; inputting the data into the first model to obtain the default probability of the first customer; in the case where the default probability of the first customer is greater than a preset value, predicting the time interval to which the default time of the first customer belongs through the second model; and predicting the default time of the first customer through the corresponding third model based on the time interval to which the default time of the first customer predicted by the second model belongs.
[0008] Based on the above technical solution, the data of the first customer obtained is input into the first model to obtain the default probability of the customer. When the default probability is greater than the preset value, the time interval to which the default time of the customer belongs is predicted through the second model, that is, how long the customer may default in the future. Further, based on the predicted time interval to which the default time of the customer belongs, the specific default time of the customer is predicted through the corresponding third model. In this way, not only can the default probability of the customer be predicted, but also the default time can be predicted for the customers who may default. Therefore, more comprehensive information can be obtained, which is convenient for financial institutions to formulate more reasonable risk control measures based on this, and is conducive to improving the efficiency of credit risk control of financial institutions.
[0009] Combined with the first aspect, in a possible implementation manner of the first aspect, the first model is a multinomial logistic model, the second model is a Gaussian mixture model (GMM), and the third model is an autoregressive integrated moving average (ARIMA) model.
[0010] Combined with the first aspect, in a possible implementation manner of the first aspect, the data includes one or more of the following: industry category, execution interest rate, loan amount, loan term, gender, age, education level, annual household income, employment situation, unit type, living situation, position, social security flag, and customer level.
[0011] Combined with the first aspect, in a possible implementation manner of the first aspect, the first model includes multiple first sub-models, and the multiple first sub-models correspond to different customer types, and the customer types are determined according to the age group or region to which the customer belongs; the step of inputting the data into the first model to obtain the default probability of the first customer includes: determining the first sub-model corresponding to the customer type of the first customer from the multiple first sub-models; inputting the data of the first customer into the first sub-model to obtain the default probability of the first customer.
[0012] Combined with the first aspect, in a possible implementation manner of the first aspect, the second model includes multiple second sub-models, and the multiple second sub-models correspond to different customer types, and the customer types are determined according to the age group or region to which the customer belongs; the step of predicting the time interval to which the default time of the first customer belongs through the second model includes: determining the second sub-model corresponding to the customer type of the first customer from the multiple second sub-models; predicting the time interval to which the default time of the first customer belongs through the second sub-model.
[0013] In combination with the first aspect, in a possible implementation manner of the first aspect, each third model includes a plurality of third sub-models, and any two third sub-models among the plurality of third sub-models correspond to different customer types, where the customer type is determined according to the age group or region to which the customer belongs; predicting the default time of the first customer through the corresponding third model based on the time interval to which the default time of the first customer predicted by the second model belongs includes: determining, in the third model corresponding to the time interval to which the default time of the first customer predicted by the second model belongs, a third sub-model corresponding to the customer type of the first customer; and predicting the default time of the first customer through the third sub-model.
[0014] In combination with the first aspect, in a possible implementation manner of the first aspect, the method further includes: obtaining a training set, where the training set includes historical data of a plurality of customers; and training the first model, the second model, and the at least one third model respectively based on the training set.
[0015] In combination with the first aspect, in a possible implementation manner of the first aspect, the method further includes: grouping the training set based on the customer types respectively corresponding to the plurality of customers to obtain multiple groups of training sets, where the customer types corresponding to the multiple groups of training sets are different, and the customer type is determined according to the age group or region to which the customer belongs; and training the first model, the second model, and the at least one third model respectively based on the training set includes: training the first model, the second model, and the at least one third model respectively based on each group of training sets to obtain a trained first sub-model, a trained second sub-model, and a plurality of trained third sub-models.
[0016] In combination with the first aspect, in a possible implementation manner of the first aspect, the method further includes: updating the training set at a preset period to obtain an updated training set; and training the first model, the second model, and the at least one third model respectively based on the updated training set.
[0017] In a second aspect, the present application provides a model training method, which can be executed by a server. The server is configured with a first model, a second model, and at least one third model. The first model is used to predict the default probability of a customer, the second model is used to predict the time interval to which the default time of the customer belongs, and the at least one third model corresponds to at least one time interval, and each third model is used to predict the number of defaults and the default time of the customer within the corresponding time interval.
[0018] Exemplarily, the method includes: obtaining data of multiple customers, where the data of each customer includes parameters for reflecting the credit risk of the customer; inputting the data of the multiple customers into the first model to obtain the default probabilities of the multiple customers; in the case where the default probability of a customer is greater than a preset value, predicting the time interval to which the default time of the customer belongs through the second model; and predicting the default time of the customer through the corresponding third model based on the time interval to which the default time of the customer predicted by the second model belongs.
[0019] Based on the above technical solution, the data of multiple customers obtained is input into the first model to obtain the default probabilities of the multiple customers. For customers with a default probability greater than the preset value, the time interval to which the default time of the customer belongs is predicted through the second model, that is, how long in the future the customer may default. Further, based on the predicted time interval to which the default time of the customer belongs, the specific default time of the customer is predicted through the corresponding third model. In this way, the first model, the second model, and at least one third model can be trained with the data of multiple customers, which is beneficial to improving the accuracy of the models.
[0020] In a third aspect, the present application provides a server, which is configured with a first model, a second model, and at least one third model. The first model is used to predict the default probability of a customer, the second model is used to predict the time interval to which the default time of the customer belongs, the at least one third model corresponds to at least one time interval, and each third model is used to predict the number of defaults and the default time of the customer within the corresponding time interval.
[0021] Exemplarily, the server includes an acquisition unit, an input unit, and a processing unit. Among them, the acquisition unit is used to acquire the data of a first customer, and the data of the first customer includes parameters for reflecting the credit risk of the first customer; the input unit is used to input the data into the first model to obtain the default probability of the first customer; the processing unit is used to, in the case where the default probability of the first customer is greater than the preset value, predict the time interval to which the default time of the first customer belongs through the second model; the processing unit is further used to predict the default time of the first customer through the corresponding third model based on the time interval to which the default time of the first customer predicted by the second model belongs.
[0022] In a fourth aspect, the present application provides a server, and the device includes a processor. The processor is coupled to the memory and can be used to execute the computer program in the memory to implement the method described in any possible implementation manner of the first aspect to the second aspect and the first aspect to the second aspect.
[0023] Optionally, the server in the fourth aspect further includes a memory.
[0024] Optionally, the server in the fourth aspect further includes a communication interface, and the processor is coupled to the communication interface.
[0025] In a fifth aspect, the present application provides a chip system, which includes at least one processor for supporting the implementation of the functions involved in any one of the possible implementation manners of the above first aspect to the second aspect, for example, receiving or processing the data involved in the above method, etc.
[0026] In a possible design, the chip system further includes a memory for storing program instructions and data, and the memory is located inside or outside the processor.
[0027] The chip system may be composed of chips or may include chips and other discrete devices.
[0028] In a sixth aspect, the present application provides a computer-readable storage medium, on which a computer program (which may also be referred to as code or instruction) is stored. When the computer program is run by a processor, the method described in any one of the above first aspect to the second aspect and any one of the possible implementation manners of the first aspect to the second aspect is executed.
[0029] In a seventh aspect, the present application provides a computer program product, which includes: a computer program (which may also be referred to as code or instruction). When the computer program is run, the method described in any one of the above first aspect to the second aspect and any one of the possible implementation manners of the first aspect to the second aspect is executed.
[0030] It should be understood that the third aspect to the seventh aspect of the present application correspond to the technical solutions of the first aspect and the second aspect of the present application, and the beneficial effects obtained by each aspect and the corresponding feasible implementation manners are similar and will not be elaborated here.
[0031] It should also be understood that a method for predicting default time and related devices provided by the present application can be applied to the field of artificial intelligence or other fields. The present application does not make any limitations in this regard. BRIEF DESCRIPTION OF THE DRAWINGS
[0032] Figure 1 is a schematic diagram of an application scenario applicable to the method provided in the embodiment of the present application;
[0033] Figure 2 is a schematic flowchart of the method for predicting default time provided in the embodiment of the present application;
[0034] Figure 3 is another schematic flowchart of the method for predicting default time provided in the embodiment of the present application;
[0035] Figure 4It is a schematic block diagram of a server provided by an embodiment of the present application;
[0036] Figure 5 It is another schematic block diagram of a server provided by an embodiment of the present application. Detailed implementation manners
[0037] The following detailed description of the embodiments of the present application provided in the accompanying drawings is not intended to limit the scope of the present application claimed, but merely represents selected embodiments of the present application. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in the present application without creative efforts shall fall within the scope of protection of the present application.
[0038] For ease of understanding the method for predicting default time provided by the embodiments of the present application, the application scenarios applicable to the embodiments of the present application will be described below. It can be understood that the application scenarios described in the embodiments of the present application are for more clearly illustrating the technical solutions of the embodiments of the present application, and do not constitute a limitation on the technical solutions provided by the embodiments of the present application.
[0039] Figure 1 It is a schematic diagram of an application scenario applicable to the method provided by the embodiments of the present application. As Figure 1 shown, technicians can input relevant data for the credit risk assessment of customers through the electronic device 110. Among them, the electronic device 110 can be communicatively connected to the server 120, and the server 120 can present a user interface through the electronic device 110. The user interface provides an interface for interaction between the technician and the server 120. The technician can send data or information to the server 120 by means of operations such as inputting or selecting on the user interface. Correspondingly, the server 120 can present the credit risk assessment result of the customer through the electronic device 110 based on the data or information input by the technician.
[0040] It should be understood that Figure 1 the scenario shown is only an example. The server 120 can be a physical device or a server cluster composed of multiple physical devices.
[0041] With the rapid development of China's economy and finance, consumer credit has gradually been accepted by people. In recent years, various credit businesses such as auto loans, education loans, small cash loans, and beauty loans have flourished. For credit businesses, the credit risk assessment of customers by financial institutions is crucial. In the big data era, although the data for credit risk assessment is becoming increasingly rich, it also brings many challenges to credit risk assessment.
[0042] The current mainstream credit risk assessment method is to use statistical models to predict whether a customer will default or calculate the customer's default probability. However, simply predicting the customer's default probability is not comprehensive enough for the risk control of financial institutions, and the efficiency of risk control is not high.
[0043] To solve the above problems, the present application provides a method for predicting the default time. The data of the first customer obtained is input into the first model to obtain the default probability of the customer. When the default probability is greater than a preset value, the time interval to which the default time of the customer belongs is predicted through the second model, that is, how long the customer may default in the future. Each time interval corresponds to a third model for predicting the customer's default time. Further, based on the predicted time interval to which the customer's default time belongs, the specific default time of the customer is predicted through the corresponding third model, so that financial institutions can formulate more reasonable risk control measures based on the default time.
[0044] The following uses specific embodiments to elaborate in detail on the technical solution of the present application and how the technical solution of the present application solves the above technical problems. These several specific embodiments can be combined with each other, and the same or similar concepts or processes may not be repeated in some embodiments. The embodiments of the present application will be described below with reference to the accompanying drawings.
[0045] Figure 2 It is a schematic flowchart of the method 200 for predicting the default time provided by the embodiment of the present application. Figure 2 The shown method 200 for predicting the default time may include step 210 and step 240. Each step in the method 200 will be described in detail below.
[0046] It should be understood that Figure 2 The shown method 200 takes the server as the execution subject, but it should not impose any limitation on the execution subject of the method. As long as a program that records the method provided by the present application can be run, the method provided by the embodiment of the present application can be executed. For example, the server can also be replaced by components configured in the server (such as chips, chip systems, etc.), or other functional modules that can call the program and execute the program.
[0047] It should also be understood that the above server is configured with a first model, a second model, and at least one third model. Among them, the first model is used to predict the default probability of the customer, the second model is used to predict the time interval to which the default time of the customer belongs, that is, to predict how long the customer may default, and at least one third model corresponds to at least one time interval, and each third model is used to predict the number of defaults and the default time of the customer within the corresponding time interval.
[0048] Step 210: Obtain the data of the first customer, where the data of the first customer includes parameters for reflecting the credit risk of the first customer.
[0049] Among them, the first customer can be an incremental customer, that is, a new customer, or a stock customer, that is, an old customer. This application embodiment does not make a limitation on this. Among them, the historical data of the old customer can be used to train the first model, the second model, and at least one third model.
[0050] A possible implementation is that the server obtains the data of the first customer in response to the input operation of the technician. That is, the technician can input the data reflecting the credit risk of the first customer through the user interface. Correspondingly, the server obtains the data of the first customer for predicting the default time of the first customer.
[0051] Another possible implementation is that the server can obtain the data of the first customer from its own platform or the partner platform. Among them, the data of the first customer is stored in the own platform or the partner platform, and the own platform or the partner platform can communicate with the server. The server can trigger the server to obtain the data of the first customer from the above platform in response to the click operation of the technician, etc. That is, the technician can click to predict the default time of the first customer through the user interface, and then trigger the server to obtain the data of the first customer from the above platform.
[0052] Optionally, the above data includes one or more of the following: industry category, execution interest rate, loan amount, loan term, gender, age, education level, annual household income, employment status, unit type, living situation, position, social security flag, and customer level.
[0053] Among them, the customer level can be used to reflect the importance of the customer. For example, the higher the customer level, the higher the importance of the customer.
[0054] In one example, the server can obtain one or more of the industry category, execution interest rate, loan amount, loan term, gender, age, education level, annual household income, employment status, unit type, living situation, position, social security flag, and customer level of the first customer in response to the input operation of the technician, so that the server can predict the default probability of the first customer based on various influences.
[0055] In another example, the server may obtain one or more of the industry category, execution interest rate, loan amount, loan term, gender, age, education level, annual household income, employment status, unit type, living situation, position, social security flag, and customer level of the first customer from its own platform or a partner platform. Among them, the above data of the first customer is stored in the own platform or the partner platform, and the own platform or the partner platform can communicate with the server. The server can trigger the server to obtain the above data of the first customer from the above platform in response to operations such as clicks by technical personnel.
[0056] It should be understood that in this application, the collection, storage, use, processing, transmission, provision, and disclosure of financial data, user personal data, and other information comply with the provisions of relevant laws and regulations and do not violate public order and good customs.
[0057] By providing the above-mentioned various data that may affect the customer's default situation in the embodiments of this application, the influence of various factors on the customer's default situation is comprehensively considered, which is beneficial to improving the accuracy of predicting the customer's default probability.
[0058] Optionally, the above first model is a Logistic model, the second model is a GMM, and the third model is an ARIMA model.
[0059] The Logistic model, GMM, and ARIMA model will be introduced in detail below.
[0060] I. Logistic model
[0061] The form of the Logistic model is similar to the following: Among them, represents the default probability of customer i, represents the data of customer i, is the parameter to be trained.
[0062] It can be understood that the above indicating that customer i defaults should not impose any limitations on the embodiments of this application. For example, can also represent that customer i does not default. Correspondingly, when the probability that the customer does not default is greater than the preset value, it means that the customer has good credit and it can be considered that the customer will abide by the contract during the loan period. In other words, the server does not need to further predict the default time of the customer; when the probability that the customer does not default is less than or equal to the preset value, it means that the customer has average credit and it can be considered that the customer will default during the loan period (it can be called a high-risk customer). In other words, the server needs to further predict the default time of the customer.
[0063] II. GMM
[0064] The form of the 1-dimensional Gaussian distribution is similar to the following: Among them, N(x|μ,σ) represents probability, σ represents the standard deviation, μ represents the mean, that is, the expectation, and x represents the customer's data, such as household income, or age, etc. The above formula represents the probability near μ. It can be understood that the closer to μ, that is, the smaller σ, the greater the probability.
[0065] The form of the d (d is a positive integer greater than 1)-dimensional Gaussian distribution is similar to the following:
[0066] Among them, d represents the dimension of x, ∑ represents the d*d covariance matrix, |∑| is the value of the determinant of the covariance, μ represents the mean, that is, the expectation, and x represents the customer's data, such as household income, age, etc.
[0067] The Gaussian mixture model is to mix multiple Gaussian models together and use weight parameters to adjust the mixing ratio of different Gaussian models (representing categories in the data samples). The form of the Gaussian mixture model is similar to the following: Among them, p(x|C j )=N(x|μ j ,∑ j ) is the conditional probability density of category or group j (subject to Gaussian distribution), p(C j )≥0, p(C j ) is the weight parameter of category j (w j =p(C j )) and The parameters of the model are p(C j ), μ j ,∑ j , where j = 1,…,k, k is the total number of categories or groups (there are a total of k categories or groups), and k is a positive integer. μ j is the mean of category j, and ∑ j is the covariance of category j.
[0068] In the embodiments of the present application, the category represents the time interval of the customer's default time, that is, how long in the future the customer will default. For example, multiple categories are: default within one year, default within two years, default within three years, and default within four years. Based on the above Gaussian mixture model, the posterior probability of the first customer can be calculated and classified into one of the Gaussian models, that is, the category to which the first customer belongs is obtained, that is, the time interval to which the default time of the first customer belongs.
[0069] It can be understood that before using the above Gaussian mixture model, parameter estimation is required to obtain a Gaussian mixture model for predicting the category to which the first customer belongs. For example, the expectation maximization (EM) algorithm can be used to estimate the parameters of the Gaussian mixture model to obtain a Gaussian mixture model for predicting the category to which the first customer belongs. The process of parameter estimation is described in detail below.
[0070] The goal of fitting the GMM is to find p(C j ), μ j , Σ j such that the following is maximized where p(x|C j ) = N(x|μ j , Σ j ). Taking the logarithm on both sides, the log-likelihood function of the GMM is obtained: The goal is to maximize this log-likelihood function, so the EM algorithm is used. The EM algorithm includes an initialization step and an iterative step. The initialization step is: Initialize K clusters: C1, …, C k , and for each cluster j, there are parameters (μ j , ∑ j ) and p(C j ). The iterative step is: Estimate the cluster to which each data point belongs p(C j |x j ) (expectation step), and calculate the expectation of the likelihood function; re-estimate the parameters (μ j , ∑ j ) and p(C j ) for each cluster j (maximization step).
[0071] Specific process of the EM algorithm:
[0072] Step 1: Let z1, …, z n represent the true sources (i.e., categories) corresponding to the data x1, …, x n . Each z i is a discrete variable taking values between j = 1, …, k, where k is the number of categories. There is the following log-likelihood function:
[0073] Step 2: Use θ to represent the model parameters, and replace logp(X, z|θ) with its expectation .
[0074] Step 3: Given the distribution of logp(X, z|θ (t) ) for the current parameter estimation, according to Bayes' rule, it satisfies
[0075] Step 4: p(z i = c|x i , θ (t) ) is called the "responsibility" borne by cluster c for data point i, and r ic = p(z i = c|x i , θ (t) ).
[0076] Step 5: The expectation step of GMM. Denote x1, …, x n as X,
[0077]
[0078] Step 6: The maximization step of GMM. Take the partial derivative of Q(θ, θ (t) ) with respect to each parameter and set these partial derivatives to zero to obtain the new parameter estimates where,
[0079] It can be understood that a more specific description of GMM and the EM algorithm can be found in known technologies and will not be elaborated here.
[0080] III. ARIMA Model
[0081] For the autoregressive model, it is first necessary to determine an order p, which represents how many periods of historical values are used to predict the current value. The formula for the p-order autoregressive model is defined as: where y t is the current value, μ is the constant term, p is the order, γ i is the autocorrelation coefficient, and e t is the error. The moving average model focuses on the accumulation of the error terms in the autoregressive model, and its formula is defined as follows: Combining the autoregressive model and the moving average model gives the autoregressive moving average model ARMA(p, q), and its calculation formula is as follows: Combining the autoregressive model, the moving average model, and the differencing method gives ARIMA(p, d, q), where d is the order of differencing required for the data.
[0082] It should be understood that the above introduction of the Logistic model, GMM, and ARIMA model is only for a clearer understanding of the method provided by the embodiments of the present application and should not constitute any limitation to the embodiments of the present application. A more detailed introduction can be found in known technologies and will not be elaborated here.
[0083] In the embodiments of the present application, the first model uses the Logistic model to predict the default probability of customers, which is beneficial to initially screen out some high-risk customers, that is, customers with a relatively high default probability, and further beneficial to the financial institution to focus on evaluating such customers and minimize the losses of the financial institution as much as possible. The second model uses GMM, which can be used to identify relatively complex distributions. By using GMM, the default situation of customers can be comprehensively considered in terms of multiple aspects such as industry, income, employment situation, etc., and further beneficial to improving the accuracy of the time interval to which the predicted default time of the customer belongs. The third model uses the ARIMA model to predict the default time of customers, which can improve the prediction accuracy.
[0084] Step 220: Input the above data into the first model to obtain the default probability of the first customer.
[0085] After the server obtains the data of the first customer, it inputs the data of the first customer into the first model to obtain the default probability of the first customer.
[0086] Exemplarily, the server inputs the industry category, execution interest rate, loan amount, loan term, gender, age, education level, annual household income, employment situation, unit type, living situation, position, social security mark, and customer level of the first customer into the Logistic model to obtain the default probability of the first customer.
[0087] It should be understood that the above description takes obtaining the default probability of the first customer as an example, but should not constitute any limitation to the embodiments of the present application. For example, the server can also calculate the probability that the first customer does not default based on the Logistic model. Correspondingly, when the probability that the customer does not default is greater than the preset value, it indicates that the customer has good credit, and it can be considered that the customer will abide by the contract during the loan period. In other words, the server does not need to further predict the default time of the customer; when the probability that the customer does not default is less than or equal to the preset value, it indicates that the customer has average credit, and it can be considered that the customer will default during the loan period (it can be called a high-risk customer). In other words, the server needs to further predict the default time of the customer.
[0088] Step 230: When the default probability of the first customer is greater than the preset value, predict the time interval to which the default time of the first customer belongs through the second model.
[0089] After the server obtains the default probability of the first customer, when the default probability of the first customer is less than or equal to the preset value, it can be considered that the first customer has good credit, that is, the first customer will abide by the contract during the loan period. In other words, the server does not need to further predict the default time of the customer.
[0090] When the default probability of the first customer is greater than the preset value, the server predicts the time interval to which the default time of the first customer belongs through the second model, that is, how long the first customer may default in the future.
[0091] Exemplarily, when the default probability of the first customer is greater than the preset value, the server predicts the time interval to which the default time of the first customer belongs through GMM. For example, the server predicts through GMM that the first customer may default within one year after the loan.
[0092] Step 240, based on the time interval to which the default time of the first customer predicted by the second model belongs, predict the default time of the first customer through the corresponding third model.
[0093] Each time interval corresponds to a third model. For example, if a customer may default within one year, it corresponds to a third model, which is used to predict the specific time when the customer defaults within one year, such as defaulting in the tenth month within one year.
[0094] After the server predicts the time interval to which the default time of the first customer belongs, it predicts the default time of the first customer through the corresponding third model.
[0095] Exemplarily, the server predicts through GMM that the first customer may default within one year after the loan. Further, through its corresponding ARIMA model, it predicts that the first customer may default in the tenth month after the loan, which is convenient for financial institutions to formulate more reasonable risk control measures.
[0096] It should be understood that the economic levels of different types of customers are different, and the probability of their default may be relatively high. Therefore, the server can select the corresponding first sub-model, second sub-model, and third sub-model according to the customer type to which the first customer belongs. How the server predicts the default time of the customer after classifying the customer according to the customer type will be described in detail below.
[0097] Optionally, the first model includes multiple first sub-models, and the multiple first sub-models correspond to different customer types, and the customer type is determined according to the age group or region to which the customer belongs; inputting the data into the first model to obtain the default probability of the first customer includes: determining the first sub-model corresponding to the customer type of the first customer from the multiple first sub-models; inputting the data of the first customer into the first sub-model to obtain the default probability of the first customer.
[0098] Exemplarily, the first model includes a first sub-model 1 and a first sub-model 2. The customer type corresponding to the first sub-model 1 is southern customers, and the customer type corresponding to the first sub-model 2 is northern customers. The server determines which first sub-model to input the data of a first customer based on the type to which the first customer belongs. For example, if the first customer belongs to southern customers, the server inputs the data of the first customer into the first sub-model 1 to obtain the default probability of the first customer.
[0099] Optionally, the second model includes a plurality of second sub-models. The plurality of second sub-models correspond to different customer types, and the customer types are determined according to the age group or region to which the customer belongs. Predicting the time interval to which the default time of the first customer belongs through the second model includes: determining the second sub-model corresponding to the customer type of the first customer from the plurality of second sub-models; predicting the time interval to which the default time of the first customer belongs through the second sub-model.
[0100] Exemplarily, the second model includes a second sub-model 1 and a second sub-model 2. The customer type corresponding to the second sub-model 1 is southern customers, and the customer type corresponding to the second sub-model 2 is northern customers. The server determines which second sub-model to input the data of a first customer based on the type to which the first customer belongs. For example, if the first customer belongs to southern customers, then if the default probability of the first customer predicted based on the first sub-model 1 is greater than a preset value, the server inputs the data of the first customer into the second sub-model 1 to predict the time interval to which the default time of the first customer belongs.
[0101] Optionally, each third model includes a plurality of third sub-models. Any two of the plurality of third sub-models correspond to different customer types, and the customer types are determined according to the age group or region to which the customer belongs. Predicting the default time of the first customer through the corresponding third model based on the time interval to which the default time of the first customer predicted by the second model belongs includes: determining the third sub-model corresponding to the customer type of the first customer in the third model corresponding to the time interval to which the default time of the first customer predicted by the second model belongs; predicting the default time of the first customer through the third sub-model.
[0102] Exemplarily, the first customer belongs to southern customers. The server predicts the time interval to which the default time of the first customer belongs based on the second sub-model 1, and different time intervals correspond to different third sub-models. For example, within one year after the loan corresponds to the third sub-model 1, within two years after the loan corresponds to the third sub-model 2, and within three years after the loan corresponds to the third sub-model 3. Assuming that the time interval to which the default time of the first customer belongs is within one year after the loan, the server predicts the default time of the first customer through the third sub-model 1, such as the first customer may default in the tenth month after the loan.
[0103] Optionally, before obtaining the data of the first customer Figure 2The method shown also includes: obtaining a training set, where the training set includes historical data of multiple customers; and based on the training set, training the first model, the second model, and at least one third model respectively.
[0104] Exemplarily, the server may obtain the historical data of multiple customers, divide the above-mentioned historical data of multiple customers into training data and validation data, that is, the data of some customers is used to train the model, and the data of some customers is used to validate the model, so as to obtain the trained first model, second model, and at least one third model.
[0105] In the embodiment of the present application, by training the first model, the second model, and at least one third model respectively through the training set, it is beneficial to improve the accuracy of the first model, the second model, and at least one third model, and further beneficial to improve the accuracy of the predicted default time of the customers.
[0106] Optionally, Figure 2 The method shown also includes: grouping the training set based on the customer types respectively corresponding to multiple customers to obtain multiple groups of training sets, where the customer types corresponding to the multiple groups of training sets are different, and the customer type is determined according to the age group or region to which the customer belongs; and, based on the training set, training the first model, the second model, and at least one third model respectively, including: training the first model, the second model, and at least one third model respectively based on each group of training sets to obtain a trained first sub-model, a second sub-model, and multiple trained third sub-models.
[0107] The server may group multiple customers according to the region where the customers are located or the age group to which the customers belong, that is, divide the training set into multiple groups of training sets, each group of training sets corresponds to a type of customer, and train the first model, the second model, and at least one third model based on each group of training sets to obtain a trained first sub-model, a second sub-model, and multiple trained third sub-models.
[0108] Exemplarily, the server obtains data of 100 customers, among which 40 customers are from the south and 60 customers are from the north. Due to the economic level differences between the south and the north, the 100 customers are divided into two groups. For example, using the data of 40 customers from the south, the first model, the second model, and at least one third model are respectively trained to obtain a trained first sub-model, a trained second sub-model, and multiple trained third sub-models; and using the data of 60 customers from the north, the first model, the second model, and at least one third model are respectively trained to obtain a trained first sub-model, a trained second sub-model, and multiple trained third sub-models. In this way, 2 first sub-models, 2 second sub-models, and several third sub-models (the number of third sub-models is twice the number of third models) can be obtained. For a new customer, the server can make predictions using the corresponding first sub-model, second sub-model, and third sub-model based on the type to which the customer belongs.
[0109] In the embodiments of the present application, classifying customers according to their types is beneficial to reducing the impact of economic differences among different types of customers on predicting the default probability of customers.
[0110] Optionally, Figure 2 The method shown further includes: updating the training set according to a preset period to obtain an updated training set; and training the first model, the second model, and at least one third model respectively based on the updated training set.
[0111] In other words, the server can update the training set periodically and train the first model, the second model, and at least one third model respectively based on the updated training set. For example, the server can periodically obtain historical data of different customers and train the first model, the second model, and at least one third model based on the obtained data.
[0112] In the embodiments of the present application, by periodically updating the training set and training the model based on the updated training set, that is, training the model multiple times, it is beneficial to improve the accuracy of the model, and further improve the accuracy of predicting the default time of customers.
[0113] Figure 3 It is another flowchart of the method for predicting the default time provided by the embodiments of the present application.
[0114] As Figure 3 shown, in step 310, the server starts.
[0115] In step 320, the server maintains or updates the model and configures the model parameters. For example, the server maintains or updates the first model, the second model, and at least one third model and configures the parameters of the above models.
[0116] Step 330, the server enables the model for customer risk monitoring. For example, the server enables the above-mentioned first model.
[0117] Step 340, the server classifies and clusters the customers. Exemplarily, the server determines whether a customer is a high-risk customer or a compliant customer based on the first model. For example, the server predicts the default probability of the customer. If the default probability of the customer is greater than a preset value, it is considered that the customer is a high-risk customer, and it is necessary to further predict the time interval to which the default time belongs and the specific default time. For the detailed process, reference can be made to Figure 2 the relevant description, which will not be elaborated here.
[0118] Step 350, the server displays the customer classification and clustering results. In this way, it can prompt the technicians which customers are high-risk for the business personnel to focus on for review. For example, for high-risk customers, their default times can be further predicted.
[0119] Step 360, the server calculates based on the historical records and displays the model monitoring effect. In other words, the server inputs the historical data of the customer into the model to determine the default time, and judges whether the default time of the customer is accurate. For example, if the actual default time of the customer is the same as the default time calculated based on the model, it indicates that the model monitoring effect is good.
[0120] It can be understood that based on the model monitoring effect, the developer can instruct the server to adjust the model to improve the accuracy of the model.
[0121] Based on the above technical solution, the server inputs the data of the first customer obtained into the first model to obtain the default probability of the customer. When the default probability is greater than the preset value, the time interval to which the default time of the customer belongs is predicted through the second model, that is, how long the customer may default in the future. Further, based on the predicted time interval to which the default time of the customer belongs, the specific default time of the customer is predicted through the corresponding third model. In this way, not only can the default probability of the customer be predicted, but also the default time can be further predicted for the customers who may default. Therefore, more comprehensive information can be obtained, which is convenient for financial institutions to formulate more reasonable risk control measures based on this, and is beneficial to improving the efficiency of credit risk control of financial institutions.
[0122] Optionally, an embodiment of the present application further provides a model training method, which can be executed by the server. The server is configured with a first model, a second model and at least one third model. The first model is used to predict the default probability of the customer, the second model is used to predict the time interval to which the default time of the customer belongs, and at least one third model corresponds to at least one time interval, and each third model is used to predict the number of default times and the default time of the customer within the corresponding time interval.
[0123] Exemplarily, the method includes: obtaining data of a plurality of customers, where the data of each customer includes parameters for reflecting the credit risk of the customer; inputting the data of the plurality of customers into a first model to obtain the default probabilities of the plurality of customers; in the case where the default probability of a customer is greater than a preset value, predicting the time interval to which the default time of the customer belongs through a second model; and predicting the default time of the customer through a corresponding third model based on the time interval to which the default time of the customer predicted by the second model belongs.
[0124] Wherein, for the specific processes of training the first model, the second model, and at least one third model, reference may be made to Figure 2 the relevant descriptions of the embodiments shown, which will not be elaborated herein.
[0125] Based on the above technical solution, the data of the obtained plurality of customers is input into the first model to obtain the default probabilities of the plurality of customers. For customers with a default probability greater than the preset value, the time interval to which the default time of the customer belongs is predicted through the second model, that is, how long the customer may default in the future. Further, based on the time interval to which the default time of the customer predicted by the second model belongs, the specific default time of the customer is predicted through the corresponding third model. In this way, the first model, the second model, and at least one third model can be trained with the data of the plurality of customers, which is beneficial to improving the accuracy of the model.
[0126] Figure 4 It is a schematic block diagram of a server provided by an embodiment of the present application.
[0127] As Figure 4 shown, the device 400 may include: an obtaining unit 410, an input unit 420, and a processing unit 430. The server 400 can be used to implement Figure 2 or Figure 3 the method described in the embodiments shown.
[0128] Exemplarily, when the device 400 is used to implement Figure 2 the method described in the embodiments shown, the obtaining unit 410 is used to obtain the data of a first customer, where the data of the first customer includes parameters for reflecting the credit risk of the first customer; the input unit 420 is used to input the data into the first model to obtain the default probability of the first customer; the processing unit 430 is used to predict the time interval to which the default time of the first customer belongs through the second model in the case where the default probability of the first customer is greater than the preset value; the processing unit 430 is further used to predict the default time of the first customer through the corresponding third model based on the time interval to which the default time of the first customer predicted by the second model belongs.
[0129] Optionally, the first model is a Logistic model, the second model is a GMM, and the third model is an ARIMA model.
[0130] Optionally, the above data includes one or more of the following: industry category, execution interest rate, loan amount, loan term, gender, age, education level, annual household income, employment status, unit type, living situation, job title, social security flag, and customer level.
[0131] Optionally, the first model includes multiple first sub-models, and the multiple first sub-models correspond to different customer types, where the customer types are determined according to the age group or region to which the customer belongs; and, the input unit 420 is specifically configured to determine, from the multiple first sub-models, the first sub-model corresponding to the customer type of the first customer; input the data of the first customer into the first sub-model to obtain the default probability of the first customer.
[0132] Optionally, the second model includes multiple second sub-models, and the multiple second sub-models correspond to different customer types, where the customer types are determined according to the age group or region to which the customer belongs; the processing unit 430 is specifically configured to determine, from the multiple second sub-models, the second sub-model corresponding to the customer type of the first customer; predict the time interval to which the default time of the first customer belongs through the second sub-model.
[0133] Optionally, each third model includes multiple third sub-models, and any two of the multiple third sub-models correspond to different customer types, where the customer types are determined according to the age group or region to which the customer belongs; the processing unit 430 is specifically configured to determine, in the third model corresponding to the time interval to which the default time of the first customer predicted by the second model belongs, the third sub-model corresponding to the customer type of the first customer; predict the default time of the first customer through the third sub-model.
[0134] Optionally, the processing unit 430 is further configured to obtain a training set, where the training set includes historical data of multiple customers; based on the training set, train the first model, the second model, and at least one third model respectively.
[0135] Optionally, the processing unit 430 is further configured to group the training set based on the customer types corresponding to multiple customers respectively to obtain multiple groups of training sets, where the customer types corresponding to the multiple groups of training sets are different, and the customer types are determined according to the age group or region to which the customer belongs; and, the processing unit 430 is specifically configured to train the first model, the second model, and at least one third model respectively based on each group of training sets to obtain a trained first sub-model, a second sub-model, and multiple trained third sub-models.
[0136] Optionally, the processing unit 430 is further configured to update the training set at a preset period to obtain an updated training set, and based on the updated training set, train the first model, the second model, and at least one third model respectively.
[0137] It should be understood that the division of units in the embodiments of the present application is illustrative, merely a logical function division, and there may be other division methods in actual implementation. In addition, each functional unit in the various embodiments of the present application may be integrated in a processor, may exist separately physically, or two or more units may be integrated in one unit. The above integrated units may be implemented in the form of hardware or in the form of software functional modules.
[0138] Exemplarily, the server may include a data management unit, a model management unit, and a risk monitoring unit. Among them, the data management unit is used to store or update loan customer data, and the loan customer data is used for machine learning model training and verification. The model management unit is used to configure and update machine learning models for data analysis and risk analysis, such as Logistic models, GMM, and ARIMA models. The risk monitoring unit is used to apply machine learning models to perform clustering analysis and risk prediction on loan customer data, etc., and display credit risks, that is, the default time and / or default probability of customers.
[0139] The data management unit may also be specifically divided into a loan customer data module, a customer group data module, and a prediction record data module.
[0140] Among them, the loan customer data module is used to store data of loan customers. The data includes industry category, execution interest rate, loan amount, loan term, gender, age, education level, family annual income, employment situation, unit type, residence situation, position, social security mark, customer level, etc. The customer group data module is used to generate customer group data (that is, a set of data of customers belonging to the same category) through clustering analysis. The prediction record data module is used to store customer risk prediction results and record data.
[0141] The model management unit may also be specifically divided into a model instance module, a model parameter module, and a model update module.
[0142] Among them, the model instance module is used to configure machine learning models for classification and clustering analysis, including Logistic models, GMM, etc. The model parameter module is used to configure and update the parameters of machine learning models. For the parameters of GMM, the Metropolis-Hastings algorithm is used to update the model parameters to convergence. The model update module is used to update existing models, or delete old models, or add new models.
[0143] The risk monitoring unit can also be specifically divided into an enabling model module, a customer clustering result module, and a model monitoring effect module.
[0144] Among them, the enabling model module is used to enable or disable the machine learning model to process risk monitoring. The customer clustering result module is used to classify and cluster loan customers using the model, display the results, and if a customer is determined to be a high-risk customer, prompt the default duration prediction result. For example, customer A will default within one year with a probability of 83%. The model monitoring effect module is used to comprehensively analyze the historical data of loan customers and display the model risk monitoring effects, such as indicators like the correct classification ratio, error ratio, and recall rate, for model developers to analyze and optimize the model.
[0145] Figure 5 It is another schematic block diagram of the server provided by the embodiments of the present application.
[0146] The server 500 can be used to implement the Figure 2 or Figure 3 methods described in the embodiments shown. The server 500 can be a chip system. In the embodiments of the present application, the chip system can be composed of chips or can include chips and other discrete devices.
[0147] As Figure 5 shown, the server 500 can include at least one processor 510.
[0148] Exemplarily, the processor 510 can be used to obtain data of a first customer, where the data of the first customer includes parameters for reflecting the credit risk of the first customer; input the data into a first model to obtain the default probability of the first customer; in the case where the default probability of the first customer is greater than a preset value, predict the time interval to which the default time of the first customer belongs through a second model; and predict the default time of the first customer through a corresponding third model based on the time interval to which the default time of the first customer predicted by the second model belongs. For specific details, refer to the detailed description in the method examples, which will not be elaborated here.
[0149] The server 500 can also include at least one memory 520, which can be used to store program instructions and / or data. The memory 520 is coupled to the processor 510. The coupling in the embodiments of the present application is an indirect coupling or communication connection between devices, units, or modules, which can be electrical, mechanical, or other forms for information interaction between devices, units, or modules. The processor 510 may cooperate with the memory 520. The processor 510 may execute the program instructions stored in the memory 520. At least one of the at least one memory may be included in the processor.
[0150] The server 500 may further include a communication interface 530 for communicating with other devices via a transmission medium, so that the server 500 can communicate with other devices. The communication interface 530 may be, for example, a transceiver, an interface, a bus, a circuit, or a device capable of implementing transceiver functions. The processor 510 may use the communication interface 530 to transmit and receive data and / or information, and is used to implement Figure 2 or Figure 3 the method described in the embodiments shown.
[0151] In the embodiments of the present application, the specific connection medium between the above-mentioned processor 510, memory 520, and communication interface 530 is not limited. In the embodiments of the present application Figure 5 it is connected by a bus 540 between the processor 510, memory 520, and communication interface 530. The bus 540 is represented by a thick line in Figure 5 The connection manners between other components are only for illustrative purposes and are not to be construed as limiting. The bus may be divided into an address bus, a data bus, a control bus, etc. For the sake of convenience of representation, Figure 5 it is only represented by a thick line in, but it does not mean that there is only one bus or one type of bus.
[0152] The present application also provides a chip system, which includes at least one processor for implementing the above Figure 2 or Figure 3 the method described in the embodiments shown.
[0153] In a possible design, the chip system further includes a memory for storing program instructions and data, and the memory is located inside or outside the processor.
[0154] The chip system may be composed of chips or may include chips and other discrete devices.
[0155] The present application also provides a computer program product, which includes: a computer program (which may also be referred to as code or instruction), when the computer program is run, it causes the computer to execute as Figure 2 or Figure 3 the method described in the embodiments shown.
[0156] The present application also provides a computer-readable storage medium, which stores a computer program (which may also be referred to as code or instruction). When the computer program is run, it causes the computer to execute as Figure 2 or Figure 3 the method described in the embodiments shown.
[0157] It should be noted that the method and related device for predicting the default time provided in the embodiments of the present application can be applied to the field of artificial intelligence, or can be applied to any field other than the field of artificial intelligence. The present application does not make any limitations in this regard.
[0158] It should be understood that the processor in the embodiments of the present application can be an integrated circuit chip with signal processing capabilities. In the implementation process, the steps of the above method embodiments can be completed by the integrated logic circuit in the hardware of the processor or instructions in the form of software. The above processor can be a general-purpose processor, a digital signal processor (DSP), an application specific integrated circuit (ASIC), a field programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components. It can implement or execute the various methods, steps, and logic block diagrams disclosed in the embodiments of the present application. The general-purpose processor can be a microprocessor, or the processor can also be any conventional processor, etc. The steps of the method disclosed in combination with the embodiments of the present application can be directly embodied as being executed and completed by a hardware decoding processor, or can be executed and completed by a combination of hardware and software modules in the decoding processor. The software module can be located in a mature storage medium in the art such as a random access memory, a flash memory, a read-only memory, a programmable read-only memory, or an electrically erasable programmable memory, a register, etc. The storage medium is located in the memory, and the processor reads the information in the memory and combines its hardware to complete the steps of the above method.
[0159] It should also be understood that the memory in the embodiments of the present application may be a volatile memory or a non-volatile memory, or may include both volatile and non-volatile memories. Among them, the non-volatile memory may be a read-only memory (ROM), a programmable ROM (PROM), an erasable programmable ROM (EPROM), an electrically erasable programmable ROM (EEPROM), or a flash memory. The volatile memory may be a random access memory (RAM), which is used as an external cache. By way of example but not limitation, many forms of RAM are available, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synchlink DRAM (SLDRAM), and directrambus RAM (DR RAM). It should be noted that the memory of the systems and methods described herein is intended to include but not be limited to these and any other suitable types of memory.
[0160] As used in this specification, terms such as "unit", "module", etc. may be used to represent computer-related entities, hardware, firmware, a combination of hardware and software, software, or software in execution.
[0161] Those of ordinary skill in the art can realize that the various illustrative logical blocks and steps described in combination with the embodiments disclosed herein can be implemented in either electronic hardware or a combination of computer software and electronic hardware. Whether these functions are executed in hardware or software depends on the specific application and design constraints of the technical solution. Skilled professionals can use different methods for each specific application to implement the described functions, but such implementation should not be considered to exceed the scope of this application. In several embodiments provided in this application, it should be understood that the disclosed devices, equipment, and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the modules is only a logical function division, and there may be other division methods in actual implementation. For example, multiple modules or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication connection between each other can be through some interfaces, and the indirect coupling or communication connection of the device or module can be in electrical, mechanical, or other forms.
[0162] The modules described as separate components may or may not be physically separated, and the components shown as modules may or may not be physical modules, that is, they may be located in one place or distributed to multiple network modules. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0163] In addition, in each embodiment of this application, the functional modules can be integrated in a processing module, or each module can exist physically alone, or two or more units can be integrated in one module.
[0164] In the above embodiments, the functions of each functional module can be implemented in whole or in part by software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions (programs). When the computer program instructions (programs) are loaded and executed on a computer, the processes or functions described in the embodiments of the present application are generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center by wire (such as coaxial cable, optical fiber, digital subscriber line (DSL)) or wireless (such as infrared, wireless, microwave, etc.). The computer-readable storage medium can be any available medium that a computer can access, or a data storage device such as a server or data center that includes one or more integrated available media. The available medium can be a magnetic medium (for example, floppy disk, hard disk, magnetic tape), an optical medium (for example, digital video disc (DVD)), or a semiconductor medium (for example, solid state disk (SSD)), etc.
[0165] If the function is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to enable a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present application. The foregoing storage medium includes: various media such as USB flash drives, mobile hard disks, ROM, RAM, magnetic disks, or optical discs that can store program codes.
[0166] As described above, the above are only the specific embodiments of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art can easily think of changes or substitutions within the technical scope disclosed in the present application, and all of them should be covered by the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.
Claims
1. A method for predicting the default time, characterized in that, Applied to a server, the server is configured with a first model, a second model, and at least one third model. The first model is used to predict the default probability of a customer. The second model is used to predict the time interval to which the default time of the customer belongs. The at least one third model corresponds to at least one time interval, and each third model is used to predict the number of defaults and the default time of the customer within the corresponding time interval. The first model is a multinomial logistic Logistic model. The second model is a Gaussian mixture model GMM. The third model is an autoregressive integrated moving average ARIMA model. The first model includes multiple first sub-models, and the multiple first sub-models correspond to different customer types, and the customer types are determined according to the age group or region to which the customer belongs. The second model includes multiple second sub-models, and the multiple second sub-models correspond to different customer types, and the customer types are determined according to the age group or region to which the customer belongs. Any two of the multiple third sub-models in each third model correspond to different customer types, and the customer types are determined according to the age group or region to which the customer belongs. The method includes: Obtain data of a first customer, where the data of the first customer includes parameters for reflecting the credit risk of the first customer; Input the data into the first model to obtain the default probability of the first customer; When the default probability of the first customer is greater than a preset value, predict the time interval to which the default time of the first customer belongs through the second model; Based on the time interval to which the default time of the first customer is predicted by the second model, predict the default time of the first customer through the corresponding third model.
2. The method according to claim 1, characterized in that, The data includes one or more of the following: industry category, execution interest rate, loan amount, loan term, gender, age, education level, annual household income, employment status, unit type, housing situation, job title, social security flag, and customer level.
3. The method according to claim 1, characterized in that, The step of inputting the data into the first model to obtain the default probability of the first customer includes: Determine a first sub-model corresponding to the customer type of the first customer from the multiple first sub-models; Input the data of the first customer into the first sub-model to obtain the default probability of the first customer.
4. The method according to claim 1, characterized in that, The step of predicting the time interval to which the default time of the first customer belongs through the second model includes: Determine a second sub-model corresponding to the customer type of the first customer from the multiple second sub-models; Predict the time interval to which the default time of the first customer belongs through the second sub-model.
5. The method according to claim 1, characterized in that, The step of predicting the default time of the first customer through the corresponding third model based on the time interval to which the default time of the first customer is predicted by the second model includes: In the third model corresponding to the time interval to which the default time of the first customer is predicted by the second model, determine a third sub-model corresponding to the customer type of the first customer; Predict the default time of the first customer through the third sub-model.
6. The method according to claim 1, characterized in that, The method further includes: Obtain a training set, where the training set includes historical data of multiple customers; Based on the training set, train the first model, the second model, and the at least one third model respectively.
7. The method according to claim 6, characterized in that, The method further includes: Based on the customer types corresponding to the multiple customers respectively, group the training set to obtain multiple groups of training sets, where the customer types corresponding to the multiple groups of training sets are different, and the customer type is determined according to the age group or region to which the customer belongs; and, The step of training the first model, the second model, and the at least one third model respectively based on the training set includes: Based on each group of training sets, train the first model, the second model, and the at least one third model respectively to obtain a trained first sub-model, a second sub-model, and multiple trained third sub-models.
8. The method according to claim 6 or 7, characterized in that, The method further includes: Update the training set at a preset period to obtain an updated training set; Based on the updated training set, train the first model, the second model, and the at least one third model respectively.
9. A server, characterized in that, The server is configured with a first model, a second model, and at least one third model. The first model is used to predict the default probability of a customer. The second model is used to predict the time interval to which the default time of the customer belongs. The at least one third model corresponds to at least one time interval, and each third model is used to predict the number of defaults and the default time of the customer within the corresponding time interval. The first model is a multinomial logistic Logistic model. The second model is a Gaussian mixture model GMM. The third model is an autoregressive integrated moving average ARIMA model. The first model includes multiple first sub-models, and the multiple first sub-models correspond to different customer types, where the customer type is determined according to the age group or region to which the customer belongs. The second model includes multiple second sub-models, and the multiple second sub-models correspond to different customer types, where the customer type is determined according to the age group or region to which the customer belongs. Any two of the multiple third sub-models in each third model correspond to different customer types, and the customer type is determined according to the age group or region to which the customer belongs. The server includes: An acquisition unit, configured to acquire data of a first customer, where the data of the first customer includes parameters for reflecting the credit risk of the first customer; An input unit, configured to input the data into the first model to obtain the default probability of the first customer; A processing unit, configured to, when the default probability of the first customer is greater than a preset value, predict the time interval to which the default time of the first customer belongs through the second model; The processing unit is further configured to, based on the time interval to which the default time of the first customer predicted by the second model belongs, predict the default time of the first customer through the corresponding third model.
10. A server, characterized in that, It includes a processor and a memory communicatively connected to the processor; The memory stores computer-executable instructions; The processor executes the computer-executable instructions stored in the memory to implement the method according to any one of claims 1 to 8.
11. A computer-readable storage medium, characterized in that, A computer program, which when run on a computer, causes the computer to execute the method according to any one of claims 1 to 8.
12. A computer program product, characterized in that, A computer program, which when run, causes a computer to execute the method according to any one of claims 1 to 8.
Citation Information
Patent Citations
Office staff energy consumption behavior prediction method and system based on MCMC
CN110490379A
User default prediction method and device and electronic equipment
CN111191825A