Landing risk prediction and evaluation method, system and device and storage medium

By acquiring and preprocessing data from multiple heterogeneous data sources, extracting and enhancing customer characteristics, and building a multi-model fusion risk prediction model, the problems of single data dimensions and limited model prediction capabilities in traditional methods are solved, and higher prediction accuracy and stability are achieved.

CN120070044AInactive Publication Date: 2025-05-30深圳优讯科技有限公司
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202510550612.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-29
Publication Date
2025-05-30
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

Traditional lending risk prediction methods have a single data dimension and limited model prediction capabilities, making it difficult to comprehensively and accurately evaluate risks.

Method used

By obtaining the original data from multiple heterogeneous data sources, data preprocessing and feature extraction are carried out, the processed feature set is constructed, and the risk prediction model is trained based on this. The model obtains the initial default probability through the features processed by multiple training models, and assigns weights based on the model performance to calculate the final default probability.

Benefits of technology

It improves prediction accuracy, can more accurately identify high-risk customers, reduces the non-performing loan rate of banks, and enhances the computing efficiency and adaptability of the model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120070044A_ABST
    Figure CN120070044A_ABST
Patent Text Reader

Abstract

According to the loan risk prediction and evaluation method, system and device and the storage medium provided by the invention, the method comprises the following steps: obtaining customer-related features from original data, carrying out feature processing on the customer-related features to obtain the processed feature set, and training the risk prediction model based on the processed feature set, so that the risk prediction accuracy is improved. Inputting to-be-predicted loan application data into the trained risk prediction model to predict the default probability of loan, calculating the default probability through a plurality of models in the risk prediction model, obtaining a final probability in combination with a corresponding weight, and integrating prediction results of the plurality of models through a model fusion technology to obtain a final probability; and the prediction accuracy and stability are further improved. Compared with a traditional single model prediction method, the method has the advantages that the prediction accuracy is improved, high-risk customers can be identified more accurately, and the bad loan rate of banks is reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of data processing, and particularly to a method, system, device, and storage medium for predicting and evaluating lending risks. Background Art

[0002] With the rapid development of fintech, the bank loan business is facing increasing complexity and competitive pressure. As one of the core businesses of financial institutions, accurately predicting lending risks is crucial for ensuring asset quality, reducing the non-performing loan ratio, and enhancing the efficiency and competitiveness of credit operations.

[0003] With the development of big data and artificial intelligence technologies, banks have accumulated a vast amount of multi-dimensional data. Traditional lending risk prediction methods have problems such as single data dimension and limited model prediction ability, making it difficult to comprehensively and accurately evaluate risks. Summary of the Invention

[0004] The technical problem to be solved by this application is that traditional lending risk prediction methods have problems such as single data dimension and limited model prediction ability, making it difficult to comprehensively and accurately evaluate risks.

[0005] To solve the above problems, or at least partially solve the above technical problems, this application provides a method, system, device, and storage medium for predicting and evaluating lending risks.

[0006] In the first aspect, the present invention discloses a method for predicting and evaluating lending risks, which includes the following steps: Obtain raw data from multiple heterogeneous data sources, perform data preprocessing on the raw data, and construct a raw data set; the raw data includes internal bank data and external data; Extract customer-related features based on the raw data set, perform enhancement and screening processing on the customer-related features, and obtain a processed feature set; Train a risk prediction model according to the processed feature set, input the lending application data to be predicted into the risk prediction model, and obtain the customer default probability; the risk prediction model obtains the initial default probability obtained by the corresponding training model by inputting the processed features into multiple training models, assigns weights according to the performance of the training models, and calculates the final default probability through the initial default probability and the weights.

[0007] Preferably, the following steps are further included: At intervals of a predetermined time period, collect a new raw data set as new data input, and retrain the risk prediction model; Collect the prediction results and business feedback of the real-time risk prediction model in actual applications in real time, conduct historical data comparison and analysis, and calculate prediction metrics; the prediction metrics include prediction accuracy, approval passing rate, and default rate. Judge the performance of the risk prediction model according to the prediction metrics. When it is judged that the performance of the risk prediction model begins to decline, retrain the risk prediction model.

[0008] Preferably, the steps of obtaining the original data from multiple heterogeneous data sources, preprocessing the original data, and constructing the original data set include the following: Collect the original data of customers from inside and outside the bank at a fixed time every day, and sort the original data of customers. Perform data cleaning on the sorted original data, process the missing values in the original data, eliminate the abnormal data in the original data set, and obtain the processed original data. Associate and aggregate the processed original data according to the key attributes of the features to obtain the original data set.

[0009] Preferably, the steps of extracting customer-related features based on the original data set, enhancing and screening the customer-related features, and obtaining the processed features specifically include the following: Extract customer features from the original data set in different directions to obtain customer-related features, and the customer-related features include basic information features, basic financial features, consumption behavior features, social relationship features, and time series features. Perform cross-processing and polynomial transformation processing on the customer-related features to obtain cross-processed features and non-linear features. Perform feature selection processing based on the customer-related features, cross-processed features, and non-linear features, remove the features with lower importance, and aggregate the qualified features to obtain the processed feature set.

[0010] Preferably, the steps of performing cross-processing and polynomial transformation processing on the customer-related features to obtain cross-processed features and non-linear features specifically include the following: For the sub-features of the customer-related features, perform cross-combination processing on at least two sub-features to obtain cross-processed features. Perform polynomial transformation processing on the numerical sub-features and time series features of the customer-related features to obtain multi-term features, and establish a connection between the multi-term features and the target variable to obtain non-linear features.

[0011] Preferably, train the risk prediction model according to the processed feature set, input the loan application data to be predicted into the risk prediction model, and obtain the customer default probability. The specific steps include the following: Input the processed feature set into multiple training models to obtain the initial default probability of the data corresponding to the customer; Input the validation data set into each model, obtain the relevant metrics of the output data of each model, and calculate the weights of each model according to the relevant metrics; Perform weighted average calculation based on the initial default probability output by each model and the calculated weights of each model to obtain the final default probability.

[0012] Preferably, the training model is constructed based on the k-means++ algorithm, the LGBMClassifier algorithm, and the logisticRegression algorithm.

[0013] In a second aspect, the present invention discloses a lending risk prediction and evaluation system, which includes the above-mentioned lending risk prediction and evaluation method.

[0014] In a third aspect, the present invention discloses a computer device, which includes a processor, a communication interface, a memory, and a communication bus. Among them, the processor, the communication interface, and the memory complete mutual communication through the communication bus; The memory is used to store computer programs; The processor is used to implement the steps of the above method when executing the program stored on the memory.

[0015] In a fourth aspect, the present invention discloses a computer-readable storage medium, on which a computer program is stored, and the computer program realizes the steps of the above method when executed by a processor.

[0016] The above technical solutions provided by the present application have the following advantages compared with the prior art: The lending risk prediction and evaluation method, system, device, and storage medium provided by the present application. The method mentioned that relevant customer features are obtained from the original data, the customer-related features are processed to obtain a processed feature set, a training risk prediction model is based on the processed feature set, and then the lending application data to be predicted is input into the trained risk prediction model to predict the default probability of the lending. In the risk prediction model, the default probability is calculated through multiple models, and the final probability is obtained by combining the corresponding weights. Through the model fusion technology, the prediction results of multiple models are integrated, further improving the accuracy and stability of the prediction. Compared with the traditional single-model prediction method, the present invention improves the prediction accuracy rate, can more accurately identify high-risk customers, and reduces the non-performing loan rate of banks. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] The accompanying drawings herein are incorporated into the specification and constitute a part of the specification, showing embodiments consistent with the present invention, and are used together with the specification to explain the principles of the present invention.

[0018] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the accompanying drawings required for the description of the embodiments or the prior art. Obviously, for those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0019] Figure 1 Schematic flow of a lending risk prediction and assessment method provided by this application Figure 1 ; Figure 2 Schematic flow of a lending risk prediction and assessment method provided by this application Figure 2 。 Detailed implementation manners

[0020] To make the objectives, technical solutions, and advantages of this application clearer, the following will clearly and completely describe the technical solutions in this application with reference to the accompanying drawings in this application. Obviously, the described embodiments are part of the embodiments of this application, rather than all of them. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in this application without creative efforts fall within the scope of protection of this application.

[0021] In a first aspect, referring to Figures 1 to 2 , the present invention discloses a lending risk prediction and assessment method, which includes the following steps: Step S1: Obtain raw data from multiple heterogeneous data sources, perform data preprocessing on the raw data, and construct a raw data set; the raw data includes internal bank data and external data; Step S2: Extract customer-related features based on the raw data set, perform enhancement and screening processing on the customer-related features, and obtain a processed feature set; Step S3: Train a risk prediction model according to the processed feature set, input the lending application data to be predicted into the risk prediction model, and obtain the customer default probability; the risk prediction model obtains the initial default probability obtained by the corresponding training model by inputting the processed features through multiple training models, assigns weights according to the performance of the training models, and calculates the final default probability through the initial default probability and the weights.

[0022] Specifically, in step S1, through various data interface technologies such as API interfaces, database connections, and file imports, the connection with different data sources is achieved. The data includes multi-dimensional data such as customer information, account information, transaction records, asset information, and behavioral characteristics within the bank, as well as external credit rating data, market data, macroeconomic indicators, and industry data. Since the formats and qualities of the data are different after acquisition, it is necessary to standardize the format of the same data to ensure data quality and availability, and data preprocessing is required.

[0023] Specifically, in step S2, features are extracted from the original data set to identify features that may have an impact on loan risk. Then, the extracted features are enhanced and screened. The most representative and predictive feature subset is selected from numerous features, redundant and irrelevant features are removed, the dimension of the model is reduced, and the computational efficiency and accuracy of the model are improved.

[0024] Specifically, in step S3, the processed feature set obtained in step S2 is input into multiple training models for training to determine the parameters of the training models, so that the default probability output by the training models is more in line with the actual situation. Then, the initial default probability output by the training models is combined with the weights assigned according to the performance of the models to calculate the final default probability. The final default probability is compared with the actual situation to determine the accuracy of the trained model. The loan application data to be predicted is input into the trained risk prediction model, and a default probability is output, which is the default probability of the customer applying for a loan.

[0025] Specifically, customer-related features are obtained from the original data, the customer-related features are processed to obtain a processed feature set, a risk prediction model is trained based on the processed feature set, and then the loan application data to be predicted is input into the trained risk prediction model to predict the default probability of the loan. In the risk prediction model, the default probability is calculated through multiple models, and the final probability is obtained by combining the corresponding weights. Through model fusion technology, the prediction results of multiple models are integrated, further improving the accuracy and stability of the prediction. Compared with the traditional single-model prediction method, the present invention improves the prediction accuracy rate, can more accurately identify high-risk customers, and reduces the non-performing loan rate of the bank.

[0026] Furthermore, the loan risk prediction method has high computational efficiency, can quickly process a large amount of loan application data, and provides timely and accurate risk assessment results for the bank's credit business. By adopting efficient algorithms and model fusion technology, the prediction speed is improved, and the credit approval time is greatly shortened. At the same time, the method of the present invention also has good scalability, can adapt to the continuous expansion of the scale of the bank's credit business, and provides strong support for the bank to improve the efficiency of the credit business.

[0027] Furthermore, it provides new technical means and ideas for bank financial risk management, promoting technological innovation and development in the field of financial risk management. By adopting advanced data mining and machine learning technologies, the present invention can evaluate lending risks more comprehensively and accurately, providing strong support for banks to formulate risk control strategies, pricing decisions, etc. At the same time, the method of the present invention also has strong interpretability, which can provide intuitive risk assessment results and decision-making basis for bank credit approval personnel, helping to improve the scientificity and transparency of bank risk management.

[0028] Furthermore, through feature selection and model fusion technologies, the generalization ability of the model is enhanced, enabling it to adapt to changes in different customer groups and market environments. During the feature selection process, relevant features strongly related to lending risks are screened out through statistical methods such as correlation analysis, chi-square test, and mutual information, avoiding the problem of overfitting.

[0029] The following steps are also included afterwards: Step S4: At intervals of a predetermined time period, collect a new set of original data as new data input and retrain the risk prediction model; Step S5: Collect the prediction results and business feedback of the risk prediction model in actual applications in real time, conduct a comparative analysis of historical data, and calculate prediction metrics; the prediction metrics include prediction accuracy, approval passing rate, and default rate; Step S6: Judge the performance of the risk prediction model according to the prediction metrics. If it is judged that the performance of the risk prediction model begins to decline, retrain the risk prediction model.

[0030] Specifically, in step S4, as time goes by and data changes, the performance of the model may fluctuate. Regular model evaluation can timely detect changes in model performance and monitor the stability of the model. If the model performance declines, timely measures can be taken for adjustment and optimization to ensure that the model always maintains good prediction ability. In order to ensure the accuracy of the model output, it is necessary to detect the performance level of the model at intervals to clarify the performance level of the model, so as to judge whether the model meets the actual needs of bank lending business, thereby accurately identifying high-risk customers and avoiding misjudgment and missed judgment.

[0031] Specifically, in step S5, by comparing and analyzing the input test data with the actual situation and historical data, prediction metrics are calculated, including various model metrics such as accuracy, precision, recall, F1-Score (F1 value), AUC (Area Under the Curve, the area under the ROC curve), KS value (Kolmogorov-Smirnov), and mean squared error. These metrics can reflect the current performance of the model.

[0032] Specifically, in step S6, the performance of the current model is judged according to the calculated metrics. When the performance deteriorates, the risk prediction model can be retrained to restore its normal use.

[0033] It can be understood that the data of bank lending business is constantly changing, and factors such as customers' behavior patterns and economic environment may change. If the model is not adjusted, it may not be able to adapt to these changes, resulting in a decline in prediction performance. By adjusting the model, the parameters and structure of the model can be updated according to the new data characteristics and distribution, so that it can better fit the current data and improve the accuracy of prediction. By evaluating the model using various metrics such as accuracy, precision, recall, F1-Score, and AUC, it is possible to comprehensively and objectively understand the performance of the model in predicting customers' lending risks.

[0034] As an embodiment, a detection and warning mechanism is set for the model performance. When the model performance metrics are lower than the preset threshold, an alarm is automatically triggered to remind relevant personnel to update and optimize the model. For example, when the accuracy of the model drops below 80%, the system automatically sends a notification to the model maintenance personnel.

[0035] Step S1 includes the following steps: Step S11: Regularly collect the original data of customers from inside and outside the bank every day, and sort the original data of customers; Step S12: Clean the sorted original data, process the missing values in the original data, and eliminate the abnormal data in the original data set to obtain the processed original data; Step S13: Associate and aggregate the processed original data according to the key attributes of the features to obtain the original data set.

[0036] Specifically, data cleaning, data integration, and data transformation are performed on the data. Among them, data cleaning includes, for the missing customer income data, if the data distribution is relatively uniform, filling it with the median income of all customers; if there is an obvious correlation, constructing a regression model using relevant features such as age and occupation for prediction and filling. For the outliers in the consumption amount, first calculate the mean and standard deviation, and detect according to the 3σ principle. If the number of outliers is small and has a large impact on the overall data, replace it with the upper and lower limits; if there are many outliers, further judgment and processing can be combined with the Isolation Forest model. Data integration includes associating and integrating the data from the core business system, credit management system, and customer relationship management system based on the customer ID. For example, through a rule-based matching algorithm, the data from different systems corresponding to the same customer ID are merged to form a comprehensive customer portrait data. Data transformation includes standardizing the numerical features and converting them into a distribution with a mean of 0 and a variance of 1, using the Z-score standardization formula: zi = (xi - μ) / σ, where xi is the original data, μ is the mean, and σ is the standard deviation. After preprocessing the data, associate the data with the customers to form a set, and obtain the corresponding original data set. During the integration process, the data from different data sources are associated and merged according to certain rules to form a complete data set, providing a basis for subsequent feature engineering and model calculation.

[0037] Step S2 specifically includes the following steps: Step S21: Extract customer features from the original data set in different directions to obtain customer-related features, where the customer-related features include basic information features, basic financial features, consumption behavior features, social relationship features, and time series features; Step S22: Perform cross-processing and polynomial transformation processing on the customer-related features to obtain cross-processed features and non-linear features; Step S23: Perform feature selection processing based on the customer-related features, cross-processed features, and non-linear features, remove the features with lower importance, and collect the qualified features to obtain the processed feature set.

[0038] Specifically, features are extracted from the original data set. Among them, the basic information features include data such as the customer's age and name. The basic financial features include the customer's debt-to-asset ratio (total debt / total assets), income-to-debt ratio (monthly income / monthly repayment amount), and credit limit utilization rate (used credit limit / total credit limit). Statistical features such as the mean, variance, maximum, and minimum of the customer's consumption amount are extracted as the customer's consumption behavior features, and the customer's consumption level and stability are understood based on these features. A customer social network graph is constructed, the degree centrality of the nodes (the number of nodes with direct social relationships with this customer) is calculated to obtain social relationship features. The social relationship features can reflect the activity of the customer in the social network. Customers with a higher degree centrality may have a wider range of social relationships. The betweenness centrality (the frequency at which the customer acts as an intermediate node to connect other customers) is calculated to evaluate the influence of the customer in the social network. Customers with a higher betweenness centrality may play a key role in information dissemination and resource allocation, and their behavioral changes may have an impact on other customers. Fourier transform is performed on the time series data to extract frequency domain features and understand the periodic and seasonal changes of the data. For example, after Fourier transform of the time series of the customer's transaction amount, the annual consumption peaks and troughs can be identified. The autocorrelation coefficient of the time series is calculated to analyze the smoothness of the data. A time series with a high autocorrelation coefficient has strong smoothness and is suitable for prediction using a linear model. Wavelet transform is applied to decompose the time series and extract features at different resolutions. Wavelet transform can decompose the time series into subsequences of different frequencies and analyze the features of the high-frequency and low-frequency parts respectively.

[0039] Step S22 specifically includes the following steps: Step S221: For the sub-features of the customer-related features, at least two sub-features are cross-combined to obtain cross-processed features; Step S222: Polynomial transformation is performed on the numerical sub-features and time series features of the customer-related features to obtain polynomial features, and non-linear features are obtained by relating the polynomial features to the target variable.

[0040] Specifically, feature crossing and polynomial feature generation are important means to improve the model performance. Feature crossing can combine two, three or more sub-features to create new features. The new features can reflect the relationship between the sub-features and reflect the customer's consumption situation under some specific conditions. Polynomial transformation can generate quadratic terms, cubic terms, etc., which can capture the non-linear relationship between the features and the loan risk. For time series features such as transaction amount, polynomial features are also generated, generating quadratic and cubic terms to enhance the model's ability to model the time series trend, enabling the model to better capture the non-linear growth or decline trend of the transaction amount over time.

[0041] Step S3 specifically includes the following steps: Step S31: Input the processed feature set into multiple training models to obtain the initial default probability of the data corresponding to the customer; Step S32: Input the validation data set into each model, obtain the relevant metrics of the output data of each model, and calculate the weights of each model according to the relevant metrics; Step S33: Perform a weighted average calculation based on the initial default probability output by each model and the calculated weights of each model to obtain the final default probability.

[0042] Specifically, weights are assigned to each model according to its performance metrics (such as accuracy, AUC, etc.). The model with a higher weight occupies a greater proportion in the fusion result. For example, the weight of LGBMClassifier is 0.6, and the weight of logisticRegression is 0.4. The prediction probabilities of the two models are combined through weighted average to obtain the final prediction result. The prediction probabilities of each model are weighted averaged, and the weights can be determined based on the performance metrics of the model or the results of cross-validation. For example, the weight of LGBMClassifier is 0.5, the weight of logisticRegression is 0.3, and the weight of the clustering result of the K-means++ algorithm is 0.2. The final fusion result is a linear combination of the prediction probabilities of each model. As an embodiment, methods such as grid search or Bayesian optimization can be used to automatically find the optimal combination of model weights to maximize the prediction performance of the fusion model.

[0043] Specifically, the loan application data to be predicted is input into the trained model after data preprocessing and feature extraction. The model will output an initial default probability. Weights are calculated according to the performance of each model. Combining the calculated weights and the initial default probability, the final default probability can be obtained. The output formula is: , where P is the final default probability, A is the calculated weight of each model, B is the initial default probability of each model, i is the number of models, and n is the actual number of models. , it can be understood that the sum of the weights is 1. According to the set threshold (such as 0.5), customers with a default probability greater than or equal to the threshold are classified as high-risk groups, and customers with a default probability less than the threshold are classified as low-risk groups. Different thresholds can also be determined through cost-benefit analysis to balance risk control and business expansion. In addition, for dynamic risks, according to the real-time behavior data of customers, such as transaction flow, credit limit usage, etc., the risk prediction results are dynamically updated. For example, when a customer makes frequent large-scale purchases in a short period of time, the model can timely adjust its risk assessment results and give an early warning of potential default risks.

[0044] It can be understood that the model fusion technology combines the advantages of multiple models, improving the adaptability of the model to different data distributions. In practical applications, different models exhibit good predictive performance on different datasets, with significantly enhanced generalization ability, enabling them to better cope with changes in the market environment and customer behavior.

[0045] As an embodiment, the training model is constructed for binary classification based on the k-means++ algorithm, DBSCAN (Density-Based Spatial Clustering of Applications with Noise), LGBMClassifier algorithm (LightGBM Classifier), XGBoost algorithm, logisticRegression algorithm (Logistic Regression algorithm), and Support Vector Machine (SVM).

[0046] In this embodiment, the k-means++ algorithm, LGBMClassifier algorithm, and logisticRegression algorithm are used to train a risk assessment model. Among them, the k-means++ algorithm first clusters the customer data and calculates the proportion of actual defaulting customers in each cluster to estimate the default probability of customers in that cluster. For example, suppose there are 5 clusters, and there are 100 customers in a certain cluster, among which 10 default. Then the estimated default probability of the customers in this cluster is 10%. In actual application, it will be used in combination with other models. The clustering result 1 is added as a new feature to the original data. For example, the cluster number to which the customer belongs is used as a new feature, and then it is input into other models (such as the LGBMClassifier model or the logisticRegression model) to assist these models in calculating the default probability more accurately. In the training stage, the LGBMClassifier model will learn based on a large amount of input customer data with labels (default or not). It is based on gradient boosting decision trees and continuously iteratively constructs multiple decision trees to mine complex non-linear relationships from the data. During the training process, the model will learn the influence patterns of different features (such as customer income, debt, etc.) on default. When predicting, the feature data of the customer to be evaluated is input into the trained LGBMClassifier model, and the decision trees inside the model will make layer-by-layer judgments on these features. Each decision tree branches according to the feature values, and finally the results of all decision trees are aggregated, and the default probability of the customer is obtained through a certain algorithm (such as a voting mechanism). The logisticRegression model is based on linear regression, and the linear regression result is mapped to between 0 and 1 through a logistic function to obtain the default probability. Before training, the customer feature data needs to be preprocessed, such as standardization, encoding, etc., to make the data more suitable for model learning. During training, the model will use methods such as maximum likelihood estimation to determine the parameters of the model (such as the coefficients of each feature) according to the input customer feature data and the corresponding default or not labels. These parameters reflect the influence degree of each feature on the default probability. When predicting, the feature data of the new customer is substituted into the trained logistic regression model, and the model will calculate a linear combination value, which is then transformed through the logistic function, and finally the default probability of the customer is output. Suppose the linear combination value calculated by the logistic regression model is z, and through the logistic function , the initial default probability is obtained . This embodiment adopts the above three models. Due to the complementarity between the three models, the k-means++ algorithm is used for data clustering to provide new feature information for subsequent models, which is helpful to explore the potential structure of the data; the LGBMClassifier model can handle complex nonlinear relationships and is good at learning complex patterns from large amounts of data; the logisticRegression model is a linear model with good interpretability. The three algorithms analyze and model data from different angles and complement each other to improve the performance and prediction accuracy of the overall model. In addition, the LGBMClassifier model is efficient in processing large-scale data and can converge quickly. The logisticRegression model is simple to calculate and has a fast training speed. The combined use can improve computing efficiency while ensuring model performance to meet the needs of practical applications.

[0047] For the performance of each model, pre-prepared test data and corresponding results, input the corresponding test data for each model, obtain the results corresponding to the test data, compare them with the pre-prepared corresponding results, judge the performance according to the test indicators, combine the performance and actual needs, allocate and calculate weights, and then combine them with the initial default probability to get the final default probability.

[0048] In a second aspect, the present invention discloses a loan risk prediction and evaluation system, which includes the above-mentioned loan risk prediction and evaluation method.

[0049] Specifically, the first aspect of the system implementation discloses a loan risk prediction and assessment method, obtains customer-related features from raw data, performs feature processing on customer-related features, obtains a processed feature set, trains a risk prediction model based on the processed feature set, and then inputs the loan application data to be predicted into the trained risk prediction model to predict the default probability of the loan. The default probability is calculated by multiple models in the risk prediction model, and the final probability is obtained by combining the corresponding weights. The prediction results of multiple models are integrated through model fusion technology, further improving the accuracy and stability of the prediction. Compared with the traditional single model prediction method, the present invention improves the prediction accuracy, can more accurately identify high-risk customers, and reduce the bank's non-performing loan rate.

[0050] In a third aspect, the present invention discloses a computer device, which includes a processor, a communication interface, a memory and a communication bus, wherein the processor, the communication interface and the memory communicate with each other via the communication bus; Memory, used to store computer programs; The processor is used to implement the steps of the above method when executing the program stored in the memory.

[0051] Specifically, the computer device can implement the lending risk prediction and assessment method disclosed in the first aspect, obtain customer-related features from the original data, perform feature processing on the customer-related features to obtain a processed feature set, train a risk prediction model based on the processed feature set, and then input the lending application data to be predicted into the trained risk prediction model to predict the default probability of the lending. In the risk prediction model, the default probability is calculated through multiple models, and the final probability is obtained by combining the corresponding weights. Through the model fusion technology, the prediction results of multiple models are integrated, further improving the accuracy and stability of the prediction. Compared with the traditional single-model prediction method, the present invention improves the prediction accuracy rate, can more accurately identify high-risk customers, and reduces the non-performing loan rate of the bank.

[0052] In a fourth aspect, the present invention discloses a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the steps of the above method are implemented.

[0053] Specifically, customer-related features are obtained from the original data, feature processing is performed on the customer-related features to obtain a processed feature set, a risk prediction model is trained based on the processed feature set, and then the lending application data to be predicted is input into the trained risk prediction model to predict the default probability of the lending. In the risk prediction model, the default probability is calculated through multiple models, and the final probability is obtained by combining the corresponding weights. Through the model fusion technology, the prediction results of multiple models are integrated, further improving the accuracy and stability of the prediction. Compared with the traditional single-model prediction method, the present invention improves the prediction accuracy rate, can more accurately identify high-risk customers, and reduces the non-performing loan rate of the bank.

[0054] In the above embodiments, the descriptions of the respective embodiments have their own emphases. For the parts not described in detail in a certain embodiment, reference may be made to the relevant descriptions of other embodiments.

[0055] In the description of the present invention, it should be understood that the orientation or positional relationship indicated by the terms "center", "longitudinal", "transverse", "length", "width", "thickness", "upper", "lower", "front", "rear", "left", "right", "vertical", "horizontal", "top", "bottom", "inner", "outer", "clockwise", "counterclockwise", etc. is based on the orientation or positional relationship shown in the drawings, and is only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and therefore should not be construed as a limitation of the present invention.

[0056] In addition, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, features defined with "first" and "second" may explicitly or implicitly include one or more of such features. In the description of the present invention, "a plurality of" means two or more unless specifically defined otherwise.

[0057] In the present invention, unless otherwise clearly specified and limited, the terms such as "mounted", "connected", "coupled", "fixed", etc. should be understood in a broad sense. For example, it may be a connection, a detachable connection, or integrated; it may be a mechanical connection or an electrical connection; it may be directly connected or indirectly connected through an intermediate medium, and it may be the internal communication of two elements or the interaction relationship between two elements. For those of ordinary skill in the art, the specific meanings of the above terms in the present invention can be understood according to specific circumstances.

[0058] In the present invention, unless otherwise clearly specified and limited, the first feature being "on" or "under" the second feature may include the direct contact between the first and second features, or may include the situation where the first and second features are not in direct contact but in contact through additional features therebetween. Moreover, the first feature being "above", "over" and "on top of" the second feature includes that the first feature is directly above and obliquely above the second feature, or merely means that the horizontal height of the first feature is higher than that of the second feature. The first feature being "under", "beneath" and "underneath" the second feature includes that the first feature is directly below and obliquely below the second feature, or merely means that the horizontal height of the first feature is lower than that of the second feature.

[0059] In the description of this specification, the descriptions with reference to the terms "one embodiment", "some embodiments", "example", "specific example", or "some examples", etc. mean that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic representations of the above terms should not be understood as necessarily referring to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described can be combined in a suitable manner in any one or more embodiments or examples. In addition, those skilled in the art can combine and combine the different embodiments or examples described in this specification.

[0060] Obviously, those skilled in the art can make various modifications and variations to the present invention without departing from the spirit and scope of the present invention. Thus, if these modifications and variations of the present invention fall within the scope of the claims of the present invention and their equivalent technologies, the present invention is also intended to include these modifications and variations.

[0061] As described above, the specific implementation manners of the present invention are provided, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention can easily think of various equivalent modifications or substitutions, and these modifications or substitutions should be covered within the protection scope of the present invention. Therefore, the protection scope of the present invention shall be subject to the protection scope of the claims.

Claims

1. A loan risk prediction and assessment method, characterized in that: The following steps are involved: Acquire raw data from multiple heterogeneous data sources, perform data preprocessing on the raw data, and construct a raw data set; the raw data includes internal bank data and external data; Extract customer-related features based on the original data set, enhance and filter the customer-related features, and obtain a processed feature set; The risk prediction model is trained according to the processed feature set, and the loan application data to be predicted is input into the risk prediction model to obtain the customer default probability; the risk prediction model obtains the initial default probability obtained by the corresponding training model through the processed features input by multiple training models, and the weight is allocated according to the performance of the training model, and the final default probability is calculated by the initial default probability and the weight.

2. The loan risk prediction and assessment method according to claim 1, characterized in that: The following steps are then included: At predetermined intervals, new raw data sets are collected as new data inputs to retrain the risk prediction model; Collect the prediction results and business feedback of the risk prediction model in actual application in real time, conduct comparative analysis of historical data, and calculate prediction indicators; the prediction indicators include prediction accuracy, approval rate, and default rate; The performance of the risk prediction model is judged according to the prediction indicators. If the performance of the risk prediction model begins to decline, the risk prediction model is retrained.

3. The loan risk prediction and assessment method according to claim 1, characterized in that: The method of obtaining raw data from multiple heterogeneous data sources, performing data preprocessing on the raw data, and constructing a raw data set includes the following steps: Collect and obtain original customer data from inside and outside the bank on a daily basis and sort the original customer data; The sorted raw data is cleaned to process missing values ​​in the raw data and eliminate abnormal data in the raw data set to obtain processed raw data; The processed raw data are associated and aggregated according to the key attributes of the features to obtain the raw data set.

4. The loan risk prediction and assessment method according to claim 1, characterized in that: The method of extracting customer-related features based on the original data set, enhancing and screening the customer-related features, and obtaining processed features specifically includes the following steps: Extracting customer features from the original data set in different directions to obtain customer-related features, wherein the customer-related features include basic information features, basic financial features, consumption behavior features, social relationship features, and time series features; Perform cross processing and polynomial transformation on customer-related features to obtain cross processing features and nonlinear features; Feature selection is performed based on customer-related features, cross-processing features and nonlinear features, features with lower importance are removed, and features that meet the requirements are grouped to obtain a processed feature set.

5. The loan risk prediction and assessment method according to claim 4, characterized in that: Cross-processing and polynomial transformation are performed on customer-related features to obtain cross-processing features and nonlinear features, which specifically includes the following steps: For the sub-features of the customer-related features, cross-combining at least two sub-features to obtain cross-processed features; Polynomial transformation is performed on the numerical sub-features and time series features of customer-related features to obtain multi-term features. The multi-term features are then linked to the target variables to obtain nonlinear features.

6. The loan risk prediction and assessment method according to claim 1, characterized in that: The risk prediction model is trained based on the processed feature set, and the loan application data to be predicted is input into the risk prediction model to obtain the customer default probability, which specifically includes the following steps: Based on the processed feature set, it is input into multiple training models to obtain the initial default probability of the customer corresponding to the data; Input validation data sets into each model, obtain relevant indicators of the output data of each model, and allocate calculation weights of each model according to relevant indicators; The final default probability is obtained by taking a weighted average of the initial default probability output by each model and the calculation weight of each model.

7. The loan risk prediction and assessment method according to claim 1, characterized in that: The training model is constructed based on the k-means++ algorithm, the LGBMClassifier algorithm, and the logisticRegression algorithm.

8. A loan risk prediction and evaluation system, characterized in that: Including the loan risk prediction and assessment method as described in any one of claims 1-7 above.

9. A computer device, characterized in that: It includes a processor, a communication interface, a memory and a communication bus, wherein the processor, the communication interface and the memory communicate with each other through the communication bus; Memory, used to store computer programs; A processor, for implementing the steps of the method described in any one of claims 1 to 6 when executing a program stored in a memory.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 6 are implemented.

Citation Information

Patent Citations

  • In-loan behavior monitoring method and system

    CN111324862A

  • Credit risk assessment method and system, terminal equipment and storage medium

    CN114240633A

  • Construction method of credit risk assessment model and credit risk assessment model

    CN119444397A