Financial risk prediction method based on artificial intelligence

By comprehensively utilizing multi-source data and gradient-enhancing decision tree models, a financial risk prediction model is constructed, which solves the problems of insufficient generalization capabilities of models and lack of effective risk warning mechanisms in the existing technology, and achieves high-precision and high-reliability financial risk prediction and real-time warning.

CN120013661APending Publication Date: 2025-05-16GUANGDONG OCEAN UNIVERSITY
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510177626.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-18
Publication Date
2025-05-16

AI Technical Summary

Technical Problem

The existing financial risk prediction methods based on artificial intelligence have problems such as insufficient generalization capabilities of model, insufficient data processing, and lack of effective risk warning mechanisms, which cannot meet the requirements of financial institutions for high accuracy, high reliability and real-time risk prediction.

Method used

By comprehensively utilizing multi-source data and gradient enhancement decision tree model, collecting and preprocessing multi-source data of borrowers, building a financial risk prediction model based on gradient enhancement decision tree, iterative training and hyperparameter tuning, and ultimately real-time risk prediction and early warning.

Benefits of technology

It improves the accuracy and reliability of financial risk prediction, provides scientific and effective risk warning and decision-making support, reduces financial risks, and ensures the stable development of financial business.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120013661A_ABST
    Figure CN120013661A_ABST
Patent Text Reader

Abstract

The invention provides a financial risk prediction method based on artificial intelligence, and the method comprises the steps: collecting and preprocessing multi-source data of a borrowing customer, the multi-source data comprising customer basic information, financial data, credit records, loan information and external macroscopic data; a financial risk prediction model based on a gradient boosting decision tree is constructed, the financial risk prediction model is formed by iteratively integrating a plurality of decision trees, and each decision tree fits a residual error between a previous round of model prediction result and a true value; training the financial risk prediction model through the preprocessed data; and the bank financial risk is predicted through the trained financial risk prediction model, and risk early warning is carried out according to a prediction result. The multi-source financial data is combined, and accurate risk prediction of financial services such as bank loan is achieved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of artificial intelligence and financial technology, and specifically relates to a financial risk prediction method based on artificial intelligence. Background Art

[0002] In the financial field, especially in the bank loan business, accurate prediction of financial risks is crucial. Traditional financial risk prediction methods mainly rely on manual experience and simple statistical analysis, which have many limitations. On the one hand, manual experience judgment is highly subjective, lacks unified standards and quantitative evaluation, and is easily affected by personal cognition and emotions, resulting in low accuracy and reliability of risk prediction. On the other hand, simple statistical analysis methods have difficulty processing complex, high-dimensional data, cannot capture potential patterns and nonlinear relationships in the data, and cannot adapt to the rapid changes in the financial market and the diversification of customer needs.

[0003] With the continuous development and innovation of the financial market, financial products and services are becoming increasingly complex, and customers’ credit status and risk characteristics are becoming more elusive. At the same time, a large amount of financial data is constantly generated. How to extract valuable information from these massive data and accurately predict financial risks has become an important challenge facing financial institutions.

[0004] In recent years, artificial intelligence technology has been widely used in various fields, providing new ideas and methods for financial risk prediction. However, the existing financial risk prediction methods based on artificial intelligence still have some problems, such as insufficient generalization ability of the model, insufficient processing of data, lack of effective risk warning mechanism, etc., which cannot meet the requirements of financial institutions for high accuracy, high reliability and real-time risk prediction.

[0005] Therefore, a more advanced, accurate and efficient AI-based financial risk prediction method is needed to improve the risk prediction capabilities of bank loan business and ensure the asset security and sound operation of financial institutions. Summary of the invention

[0006] The purpose of the present invention is to provide a financial risk prediction method based on artificial intelligence, which improves the accuracy and reliability of financial risk prediction by comprehensively utilizing multi-source data and gradient boosting decision tree models, provides scientific and effective risk warning and decision support for banks and other financial institutions, reduces financial risks, and ensures the stable development of financial business.

[0007] The technical solution of the present invention is as follows:

[0008] A financial risk prediction method based on artificial intelligence, the method comprising:

[0009] Collect and pre-process the multi-source data of borrowers, including basic customer information, financial data, credit records, loan information, and external macro data;

[0010] Constructing a financial risk prediction model based on a gradient boosting decision tree, wherein the financial risk prediction model is formed by iterative integration of multiple decision trees, and each decision tree fits the residual between the prediction result of the previous round of model and the true value;

[0011] Training the financial risk prediction model through the preprocessed data;

[0012] Bank financial risks are predicted through trained financial risk prediction models, and risk warnings are issued based on the prediction results.

[0013] Furthermore, the collecting of multi-source data of borrowers and preprocessing specifically includes:

[0014] Collect data from internal bank systems, credit reporting agencies, and public data sources;

[0015] Clean the collected data, including processing missing values, outliers and duplicate data;

[0016] Standardize and encode the cleaned data to eliminate the dimension effects between features and convert categorical variables into numerical types;

[0017] The preprocessed data is divided into training set, validation set and test set according to a certain ratio.

[0018] Furthermore, when processing missing values, for numerical data, the mean is used for filling according to the missing ratio, and for categorical data, the mode is used for filling; the processing of outliers adopts statistical methods; the standardization adopts the Z-score standardization method; and the encoding adopts the one-hot encoding method.

[0019] Furthermore, the construction of a financial risk prediction model based on a gradient boosting decision tree specifically includes:

[0020] The initial model prediction value is the mean of the target values ​​of all training samples;

[0021] In each iteration, the residual between the model prediction value and the true value in the previous round is calculated;

[0022] The residual is used as the target value to train a new decision tree. The decision tree contains a root node, internal nodes, and leaf nodes. The internal nodes divide the samples based on the feature threshold.

[0023] Multiply the prediction results of the newly trained decision tree by the learning rate and add them to the prediction results of the previous round of model to update the financial risk prediction model.

[0024] Furthermore, when training a decision tree, hyperparameters need to be determined, including the maximum depth of the tree, the minimum number of sample splits, and the minimum number of sample leaves.

[0025] Furthermore, the method of predicting bank financial risks through the trained prediction model specifically includes:

[0026] Obtain the latest information of borrowers in real time and perform the same preprocessing as training data;

[0027] Input the preprocessed real-time data into the trained financial risk prediction model and output the loan default probability;

[0028] Compare the probability of default with the set risk threshold to determine the loan risk level.

[0029] Furthermore, the risk warning according to the prediction results specifically includes:

[0030] When the prediction result is a high-risk loan, an early warning signal containing risk type and level information will be sent to relevant credit managers through the bank's internal system.

[0031] Furthermore, the prediction results are used for financial business decision-making, including but not limited to adjusting loan amounts, interest rate pricing, and loan term setting; stratifying loan customers according to risk levels and formulating differentiated risk management strategies for customers of different risk levels.

[0032] Compared with the prior art, the present invention has the following advantages:

[0033] By comprehensively utilizing multi-source data and the gradient boosting decision tree model, the present invention can more comprehensively and deeply mine the potential information and patterns in the data and capture the nonlinear relationships in the data, thereby improving the accuracy and reliability of financial risk prediction. BRIEF DESCRIPTION OF THE DRAWINGS

[0034] The accompanying drawings generally illustrate various embodiments by way of example and not limitation, and together with the description and claims, serve to illustrate the embodiments of the invention. Where appropriate, the same reference numerals are used throughout the drawings to refer to the same or similar parts. Such embodiments are illustrative and are not intended to be exhaustive or exclusive embodiments of the present apparatus or method.

[0035] Figure 1 A schematic flow chart of the method of the present invention is shown. DETAILED DESCRIPTION

[0036] It should be noted that, in the absence of conflict, the embodiments and features in the embodiments of the present application can be combined with each other. The present application will be described in detail below with reference to the accompanying drawings and in combination with the embodiments.

[0037] like Figure 1 As shown, an embodiment of the present invention provides a financial risk prediction method based on artificial intelligence, comprising:

[0038] Data collection and preprocessing

[0039] Data Collection

[0040] Internal bank systems: Regularly extract customer information from the bank's core business system through a dedicated data interface. For example, at the beginning of each month, extract basic information about new borrowers from the previous month, including personal identity information, occupational status, income flow, and detailed loan-related data, such as loan amount, term, purpose, repayment plan, and past repayment records.

[0041] Credit reporting agencies: According to the established cooperation agreement, the credit report of the customer is obtained from the credit reporting agency on a regular basis. For example, once a quarter, the content covers the customer's borrowing and lending situation in other financial institutions, credit card usage details, credit score, and any bad credit records such as overdue payments and defaults.

[0042] Public data sources: Use professional data collection tools to obtain macroeconomic data from government statistical department websites, such as monthly collection of GDP growth rate, inflation rate and other data; obtain interest rate levels, industry prosperity index and other information from financial information websites, and update at least once a week. At the same time, through public opinion monitoring tools, capture dynamic information related to the borrower's industry from social media and news websites, such as policy changes, market trends, etc., and summarize and analyze them daily.

[0043] Data cleaning

[0044] Missing value processing: For numerical data, the missing ratio is calculated through statistical analysis. If the missing ratio of a feature is less than 20%, the missing value is directly filled with the mean of the feature; if the missing ratio is higher, an expert team can be organized to select appropriate interpolation methods or use machine learning algorithms for prediction and filling based on business experience and data characteristics. For categorical data, the most frequent category (mode) is directly used to fill the missing value.

[0045] Outlier processing: Use the Z-score method to detect outliers on numerical data. Calculate the Z-score value of each data point, and consider data points with an absolute Z-score greater than 3 as outliers. According to the cause of the outlier, if it is caused by data entry errors, correct it directly; if it is a real abnormal situation, communicate with the business department to determine whether to keep or delete it.

[0046] Duplicate data processing: Use data processing tools to compare the various fields of each data record, filter out and delete the identical duplicate records, and keep only one to ensure the uniqueness and validity of the data.

[0047] Data standardization and coding

[0048] Standardization: The Z-score standardization method is used to convert the original data into standardized data with a mean of 0 and a standard deviation of 1 by calculating the mean and standard deviation of each numerical feature, thereby eliminating the impact of dimensions between different features.

[0049] Encoding: For categorical data, one-hot encoding is used. The different values ​​of each categorical variable are converted into corresponding binary vectors so that the model can effectively process these categorical information.

[0050] Data partitioning: The preprocessed data is divided into training set, validation set and test set in the ratio of 70%, 20% and 10% respectively. During the partitioning process, random sampling is adopted to ensure that the data distribution of each subset is representative and avoid the influence of data bias on model training and evaluation.

[0051] Constructing a financial risk prediction model based on gradient boosting decision tree

[0052] Model initialization

[0053] Before starting training, the model needs to be initialized. There is a training data set containing multiple borrower samples, each of which has corresponding feature information and real risk labels (such as whether it is a default). The initialized model will take the average of the risk labels of all samples to get an initial prediction value. This initial prediction value can be regarded as a preliminary estimate of the risk situation of all samples by the model without any additional information.

[0054] Iteratively train decision trees

[0055] Next, we enter the iterative training phase. We set a maximum number of iterations, such as 100, and then gradually build and train the decision tree starting from the first iteration. The specific steps are as follows:

[0056] Calculate residuals: In each iteration, the first thing to do is to calculate the difference between the previous model prediction and the true risk label, that is, the residual. The residual reflects the prediction error of the previous model on each sample. The new decision tree can learn these errors, thereby improving the prediction ability of the overall model. For example, if the previous model predicts that a customer will not default, but the customer actually defaults, then the residual of this sample will be relatively large, and the new decision tree will focus on such samples with incorrect predictions.

[0057] Train a new decision tree: Use the calculated residual as the new target value to train a new decision tree. The decision tree construction process is a recursive partitioning process. Starting from the root node, appropriate features and thresholds are selected based on the characteristic information of the sample to divide the sample into different subsets to form different branches. Each branch continues to divide until a certain stopping condition is met, such as the number of samples in the node is less than a certain threshold, or the residual difference of the samples in the node is small enough. Ultimately, the decision tree will form multiple leaf nodes, each of which corresponds to a specific prediction value. This prediction value is obtained by statistically analyzing the residuals of all samples in the leaf node, usually the mean or median of these residuals.

[0058] After training a new decision tree in each round, the prediction results of the new tree need to be integrated into the previous model to update the model. The specific method is to multiply the prediction results of the new decision tree by a learning rate and then add it to the model of the previous round. The learning rate is a positive number less than 1, which plays a role in controlling the contribution of the decision tree to the model update in each round of iteration. A smaller learning rate can make the model training more stable and avoid excessive fluctuations in the model during training, but it may increase the training time; a larger learning rate can speed up the convergence of the model, but it may cause the model to skip the optimal solution and overfit. Through continuous iterative training and model updates, the model will gradually learn the complex patterns and laws in the data and improve the accuracy of predicting financial risks.

[0059] When training a financial risk prediction model, there are multiple hyperparameters that need to be adjusted, which will have a significant impact on the performance of the model. Hyperparameters include the maximum depth of the tree, the minimum number of sample splits, the minimum number of sample leaves, the learning rate, and the maximum number of iterations. For example, the maximum depth of the tree limits the growth depth of the decision tree. A tree that is too deep may cause the model to overfit, while a tree that is too shallow may not be able to learn the complex features of the data; the learning rate determines the step size of the model update in each round of iteration and needs to be adjusted according to the specific situation.

[0060] Use grid search to find the optimal hyperparameter combination. Grid search will traverse the pre-defined hyperparameter value combinations, train and verify the model for each combination, and then select the combination with the best performance on the verification set as the final hyperparameter.

[0061] After multiple iterations of training and hyperparameter tuning, the final financial risk prediction model was obtained.

[0062] Risk prediction and early warning

[0063] Real-time risk prediction

[0064] Build a real-time data collection system and maintain real-time connection with the bank's internal business system, credit reporting agencies and other data sources. When new borrower information is generated, obtain it immediately and clean, standardize and encode it according to the previous preprocessing process. Input the processed real-time data into the trained financial risk prediction model, and convert the model output into the predicted probability of loan default through the logic function:

[0065] The default probability is compared with the pre-set risk threshold. If the default probability is greater than the pre-set risk threshold, it is judged as a high-risk loan; if the default probability is less than or equal to the pre-set risk threshold, it is judged as a low-risk loan.

[0066] Risk Warning

[0067] When the model predicts a high-risk loan, an early warning message is sent to the relevant credit management personnel through the bank's internal risk management system. The early warning message contains detailed information about the risk type (such as default risk), risk level (high, medium, low), the possibility of risk occurrence (expressed as default probability), and the potential impact that the loan may have on the bank, so that the credit management personnel can take appropriate risk control measures in a timely manner.

[0068] The above description is only a preferred specific implementation manner of the present invention, but the protection scope of the present invention is not limited thereto. Any technician familiar with the technical field can make equivalent replacements or changes according to the technical scheme and inventive concept of the present invention within the technical scope disclosed by the present invention, which should be covered by the protection scope of the present invention.

Claims

1. A financial risk prediction method based on artificial intelligence, characterized in that: The method comprises: Collect and pre-process the multi-source data of borrowers, including basic customer information, financial data, credit records, loan information, and external macro data; Constructing a financial risk prediction model based on a gradient boosting decision tree, wherein the financial risk prediction model is formed by iterative integration of multiple decision trees, and each decision tree fits the residual between the prediction result of the previous round of model and the true value; Training the financial risk prediction model through the preprocessed data; Bank financial risks are predicted through trained financial risk prediction models, and risk warnings are issued based on the prediction results.

2. The financial risk prediction method based on artificial intelligence according to claim 1, characterized in that: The collecting of multi-source data of borrowers and preprocessing specifically include: Collect data from internal bank systems, credit reporting agencies, and public data sources; Clean the collected data, including processing missing values, outliers and duplicate data; Standardize and encode the cleaned data to eliminate the dimension effects between features and convert categorical variables into numerical types; The preprocessed data is divided into training set, validation set and test set according to a certain ratio.

3. The financial risk prediction method based on artificial intelligence according to claim 1, characterized in that: When processing missing values, for numerical data, the mean is used for filling according to the missing ratio, and for categorical data, the mode is used for filling; the processing of outliers adopts statistical methods; the standardization adopts the Z-score standardization method; the encoding adopts the one-hot encoding method.

4. The financial risk prediction method based on artificial intelligence according to claim 1, characterized in that: The construction of a financial risk prediction model based on a gradient boosting decision tree specifically includes: The initial model prediction value is the mean of the target values ​​of all training samples; In each iteration, the residual between the model prediction value and the true value in the previous round is calculated; The residual is used as the target value to train a new decision tree. The decision tree contains a root node, internal nodes, and leaf nodes. The internal nodes divide the samples based on the feature threshold. Multiply the prediction results of the newly trained decision tree by the learning rate and add them to the prediction results of the previous round of model to update the financial risk prediction model.

5. The financial risk prediction method based on artificial intelligence according to claim 1, characterized in that: When training a decision tree, hyperparameters need to be determined, including the maximum depth of the tree, the minimum number of sample splits, and the minimum number of sample leaves.

6. The financial risk prediction method based on artificial intelligence according to claim 1 is characterized in that: The method of predicting bank financial risks through the trained prediction model specifically includes: Obtain the latest information of borrowers in real time and perform the same preprocessing as training data; Input the preprocessed real-time data into the trained financial risk prediction model and output the loan default probability; Compare the probability of default with the set risk threshold to determine the loan risk level.

7. The method for predicting bank loan risks based on artificial intelligence according to claim 1, characterized in that: The risk warning according to the prediction results specifically includes: When the prediction result is a high-risk loan, an early warning signal containing risk type and level information will be sent to relevant credit managers through the bank's internal system.

8. The method for predicting bank loan risks based on artificial intelligence according to claim 1, characterized in that: The prediction results are used for financial business decision-making, including but not limited to adjusting loan amounts, interest rate pricing, and loan term setting; stratifying loan customers according to risk levels and formulating differentiated risk management strategies for customers of different risk levels.