Financial credit risk analysis early warning method and system based on artificial intelligence

By standardizing the financial indicator data of the XGBoost model and dynamic penalty factor optimization, the accuracy problem of the traditional XGBoost model in evaluating corporate credit risks is solved, and a more efficient credit risk assessment is achieved.

CN120278837APending Publication Date: 2025-07-08SHANDONG CHINA SOFT FINTECH INFORMATION TECH CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510465504.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-15
Publication Date
2025-07-08

AI Technical Summary

Technical Problem

When evaluating corporate credit risks, the fixed punishment factor cannot adapt to data changes in different industries and the same industry, resulting in a decrease in evaluation accuracy.

Method used

By standardizing the financial indicator data of the target enterprise, dividing the training set and the hyperparameter tuning optimization set, using the training set to obtain the importance of the initial punishment factor and financial indicators, combining the data distribution and correlation, dynamically optimize the punishment factor of each decision tree, forming an adaptive punishment factor, and building a tuned XGBoost model.

Benefits of technology

The XGBoost model's assessment accuracy of corporate credit risk is improved, overfitting and underfitting are prevented, and the model's prediction ability in different data environments is enhanced.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120278837A_ABST
    Figure CN120278837A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of data processing, in particular to a financial credit risk analysis early warning method and system based on artificial intelligence, and the method comprises the steps: dividing the historical data of each financial index of a target enterprise into a hyper-parameter adjustment optimal set and a training set; training an XGBoost model by using the training set to obtain the importance degree of each financial index, and obtaining the dynamic relative importance degree according to the importance degree of each financial index in the feature subset corresponding to any decision tree; obtaining a data distribution confusion degree according to the distribution condition of the historical data in the hyper-parameter tuning set, and obtaining a complexity degree according to the correlation between the financial indexes and the data distribution confusion degree; the self-adaptive penalty factor is obtained by combining the dynamic relative importance degree and the complexity degree, and the XGBoost model is optimized according to the self-adaptive penalty factor of each decision tree, so that the credit risk of the target enterprise is pre-warned, and the accuracy of pre-warning the credit risk of the enterprise is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of data processing, and particularly to a method and system for financial credit risk analysis and early warning based on artificial intelligence. Background Art

[0002] Currently, with the expansion and business development of enterprises, in order to maintain a good cash flow to meet daily operating needs and optimize the capital structure, they need to borrow from financial institutions. And financial institutions need to analyze the credit risks of enterprises through multi-dimensional corporate financial data, evaluate their repayment capabilities, and predict possible credit risks to reduce their own losses. Financial credit risk refers to the loss risk that financial institutions may face when providing loans or credit services to customers due to the default of the borrower or other factors. The analysis of financial credit risk mainly focuses on analyzing whether the borrower has credit risks and fraud risks. Among them, the most important is to analyze the credit risk of the borrower to directly evaluate whether the borrower has sufficient repayment ability.

[0003] Under the traditional method, the XGBoost model is used to evaluate the credit risk of enterprises. The XGBoost model is an efficient ensemble learning method that gradually optimizes the model's prediction ability by constructing multiple decision trees, automatically selects important features, has strong robustness and high prediction accuracy, and can effectively identify risk points when facing high-dimensional and massive complex data, helping financial institutions detect the credit risks of enterprises to dynamically adjust loan strategies and reduce default risks.

[0004] The core idea of the XGBoost model is to train a weak estimator (usually a decision tree model), and then adjust the training process of the subsequent decision tree model according to the wrong predictions of the decision tree model, so as to gradually reduce the prediction error. That is, after the first evaluation of the decision tree model, the samples predicted wrongly by the decision tree model are fed back to the data set, and then the next decision tree model is constructed to correct the error of the previous decision tree model. This is repeated iteratively to obtain the final predicted value. The XGBoost model adopts the Boosting framework, which consists of a loss function (used to measure the error between the true value and the predicted value. The XGBoost model uses the second-order Taylor expansion approximation to improve the optimization efficiency) and a regularization term (used to control the complexity of the decision tree model and prevent overfitting), ensuring the accuracy of the decision tree model and reducing overfitting.

[0005] Among them, the reasonable setting of the regularization term is crucial for improving the prediction accuracy, stability, and generalization ability of the decision tree model. In the regularization term, a penalty factor is required to control the splitting of leaf nodes in the decision tree model to prevent overfitting, while in the traditional method, a fixed The value is used to control the splitting of leaf nodes in all decision tree models. However, due to industry differences (the financial index characteristics vary greatly among different industries, such as the completely different profit models of manufacturing and Internet enterprises, showing significant differences in financial indicators) and the cold and warm periods within the same industry (each industry has its own cold and warm periods, and with the alternation of the off-season and peak season, the financial data fluctuates greatly at different stages), a fixed value cannot adapt to this large-difference data change pattern. For different decision tree models, there are problems of being relatively too large or too small, resulting in the decision tree model being too simple to capture complex relationships, or the decision tree model being too complex and overfitting, thereby reducing the accuracy of evaluating the credit risk of enterprises.

[0006] Therefore, how to adaptively obtain the penalty factor to improve the accuracy of using the XGBoost model to evaluate the credit risk of enterprises has become an urgent problem to be solved. Summary of the Invention

[0007] In view of this, the embodiments of the present invention provide a method and system for financial credit risk analysis and early warning based on artificial intelligence to solve the problem of how to adaptively obtain the penalty factor to improve the accuracy of using the XGBoost model to evaluate the credit risk of enterprises.

[0008] In a first aspect, the embodiments of the present invention provide a method for financial credit risk analysis and early warning based on artificial intelligence. The method includes the following steps: Standardize the data of each financial indicator of the target enterprise at different time nodes to obtain the corresponding historical data, and divide the historical data of all financial indicators into a hyperparameter tuning set and a training set according to a preset ratio; Use the training set to train the XGBoost model to obtain the initial penalty factor of at least one decision tree and the importance degree of each financial indicator. For any decision tree, form a feature subset with all the financial indicators corresponding to the any decision tree, and obtain the dynamic relative importance degree of the feature subset according to the importance degree of each financial indicator in the feature subset; Obtain the data distribution chaos degree according to the data distribution and data fluctuation of the historical data of each financial indicator in the hyperparameter tuning set, and obtain the complexity degree of the feature subset according to the correlation between every two financial indicators in the feature subset and the data distribution chaos degree; Optimize the initial penalty factor of any decision tree according to the dynamic relative importance and the complexity to obtain an adaptive penalty factor. Replace the initial penalty factor of each decision tree with the corresponding adaptive penalty factor to obtain an optimized XGBoost model. Warn of the credit risk of the target enterprise according to the optimized XGBoost model.

[0009] Preferably, the method for obtaining the importance of each financial indicator includes: For any financial indicator, during the process of training the XGBoost model using the training set, obtain the information gain after each decision tree split using the any financial indicator, and obtain the average information gain of the any financial indicator. Take the average information gain of the any financial indicator as the independent variable of the hyperbolic tangent function to obtain the information gain eigenvalue of the any financial indicator. Obtain the importance score of the any financial indicator, and calculate the average value of the information gain eigenvalue and the importance score of the any financial indicator to obtain the importance of the any financial indicator.

[0010] Preferably, obtaining the dynamic relative importance of the feature subset according to the importance of each financial indicator in the feature subset includes: Calculate the average value of the importance of each financial indicator in the feature subset to obtain the average importance of the feature subset. Combine all financial indicators into a total feature set, calculate the average value of the importance of each financial indicator in the total feature set to obtain the average importance of the total feature set. Take the difference between the average importance of the feature subset and the average importance of the total feature set as the independent variable of the activation function to obtain the dynamic relative importance of the feature subset.

[0011] Preferably, obtaining the degree of data distribution chaos according to the data distribution and data fluctuation of the historical data of each financial indicator in the hyperparameter tuning set includes: In the hyperparameter tuning set, form a data subset from the historical data of each financial indicator in the feature subset; Take the difference between the maximum value and the minimum value in the data subset as the independent variable of the hyperbolic tangent function to obtain the fluctuation range index of the data subset; Calculate the deviation from the mean of each data in the data subset, and take the cumulative result of all deviations from the mean as the independent variable of the hyperbolic tangent function to obtain the degree of fluctuation of the data subset; Calculate the sum of the fluctuation range index and the degree of fluctuation of the data subset to obtain the degree of data distribution chaos.

[0012] Preferably, obtaining the complexity of the feature subset according to the correlation between every two financial indicators in the feature subset and the degree of data distribution chaos includes: In the hyperparameter tuning set, respectively obtain the historical data of each financial indicator in the feature subset to obtain the data sequence corresponding to each financial indicator in the feature subset, calculate the absolute value of the Spearman coefficient between the data sequences corresponding to every two financial indicators in the feature subset, and calculate the average value of all the absolute values of the Spearman coefficients to obtain the average correlation degree of the feature subset; Calculate the difference between the preset correlation threshold and the average correlation degree, and use the sum of the difference and the degree of data distribution chaos as the independent variable of the hyperbolic tangent function to obtain the complexity of the feature subset.

[0013] Preferably, optimizing the initial penalty factor of any decision tree according to the dynamic relative importance and the complexity to obtain an adaptive penalty factor includes: Calculate the average value between the dynamic relative importance and the complexity to obtain the comprehensive eigenvalue of any decision tree; If the comprehensive eigenvalue is greater than the preset feature threshold, then subtract the comprehensive eigenvalue from the constant 1 to obtain an adjustment coefficient, and calculate the product of the initial penalty factor and the adjustment coefficient to obtain the adaptive penalty factor of any decision tree; If the comprehensive eigenvalue is less than or equal to the preset feature threshold, then calculate the sum of the constant 1 and the comprehensive eigenvalue to obtain an adjustment coefficient, and calculate the product of the initial penalty factor and the adjustment coefficient to obtain the adaptive penalty factor of any decision tree.

[0014] In a second aspect, an embodiment of the present invention further provides an artificial intelligence-based financial credit risk analysis and early warning system, including a memory, a processor, and a computer program stored in the memory and running on the processor. When the processor executes the computer program, it implements an artificial intelligence-based financial credit risk analysis and early warning method as described in the first aspect.

[0015] The beneficial effects of the embodiments of the present invention compared with the prior art are: The present invention standardizes the data of various financial indicators of the target enterprise at different time nodes to obtain corresponding historical data, and divides the historical data of all financial indicators into a hyperparameter tuning set and a training set according to a preset ratio; trains the XGBoost model using the training set to obtain the initial penalty factor of at least one decision tree and the importance degree of each financial indicator. For any decision tree, all the financial indicators corresponding to the any decision tree are used to form a feature subset, and according to the importance degree of each financial indicator in the feature subset, the dynamic relative importance degree of the feature subset is obtained; according to the data distribution and data fluctuation of the historical data of each financial indicator in the hyperparameter tuning set, the degree of data distribution chaos is obtained, and according to the correlation between every two financial indicators in the feature subset and the degree of data distribution chaos, the complexity of the feature subset is obtained; according to the dynamic relative importance degree and the complexity, the initial penalty factor of the any decision tree is optimized to obtain an adaptive penalty factor, and the initial penalty factor of each decision tree is replaced with the corresponding adaptive penalty factor to obtain a tuned XGBoost model, and the credit risk of the target enterprise is warned according to the tuned XGBoost model. Among them, by analyzing the importance degree of each feature (each financial indicator) used to construct the decision tree and the data distribution in the data subset (the degree of data distribution chaos), and combining the correlation between every two features, the adaptive penalty factor of each decision tree in the XGBoost model is comprehensively obtained, and a tuned XGBoost model is obtained, so that the tuned XGBoost model can make an optimal decision according to different data environments, capture more effective information, prevent overfitting and underfitting, and effectively improve the accuracy of financial institutions in warning the credit risk of enterprises. Description of the Drawings

[0016] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the drawings required for use in the embodiments or the description of the prior art. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0017] Figure 1 It is a flowchart of a method for analyzing and warning financial credit risk based on artificial intelligence provided in Embodiment 1 of the present invention. Detailed Embodiments

[0018] The following details the embodiments of the present disclosure. The examples of the embodiments are shown in the drawings, where the same or similar reference numerals represent the same or similar elements or elements with the same or similar functions throughout. The embodiments described below with reference to the drawings are exemplary and are intended to explain the present disclosure and should not be construed as a limitation of the present disclosure.

[0019] It should be noted that the terms "first", "second", etc. in the specification of the present disclosure and the above-mentioned drawings are used to distinguish similar objects and do not necessarily have to be used to describe a specific order or sequence. It should be understood that the data used in this way can be interchanged under appropriate circumstances so that the embodiments of the present disclosure described here can be implemented in an order different from those illustrated or described here. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the present disclosure. On the contrary, they are only examples of devices and methods consistent with some aspects of the present disclosure.

[0020] In order to illustrate the technical solution of the present invention, specific embodiments will be used for illustration below.

[0021] See Figure 1 , which is a method flowchart of a financial credit risk analysis and early warning method based on artificial intelligence provided in the first embodiment of the present invention. As Figure 1 shown, the method may include: Step S101, standardize the data of each financial indicator of the target enterprise at different time nodes to obtain the corresponding historical data, and divide the historical data of all financial indicators into a hyperparameter tuning set and a training set according to a preset ratio.

[0022] When a financial institution provides loans or credit services to an enterprise, it needs to conduct a credit risk analysis of the enterprise based on the financial data of the enterprise to directly evaluate whether the enterprise has sufficient repayment ability. Among them, there are multiple indicators in the financial data of the enterprise, usually including profitability indicators (including: net profit, net interest rate, etc.), solvency indicators (asset-liability ratio, current ratio, etc.), operation ability indicators, cash flow indicators, etc. In the embodiments of the present invention, the enterprise to be analyzed is used as the target enterprise. During the process of the financial institution's credit risk assessment of the target enterprise, the financial institution screens the financial indicators that need to be analyzed and the historical time range that needs to be analyzed. Assume that it is necessary to analyze the data of the net profit, net interest rate, asset-liability ratio, revenue growth rate, current ratio, operation ability indicators, and cash flow indicators of the target enterprise on a monthly basis within three years. There is no limit here, and the implementer can set according to the specific scenario, obtain the data of each financial indicator on a monthly basis within three years, and standardize all the data to eliminate the dimension difference for subsequent analysis, obtain the historical data corresponding to each financial indicator. At the same time, the financial institution needs to analyze all the financial indicators of the target enterprise to form a total feature set, and the historical data corresponding to each financial indicator forms a total data set. Among them, the standardization process is a prior art and will not be elaborated here.

[0023] In the traditional method, the XGBoost model constructs decision trees based on the historical data of various financial indicators of the target enterprise to evaluate the credit risk of the target enterprise. The core idea of the XGBoost model is to train a weak estimator (usually a decision tree), and then adjust the subsequent decision trees according to the wrong predictions of the decision tree. That is, after the first evaluation of the decision tree, the samples with wrong predictions of the decision tree are fed back to the data set, and then the next decision tree is constructed to correct the error of the previous decision tree. This iterative process is repeated to obtain the final predicted value, thereby gradually reducing the prediction error.

[0024] The objective function of the XGBoost model consists of a loss function (used to measure the error between the true value and the predicted value) and a regularization term (used to control the complexity of the decision tree and prevent overfitting). Among them, the reasonable setting of the regularization term is crucial for improving the prediction accuracy, stability and generalization ability of the decision tree model. A penalty factor is required in the regularization term to control the splitting of leaf nodes in the decision tree and prevent overfitting: the larger the penalty factor, the simpler the decision tree, the shallower the tree depth, the stronger the anti-overfitting and generalization ability; on the contrary, the smaller the penalty factor, the more complex the decision tree, the deeper the tree depth, the easier it is to overfit, and it can capture more details. In the traditional method, the same value is used to control the splitting of leaf nodes. However, due to the cold and warm periods in the same industry (each industry has its own cold and warm periods, and with the alternation of the off-season and peak season, there are large fluctuations in financial data at different stages), the fixed value cannot adapt to this large difference in data change patterns, and there is a problem of being relatively too large or too small for different decision trees, resulting in the decision tree being too simple to capture complex relationships, or the decision tree being too complex and overfitting, thereby reducing the accuracy of evaluating the credit risk of the target enterprise.

[0025] Therefore, in the embodiments of the present invention, the historical data of the net profit, net profit rate, asset-liability ratio, revenue growth rate, current ratio, operation ability indicators, and cash flow indicators of the target enterprise in each month within three years are divided into a hyperparameter tuning set and a training set according to a ratio of 8:2. There is no limitation here, and the implementer can set the financial indicators and the division ratio according to the specific scenario. Among them, the training set is used to train the XGBoost model to obtain various parameters in the XGBoost model, such as the number of decision trees, learning rate, tree depth, data sampling ratio, feature sampling ratio, and penalty factor. The hyperparameter tuning set is used to optimize the penalty factor in the XGBoost model.

[0026] Step S102: Use the training set to train the XGBoost model to obtain the initial penalty factors of at least one decision tree and the importance degrees of various financial indicators. For any decision tree, form a feature subset with all the financial indicators corresponding to the said any decision tree, and obtain the dynamic relative importance degree of the feature subset according to the importance degrees of the financial indicators in the feature subset.

[0027] It is known that before training the XGBoost model, it is necessary to first determine the hyperparameters in the XGBoost model. The selection of hyperparameters has a great impact on the performance of the XGBoost model. For example, the learning rate is used to control the step size of updating the weights in each iteration. If the learning rate is too small, the model convergence speed is slow; if the learning rate is too large, the model may oscillate near the optimal solution and cannot converge. The depth of the tree determines the complexity of the decision tree. If the depth is too deep, it is easy to cause overfitting, and if the depth is too shallow, it will cause underfitting. Therefore, in the embodiments of the present invention, before using the training set to train the XGBoost model, first use methods such as grid search and random search to find the optimal combination of hyperparameters to obtain the initial values of each hyperparameter in the XGBoost model. Preferably, in the embodiments of the present invention, the number of decision trees is set to 500, the learning rate is 0.05, the tree depth is 8, the data sampling ratio is 0.7, the feature sampling ratio is 0.6, and the penalty factor is 0.5. Among them, grid search and random search are existing technologies and will not be elaborated here. Then use the training set to train the XGBoost model. During the training process, use the mean squared error (MSE) loss function to calculate the prediction loss, and reverse correct the hyperparameters in the XGBoost model according to the gradient descent method until the prediction loss converges to obtain the trained XGBoost model.

[0028] At this time, the trained XGBoost model includes at least one decision tree, and the penalty factors of all decision trees are the same. Since each industry has its own warm and cold periods, that is, the data of various financial indicators at different time nodes fluctuate greatly. Therefore, when the penalty factors of all decision trees are the same, some decision trees will be too complex or too simple, resulting in a decrease in the accuracy of evaluating the credit risk of the target enterprise. Therefore, in the embodiments of the present invention, the penalty factor of each decision tree in the XGBoost model trained with the training set is used as the initial penalty factor, and the hyperparameter tuning set is used to tune the initial penalty factor of each decision tree to obtain the adaptive penalty factor of each decision tree, thereby improving the accuracy of evaluating the credit risk of the target enterprise.

[0029] The XGBoost model obtains the final predicted value by iteratively constructing decision trees and accumulating the prediction results of each decision tree. Before constructing each decision tree, a part is randomly sampled with replacement from the total feature set as the feature subset for constructing the decision tree according to the feature sampling ratio (the feature sampling ratio in the embodiments of the present invention is 0.6). Then, the splitting of the decision tree is controlled according to the characteristics of the data corresponding to each feature in the feature subset. The core principle of the decision tree is to select the feature with the best splitting effect, that is, the feature with the largest information gain, for each split. At the same time, financial institutions have different importance evaluations for different financial indicators based on business experience. For example, financial institutions will conduct importance evaluations on various financial indicators according to the industry characteristics of the target enterprise. If the target enterprise is a technology enterprise, the financial institution will score the importance of each financial indicator of the target enterprise according to the principle of giving priority to growth and putting profitability second. That is, the importance scores of the financial indicators of the target enterprise by the financial institution may be: the importance score of the revenue growth rate is 0.4, the importance score of the cash flow indicator is 0.25, and the importance score of the net profit is 0.15. The higher the importance score of a financial indicator, the more attention the financial institution pays to it, that is, the greater the contribution of this financial indicator to the XGBoost model. Therefore, in the embodiments of the present invention, during the training of the XGBoost model according to the training set, the information gain of each financial indicator and the importance score of each financial indicator by the financial institution (the value range is ) are comprehensively used to obtain the importance degree of each financial indicator, which is used to reflect the splitting effect of each financial indicator: the greater the information gain, the better the splitting effect. The specific method for obtaining the importance degree of each financial indicator is as follows: For any financial indicator, during the training of the XGBoost model using the training set, the information gain after each decision tree split using the any financial indicator is obtained, and the average information gain of the any financial indicator is obtained; Taking the average information gain of the any financial indicator as the independent variable of the hyperbolic tangent function, the information gain eigenvalue of the any financial indicator is obtained, the importance score of the any financial indicator is obtained, and the average value of the information gain eigenvalue and the importance score of the any financial indicator is calculated to obtain the importance degree of the any financial indicator.

[0030] In an embodiment, taking the nth financial indicator as an example, the calculation formula for the importance degree of the nth financial indicator is:

[0031] Wherein, represents the importance degree of the nth financial indicator, represents the average information gain of the nth financial indicator, represents the importance score of the nth financial indicator, represents the hyperbolic tangent function, which is used to limit the output result to .

[0032] It should be noted that the higher the average information gain of the nth financial indicator, the higher the importance score, the greater the contribution of the nth financial indicator to the XGBoost model, and the greater the degree of importance.

[0033] So far, the parameters in the XGBoost model and the importance levels of various financial indicators have been obtained based on the training set. Further, the initial penalty factor in the XGBoost model is tuned using the hyperparameter tuning set. Taking the t-th decision tree obtained by training the XGBoost model with the training set as an example, all the financial indicators corresponding to the t-th decision tree form a feature subset. According to the importance levels of the financial indicators in the feature subset, the dynamic relative importance level of the feature subset is obtained to evaluate the contribution degree of the financial indicators in the feature subset to the XGBoost model: the greater the contribution degree, the smaller the penalty factor should be, so that the decision tree has more splitting times, can capture more effective information, and the final prediction result is more accurate, and the credit risk assessment of the target enterprise is more accurate. Among them, the specific method for obtaining the dynamic relative importance level of the feature subset is as follows: Calculate the average value of the importance levels of the financial indicators in the feature subset to obtain the average importance level of the feature subset. Calculate the average value of the importance levels of the financial indicators in the total feature set (the set composed of all the financial indicators that the financial institution needs to analyze for the target enterprise) to obtain the average importance level of the total feature set. Take the difference between the average importance level of the feature subset and the average importance level of the total feature set as the independent variable of the activation function to obtain the dynamic relative importance level of the feature subset.

[0034] In an implementation manner, the calculation formula for the dynamic relative importance level of the feature subset corresponding to the t-th decision tree is:

[0035] Among them, represents the dynamic relative importance level of the feature subset corresponding to the t-th decision tree, represents the number of all financial indicators in the feature subset corresponding to the t-th decision tree, N represents the number of all financial indicators in the total feature set, represents the importance level of the th financial indicator in the feature subset corresponding to the t-th decision tree, represents the importance level of the nth financial indicator in the total feature set, represents the activation function, which is used to limit the output result to .

[0036] It should be noted that is the average importance of the feature subset corresponding to the t-th decision tree, is the average importance of the total feature set. When the average importance of the feature subset far exceeds that of the total feature set, that is when it is a positive number, the greater the difference in the average importance between the feature subset and the total feature set, the greater it is, the greater the contribution of the feature subset corresponding to the t-th decision tree to the XGBoost model. At this time, the penalty factor needs to be reduced to allow the decision tree to split more fully on the financial indicators in the feature subset, preventing underfitting and thus capturing more effective information and improving the accuracy of the XGBoost model in evaluating the credit risk of the target enterprise; when the average importance of the feature subset is less than that of the total feature set, that is when it is a negative number, the greater the difference in the average importance between the feature subset and the total feature set, the fewer important features are contained in the feature subset corresponding to the t-th decision tree, and the activation function is closer to 0, and thus the smaller it is, the less the contribution of the feature subset corresponding to the t-th decision tree to the XGBoost model. At this time, the penalty factor needs to be increased to prevent the decision tree from splitting too much on the financial indicators in the feature subset and causing overfitting, and to improve the accuracy of the XGBoost model in evaluating the credit risk of the target enterprise.

[0037] Step S103: Obtain the degree of data distribution chaos according to the data distribution and data fluctuation of the historical data of each financial indicator in the hyperparameter tuning set, and obtain the complexity of the feature subset according to the correlation between every two financial indicators in the feature subset and the degree of data distribution chaos.

[0038] Since the data distribution will affect the effectiveness and necessity of the splitting of each decision tree in the XGBoost model, if the difference between the data is large, more splitting times are required to effectively capture the complex patterns in the data, that is, the penalty factor should be appropriately reduced. On the contrary, if the difference between the data is small and the data distribution is stable, the penalty factor should be appropriately increased to prevent overfitting and improve the accuracy of evaluating the credit risk of the target enterprise. Therefore, in the embodiment of the present invention, the historical data of each financial indicator in the feature subset corresponding to the t-th decision tree in the hyperparameter tuning set is formed into a data subset, and then according to the data distribution and data fluctuation in the data subset, the degree of data distribution chaos of the data subset corresponding to the t-th decision tree is obtained to adjust the size of the penalty factor. The specific method for obtaining the degree of data distribution chaos of the data subset corresponding to the t-th decision tree is as follows: Take the difference between the maximum value and the minimum value in the data subset as the independent variable of the hyperbolic tangent function to obtain the fluctuation range index of the data subset; Calculate the deviation from the mean of each data in the data subset, and use the cumulative result of all deviations from the mean as the independent variable of the hyperbolic tangent function to obtain the degree of fluctuation of the data subset; Calculate the sum between the fluctuation range index and the degree of fluctuation of the data subset to obtain the degree of data distribution chaos.

[0039] In one embodiment, the calculation formula for the degree of data distribution chaos of the data subset corresponding to the t-th decision tree is:

[0040] where, represents the degree of data distribution chaos of the data subset corresponding to the t-th decision tree, represents the maximum value in the data subset corresponding to the t-th decision tree, represents the minimum value in the data subset corresponding to the t-th decision tree, m represents the number of all elements in the data subset, represents the i-th data in the data subset corresponding to the t-th decision tree, represents the hyperbolic tangent function, which is used to limit the output result within , represents the absolute value symbol.

[0041] It should be noted that is the fluctuation range index of the data subset, The larger it is, the larger the fluctuation range of the data in the data subset, and thus the larger it is, the more chaotic the data distribution in the data subset; is the deviation from the mean. The larger the deviation from the mean, the greater the data fluctuation and the more discrete the data distribution, and thus the larger it is, the more chaotic the data distribution in the data subset.

[0042] Also considering that there may be a strong correlation between any two financial indicators in the feature subset corresponding to the t-th decision tree. If there is a strong correlation between any two financial indicators, it means that these two financial indicators may provide duplicate information. When the decision tree uses these two financial indicators for splitting, overfitting may occur, which may lead to a decrease in the accuracy of evaluating the credit risk of the target enterprise. Therefore, in the embodiments of the present invention, by combining the correlation between every two financial indicators in the feature subset corresponding to the t-th decision tree and the degree of data distribution chaos, the complexity of the feature subset corresponding to the t-th decision tree is obtained, which is used to adjust the size of the penalty factor. The greater the complexity, the smaller the penalty factor. Conversely, the smaller the complexity, the greater the penalty factor. To ensure that in the case of chaotic data distribution and high correlation, the decision tree will not split too much and cause overfitting, nor will it split too little and lead to inaccurate prediction results due to stable data distribution and low correlation. The specific method for obtaining the complexity of the feature subset corresponding to the t-th decision tree is as follows: In the hyperparameter tuning set, historical data of each financial indicator in the feature subset are respectively obtained to obtain the data sequences corresponding to each financial indicator in the feature subset. The absolute value of the Spearman coefficient between the data sequences corresponding to every two financial indicators in the feature subset is calculated, and the average value of all the absolute values of the Spearman coefficients is calculated to obtain the average correlation degree of the feature subset. The Spearman coefficient is prior art and will not be elaborated here; Calculate the difference between the preset correlation threshold and the average correlation degree, and take the sum of the difference and the degree of data distribution chaos as the independent variable of the hyperbolic tangent function to obtain the complexity of the feature subset.

[0043] In one embodiment, the calculation formula for the complexity of the feature subset corresponding to the t-th decision tree is:

[0044] wherein, represents the complexity of the feature subset corresponding to the t-th decision tree, represents the degree of data distribution chaos of the data subset corresponding to the t-th decision tree, represents the preset correlation threshold, represents the number of all financial indicators in the feature subset corresponding to the t-th decision tree, j represents the data sequence corresponding to the j-th financial indicator in the feature subset corresponding to the t-th decision tree, k represents the data sequence corresponding to the k-th financial indicator in the feature subset corresponding to the t-th decision tree, represents the Spearman coefficient between the data sequence corresponding to the j-th financial indicator and the data sequence corresponding to the k-th financial indicator, and P represents the number of combinations of all pairs of financial indicators in the feature subset corresponding to the t-th decision tree, that is , represents the hyperbolic tangent function, which is used to limit the output result to , represents the absolute value symbol.

[0045] It should be noted that the larger the [[value]], the more chaotic the data distribution in the data subset, and thus the larger the [[value]], the more splitting times are required to effectively capture the complex patterns in the data subset, that is, the penalty factor should be appropriately reduced; is the average correlation degree. Among them, since the value range of the Spearman coefficient is in , the closer the Spearman coefficient is to , it indicates that the monotonicity between the two groups of data is stronger, and there is more correlation between the two groups of data. Therefore, in the embodiments of the present invention, the average correlation degree of the feature subset corresponding to the t-th decision tree is calculated through the absolute value of the Spearman coefficient between the data sequences corresponding to every two financial indicators in the feature subset corresponding to the t-th decision tree; represents the preset correlation threshold. According to experimental statistics, [[a value]] is set , which is not limited here. The implementer can set it according to the specific scenario. When the average correlation degree is less than 0.4, it indicates that the correlation between every two financial indicators in the feature subset corresponding to the t-th decision tree is relatively low, and the possibility of providing duplicate information is relatively small. At this time, the penalty factor should be appropriately increased to effectively capture the complex patterns in the data subset. Therefore, when the average correlation degree is less than 0.4, the smaller the average correlation degree, the larger the [[value]], and thus the larger the [[value]], the more complex the data distribution of each financial indicator in the feature subset corresponding to the t-th decision tree, and the splitting times of the decision tree need to be increased to prevent underfitting. When the average correlation degree is greater than or equal to 0.4, it indicates that the correlation between every two financial indicators in the feature subset corresponding to the t-th decision tree is relatively high, and the possibility of providing duplicate information is relatively large. At this time, the penalty factor should be increased to prevent overfitting. Therefore, when the average correlation degree is greater than or equal to 0.4, the larger the average correlation degree, the smaller the [[value]], and thus the smaller the [[value]], the more stable the data distribution of each financial indicator in the feature subset corresponding to the t-th decision tree, and the splitting times of the decision tree need to be reduced to prevent overfitting.

[0046] Step S104, optimize the initial penalty factor of any one of the decision trees according to the dynamic relative importance degree and the complexity degree to obtain an adaptive penalty factor, replace the initial penalty factor of each decision tree with the corresponding adaptive penalty factor to obtain an optimized XGBoost model, and perform early warning on the credit risk of the target enterprise according to the optimized XGBoost model.

[0047] Through steps S102 and S103, the feature subset and data subset corresponding to the t-th decision tree are analyzed to obtain the dynamic relative importance and complexity of the feature subset. Then, based on the dynamic relative importance and complexity, the initial penalty factor of the t-th decision tree obtained by training the XGBoost model with the training set is optimized to obtain the adaptive penalty factor of the t-th decision tree. Specifically: Calculate the average value between the dynamic relative importance and the complexity to obtain the comprehensive feature value of the t-th decision tree; If the comprehensive feature value is greater than the preset feature threshold, subtract the comprehensive feature value from the constant 1 to obtain the adjustment coefficient, and calculate the product between the initial penalty factor and the adjustment coefficient to obtain the adaptive penalty factor of the t-th decision tree; If the comprehensive feature value is less than or equal to the preset feature threshold, calculate the sum between the constant 1 and the comprehensive feature value to obtain the adjustment coefficient, and calculate the product between the initial penalty factor and the adjustment coefficient to obtain the adaptive penalty factor of the t-th decision tree.

[0048] In an embodiment, the calculation formula for the adaptive penalty factor of the t-th decision tree is:

[0049] Wherein, represents the adaptive penalty factor of the t-th decision tree, represents the initial penalty factor of the t-th decision tree, represents the dynamic relative importance of the feature subset corresponding to the t-th decision tree, represents the complexity of the feature subset corresponding to the t-th decision tree, represents the preset feature threshold, and 1 represents a constant.

[0050] It should be noted that according to experimental statistics, set , which is not limited here, and the implementer can set it according to the specific scenario. is the comprehensive feature value of the t-th decision tree. The larger the comprehensive feature value, the higher the importance of the corresponding financial indicators in the t-th decision tree and the lower the correlation. At the same time, the more complex the data distribution of the data subset corresponding to the t-th decision tree in the hyperparameter tuning set, the penalty factor should be appropriately reduced, that is The smaller it is, the more times the t-th decision tree is split to prevent underfitting, capture more effective information, and improve the accuracy of using the XGBoost model to evaluate the credit risk of the target enterprise; conversely, the smaller the comprehensive eigenvalue is, it indicates that the importance of the corresponding financial indicators in the t-th decision tree is relatively low and the correlation is relatively strong. At the same time, the data distribution of the dataset corresponding to the t-th decision tree in the hyperparameter tuning set is more stable, and the penalty factor should be appropriately increased, that is the larger it is, the fewer times the t-th decision tree is split, avoid meaningless splitting, prevent overfitting, and improve the accuracy of using the XGBoost model to evaluate the credit risk of the target enterprise.

[0051] Similarly, according to the data in the hyperparameter tuning set, the adaptive penalty factor of each decision tree is obtained, and the initial penalty factor of each decision tree is replaced with the corresponding adaptive penalty factor to obtain the tuned XGBoost model. The training set and the hyperparameter tuning set are combined into an input set, and the input set is input into the tuned XGBoost model to obtain the prediction result for evaluating the credit risk of the target enterprise. The financial institution sets relevant judgment thresholds. If the prediction result exceeds the relevant judgment thresholds set by the financial institution, a warning is issued for the credit risk of the target enterprise. The financial institution has the right to refuse to provide credit services or additional guarantees and other related operations to the target enterprise. Among them, warning of the credit risk of the target enterprise according to the evaluation result is prior art and will not be elaborated here.

[0052] In summary, the present invention standardizes the data of various financial indicators of the target enterprise at different time nodes to obtain corresponding historical data, and divides the historical data of all financial indicators into a hyperparameter tuning set and a training set according to a preset ratio; uses the training set to train the XGBoost model to obtain the initial penalty factors of at least one decision tree and the importance degrees of various financial indicators. For any decision tree, all the financial indicators corresponding to the any decision tree are used to form a feature subset, and according to the importance degrees of the financial indicators in the feature subset, the dynamic relative importance degree of the feature subset is obtained; according to the data distribution and data fluctuation of the historical data of the financial indicators in the hyperparameter tuning set, the data distribution chaos degree is obtained, and according to the correlation between every two financial indicators in the feature subset and the data distribution chaos degree, the complexity of the feature subset is obtained; according to the dynamic relative importance degree and the complexity, the initial penalty factor of the any decision tree is optimized to obtain an adaptive penalty factor, and the initial penalty factor of each decision tree is replaced with the corresponding adaptive penalty factor to obtain a tuned XGBoost model, and the credit risk of the target enterprise is warned according to the tuned XGBoost model. Among them, by analyzing the importance degree of each feature (each financial indicator) used to construct the decision tree and the data distribution situation (data distribution chaos degree) in the data subset, and combining the correlation between every two features, the adaptive penalty factor of each decision tree in the XGBoost model is comprehensively obtained, and a tuned XGBoost model is obtained, so that the tuned XGBoost model can make optimal decisions according to different data environments, capture more effective information, prevent overfitting and underfitting, and effectively improve the accuracy of financial institutions in warning the credit risk of enterprises.

[0053] Based on the same inventive concept as the above method, an embodiment of the present invention further provides an artificial intelligence-based financial credit risk analysis and warning system, including a memory, a processor, and a computer program stored in the memory and running on the processor. When the processor executes the computer program, the steps of any one of the above artificial intelligence-based financial credit risk analysis and warning methods are implemented.

[0054] The above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present invention, and should all be included in the protection scope of the present invention.

Claims

1. An artificial intelligence-based financial credit risk analysis and early warning method, characterized in that, The financial credit risk analysis and early warning method based on artificial intelligence includes: Standardize the data of each financial indicator of the target enterprise at different time nodes to obtain corresponding historical data, and divide the historical data of all financial indicators into a hyperparameter tuning set and a training set according to a preset ratio; Use the training set to train the XGBoost model to obtain the initial penalty factor of at least one decision tree and the importance degree of each financial indicator. For any decision tree, form a feature subset with all the financial indicators corresponding to the any decision tree, and obtain the dynamic relative importance degree of the feature subset according to the importance degree of each financial indicator in the feature subset; Obtain the data distribution chaos degree according to the data distribution and data fluctuation of the historical data of each financial indicator in the hyperparameter tuning set, and obtain the complexity degree of the feature subset according to the correlation between every two financial indicators in the feature subset and the data distribution chaos degree; Optimize the initial penalty factor of the any decision tree according to the dynamic relative importance degree and the complexity degree to obtain an adaptive penalty factor, replace the initial penalty factor of each decision tree with the corresponding adaptive penalty factor to obtain a tuned XGBoost model, and early warn the credit risk of the target enterprise according to the tuned XGBoost model.

2. The method for analyzing and warning financial credit risks based on artificial intelligence according to claim 1, characterized in that, The method for obtaining the importance degree of each financial indicator includes: For any financial indicator, during the process of training the XGBoost model using the training set, obtain the information gain after each decision tree split using the any financial indicator to obtain the average information gain of the any financial indicator; Take the average information gain of the any financial indicator as the independent variable of the hyperbolic tangent function to obtain the information gain eigenvalue of the any financial indicator, obtain the importance score of the any financial indicator, and calculate the average value of the information gain eigenvalue and the importance score of the any financial indicator to obtain the importance degree of the any financial indicator.

3. The method for analyzing and warning financial credit risks based on artificial intelligence according to claim 1, wherein The obtaining of the dynamic relative importance degree of the feature subset according to the importance degree of each financial indicator in the feature subset includes: Calculate the average value of the importance degrees of each financial indicator in the feature subset to obtain the average importance degree of the feature subset, form a total feature set with all financial indicators, calculate the average value of the importance degrees of each financial indicator in the total feature set to obtain the average importance degree of the total feature set, and take the difference between the average importance degree of the feature subset and the average importance degree of the total feature set as the independent variable of the activation function to obtain the dynamic relative importance degree of the feature subset.

4. The method for analyzing and warning financial credit risks based on artificial intelligence according to claim 1, characterized in that, The obtaining of the data distribution chaos degree according to the data distribution and data fluctuation of the historical data of each financial indicator in the hyperparameter tuning set includes: In the hyperparameter tuning set, form a data subset with the historical data of each financial indicator in the feature subset; Take the difference between the maximum value and the minimum value in the data subset as the independent variable of the hyperbolic tangent function to obtain the fluctuation range index of the data subset; Calculate the deviation from the mean of each data in the data subset, and use the cumulative result of all deviations from the mean as the independent variable of the hyperbolic tangent function to obtain the degree of fluctuation of the data subset; Calculate the sum between the fluctuation range index and the degree of fluctuation of the data subset to obtain the degree of data distribution chaos.

5. The method for analyzing and warning financial credit risks based on artificial intelligence according to claim 1, characterized in that The complexity of the feature subset is obtained according to the correlation between every two financial indicators in the feature subset and the degree of data distribution chaos, including: In the hyperparameter tuning set, respectively obtain the historical data of each financial indicator in the feature subset to obtain the data sequence corresponding to each financial indicator in the feature subset, calculate the absolute value of the Spearman coefficient between the data sequences corresponding to every two financial indicators in the feature subset, and calculate the average value of all absolute values of the Spearman coefficients to obtain the average degree of correlation of the feature subset; Calculate the difference between the preset correlation threshold and the average degree of correlation, and use the sum of the difference and the degree of data distribution chaos as the independent variable of the hyperbolic tangent function to obtain the complexity of the feature subset.

6. The method for analyzing and warning financial credit risks based on artificial intelligence according to claim 1, characterized in that The initial penalty factor of any decision tree is optimized according to the dynamic relative importance and the complexity to obtain an adaptive penalty factor, including: Calculate the average value between the dynamic relative importance and the complexity to obtain the comprehensive eigenvalue of any decision tree; If the comprehensive eigenvalue is greater than the preset eigenvalue threshold, subtract the comprehensive eigenvalue from the constant 1 to obtain an adjustment coefficient, and calculate the product between the initial penalty factor and the adjustment coefficient to obtain the adaptive penalty factor of any decision tree; If the comprehensive eigenvalue is less than or equal to the preset eigenvalue threshold, calculate the sum between the constant 1 and the comprehensive eigenvalue to obtain an adjustment coefficient, and calculate the product between the initial penalty factor and the adjustment coefficient to obtain the adaptive penalty factor of any decision tree.

7. An artificial intelligence-based financial credit risk analysis and early warning system, comprising a memory, a processor, and a computer program stored in the memory and running on the processor, characterized in that, When the processor executes the computer program, it implements the steps of the artificial intelligence-based financial credit risk analysis and warning method according to any one of claims 1-6.

Citation Information

Cited By

  • Enterprise credit risk dynamic early warning system and method based on behavior change

    CN122264921A