Financial asset rating method and device based on gradient boosting tree, equipment and medium

Through the financial asset rating method based on the gradient boosting tree, the problems of insufficient model performance, data processing capabilities and interpretability of traditional rating methods are solved, and a more accurate, reliable and transparent rating service is achieved, which is suitable for intelligent financial assessment tools.

CN120597113APending Publication Date: 2025-09-05PING AN HEALTH INSURANCE CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510727054.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-30
Publication Date
2025-09-05

AI Technical Summary

Technical Problem

When faced with complex, high-dimensional, and nonlinear financial data, traditional financial asset rating methods have problems such as limited model performance, insufficient data processing capabilities, poor model stability, and insufficient interpretability.

Method used

A financial asset rating method based on gradient boosting tree is adopted. By preprocessing financial data, financial rating features are extracted by combining statistical analysis, domain knowledge and automated feature generation methods. The rating model is constructed using the gradient boosting tree algorithm, and optimized through parameter tuning and model integration technology to generate an understandable rating report.

Benefits of technology

It improves the accuracy and reliability of financial asset ratings, enhances the stability of the model, improves the transparency and interpretability of rating results, meets regulatory requirements, supports automated feature engineering, and flexibly adapts to the needs of different financial scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120597113A_ABST
    Figure CN120597113A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of intelligent financial evaluation, and discloses a financial asset rating method and device based on a gradient boosting tree, equipment and a medium, and the method comprises the steps: preprocessing financial data; financial rating features related to financial asset rating in the preprocessed financial data are determined based on statistical analysis, domain knowledge and an automatic feature generation method, and the financial rating features are extracted; on the basis of a gradient boosting tree algorithm, a financial asset rating model is constructed according to financial rating characteristics, and the financial asset rating model is trained and optimized through a parameter tuning and model integration technology; and obtaining a prediction result and a result explanation of the financial asset rating model, recording the result explanation and generating an understandable rating report. The method can be applied to the development of business systems such as financial science and technology, medical health, old-age care and the like, and the accuracy and reliability of financial asset rating are remarkably improved through nonlinear modeling and high-dimensional data processing capacity.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of intelligent financial evaluation, and in particular to a financial asset rating method, device, equipment and medium based on a gradient boosting tree. Background Art

[0002] Financial asset ratings are an important basis for financial institutions to assess investment risks and formulate investment strategies. Traditional financial asset rating methods rely primarily on expert experience, simple statistical models, or single machine learning algorithms (such as logistic regression and support vector machines).

[0003] However, these methods have the following shortcomings when faced with complex, high-dimensional, and nonlinear financial data: limited model performance: traditional models have weak ability to capture nonlinear relationships and cannot fully explore the potential patterns in financial data; insufficient data processing capabilities: financial data are usually high-dimensional, unstructured, and noisy, and traditional methods cannot effectively process these complex data; poor model stability: in the drastic fluctuations of the financial market, the prediction results of traditional models are often not stable enough to meet actual needs; insufficient interpretability: although many machine learning models (such as deep neural networks) have excellent performance, their "black box" characteristics make model results difficult to explain, which is particularly important in the financial field. Summary of the Invention

[0004] The present invention provides a financial asset rating method, device, computer equipment and medium based on a gradient boosting tree to solve the technical problems in the prior art of financial asset rating processing, such as limited model performance, insufficient data processing capabilities, poor model stability and insufficient interpretability of traditional models.

[0005] In a first aspect, a financial asset rating method based on a gradient boosting tree is provided, comprising:

[0006] Preprocessing financial data, which at least includes financial indicators, market data, and macroeconomic indicators;

[0007] Determine the financial rating features related to financial asset ratings in the pre-processed financial data based on statistical analysis, domain knowledge, and automated feature generation methods, and extract the financial rating features;

[0008] Based on the gradient boosting tree algorithm, a financial asset rating model is constructed according to the financial rating characteristics. Through parameter tuning and model integration technology, the financial asset rating model is trained and optimized;

[0009] Obtain the prediction results and explanations of the financial asset rating model, record the explanations and generate understandable rating reports.

[0010] In a second aspect, a financial asset rating device based on a gradient boosting tree is provided, comprising:

[0011] A preprocessing module is used to preprocess financial data, which includes at least financial indicators, market data, and macroeconomic indicators;

[0012] An extraction module is used to determine and extract financial rating features related to financial asset ratings from pre-processed financial data based on statistical analysis, domain knowledge, and automated feature generation methods;

[0013] The training module is used to build a financial asset rating model based on the gradient boosting tree algorithm and financial rating characteristics. The financial asset rating model is trained and optimized through parameter tuning and model integration technology.

[0014] The generation module is used to obtain the prediction results and result explanations of the financial asset rating model, record the result explanations and generate an understandable rating report.

[0015] In a third aspect, a computer device is provided, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the steps of the above-mentioned method for rating financial assets based on a gradient boosting tree are implemented.

[0016] In a fourth aspect, a computer-readable storage medium is provided, which stores a computer program. When the computer program is executed by a processor, the steps of the above-mentioned financial asset rating method based on the gradient boosting tree are implemented.

[0017] The solution implemented by the aforementioned gradient boosting tree-based financial asset rating method, apparatus, computer equipment, and storage medium can preprocess financial data, which includes at least financial indicators, market data, and macroeconomic indicators; determine and extract financial rating features related to financial asset ratings from the preprocessed financial data based on statistical analysis, domain knowledge, and automated feature generation methods; construct a financial asset rating model based on the financial rating features using the gradient boosting tree algorithm; train and optimize the financial asset rating model through parameter tuning and model integration techniques; obtain prediction results and explanations of the financial asset rating model, record the explanations, and generate an understandable rating report. In this invention, by combining the high interpretability, strong nonlinear modeling capabilities, and high stability of the GBDT algorithm, it effectively addresses the shortcomings of traditional methods in model performance, data processing capabilities, and interpretability. This invention has important theoretical value and practical application prospects, and can provide financial institutions with more accurate, reliable, and transparent rating services. Benefits of this improvement include improved rating accuracy: through nonlinear modeling and high-dimensional data processing capabilities, the accuracy and reliability of financial asset ratings are significantly improved. Enhance model stability: Regularization and random sampling techniques are used to improve model performance in the face of financial market fluctuations. Improve interpretability: Feature importance analysis and model interpretation techniques are used to generate transparent and understandable rating results to meet regulatory requirements. Automation and flexibility: Support for automated feature engineering and model training allows for flexible adaptation to the needs of different financial scenarios. BRIEF DESCRIPTION OF THE DRAWINGS

[0018] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments of the present invention. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative labor.

[0019] Figure 1 1 is a schematic diagram of an application environment of a financial asset rating method based on a gradient boosting tree in one embodiment of the present invention;

[0020] Figure 2 This is a flow chart of a method for rating financial assets based on a gradient boosting tree in one embodiment of the present invention;

[0021] Figure 3 1 is a schematic structural diagram of a financial asset rating device based on a gradient boosting tree in one embodiment of the present invention;

[0022] Figure 4 is a structural diagram of a computer device in one embodiment of the present invention;

[0023] Figure 5FIG. 2 is another structural diagram of a computer device according to an embodiment of the present invention. DETAILED DESCRIPTION

[0024] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of them. All other embodiments obtained by ordinary technicians in this field based on the embodiments of the present invention without making any creative efforts shall fall within the scope of protection of the present invention.

[0025] The financial asset rating method based on gradient boosting tree provided by the embodiment of the present invention can be applied in Figure 1 In an application environment, the client preprocesses financial data, which includes at least financial indicators, market data, and macroeconomic indicators. Based on statistical analysis, domain knowledge, and automated feature generation methods, the client determines and extracts financial rating features relevant to financial asset ratings from the preprocessed financial data. Based on the financial rating features, the client constructs a financial asset rating model using the gradient boosting tree algorithm. The model is trained and optimized through parameter tuning and model integration techniques. The client obtains the prediction results and explanations of the financial asset rating model, records the explanations, and generates an understandable rating report. This invention effectively addresses the shortcomings of traditional methods in model performance, data processing capabilities, and interpretability by combining the high interpretability, strong nonlinear modeling capabilities, and high stability of the GBDT algorithm. This invention has significant theoretical value and practical application prospects, and can provide financial institutions with more accurate, reliable, and transparent rating services. Benefits of this improvement include improved rating accuracy: Through nonlinear modeling and high-dimensional data processing capabilities, the accuracy and reliability of financial asset ratings are significantly improved. Enhanced model stability: Through techniques such as regularization and random sampling, the model's performance in financial market fluctuations is enhanced. Improved interpretability: Through feature importance analysis and model interpretation techniques, transparent and understandable rating results are generated to meet regulatory requirements. Automation and flexibility: Support for automated feature engineering and model training can flexibly adapt to the needs of different financial scenarios. Clients can include, but are not limited to, various personal computers, laptops, smartphones, tablets, and portable wearable devices. The server can be implemented as a standalone server or a server cluster consisting of multiple servers. The present invention is described in detail below through specific embodiments.

[0026] See also Figure 2 As shown, Figure 2 A flowchart of a method for rating financial assets based on a gradient boosting tree according to an embodiment of the present invention includes the following steps:

[0027] S10: Preprocessing financial data, where the financial data includes at least financial indicators, market data, and macroeconomic indicators;

[0028] The gradient boosting tree-based financial asset rating method provided by this invention can be applied to intelligent financial assessment tools, such as smart devices or mobile devices, in various application scenarios. These intelligent financial assessment tools are typically implemented through a server, which can use a gradient boosting tree to perform financial asset ratings on financial data. This gradient boosting tree-based financial asset rating method can be applied to multiple fields, including intelligent financial assessments in both financial and medical scenarios, achieving improved rating accuracy: through nonlinear modeling and high-dimensional data processing capabilities, it significantly improves the accuracy and reliability of financial asset ratings.

[0029] Financial indicator data typically comes from a company's financial statements, such as balance sheets, income statements, and cash flow statements. Financial indicator data reflects a company's financial status and operating results. Financial indicator data is collected from a company's financial statements (such as balance sheets, income statements, and cash flow statements). Data from different sources is consolidated to ensure consistency and accuracy, for example, by standardizing currency units and time periods.

[0030] Specifically, for financial data processing of financial indicators, financial data is preprocessed and cleaned. Data cleaning includes: processing missing values, correcting erroneous data, and removing duplicate data; normalizing and standardizing the cleaned financial data; and constructing new financial data features from the processed financial data based on business needs and analysis purposes.

[0031] There may be some missing indicators in the financial data. For example, some small businesses may not disclose certain non-mandatory financial data. For missing values, the following methods can be used: delete missing values. If the sample size of missing data is small and these samples have little impact on the overall analysis, you can choose to delete the records containing missing values; fill missing values. If deleting missing values ​​will result in a significant reduction in the sample size, you can use a filling method. For example, you can use the mean, median or mode of the indicator to fill the missing values. For time series data, you can also use forward filling (filling with the value of the previous time point) or backward filling (filling with the value of the next time point) methods.

[0032] Financial data may contain entry or calculation errors. For example, a company's net profit may be mistakenly recorded as a negative value when the company is actually profitable. These errors can be discovered and corrected by comparing historical data, industry averages, or performing logical checks with the company's other financial indicators. During the data collection process, financial data for the same company may be recorded repeatedly. Duplicate data can be identified and deleted by checking the company's unique identifier (such as company name, stock code, etc.).

[0033] The numerical ranges of financial indicators can vary widely. For example, total assets may reach billions, while net profit may only be a few million. To facilitate analysis and modeling, data can be normalized, scaling all indicator values ​​to between 0 and 1. This helps eliminate the impact of different scales on the model. Another common method is standardization, which converts financial indicator data into a standard normal distribution with a mean of 0 and a standard deviation of 1. This is suitable for comparing financial indicators of different scales and magnitudes.

[0034] In addition to raw financial indicators, derived indicators can be calculated to enrich the dataset. For example, the debt-to-asset ratio (total liabilities / total assets), the current ratio (current assets / current liabilities), and the gross profit margin (gross profit / operating income) can be calculated. These derived indicators can better reflect a company's financial status and operational capabilities. Financial indicators can be grouped based on factors such as the company's industry and size. For example, companies in different industries can be analyzed separately to better compare the financial performance of companies in the same industry.

[0035] Specifically, for financial data processing of market data, market data includes stock prices, trading volumes, market indices, etc., and there may also be missing values ​​in market data. For example, some stocks may have no trading records on certain trading days. For missing trading data, the following methods can be used: forward filling or backward filling. For stock price data, missing values ​​can be filled with the price of the previous trading day, or with the price of the next trading day. For trading volume data, it can be filled with 0 to indicate that there was no trading on that day. There may be outliers in market data, such as abnormal fluctuations in stock prices. By calculating the mean and standard deviation of the data, outliers that exceed a certain range (such as the mean ± 3 times the standard deviation) can be identified and processed. For example, outliers can be replaced with the values ​​of adjacent data points, or outliers can be directly deleted.

[0036] Check market data for duplicate transaction records, such as duplicate price records for the same stock on the same day, and delete them.

[0037] For time series data such as stock prices, moving averages can be used to smooth the data and reduce the impact of short-term fluctuations. For example, a 5-day moving average price is calculated by dividing the sum of the prices over five consecutive trading days by 5. This method can more clearly reflect the long-term trend of stock prices. Exponential smoothing is a more advanced time series smoothing method that places greater weight on recent data and can better capture data trends.

[0038] In financial market analysis, commonly used technical indicators can be used as features. For example, the relative strength index (RSI) is calculated to measure the degree of overbought or oversold stock prices, and Bollinger Bands are calculated to determine the range of stock price fluctuations. Market sentiment indicators can also be constructed by analyzing market data such as trading volume and price fluctuations. For example, a significant increase in trading volume and rising prices may indicate optimistic market sentiment; conversely, an increase in trading volume but falling prices may indicate pessimistic market sentiment.

[0039] Specifically, for financial data processing of macroeconomic indicators, such as gross domestic product (GDP), inflation rates, interest rates, and unemployment rates, there may be missing data. For example, some countries or regions may not release certain economic indicators in a timely manner. For missing values, the following methods can be used: For time series data, linear interpolation, polynomial interpolation, or spline interpolation can be used to fill in missing values. For example, if the inflation rate data for a particular month is missing, linear interpolation can be performed based on the inflation rates of the preceding and following months.

[0040] Macroeconomic data may contain outliers due to statistical errors or data entry errors. For example, the GDP growth rate for a particular year may be mistakenly recorded as negative, when in fact the economy grew. These errors can be detected and corrected by comparing historical data, international data, or relevant economic models. Check for duplicate records in macroeconomic data, such as duplicate indicator data at the same time point, and remove them.

[0041] Macroeconomic indicators have widely varying numerical ranges. For example, GDP may reach trillions of dollars, while the unemployment rate may be only a few percentage points. To facilitate analysis and modeling, data can be normalized, scaling all indicator values ​​to between 0 and 1. Standardization can also be applied to macroeconomic indicators, converting the data to a distribution with a mean of 0 and a standard deviation of 1. This approach eliminates the influence of different indicator dimensions and makes the data more consistent with the assumption of a normal distribution.

[0042] Derivative indicators of some macroeconomic indicators can be calculated. For example, real GDP growth (GDP growth after deducting inflation) can be calculated to more accurately reflect the actual level of economic growth. Inflation expectations can also be calculated to predict future inflation levels by analyzing the changing trends of inflation rates. Economic cycle indicators can also be constructed by analyzing the time series data of macroeconomic indicators. For example, based on the changing trends of indicators such as GDP growth rate and unemployment rate, the economy can be divided into expansion and recession periods, and the economic cycle status can be represented by a binary variable.

[0043] S20: Determine financial rating features related to financial asset ratings in the preprocessed financial data based on statistical analysis, domain knowledge, and automated feature generation methods, and extract financial rating features;

[0044] Calculate the correlation coefficients between pre-processed financial indicators, market data, and macroeconomic indicators and financial asset ratings. Based on the correlation coefficients, select financial rating features with high correlation to financial asset ratings. Identify and extract the relevant financial rating features from the pre-processed financial data. Statistical analysis uses statistical indicators such as data distribution and correlation to identify features related to financial asset ratings.

[0045] Calculate the correlation coefficient between the pre-processed financial indicators, market data, and macroeconomic indicators and the financial asset rating. For example, calculate the Pearson correlation coefficient or Spearman rank correlation coefficient between the company's debt-to-asset ratio, net profit margin, stock price volatility and other indicators and the credit rating. Screening for highly correlated features: Based on the size of the correlation coefficient, screen out features with a high correlation with the financial asset rating. Generally, features with an absolute value of the correlation coefficient greater than 0.3 or 0.4 can be considered to have a strong correlation. In one embodiment, the correlation coefficient between the debt-to-asset ratio and the financial asset rating is -

[0046] 0.5, indicating that the higher the debt-to-asset ratio, the lower the rating of the financial asset, which is a feature that is strongly correlated with the rating of the financial asset.

[0047] If the financial asset rating is a categorical variable (e.g., AAA, AA, A, etc.), analysis of variance (ANOVA) can be used to test whether certain continuous variables (e.g., financial indicators) differ significantly across rating categories. In one example, the goal is to test whether there are significant differences in net profit margins for companies with different credit ratings. If the ANOVA analysis reveals significant differences in net profit margins across rating categories (p-value less than 0.05), the net profit margin can be considered a characteristic associated with the financial asset rating.

[0048] Among them, a regression model is constructed with financial asset rating as the dependent variable and pre-processed financial indicators, market data and macroeconomic indicators as independent variables; the influence of each feature on the financial asset rating is evaluated based on the coefficient of the regression model or the feature importance index; when the influence degree is greater than the preset influence threshold, the financial rating features related to the financial asset rating in the pre-processed financial data are determined, and the financial rating features are extracted.

[0049] A regression model (such as linear regression or logistic regression) is constructed using the financial asset rating as the dependent variable and preprocessed financial indicators, market data, and macroeconomic indicators as the independent variables. The regression model's coefficients or feature importance indicators (such as the penalty term coefficient in Lasso regression) are used to assess the impact of each feature on the financial asset rating. The larger the absolute value of the coefficient, the more important the feature is influencing the financial asset rating. In a logistic regression model, if the regression coefficient for a company's current ratio is 0.8, this means that for every unit increase in the current ratio, the probability of the company receiving a high rating increases significantly. The current ratio is a key feature in financial ratings.

[0050] Domain knowledge refers to the experience and expertise of financial industry experts, which can help identify important characteristics closely related to financial asset ratings. This domain knowledge indicates that a company's debt repayment ability is a key factor influencing financial asset ratings. For example, indicators such as the debt-to-asset ratio, current ratio, and quick ratio reflect a company's short- and long-term debt repayment ability. A high debt-to-asset ratio may indicate a company faces higher financial risk, thus affecting its financial asset rating. Indicators such as net profit margin, gross profit margin, and return on equity reflect a company's profitability. Companies with strong profitability typically have higher credit ratings because they are more likely to repay debt and pay interest on time. Indicators such as total asset turnover and inventory turnover reflect a company's operational efficiency. Companies with high operational efficiency typically have better cash flow and lower operating risk, which also affects their financial asset ratings.

[0051] Market data such as stock price volatility (such as standard deviation) and stock returns can reflect a company's market recognition and risk profile. For example, companies with higher stock price volatility may face higher market risk, which can affect their financial asset ratings. The correlation between a company's stock and a market index (such as the S&P 500) can also serve as a financial rating feature. Companies with higher correlations with market indices may be more susceptible to macroeconomic factors, and their financial asset ratings may need to consider macroeconomic factors. The economic growth rate reflects the overall macroeconomic situation. During periods of rapid economic growth, companies' operating environment is favorable, and their financial asset ratings may be higher. During periods of economic recession, companies face increased risks, and their financial asset ratings may decline. Interest rate levels affect a company's financing costs and debt burden. In a high interest rate environment, companies face increased debt repayment pressure, which may lead to a decline in their financial asset ratings. Inflation rates affect a company's costs and revenue. High inflation rates can lead to higher costs and lower profits for companies, thus affecting their financial asset ratings.

[0052] Automated feature generation methods utilize machine learning algorithms or data mining techniques to automatically discover features related to financial asset ratings from large amounts of data. Principal component analysis is a dimensionality reduction technique that combines multiple related variables into a small number of principal components that can explain most of the variance in the original data. In one embodiment, there are multiple financial indicators (such as debt-to-asset ratio, current ratio, net profit margin, etc.), and several principal components can be extracted through PCA. These principal components can serve as new financial rating features. For example, the first principal component may mainly reflect the company's debt repayment ability, and the second principal component may mainly reflect the company's profitability. PCA can reduce the dimensionality of features while retaining important information, improving the computational efficiency and interpretability of the model.

[0053] Use the feature importance assessment capabilities of machine learning models (such as random forests and gradient boosted trees) to automatically select features relevant to financial asset ratings. For example, random forests can assess feature importance by calculating the Gini index or mean squared error. In a random forest model, if a financial indicator (such as the debt-to-asset ratio) has a high feature importance score, it indicates that the indicator has a significant impact on the financial asset rating and can be used as an important financial rating feature.

[0054] Recursive feature elimination is a method for progressively filtering features. It starts with a complete feature set and gradually removes features that have the least impact on model performance until the desired number of features is reached. In one example, with 100 features, RFE can be used to gradually filter out the 10 features most relevant to the financial asset rating. These features can then be used for subsequent modeling and analysis.

[0055] Deep learning models (such as convolutional neural networks and recurrent neural networks) can automatically extract complex features from data through multi-layer neural network structures. For example, for time series market data, recurrent neural networks can automatically learn temporal dependencies and patterns within the data. In financial asset rating tasks, deep learning models can be used to model market data such as stock prices and trading volume. The model's hidden layers can extract features relevant to the financial asset rating. Deep learning models can process large amounts of data and automatically learn complex relationships and patterns within the data, making them suitable for feature extraction from complex financial data.

[0056] Based on the results of correlation analysis, variance analysis, and regression analysis, domain knowledge and automated feature generation methods are listed to determine the financial rating features with high correlation with financial asset ratings in the preprocessed financial data; combined with the experience and expertise of financial industry experts, important financial indicators, market data, and macroeconomic indicators are listed; and financial rating features are extracted based on principal component analysis, feature selection algorithms, and deep learning models.

[0057] Based on the results of correlation analysis, variance analysis, and regression analysis, the characteristics with a high correlation with the ratings of financial assets are listed.

[0058] Combining the experience and expertise of financial industry experts, it lists important financial indicators, market data and macroeconomic indicators.

[0059] Financial rating features are extracted from principal component analysis, feature selection algorithm and deep learning model, and the results are generated by automatic feature generation. The features extracted by the above three methods are deduplicated and merged to form a complete feature list. The importance of the merged features is evaluated by machine learning models (such as random forest, gradient boosting tree, etc.), and the features that have the greatest impact on the financial asset rating are further screened out. In one embodiment, features such as debt-to-asset ratio, net profit margin, and stock price volatility are extracted through statistical analysis and domain knowledge, and features such as principal component 1 and principal component 2 are extracted through automatic feature generation. The importance of these features is evaluated by the random forest model, and the debt-to-asset ratio, net profit margin and principal component 1 are finally screened out as the most important financial rating features.

[0060] The extracted financial rating features are verified using historical data to ensure that these features can effectively predict the ratings of financial assets. The predictive ability of the features can be evaluated by building a prediction model (such as logistic regression, support vector machine, etc.). Based on the verification results, the feature combination is further optimized. For example, if the performance of a certain feature on different data sets is unstable, you can consider removing the feature; if there is a high correlation between certain features, you can consider merging or selecting one of the features. In one embodiment, a financial asset rating model based on the extracted features is constructed using historical data. It is found that the debt-to-asset ratio and net profit margin contribute the most to the predictive ability of the model, while the contribution of stock price volatility is relatively small. Therefore, the feature combination can be optimized, stock price volatility can be removed, and only the debt-to-asset ratio and net profit margin can be retained as financial rating features.

[0061] S30: Based on the gradient boosting tree algorithm, a financial asset rating model is constructed according to the financial rating characteristics. The financial asset rating model is trained and optimized through parameter tuning and model integration technology.

[0062] Building a financial asset rating model based on the GBDT (Gradient Boosting Decision Tree) algorithm is an effective approach. The GBDT algorithm combines multiple weak learners (typically decision trees) to form a strong learner. It is capable of handling high-dimensional and missing data and can capture nonlinear relationships in the data.

[0063] The GBDT algorithm combines multiple weak learners (typically decision trees) to form a strong learner, suitable for classification and regression tasks. In financial asset rating, GBDT can be used to predict the credit rating or default probability of a financial asset. The GBDT model consists of multiple decision trees, and the model structure includes parameters such as the number of decision trees, tree depth, and learning rate. The choice of these parameters has a significant impact on model performance. Typically, the initial prediction value of a GBDT model is the mean or median of the data. For classification problems, this can be the prior probability of each class.

[0064] Specifically, the financial rating feature data and the label data of the financial asset rating are divided into a training set, a validation set, and a test set, and the initial parameters of the gradient boosting tree model are set; for the training set data, multiple decision tree models are iteratively trained according to the principle of the gradient boosting tree algorithm to construct an initial gradient boosting tree model; based on grid search, random search, or Bayesian optimization method, the key parameters of the gradient boosting tree model are tuned on the validation set, and the parameter combination with the best performance is selected; the model integration technology is selected according to the needs, and the tuned gradient boosting tree model is integrated through the model integration technology; the test set data is used to evaluate the final trained and optimized financial asset rating model. If the evaluation result does not meet the preset requirements, return to the parameter tuning or model integration step for further optimization.

[0065] The extracted financial rating features are used as input variables, and the actual rating of the financial asset is used as the output variable (the rating can usually be numerically coded, such as 5 for AAA and 4 for AA). The dataset is divided into a training set, a validation set, and a test set. For example, in one embodiment, the division ratio is 70% (training set), 15% (validation set), and 15% (test set). The training set is used to train the model, the validation set is used for parameter tuning, and the test set is used to evaluate the final performance of the model.

[0066] In the GBDT algorithm, the initial model is usually a constant model whose predicted value is the mean of the output variable of the training set (for regression problems) or the most likely category (for classification problems). For example, in the regression problem of financial asset ratings, the predicted value of the initial model is the average of all financial asset rating scores in the training set.

[0067] For each iteration, the residual of the current model is calculated. In the regression problem, the residual rti = yi - Ft-1(xi), where yi is the true rating score of the i-th sample and Ft-1(xi) is the predicted value of the model obtained in the previous t-1 iterations.

[0068] Use the residual as the new target variable to train a new decision tree model ht(x). The decision tree can be trained using the CART (Classification and Regression Trees) algorithm, which recursively divides the feature space to find the optimal partitioning point and partitioning method so that the sum of squared residuals of the partitioned subsets is minimized. Calculate the weight γt of the new decision tree model so that the new model Ft(x) = Ft-1(x) + γtht(x) with this weight minimizes the loss function on the training set. For regression problems, the loss function is usually the mean squared error (MSE), which can be solved by taking the derivative of the loss function with respect to γt and setting the derivative to zero. After multiple rounds of iteration, all trained decision tree models are weighted and combined to obtain the final GBDT model.

[0069] Grid search is an exhaustive search method that traverses given parameter combinations, evaluates the performance of the model under each parameter combination on the validation set, and selects the parameter combination with the best performance. For example, we can set the learning rate value range to [0.01, 0.05, 0.1, 0.2] and the maximum depth of the decision tree to [3, 5, 7, 9]. Then, all possible parameter combinations are traversed, and the accuracy (for classification problems) or mean square error (for regression problems) of the model under each combination on the validation set is calculated, and the parameter combination with the best performance is selected.

[0070] The advantage is that it can find the globally optimal parameter combination. The disadvantage is that it is computationally expensive. When there are many parameters or a wide range of values, the search time can be very long. Random search randomly samples parameter combinations within a given parameter range and then evaluates the model's performance on a validation set. Compared to grid search, random search does not require traversing all possible parameter combinations and is therefore more computationally efficient. For example, we can randomly sample 100 values ​​within the range of the learning rate and 100 values ​​within the range of the maximum depth of the decision tree, and then combine these into 100 parameter combinations for evaluation. The advantage is its high computational efficiency and the ability to find a good parameter combination in a short time. The disadvantage is that it may not find the globally optimal parameter combination.

[0071] Bayesian optimization is a global optimization method based on Bayes' theorem. It constructs a probabilistic model of the objective function (model performance on a validation set) and then selects the next most promising parameter combination for evaluation based on this probabilistic model. Bayesian optimization leverages information from previously evaluated parameter combinations to more efficiently search for optimal parameters. Its advantage is that it can find optimal parameter combinations with fewer evaluations, making it suitable for high-dimensional parameter spaces and computationally expensive model training processes. However, its disadvantage is that its implementation is relatively complex.

[0072] Bagging (Bootstrap Aggregating) is a parallel ensemble learning method that generates multiple training subsets by sampling with replacement from the original training set. A base model (such as a GBDT model) is then trained independently on each training subset. The predictions of all base models are then averaged (for regression problems) or voted (for classification problems) to obtain the final prediction. For example, suppose the original training set contains 100 examples. We sample with replacement to generate 10 training subsets, each containing 100 examples (possibly with duplicates). A GBDT model is then trained on each subset. For a new test example, each of the 10 GBDT models is asked to make a prediction, and the final prediction is averaged. This reduces model variance and improves model stability. Because different training subsets contain different examples, the correlation between base models is low. Averaging or voting can reduce the prediction error of individual models. This approach is robust to noisy data and outliers.

[0073] Stacking is a hierarchical ensemble learning method that typically consists of two layers of models. The first layer uses multiple base models (such as GBDT, random forest, and support vector machine) to train the original data and obtain prediction results for each base model on the validation and test sets. These prediction results are then used as new features, along with the original features, to form the input data for the second layer model (such as logistic regression or linear regression). The second layer model ultimately produces the final prediction results.

[0074] In the first layer, we trained three base models, GBDT, random forest, and support vector machine, on the financial asset rating data, obtaining prediction results on the validation and test sets. These prediction results were then combined with the original financial rating features as input data for the second-layer logistic regression model. This model was then trained to obtain the final financial asset rating prediction results. This method leverages the strengths of different base models to improve model prediction performance. Because different base models may learn data features from different perspectives, stacking ensembles can combine this information, resulting in better generalization.

[0075] Initialization: Prepare financial rating feature data and financial asset rating label data, divide the data into training, validation, and test sets, and set the initial parameters of the GBDT model. Model construction: Use the training set data to iteratively train multiple decision tree models according to the principles of the GBDT algorithm to build an initial GBDT model. Parameter tuning: Use grid search, random search, or Bayesian optimization to tune key parameters of the GBDT model (such as the learning rate, maximum depth of the decision tree, subsampling ratio, and feature subsampling ratio) on the validation set and select the optimal parameter combination. Model ensemble: Select model ensemble techniques such as bagging or stacking based on your needs to ensemble the tuned GBDT models to further improve model performance. Model evaluation: Use the test set data to evaluate the final trained and optimized financial asset rating model. Common evaluation metrics include accuracy, precision, recall, F1 score (for classification problems), or mean squared error (MSE) and mean absolute error (MAE) (for regression problems). If the evaluation results do not meet your requirements, return to the parameter tuning or model ensemble steps for further optimization.

[0076] S40: Obtain the prediction results and explanations of the financial asset rating model, record the explanations and generate an understandable rating report.

[0077] Input financial rating features into the financial asset rating model, use the prediction function of the financial asset rating model to make predictions, and save the prediction results in a suitable data structure or file; analyze the importance of each feature in the financial asset rating model and record the explanation of the results; present the predicted rating results of each financial asset in the form of tables, charts, etc., and generate an understandable rating report for the predicted rating results of each financial asset.

[0078] Analyze the importance of each feature in the model to understand which features have the greatest impact on the predictions. This can be achieved using the model's built-in feature importance score. For each prediction, you can use methods such as the SHAP (SHapley Additive exPlanations) value to interpret the model's predictions. The SHAP value quantifies the contribution of each feature to the prediction. Analyze the overall behavior of the model, including the relationship between features and predictions. You can plot feature importance graphs and feature dependency graphs on predictions.

[0079] Design a clear report structure, typically consisting of the following sections: report summary, forecast results, explanation of results, rating rationale, recommendations, and conclusion. Briefly describe the background, purpose, and model used for the rating. Clearly present the forecast rating results for each financial asset in tables, charts, and other formats. For classification problems, list the predicted category for each asset; for regression problems, list the score. Explain the reasoning behind the forecast results in plain language. Cite the results of feature importance analysis and individual forecast explanations to explain which factors positively or negatively influenced the rating. Provide a detailed description of the model, data, and methodology underlying the rating. This should include the model training process, features used, and parameter settings. Provide recommendations based on the rating results, such as recommending risk control measures for high-risk assets and increasing investment in low-risk assets. Summarize the overall situation and trends of the ratings. Charts and tables can make complex data and results more intuitive and understandable. For example, use a bar chart to display the score distribution of different assets, or a pie chart to display the proportion of different rating categories. Use plain language and avoid excessive technical terminology. If technical terminology is necessary, provide explanations or annotations. Provide specific examples in the report to help readers better understand the rating results and explanations. For example, select a representative asset and explain its rating results and influencing factors in detail. Organize the report content into layers so that readers can quickly find the sections they are interested in. Use headings, subheadings, paragraphs, etc. to organize the content.

[0080] As can be seen, in the above scheme, for intelligent financial assessment, financial data is first preprocessed, including at least financial indicators, market data, and macroeconomic indicators. Financial rating features relevant to financial asset ratings are then identified and extracted from the preprocessed financial data using statistical analysis, domain knowledge, and automated feature generation methods. A financial asset rating model is constructed based on these features using the gradient boosting tree algorithm. The model is trained and optimized through parameter tuning and model integration techniques. The prediction results and explanations of the financial asset rating model are obtained, recorded, and an understandable rating report is generated. This invention, by combining the high interpretability, strong nonlinear modeling capabilities, and high stability of the GBDT algorithm, effectively addresses the shortcomings of traditional methods in model performance, data processing capabilities, and interpretability. This invention has significant theoretical value and practical application prospects, and can provide financial institutions with more accurate, reliable, and transparent rating services. Benefits of this improvement include improved rating accuracy: Through nonlinear modeling and high-dimensional data processing capabilities, the accuracy and reliability of financial asset ratings are significantly improved. Enhanced model stability: Through techniques such as regularization and random sampling, the model's performance in financial market fluctuations is enhanced. Improved explainability: Through feature importance analysis and model interpretation techniques, transparent and understandable rating results are generated to meet regulatory requirements. Automation and flexibility: Supports automated feature engineering and model training, which can flexibly adapt to the needs of different financial scenarios.

[0081] It should be understood that the size of the serial numbers of the steps in the above embodiments does not mean the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present invention.

[0082] In one embodiment, a financial asset rating device based on a gradient boosting tree is provided. The financial asset rating device based on a gradient boosting tree corresponds one-to-one to the financial asset rating method based on a gradient boosting tree in the above embodiment. Figure 3 As shown, the financial asset rating device based on the gradient boosting tree includes a preprocessing module 101, an extraction module 102, a training module 103, and a generation module 104. The functional modules are described in detail as follows:

[0083] A preprocessing module 101 is used to preprocess financial data, where the financial data includes at least financial indicators, market data, and macroeconomic indicators;

[0084] Extraction module 102, configured to determine financial rating features related to financial asset ratings in the pre-processed financial data based on statistical analysis, domain knowledge, and automated feature generation methods, and extract the financial rating features;

[0085] A training module 103 is used to construct a financial asset rating model based on the gradient boosting tree algorithm and the financial rating characteristics, and to train and optimize the financial asset rating model through parameter tuning and model integration technology;

[0086] The generation module 104 is used to obtain the prediction results and result explanations of the financial asset rating model, record the result explanations and generate an understandable rating report.

[0087] In one embodiment, the configuration module 101 is specifically configured to:

[0088] Preprocessing and cleaning financial data, including handling missing values, correcting erroneous data, and removing duplicate data;

[0089] Normalization and standardization of cleaned financial data;

[0090] Based on business needs and analysis purposes, new financial data features are constructed from the processed financial data.

[0091] In one embodiment, the configuration module 102 is specifically configured to:

[0092] Calculate the correlation coefficients between pre-processed financial indicators, market data and macroeconomic indicators and financial asset ratings;

[0093] According to the size of the correlation coefficient, the financial rating features with high correlation with the financial asset rating are screened out, the financial rating features related to the financial asset rating in the preprocessed financial data are determined, and the financial rating features are extracted.

[0094] In one embodiment, the configuration module 102 is further configured to:

[0095] A regression model is constructed with financial asset ratings as the dependent variable and pre-processed financial indicators, market data and macroeconomic indicators as independent variables;

[0096] Evaluate the impact of each feature on the financial asset rating based on the coefficient of the regression model or the feature importance index;

[0097] When the degree of influence is greater than a preset influence threshold, the financial rating features related to the financial asset rating in the preprocessed financial data are determined and the financial rating features are extracted.

[0098] In one embodiment, the configuration module 102 is further configured to:

[0099] Based on the results of correlation analysis, variance analysis, and regression analysis, domain knowledge and automated feature generation methods are listed to identify financial rating features with high correlation with financial asset ratings in preprocessed financial data;

[0100] Combining the experience and expertise of financial industry experts, list important financial indicators, market data and macroeconomic indicators;

[0101] Financial rating features are extracted based on principal component analysis, feature selection algorithm and deep learning model.

[0102] In one embodiment, the interface module 103 is specifically configured to:

[0103] Divide the financial rating feature data and the labeled data of financial asset ratings into training sets, validation sets, and test sets, and set the initial parameters of the gradient boosting tree model;

[0104] Training set data, iteratively train multiple decision tree models according to the principle of gradient boosting tree algorithm to build the initial gradient boosting tree model;

[0105] Based on grid search, random search, or Bayesian optimization methods, the key parameters of the gradient boosting tree model are tuned on the validation set to select the parameter combination with the best performance;

[0106] Select model integration technology based on needs, and integrate the tuned gradient boosting tree model through model integration technology;

[0107] Use the test set data to evaluate the final trained and optimized financial asset rating model. If the evaluation result does not meet the preset requirements, return to the parameter tuning or model integration step for further optimization.

[0108] In one embodiment, the replacement module 104 is specifically configured to:

[0109] Input the financial rating features into the financial asset rating model, use the prediction function of the financial asset rating model to make predictions, and save the prediction results into a suitable data structure or file;

[0110] Analyze the importance of various features in financial asset rating models and document the interpretation of the results;

[0111] The predicted rating results of each financial asset are presented in the form of tables, charts, etc., and an understandable rating report is generated based on the predicted rating results of each financial asset.

[0112] This invention provides a financial asset rating device based on a gradient boosting tree algorithm. The device preprocesses financial data, which includes at least financial indicators, market data, and macroeconomic indicators. The device then uses statistical analysis, domain knowledge, and automated feature generation methods to identify and extract financial rating features relevant to financial asset ratings from the preprocessed financial data. Based on these features, the device constructs a financial asset rating model using the gradient boosting tree algorithm. The model is trained and optimized through parameter tuning and model integration techniques. The device then obtains prediction results and explanations of the model, records these explanations, and generates an understandable rating report. By combining the high interpretability, strong nonlinear modeling capabilities, and high stability of the GBDT algorithm, the device effectively addresses the shortcomings of traditional methods in model performance, data processing capabilities, and interpretability. This invention has significant theoretical value and practical application prospects, enabling it to provide financial institutions with more accurate, reliable, and transparent rating services. Benefits of this improvement include improved rating accuracy: Through nonlinear modeling and high-dimensional data processing capabilities, the accuracy and reliability of financial asset ratings are significantly improved. Enhanced model stability: Through techniques such as regularization and random sampling, the model's performance in the face of financial market fluctuations is enhanced. Improved explainability: Through feature importance analysis and model interpretation techniques, transparent and understandable rating results are generated to meet regulatory requirements. Automation and flexibility: Supports automated feature engineering and model training, which can flexibly adapt to the needs of different financial scenarios.

[0113] The specific limitations of the gradient boosting tree-based financial asset rating device can be found in the limitations of the gradient boosting tree-based financial asset rating method described above and will not be further elaborated here. Each module in the aforementioned gradient boosting tree-based financial asset rating device can be implemented in whole or in part through software, hardware, or a combination thereof. Each of these modules can be embedded in or independent of a processor in a computer device in hardware form, or stored in a computer device memory in software form, so that the processor can call and execute the corresponding operations of each of these modules.

[0114] In one embodiment, a computer device is provided. The computer device may be a server, and its internal structure diagram may be as follows: Figure 4As shown. The computer device includes a processor, memory, network interface and database connected via a system bus. The processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes non-volatile and / or volatile storage media and internal memory. The non-volatile storage medium stores an operating system, a computer program and a database. The internal memory provides an environment for the operation of the operating system and computer program in the non-volatile storage medium. The network interface of the computer device is used to communicate with an external client via a network connection. When the computer program is executed by the processor, it implements the functions or steps on the service side of a financial asset rating method based on a gradient boosting tree.

[0115] In one embodiment, a computer device is provided. The computer device may be a client, and its internal structure diagram may be as follows: Figure 5 As shown. The computer device includes a processor, memory, network interface, display screen and input device connected via a system bus. The processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and computer program in the non-volatile storage medium. The network interface of the computer device is used to communicate with an external server via a network connection. When the computer program is executed by the processor, it implements the functions or steps on the client side of a financial asset rating method based on a gradient boosting tree.

[0116] In one embodiment, a computer device is provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the following steps are performed:

[0117] Preprocessing financial data, which at least includes financial indicators, market data, and macroeconomic indicators;

[0118] Determine the financial rating features related to financial asset ratings in the pre-processed financial data based on statistical analysis, domain knowledge, and automated feature generation methods, and extract the financial rating features;

[0119] Based on the gradient boosting tree algorithm, a financial asset rating model is constructed according to the financial rating characteristics. Through parameter tuning and model integration technology, the financial asset rating model is trained and optimized;

[0120] Obtain the prediction results and explanations of the financial asset rating model, record the explanations and generate understandable rating reports.

[0121] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the following steps are implemented:

[0122] Preprocessing financial data, which at least includes financial indicators, market data, and macroeconomic indicators;

[0123] Determine the financial rating features related to financial asset ratings in the pre-processed financial data based on statistical analysis, domain knowledge, and automated feature generation methods, and extract the financial rating features;

[0124] Based on the gradient boosting tree algorithm, a financial asset rating model is constructed according to the financial rating characteristics. Through parameter tuning and model integration technology, the financial asset rating model is trained and optimized;

[0125] Obtain the prediction results and explanations of the financial asset rating model, record the explanations and generate understandable rating reports.

[0126] It should be noted that the above functions or steps that can be implemented by the computer-readable storage medium or computer device can be found in the relevant descriptions of the server side and the client side in the aforementioned method embodiment. To avoid repetition, they will not be described one by one here.

[0127] Those skilled in the art will appreciate that all or part of the processes in the above-mentioned embodiments can be implemented by instructing the relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, any reference to memory, storage, database or other media used in the embodiments provided in this application can include non-volatile and / or volatile memory. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM) or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link (Synchlink) DRAM (SLDRAM), memory bus (Rambus) direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM).

[0128] Those skilled in the art will clearly understand that for the sake of convenience and brevity of description, only the division of the above-mentioned functional units and modules is used as an example. In actual applications, the above-mentioned functions can be distributed and completed by different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above.

[0129] The embodiments described above are only used to illustrate the technical solutions of the present invention, rather than to limit the same. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. These modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the various embodiments of the present invention, and should all be included in the scope of protection of the present invention.

Claims

1. A financial asset rating method based on gradient boosting tree, characterized in that: include: Preprocessing financial data, which at least includes financial indicators, market data, and macroeconomic indicators; Determine the financial rating features related to financial asset ratings in the pre-processed financial data based on statistical analysis, domain knowledge, and automated feature generation methods, and extract the financial rating features; Based on the gradient boosting tree algorithm, a financial asset rating model is constructed according to the financial rating characteristics. Through parameter tuning and model integration technology, the financial asset rating model is trained and optimized; Obtain the prediction results and explanations of the financial asset rating model, record the explanations and generate understandable rating reports.

2. The financial asset rating method based on the gradient boosting tree according to claim 1, characterized in that: The steps of pre-processing financial data, where the financial data at least includes financial indicators, market data, and macroeconomic indicators, include: Preprocessing and cleaning financial data, including handling missing values, correcting erroneous data, and removing duplicate data; Normalization and standardization of cleaned financial data; Based on business needs and analysis purposes, new financial data features are constructed from the processed financial data.

3. The financial asset rating method based on gradient boosting tree according to claim 1, characterized in that: The step of determining financial rating features related to financial asset ratings in the pre-processed financial data based on statistical analysis, domain knowledge, and an automated feature generation method, and extracting financial rating features includes: Calculate the correlation coefficients between pre-processed financial indicators, market data and macroeconomic indicators and financial asset ratings; According to the size of the correlation coefficient, the financial rating features with high correlation with the financial asset rating are screened out, the financial rating features related to the financial asset rating in the preprocessed financial data are determined, and the financial rating features are extracted.

4. The financial asset rating method based on gradient boosting tree according to claim 1, characterized in that: The step of determining the financial rating features related to the financial asset rating in the pre-processed financial data based on statistical analysis, domain knowledge and automated feature generation method, and extracting the financial rating features further includes: A regression model is constructed with financial asset ratings as the dependent variable and pre-processed financial indicators, market data and macroeconomic indicators as independent variables; Evaluate the impact of each feature on the financial asset rating based on the coefficient of the regression model or the feature importance index; When the degree of influence is greater than a preset influence threshold, the financial rating features related to the financial asset rating in the preprocessed financial data are determined and the financial rating features are extracted.

5. The financial asset rating method based on gradient boosting tree according to claim 1, characterized in that: The step of determining the financial rating features related to the financial asset rating in the pre-processed financial data based on statistical analysis, domain knowledge and automated feature generation method, and extracting the financial rating features further includes: Based on the results of correlation analysis, variance analysis, and regression analysis, domain knowledge and automated feature generation methods are listed to identify financial rating features with high correlation with financial asset ratings in pre-processed financial data; Combining the experience and expertise of financial industry experts, list important financial indicators, market data and macroeconomic indicators; Financial rating features are extracted based on principal component analysis, feature selection algorithm and deep learning model.

6. The financial asset rating method based on gradient boosting tree according to claim 1, characterized in that: The steps of constructing a financial asset rating model based on the gradient boosting tree algorithm according to financial rating characteristics and training and optimizing the financial asset rating model through parameter tuning and model integration technology include: Divide the financial rating feature data and the labeled data of financial asset ratings into training sets, validation sets, and test sets, and set the initial parameters of the gradient boosting tree model; Training set data, iteratively train multiple decision tree models according to the principle of gradient boosting tree algorithm to build the initial gradient boosting tree model; Based on grid search, random search, or Bayesian optimization methods, the key parameters of the gradient boosting tree model are tuned on the validation set to select the parameter combination with the best performance; Select model integration technology based on needs, and integrate the tuned gradient boosting tree model through model integration technology; Use the test set data to evaluate the final trained and optimized financial asset rating model. If the evaluation result does not meet the preset requirements, return to the parameter tuning or model integration step for further optimization.

7. The financial asset rating method based on gradient boosting tree according to claim 1, characterized in that: The steps of obtaining the prediction results and explanations of the financial asset rating model, recording the explanations of the results, and generating an understandable rating report include: Input the financial rating features into the financial asset rating model, use the prediction function of the financial asset rating model to make predictions, and save the prediction results into a suitable data structure or file; Analyze the importance of various features in financial asset rating models and document the interpretation of the results; The predicted rating results of each financial asset are presented in the form of tables, charts, etc., and an understandable rating report is generated based on the predicted rating results of each financial asset.

8. A financial asset rating device based on gradient boosting tree, characterized in that: include: A preprocessing module is used to preprocess financial data, which includes at least financial indicators, market data, and macroeconomic indicators; An extraction module is used to determine and extract financial rating features related to financial asset ratings from pre-processed financial data based on statistical analysis, domain knowledge, and automated feature generation methods; The training module is used to build a financial asset rating model based on the gradient boosting tree algorithm and financial rating characteristics. The financial asset rating model is trained and optimized through parameter tuning and model integration technology. The generation module is used to obtain the prediction results and result explanations of the financial asset rating model, record the result explanations and generate an understandable rating report.

9. A computer device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the computer program, the steps of the financial asset rating method based on the gradient boosting tree are implemented as described in any one of claims 1 to 7.

10. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the steps of the financial asset rating method based on the gradient boosting tree are implemented as described in any one of claims 1 to 7.

Citation Information

Cited By

  • AI financial intelligent analysis system

    CN121458470A