Integrated model-based listed company credit risk assessment method and system

By fine-grained division of listed companies and screening characteristic variables, combined with the integrated model to integrate the prediction results of multiple classifiers, the traditional credit risk assessment methods are solved in terms of accuracy and timeliness, and more efficient credit risk assessment is achieved.

CN120163641APending Publication Date: 2025-06-17BEIJING CHINA EVERBRIGHT TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510214173.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-26
Publication Date
2025-06-17

AI Technical Summary

Technical Problem

When evaluating the credit risk of listed companies, the traditional credit risk assessment method has low accuracy and timeliness, especially the non-annual report financial data such as quarterly and interim reports, making it difficult for the model to capture seasonal fluctuations and immediate changes in the company's financial status.

Method used

The credit risk assessment method of listed companies based on an integrated model is adopted, and the financial statement data is divided into fine-grained areas (including first-quarter reports, interim reports, third-quarter reports and annual reports), combined with chi-square filtering, F-test and feature variable filters, feature variables with a greater impact on the prediction results, and the prediction results of multiple classifiers are integrated through a single-layer mean integration model to improve the accuracy and timeliness of credit risk assessment.

Benefits of technology

Through fine-grained financial data processing and optimized screening of characteristic variables, the accuracy and timeliness of credit risk assessment are improved, and the seasonal and immediate changes in the financial status of listed companies can be captured more accurately.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120163641A_ABST
    Figure CN120163641A_ABST
Patent Text Reader

Abstract

The invention provides a listed company credit risk assessment method and system based on an integrated model, and relates to the field of electric digital data processing, and the method comprises the steps: dividing the financial statement data of a listed company into four statement types of one-quarter reports, middle reports, three-quarter reports and annual reports, and corresponding sample data; all the feature variables of the sample data are subjected to chi-square filtering and F testing, and the number of to-be-selected feature variables is determined; determining a target feature variable and a corresponding training sample data set in combination with an importance score sorting result of the feature variable filter; and carrying out random down-sampling processing on the training sample data set by a balanced Bagging classifier based on a decision tree to obtain a balanced sample subset, and training to obtain a single-layer mean value integration model for carrying out credit risk assessment. By implementing the method, the accuracy and timeliness of credit risk assessment on listed companies are improved by making the fine granularity of the financial statement data accurate to quarters and optimizing and screening the characteristic variables trained by the model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of electronic digital data processing, and in particular, to a method and system for evaluating the credit risk of listed companies based on an integrated model. Background Art

[0002] In the modern financial system, credit risk assessment is an important tool for banks and investment institutions to manage risks. Accurately evaluating the credit risk of listed companies can help these institutions make more informed decisions when providing loans, making investment decisions, and controlling risks. With the rapid development of the financial market and the increasing amount of financial data, traditional credit assessment methods are gradually unable to meet the efficient and dynamic market demands.

[0003] In related technologies, most credit risk assessment methods rely on statistical or machine learning technologies. They mainly perform feature analysis on the characteristic indicators in the annual financial statements of listed companies by constructing models, and then evaluate the future credit risk situation of listed companies.

[0004] However, traditional credit risk assessment methods mainly focus on analyzing annual report data and ignore the financial data of other reporting periods (such as quarterly reports and semi-annual reports). These financial data may contain key information for evaluating the short-term financial condition changes of listed companies, making it difficult for the model to capture the seasonal fluctuations and immediate changes in the company's financial condition, thus affecting the timeliness and accuracy of the assessment. Summary of the Invention

[0005] This application provides a method and system for evaluating the credit risk of listed companies based on an integrated model to address the problem of low accuracy and timeliness of traditional credit risk assessment methods when evaluating the credit risk of listed companies.

[0006] In a first aspect, this application provides a method for evaluating the credit risk of listed companies based on an integrated model, which is applied to a risk assessment system. The method includes: Dividing the financial statement data of a listed company according to 4 types of statement types to obtain sample data corresponding to different statement types. The statement types include the first quarterly report, semi-annual report, third quarterly report, and annual report. The sample data includes multiple indicator categories and one or more characteristic variables corresponding to each such indicator category; Performing chi-square filtering and F-tests on all the characteristic variables of the sample data to determine the number of characteristic variables to be selected; Calculating and ranking the importance scores of all the characteristic variables using a preset characteristic variable selector, and determining target characteristic variables based on the number of characteristic variables. The characteristic variable selector includes Random Forest, XGBoost, GBDT, and LightGBM. The target characteristic variables are used to construct an index system for evaluating the credit risk of listed companies; Standardize the sample data of the target feature variable according to the positive and negative calculation formula of the credit risk indicator to obtain a training sample data set. The positive and negative calculation formula of the credit risk indicator includes a positive indicator standardization formula and a negative indicator standardization formula; If it is detected that the training sample data set is unbalanced, use a decision tree-based balanced Bagging classifier to perform random downsampling on the training sample data set to obtain multiple balanced sample subsets; Use the multiple balanced sample subsets to train classifiers Random Forest, XGBoost, and LightGBM respectively, and integrate them to obtain a single-layer mean integration model; Input the financial statement data of the listed company to be evaluated into the single-layer mean integration model to obtain the corresponding credit risk assessment result. The credit risk assessment result includes two types: the financial condition is at risk and the financial state is normal.

[0007] Through the above embodiments, the risk assessment system evaluates by dividing the financial data into 4 report types (first-quarter report, mid-year report, third-quarter report, and annual report), enabling the model to more accurately capture the financial performance in different time periods. Further, through chi-square filtering, F-test, and using a feature variable filter to calculate and sort the importance scores, the feature variables with a greater impact on the prediction result are selected, and an equilibrium data set is constructed based on the financial data corresponding to the feature variables to train the single-layer mean integration model. The final credit risk assessment result is determined by integrating the prediction results of multiple classifiers. This method improves the accuracy and timeliness of credit risk assessment of listed companies by refining the financial statement data to the quarterly level and optimizing the feature variables for screening model training.

[0008] In some embodiments, the step of performing chi-square filtering and F-test on all feature variables of the sample data respectively to determine the number of feature variables to be selected specifically includes: Perform chi-square filtering and F-test on all feature variables of the sample data respectively to obtain the chi-square test scores and F-test scores of different report types under different numbers of feature variables in different years; Statistically count the top 3 feature quantity values with the highest chi-square test scores and F-test scores and their corresponding frequencies that appear in different years of different report types respectively to obtain a table of the frequencies of the optimal feature quantity values; Determine the number of target feature variables based on the table of the frequencies of the optimal feature quantity values.

[0009] Through the above embodiments, the risk assessment system determines the chi-square test score and the F-test score through chi-square filtering, and determines the number of feature variables participating in model training according to the variation laws of the two with the number of feature variable values in different report types in the same year, so that the number of feature variables selected to participate in model training is more representative, enabling the model to better predict the credit risk of listed companies based on a small number of feature variables, improving the reliability and efficiency of prediction.

[0010] In some embodiments, the step of determining the target number of feature variables according to the optimal feature number value frequency table specifically includes: Determine the feature number value with the highest frequency in the chi-square test as the candidate feature number value; According to the optimal feature number value frequency table, determine the target frequency corresponding to the candidate feature number value in the F-test; If the target frequency is within the preset frequency range, determine the candidate feature number value as the target number of feature variables; If the target frequency is not within the preset frequency range, give an early warning about the accuracy of the chi-square test.

[0011] Through the above embodiments, the risk assessment system creates an optimal feature number value frequency table, verifies and compares the chi-square filtering results with the F-test results, so that the number of feature variables selected to participate in model training is more representative and improves the reliability of prediction.

[0012] In some embodiments, the step of calculating the importance scores and sorting all feature variables using a preset feature variable filter and determining the target feature variables according to the number of feature variables specifically includes: Based on the preset feature variable filter, calculate the importance scores of the feature variables in different report types respectively, and obtain the importance scores of each feature variable in different report types under different feature variable filters; In the same report type, sum up the multiple importance scores corresponding to each feature variable to obtain the single-table importance sum corresponding to each feature variable; Sum up and sort the single-table importance sums corresponding to the same feature variable in different report types to obtain the total-table importance sum and the corresponding sorting result; In the sorting result, select the first second preset number of candidate feature variables from the candidate feature variables ranked in the top first preset number, and determine them as the target feature variables. The sorting result is sorted according to the rule of the total-table importance sum from large to small, and the second preset number is equal to the target number of feature variables.

[0013] Through the above embodiments, the risk assessment system calculates the importance scores of feature variables in different report types through the feature variable filter, combines the importance scores of each report type to obtain the total importance sum of the summary table, sorts them, and then screens out the feature variables that have the greatest impact on the prediction output, optimizing the performance of the model and reducing the unnecessary computational burden, enabling the model to more effectively process a large amount of data and improve the operation efficiency.

[0014] In some embodiments, the step of standardizing the sample data of the target feature variable according to the positive and negative formula of the credit risk indicator to obtain the training sample data set specifically includes: Determine the orientation category of each target feature variable according to the credit risk indicator orientation category table, where the orientation category includes positive and negative, and the credit risk indicator orientation category table records the orientation categories corresponding to different target feature variables; Standardize the sample data of the target feature variable according to the positive and negative formula of the credit risk indicator to obtain standardized sample data, where the standardized sample data includes positive index values and negative index values, and the positive index standardization formula is: ; The negative index standardization formula is: ; Among them, is the positive index value, is the negative index value, is the sample data mean, is the sample data maximum value, is the sample minimum value; Construct the training sample data set according to the standardized sample data and the corresponding historical credit risk assessment results.

[0015] Through the above embodiments, the risk assessment system standardizes the data according to the positive and negative of the credit risk indicator, ensuring the consistency and comparability of the data in the training sample data set, and then helping the model to more accurately identify and interpret the impacts of various risk indicators, improving the prediction accuracy and reliability of the model.

[0016] In some embodiments, the step of using the multiple balanced sample subsets to train the classifiers Random Forest, XGBoost, and LightGBM respectively and integrating them to obtain a single-layer mean integration model specifically includes: Divide the balanced sample subset into a training set and a test set according to a preset ratio; The training set is used to fit the classifiers Random Forest, XGBoost, and LightGBM respectively to obtain the prediction results corresponding to each classifier, and each of these classifiers is trained through 5-fold cross-validation. The test results are averaged to obtain the integrated prediction results of the single-layer mean integration model.

[0017] Through the above embodiments, the risk assessment system uses multiple balanced sample subsets to train and test classifiers, and integrates the prediction results of each classifier through an integration method. This integration method can balance the biases of individual models when dealing with complex and variable financial data, improving the stability and accuracy of the overall prediction system.

[0018] In some embodiments, after the step of averaging the test results to obtain the integrated prediction results of the single-layer mean integration model, the following steps are further included: Under different classifier integration methods, calculate the evaluation metrics of each classifier on the balanced sample subset. The classifier integration methods include Stacking, Blending, and SLME, and the evaluation metrics include accuracy, precision, recall, F1 score, and AUC. By comparing the evaluation metrics under different classifier integration methods, determine the optimal classifier integration method and use it for subsequent integration of the classifiers.

[0019] Through the above embodiments, the risk assessment system can select the optimal classifier integration strategy by comparing the evaluation metrics under different classifier integration methods, further enhancing the robustness and accuracy of the model in practical applications.

[0020] In a second aspect, the present application provides a risk assessment system, which includes: one or more processors and a memory; The memory is coupled to the one or more processors. The memory is used to store computer program code, and the computer program code includes computer instructions. The one or more processors call the computer instructions so that the risk assessment system can implement a method for evaluating the credit risk of listed companies based on an integrated model provided in the above embodiments, which will not be elaborated here.

[0021] In a third aspect, the present application provides a computer-readable storage medium, including instructions, which when running on the risk assessment system, enable the risk assessment system to implement a method for evaluating the credit risk of listed companies based on an integrated model provided in the above embodiments, which will not be elaborated here.

[0022] Fourthly, the present application provides a computer program product. When the computer program product runs on a risk assessment system, the risk assessment system can implement a method for assessing the credit risk of listed companies based on an integrated model provided in the above embodiments, which will not be elaborated here.

[0023] One or more technical solutions provided in the embodiments of the present application have at least the following technical effects or advantages: 1. The risk assessment system subdivides financial data into four types: quarterly reports, semi-annual reports, three-quarter reports, and annual reports. This method improves the model's ability to capture financial performance in different time periods. By combining the chi-square filter and the F-test with a feature variable selector to calculate and rank importance scores, the system can accurately screen out feature variables that have a significant impact on predicting credit risk. This fine-grained data processing and optimized screening of feature variables not only improve the accuracy of credit risk assessment but also enhance the timeliness of the model in dealing with seasonal and phased financial changes.

[0024] 2. The risk assessment system trains and tests multiple classifiers including Random Forest, XGBoost, and LightGBM using multiple balanced sample subsets, and integrates the prediction results of each classifier through a single-layer mean ensemble model. This ensemble learning method effectively integrates the advantages of different models, improves the stability and accuracy of the overall prediction system in dealing with complex and variable financial data by balancing the biases of each single model. In addition, the system further enhances the robustness and adaptability of the model in practical applications by comparing evaluation metrics under different ensemble methods and selecting the optimal classifier ensemble strategy.

[0025] 3. The risk assessment system performs random sampling with replacement on the imbalanced training sample dataset through the random downsampling method to obtain 9 balanced sample subsets, making the number of samples in the positive and negative classes in the balanced sample subsets balanced, thereby improving the accuracy and reliability of model prediction. BRIEF DESCRIPTION OF THE DRAWINGS

[0026] Figure 1 is a flowchart of a method for assessing the credit risk of listed companies based on an integrated model in an embodiment of the present application; Figure 2 is another flowchart of a method for assessing the credit risk of listed companies based on an integrated model in an embodiment of the present application; Figure 3 (a) is a curve showing the change of chi-square test scores of quarterly reports from 2005 to 2010 with the increase of feature variables in an embodiment of the present application; Figure 3 (b) is a curve showing the change of chi-square test scores of quarterly reports from 2011 to 2016 with the increase of feature variables in an embodiment of the present application;Figure 3 (c) is a curve showing the change of the chi-square test score of the first quarterly report with the increase of feature variables from 2017 to 2022 in the embodiment of the present application; Figure 4 are curves showing the change of the F-test score of 4 types of financial statements with the increase of feature variables from 2017 to 2022 in the embodiment of the present application; Figure 5 is an exemplary schematic diagram of the importance ranking result of the annual report feature variables of the feature selector in the embodiment of the present application; Figure 6 (a) is an exemplary schematic diagram of the total importance of feature variables in the embodiment of the present application; Figure 6 (b) is an exemplary schematic diagram of the credit risk index system of listed enterprises in the embodiment of the present application; Figure 7 is a schematic diagram of the entity device structure of the risk assessment system in the embodiment of the present application. Detailed implementation manners

[0027] The terms used in the following embodiments of the present application are only for the purpose of describing specific embodiments, and are not intended to limit the present application. As used in the specification and appended claims of the present application, the singular forms "a", "an", "the", "above-mentioned", "said", and "this" are also intended to include the plural forms, unless the context clearly indicates otherwise. It should also be understood that the term "and / or" used in the present application refers to any or all possible combinations including one or more of the listed items.

[0028] Hereinafter, the terms "first" and "second" are only used for descriptive purposes, and cannot be understood as implying or indicating relative importance or implicitly specifying the quantity of the indicated technical features. Thus, the features defined with "first" and "second" may explicitly or implicitly include one or more of such features. In the description of the embodiments of the present application, unless otherwise stated, the meaning of "a plurality" is two or more.

[0029] For ease of understanding, the method provided in this embodiment is described in the following process. Please refer to Figure 1 , which is a schematic flow chart of a method for assessing the credit risk of listed companies based on an integrated model in the embodiment of the present application.

[0030] S101. Divide the financial statement data of the listed company according to 4 types of financial statements to obtain sample data corresponding to different types of financial statements.

[0031] The risk assessment system first collects and preprocesses the financial data of the target listed company. Specifically, the target listed company can upload the financial statement data of its own company to the risk assessment system, and then the risk assessment system divides the received financial statement data according to seasons (including the first quarterly report, the second quarterly report, the third quarterly report, and the fourth quarterly report) to obtain the sample data corresponding to different statement types, namely the first quarterly report, the interim report, the third quarterly report, and the annual report.

[0032] Among them, the financial data of the target listed company includes, but is not limited to, index categories such as company information, risk identification, per-share indicators, profitability, solvency, growth ability, operation ability, cash flow, dividend ability, capital structure, and earnings quality. Each index category can include one or more characteristic scalar quantities. For example, in a specific embodiment, the index category "growth ability" includes multiple characteristic variables such as "year-on-year growth of total operating costs", "year-on-year growth of total liabilities", and "year-on-year growth of net cash flow". Another example is that the index category "per-share indicators" includes multiple characteristic variables such as "earnings per share (TTM)" and "total operating revenue per share", which will not be elaborated here.

[0033] S102. Perform chi-square filtering and F-test on all characteristic variables of the sample data respectively to determine the number of characteristic variables to be selected.

[0034] Specifically, the risk assessment system performs chi-square filtering and F-test on the sample data of different statement types respectively to obtain the chi-square test scores and F-test scores under different characteristic variable quantity values in different years; then, by analyzing the changing trends of the chi-square test scores and F-test scores with the characteristic variable quantity values, respectively determine the characteristic quantity values corresponding to the top 3 highest chi-square test scores and the frequency of occurrence of this characteristic quantity value, and the characteristic quantity values corresponding to the top 3 highest F-test scores and the frequency of occurrence of this characteristic quantity value, so as to obtain the optimal characteristic quantity value occurrence frequency table, and then determine the target characteristic variable quantity according to the optimal characteristic quantity value occurrence frequency table.

[0035] In a specific embodiment, the optimal characteristic quantity value occurrence frequency table is as follows: ; It should be noted that in the embodiment of the present application, random forest can be selected as the classification algorithm for chi-square filtering to filter the sample data, obtain the chi-square values of all samples and the p-values of all samples, and then use the K value and p value for filtering, and select the cross-validation evaluation method to perform multiple cross-validations, and finally select the average score as the tool for statistically scoring the sample labels before and after filtering. The chi-square distribution ( distribution) is defined as follows: , where, is the observed frequency, is the expected frequency, r is the number of rows, and c is the number of columns. is the chi-square test score. The larger it is, the greater the difference between the observed value and the theoretical value. When is greater than the preset critical value, a conclusion of statistical significance can be obtained.

[0036] The calculation formula for the F-test score is: ; ; ; ; ; ; where is the j-th sample of the i-th category, is the sample mean of the j-th category, is the total sample mean, and SST, SSE, and SSB are all intermediate variables.

[0037] In the above embodiment, the risk assessment system determines the chi-square test score and the F-test score through chi-square filtering, and determines the number of feature variables participating in model training according to the variation laws of the two with the number of feature variable values in different report types and in different years of the same type, so that the number of feature variables selected to participate in model training is more representative, enabling the model to better predict the credit risk of listed companies based on a small number of feature variables, improving the reliability and efficiency of the prediction.

[0038] S103. Use a preset feature variable filter to calculate and rank the importance scores of all feature variables, and determine the target feature variables according to the number of feature variables.

[0039] The risk assessment system further optimizes the screening of feature variables by introducing feature selection methods in machine learning. Optional feature selection algorithms include feature importance evaluation based on tree models, such as Random Forest, XGBoost, GBDT, and LightGBM, etc. During the training process, these algorithms can automatically calculate the contribution degree of each feature to the prediction target. The higher the contribution degree of a feature, the greater its impact on the performance of the model, and the higher its importance.

[0040] Specifically, the risk assessment system uses a variety of mainstream tree models as feature variable filters to perform importance scoring on all candidate feature variables. Taking Random Forest as an example, the risk assessment system inputs all feature variables into a pre-trained Random Forest model, and the model calculates an importance score for each feature variable. This score synthesizes the contribution of the feature on all decision trees in the forest, and the higher the value, the stronger the discriminative power of the feature for the entire forest. The risk assessment system performs importance scoring on all feature variables of the same report respectively based on different feature variable filters, and sorts them based on the importance scores of each feature variable in different feature variable filters. Specifically, as Figure 5 shown, it is an exemplary schematic diagram of the importance ranking result of the feature variables in the annual report of the feature selector in the embodiment of the present application. It specifically shows the importance scores of different feature variables under different feature variable filters (Random Forest, GBDT, XGBoost, and LightGBM), as well as the comprehensive importance scores corresponding to each feature variable calculated based on the importance scores, and sorts them according to the comprehensive importance scores.

[0041] Furthermore, the risk assessment system selects the feature variables with the number determined in step S103 from the sorted result of the comprehensive importance scores in descending order and sets them as target feature variables, and constructs an index system for evaluating the credit risk of listed companies. Specifically, as Figure 6 shown in (a) below, it is an exemplary schematic diagram of the total importance of feature variables in the embodiment of the present application. Given that the number of feature variables determined in step S103 is 26, the risk assessment system selects Figure 5 26 feature variables from the corresponding table from top to bottom and sets them as target feature variables, and then constructs an index system for evaluating the credit risk of listed companies based on the target feature variables. Specifically, as Figure 6 shown in (b) is an exemplary schematic diagram of the credit risk index system of listed enterprises in the embodiment of the present application.

[0042] S104. Standardize the sample data of the target feature variables according to the positive and negative calculation formula of the credit risk index to obtain a training sample data set.

[0043] Specifically, for each feature variable (i.e., target feature variable) in the credit risk index system of listed enterprises, the risk assessment system determines the orientation category of each target feature variable according to the preset credit risk index orientation category table.

[0044] For example, in a specific embodiment, the credit risk index orientation category table is as follows: ; Based on the above table, the risk assessment system can determine the directional category of each target feature variable, and then standardize the sample data of each target feature variable according to the positive and negative calculation formula of the credit risk indicator to obtain the standardized sample data, and construct the training sample data set based on the standardized sample data and the corresponding historical credit risk assessment results.

[0045] It should be noted that the historical credit risk assessment results include two types of listing statuses, namely, the financial condition is at risk (represented by "1") and the financial condition is normal (represented by "0").

[0046] In the above embodiment, the risk assessment system standardizes the data according to the positive and negative of the credit risk indicator, ensuring the consistency and comparability of the data in the training sample data set, and then helping the model to more accurately identify and interpret the impacts of various risk indicators, improving the prediction accuracy and reliability of the model.

[0047] S105. If it is detected that the training sample data set is unbalanced, use the balanced Bagging classifier based on decision tree to perform random downsampling on the training sample data set to obtain multiple balanced sample subsets.

[0048] Based on the credit risk indicator directional category table in step S104, the risk assessment system can respectively count the proportion of the positive indicator value set and the negative indicator value set in the training sample data set. If the difference between the ratios of the two exceeds the preset ratio threshold (such as 10), it is determined that the training sample data set is unbalanced.

[0049] Furthermore, the risk assessment system can use the balanced Bagging classifier based on decision tree to perform random downsampling on the training sample data set, and generate multiple balanced sample subsets through random sampling with replacement, so that the proportions of the positive indicator value set and the negative indicator value set in the sampled balanced sample subsets are equal.

[0050] Among them, to ensure that the sample data volume of all balanced sample subsets is close to the sample data volume of the initial training sample data set, therefore, the number of selected balanced sample subsets is 9, and the sample ratio of the training set and the test set for training is 8:2.

[0051] S106. Use multiple balanced sample subsets to train the classifiers Random Forest, XGBoost, and LightGBM respectively, and integrate them to obtain a single-layer mean integration model.

[0052] Specifically, for each balanced sample subset, the risk assessment system divides it into a training set and a test set according to a preset ratio (such as 8:2). The training set is used to train the classifier, and the test set is used to evaluate the performance of the classifier. The risk assessment system uses the training set to perform fitting training on Random Forest, XGBoost, and LightGBM respectively, and optimizes the model hyperparameters through 5-fold cross-validation to make the fitting effect of the model optimal on the training set. Among them, cross-validation can effectively avoid the problem of model overfitting and improve the generalization ability of the model. Finally, the risk assessment system integrates the three trained classifiers to obtain a single-layer mean ensemble model SLME (Single layer mean ensemble model), that is, the prediction results of the three trained classifiers are averaged to obtain the prediction result of the single-layer mean ensemble model.

[0053] In the above embodiment, the risk assessment system uses multiple balanced sample subsets to train and test the classifier, and integrates the prediction results of each classifier through an ensemble method. This ensemble method can balance the biases of each single model when dealing with complex and variable financial data, and improve the stability and accuracy of the overall prediction system.

[0054] S107. Input the financial statement data of the listed company to be evaluated into the single-layer mean ensemble model to obtain the corresponding credit risk assessment result.

[0055] After the risk assessment system obtains the financial statement data of the listed company to be evaluated, it standardizes the statement data and constructs a balanced sample data set. Then, it inputs the balanced sample data set into each classifier in the single-layer mean ensemble model to obtain the prediction results output by each classifier. Then, these prediction results are simply averaged to obtain the prediction output of the single-layer mean ensemble model on this balanced sample subset, that is, the credit risk assessment result.

[0056] In the above embodiment, the risk assessment system evaluates by dividing the financial data into 4 report types (first-quarter report, mid-year report, third-quarter report, and annual report), enabling the model to more accurately capture the financial performance in different time periods. Further, through chi-square filtering, F-test, and using a feature variable filter to calculate and sort the importance scores, the feature variables that have a greater impact on the prediction results are screened out, and an equilibrium data set is constructed based on the financial data corresponding to the feature variables to train the single-layer mean ensemble model. The final credit risk assessment result is determined by integrating the prediction results of multiple classifiers. This method improves the accuracy and timeliness of credit risk assessment of listed companies by refining the financial statement data to the quarterly level and optimizing the screening of feature variables for model training.

[0057] The following is a further and more specific process description of the method provided in this embodiment. Please refer to Figure 2 , which is another process schematic diagram of a listed company credit risk assessment method based on an integrated model in the embodiments of the present application.

[0058] S201. Perform chi-square filtering and F-test on all feature variables of the sample data respectively to obtain the chi-square test scores and F-test scores of different statement types under different feature variable quantity values in different years.

[0059] Specifically Figure 3 As shown in (a), (b), and (c) respectively, they are the change curves of the chi-square test scores of the first-quarter reports in the embodiments of the present application with the increase of feature variables from 2005 to 2010; the change curves of the chi-square test scores of the first-quarter reports in the embodiments of the present application with the increase of feature variables from 2011 to 2016; the change curves of the chi-square test scores of the first-quarter reports in the embodiments of the present application with the increase of feature variables from 2017 to 2022; Among them, the abscissa represents the number of feature variables, the ordinate represents the chi-square test score, and the color of the curve represents the corresponding year. For example, the green curve represents the change curve of the chi-square test score with the increase of feature variables in 2005. By Figure 3 It is possible to determine the change trend of the chi-square test score with the increase of feature variables for different statement types and in different years, and then determine the range where the highest score appears. For example, as Figure 3 shown in (a), it can be seen that when there are 0 - 20 feature variables, the chi-square test score rises with the increase of variables. When there are 20 - 30 feature variables, the highest score appears in most years, and the increase in the chi-square test score fluctuates near the peak. When there are 30 - 80 feature variables, the chi-square test score begins to decline with the increase of feature variables and shows a fluctuating decline. Therefore, when selecting the number of feature variables, the general range is between 20 - 30. Similarly, the number of feature variables corresponding to the peak of other statement types can be determined, which will not be elaborated here.

[0060] Similar to chi-square filtering, calculate the F-test score and determine the number of feature variables corresponding to the peak. Specifically, as Figure 4 shown, they are the change curves of the F-test scores of 4 statement types in the embodiments of the present application with the increase of feature variables from 2017 to 2022 respectively, which will not be elaborated here.

[0061] S202. Respectively count the top 3 feature quantity values with the highest chi-square test scores and F-test scores and the corresponding frequencies that appear in different years for different statement types to obtain the optimal feature quantity value frequency table.

[0062] For each report type, the risk assessment system counts the chi-square test scores and F-test scores obtained in step S201 by year. For each year, the risk assessment system finds the top 3 feature quantities with the highest chi-square scores and their corresponding occurrence frequencies; at the same time, it also finds the top 3 feature quantities and frequencies with the highest F-test scores. In this way, the risk assessment system obtains a table of the occurrence frequencies of the optimal feature quantity values. For a specific table example, see step S102, which will not be elaborated here.

[0063] S203. Determine the feature quantity value with the highest frequency in the chi-square test as the candidate feature quantity value.

[0064] After obtaining the table of the occurrence frequencies of the optimal feature quantity values, the risk assessment system selects a best feature quantity from it as the feature dimension for constructing the subsequent risk assessment model. Considering that chi-square filtering and F-test are two different statistical methods, there may be certain differences in their evaluation results of feature importance. To take into account the evaluation results of both methods as much as possible, the risk assessment system selects the feature quantity with the highest occurrence frequency in the chi-square filtering score as the candidate feature quantity value.

[0065] For example, in the table of the occurrence frequencies of the optimal feature quantity values in step S102, the chi-square filtering result shows that the occurrence frequency of "26 features" is the highest, reaching 8 times; the second-ranked "23 features" is 6 times, and the third-ranked " features" is 5 times. Then the risk assessment system takes "26 features" as the candidate feature quantity value.

[0066] S204. Determine the target frequency corresponding to the candidate feature quantity value in the F-test according to the table of the occurrence frequencies of the optimal feature quantity values.

[0067] After selecting the feature quantity with the highest occurrence frequency in the chi-square filtering as the candidate value, the risk assessment system evaluates the performance of this feature quantity in the F-test to further verify whether it has sufficient stability and representativeness.

[0068] Specifically, the risk assessment system finds the record in the table of the occurrence frequencies of the optimal feature quantity values that is the same as the candidate feature quantity value in the F-test result. If such a record exists, take its corresponding occurrence frequency as the "target frequency"; if there is no exactly the same feature quantity value, select a feature quantity closest to the candidate value and take its frequency as the "target frequency".

[0069] S205. When the target frequency is within the preset frequency range, determine the candidate feature quantity value as the target feature variable quantity.

[0070] After determining the target frequency corresponding to the number of candidate features to be filtered by chi-square in the F-test, the risk assessment system further determines whether this target frequency is high enough to prove the stability and reliability of the candidate feature combination. Among them, the preset frequency range is a reasonable range of occurrence frequencies set by relevant technical personnel on the risk assessment system based on experience and practice. If the target frequency falls within this range, it is determined that the candidate feature combination is not only the best in chi-square filtering but also highly recognized in the F-test, and it is a robust choice; conversely, if the target frequency is lower than this range, it can be prompted in the form of voice or text message that the superiority of the candidate feature combination may not be significant enough and there is a certain degree of contingency.

[0071] For example, assume that the preset frequency range of the risk assessment system is 3 - 5 times. If in step S204, the target frequency corresponding to the number of candidate features "26 features" is exactly 4 times, then the system determines that this feature combination is stable and reliable and determines it as the final number of target feature variables. However, if the target frequency is only 1 time, the system will generate a warning, indicating that the performance of this candidate value in the F-test is not ideal and it may be necessary to reconsider the result of chi-square filtering or further adjust the setting of the preset frequency range.

[0072] In the above embodiment, the risk assessment system creates a frequency table of the occurrence of the optimal number of feature values, verifies and compares the chi-square filtering result with the F-test result, making the number of feature variables participating in model training more representative and improving the reliability of prediction.

[0073] S206. Calculate the importance scores and sort all feature variables using a preset feature variable filter.

[0074] This step has been explained in step S103 and will not be elaborated here.

[0075] S207. Among the sorting results, select the first second preset number of candidate feature variables from the candidate feature variables ranked in the top first preset number to determine the target feature variables.

[0076] After determining the number of feature variables, the risk assessment system determines the first preset number based on a preset rule, and then screens out the first second preset number (the number of feature variables) of candidate feature variables from the sorting results from top to bottom. For example, if the number of feature variables is 26, the first preset number can be set to 30 or 40 for display, thus reducing the data processing volume of the system.

[0077] In the above embodiments, the risk assessment system calculates the importance scores of feature variables in different report types through a feature variable filter, combines the importance scores of each report type to obtain the total importance sum of the master table, sorts them, and then screens out the feature variables that have the greatest impact on the prediction output, optimizing the performance of the model and reducing the unnecessary computational burden, enabling the model to more effectively process a large amount of data and improve the operation efficiency.

[0078] S208. Construct a balanced sample subset based on the target feature variables to train a single-layer mean ensemble model for credit risk assessment.

[0079] This step has been explained in steps S104 to S107 and will not be elaborated here.

[0080] S209. Calculate the evaluation metrics of each classifier on the balanced sample subset under different classifier ensemble methods.

[0081] In the previous steps, the risk assessment system has trained multiple base classifiers, including Random Forest, XGBoost, and LightGBM, using the balanced sample subset. To further improve the accuracy of risk assessment, the risk assessment system integrates the prediction results of these classifiers to reduce the bias and variance of a single learner, thereby obtaining a more robust and stable prediction effect.

[0082] Furthermore, in the embodiments of the present application, the risk assessment system integrates the prediction results of the above three classifiers, namely Random Forest, XGBoost, and LightGBM, using multiple ensemble methods (including but not limited to Stacking, Blending, and SLME) respectively to determine the final risk assessment result. Among them, Stacking is an ensemble learning method based on model stacking. The prediction results of each base classifier are used as a new feature matrix, and then a secondary learner (usually logistic regression) is trained to make the final prediction decision; Blending is an ensemble method based on linear weighting of prediction results. It performs a weighted average on the prediction results of each base classifier to obtain the final prediction probability, where the weight of each classifier is determined by its performance on the validation set; SLME is the single-layer mean ensemble model used in the present application and will not be elaborated here.

[0083] Under each ensemble method, the risk assessment system uses the balanced sample subset to evaluate the performance of each base classifier. Specifically, the following five commonly used classification model evaluation metrics are calculated, including accuracy, precision, recall, F1 score, and AUC.

[0084] S210. By comparing the evaluation metrics under different classifier ensemble methods, determine the optimal classifier ensemble method and use it for subsequent ensemble of classifiers.

[0085] After obtaining the evaluation metrics of each base classifier under different ensemble strategies, the risk assessment system selects an optimal ensemble method from them for subsequent ensemble of classifiers. For example, in a specific embodiment, by calculating the evaluation metrics, it can be known that the results of SLME are very close to those of Stacking in terms of accuracy, precision, recall, and F1 score, but SLME is superior to Stacking in terms of the value of AUC. Therefore, compared with Stacking, the ensemble method of SLME is significantly better, so SLME is selected to ensemble the prediction results of classifiers.

[0086] In the above embodiment, the risk assessment system can select the optimal classifier ensemble strategy by comparing the evaluation metrics under different classifier ensemble methods, further enhancing the robustness and accuracy of the model in practical applications.

[0087] The risk assessment system of the embodiment of the present invention is applied to an electronic device. Figure 7 The schematic diagram of the architecture of the electronic device suitable for implementing the embodiment of the present invention is shown.

[0088] It should be noted that Figure 7 The shown electronic device is only an example and should not bring any limitation to the functions and usage scope of the embodiment of the present invention.

[0089] Those of ordinary skill in the art can understand that all or part of the steps in the various methods of the above embodiments can be completed by instructions (computer programs), or the relevant hardware can be controlled by instructions (computer programs). The instructions can be stored in a computer-readable storage medium and loaded and executed by a processor. The electronic device of this embodiment includes a storage medium and a processor. Among them, multiple instructions are stored in the storage medium, and the instructions can be loaded by the processor to execute any step of the method provided by the embodiment of the present invention.

[0090] Specifically, the storage medium and the processor are electrically connected directly or indirectly to achieve data transmission or interaction. For example, these components can be electrically connected to each other through one or more signal lines. The storage medium stores computer-executable instructions for implementing the data access control method, including at least one software function module that can be stored in the storage medium in the form of software or firmware. The processor executes various functional applications and data processing by running the software programs and modules stored in the storage medium. The storage medium can be, but is not limited to, a random access memory (RAM), a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), etc. Among them, the storage medium is used to store programs, and the processor executes the programs after receiving the execution instructions.

[0091] Furthermore, the software programs and modules in the above storage medium may further include an operating system, which may include various software components and / or drivers for managing system tasks (such as memory management, storage device control, power management, etc.), and may communicate with various hardware or software components to provide a running environment for other software components. The processor can be an integrated circuit chip with signal processing capabilities. The above processor can be a general-purpose processor, including a central processing unit (CPU), a network processor (NP), etc., which can implement or execute the various methods, steps, and logic flow block diagrams disclosed in this embodiment. The general-purpose processor can be a microprocessor or the processor can also be any conventional processor, etc.

[0092] Since the instructions stored in the storage medium can execute the steps in any of the methods provided in the embodiments of the present invention, the beneficial effects of any of the methods provided in the embodiments of the present invention can be achieved. For details, see the previous embodiments and will not be repeated here.

[0093] As described above, it is only the preferred specific embodiment of the present invention, but the protection scope of the present invention is not limited thereto. Any changes or substitutions that can be easily thought of by those skilled in the art within the technical scope disclosed by the present invention should be covered within the protection scope of the present invention. Therefore, the protection scope of the present invention should be subject to the protection scope of the claims.

Claims

1. A credit risk assessment method for listed companies based on an integrated model, applied to a risk assessment system, characterized in that: The method comprises: The financial statement data of listed companies are divided into four report types to obtain sample data corresponding to different report types, wherein the report types include quarterly report, mid-year report, quarterly report and annual report, and the sample data includes multiple indicator categories and one or more characteristic variables corresponding to each indicator category; Performing chi-square filtering and F-test on all characteristic variables of the sample data to determine the number of characteristic variables to be selected; Use a preset feature variable filter to calculate and sort the importance scores of all the feature variables, and determine the target feature variables according to the number of the feature variables. The feature variable filter includes Random Forest, XGBoost, GBDT and LightGBM. The target feature variables are used to construct an indicator system for evaluating the credit risk of listed companies; Standardizing the sample data of the target characteristic variable according to the credit risk indicator positivity calculation formula to obtain a training sample data set, wherein the credit risk indicator positivity calculation formula includes a positive indicator standardization formula and a negative indicator standardization formula; If it is detected that the training sample data set is unbalanced, a balanced Bagging classifier based on a decision tree is used to perform random downsampling processing on the training sample data set to obtain multiple balanced sample subsets; Using the multiple balanced sample subsets to train classifiers Random Forest, XGBoost and LightGBM respectively, and integrating them to obtain a single-layer mean ensemble model; The financial statement data of the listed company to be evaluated is input into the single-layer mean integration model to obtain the corresponding credit risk assessment results, which include two types: financial status at risk and financial status normal.

2. The method according to claim 1, characterized in that The step of performing chi-square filtering and F-test on all characteristic variables of the sample data to determine the number of characteristic variables to be selected specifically includes: Perform chi-square filtering and F-test on all characteristic variables of the sample data to obtain chi-square test scores and F-test scores for different report types under different characteristic variable quantity values ​​in different years; The top three feature quantity values ​​and corresponding frequencies of the chi-square test scores and the F-test scores that appeared in different report types and different years are counted respectively to obtain a frequency table of the optimal feature quantity values; The target feature variable quantity is determined according to the occurrence frequency table of the optimal feature quantity value.

3. The method according to claim 2, characterized in that The step of determining the number of target feature variables according to the occurrence frequency table of the optimal feature quantity value specifically includes: The characteristic quantity value with the highest frequency in the chi-square test is determined as the characteristic quantity value to be selected; Determine the target frequency corresponding to the feature quantity value to be selected in the F test according to the frequency table of occurrence of the optimal feature quantity value; If the target frequency is within the preset frequency range, the value of the number of features to be selected is determined as the target number of feature variables; If the target frequency is not within the preset frequency range, a warning is issued on the accuracy of the chi-square test.

4. The method according to claim 1, characterized in that: The step of using a preset feature variable filter to calculate and sort the importance scores of all the feature variables, and determining the target feature variable according to the number of the feature variables, specifically includes: Based on the preset characteristic variable filters, importance scores of characteristic variables in different report types are calculated respectively to obtain importance scores of each characteristic variable in different report types under different characteristic variable filters; In the same report type, multiple importance scores corresponding to each feature variable are summed to obtain the sum of the importance of each feature variable in a single table. Sum and sort the importance sums of the single tables corresponding to the same characteristic variable of different report types to obtain the importance sum of the total table and the corresponding sorting results; In the sorting result, the first second preset number of candidate feature variables are selected from the first preset number of candidate feature variables and determined as target feature variables. The sorting result is sorted from large to small according to the rule of the total importance of the total table, and the second preset number is equal to the number of target feature variables.

5. The method according to claim 1, characterized in that The step of performing standardization processing on the sample data of the target characteristic variable according to the credit risk indicator positivity calculation formula to obtain a training sample data set specifically includes: Determining the tropism category of each target characteristic variable according to the credit risk indicator tropism category table, wherein the tropism category includes positive and negative, and the credit risk indicator tropism category table records the tropism categories corresponding to different target characteristic variables; The sample data of the target characteristic variable is standardized according to the credit risk indicator positivity calculation formula to obtain standardized sample data, wherein the standardized sample data includes positive indicator values ​​and negative indicator values, and the positive indicator standardization formula is: ; The negative indicator standardization formula is: ; in, is the positive indicator value, is a negative indicator value. is the sample data mean, is the maximum value of the sample data, is the minimum value of the sample; The training sample data set is constructed based on the standardized sample data and the corresponding historical credit risk assessment results.

6. The method according to claim 1, characterized in that The step of using the multiple balanced sample subsets to respectively train the classifiers Random Forest, XGBoost and LightGBM, and integrating them to obtain a single-layer mean ensemble model specifically includes: Dividing the balanced sample subset into a training set and a test set according to a preset ratio; The classifiers Random Forest, XGBoost and LightGBM are fitted using the training set to obtain the prediction results corresponding to each classifier. Each classifier is trained by 5-fold cross validation. The test results are averaged to obtain the integrated prediction results of the single-layer mean integrated model.

7. The method according to claim 6, characterized in that After the step of averaging the test results to obtain the integrated prediction results of the single-layer mean integrated model, the method further includes: Calculate the evaluation index of each classifier on the balanced sample subset under different classifier integration modes, the classifier integration modes include Stacking, Blending and SLME, and the evaluation index includes accuracy, precision, recall, F1 score and AUC; By comparing the evaluation indicators under different classifier integration methods, the optimal classifier integration method is determined and used for subsequent integration of the classifiers.

8. A risk assessment system, characterized in that: The risk assessment system includes: one or more processors and memory; The memory is coupled to the one or more processors, and the memory is used to store computer program codes, wherein the computer program codes include computer instructions, and the one or more processors call the computer instructions to enable the risk assessment system to perform the method according to any one of claims 1 to 7.

9. A computer-readable storage medium comprising instructions, characterized in that: When the instructions are executed on a risk assessment system, the risk assessment system is caused to perform the method according to any one of claims 1 to 7.

10. A computer program product, characterized in that When the computer program product is run on a risk assessment system, the risk assessment system is caused to perform the method according to any one of claims 1 to 7.