Enterprise financial risk classification labeling method based on Transform improved algorithm
By combining Transformer and random forest algorithm, the corporate financial risk classification and annotation method is optimized, and the problem of poor performance in risk classification and annotation is solved in the existing technology, achieving higher classification accuracy and more efficient risk judgment.
Patent Information
- Application Number
- CN202510035521.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-09
- Publication Date
- 2025-05-06
AI Technical Summary
The existing technology is difficult to focus on domain knowledge in corporate financial risk classification annotation, general big models perform poorly, and random forest algorithms have limitations in processing text data.
Combining Transformer and random forest algorithm, by preprocessing text data, we build an enhanced feature matrix, remove the last full connection layer of the original Transformer model and introduce a new full connection layer, optimize and improve the output layer and model, use the Softmax function to calculate the classification probability distribution, and build an annotated data set training and optimization and improvement Transformer model.
It improves the accuracy and practicality of risk labeling classification of corporate financial business, significantly improves the efficiency of risk judgment, enhances the model's ability to capture domain-specific features, and reduces training costs.
Smart Images

Figure CN119939430A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of financial risk identification, and in particular to a method for classifying and labeling enterprise financial risks based on an improved Transformer algorithm. Background Art
[0002] Enterprise financial risk refers to the possibility of loss caused by various factors when an enterprise engages in financial activities. Specifically, enterprise financial risks mainly include: exchange rate risk. When an enterprise conducts cross-border transactions or holds foreign currency assets, fluctuations in exchange rates may lead to changes in asset values, resulting in losses. Interest rate risk. Changes in market interest rates will affect the borrowing costs and investment returns of enterprises, which may lead to unstable financial conditions. Credit risk refers to the possibility that the counterparty will not be able to fulfill its contractual obligations, which may cause the enterprise to face losses. Liquidity risk. When an enterprise needs funds, it cannot obtain the required funds in a timely manner at a reasonable cost, which may lead to financial difficulties. Market risk. Fluctuations in market prices, such as changes in the prices of financial instruments such as stocks and bonds, may affect the investment value and financing costs of enterprises. These risks not only affect the financial status of enterprises, but may also have a significant impact on their business strategies and long-term development. Therefore, enterprises need to reduce these risks through effective risk management measures to ensure financial stability and sustainable development.
[0003] In risk management, enterprise financial risk classification and labeling can help enterprises better understand and deal with different types of financial risks, thereby improving the efficiency and effectiveness of risk management. In the existing technology, although there are text processing methods based on large models, these general large models are difficult to focus on domain knowledge, and their performance in the specific field of enterprise financial business risk labeling is not ideal, and their deployment is also restricted by data privacy. On the other hand, although the random forest algorithm shows significant advantages in data classification, it is highly dependent on feature input and has limitations in its ability to process text data. Summary of the invention
[0004] In order to overcome the shortcomings of the prior art, the purpose of the present invention is to provide an enterprise financial risk classification and labeling method based on the improved Transformer algorithm. By combining the advantages of Transformer and random forest algorithms, the accuracy and practicality of enterprise financial business risk labeling and classification are improved, thereby significantly improving the efficiency of risk judgment.
[0005] To achieve the above object, the present invention provides the following solution: a method for classifying and labeling enterprise financial risks based on an improved Transformer algorithm, comprising the following steps:
[0006] Collecting text data related to the economic activities of the enterprise, preprocessing the text data, and obtaining a text input representation;
[0007] Encoding and decoding the text input representation to obtain a feature matrix, and constructing an enhanced feature matrix according to the importance of the feature matrix;
[0008] Removing the last fully connected layer of the original Transformer model and introducing a new fully connected layer to obtain an improved output layer and an improved Transformer model, and then optimizing the improved output layer and the improved Transformer model;
[0009] Inputting the enhanced feature matrix into the improved output layer, and calculating the classification probability distribution in combination with the Softmax function;
[0010] Construct a labeled data set, use the labeled data set to train and optimize the improved Transformer model to obtain a Transformer classification and labeling model, and use the Transformer classification and labeling model to classify and label corporate financial risks.
[0011] Optionally, text data related to the enterprise's economic activities are collected, and the text data are preprocessed to obtain a text input representation, including: collecting text data related to the enterprise's economic activities, performing data cleaning on the text data to remove invalid characters, standardize text formats and non-financial field vocabulary, and then performing word segmentation and word embedding operations on the text data to obtain a text input representation.
[0012] Optionally, encoding and decoding the text input representation to obtain a feature matrix, and then constructing an enhanced feature matrix according to the importance of the feature matrix, including:
[0013] Extracting a feature matrix in the last hidden state of the Transformer encoder and decoder layers, and normalizing the feature matrix to obtain a normalized feature matrix;
[0014] Constructing a financial risk category label corresponding to the text input representation, training a random forest model using the standardized feature matrix and the financial risk category label, and introducing a random seed and a dynamic parameter into the random forest model to obtain a feature importance model;
[0015] Based on the feature importance model, the Gini index is used to evaluate the importance of each feature to obtain an importance score vector, and then the first k features are selected as key features according to the importance score vector;
[0016] Using the index set of the key features, corresponding columns are extracted from the standardized feature matrix, and the extracted columns are integrated into an enhanced feature matrix.
[0017] Optionally, the last fully connected layer of the original Transformer model is removed, and a new fully connected layer is introduced to obtain an improved output layer and an improved Transformer model, and then the improved output layer and the improved Transformer model are optimized, including:
[0018] Removing the last fully connected layer of the original Transformer model and introducing a new fully connected layer for mapping the enhanced feature matrix to the category space of the financial risk category label;
[0019] A Dropout layer is added after the new fully connected layer to prevent overfitting, and during the model training process, the Dropout layer is used to randomly discard the outputs of some neurons to enhance the generalization ability of the model;
[0020] A BatchNormalization layer for stabilizing feature distribution is added after the new fully connected layer and the Dropout layer to normalize the small batch data to obtain an improved output layer and an improved Transformer model;
[0021] The weight matrix of the new connection layer is initialized using the original Transformer model, and the improved Transformer model is optimized using an objective function.
[0022] Optionally, the weight matrix can be expressed as
[0023] W new ∈R k×C
[0024] Among them, K is the number of features selected after feature selection and screening, and C is the number of categories of financial risk classification;
[0025] The expression of the objective function is:
[0026]
[0027] Among them, L task is the Dropout function, is a function, and λ is a regularization strength hyperparameter.
[0028] Optionally, the enhanced feature matrix is input into the improved output layer, and a classification probability distribution is obtained by combining with a Softmax function, including:
[0029] The expression of the improved output layer is:
[0030] S = BatchNorm(Dropout(X selected W new ))
[0031] The expression of the improved output layer combined with the Softmax function is:
[0032] P = softmax(S), P∈R n×C
[0033] Where n is the number of samples.
[0034] Optionally, constructing a labeled data set, using the labeled data set to train and optimize the improved Transformer model to obtain a Transformer classification and labeling model, and using the Transformer classification and labeling model to classify and label the financial risks of the enterprise, including:
[0035] Based on the collected text input representation, construct a labeled data set including a training set and a validation set, and perform desensitization processing on the labeled data set;
[0036] Use the cross entropy loss function to train the model, use the optimizer to adjust the model parameters, and during the training process, monitor the model loss value and accuracy index to evaluate and adjust the model. Use the learning rate decay strategy to control the learning rate change and complete the model training.
[0037] Data augmentation technology is used to increase the diversity of training samples and perform model optimization. During the optimization process, the validation set is used to evaluate the model. The hyperparameters of the model are adjusted according to the evaluation results to obtain the Transformer classification and labeling model for actual risk classification and labeling tasks.
[0038] The present invention discloses the following technical effects by providing a method for classifying and labeling enterprise financial risks based on an improved Transformer algorithm: By combining the advantages of both Transformer and random forest, the present invention not only improves the classification accuracy, but also enhances the model's ability to capture domain-specific features, making risk labeling more in line with actual business needs. In addition, the present invention also adopts a targeted output layer improvement optimization strategy, which significantly improves the convergence speed of the model and reduces the training cost.
[0039] The technical solution of the present invention is further described in detail below through the accompanying drawings and embodiments. BRIEF DESCRIPTION OF THE DRAWINGS
[0040] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative labor.
[0041] Figure 1 A schematic diagram of a method flow chart provided by an embodiment of the present invention;
[0042] Figure 2 A schematic diagram of a flow chart of an improved model output layer provided by an embodiment of the present invention;
[0043] Figure 3 A schematic diagram of ROC curve comparison before and after the model improvement provided in an embodiment of the present invention. DETAILED DESCRIPTION
[0044] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0045] In order to make the above-mentioned objects, features and advantages of the present invention more obvious and easy to understand, the present invention is further described in detail below with reference to the accompanying drawings and specific embodiments.
[0046] like Figure 1 As shown, the present invention provides a method for classifying and labeling enterprise financial risks based on the improved Transformer algorithm.
[0047] 1. The present invention comprises the following steps:
[0048] 1. Collect text data related to the enterprise's economic activities, pre-process the text data, and obtain text input representation. Including:
[0049] Collect text data related to corporate economic activities. These data can come from corporate announcements, financial reports, market research reports, judgment documents on economic disputes and third parties, etc., to ensure the diversity and comprehensiveness of the data to cover different financial risk categories and scenarios.
[0050] The text data is cleaned to remove invalid characters, standardize text formats and non-financial domain vocabulary, and then the text data is segmented and embedded to obtain a text input representation to better capture domain-related features.
[0051] 2. Encode and decode the text input representation to obtain a feature matrix, and construct an enhanced feature matrix based on the importance of the feature matrix. Including:
[0052] The feature matrix (H) is composed of the last hidden state of the Transformer encoder-decoder layer and contains rich text feature information. In order to effectively screen the feature labels of domain characteristics, a certain number of corporate finance-related texts and corresponding financial risk category labels (Y) are prepared. These labels are used to train the random forest model and serve as the basis for feature screening.
[0053] 2.1 Feature Standardization:
[0054] Extract the feature matrix H∈R from the last hidden state of the Transformer encoder-decoder layer n×d , where n is the sequence length and d is the hidden state dimension. H is normalized with zero mean and unit variance to obtain the standardized feature matrix Z, Z∈R n×d .
[0055] 2.2 Feature Importance Modeling
[0056] Construct a financial risk category label Y corresponding to the text input representation, and train a random forest model using the standardized feature matrix and the financial risk label, ensuring that the random forest input matrix and the number of rows n of the label correspond one to one. The results are reproducible by setting a random seed (such as the random_state parameter). Adjust parameters (dynamic parameters) to adapt to data characteristics (including the number of trees n_estimators and the maximum depth max_depth)
[0057] 2.3 Evaluating features and selecting key features
[0058] Based on the feature importance model, the Gini index is used to evaluate the importance of each feature and obtain the importance score vector IM∈R d , and then select the first k features as key features according to the importance score vector IM.
[0059] 2.4 Constructing the enhanced feature matrix
[0060] Using the index set of the key features, the corresponding columns are extracted from the standardized feature matrix, and the extracted columns are integrated into the enhanced feature matrix X selected ∈R n×k , this matrix will be used as input to improve the output layer of the Transformer.
[0061] 3. If Figure 2 As shown, the last fully connected layer of the original Transformer model is removed, and a new fully connected layer is introduced to obtain an improved output layer and an improved Transformer model, and then the improved output layer and the improved Transformer model are optimized.
[0062] 3.1 Remove the last fully connected layer of the original Transformer model and introduce a new fully connected layer for mapping the enhanced feature matrix to the category space of the financial risk category label; the weight matrix of the new fully connected layer is W new ∈R k×C , where K is the number of features selected after feature selection and screening, and C is the number of categories for financial risk classification.
[0063] 3.2 A Dropout layer is added after the new fully connected layer to randomly shield some neurons to prevent overfitting, and during the model training process, the Dropout layer is used to randomly discard the outputs of some neurons to enhance the generalization ability of the model.
[0064] 3.3 A BatchNormalization layer for stabilizing feature distribution is added after the new fully connected layer and the Dropout layer to normalize the small batch data to obtain an improved output layer and an improved Transformer model;
[0065] 3.4 Initialize the weight matrix of the new connection layer using the original Transformer model, and optimize the improved Transformer model using the objective function.
[0066] The expression of the weight matrix is:
[0067] W new ∈R k×C
[0068] Among them, K is the number of features selected after feature selection and screening, and C is the number of categories of financial risk classification;
[0069] The expression for initializing the weight matrix is:
[0070] W new ∈W pertrain
[0071] Among them, W pertrain The weights related to the output layer are extracted from the pre-trained model to initialize the new weight matrix. In this way, the weight initialization is more in line with the characteristics of the domain, which helps the model better adapt to the classification and identification of corporate financial risks.
[0072] The expression of the objective function is:
[0073]
[0074] Among them, L task is the Dropout function (which enhances the randomness of model learning and avoids excessive dependence between neurons), is the L2 function (controls the parameter amplitude and further optimizes the limit of model complexity), λ is the regularization strength hyperparameter, and the appropriate value can be selected through cross-validation.
[0075] 4. Input the enhanced feature matrix into the improved output layer, and calculate the classification probability distribution in combination with the Softmax function.
[0076] The expression of the improved output layer is:
[0077] S = BatchNorm(Dropout(X selected W new ))
[0078] The expression of the improved output layer combined with the Softmax function is:
[0079] P = softmax(S), P∈R n×C
[0080] Where n is the number of samples.
[0081] 5. Construct a labeled data set, use the labeled data set to train and optimize the improved Transformer model to obtain a Transformer classification and labeling model, and use the Transformer classification and labeling model to classify and label corporate financial risks.
[0082] 5.1 Dataset Construction
[0083] Based on the collected text input representation, a labeled data set including a training set and a validation set is constructed. The labeled data can be obtained through manual labeling or integration using existing labeling resources. The quality and diversity of the data set are ensured to cover different financial risk categories and scenarios. At the same time, in order to protect data privacy and security, the labeled data set is desensitized.
[0084] 5.2 Model Training
[0085] After constructing the data set, the cross entropy loss function is used to train the model. The cross entropy loss function is used to measure the difference between the model prediction result and the true label. The model parameters are adjusted through the optimizer (Adam) to minimize the loss function value. During the training process, indicators such as the model loss value and accuracy can be monitored to evaluate the performance of the model and make adjustments. At the same time, the learning rate decay strategy is used to control the learning rate change during the training process.
[0086] 5.3 Model Optimization
[0087] In order to further improve the performance of the model, L2 regularization is used to prevent overfitting, and the training process is controlled by adjusting the learning rate and learning rate decay strategy. In addition, data augmentation technology is used to increase the diversity of training samples to improve the generalization ability of the model. Data augmentation technology can include text replacement, synonym replacement, sentence reorganization and other methods. In the process of model optimization, the trained model is evaluated using the validation set data, and the hyperparameters of the model are adjusted according to the evaluation results, including but not limited to the learning rate, Dropout rate, parameters of the BatchNormalization layer, and the weight matrix of the new fully connected layer in the Transformer output layer. After multiple iterations of tuning, the optimal model parameters and structure are obtained for the actual risk classification and labeling task, completing the construction of the Transformer classification and labeling model.
[0088] 2. Take the text content published by a certain enterprise as an example
[0089] Summary of the text (full text omitted): A listed manufacturing company recently issued an announcement stating that due to the double blow of rising raw material prices, the company's operating costs have risen sharply, coupled with weak market demand, resulting in a sharp decline in sales. In order to maintain operations, the company has to borrow a lot of money, but its financial situation continues to deteriorate, and it is now unable to repay some short-term debts on time. At the same time, the company's announcement also mentioned that due to the increasing uncertainty of the market environment, the company's future market performance is also subject to great uncertainty.
[0090] 1. Labeling results (using the method of the present invention):
[0091] Credit risk: 0.1
[0092] Liquidity risk: 0.85
[0093] Market risk: 0.03
[0094] Operational risk: 0.01
[0095] Legal and compliance risk: 0.01
[0096] No risk: 0
[0097] This example demonstrates how the present invention can accurately identify and label specific financial risk categories contained in text.
[0098] 2. Application of classification annotation in actual business
[0099] In a financial institution, the present invention can be used to classify and label a large amount of text data such as corporate announcements and financial reports, which can help the risk control department quickly identify potential risk points and take corresponding measures to prevent and control risks. For example, for enterprises marked as having high "liquidity risk", financial institutions can strengthen the monitoring of their capital flows to prevent potential default events.
[0100] 3. Performance evaluation
[0101] 3.1. Improved accuracy
[0102] In order to evaluate the performance of the invention in the task of financial risk classification, we selected a text dataset related to corporate economic activities covering various financial risk categories (including credit risk, liquidity risk, market risk, operational risk, legal and compliance risk, etc.) for experiments. We evaluated the model by drawing the receiver operating characteristic (ROC) curve and calculating the area under the curve (AUC) value.
[0103] The ROC curve is used to show the trade-off between the true positive rate (TPR) and the false positive rate (FPR) of the classifier at different thresholds. The AUC value is a quantitative indicator of the overall performance of the ROC curve, and its value ranges from 0.5 (random guessing) to 1 (perfect classification). The higher the value, the better the performance of the classifier.
[0104] The experimental results show that compared with the traditional method, the AUC value of the present invention (the improved method based on transformer) is significantly improved from 86% to 94%. This improvement is not only reflected in the numerical value, but also more intuitively reflected in the ROC curve, that is, the improved model can achieve a higher TPR at the same FPR, or maintain a lower FPR at the same TPR. This fully proves the significant advantages of the present invention in capturing the financial risk characteristics in corporate texts, performing fine classification, and improving the generalization ability of the model.
[0105] The improvement in AUC value means that the present invention can more accurately identify financial risk categories, especially liquidity risk and market risk, in practical applications, and provide more accurate and reliable decision support for the risk management of financial institutions. At the same time, this also further verifies the effectiveness of the transformer architecture in processing complex text data and capturing deep features.
[0106] Figure 3 This is a comparison chart of the ROC curves before and after the model improvement, such as Figure 3 As shown in the figure, by comparison, we can clearly see the significant improvement in classification performance of the improved model.
[0107] Therefore, the present invention provides an enterprise financial risk classification and labeling method based on the improved Transformer algorithm, which improves the accuracy and practicality of enterprise financial business risk labeling and classification by combining the advantages of Transformer and random forest algorithms, thereby significantly improving the efficiency of risk judgment.
[0108] The various embodiments in this specification are described in a progressive manner, and each embodiment focuses on the differences from other embodiments. The same or similar parts between the various embodiments can be referenced to each other.
[0109] The principles and implementation methods of the present invention are described in this article using specific examples. The description of the above embodiments is only used to help understand the method and core idea of the present invention. At the same time, for those skilled in the art, according to the idea of the present invention, there will be changes in the specific implementation methods and application scope. In summary, the content of this specification should not be understood as limiting the present invention.
Claims
1. A method for classifying and labeling enterprise financial risks based on the improved Transformer algorithm, characterized in that: The following steps are involved: Collecting text data related to the economic activities of the enterprise, preprocessing the text data, and obtaining a text input representation; Encoding and decoding the text input representation to obtain a feature matrix, and constructing an enhanced feature matrix according to the importance of the feature matrix; Removing the last fully connected layer of the original Transformer model and introducing a new fully connected layer to obtain an improved output layer and an improved Transformer model, and then optimizing the improved output layer and the improved Transformer model; Inputting the enhanced feature matrix into the improved output layer, and calculating the classification probability distribution in combination with the Softmax function; Construct a labeled data set, use the labeled data set to train and optimize the improved Transformer model to obtain a Transformer classification and labeling model, and use the Transformer classification and labeling model to classify and label corporate financial risks.
2. According to claim 1, a method for classifying and labeling enterprise financial risks based on the improved Transformer algorithm is characterized in that: Collecting text data related to the economic activities of the enterprise, preprocessing the text data, and obtaining a text input representation, including: collecting text data related to the economic activities of the enterprise, cleaning the text data to remove invalid characters, standardize text formats and non-financial field vocabulary, and then performing word segmentation and word embedding operations on the text data to obtain a text input representation.
3. According to claim 2, a method for classifying and labeling enterprise financial risks based on the improved Transformer algorithm is characterized in that: encoding and decoding said text input representation, A feature matrix is obtained, and then an enhanced feature matrix is constructed according to the importance of the feature matrix, including: Extracting a feature matrix in the last hidden state of the Transformer encoder and decoder layers, and normalizing the feature matrix to obtain a normalized feature matrix; Constructing a financial risk category label corresponding to the text input representation, training a random forest model using the standardized feature matrix and the financial risk category label, and introducing a random seed and a dynamic parameter into the random forest model to obtain a feature importance model; Based on the feature importance model, the Gini index is used to evaluate the importance of each feature to obtain an importance score vector, and then the first k features are selected as key features according to the importance score vector; Using the index set of the key features, corresponding columns are extracted from the standardized feature matrix, and the extracted columns are integrated into an enhanced feature matrix.
4. According to claim 3, a method for classifying and labeling enterprise financial risks based on the improved Transformer algorithm is characterized in that: The last fully connected layer of the original Transformer model is removed, and a new fully connected layer is introduced to obtain an improved output layer and an improved Transformer model, and then the improved output layer and the improved Transformer model are optimized, including: Removing the last fully connected layer of the original Transformer model and introducing a new fully connected layer for mapping the enhanced feature matrix to the category space of the financial risk category label; A Dropout layer is added after the new fully connected layer to prevent overfitting, and during the model training process, the Dropout layer is used to randomly discard the outputs of some neurons to enhance the generalization ability of the model; A BatchNormalization layer for stabilizing feature distribution is added after the new fully connected layer and the Dropout layer to normalize the small batch data to obtain an improved output layer and an improved Transformer model; The weight matrix of the new connection layer is initialized using the original Transformer model, and the improved Transformer model is optimized using an objective function.
5. According to claim 4, a method for classifying and labeling enterprise financial risks based on the improved Transformer algorithm is characterized by: The expression of the weight matrix is: IN new ∈R k×C Among them, K is the number of features selected after feature selection and screening, and C is the number of categories of financial risk classification; The expression of the objective function is: Among them, L task is the Dropout function, is a function, and λ is a regularization strength hyperparameter.
6. According to claim 5, a method for classifying and labeling enterprise financial risks based on the improved Transformer algorithm is characterized in that: The enhanced feature matrix is input into the improved output layer, and the classification probability distribution is obtained by combining with the Softmax function, including: The expression of the improved output layer is: S=BatchNorm(Dropout(X selected W new )) The expression of the improved output layer combined with the Softmax function is: P=softmax(S),P∈R n×C Where n is the number of samples.
7. According to claim 6, a method for classifying and labeling enterprise financial risks based on the improved Transformer algorithm is characterized in that: Constructing a labeled data set, using the labeled data set to train and optimize the improved Transformer model to obtain a Transformer classification and labeling model, and using the Transformer classification and labeling model to classify and label corporate financial risks, including: Based on the collected text input representation, construct a labeled data set including a training set and a validation set, and perform desensitization processing on the labeled data set; Use the cross entropy loss function to train the model, use the optimizer to adjust the model parameters, and during the training process, monitor the model loss value and accuracy index to evaluate and adjust the model. Use the learning rate decay strategy to control the learning rate change and complete the model training. Data augmentation technology is used to increase the diversity of training samples and perform model optimization. During the optimization process, the validation set is used to evaluate the model. The hyperparameters of the model are adjusted according to the evaluation results to obtain the Transformer classification and labeling model for actual risk classification and labeling tasks.
Citation Information
Cited By
Nuclear power plant radiation risk classification method and device based on random forest algorithm
CN121919707A