Enterprise Credit Risk Early Warning Method and System Integrating Temporal Features of Forum Texts

By integrating time-sensitive forum text features and residual deep-gated recurrent neural networks, the method improves the accuracy of enterprise credit risk warnings, addressing the limitations of existing systems in handling time-sensitive data and feature extraction granularity.

CN115496357BActive Publication Date: 2025-07-15HEFEI UNIV OF TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211140585.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-09-20
Publication Date
2025-07-15
Estimated Expiration
2042-09-20

AI Technical Summary

Technical Problem

In the existing enterprise credit risk warning methods, the time characteristics of the forum text data are not fully considered, resulting in low warning accuracy.

Method used

By obtaining enterprise credit data and forum text data, a feature screening method based on mutual information is adopted, combining time fine-grained division and residual deep-gated gated recurrent network model, integrating forum text timing characteristics, and early warning is used using Adaboost integrated learning model.

Benefits of technology

It improves the accuracy of enterprise credit risk warning and solves the problem of insufficient warning accuracy caused by time factors in the existing technology.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115496357B_ABST
    Figure CN115496357B_ABST
Patent Text Reader

Abstract

The present invention provides an enterprise credit risk early warning method, system, storage medium and electronic device that integrates the time series characteristics of forum texts, and relates to the field of enterprise credit risk early warning. It mainly includes obtaining enterprise credit data and text data in enterprise forums; performing feature screening according to the obtained enterprise credit data; statistically counting the data corresponding to each feature in chronological order according to the obtained forum text data; dividing the obtained serialized statistical data of forum texts using different time granularities, respectively inputting the data divided by time granularity into the constructed residual deep gated recurrent model, selecting the one with the smallest error as the time series feature to construct the model, and finally integrating the credit risk feature and the time series feature, and using the Adaboost model to early warn the enterprise credit risk. It solves the problems that the existing text data processing methods do not consider time factors and the accuracy of the early warning model is not high.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of enterprise credit risk early warning, and in particular to an enterprise credit risk early warning method, system, storage medium and electronic device integrating the temporal characteristics of forum texts. Background Art

[0002] In the context of the global financial crisis, it is even more important for enterprises to establish a scientific early warning system for corporate credit risk. The work involved in corporate credit risk management is numerous and complex, and corporate credit risk managers need to reasonably plan and deploy corporate credit risk early warning management content according to the development of the market economy to ensure the smooth progress of corporate activities. Therefore, how to detect and correctly predict financial crises as early as possible is of great practical significance for reducing corporate operating risks, protecting the interests of investors and creditors, government supervision of enterprises, and preventing financial crises.

[0003] The existing corporate credit risk warning uses the traditional Logit model and machine learning model, which has achieved good results in warning, but the accuracy of the warning is not high. The integrated learning model trains multiple weak classifiers and continuously updates the weight of each weak classifier to obtain the optimal classifier combination. It has good prediction effects in many practical activities.

[0004] The forum is an important platform that can provide investors, experts, etc. with online communication and information sharing. It brings together relevant information collected from various channels, and has the characteristics of wide coverage, large amount of information, and high update frequency. Compared with obtaining information from other channels, it has higher timeliness and professionalism. For the large amount of text data in the forum, existing studies directly collect text data and use the entire document corpus as the information extraction source to extract sentiment features, semantic features, topic features, etc. However, the release of text data has time tags, and different time tags have different effects on the credit risk level of enterprises. Directly using text data as the entire corpus for feature extraction research is not fine-grained enough. At the same time, the existing deep learning model can learn serialized data well, but there is still a problem of low accuracy. Summary of the invention

[0005] 1. Technical issues to be resolved

[0006] In view of the deficiencies in the prior art, the present invention provides a method, system, storage medium and electronic device for early warning of enterprise credit risk integrating the temporal characteristics of forum text, which solves the technical problem of low accuracy of early warning of enterprise credit risk.

[0007] (II) Technical solution

[0008] To achieve the above objectives, the present invention is implemented through the following technical solutions:

[0009] An enterprise credit risk early warning method integrating the time series characteristics of forum texts, comprising:

[0010] Obtain enterprise credit data and text data in enterprise forums;

[0011] According to the enterprise credit data, adopt a feature screening method based on mutual information to obtain credit risk features for enterprise risk early warning;

[0012] According to the text data, count the number of text data published in the forum, the number of replies corresponding to each piece of text data, and the positive or negative sentiment corresponding to each piece of text data;

[0013] Perform serialization processing on the obtained statistical information according to time granularity, divide it by day, week, and month respectively, and obtain three kinds of serialized text data with fine granularity;

[0014] Input the three kinds of serialized data with fine granularity into a pre-constructed residual deep gated recurrent model respectively, and use the vector of the last hidden layer of the model as the time series feature;

[0015] Fuse the credit risk features and time series features, and adopt a pre-constructed Adaboost ensemble learning model to obtain the early warning result of the enterprise credit risk level.

[0016] Preferably, the process of obtaining the credit risk features includes:

[0017] S21. Standardize the enterprise credit data:

[0018]

[0019] where a i,j represents the j-th enterprise credit risk feature value of the i-th enterprise, max(a i,j ), min(a i,j ) represent the maximum and minimum values among the j-th enterprise credit risk feature values respectively, and x i,j represents the standardized enterprise credit risk feature;

[0020] S22. Solve the mutual information I(X; Y) between the standardized enterprise credit risk features and the enterprise credit risk level,

[0021]

[0022] where the variables follow the distributions X ∼ p(x), Y ∼ p(y), (X, Y) ∼ p(x, y);

[0023] X represents the standardized enterprise credit risk feature, There are I rows and J columns in total;

[0024] Y represents the enterprise credit risk level;

[0025] p(x, y) represents the joint distribution probability, p(x) represents the probability distribution that the enterprise credit risk characteristics follow, and p(y) represents the probability distribution of the enterprise credit risk level;

[0026] S23. Obtain the credit risk characteristics for enterprise risk early warning according to the size relationship of I(X; Y).

[0027] Preferably, the process of obtaining the positive or negative sentiment corresponding to each piece of text data includes:

[0028] Perform regularization processing, word segmentation, and stop word removal on the text data, convert the words into corresponding scores based on the positive sentiment dictionary, negative sentiment dictionary, negation word dictionary, and degree word dictionary, obtain the sentiment score of each piece of text data, and determine whether the text data belongs to positive sentiment or negative sentiment according to the score. If it belongs to positive sentiment, mark 1 in the positive sentiment quantity feature and 0 in the negative sentiment quantity feature. If it belongs to negative sentiment, mark 0 in the positive sentiment quantity feature and 1 in the negative sentiment quantity feature, so as to complete the sentiment classification feature work.

[0029] Preferably, the serialization process of the obtained statistical information according to the time granularity is carried out by dividing it according to days, weeks, and months respectively to obtain three kinds of serialized text data with fine granularity, which specifically includes:

[0030] A. Serialized text data divided by day as the fine granularity:

[0031] According to the time label, perform quantity feature statistics on each piece of text data generated in a day, and obtain the serialized data x_day in units of days; expressed as

[0032] x_day = {D(total_num, reply_num, postive_num, negtive_num)1,

[0033] …, D(total_num, reply_num, postive_num, negtive_num) l ,

[0034] …, D(total_num, reply_num, postive_num, negtive_num) L}

[0035] Among them, D(total_num, reply_num, postive_num, negtive_num)l It represents the total number of text data total_num, the number of replies reply_num corresponding to the text data, the number of positive emotions postive_num, and the number of negative emotions negtive_num released on the l-th day of the enterprise forum with a daily time granularity; l = 1,..., L;

[0036] B. Serialized text data divided by week as the granularity:

[0037] According to the time tag, the quantity characteristics of the text data generated in a week are statistically analyzed, and serialized data x_week is obtained in units of weeks; expressed as

[0038] x_week = {W(total_num, reply_num, postive_num, negtive_num)1,

[0039] ..., W(total_num, reply_num, postive_num, negtive_num) m ,

[0040] ..., W(total_num, reply_num, postive_num, negtive_num) M}

[0041] Among them, W(total_num, reply_num, postive_num, negtive_num) m It represents the total number of text data total_num, the number of replies reply_num corresponding to the text data, the number of positive emotions postive_num, and the number of negative emotions negtive_num released on the m-th week of the enterprise forum with a weekly time granularity; m = 1,..., M;

[0042] C. Obtain serialized text data divided by month as the granularity:

[0043] According to the time tag, the quantity characteristics of the text data generated in a month are statistically analyzed, and serialized data x_month is obtained in units of months; expressed as

[0044] x_month = {M(total_num, reply_num, postive_num, negtive_num)1,

[0045] ..., M(total_num, reply_num, postive_num, negtive_num) n ,

[0046] …, M(total_num, reply_num, postive_num, negtive_num) N}

[0047] Among them, M(total_num, reply_num, postive_num, negtive_num) n represents that with a monthly time granularity division, the total number of text data total_num, the number of replies reply_num corresponding to the text data, the number of positive sentiment postive_num, and the number of negative sentiment negtive_num published on the enterprise forum in the nth month; n = 1, …, N.

[0048] Preferably, the construction process of the residual deep gated recurrent model includes:

[0049] S10. Data division: According to the method of K-fold stratified sampling, divide the training set and the test set according to the enterprise number;

[0050] S20. Input layer: According to the divided training set and test set, respectively use the three kinds of serialized text data x ∈ {x_day, x_week, x_month} as the model input;

[0051] S30. Encoding layer: Input the serialized text data into the residual deep gated recurrent network model respectively. The residual deep gated recurrent network model consists of three stacked residual deep gated recurrent networks. Each layer passes through the deep gated recurrent network, and then the output of the first layer is changed through the residual structure. The calculation result of the first-layer residual network is used as the input of the second layer, and similarly, the calculation result of the second-layer residual network is used as the input of the third layer;

[0052] S40. Output layer: According to the enterprise credit risk level, set the output parameter of the last layer of the DNN to the number of enterprise credit risk levels;

[0053]

[0054] Among them, c = 1, 2, …, k, k represents the level of enterprise credit risk, yc represents the probability that the actual enterprise credit risk level is c, represents the output vector after passing through three stacked residual deep gated recurrent networks. DNN is a multi-layer perceptron module, which consists of multiple multi-layer perceptron layers;

[0055] S50. According to the training effects of the three kinds of fine-grained models, select the fine-grained model with the smallest error as the forum text time series feature extraction model;

[0056]

[0057] Among them, R represents the number of training samples, represents the probability that the predicted enterprise credit risk level is c.

[0058] Preferably, the encoding layer specifically includes:

[0059] Regarding the first-layer residual depth gated recurrent network,

[0060] First, calculate through the depth gated recurrent network:

[0061] x 1,t h 1,t = GRU(x)

[0062] Among them, x 1,t is the output vector after passing through the depth gated recurrent network, h 1,t is the output vector of the last hidden layer of the depth gated recurrent network, and x is the serialized text input data;

[0063] Secondly, calculate through the residual network:

[0064]

[0065] Among them, x 1,t is the output vector after passing through the depth gated recurrent network, and the two are added to form a residual structure;

[0066] Take the calculation result of the first-layer residual depth gated recurrent network as the input of the second-layer depth gated recurrent network;

[0067] Regarding the second-layer residual depth gated recurrent network,

[0068] First, calculate through the depth gated recurrent network:

[0069]

[0070] Among them, x 2,t is the output vector after passing through the depth gated recurrent network, h 2,t is the output vector of the last hidden layer of the depth gated recurrent network;

[0071] Secondly, calculate through the residual network:

[0072]

[0073] Among them, is the output vector of the first-layer residual depth gated recurrent network, x 2,tis the output vector of the deep gated recurrent network, and the calculation result of the first-layer residual deep gated recurrent network is added to the calculation result of the second pass through the deep gated recurrent network to form the residual structure of the second layer;

[0074] The result after the transformation of the second layer through the residual network is used as the input of the third-layer deep gated recurrent network;

[0075] Regarding the third-layer residual deep gated recurrent network,

[0076] First, perform calculations through the deep gated recurrent network:

[0077]

[0078] where x 3,t is the output vector of the deep gated recurrent network, and h 2,t is the output vector of the last hidden layer of the deep gated recurrent network;

[0079] Secondly, perform calculations through the residual network:

[0080]

[0081] where, is the output vector of the second-layer residual deep gated recurrent network, x 3,t is the output vector of the deep gated recurrent network, and the calculation result of the second-layer residual deep gated recurrent network is added to the calculation result of the third pass through the deep gated recurrent network to form the residual structure of the third layer.

[0082] Preferably, the construction process of the Adaboost ensemble learning model includes:

[0083] S61. Initialize the sample weights, and the initial value is

[0084] where N represents the number of samples during training;

[0085] S62. Train the weak classifier and calculate the prediction error:

[0086]

[0087] where P is 1 when the inequality in the parentheses holds and 0 when it does not hold, and y n represents the actual credit risk level corresponding to the nth enterprise, represents the credit risk level corresponding to the nth enterprise predicted by the weak classifier;

[0088] S63. Calculate the weight of each weak classifier according to the prediction error:

[0089]

[0090] S64. Calculate the weights of the samples for the next round of training according to the obtained weights:

[0091]

[0092] where Z n represents the normalization factor;

[0093] S65. After multiple rounds of iteration, when the training error reaches the minimum, obtain the final strong classifier.

[0094] An enterprise credit risk early warning system integrating the time series features of forum texts, comprising:

[0095] An acquisition module, configured to acquire enterprise credit data and text data in enterprise forums;

[0096] A screening module, configured to obtain credit risk features for enterprise risk early warning by using a feature screening method based on mutual information according to the enterprise credit data;

[0097] A statistics module, configured to count the number of text data published in the forum, the number of replies corresponding to each piece of text data, and the positive or negative sentiment corresponding to each piece of text data according to the text data;

[0098] A serialization module, configured to perform serialization processing on the obtained statistical information according to time granularity, divide it by day, week, and month respectively, and obtain three types of serialized text data with fine granularity;

[0099] An extraction module, configured to input the three types of serialized data with fine granularity into a pre-constructed residual deep gated recurrent model respectively, and use the hidden layer vector of the last layer of the model as the time series feature;

[0100] An early warning module, configured to integrate the credit risk features and time series features, and use a pre-constructed Adaboost ensemble learning model to obtain the early warning result of the enterprise credit risk level.

[0101] A storage medium stores a computer program for enterprise credit risk early warning integrating the time series features of forum texts, wherein the computer program enables a computer to execute the enterprise credit risk early warning method as described above.

[0102] An electronic device, comprising:

[0103] One or more processors;

[0104] A memory; and

[0105] One or more programs, wherein the one or more programs are stored in the memory and configured to be executed by the one or more processors, and the programs include those for executing the enterprise credit risk warning method as described above.

[0106] (III) Advantageous Effects

[0107] The present invention provides an enterprise credit risk warning method, system, storage medium and electronic device that integrate the time series characteristics of forum texts. Compared with the prior art, the following advantageous effects are achieved:

[0108] Starting from improving the accuracy of enterprise credit risk warning, the present invention considers a method of integrating the time series characteristics of forum texts, mainly including obtaining enterprise credit data and text data in enterprise forums; performing feature screening based on the obtained enterprise credit data; statistically counting the data corresponding to each feature in the obtained forum text data in chronological order; dividing the obtained serialized statistical data of forum texts using different time granularities, respectively inputting the data divided by the time granularity into the constructed residual deep gated recurrent model, selecting the one with the smallest error as the time series feature to construct the model, and finally integrating the credit risk features and time series features, and using the Adaboost model to warn of enterprise credit risks. It solves the problems that the existing text data processing methods do not consider time factors and the accuracy of the warning model is not high. BRIEF DESCRIPTION OF THE DRAWINGS

[0109] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0110] Figure 1 It is a block diagram of an enterprise credit risk warning method that integrates the time series characteristics of forum texts provided by an embodiment of the present invention;

[0111] Figure 2 It is a flowchart of an enterprise credit risk warning method that integrates the time series characteristics of forum texts provided by an embodiment of the present invention;

[0112] Figure 3 It is a schematic diagram of the extraction of the time series characteristics of forum texts provided by an embodiment of the present invention;

[0113] Figure 4 It is a schematic diagram of a residual deep gated recurrent model provided by an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0114] To make the objectives, technical solutions and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described. Apparently, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without making creative efforts shall fall within the protection scope of the present invention.

[0115] By providing an enterprise credit risk early warning method, system, storage medium and electronic device that integrates the temporal characteristics of forum texts, the embodiments of the present application solve the technical problem of low accuracy in enterprise credit risk early warning.

[0116] The general idea of the technical solutions in the embodiments of the present application to solve the above technical problems is as follows:

[0117] Starting from improving the accuracy of enterprise credit risk early warning, the embodiments of the present invention consider a method of integrating the temporal characteristics of forum texts, which specifically includes the following aspects:

[0118] First, in enterprise credit risk early warning, not only the credit characteristics of the enterprise are considered, but also the temporal characteristics of a kind of forum text data are integrated; second, for the temporal characteristics of forum texts, the influence of time granularity on the construction of temporal characteristics is more comprehensively considered. Not only time series are considered, but also a residual deep gated recurrent network is constructed, and further trained and precise temporal characteristics are extracted; finally, a more accurate enterprise credit risk early warning is realized based on the Adaboost ensemble learning model.

[0119] To better understand the above technical solutions, the above technical solutions will be described in detail below in conjunction with the accompanying drawings of the specification and specific implementation manners.

[0120] Embodiment:

[0121] As Figures 1 - 2 shown, the embodiments of the present invention provide an enterprise credit risk early warning method that integrates the temporal characteristics of forum texts, including:

[0122] S1. Obtain enterprise credit data and text data in the enterprise forum;

[0123] S2. According to the enterprise credit data, use a feature screening method based on mutual information to obtain credit risk features for enterprise risk early warning;

[0124] S3. According to the text data, count the number of text data published in the forum, the number of replies corresponding to each text data, and the positive or negative sentiment corresponding to each text data;

[0125] S4. Serialize the obtained statistical information according to time granularity, divide it by day, week, and month respectively, and obtain serialized text data of three granularities;

[0126] S5. Input the serialized data of the three granularities into a pre-constructed residual deep gated recurrent model respectively, and use the vector of the last hidden layer of the model as the time series feature;

[0127] S6. Integrate the credit risk feature and the time series feature, and use a pre-constructed Adaboost ensemble learning model to obtain the early warning result of the enterprise credit risk level.

[0128] The embodiment of the present invention solves the problems that the existing text data processing method does not consider time and the fine granularity of model data input division, and improves the accuracy of enterprise credit risk prediction.

[0129] The following will introduce each step of the above technical solution in detail:

[0130] In step S1, enterprise credit data and text data in the enterprise forum are obtained.

[0131] In step S2, according to the enterprise credit data, a feature screening method based on mutual information is used to obtain credit risk features for enterprise risk early warning.

[0132] The obtaining process of the credit risk feature includes:

[0133] S21. Standardize the enterprise credit data:

[0134]

[0135] where a i,j represents the jth enterprise credit risk feature value of the ith enterprise, max(a i,j ), min(a i,j ) represent the maximum and minimum values of the jth enterprise credit risk feature value respectively, and x i,j represents the standardized enterprise credit risk feature;

[0136] S22. Solve the mutual information I(X; Y) between the standardized enterprise credit risk feature and the enterprise credit risk level,

[0137]

[0138] where the variables follow the distributions X ~ p(x), Y ~ p(y), (X, Y) ~ p(x, y);

[0139] X represents the standardized enterprise credit risk feature, with a total of I rows and J columns;

[0140] Y represents the enterprise credit risk level;

[0141] p(x, y) represents the joint distribution probability, p(x) represents the probability distribution that the enterprise credit risk characteristics follow, and p(y) represents the probability distribution of the enterprise credit risk level;

[0142] S23. Obtain the credit risk characteristics for enterprise risk warning according to the size relationship of I(X; Y).

[0143] In step S3, according to the text data, count the number of text data published in the forum, the number of replies corresponding to each piece of text data, and the positive or negative sentiment corresponding to each piece of text data.

[0144] A large number of users participate in the forum and express their opinions. One piece of text data is generated according to the expressed opinions. The number of replies corresponding to each piece of text data is different, and the number of replies corresponding to each piece of text data is counted respectively; the sentiment characteristics shown by each piece of text data are different. Using the method of sentiment dictionary, the sentiment characteristics corresponding to each piece of text data are classified.

[0145] The process of obtaining the positive or negative sentiment corresponding to each piece of text data includes: performing regularization processing, word segmentation, and stop word removal on the text data, converting the words into corresponding scores based on the positive sentiment dictionary, negative sentiment dictionary, negation word dictionary, and degree word dictionary, obtaining the sentiment score of each piece of text data, and determining whether the piece of text data belongs to positive sentiment or negative sentiment according to the score. If it belongs to positive sentiment, mark 1 in the positive sentiment quantity feature and 0 in the negative sentiment quantity feature; if it belongs to negative sentiment, mark 0 in the positive sentiment quantity feature and 1 in the negative sentiment quantity feature, thus completing the sentiment classification feature work.

[0146] In step S4, as Figure 3 shown, perform serialization processing on the obtained statistical information according to the time granularity, divide it by day, week, and month respectively, and obtain three types of serialized text data with fine granularity.

[0147] A. Serialized text data divided by day as the fine granularity:

[0148] According to the time tag, perform quantity feature statistics on each piece of text data generated in a day, and obtain the serialized data x_day in units of days; expressed as

[0149] x_day = {D(total_num, reply_num, postive_num, negtive_num)1,

[0150] …, D(total_num, reply_num, postive_num, negtive_num) l ,

[0151] …, D(total_num, reply_num, postive_num, negtive_num) L}

[0152] Among them, D(total_num, reply_num, postive_num, negtive_num) l represents the total number of text data total_num, the number of replies reply_num corresponding to the text data, the number of positive sentiment postive_num, and the number of negative sentiment negtive_num published on the enterprise forum on the l-th day, with the time granularity divided by day; l = 1,..., L;

[0153] B. Serialized text data divided by week as the fine granularity:

[0154] According to the time tag, the quantity characteristics of the text data generated in a week are statistically counted, and serialized data x_week is obtained with a week as the unit; expressed as

[0155] x_week = {W(total_num, reply_num, postive_num, negtive_num)1,

[0156] …, W(total_num, reply_num, postive_num, negtive_num) m ,

[0157] …, W(total_num, reply_num, postive_num, negtive_num) M}

[0158] Among them, W(total_num, reply_num, postive_num, negtive_num) m represents the total number of text data total_num, the number of replies reply_num corresponding to the text data, the number of positive sentiment postive_num, and the number of negative sentiment negtive_num published on the enterprise forum in the m-th week, with the time granularity divided by week; m = 1,..., M;

[0159] C. Obtain serialized text data divided by month as the fine granularity:

[0160] Statistically analyze the quantitative features of the text data generated in a month according to the time tags. Taking a month as a unit, obtain the serialized data x_month, which is expressed as

[0161] x_month = {M(total_num, reply_num, postive_num, negtive_num)1,

[0162] …, M(total_num, reply_num, postive_num, negtive_num) n ,

[0163] …, M(total_num, reply_num, postive_num, negtive_num) N}

[0164] Among them, M(botal_num, reply_num, postive_num, negtive_num) n represents the total number of text data total_num, the number of replies reply_num corresponding to the text data, the number of positive sentiment quantities postive_num, and the number of negative sentiment quantities negtive_num published on the enterprise forum in the nth month divided by the time granularity of a month; n = 1, …, N.

[0165] In step S5, input the serialized data of the three granularities into a pre-constructed residual deep gated recurrent model respectively, and use the vector of the last hidden layer of the model as the time series feature.

[0166] According to the constructed serialized text data of the three granularities, the embodiments of the present invention do not directly use the serialized text data as the time series feature, but adopt the superiority of the deep gated recurrent network model combined with the residual network. It can not only retain the original information, but also further aggregate the time series information trained by the deep gated recurrent network. As Figure 3 shown, put the original serialized text data of the three granularities into the model for training. According to the training errors of the three granularity divisions, select the granularity model with the smallest error as the time series feature extraction model, and use the vector of the last hidden layer after model training and processing as the time series feature of the text.

[0167] As Figure 4 shown, the construction process of the residual deep gated recurrent model includes:

[0168] S10. Data division: Divide the training set and the test set according to the enterprise number by the method of K-fold stratified sampling.

[0169] S20. Input layer: According to the divided training set and test set, the three types of fine-grained serialized text data \(x\in\{x_{day}, x_{week}, x_{month}\}\) are respectively used as the model input.

[0170] S30. Encoding layer: The serialized text data is respectively input into the residual deep gated recurrent network model. The residual deep gated recurrent network model consists of three stacked residual deep gated recurrent networks. Each layer passes through the deep gated recurrent network, and then the output of the first layer is transformed through the residual structure. The calculation result of the first-layer residual network is used as the input of the second layer. Similarly, the calculation result of the second-layer residual network is used as the input of the third layer; thus, the accuracy of text temporal feature construction is improved.

[0171] Among them, regarding the first-layer residual deep gated recurrent network,

[0172] First, it is calculated through the deep gated recurrent network:

[0173] x 1,t , h 1,t =GRU(x)

[0174] Among them, x 1,t is the output vector passing through the deep gated recurrent network, h 1,t is the output vector of the last hidden layer of the deep gated recurrent network, and x is the serialized text input data;

[0175] Secondly, it is calculated through the residual network:

[0176]

[0177] Among them, x 1,t is the output vector passing through the deep gated recurrent network, and the two are added to form a residual structure;

[0178] The calculation result of the first-layer residual deep gated recurrent network is used as the input of the second-layer deep gated recurrent network;

[0179] Regarding the second-layer residual deep gated recurrent network,

[0180] First, it is calculated through the deep gated recurrent network:

[0181]

[0182] Among them, x 2,t is the output vector passing through the deep gated recurrent network, h 2,t is the output vector of the last hidden layer of the deep gated recurrent network;

[0183] Secondly, it is calculated through the residual network:

[0184]

[0185] Among them, is the output vector of the first-layer residual depth gated recurrent network, and x 2,t is the output vector after passing through the depth gated recurrent network. The result of the calculation through the first-layer residual depth gated recurrent network is added to the result of the second calculation through the depth gated recurrent network to form the residual structure of the second layer;

[0186] The result after the transformation of the second layer through the residual network is used as the input of the third-layer depth gated recurrent network;

[0187] Regarding the third-layer residual depth gated recurrent network,

[0188] First, calculate through the depth gated recurrent network:

[0189]

[0190] Among them, x 3,t is the output vector after passing through the depth gated recurrent network, and h 2,t is the output vector of the last hidden layer of the depth gated recurrent network;

[0191] Secondly, calculate through the residual network:

[0192]

[0193] Among them, is the output vector of the second-layer residual depth gated recurrent network, and x 3,t is the output vector after passing through the depth gated recurrent network. The result of the calculation through the second-layer residual depth gated recurrent network is added to the result of the third calculation through the depth gated recurrent network to form the residual structure of the third layer.

[0194] S40, Output layer: According to the enterprise credit risk level, set the output parameter of the last layer of the DNN to the number of enterprise credit risk levels;

[0195]

[0196] Among them, c = 1, 2,..., k, where k represents the level of enterprise credit risk, and yc represents the probability that the actual enterprise credit risk level is c, represents the output vector after passing through the three-layer stacked residual depth gated recurrent network. The DNN is a multi-layer perceptron module composed of multiple multi-layer perceptron layers.

[0197] S50. According to the training effects of the three fine-grained models, select the fine-grained model with the smallest error as the forum text time series feature extraction model;

[0198]

[0199] Among them, R represents the number of training samples, represents the probability that the predicted enterprise credit risk level is c.

[0200] In step S6, fuse the credit risk features and time series features, and use the pre-constructed Adaboost ensemble learning model to obtain the early warning result of the enterprise credit risk level.

[0201] It can be seen from the above that according to the data set division method in the time series feature extraction, the enterprise credit risk features are divided according to the same enterprise number. In this step, the divided data set is used to obtain the time series features through step S4, and then fused with the enterprise credit risk features. Each fold of data is put into the Adaboost model in turn to conduct early warning on the enterprise credit risk.

[0202] The construction process of the Adaboost ensemble learning model includes:

[0203] S61. Initialize the sample weights, and the initial value is

[0204] Among them, N represents the number of samples during training;

[0205] S62. Train the weak classifier and calculate the prediction error:

[0206]

[0207] Among them, P is 1 when the inequality in the parentheses holds and 0 when it does not hold, and y n represents the actual corresponding credit risk level of the nth enterprise, represents the credit risk level corresponding to the nth enterprise predicted by the weak classifier;

[0208] S63. Calculate the weight of each weak classifier according to the prediction error:

[0209]

[0210] S64. Calculate the weights of the samples included in the next round of training according to the obtained weights:

[0211]

[0212] Among them, Z n represents the normalization factor;

[0213] S65. After multiple rounds of iteration, when the training error reaches the minimum, obtain the final strong classifier.

[0214] An embodiment of the present invention provides an enterprise credit risk early warning system integrating the time series characteristics of forum texts, including:

[0215] An acquisition module, configured to acquire enterprise credit data and text data in enterprise forums;

[0216] A screening module, configured to obtain credit risk characteristics for enterprise risk early warning by using a feature screening method based on mutual information according to the enterprise credit data;

[0217] A statistics module, configured to count the number of text data published in the forum, the number of replies corresponding to each piece of text data, and the positive or negative sentiment corresponding to each piece of text data according to the text data;

[0218] A serialization module, configured to perform serialization processing on the obtained statistical information according to time granularity, divide it by day, week, and month respectively, and obtain three kinds of serialized text data with fine granularity;

[0219] An extraction module, configured to input the three kinds of serialized data with fine granularity into a pre-constructed residual deep gated recurrent model respectively, and use the vector of the last hidden layer of the model as the time series feature;

[0220] An early warning module, configured to fuse the credit risk characteristics and time series characteristics, and use a pre-constructed Adaboost integrated learning model to obtain an early warning result of enterprise credit risk level.

[0221] An embodiment of the present invention provides a storage medium, which stores a computer program for enterprise credit risk early warning integrating the time series characteristics of forum texts, wherein the computer program enables a computer to execute the enterprise credit risk early warning method as described above.

[0222] An embodiment of the present invention provides an electronic device, including:

[0223] One or more processors;

[0224] A memory; and

[0225] One or more programs, wherein the one or more programs are stored in the memory and are configured to be executed by the one or more processors, and the programs include those for executing the enterprise credit risk early warning method as described above.

[0226] It is understandable that the enterprise credit risk early warning system, storage medium, and electronic device provided by the embodiments of the present invention for integrating the temporal characteristics of forum texts correspond to the enterprise credit risk early warning method provided by the embodiments of the present invention. For the explanations, examples, beneficial effects, and other parts of the relevant content, reference can be made to the corresponding parts in the enterprise credit risk early warning method based on blockchain, which will not be elaborated here.

[0227] In summary, compared with the prior art, the following beneficial effects are achieved:

[0228] 1. First, in enterprise credit risk early warning, not only the credit characteristics of the enterprise are considered, but also the temporal characteristics of a kind of forum text data are incorporated. Compared with not considering the time sequence, the early warning effect is significantly improved after incorporating the temporal characteristics.

[0229] 2. Second, for the temporal characteristics of forum texts, the influence of time granularity on the construction of temporal characteristics is more comprehensively considered. Not only the time series is considered, but also a residual deep gated recurrent network is constructed to further train and extract precise temporal characteristics.

[0230] 3. Finally, a more accurate enterprise credit risk early warning is realized based on the Adaboost ensemble learning model.

[0231] It should be noted that in this article, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device comprising a series of elements not only includes those elements, but also includes other elements not expressly listed, or elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "comprising one..." does not exclude the presence of additional identical elements in the process, method, article or device comprising the element.

[0232] The above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. An enterprise credit risk early warning method integrating the time series characteristics of forum texts, characterized in that, Including: Obtaining enterprise credit data and text data in the enterprise forum; According to the enterprise credit data, adopting a feature screening method based on mutual information to obtain credit risk features for enterprise risk early warning; According to the text data, counting the number of text data published in the forum, the number of replies corresponding to each piece of text data, and the positive or negative sentiment corresponding to each piece of text data; Performing serialization processing on the obtained statistical information according to time granularity, dividing it by day, week, and month respectively to obtain three types of serialized text data with fine granularity; Inputting the three types of serialized data with fine granularity into a pre-constructed residual deep gated recurrent model respectively, and using the hidden layer vector of the last layer of the model as the time series feature; Fusing the credit risk features and time series features, and adopting a pre-constructed Adaboost ensemble learning model to obtain the early warning result of the enterprise credit risk level; The construction process of the residual deep gated recurrent model includes: S10. Data division: Divide the training set and the test set according to the enterprise number by the method of K-fold stratified sampling; S20. Input layer: According to the divided training set and test set, respectively use the three types of serialized text data x∈{x_day, x_week, x_month} as the model input; S30. Encoding layer: Input the serialized text data into the residual deep gated recurrent network model respectively. The residual deep gated recurrent network model consists of three stacked residual deep gated recurrent networks. Each layer passes through the deep gated recurrent network, and then the output of the first layer is changed through the residual structure. The calculation result of the first-layer residual network is used as the input of the second layer, and similarly, the calculation result of the second-layer residual network is used as the input of the third layer; S40. Output layer: According to the enterprise credit risk level, set the output parameter of the last layer of the DNN to the number of enterprise credit risk levels; where \(c = 1, 2, \ldots, k\), \(k\) represents the level of enterprise credit risk, and \(y\) c represents the probability that the actual enterprise credit risk level is \(c\). represents the output vector of the residual deep gated recurrent network after three - layer stacking. DNN is a multi - layer perceptron module composed of multiple multi - layer perceptron layers; S50. According to the training effects of the three types of fine-grained models, select the fine-grained model with the smallest error as the forum text time series feature extraction model; where R represents the number of training samples, represents the probability that the predicted enterprise credit risk level is c.

2. The enterprise credit risk early warning method according to claim 1, characterized in that The obtaining process of the credit risk features includes: S21. Standardize the enterprise credit data: Among them, a i,j represents the j-th enterprise credit risk eigenvalue of the i-th enterprise, max(a i,j ), min(a i,j ) respectively represent the maximum and minimum values among the j-th enterprise credit risk eigenvalues, and x i,j represents the standardized enterprise credit risk feature; S22. Solve the mutual information I(X; Y) between the standardized enterprise credit risk features and the enterprise credit risk level, where the variables follow the distributions X~p(x), Y~p(y), (X, Y)~p(x, y); X represents the standardized enterprise credit risk characteristics, with a total of I rows and J columns; Y represents the enterprise credit risk level; p(x, y) represents the joint distribution probability, p(x) represents the probability distribution that the enterprise credit risk features follow, and p(y) represents the probability distribution of the enterprise credit risk level; S23. According to the size relationship of I(X; Y), obtain the credit risk features for enterprise risk early warning.

3. The enterprise credit risk early warning method according to claim 1, characterized in that The obtaining process of the positive or negative sentiment corresponding to each piece of text data includes: Regularize, tokenize, and remove stop words from the text data. Convert words into corresponding scores based on positive sentiment dictionaries, negative sentiment dictionaries, negation word dictionaries, and degree word dictionaries, and obtain the sentiment scores of each piece of text data. Determine whether the text data belongs to positive sentiment or negative sentiment based on the scores. If it belongs to positive sentiment, mark 1 for the positive sentiment quantity feature and 0 for the negative sentiment quantity feature; if it belongs to negative sentiment, mark 0 for the positive sentiment quantity feature and 1 for the negative sentiment quantity feature, thus completing the sentiment classification feature work.

4. The enterprise credit risk early warning method according to claim 1, characterized in that Serializing the obtained statistical information according to time granularity, dividing it by day, week, and month respectively, and obtaining three types of serialized text data with fine granularity, specifically including: A. Serialized text data divided by day as the fine granularity: According to the time label, conduct quantity feature statistics on each piece of text data generated in a day, and obtain serialized data x_day in units of days; expressed as x_day = {D(total_num, reply_num, postive_num, negtive_num)1, …, D(total_num, reply_num, postive_num, negtive_num) l , …, D(total_num, reply_num, postive_num, negtive_num) L} Among them, D(total_num, reply_num, postive_num, negtive_num) l represents the total number of text data total_num, the number of replies reply_num corresponding to the text data, the number of positive sentiment postive_num, and the number of negative sentiment negtive_num of the text data published on the enterprise forum on the l-th day with a daily time granularity; l = 1,..., L; B. Serialized text data divided by week as the fine granularity: According to the time label, conduct quantity feature statistics on the text data generated in a week, and obtain serialized data x_week in units of weeks; expressed as x_week = {W(total_num, reply_num, postive_num, negtive_num)1, …, W(total_num, reply_num, postive_num, negtive_num) m , …, W(total_num, reply_num, postive_num, negtive_num) M} Among them, W(total_num, reply_num, postive_num, negtive_num) m represents the total number of text data total_num, the number of replies reply_num corresponding to the text data, the number of positive sentiment postive_num, and the number of negative sentiment negtive_num of the text data released by the enterprise forum in the m-th week with a weekly time granularity; m = 1,..., M; C. Obtain serialized text data divided by month as the fine granularity: According to the time label, conduct quantity feature statistics on the text data generated in a month, and obtain serialized data x_month in units of months; expressed as x_month = {M(total_num, reply_num, postive_num, negtive_num)1, …, M(total_num, reply_num, postive_num, negtive_num) n , …, M(total_num, reply_num, postive_num, negtive_num) N} Among them, M(total_num, reply_num, postive_num, negtive_num) n represents the total number of text data total_num, the number of replies reply_num corresponding to the text data, the number of positive sentiment postive_num, and the number of negative sentiment negtive_num of the text data published on the enterprise forum in the nth month with a monthly time granularity; n = 1,..., N.

5. The enterprise credit risk early warning method according to claim 1, characterized in that The encoding layer specifically includes: Regarding the first-layer residual deep gated recurrent network, First, calculate through the deep gated recurrent network: x 1,t ,h 1,t = GRU(x) where x 1,t is the output vector of the deep gated recurrent network, h 1,t is the output vector of the last hidden layer of the deep gated recurrent network, and x is the serialized text input data; Secondly, calculate through the residual network: where x 1,t is the output vector of the deep gated recurrent network, and the two are added to form a residual structure; Take the calculation result of the first-layer residual deep gated recurrent network as the input of the second-layer deep gated recurrent network; Regarding the second-layer residual deep gated recurrent network, First, calculate through the deep gated recurrent network: where x 2,t is the output vector of the deep gated recurrent network, and h 2,t is the output vector of the last hidden layer of the deep gated recurrent network; Secondly, calculate through the residual network: Among them, is the output vector of the first-layer residual depth gated recurrent network, and x 2,t is the output vector passing through the depth gated recurrent network. The result of the first-layer residual depth gated recurrent network calculation and the result of the second pass through the depth gated recurrent network are added together to form the residual structure of the second layer; Take the result after the transformation of the second layer through the residual network as the input of the third-layer deep gated recurrent network; Regarding the third-layer residual deep gated recurrent network, First, calculate through the deep gated recurrent network: where x 3,t is the output vector of the deep gated recurrent network, and h 2,t is the output vector of the last hidden layer of the deep gated recurrent network; Secondly, calculate through the residual network: Among them, is the output vector of the second-layer residual depth gated recurrent network, and x 3,t is the output vector passing through the depth gated recurrent network. The result of the calculation through the second-layer residual depth gated recurrent network is added to the result of the third calculation through the depth gated recurrent network to form the residual structure of the third layer.

6. The enterprise credit risk early warning method according to claim 1, wherein The construction process of the Adaboost ensemble learning model includes: S61. Initialize the sample weights, with the initial value being Among them, N represents the number of samples during training; S62. Train weak classifiers and calculate prediction errors: where P takes the value of 1 when the inequality in the parentheses holds and 0 when it does not, and y n represents the actual credit risk level corresponding to the nth enterprise, and represents the credit risk level corresponding to the nth enterprise predicted by the weak classifier; S63. Calculate the weight of each weak classifier according to the prediction error: S64. Calculate the weights of the samples included in the next round of training according to the obtained weights: where Z n represents a normalization factor; S65. After multiple rounds of iteration, when the training error reaches the minimum, obtain the final strong classifier.

7. An enterprise credit risk early warning system integrating the temporal characteristics of forum texts, characterized in that, For performing the enterprise credit risk warning method as described in claim 1, including: An acquisition module, configured to acquire enterprise credit data and text data in enterprise forums; A screening module, configured to, according to the enterprise credit data, adopt a feature screening method based on mutual information to obtain credit risk features for enterprise risk warning; A statistics module, configured to, according to the text data, count the number of text data published in the forum, the number of replies corresponding to each piece of text data, and the positive or negative sentiment corresponding to each piece of text data; A serialization module, configured to perform serialization processing on the obtained statistical information according to time granularity, divide it by day, week, and month respectively, and obtain three types of serialized text data with fine granularity; An extraction module, configured to respectively input the three types of serialized data with fine granularity into a pre-constructed residual deep gated recurrent model, and use the vector of the last hidden layer of the model as a time series feature; A warning module, configured to fuse the credit risk features and time series features, and adopt a pre-constructed Adaboost integrated learning model to obtain an enterprise credit risk level warning result.

8. A storage medium, characterized in that, It stores a computer program for enterprise credit risk warning that fuses forum text time series features, wherein the computer program enables a computer to execute the enterprise credit risk warning method as described in any one of claims 1 to 6.

9. An electronic device, characterized in that, Including: One or more processors; A memory; And one or more programs, wherein the one or more programs are stored in the memory and are configured to be executed by the one or more processors, and the programs include those for executing the enterprise credit risk warning method as described in any one of claims 1 to 6.

Citation Information

Patent Citations

  • Enterprise credit risk evaluation method and device and electronic equipment

    CN112734559A

  • Credit risk assessment method based on time sequence deep learning and legal document information

    CN114519508A