A risk quantification and assessment method and system based on a subject association network

By automatically identifying and evaluating credit risks using BERT-base model and a two-layer graph attention network, the problem of dependence on the initial network in the prior art is solved, and efficient and accurate quantitative assessment of credit risks is achieved.

CN119006141BActive Publication Date: 2025-07-25E FUND MANAGEMENT CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202411022061.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-07-29
Publication Date
2025-07-25
Estimated Expiration
2044-07-29

AI Technical Summary

Technical Problem

The prior art relies heavily on the initially defined subject-related network in credit risk transmission calculation, resulting in low accuracy and efficiency of risk quantitative assessment, and manual data collection depends on the experience of technicians, making it difficult to ensure the accuracy of assessment.

Method used

By obtaining public opinion text data and subject association network, using the credit risk scoring model based on the BERT-base model and the two-layer graph attention network, we automatically identify the risk subjects and generate the initial credit risk score, predict the association relationship and strength, realize risk transmission, and avoid high dependence on the initial network.

Benefits of technology

It improves the accuracy and efficiency of risk quantitative assessment, can monitor changes in subject status in real time, reduce computing resources, and enhances the accuracy and rationality of risk transmission.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119006141B_ABST
    Figure CN119006141B_ABST
Patent Text Reader

Abstract

The present application provides a risk quantification and assessment method and system based on a subject association network. The method includes: obtaining public opinion text data and a subject association network; inputting the public opinion text data into a preset credit risk scoring model, so that the credit risk scoring model extracts a number of risk subjects from the public opinion text data and generates an initial credit risk score corresponding to each of the risk subjects; inputting the subject association network into a pre-trained double-layer graph attention network, so that the double-layer graph attention network predicts the association relationship and association strength between each of the subjects to obtain an enhanced subject association network; according to the enhanced subject association network and a preset risk transmission rule, performing risk transmission on each of the initial credit risk scores starting from their respective corresponding risk subjects to obtain a risk quantification and assessment result for each subject, improving the accuracy and efficiency of risk quantification and assessment.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of credit risk transmission calculation, and particularly to a risk quantification assessment method and system based on a subject association network. Background Art

[0002] Credit risk does not exist independently, but there is a transmission effect. With the frequent occurrence of credit risk and the improvement of regulatory standards, the market's demand for credit early warning is increasing day by day. Merely focusing on the credit status of a single enterprise without considering the spread of risk among key enterprises may lead to misjudgment of the enterprise's credit level and may miss the golden opportunity to handle related risks. The risk quantification assessment scheme based on the subject association network attempts to effectively transmit risk indicators among different credit subjects in a network-based manner, thereby improving the accuracy and interpretability of enterprise risk identification. At the same time, it can monitor the bond market from a systematic perspective, achieving more accurate, comprehensive, and timely monitoring.

[0003] However, due to the intricate association relationships of financial subjects, there are often multiple paths for credit risk transmission. Currently, the calculation method for credit risk transmission highly depends on the initially defined subject association network. When the subject status changes, the entire network needs to be reconstructed to update the nodes. Real-time monitoring of subject status changes and path updates may require a large amount of computing resources. In addition, the subject association network is mainly constructed by manually collecting data related to target enterprises and risks. The construction process highly depends on the professional experience and level of technicians. When technicians make misjudgments, the constructed subject association network is difficult to ensure the accuracy of risk quantification assessment. At the same time, the news data related to financial subjects may contain risk information of the financial subject, which helps us accurately evaluate the credit level of the financial subject. Therefore, how to quickly identify and extract the risk information of the target financial subject from a large amount of text data to provide a factual basis for the credit risk assessment of the target financial subject is also a key problem to be solved in this field. Summary of the Invention

[0004] In view of the above technical problems, this application provides a risk quantification assessment method and system based on a subject association network, which avoids the problem of high dependence on the initial network by enhancing the subject association network, and improves the accuracy and efficiency of risk quantification assessment.

[0005] In a first aspect, this application provides a risk quantification assessment method based on a subject association network, including:

[0006] Obtain public opinion text data and a subject association network, where the subject association network includes several subjects to be evaluated and the initial association relationships between each subject;

[0007] Input the public opinion text data into a preset credit risk scoring model, so that the credit risk scoring model extracts several risk entities from the public opinion text data, and conducts risk assessment on each risk entity according to the public opinion content in the public opinion text data, and generates an initial credit risk score corresponding to each risk entity. Among them, the risk scoring model is obtained by training based on the BERT-base model through a number of historical news data;

[0008] Input the entity association network into a pre-trained double-layer graph attention network, so that the double-layer graph attention network predicts the association relationship and association strength between each entity, and obtains an enhanced entity association network. Among them, the double-layer graph attention network is obtained by training through a number of historical research reports containing each entity;

[0009] According to the enhanced entity association network and the preset risk transfer rules, transfer the initial credit risk scores from their respective corresponding risk entities, and obtain the risk quantitative assessment results of each entity.

[0010] The embodiment of the present application provides a risk quantitative assessment method based on an entity association network. The risk entities in the public opinion text data are automatically identified through a credit risk scoring model, and the initial credit risk scores of each risk entity are generated according to the public opinion text data. Finally, risk transfer is carried out based on the enhanced entity association network to obtain the risk quantitative assessment results of each entity, improving the accuracy of risk quantitative assessment. At the same time, in order to avoid the problem of high dependence on the initial network in risk transfer, the embodiment of the present application uses a double-layer graph attention network to further predict the association relationship between each entity and obtains an enhanced entity association network. In the embodiment of the present application, even if the entity state changes, such as when new link edges or new entities are added between entities, there is no need to reconstruct the entire network. Only corresponding modifications need to be made in the entity association network and then input into the double-layer graph attention network, and the link edges between the new entity and other entities can be automatically predicted, improving the efficiency of risk quantitative assessment.

[0011] In a possible implementation manner, the risk scoring model obtained by training the BERT-base model through a number of historical news data includes:

[0012] Add a linear layer or a softmax layer on top of the BERT-base model, and set the loss function as the cross-entropy loss function to construct an initial risk scoring model;

[0013] Obtain a number of historical news data and perform annotation;

[0014] Perform word segmentation, stop word removal, and modal particle preprocessing operations on each historical news data to obtain corresponding training data;

[0015] Input each of the training data into the initial risk scoring model, so that the initial risk scoring model fine-tunes the model parameters through backpropagation according to the calculation result of the loss function to obtain the risk scoring model.

[0016] The embodiment of the present application provides a training method for a risk scoring model. First, an initial risk scoring model is constructed based on the BERT-base model, and a linear layer or a softmax layer is added for outputting risk categories. Then, a number of historical data are labeled and preprocessed to obtain a number of training data. The training data are used to train the initial risk scoring model, and the model parameters are fine-tuned through backpropagation, so that the model has the ability of text recognition and risk scoring. Furthermore, the risk scoring model can be used to monitor the public opinion text data in real time and perform credit risk scoring on the risk entities appearing in the public opinion text data, improving the efficiency of real-time risk monitoring and providing data support for subsequent risk quantification assessment.

[0017] Further, the double-layer graph attention network predicts the association relationship and association strength between each of the entities to obtain an enhanced entity association network, including:

[0018] Encode the entity association network through an encoder to obtain a feature vector of a preset dimension, where the encoder is composed of double-layer graph attention layers;

[0019] According to the feature vector, calculate the feature dot product between each entity through a decoder;

[0020] Construct a probability adjacency matrix for representing the association probability between each entity according to each feature dot product;

[0021] Generate a number of link edges between each entity according to the probability adjacency matrix, and determine the association strength of each link edge to obtain the enhanced entity association network.

[0022] This application embodiment discloses the specific process of enhancing the subject association network using a double-layer graph attention network. First, a feature vector in the subject association network is extracted by an encoder composed of double-layer graph attention layers. Then, according to the feature vector, a decoder is used to calculate the feature dot product between each subject, and further a probability adjacency matrix for representing the association probability between each subject is constructed. Based on the probability adjacency matrix, link edges between each subject can be predicted and generated. For example, if the generation threshold of the link edge is set to 0.5 in advance, the double-layer graph attention network will generate link edges for two subjects corresponding to the matrix cells in the probability adjacency matrix that are greater than 0.5. At the same time, the association strength of each link edge is determined according to the probability adjacency matrix, and finally the enhanced subject association network is generated. By enhancing the subject association network, this application embodiment avoids the problem of high dependence on the initial network in the prior art and provides data support for subsequent risk transmission.

[0023] In a possible implementation manner, training the double-layer graph attention network through a number of historical research reports including each of the subjects includes:

[0024] Construct a training network according to each historical research report;

[0025] Set negative link edges and positive link edges in the training network;

[0026] Construct an initial double-layer graph attention network, and use the binary cross-entropy between the positive link edges and negative link edges generated by the initial double-layer graph attention network as the loss function;

[0027] Input the training network into the initial double-layer graph attention network, so that the initial double-layer graph attention network updates its own parameters through backpropagation according to the loss function to obtain the double-layer graph attention network.

[0028] This application embodiment provides a training method for a double-layer graph attention network. By mining the association relationships of each subject from each historical research report, constructing a training network, setting negative link edges and positive link edges in the training network as sample labels of training data, and training the link edge prediction of the initial double-layer graph attention network in combination with the loss function, the initial double-layer graph attention network continuously updates its own parameters during the training process, and finally obtains the double-layer graph attention network. By updating the model's own parameters through backpropagation during the training process, this application embodiment improves the efficiency of model training. At the same time, a training network with negative link edges and positive link edges is used to train the model, which improves the accuracy of subsequent model prediction.

[0029] Further, according to the enhanced entity association network and the preset risk transmission rules, the initial credit risk scores are transmitted from their respective corresponding risk entities to obtain the risk quantification evaluation results of each entity, including:

[0030] Determine several secondary risk entities corresponding to each risk entity according to the enhanced entity association network and the preset number of risk transmission layers, where a certain risk entity can be simultaneously used as the secondary risk entity of other risk entities;

[0031] Calculate the credit risk scores of each secondary risk entity by weighted summation according to the association strength between each risk entity and its corresponding several secondary risk entities, and then obtain the risk quantification evaluation results of each entity.

[0032] In the embodiment of the present application, the preset number of risk transmission layers determines the risk transmission range of each risk entity, and the enhanced entity association network determines the risk transmission path of each risk entity. Therefore, first, several secondary risk entities corresponding to each risk entity are determined according to the enhanced entity association network and the preset number of risk transmission layers, and the initial credit risk scores will be transmitted from each risk entity to its corresponding secondary risk entity. Then, in the embodiment of the present application, the association strength between the risk entity and the corresponding secondary risk entity is used as the weight, and the credit risk scores of each secondary risk entity are calculated by weighted summation to achieve the transmission of risks and obtain the risk quantification evaluation results of each entity. It should be noted that in the embodiment of the present application, a risk entity can be simultaneously used as the secondary risk entity of other risk entities, that is, when two risk entities are within the preset number of risk transmission layers, it indicates that there is a relatively close connection between the two risk entities in reality. Therefore, in the risk quantification evaluation process of the present application, the two entities will affect each other and transmit credit risks to each other, further improving the accuracy and rationality of the risk quantification evaluation.

[0033] In the second aspect, the present application provides a risk quantification evaluation system based on an entity association network, including an acquisition module, a risk evaluation module, a network enhancement module, and a risk transmission module;

[0034] The acquisition module is used to acquire public opinion text data and an entity association network, where the entity association network includes several entities to be evaluated and the initial association relationships between each entity;

[0035] The risk assessment module is used to input the public opinion text data into a preset credit risk scoring model, so that the credit risk scoring model extracts several risk entities from the public opinion text data, and conducts risk assessment on each of the risk entities according to the public opinion content in the public opinion text data, generating an initial credit risk score corresponding to each of the risk entities. Among them, the risk scoring model is obtained by training based on the BERT-base model through a number of historical news data;

[0036] The network enhancement module is used to input the subject association network into a pre-trained double-layer graph attention network, so that the double-layer graph attention network predicts the association relationship and association strength between each of the subjects, obtaining an enhanced subject association network. Among them, the double-layer graph attention network is obtained by training through a number of historical research reports containing each of the subjects;

[0037] The risk transmission module is used to perform risk transmission on each of the initial credit risk scores starting from their respective corresponding risk entities according to the enhanced subject association network and a preset risk transmission rule, obtaining a risk quantitative assessment result for each subject.

[0038] In a possible implementation manner, the risk quantitative assessment system further includes a model training module, and the model training module is used to obtain the risk scoring model by training based on the BERT-base model through a number of historical news data, including a model construction unit, an acquisition unit, a data preprocessing unit, and a model training unit;

[0039] Among them, the model construction unit is used to add a linear layer or a softmax layer on the top of the BERT-base model, and set the loss function as the cross-entropy loss function to construct an initial risk scoring model;

[0040] The acquisition unit is used to acquire a number of historical news data and perform annotation;

[0041] The data preprocessing unit is used to perform word segmentation, stop word removal, and modal particle preprocessing operations on each of the historical news data to obtain corresponding training data;

[0042] The model training unit is used to input each of the training data into the initial risk scoring model, so that the initial risk scoring model fine-tunes the model parameters through backpropagation according to the calculation result of the loss function, obtaining the risk scoring model.

[0043] Furthermore, the network enhancement module includes an encoding unit, a decoding unit, a probability adjacency matrix construction unit, and a network enhancement unit;

[0044] Among them, the encoding unit is used to encode the main body association network through an encoder to obtain a feature vector of a preset dimension, where the encoder is composed of a two-layer graph attention layer;

[0045] The decoding unit is used to calculate the feature dot product between each main body according to the feature vector through a decoder;

[0046] The probability adjacency matrix construction unit is used to construct a probability adjacency matrix for representing the association probability between each main body according to each feature dot product;

[0047] The network enhancement unit is used to generate a number of link edges between each main body according to the probability adjacency matrix and determine the association strength of each link edge to obtain the enhanced main body association network.

[0048] In a possible implementation manner, the two-layer graph attention network is obtained by training with a number of historical research reports each containing each of the main bodies, including:

[0049] Construct a training network according to each historical research report;

[0050] Set negative link edges and positive link edges in the training network;

[0051] Construct an initial two-layer graph attention network, and use the binary cross-entropy between the positive link edges and negative link edges generated by the initial two-layer graph attention network as the loss function;

[0052] Input the training network into the initial two-layer graph attention network, so that the initial two-layer graph attention network updates its own parameters through backpropagation according to the loss function to obtain the two-layer graph attention network.

[0053] Furthermore, the risk transfer module transfers the initial credit risk scores of each of the main bodies starting from their respective corresponding risk main bodies according to the enhanced main body association network and a preset risk transfer rule to obtain the risk quantification evaluation results of each main body, including:

[0054] Determine a number of secondary risk main bodies corresponding to each risk main body according to the enhanced main body association network and a preset risk transfer layer number, where a certain risk main body can simultaneously be a secondary risk main body of other risk main bodies;

[0055] Calculate the credit risk scores of each of the secondary risk main bodies by means of weighted summation according to the association strength between each risk main body and its corresponding number of secondary risk main bodies, and further obtain the risk quantification evaluation results of each main body. Description of the Drawings

[0056] Figure 1: Schematic flowchart of a risk quantification and assessment method based on a subject association network provided by an embodiment of the present application.

[0057] Figure 2 : Schematic flowchart of a brief risk quantification and assessment method based on a subject association network provided by an embodiment of the present application.

[0058] Figure 3 : Schematic flowchart of enhancing a subject association network using a two - layer graph attention network in a risk quantification and assessment method based on a subject association network provided by an embodiment of the present application.

[0059] Figure 4 : Schematic diagram of the structure of a two - layer graph attention network in a risk quantification and assessment method based on a subject association network provided by an embodiment of the present application.

[0060] Figure 5 : Schematic diagram of risk transmission based on an enhanced subject association network in a risk quantification and assessment method based on a subject association network provided by an embodiment of the present application.

[0061] Figure 6 : Schematic diagram of the experimental results of a risk quantification and assessment method based on a subject association network provided by an embodiment of the present application.

[0062] Figure 7 : Schematic diagram of the structure of a risk quantification and assessment system based on a subject association network provided by an embodiment of the present application.

[0063] Figure 8 : Schematic diagram of the structure of the model training module of a risk quantification and assessment system based on a subject association network provided by an embodiment of the present application.

[0064] Figure 9 : Schematic diagram of the structure of the network enhancement module of a risk quantification and assessment system based on a subject association network provided by an embodiment of the present application. Detailed implementation manners

[0065] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art without making creative efforts based on the embodiments of the present application belong to the scope of protection of the present application.

[0066] It should be noted that the step numbers in the text are only for the convenience of explaining specific embodiments and do not serve to limit the execution order of the steps. In the description of the present application, the terms "first" and "second" are only used for descriptive purposes and cannot be construed as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, features defined with "first" and "second" may explicitly or implicitly include one or more of such features.

[0067] Throughout this specification, the public opinion text data described in this specification refers to text data such as news, financial research reports, and article reviews published publicly within a certain period of time. The risk entities described in this specification are related enterprises involved in the public opinion text data and are simultaneously one of several entities to be evaluated. The double-layer graph attention network (Graph Attention Network, GAT) described in this specification is a neural network model for graph-structured data that can capture complex relationships between nodes. The entire network structure consists of an encoder and a decoder. The encoder structure is mainly composed of two main graph attention layers (GATConv), and each layer uses different attention heads and parameter settings.

[0068] Embodiment 1:

[0069] As Figure 1 shown, Embodiment 1 provides a risk quantification and assessment method based on a subject association network, including steps S1 - S4:

[0070] Step S1, obtain public opinion text data and a subject association network, where the subject association network includes several subjects to be evaluated and the initial association relationships between each subject;

[0071] Step S2, input the public opinion text data into a preset credit risk scoring model, so that the credit risk scoring model extracts several risk entities from the public opinion text data, and conducts risk assessments on each of the risk entities according to the public opinion content in the public opinion text data, generating an initial credit risk score corresponding to each of the risk entities. Among them, the risk scoring model is obtained by training based on the BERT-base model through several historical news data;

[0072] Step S3, input the subject association network into a pre-trained double-layer graph attention network, so that the double-layer graph attention network predicts the association relationships and association intensities between each of the subjects, obtaining an enhanced subject association network. Among them, the double-layer graph attention network is obtained by training through several historical research reports including each of the subjects;

[0073] Step S4: Based on the enhanced subject association network and the preset risk transfer rules, each of the initial credit risk scores is transferred from the corresponding risk subject to obtain a quantitative risk assessment result for each subject.

[0074] For ease of understanding, a brief flowchart of an embodiment of the present application is shown in the following figure. Figure 2 shown.

[0075] The embodiment of the present application provides a risk quantification assessment method based on a subject association network, which automatically identifies risk subjects in public opinion text data through a credit risk scoring model, generates an initial credit risk score for each risk subject based on the public opinion text data, and finally performs risk transfer based on an enhanced subject association network to obtain risk quantification assessment results for each subject, thereby improving the accuracy of risk quantification assessment. At the same time, in order to avoid the problem of high dependence on the initial network in risk transfer, the embodiment of the present application uses a two-layer graph attention network to further predict the association relationship between each subject and obtain an enhanced subject association network. In the embodiment of the present application, even if the subject state changes, such as when a new link edge is added between subjects or a new subject is added, there is no need to rebuild the entire network. It only needs to make corresponding modifications in the subject association network and then input the two-layer graph attention network to automatically predict the link edge between the new subject and other subjects, thereby improving the efficiency of risk quantification assessment.

[0076] Among them, the subject association network described in step S1 can be the initially defined subject association network, or it can be the subject association network after the change. For example, in a preferred embodiment, the subject association network is obtained by constructing the research reports from January to March 2024, and a total of 1932 bond issuers and 9163 associations are extracted, with a link coverage rate of 0.2%. At this time, in order to explore potential associations and make full use of subject associations, we use a double-layer graph attention network to enhance the subject association network and mine more subject associations. In another scenario, when the subject association network that has been fully mined has undergone structural changes (for example, two companies have increased cooperation, or a new company has been added), it is only necessary to add edges or nodes to the existing network, and the double-layer graph attention network is also used for enhancement to predict and generate more related edges. Therefore, this application does not depend on whether the initially defined subject association network is complete, and can realize the automatic enhancement and update of the subject association network. Accordingly, in order to improve the accuracy of enhancing the subject association network, the double-layer graph attention network will be retrained with new training data at intervals, and the interval is generally one month.

[0077] In step S2, the credit risk scoring model extracts a number of risk entities from the public opinion text data. The specific process is as follows: First, perform word segmentation, stop word removal, and modal particle removal preprocessing operations on each piece of the public opinion text data to obtain preprocessed text data. Then, based on a preset text recognition algorithm, according to the names of each entity in the entity association network, extract the corresponding number of risk entities from the text data. To avoid extracting a large number of irrelevant entity names, generally, at most one risk entity is extracted from each piece of text data, and those skilled in the art can set it according to the actual situation.

[0078] In a possible implementation manner, in step S2, the risk scoring model is obtained by training the BERT-base model with a number of historical news data, including:

[0079] Add a linear layer or a softmax layer on top of the BERT-base model, and set the loss function as the cross-entropy loss function to construct an initial risk scoring model;

[0080] Obtain a number of historical news data and perform annotation;

[0081] Perform word segmentation, stop word removal, and modal particle removal preprocessing operations on each piece of the historical news data to obtain corresponding training data;

[0082] Input each of the training data into the initial risk scoring model, so that the initial risk scoring model fine-tunes the model parameters through backpropagation according to the calculation result of the loss function to obtain the risk scoring model.

[0083] In a preferred embodiment, for credit risk, we annotated 200,000 pieces of Caihui historical news data, and performed word segmentation, stop word removal, and modal particle removal preprocessing operations on the data. The pre-trained model of BERT-base is adopted, and a linear layer or a softmax layer is added on top of the BERT-base model to output the risk category. The loss function of the model is defined as cross-entropy loss. The training process first loads the cleaned and annotated data into the training environment, and after batch processing, fine-tunes the parameters through backpropagation as the final BERT model.

[0084] The embodiment of the present application provides a training method for a risk scoring model. First, an initial risk scoring model is constructed based on the BERT-base model, and a linear layer or a softmax layer is added to output risk categories. Then, a number of historical data are labeled and preprocessed to obtain a number of training data. The initial risk scoring model is trained using the training data, and the model parameters are fine-tuned through backpropagation, so that the model has the ability of text recognition and risk scoring. Furthermore, the risk scoring model can be used to monitor public opinion text data in real time and perform credit risk scoring on the risk subjects that appear in the public opinion text data, improving the efficiency of real-time risk monitoring and providing data support for subsequent risk quantification assessment.

[0085] Further, in step S3, the double-layer graph attention network predicts the association relationship and association strength between each of the subjects, and obtains an enhanced subject association network, as Figure 3 shown, including steps S301-S304:

[0086] Step S301, encode the subject association network through an encoder to obtain a feature vector of a preset dimension, where the encoder is composed of double-layer graph attention layers;

[0087] Step S302, calculate the feature dot product between each subject according to the feature vector through a decoder;

[0088] Step S303, construct a probability adjacency matrix for representing the association probability between each subject according to each feature dot product;

[0089] Step S304, generate a number of link edges between each subject according to the probability adjacency matrix, and determine the association strength of each link edge to obtain the enhanced subject association network.

[0090] The embodiment of the present application discloses the specific process of enhancing the subject association network using a double-layer graph attention network. First, a feature vector in the subject association network is extracted through an encoder composed of double-layer graph attention layers, and then the feature dot product between each subject is calculated using a decoder according to the feature vector, and further a probability adjacency matrix for representing the association probability between each subject is constructed. Based on the probability adjacency matrix, link edges between each subject can be predicted and generated. For example, if the generation threshold of the link edge is set to 0.5 in advance, the double-layer graph attention network will generate link edges for two subjects corresponding to the matrix cells in the probability adjacency matrix that are greater than 0.5, and at the same time determine the association strength of each link edge according to the probability adjacency matrix, and finally generate the enhanced subject association network. The embodiment of the present application avoids the problem of high dependence on the initial network in the prior art by enhancing the subject association network, and provides data support for subsequent risk transmission.

[0091] In a possible implementation manner, in step S3, the training of the double-layer graph attention network by using a plurality of historical research reports each including each of the subjects includes:

[0092] Construct a training network according to each historical research report;

[0093] Set negative link edges and positive link edges in the training network;

[0094] Construct an initial double-layer graph attention network, and use the binary cross-entropy between the positive link edges and negative link edges generated by the initial double-layer graph attention network as a loss function;

[0095] Input the training network into the initial double-layer graph attention network, so that the initial double-layer graph attention network updates its own parameters through backpropagation according to the loss function, and obtains the double-layer graph attention network.

[0096] The embodiment of the present application provides a training method for a double-layer graph attention network. By mining the association relationships of each subject from each historical research report, a training network is constructed, and negative link edges and positive link edges in the training network are set as sample labels of training data. Combining with the loss function, the link edge prediction of the initial double-layer graph attention network is trained, so that the initial double-layer graph attention network continuously updates its own parameters during the training process, and finally obtains the double-layer graph attention network. By updating the model's own parameters through backpropagation during the training process, the embodiment of the present application improves the efficiency of model training. At the same time, a training network with negative link edges and positive link edges is used to train the model, which improves the accuracy of the subsequent prediction of the model.

[0097] In a preferred embodiment, as Figure 4 shown, the structure and parameter settings of the double-layer graph attention network are as follows:

[0098] 1. The first-layer encoder (conv1):

[0099] Number of input channels: N*N (number of subject nodes)

[0100] Number of output channels: 68

[0101] Number of attention heads: 8

[0102] Dropout probability: 0.6

[0103] 2. The second-layer encoder (conv2):

[0104] Number of input channels: the number of output channels of the first layer multiplied by the number of attention heads (i.e., 68*8)

[0105] Number of output channels: 16

[0106] Number of attention heads: 1

[0107] Whether to splice: True

[0108] Dropout probability: 0.6

[0109] The above parameters are the optimal values in the optional parameter set found through grid search and combined with cross - validation during the training process.

[0110] Decoder: Calculate the dot product of the low - dimensional representations of the nodes in the encoder output, and then sum to obtain the score of the link. By calculating the dot product of the feature vectors of all node pairs, a probability adjacency matrix is obtained. Then, according to the threshold of 0.5, the predicted edges are generated, and the predicted network result, that is, the enhanced subject association network, is obtained. The specific process is as follows: For each pair of nodes in the graph, calculate the dot product of their feature vectors. The result of the dot product can reflect the similarity between two nodes. Normalize the result of the dot product so that it falls within the interval [0, 1]. This is achieved by dividing by the norm (i.e., Euclidean norm) of the two feature vectors. Use the logistic function to map the dot product result to the interval [0, 1]. If the probability after mapping the dot product result of two nodes exceeds 0.5, record this probability value at the corresponding position in the adjacency matrix; otherwise, record 0. After traversing all dot product results, the probability adjacency matrix can be constructed. Traverse the probability adjacency matrix. If there is a value at a certain position in the probability adjacency matrix, generate a link edge between the corresponding two nodes, and the corresponding probability value is the association strength of the link edge, thereby generating the enhanced subject association network.

[0111] Loss function: First, set the links between significantly unrelated subjects as negative link edges, and sample some edges of the input network as positive sample edges. Execute the link score after the encoder and the encoder output, calculate the binary cross - entropy between the generated links and the positive and negative link edges as the loss function, and update the model parameters through backpropagation.

[0112] Further, in step S4, according to the enhanced subject association network and the preset risk transfer rules, the initial credit risk scores are transferred from their respective corresponding risk subjects to obtain the risk quantification evaluation results of each subject, including:

[0113] According to the enhanced subject association network and the preset number of risk transfer layers, determine several secondary risk subjects corresponding to each risk subject, where a certain risk subject can be a secondary risk subject of other risk subjects at the same time;

[0114] According to the association strength between each of the risk entities and several corresponding secondary risk entities, calculate the credit risk scores of each of the secondary risk entities by means of weighted summation, and then obtain the risk quantification evaluation results of each entity.

[0115] In the embodiment of the present application, the preset number of risk transmission layers determines the risk transmission range of each risk entity, and the enhanced entity association network determines the risk transmission path of each risk entity. Therefore, first determine several secondary risk entities corresponding to each risk entity according to the enhanced entity association network and the preset number of risk transmission layers, and the initial credit risk scores will be transmitted from each risk entity to its corresponding secondary risk entities. Then, in the embodiment of the present application, the association strength between the risk entity and the corresponding secondary risk entity is used as the weight, and the credit risk scores of each secondary risk entity are calculated by means of weighted summation to achieve the transmission of risks and obtain the risk quantification evaluation results of each entity. It should be noted that in the embodiment of the present application, a risk entity can also be a secondary risk entity of other risk entities at the same time, that is, when two risk entities are within the preset number of risk transmission layers, it indicates that there is a relatively close connection between these two risk entities in reality. Therefore, in the risk quantification evaluation process of the present application, the two entities will affect each other and transmit credit risks to further improve the accuracy and rationality of the risk quantification evaluation.

[0116] In a preferred embodiment, the risk transmission process based on the enhanced entity association network is as Figure 5 shown. Figure 5 shows a part of the enhanced entity association network, which includes three enterprises and the association relationships and association strengths between each enterprise. The preset number of risk transmission layers is 1. The risk entities extracted from the two pieces of public opinion text data are enterprise A and enterprise B respectively, and the corresponding initial credit risk scores are 0.89 and 0.09. At this time, the secondary risk entities corresponding to enterprise A are enterprise B and enterprise C, and the secondary risk entity corresponding to enterprise B is enterprise A. Therefore, in the process of risk transmission, enterprise A is affected by the risk transmission of enterprise B on the basis of its own initial credit risk score, and enterprise B is also affected by the risk transmission of enterprise A on the basis of its own initial credit risk score. Even though the two pieces of public opinion text data do not involve enterprise C, since enterprise C has an association relationship with enterprise A, enterprise C is also affected by the risk transmission of enterprise A. The finally calculated risk quantification evaluation results are as Figure 5 shown.

[0117] Furthermore, the embodiment of the present application can construct enterprise portraits corresponding to each entity according to the risk quantification evaluation results of each entity, and at the same time classify the credit grades of each entity in combination with specific business scenarios to provide data support for the development and decision-making of specific businesses.

[0118] Further, in order to verify the effect of the method provided in this application, a downstream default prediction experiment was conducted in the embodiments of this application. First, a risk conduction model was constructed based on the risk transfer rules used in this application. The risk conduction model realized the risk transfer of the initial credit risk score in the entity association network or the enhanced entity association network. Then, a time series model for default prediction based on public opinion quantification indicators internally was adopted to test the risk conduction model, and it was found that the risk conduction model based on the entity association network (referred to as LLM-net for short) could increase the prediction AUC from 0.91 to 0.92. While the AUC of the risk conduction model based on the enhanced entity association network (referred to as GAT-net for short) could be increased to 0.95, exceeding the effect of the manually labeled industrial chain network (referred to as IC-net for short). The experimental results are as Figure 6 shown. It shows that in addition to the data of the industrial chain, the research report information contains more dimensions of association information, and this information helps the downstream task of default prediction. The double-layer graph attention network provided in this application can effectively mine these potential information.

[0119] Embodiment 2:

[0120] As Figure 7 shown, Embodiment 2 provides a risk quantification and assessment system based on an entity association network, including an acquisition module 10, a risk assessment module 20, a network enhancement module 30, and a risk transfer module 40;

[0121] Among them, the acquisition module 10 is used to acquire public opinion text data and an entity association network, where the entity association network includes a number of entities to be evaluated and the initial association relationships between each entity;

[0122] The risk assessment module 20 is used to input the public opinion text data into a preset credit risk scoring model, so that the credit risk scoring model extracts a number of risk entities from the public opinion text data, and conducts risk assessment on each of the risk entities according to the public opinion content in the public opinion text data, and generates an initial credit risk score corresponding to each of the risk entities. Among them, the risk scoring model is obtained by training through a number of historical news data based on the BERT-base model;

[0123] The network enhancement module 30 is used to input the entity association network into a pre-trained double-layer graph attention network, so that the double-layer graph attention network predicts the association relationships and association strengths between each entity, and obtains an enhanced entity association network. Among them, the double-layer graph attention network is obtained by training through a number of historical research reports including each entity;

[0124] The risk transfer module 40 is used to perform risk transfer on each of the initial credit risk scores starting from their respective corresponding risk entities according to the enhanced entity association network and a preset risk transfer rule, and obtain the risk quantification evaluation results of each entity.

[0125] In a possible implementation manner, as Figure 8 shown, the risk quantification evaluation system further includes a model training module 50. The model training module 50 is used to train the risk scoring model based on the BERT-base model through a number of historical news data, including a model construction unit 501, an acquisition unit 502, a data preprocessing unit 503, and a model training unit 504;

[0126] Among them, the model construction unit 501 is used to add a linear layer or a softmax layer on top of the BERT-base model, and set the loss function as the cross-entropy loss function to construct an initial risk scoring model;

[0127] The acquisition unit 502 is used to acquire a number of historical news data and perform annotation;

[0128] The data preprocessing unit 503 is used to perform word segmentation, stop word removal, and modal particle preprocessing operations on each piece of the historical news data to obtain corresponding training data;

[0129] The model training unit 504 is used to input each of the training data into the initial risk scoring model, so that the initial risk scoring model fine-tunes the model parameters through backpropagation according to the calculation result of the loss function, and obtains the risk scoring model.

[0130] Furthermore, as Figure 9 shown, the network enhancement module 30 includes an encoding unit 301, a decoding unit 302, a probability adjacency matrix construction unit 303, and a network enhancement unit 304;

[0131] Among them, the encoding unit 301 is used to encode the entity association network through an encoder to obtain a feature vector of a preset dimension, where the encoder is composed of a double-layer graph attention layer;

[0132] The decoding unit 302 is used to calculate the feature dot product between each entity according to the feature vector;

[0133] The probability adjacency matrix construction unit 303 is used to construct a probability adjacency matrix for representing the association probability between each entity according to each feature dot product;

[0134] The network enhancement unit 304 is used to generate a number of link edges between each subject according to the probability adjacency matrix, determine the association strength of each link edge, and obtain the enhanced subject association network.

[0135] In a possible implementation manner, the training of the double-layer graph attention network through a number of historical research reports including each of the subjects includes:

[0136] Construct a training network according to each historical research report;

[0137] Set negative link edges and positive link edges in the training network;

[0138] Construct an initial double-layer graph attention network, and use the binary cross-entropy between the positive link edges and negative link edges generated by the initial double-layer graph attention network as the loss function;

[0139] Input the training network into the initial double-layer graph attention network, so that the initial double-layer graph attention network updates its own parameters through backpropagation according to the loss function, and obtain the double-layer graph attention network.

[0140] Furthermore, the risk transfer module 40 transfers risks starting from each of the initial credit risk scores from their respective corresponding risk subjects according to the enhanced subject association network and a preset risk transfer rule, and obtains the risk quantification evaluation results of each subject, including:

[0141] Determine a number of secondary risk subjects corresponding to each of the risk subjects according to the enhanced subject association network and a preset risk transfer layer number, where a certain risk subject can be used as a secondary risk subject of other risk subjects at the same time;

[0142] Calculate the credit risk scores of each of the secondary risk subjects by means of weighted summation according to the association strength between each of the risk subjects and their respective corresponding a number of secondary risk subjects, and further obtain the risk quantification evaluation results of each subject.

[0143] The embodiment of the present application provides a risk quantification evaluation system based on a subject association network, which automatically identifies risk subjects in public opinion text data through a credit risk scoring model, generates initial credit risk scores for each risk subject according to the public opinion text data, and finally performs risk transmission based on an enhanced subject association network to obtain the risk quantification evaluation results of each subject, improving the accuracy of risk quantification evaluation. At the same time, in order to avoid the problem of high dependence on the initial network in risk transmission, the embodiment of the present application uses a double-layer graph attention network to further predict the association relationship between each subject and obtains an enhanced subject association network. In the embodiment of the present application, even when the subject state changes, such as when new link edges or new subjects are added between subjects, there is no need to reconstruct the entire network. Only corresponding modifications need to be made in the subject association network and then input into the double-layer graph attention network, and the link edges between the new subject and other subjects can be automatically predicted, improving the efficiency of risk quantification evaluation.

[0144] The more detailed working principle and step flow of this embodiment can, but are not limited to, refer to the relevant records of Embodiment 1.

[0145] The specific embodiments described above have further elaborated on the purpose, technical solutions, and beneficial effects of the present application. It should be understood that the above description is only for the specific embodiments of the present application and is not used to limit the protection scope of the present application. In particular, for those skilled in the art, any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present application shall be included in the protection scope of the present application.

Claims

1. A risk quantification and assessment method based on a subject association network, characterized in that, Including: Obtain public opinion text data and a subject association network, where the subject association network includes a number of subjects to be evaluated and initial association relationships between each subject; Input the public opinion text data into a preset credit risk scoring model, so that the credit risk scoring model extracts a number of risk subjects from the public opinion text data, and conducts risk assessment on each of the risk subjects according to the public opinion content in the public opinion text data, generating an initial credit risk score corresponding to each of the risk subjects. Among them, the risk scoring model is obtained by training based on the BERT-base model through a number of historical news data; Input the subject association network into a pre-trained double-layer graph attention network, so that the double-layer graph attention network predicts the association relationships and association strengths between each subject, obtaining an enhanced subject association network. Among them, the double-layer graph attention network is obtained by training through a number of historical research reports including each subject; The step of obtaining the double-layer graph attention network by training through a number of historical research reports including each subject includes: constructing a training network according to each historical research report; setting negative link edges and positive link edges in the training network; constructing an initial double-layer graph attention network, and taking the binary cross-entropy between the positive link edges and negative link edges generated by the initial double-layer graph attention network as the loss function; inputting the training network into the initial double-layer graph attention network, so that the initial double-layer graph attention network updates its own parameters through backpropagation according to the loss function, obtaining the double-layer graph attention network; According to the enhanced subject association network and a preset risk transmission rule, perform risk transmission on each of the initial credit risk scores starting from their corresponding risk subjects, obtaining a risk quantitative assessment result for each subject.

2. The risk quantification and assessment method based on a subject association network according to claim 1, wherein The step of obtaining the risk scoring model by training based on the BERT-base model through a number of historical news data includes: Adding a linear layer or a softmax layer on top of the BERT-base model, and setting the loss function as the cross-entropy loss function to construct an initial risk scoring model; Obtain a number of historical news data and perform annotation; Perform word segmentation, stop word removal, and modal particle preprocessing operations on each piece of the historical news data to obtain corresponding training data; Input each of the training data into the initial risk scoring model, so that the initial risk scoring model fine-tunes the model parameters through backpropagation according to the calculation result of the loss function, obtaining the risk scoring model.

3. A risk quantification and assessment method based on a subject association network according to claim 1, characterized in that The step that the double-layer graph attention network predicts the association relationships and association strengths between each subject, obtaining an enhanced subject association network, includes: Encoding the subject association network through an encoder to obtain a feature vector of a preset dimension, where the encoder is composed of double-layer graph attention layers; According to the feature vector, calculating the feature dot product between each subject through a decoder; Constructing a probability adjacency matrix for representing the association probability between each subject according to each feature dot product; Generate a number of link edges between each subject according to the probability adjacency matrix, and determine the association strength of each link edge to obtain the enhanced subject association network.

4. A risk quantification and assessment method based on a subject association network according to any one of claims 1-3, characterized in that, According to the enhanced subject association network and the preset risk transfer rules, perform risk transfer on each of the initial credit risk scores starting from their respective corresponding risk subjects to obtain the risk quantitative assessment results of each subject, including: Determine a number of secondary risk subjects corresponding to each risk subject according to the enhanced subject association network and the preset number of risk transfer layers, where a certain risk subject can simultaneously be a secondary risk subject of other risk subjects; Calculate the credit risk scores of each of the secondary risk subjects by weighted summation according to the association strength between each risk subject and its corresponding number of secondary risk subjects, and then obtain the risk quantitative assessment results of each subject.

5. A risk quantification and assessment system based on a subject association network, characterized in that, It includes an acquisition module, a risk assessment module, a network enhancement module, and a risk transfer module; Among them, the acquisition module is used to acquire public opinion text data and a subject association network, where the subject association network includes a number of subjects to be evaluated and the initial association relationships between each subject; The risk assessment module is used to input the public opinion text data into a preset credit risk scoring model, so that the credit risk scoring model extracts a number of risk subjects from the public opinion text data, and performs risk assessment on each of the risk subjects according to the public opinion content in the public opinion text data, and generates an initial credit risk score corresponding to each of the risk subjects. Among them, the risk scoring model is obtained by training through a number of historical news data based on the BERT-base model; The network enhancement module is used to input the subject association network into a pre-trained double-layer graph attention network, so that the double-layer graph attention network predicts the association relationships and association strengths between each subject to obtain an enhanced subject association network. Among them, the double-layer graph attention network is obtained by training through a number of historical research reports including each subject; The double-layer graph attention network is obtained by training through a number of historical research reports including each subject, including: constructing a training network according to each historical research report; setting negative link edges and positive link edges in the training network; constructing an initial double-layer graph attention network, and using the binary cross-entropy between the positive link edges and negative link edges generated by the initial double-layer graph attention network as the loss function; inputting the training network into the initial double-layer graph attention network, so that the initial double-layer graph attention network updates its own parameters through backpropagation according to the loss function to obtain the double-layer graph attention network; The risk transfer module is used to perform risk transfer on each of the initial credit risk scores starting from their respective corresponding risk subjects according to the enhanced subject association network and the preset risk transfer rules to obtain the risk quantitative assessment results of each subject.

6. The risk quantification and assessment system based on the subject association network according to claim 5, characterized in that, The risk quantification assessment system further includes a model training module, which is used to obtain the risk scoring model by training based on the BERT-base model through a number of historical news data, including a model construction unit, an acquisition unit, a data preprocessing unit, and a model training unit; Among them, the model construction unit is used to add a linear layer or a softmax layer on top of the BERT-base model, and set the loss function as the cross-entropy loss function to construct an initial risk scoring model; The acquisition unit is used to acquire a number of historical news data and perform annotation; The data preprocessing unit is used to perform word segmentation, stop word removal, and modal particle preprocessing operations on each piece of the historical news data to obtain corresponding training data; The model training unit is used to input each piece of the training data into the initial risk scoring model, so that the initial risk scoring model fine-tunes the model parameters through backpropagation according to the calculation result of the loss function to obtain the risk scoring model.

7. The risk quantification and assessment system based on the subject association network according to claim 5, wherein The network enhancement module includes an encoding unit, a decoding unit, a probability adjacency matrix construction unit, and a network enhancement unit; Among them, the encoding unit is used to encode the subject association network through an encoder to obtain a feature vector of a preset dimension, where the encoder is composed of a double-layer graph attention layer; The decoding unit is used to calculate the feature dot product between each subject according to the feature vector through a decoder; The probability adjacency matrix construction unit is used to construct a probability adjacency matrix for representing the association probability between each subject according to each feature dot product; The network enhancement unit is used to generate a number of link edges between each subject according to the probability adjacency matrix and determine the association strength of each link edge to obtain the enhanced subject association network.

8. A risk quantification and assessment system based on a subject association network according to any one of claims 5-7, characterized in that The risk transfer module transfers the initial credit risk scores from their respective corresponding risk subjects according to the enhanced subject association network and a preset risk transfer rule to obtain the risk quantification assessment results of each subject, including: Determine a number of secondary risk subjects corresponding to each risk subject according to the enhanced subject association network and a preset risk transfer layer number, where a certain risk subject can be used as a secondary risk subject of other risk subjects at the same time; Calculate the credit risk scores of each secondary risk subject by means of weighted summation according to the association strength between each risk subject and its corresponding number of secondary risk subjects, and then obtain the risk quantification assessment results of each subject.

Citation Information

Patent Citations

  • Public opinion early warning and risk propagation analysis method, system and device and storage medium

    CN111241300A

  • Network alignment method based on double-layer graph attention neural network

    CN111931903A

  • Business risk early warning method, equipment, storage medium and device

    CN116563006A