A method and device for analyzing enterprise risk transmission based on knowledge graph
Through the enterprise risk transmission analysis method based on the knowledge graph, combined with public opinion crawler, semantic analysis and CatBoost model, the problems of low efficiency and low accuracy of enterprise risk transmission analysis in the existing technology are solved, and more efficient and accurate risk warning and analysis are achieved.
Patent Information
- Application Number
- CN202111446656.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-11-30
- Publication Date
- 2025-06-06
- Estimated Expiration
- 2041-11-30
AI Technical Summary
The existing technology is inefficient and accurate when analyzing enterprise risk transmission, lacks timeliness and accuracy, making it difficult to effectively locate risk categories and predict the impact of risk event transmission on the target of attention.
The enterprise risk transmission analysis method based on knowledge graph is adopted, and the positioning and transmission path of enterprise risks is predicted through public opinion crawlers, semantic analysis, knowledge graph construction and CatBoost model prediction.
It improves the efficiency and accuracy of enterprise risk transmission analysis, can identify risk categories more quickly and predict the impact of risk events on the target company, providing a more accurate and timely risk warning.
Smart Images

Figure CN114328949B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of enterprise risk data analysis, and in particular to a method and device for enterprise risk conduction analysis based on a knowledge graph. Background Art
[0002] With the rapid development and constant changes of the global economy, there are a lot of uncertainties, and corporate risks are becoming more and more complex and diverse. At the same time, corporate risks will restrict the production and development of the company itself. The emergence of corporate risks is accompanied by various production and operation activities of the company, and it has strong transmission. How to calculate and locate the birth and transmission of corporate risks is particularly important.
[0003] Deng Mingran, Xia Zhe and others studied the dynamic transmission law of risks within enterprises by analyzing the mutual coupling of risks in the process of risk transmission and calculating the amount of risk transfer and the risk transmission coefficient in the process of risk transmission; Zhou Jiamu analyzed the influence range of risk sources through unsupervised learning, designed a hierarchical weight decay algorithm to calculate the risk factors of neighbor nodes in multiple layers, and obtained the risk transmission path.
[0004] Traditional risk transmission analysis requires a high amount of manpower to analyze the various relationships and upstream and downstream transmission directions of the target of attention. When the relationship hierarchy is complex and there are many related companies, the efficiency of manual analysis is very limited and the accuracy is not high. It is mainly based on the subjective judgment and qualitative analysis of analysts, and lacks timeliness and accuracy.
[0005] With respect to the technical problems existing in the above-mentioned prior art, no effective solutions have been proposed yet. Summary of the invention
[0006] The embodiments of the present disclosure provide a method and device for analyzing enterprise risk transmission based on knowledge graphs to at least solve the technical problems existing in the prior art. The problem to be solved by this application is how to better locate the risk category and predict the impact of risk event transmission on the target of interest when various risk events occur in the associated enterprises of the target of interest.
[0007] According to one aspect of an embodiment of the present disclosure, a method for analyzing enterprise risk transmission based on a knowledge graph is provided, including:
[0008] The public opinion crawler step uses a crawler tool to obtain news structured data from the Internet, performs word segmentation on the text content of the data, and removes high-frequency commonly used words;
[0009] The step of creating a pool of targets to be watched is to establish a pool of targets to be watched, store the information of targets to be watched in a relational database, and support the subscription and modification requirements of different business departments;
[0010] The public opinion semantic analysis step is to identify emotional and risk event labels based on the list of companies in the target pool and the structured news data;
[0011] The enterprise knowledge graph step is to construct an enterprise knowledge graph based on the enterprise list and news structured data, wherein the enterprise knowledge graph is composed of nodes and relationships; based on the enterprise list of the target pool and the subscription relationship type of the business department, a subject penetration calculation of a specific relationship and a specific number of layers is performed to obtain a subject penetration relationship link;
[0012] The risk transmission calculation step uses the CatBoost model to predict the impact of a risk event occurring at a node on the subject penetration relationship link on the target enterprise's stock price change after the risk event occurs;
[0013] The risk warning push step is to push relevant public opinion information, risk category information and risk transmission path based on the business department's subscription needs for the target company and the choice of risk category.
[0014] Furthermore, the use of a crawler tool to obtain news structured data from the Internet includes:
[0015] By using Python's web crawling framework Scrapy, structured data is obtained from mainstream financial websites and financial public accounts on a daily basis, including the website, title, content, author and news release time of news information.
[0016] Furthermore, the identification of emotional and risk event labels based on the list of enterprises in the target pool and the news structured data includes:
[0017] The weight of the news title is increased and combined with its main text as the input item for classification reasoning. The positive and negative sentiment bias model and risk event label model pre-trained by the RoBERTa model are used to infer the positive and negative sentiment direction and risk event label classification of each input information to obtain the analytical reasoning results.
[0018] Furthermore, the nodes are classified into enterprise nodes, natural person nodes, and product nodes, and the node attributes include enterprise name, enterprise status, organization code, enterprise registered capital, registration time, and registration number.
[0019] Relationships are classified into investment relationships, employment relationships, customer relationships, supplier relationships, guarantee relationships, etc. Relationship attributes include subscribed capital ratio, subscribed capital amount, job position information, customer income ratio of total revenue, purchase amount ratio of total purchase amount, guarantee period, and guarantee amount.
[0020] Furthermore, the processing process of the CatBoost model is as follows:
[0021] 1) Randomly sort the input sample set and generate multiple groups of random permutations;
[0022] 2) Convert the floating point type or attribute value tag to an integer;
[0023] 3) All categorical feature value results are converted into numerical results according to the following formula:
[0024]
[0025] Where countInClass indicates how many samples have a label value of 1 in the current category feature value; prior is the initial value of the numerator, which is determined according to the initial parameters; totalClass is the number of samples in all samples that have the same category feature value as the current sample.
[0026] Furthermore, the features used by the CatBoost model include node characteristics, risk event characteristics, link characteristics and other combined characteristics:
[0027] Node characteristics include the industry to which the start node and the target node belong, the nature of the enterprise, the registered capital of the enterprise, whether it has issued bonds, the subject rating, the total number of judicial risks in the last six months, the total number of operating risks in the last six months, the total number of negative public opinions in the last six months, the debt-to-asset ratio, and the ROE characteristics;
[0028] Risk event characteristics include risk event category, risk sentiment, and the number of enterprises associated with risk events;
[0029] Link characteristics include relationship penetration depth, number of relationship types, intermediate node attribute characteristics, and link neighborhood characteristics;
[0030] Other combined features include various extended feature vectors formed by high-order combination calculations of basic labels.
[0031] According to another aspect of an embodiment of the present disclosure, a storage medium is further provided, the storage medium including a stored program, wherein when the program is running, a processor executes any one of the methods described above.
[0032] According to another aspect of the embodiment of the present disclosure, there is also provided an enterprise risk conduction analysis device based on a knowledge graph, comprising:
[0033] The public opinion crawler module uses crawler tools to obtain news structured data from the Internet, performs word segmentation on the text content of the data, and removes high-frequency commonly used words;
[0034] The target pool module builds a target pool, stores the target information in a relational database, and supports the subscription and modification requirements of different business departments;
[0035] The public opinion semantic analysis module identifies emotional and risk event labels based on the list of companies in the target pool and the structured news data;
[0036] The enterprise knowledge graph module constructs an enterprise knowledge graph based on the enterprise list and news structured data, wherein the enterprise knowledge graph is composed of nodes and relationships; performs subject penetration calculation of a specific relationship and a specific number of layers based on the enterprise list of the target pool and the subscription relationship type of the business department, and obtains a subject penetration relationship link;
[0037] The risk transmission calculation module uses the CatBoost model to predict the impact of a risk event occurring at a node on the subject penetration relationship link on the target enterprise's stock price change after the risk event occurs;
[0038] The risk warning push module pushes relevant public opinion information, risk category information and risk transmission paths based on the business department's subscription needs for the target companies and the choice of risk categories.
[0039] According to another aspect of the embodiment of the present disclosure, there is also provided an enterprise risk conduction analysis device based on a knowledge graph, comprising:
[0040] a first processor; and
[0041] A first memory is connected to the first processor and is used to provide the first processor with instructions for processing the following processing steps:
[0042] The public opinion crawler step uses a crawler tool to obtain news structured data from the Internet, performs word segmentation on the text content of the data, and removes high-frequency commonly used words;
[0043] The step of creating a pool of targets to be watched is to establish a pool of targets to be watched, store the information of targets to be watched in a relational database, and support the subscription and modification requirements of different business departments;
[0044] The public opinion semantic analysis step is to identify emotional and risk event labels based on the list of companies in the target pool and the structured news data;
[0045] The enterprise knowledge graph step is to construct an enterprise knowledge graph based on the enterprise list and news structured data, wherein the enterprise knowledge graph is composed of nodes and relationships; based on the enterprise list of the target pool and the subscription relationship type of the business department, a subject penetration calculation of a specific relationship and a specific number of layers is performed to obtain a subject penetration relationship link;
[0046] The risk transmission calculation step uses the CatBoost model to predict the impact of a risk event occurring at a node on the subject penetration relationship link on the target enterprise's stock price change within ten days after the risk event occurs;
[0047] The risk warning push step is to push relevant public opinion information, risk category information and risk transmission path based on the business department's subscription needs for the target company and the choice of risk category.
[0048] The technical solution of this application has the following beneficial effects:
[0049] This application adopts a corporate risk transmission analysis method based on knowledge graph, uses the crawler framework Scrapy to crawl public opinion information, classifies risk event information by risk and sentiment through the pre-trained public opinion model RoBERTa, and performs link penetration calculation on the specific relationships of risky companies based on the knowledge graph database Neo4j with added reverse link indexes, and then uses the CatBoost model to predict the impact of risk events on the link on the target company's stock price fluctuations, thereby realizing corporate risk analysis and early warning functions. BRIEF DESCRIPTION OF THE DRAWINGS
[0050] The drawings described herein are used to provide a further understanding of the present disclosure and constitute a part of the present application. The illustrative embodiments of the present disclosure and their descriptions are used to explain the present disclosure and do not constitute an improper limitation on the present disclosure. In the drawings:
[0051] Figure 1 is a hardware structure block diagram of a computing device for implementing the method according to Embodiment 1 of the present disclosure;
[0052] Figure 2 is a schematic diagram of a system for implementing enterprise risk conduction analysis based on a knowledge graph according to Embodiment 1 of the present disclosure;
[0053] Figure 3 It is a flow chart of a method for implementing enterprise risk conduction analysis based on knowledge graph according to the first aspect of Embodiment 1 of the present disclosure;
[0054] Figure 4 is a schematic diagram of a device for implementing enterprise risk conduction analysis based on a knowledge graph according to the first aspect of Embodiment 2 of the present disclosure;
[0055] Figure 5 It is a schematic diagram of a device for implementing enterprise risk conduction analysis based on a knowledge graph according to the second aspect of Example 2 of the present disclosure. DETAILED DESCRIPTION
[0056] In order to enable those skilled in the art to better understand the technical solutions of the present disclosure, the technical solutions in the embodiments of the present disclosure will be clearly and completely described below in conjunction with the drawings in the embodiments of the present disclosure. Obviously, the described embodiments are only embodiments of a part of the present disclosure, not all of the embodiments. Based on the embodiments in the present disclosure, all other embodiments obtained by ordinary technicians in this field without making creative work should fall within the scope of protection of the present disclosure.
[0057] It should be noted that the terms "first", "second", etc. in the specification and claims of the present disclosure and the above-mentioned drawings are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence. It should be understood that the data used in this way can be interchangeable where appropriate, so that the embodiments of the present disclosure described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusions, for example, a process, method, system, product, or device that includes a series of steps or units is not necessarily limited to those steps or units that are clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products, or devices.
[0058] This application adopts a risk conduction analysis method based on knowledge graph, labels risk events by risk and sentiment through the pre-trained public opinion model RoBERTa, calculates links for specific relationships of risky companies based on the knowledge graph database with added reverse link indexes, and uses the CatBoost model to predict the impact of risk events on the link on the target company's stock price fluctuations, thereby realizing risk analysis and early warning functions.
[0059] Example 1
[0060] According to this embodiment, an embodiment of an enterprise risk conduction analysis method based on a knowledge graph is provided. It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer executable instructions, and although a logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in an order different from that shown here.
[0061] The method embodiment provided in this embodiment can be executed in a mobile terminal, a computer terminal, a server or a similar computing device. Figure 1 FIG. 1 shows a hardware structure block diagram of a computing device for implementing an enterprise risk conduction analysis method based on a knowledge graph. Figure 1As shown, the computing device may include one or more processors (the processor may include but is not limited to a processing device such as a microprocessor MCU or a programmable logic device FPGA), a memory for storing data, and a transmission device for communication functions. In addition, it may also include: a display, an input / output interface (I / O interface), a universal serial bus (USB) port (which may be included as one of the ports of the I / O interface), a network interface, a power supply and / or a camera. It can be understood by those skilled in the art that Figure 1 The structure shown is only for illustration and does not limit the structure of the above electronic device. Figure 1 More or fewer components as shown, or with Figure 1 Different configurations are shown.
[0062] It should be noted that the one or more processors and / or other data processing circuits described above may generally be referred to herein as "data processing circuitry". The data processing circuitry may be embodied in whole or in part as software, hardware, firmware, or any other combination thereof. In addition, the data processing circuitry may be a single independent processing module, or may be incorporated in whole or in part into any of the other components of the computing device. As involved in the embodiments of the present disclosure, the data processing circuitry acts as a processor control (e.g., selection of a variable resistor terminal path connected to an interface).
[0063] The memory can be used to store software programs and modules of application software, such as the program instructions / data storage device corresponding to the enterprise risk conduction analysis method based on knowledge graph in the embodiment of the present disclosure. The processor executes various functional applications and data processing by running the software programs and modules stored in the memory, that is, the enterprise risk conduction analysis method based on knowledge graph of the above-mentioned application program is realized. The memory may include a high-speed random access memory, and may also include a non-volatile memory, such as one or more part-of-speech storage devices, flash memory, or other non-volatile solid-state memory. In some instances, the memory may further include a memory remotely arranged relative to the processor, and these remote memories may be connected to the computing device via a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.
[0064] The transmission device is used to receive or send data via a network. The specific example of the above network may include a wireless network provided by a communication provider of the computing device. In one example, the transmission device includes a network adapter (Network Interface Controller, NIC), which can be connected to other network devices through a base station so as to communicate with the Internet. In one example, the transmission device can be a radio frequency (Radio Frequency, RF) module, which is used to communicate with the Internet wirelessly.
[0065] The display may be, for example, a touch screen liquid crystal display (LCD) that may enable a user to interact with a user interface of the computing device.
[0066] It should be noted that, in some optional embodiments, the above Figure 1 The computing device shown may include hardware elements (including circuits), software elements (including computer code stored on a computer-readable medium), or a combination of both hardware elements and software elements. Figure 1 This is merely one example of a particular embodiment and is intended to illustrate the types of components that may be present in the computing devices described above.
[0067] Figure 2 is a schematic diagram of a system for implementing enterprise risk conduction analysis based on knowledge graph according to this embodiment. Figure 2 As shown, the system includes: a front-end portable electronic terminal 100 (such as a laptop), a computing device 200 for implementing an enterprise risk conduction analysis method based on a knowledge graph, and a cloud server 300. It should be noted that the computing device 200 for implementing an enterprise risk conduction analysis method based on a knowledge graph in the system can be applicable to the hardware structure described above.
[0068] In the above operating environment, according to the first aspect of this embodiment, a method for implementing enterprise risk conduction analysis based on knowledge graph is provided. The method is composed of Figure 2 A computing device 200 for implementing an enterprise risk conduction analysis method based on a knowledge graph is shown in FIG. Figure 3 A schematic diagram showing the process of the method is shown in FIG. Figure 3 As shown, the method includes:
[0069] S302: Public opinion crawler steps
[0070] By using Python's web crawling framework Scrapy, we obtain structured data from mainstream financial websites and public financial accounts on a daily basis, including the website, title, content, author, and news release time of news information. We also perform word segmentation on the text content of the news and remove high-frequency common words to optimize subsequent sample processing and training.
[0071] S304: Focus on the target pool step
[0072] In order to support the project risk management work of securities companies' risk control, investment banking and other business departments, continue to pay attention to various public information and external clues of project-related targets, and take timely prevention and control of sudden risk events, a special target pool is established, and the information of the target is stored in a relational database, which can support various subscription and modification needs of different business departments.
[0073] S306: Public opinion semantic analysis steps
[0074] Based on the list of companies in the target pool, financial news information is crawled daily and the public opinion semantic analysis step is entered to identify emotional and risk event labels.
[0075] The public opinion semantic analysis step increases the weight of the news title and combines it with the text as the input item for classification reasoning. The positive and negative sentiment bias model and risk event label model pre-trained by the RoBERTa model are used to infer the positive and negative sentiment direction and risk event label classification of each input information to obtain the analysis and reasoning results.
[0076] The specific positive and negative sentiment labels output by the public opinion model are divided into -3, -2, -1, 0, 1, 2, and 3. Negative numbers represent negative news, positive numbers represent positive news, and the size of the number indicates the degree of impact. Risk event labels are divided into credit risk, operating risk, financial risk, securities market risk, governance and management risk, and force majeure risk.
[0077] S308: Enterprise Knowledge Graph Steps
[0078] The data of the enterprise knowledge graph is stored in the Neo4j graph database, which supports Cypher statement query and API interface batch query. The stored information includes node information and relationship information, and both nodes and relationships can have multiple attribute information.
[0079] Nodes are classified into enterprise nodes, natural person nodes, product nodes, etc. Node attributes include enterprise name, enterprise status, organization code, enterprise registered capital (10,000 yuan), registration time, registration number, etc.
[0080] Relationships are classified into investment relationships, employment relationships, customer relationships, supplier relationships, guarantee relationships, etc. Relationship attributes include subscribed capital ratio, subscribed capital amount, job position information, customer income ratio of total revenue, purchase amount ratio of total purchase amount, guarantee period, guarantee amount (10,000 yuan), etc.
[0081] In the graph database, through the pre-calculation method, taking equity and capital attributes as an example, a "link" is added in the reverse direction on the graph to represent the equity relationship between the parent enterprise entity and the grandchild enterprise entity after multi-level investment, forming a "ring structure" subgraph on the graph. The reverse links of this structure are indexed by node degree and weight, which can greatly save query and analysis time.
[0082] Based on the list of enterprises in the target pool of attention and the subscription relationship types of the business departments, the subject penetration calculation of specific relationships and specific layers is performed, and the subject penetration results are sent to the risk transmission calculation step.
[0083] S310: Risk transmission calculation steps
[0084] The risk transmission calculation step uses the CatBoost model to predict the impact of a risk event occurring at a node in the subject penetration relationship link on the target company's stock price changes within ten days after the risk event occurs.
[0085] The CatBoost model predicts the stock price fluctuations of the target company on the link ten days later by analyzing various risk events that occurred at the starting nodes of the historical links. The CatBoost model is a machine learning library that Yandex opened in 2017 and is a type of Boosting family algorithm. CatBoost, XGBoost, and LightGBM are all improved implementations under the GBDT algorithm framework. XGBoost is widely used in the industry, LightGBM effectively improves the computational efficiency of GBDT, and CatBoost is an algorithm that performs better than XGBoost and LightGBM in terms of algorithm accuracy.
[0086] CatBoost is a GBDT framework based on symmetric decision trees (Oblivious Trees) as the base learner with fewer parameters, support for categorical variables and high accuracy. The main pain point it solves is to efficiently and reasonably handle categorical features. In addition, CatBoost also solves the problems of gradient bias and prediction shift, thereby reducing the occurrence of overfitting and improving the accuracy and generalization ability of the algorithm.
[0087] CatBoost is very flexible in handling categorical features. You can directly pass in the column identifier of the categorical feature, and the model will automatically use one-hot encoding for it. You can also set the one_hot_max_size parameter to limit the length of the one-hot feature vector. If you do not pass in the column identifier of the categorical feature, CatBoost will treat all columns as numerical features. For features whose one-hot encoding exceeds the set one_hot_max_size value, CatBoost will use an efficient encoding method. The processing process is as follows:
[0088] 1) Randomly sort the input sample set and generate multiple groups of random permutations;
[0089] 2) Convert the floating point type or attribute value tag to an integer;
[0090] 3) All categorical feature value results are converted into numerical results according to the following formula.
[0091]
[0092] Where countInClass indicates how many samples have a label value of 1 in the current class feature value; prior is the initial value of the numerator, which is determined according to the initial parameters. TotalClass is the number of samples in all samples (including the current sample) that have the same class feature value as the current sample.
[0093] The features used in the model mainly include node characteristics, risk event characteristics, link characteristics and other combined characteristics:
[0094] 1) Node characteristics include the industry to which the starting node and the target node belong, the nature of the enterprise, the registered capital of the enterprise, whether it has issued bonds, the subject rating, the total number of judicial risks in the last 6 months, the total number of operating risks in the last 6 months, the total number of negative public opinions in the last 6 months, the debt-to-asset ratio, and the return on equity (ROE).
[0095] 2) Risk event characteristics include risk event category, risk sentiment, number of risk event-related enterprises, etc.;
[0096] 3) Link characteristics include relationship penetration depth, number of relationship types, intermediate node attribute characteristics, link neighborhood characteristics, etc.;
[0097] 4) Other combined features include various extended feature vectors formed by high-order combination calculations of basic labels.
[0098] Model parameter tuning uses GridSearchCV to automatically search for optimal parameters. The relevant tuning parameters are listed as follows:
[0099]
[0100]
[0101] Model evaluation indicators using R 2 , the calculation formula is as follows. is the model’s prediction of stock price, y (i) is the actual stock price, y (i) is the average stock price. 2 The closer it is to 1, the better the model fit is.
[0102]
[0103] S312: Risk warning push steps
[0104] According to the subscription needs of risk control, investment banking and other business departments for the target companies and the choice of risk categories, relevant public opinion information, risk category information and risk transmission paths are pushed.
[0105] Therefore, according to the first aspect of this embodiment, the following beneficial effects are achieved:
[0106] This application adopts a corporate risk transmission analysis method based on knowledge graph, uses the crawler framework Scrapy to crawl public opinion information, classifies risk event information by risk and sentiment through the pre-trained public opinion model RoBERTa, and performs link penetration calculation on the specific relationships of risky companies based on the knowledge graph database Neo4j with added reverse link indexes, and then uses the CatBoost model to predict the impact of risk events on the link on the target company's stock price fluctuations, thereby realizing corporate risk analysis and early warning functions.
[0107] In addition, reference Figure 1 As shown, according to the second aspect of this embodiment, a storage medium is provided, wherein the storage medium includes a stored program, wherein when the program is run, a processor executes any one of the above methods.
[0108] It should be noted that, for the aforementioned method embodiments, for the sake of simplicity, they are all expressed as a series of action combinations, but those skilled in the art should be aware that the present application is not limited by the described order of actions, because according to the present application, certain steps can be performed in other orders or simultaneously. Secondly, those skilled in the art should also be aware that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily required by the present application.
[0109] Through the description of the above implementation methods, those skilled in the art can clearly understand that the method according to the above embodiment can be implemented by means of software plus a necessary general hardware platform, and of course by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of the present application, or the part that contributes to the prior art, can be embodied in the form of a software product, which is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk), and includes a number of instructions for a terminal device (which can be a mobile phone, computer, server, or network device, etc.) to execute the methods described in each embodiment of the present application.
[0110] Example 2
[0111] Figure 4 The enterprise risk conduction analysis device based on knowledge graph according to the first aspect of this embodiment is shown, and the device corresponds to the method according to the first aspect of embodiment 1. Figure 4 As shown, the device comprises:
[0112] S01 Public Opinion Crawler Module
[0113] By using Python's web crawling framework Scrapy, we obtain structured data from mainstream financial websites and public financial accounts on a daily basis, including the website, title, content, author, and news release time of news information. We also perform word segmentation on the text content of the news and remove high-frequency common words to optimize subsequent sample processing and training.
[0114] S02 Focus on the target pool module
[0115] In order to support the project risk management work of securities companies' risk control, investment banking and other business departments, continue to pay attention to various public information and external clues of project-related targets, and take timely prevention and control of sudden risk events, a special target pool is established, and the information of the target is stored in a relational database, which can support various subscription and modification needs of different business departments.
[0116] S03 Public Opinion Semantic Analysis Module
[0117] Based on the list of companies in the target pool, the financial news information is crawled daily and enters the public opinion semantic analysis module to identify emotional and risk event labels.
[0118] The public opinion semantic analysis module increases the weight of the news title and combines it with the text as the input for classification reasoning. Using the positive and negative sentiment bias model and risk event label model pre-trained by the RoBERTa model, the positive and negative sentiment direction and risk event label classification of each input news are inferred to obtain the analysis and reasoning results.
[0119] The specific positive and negative sentiment labels output by the public opinion model are divided into -3, -2, -1, 0, 1, 2, and 3. Negative numbers represent negative news, positive numbers represent positive news, and the size of the number indicates the degree of impact. Risk event labels are divided into credit risk, operating risk, financial risk, securities market risk, governance and management risk, and force majeure risk.
[0120] S04 Enterprise Knowledge Graph Module
[0121] The data of the enterprise knowledge graph is stored in the Neo4j graph database, which supports Cypher statement query and API interface batch query. The stored information includes node information and relationship information, and both nodes and relationships can have multiple attribute information.
[0122] Nodes are classified into enterprise nodes, natural person nodes, product nodes, etc. Node attributes include enterprise name, enterprise status, organization code, enterprise registered capital (10,000 yuan), registration time, registration number, etc.
[0123] Relationships are classified into investment relationships, employment relationships, customer relationships, supplier relationships, guarantee relationships, etc. Relationship attributes include subscribed capital ratio, subscribed capital amount, job position information, customer income ratio of total revenue, purchase amount ratio of total purchase amount, guarantee period, guarantee amount (10,000 yuan), etc.
[0124] In the graph database, through the pre-calculation method, taking equity and capital attributes as an example, a "link" is added in the reverse direction on the graph to represent the equity relationship between the parent enterprise entity and the grandchild enterprise entity after multi-level investment, forming a "ring structure" subgraph on the graph. The reverse links of this structure are indexed by node degree and weight, which can greatly save query and analysis time.
[0125] Based on the list of enterprises in the target pool of attention and the subscription relationship types of the business departments, the subject penetration calculation of specific relationships and specific layers is performed, and the subject penetration results are sent to the risk transmission calculation module.
[0126] Contents of S05 Risk Transmission Calculation Module
[0127] The risk transmission calculation module uses the CatBoost model to predict the impact of a risk event occurring at a node in the subject penetration relationship link on the target company's stock price changes within ten days after the risk event occurs.
[0128] The CatBoost model predicts the stock price fluctuations of the target company on the link ten days later by taking into account various risk events that occurred at the starting nodes of the historical link.
[0129] CatBoost is very flexible in handling categorical features. You can directly pass in the column identifier of the categorical feature, and the model will automatically use one-hot encoding for it. You can also set the one_hot_max_size parameter to limit the length of the one-hot feature vector. If you do not pass in the column identifier of the categorical feature, CatBoost will treat all columns as numerical features. For features whose one-hot encoding exceeds the set one_hot_max_size value, CatBoost will use an efficient encoding method. The processing process is as follows:
[0130] 1) Randomly sort the input sample set and generate multiple groups of random permutations;
[0131] 2) Convert the floating point type or attribute value tag to an integer;
[0132] 3) All categorical feature value results are converted into numerical results according to the following formula.
[0133]
[0134] Where countInClass indicates how many samples have a label value of 1 in the current class feature value; prior is the initial value of the numerator, which is determined according to the initial parameters. TotalClass is the number of samples in all samples (including the current sample) that have the same class feature value as the current sample.
[0135] The features used in the model mainly include node characteristics, risk event characteristics, link characteristics and other combined characteristics:
[0136] 1) Node characteristics include the industry to which the starting node and the target node belong, the nature of the enterprise, the registered capital of the enterprise, whether it has issued bonds, the subject rating, the total number of judicial risks in the last 6 months, the total number of operating risks in the last 6 months, the total number of negative public opinions in the last 6 months, the debt-to-asset ratio, and the return on equity (ROE).
[0137] 2) Risk event characteristics include risk event category, risk sentiment, number of risk event-related enterprises, etc.;
[0138] 3) Link characteristics include relationship penetration depth, number of relationship types, intermediate node attribute characteristics, link neighborhood characteristics, etc.;
[0139] 4) Other combined features include various extended feature vectors formed by high-order combination calculations of basic labels.
[0140] Model parameter tuning uses GridSearchCV to automatically search for optimal parameters. The relevant tuning parameters are listed as follows:
[0141]
[0142] Model evaluation indicators using R 2 , the calculation formula is as follows. is the model’s prediction of stock price, y (i) is the actual stock price, y (i) is the average stock price. 2 The closer it is to 1, the better the model fit is.
[0143]
[0144] S06 Risk warning push module
[0145] According to the subscription needs of risk control, investment banking and other business departments for the target companies and the choice of risk categories, relevant public opinion information, risk category information and risk transmission paths are pushed.
[0146] also, Figure 5 The enterprise risk conduction analysis device 600 based on knowledge graph according to the second aspect of this embodiment is shown, and the device 600 corresponds to the method according to the second aspect of embodiment 1. Figure 5 As shown, the device 600 includes: an enterprise risk conduction analysis device based on knowledge graph, characterized in that it includes:
[0147] A first processor 610; and
[0148] The first memory 620 is connected to the first processor and is used to provide the first processor with instructions for processing the following processing steps:
[0149] The public opinion crawler step uses a crawler tool to obtain news structured data from the Internet, performs word segmentation on the text content of the data, and removes high-frequency commonly used words;
[0150] The step of creating a pool of targets to be watched is to establish a pool of targets to be watched, store the information of targets to be watched in a relational database, and support the subscription and modification requirements of different business departments;
[0151] The public opinion semantic analysis step is to identify emotional and risk event labels based on the list of companies in the target pool and the structured news data;
[0152] The enterprise knowledge graph step is to construct an enterprise knowledge graph based on the enterprise list and news structured data, wherein the enterprise knowledge graph is composed of nodes and relationships; based on the enterprise list of the target pool and the subscription relationship type of the business department, a subject penetration calculation of a specific relationship and a specific number of layers is performed to obtain a subject penetration relationship link;
[0153] The risk transmission calculation step uses the CatBoost model to predict the impact of a risk event occurring at a node on the subject penetration relationship link on the target enterprise's stock price change within ten days after the risk event occurs;
[0154] The risk warning push step is to push relevant public opinion information, risk category information and risk transmission path based on the business department's subscription needs for the target company and the choice of risk category.
[0155] Therefore, according to this embodiment, this application adopts a corporate risk conduction analysis method based on knowledge graph, uses the crawler framework Scrapy to crawl public opinion information, classifies risk event information by risk and sentiment through the pre-trained public opinion model RoBERTa, and performs link penetration calculations on the specific relationships of risky companies based on the knowledge graph database Neo4j with an added reverse link index, and then uses the CatBoost model to predict the impact of risk events on the link on the target company's stock price fluctuations, thereby realizing enterprise risk analysis and early warning functions.
[0156] The serial numbers of the above-mentioned embodiments of the present application are for description only and do not represent the advantages or disadvantages of the embodiments.
[0157] In the above embodiments of the present application, the description of each embodiment has its own emphasis. For parts that are not described in detail in a certain embodiment, please refer to the relevant description of other embodiments.
[0158] In the several embodiments provided in this application, it should be understood that the disclosed technical content can be implemented in other ways. Among them, the device embodiments described above are only schematic. For example, the division of the units is only a logical function division. There may be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of units or modules, which can be electrical or other forms.
[0159] The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed on multiple network units. Some or all of the units may be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0160] In addition, each functional unit in each embodiment of the present application may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit. The above-mentioned integrated unit may be implemented in the form of hardware or in the form of software functional units.
[0161] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application is essentially or the part that contributes to the prior art or all or part of the technical solution can be embodied in the form of a software product, and the computer software product is stored in a storage medium, including a number of instructions to enable a computer device (which can be a personal computer, server or network device, etc.) to perform all or part of the steps of the method described in each embodiment of the present application. The aforementioned storage medium includes: U disk, read-only memory (ROM, Read-Only Memory), random access memory (RAM, RandomAccess Memory), mobile hard disk, disk or CD-ROM and other media that can store program codes.
[0162] The above is only a preferred implementation of the present application. It should be pointed out that for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the principles of the present application. These improvements and modifications should also be regarded as the scope of protection of the present application.
Claims
1. A method for analyzing enterprise risk transmission based on knowledge graph, It is characterized in that include: The public opinion crawler step uses a crawler tool to obtain news structured data from the Internet, performs word segmentation on the text content of the data, and removes high-frequency commonly used words; The step of creating a pool of targets to be watched is to establish a pool of targets to be watched, store the information of targets to be watched in a relational database, and support the subscription and modification requirements of different business departments; The public opinion semantic analysis step is to identify emotional and risk event labels based on the list of companies in the target pool and the structured news data; The enterprise knowledge graph step is to construct an enterprise knowledge graph based on the enterprise list and news structured data, wherein the enterprise knowledge graph is composed of nodes and relationships; and subject penetration calculation is performed based on the enterprise list of the target pool and the subscription relationship type of the business department to obtain the subject penetration relationship link; The risk transmission calculation step uses the CatBoost model to predict the impact of a risk event occurring at a node on the subject penetration relationship link on the target enterprise's stock price change after the risk event occurs; The risk warning push step pushes relevant public opinion information, risk category information and risk transmission path according to the subscription needs of the business department for the target company and the choice of risk category; the processing process of the CatBoost model is as follows: 1) Randomly sort the input sample set and generate multiple sets of random permutations; 2) Convert floating point or attribute value tags to integers; 3) Convert all categorical feature value results into numerical results according to the following formula: ; in Indicates how many samples have a label value of 1 in the current category feature value; is the initial value of the numerator, determined according to the initial parameters; It is the number of samples in all samples that have the same categorical feature value as the current sample.
2. The method according to claim 1, It is characterized in that The crawler tool is used to obtain news structured data from the Internet, including: By using Python's web crawling framework Scrapy, structured data is obtained from mainstream financial websites and financial public accounts on a daily basis, including the website, title, content, author and news release time of news information.
3. The method according to claim 1, It is characterized in that The identification of sentiment and risk event labels based on the list of companies in the target pool and the news structured data includes: The weight of the news title is increased and combined with its main text as the input item for classification reasoning. The positive and negative sentiment bias model and risk event label model pre-trained by the RoBERTa model are used to infer the positive and negative sentiment direction and risk event label classification of each input information to obtain the analytical reasoning results.
4. The method according to claim 1, It is characterized in that The nodes are classified into enterprise nodes, natural person nodes, and product nodes, and the node attributes include enterprise name, enterprise status, organization code, enterprise registered capital, registration time, and registration number; Relationships are classified into investment relationships, employment relationships, customer relationships, supplier relationships, guarantee relationships, etc. Relationship attributes include subscribed capital ratio, subscribed capital amount, job position information, customer income ratio of total revenue, purchase amount ratio of total purchase amount, guarantee period, and guarantee amount.
5. The method according to claim 1, It is characterized in that The features used by the CatBoost model include node characteristics, risk event characteristics, link characteristics and other combined characteristics: Node characteristics include the industry to which the start node and the target node belong, the nature of the enterprise, the registered capital of the enterprise, whether it has issued bonds, the subject rating, the total number of judicial risks in the last six months, the total number of operating risks in the last six months, the total number of negative public opinions in the last six months, the debt-to-asset ratio, and the ROE characteristics; Risk event characteristics include risk event category, risk sentiment, and the number of enterprises associated with risk events; Link characteristics include relationship penetration depth, number of relationship types, intermediate node attribute characteristics, and link neighborhood characteristics; Other combined features include various extended feature vectors formed by high-order combination calculations of basic labels.
6. A storage medium, It is characterized in that The storage medium includes a stored program, wherein when the program is run, the processor executes the method according to any one of claims 1 to 5.
7. An enterprise risk transmission analysis device based on knowledge graph, It is characterized in that include: The public opinion crawler module uses crawler tools to obtain news structured data from the Internet, performs word segmentation on the text content of the data, and removes high-frequency commonly used words; The target pool module builds a target pool, stores the target information in a relational database, and supports the subscription and modification requirements of different business departments; The public opinion semantic analysis module identifies emotional and risk event labels based on the list of companies in the target pool and the structured news data; The enterprise knowledge graph module constructs an enterprise knowledge graph based on the enterprise list and news structured data, wherein the enterprise knowledge graph is composed of nodes and relationships; performs subject penetration calculation based on the enterprise list of the target pool and the subscription relationship type of the business department to obtain the subject penetration relationship link; The risk transmission calculation module uses the CatBoost model to predict the impact of a risk event occurring at a node on the subject penetration relationship link on the target enterprise's stock price change after the risk event occurs; The risk warning push module pushes relevant public opinion information, risk category information and risk transmission path according to the subscription needs of the business department for the target company and the selection of risk category; the processing process of the CatBoost model is as follows: 1) Randomly sort the input sample set and generate multiple sets of random permutations; 2) Convert floating point or attribute value tags to integers; 3) Convert all categorical feature value results into numerical results according to the following formula: ; in Indicates how many samples have a label value of 1 in the current category feature value; is the initial value of the numerator, determined according to the initial parameters; It is the number of samples in all samples that have the same categorical feature value as the current sample.
8. The device according to claim 7, It is characterized in that The crawler tool is used to obtain news structured data from the Internet, including: By using Python's web crawling framework Scrapy, structured data is obtained from mainstream financial websites and financial public accounts on a daily basis, including the website, title, content, author and news release time of news information.
9. An enterprise risk transmission analysis device based on knowledge graph, It is characterized in that include: a first processor; as well as A first memory is connected to the first processor and is used to provide the first processor with instructions for processing the following processing steps: The public opinion crawler step uses a crawler tool to obtain news structured data from the Internet, performs word segmentation on the text content of the data, and removes high-frequency commonly used words; The step of creating a pool of targets to be watched is to establish a pool of targets to be watched, store the information of targets to be watched in a relational database, and support the subscription and modification requirements of different business departments; The public opinion semantic analysis step is to identify emotional and risk event labels based on the list of companies in the target pool and the structured news data; The enterprise knowledge graph step is to construct an enterprise knowledge graph based on the enterprise list and news structured data, wherein the enterprise knowledge graph is composed of nodes and relationships; and subject penetration calculation is performed based on the enterprise list of the target pool and the subscription relationship type of the business department to obtain the subject penetration relationship link; The risk transmission calculation step uses the CatBoost model to predict the impact of a risk event occurring at a node on the subject penetration relationship link on the target enterprise's stock price change within ten days after the risk event occurs; The risk warning push step pushes relevant public opinion information, risk category information and risk transmission path according to the subscription needs of the business department for the target company and the choice of risk category; the processing process of the CatBoost model is as follows: 1) Randomly sort the input sample set and generate multiple sets of random permutations; 2) Convert floating point or attribute value tags to integers; 3) Convert all categorical feature value results into numerical results according to the following formula: ; in Indicates how many samples have a label value of 1 in the current category feature value; is the initial value of the numerator, determined according to the initial parameters; It is the number of samples in all samples that have the same categorical feature value as the current sample.
Citation Information
Patent Citations
Knowledge graph-based enterprise risk prediction method and system
CN108596439A
Stock risk prediction method and system based on user portrait and knowledge graph
CN112734569A