Risk estimation method based on knowledge graph
Through the risk estimate method based on the knowledge graph, the knowledge graph and risk estimate model are dynamically updated, and the problem of inaccurate corporate credit risk warning is solved, accurate and reliable risk warning is achieved, and a new risk management perspective is provided.
Patent Information
- Application Number
- CN202411983057.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-31
- Publication Date
- 2025-06-03
AI Technical Summary
Under the cross-platform and cross-field credit service model, the sources of credit information of enterprise management are uneven, corporate credit risks are everywhere, and the graph structure based on historical data cannot be adjusted and updated in a timely manner, resulting in inaccurate risk warnings.
The knowledge graph-based risk estimate method is adopted, and the knowledge graph and risk estimate model are dynamically updated through steps such as data collection, entity recognition, graph construction, feature extraction and risk model training to achieve accurate warnings on corporate credit risks.
The transformation from static data analysis to dynamic behavioral model mining has been achieved, the accuracy and reliability of risk estimates have been improved, concept drifting is avoided, and new perspectives and technical support for risk management in the financial industry.
Smart Images

Figure CN120087746A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of risk identification, and particularly to a risk prediction method based on a knowledge graph. Background Art
[0002] As a structured semantic network, a knowledge graph can connect entities and their relationships from different fields to form a comprehensive enterprise intelligence platform. A multimodal knowledge graph refers to a knowledge graph that can fully integrate and utilize data from multiple modal sources such as language, vision, and audition. It should be noted that any original data containing knowledge can be used as the data source for constructing a knowledge graph.
[0003] In related technologies, when warning of enterprise reputation risks, it is necessary to quickly discover the risks that occur in an enterprise through Internet data. Dynamic risk identification generally uses collected time case information to extract and process the event content data of the information, construct a basic knowledge base, label according to the content, and the labels include industry classification, time classification, risk control attributes, etc. The labels are associated with the corresponding enterprises, and other information related to this time is retrieved according to the pre-constructed knowledge graph, so as to connect the event with other content, and obtain the industry information concerned by the enterprise to form a relationship graph. A score is given for the dimensions of established modules such as the degree of public opinion influence, sensitive elements, and the development stage of public opinion of the event. According to the score and the label, a content risk warning rule is constructed. When a new event case is added or updated in the system, the content risk warning rule is used to perform risk warning on the customer system. Generally, the enterprise risk conduction analysis based on a knowledge graph is generally an opinion crawler, focusing on the target, opinion semantic analysis, then the enterprise knowledge graph, risk conduction calculation, and then risk warning push; However, in the current cross-platform and cross-domain credit large service model, the sources of enterprise management credit information are uneven, enterprise credit risks are everywhere, and enterprise credit risks do not exist in isolation, but there is a transmission effect. The correlation between enterprise credit subjects reflects the intricate connections between credit subjects in social production and life. The transmission intensity of enterprise credit risks is mainly related to the correlation density of the correlation structure of credit subjects and the correlation intensity between enterprise credit subjects. And for the graph structure trained based on historical data, with the arrival of new information, the original entity recognition and relationship extraction rules will become inapplicable, and it is impossible to adjust and update the rules in a timely manner according to the latest industry standards and technological progress.
[0004] Therefore, it is necessary to provide a new risk prediction method based on a knowledge graph to solve the above technical problems. Summary of the Invention
[0005] To solve the above technical problems, the present invention provides a risk prediction method based on a knowledge graph. The risk prediction method based on a knowledge graph provided by the present invention includes the following steps: S1. Data collection, collecting data from multiple dimensions including supply chain data, market competition data, enterprise credit history data, and enterprise financial health; S2. Entity recognition, using natural language processing technology to identify key entities in the relevant data and extract relationships between the key entities. Among them, the features useful for identifying key entities are as follows: for supply chain data, extracting features such as supplier name, product name, and purchase quantity; for market competition data, extracting features such as competitor name and market share; for enterprise credit history data, extracting features such as loan amount, repayment date, and credit rating; for enterprise financial health data, extracting features such as financial indicators and cash flow; S3. Graph construction, constructing a knowledge graph in a graphical structure based on the identified key entities and the relationships between them; S4. Feature extraction, extracting dynamic interaction key feature information from the time series data included in the existing knowledge graph based on time-frequency analysis, and adding the new features as additional information to the corresponding entities or relationships to update the knowledge graph; S5. Risk model training, using the updated knowledge graph data to evaluate the importance of the dynamic interaction key feature information to obtain a dynamic interaction importance score, constructing a risk prediction model based on the neural network learning algorithm, and extracting a data set for training the model from the updated knowledge graph; S6. Risk prediction, applying the trained risk prediction model to new relevant data, analyzing the output of the model, and integrating the analysis results to form an overall risk assessment report to predict potential risk points; S7. Setting a threshold for the error rate of the risk prediction model, continuously monitoring and dynamically updating. When the error rate exceeds the set threshold, it is confirmed that concept drift has occurred, and the knowledge graph and the risk prediction model need to be updated again.
[0006] As a further method, before the feature extraction, cross-modal shared features are also included, specifically including the following steps: Input modalities: text modality, numerical modality, and graph modality; Vector transformation: using word embedding to convert text into a vector, using linear transformation to convert numerical data into a vector, and using graph embedding method to convert graph data into a vector; Shared fusion: splicing the embedding vectors of different modalities together to form multi-modal features.
[0007] As a further method, the specific steps for extracting dynamic interaction key feature information based on the time-frequency analysis are as follows: The long - time signal is segmented into multiple short - time frames, and a window function is applied to each frame to make the signal exhibit pseudo - stationary characteristics within each time frame; Within each time frame, the Fourier transform is performed on the windowed signal to obtain the frequency components of the signal within that time frame; The Fourier transform results of each time frame are arranged into a matrix to form a time - frequency diagram. In the time - frequency diagram, the horizontal axis represents time, the vertical axis represents frequency, and the intensity of each point represents the amplitude of the corresponding frequency at the corresponding moment, observing the changes of the dynamic interaction feature signal in time and frequency.
[0008] As a further method, after collecting the relevant data of various data modalities in the step S1, it further includes cleaning the data to remove the noise and irrelevant information in the text.
[0009] Compared with the related technologies, the risk prediction method based on the knowledge graph provided by the present invention has the following beneficial effects: 1. Based on time - frequency analysis, the present invention extracts dynamic interaction key feature information from the time - series data contained in the existing knowledge graph, and adds the new features as additional information to the corresponding entities or relationships, updating the knowledge graph, realizing the transformation from static data analysis to dynamic behavior pattern mining. A threshold is set for the error rate of the risk prediction model, and continuous monitoring and dynamic update are carried out. When the error rate exceeds the set threshold, it is confirmed that concept drift has occurred, and the knowledge graph and the risk prediction model need to be updated again to avoid the situation that the meaning of some risk indicators may change over time due to concept drift, providing a new perspective and technical support for risk management in the financial industry.
[0010] 2. Through the effective integration and in - depth analysis of multi - source heterogeneous data, the present invention makes the risk prediction more accurate and reliable. The analysis results are integrated to form an overall risk assessment report, and the risk assessment results are applied to multiple scenarios such as credit approval, supply chain management, and investment decision - making, thereby improving the fineness and accuracy of risk prediction.
[0011] 3. By constructing supply chain, market competition, enterprise credit, and enterprise financial health graphs, the present invention forms a transparent industrial chain and supply chain transaction relationship and enterprise credit network, enabling capital providers such as banks to deeply understand the upstream and downstream industrial chain relationships and supply chain relationships of leading enterprises in key industrial chains during the credit approval process, effectively verifying the true trade background between enterprises, and improving the credit risk control level. BRIEF DESCRIPTION OF THE DRAWINGS
[0012] Figure 1 is a flowchart of the risk prediction method based on the knowledge graph provided by the present invention; Figure 2 is a flowchart of the shared features provided by the present invention; Figure 3 A flowchart for extracting key feature information of dynamic interaction provided by the present invention. Specific implementation manners
[0013] The present invention will be further described below in conjunction with the accompanying drawings and implementation manners.
[0014] Please refer to Figure 1 , Figure 2 and Figure 3 , where Figure 1 is a flowchart of a risk prediction method based on a knowledge graph provided by the present invention; Figure 2 is a flowchart of shared features provided by the present invention; Figure 3 is a flowchart for extracting key feature information of dynamic interaction provided by the present invention.
[0015] In the specific implementation process, as Figures 1 - 3 shown, the risk prediction method based on a knowledge graph includes the following steps: S1. Data collection: Collect relevant data of various data modalities from various sources, where the relevant data includes supply chain data, market competition data, enterprise credit history data, and enterprise financial health data. After collecting the relevant data of various data modalities, it also includes cleaning the data to remove noise and irrelevant information in the text; Data collection includes collecting supplier and customer information from channels such as enterprise annual reports, purchase orders, and sales orders; Collecting enterprise market share and competitor information from channels such as industry reports, market research reports, and news reports, and collecting market dynamic data such as price changes and new product releases; Collecting enterprise loan records, repayment records, and default records from channels such as financial institutions and credit rating agencies, and collecting credit rating data; Extracting data such as balance sheets, income statements, and cash flow statements from enterprise financial statements, and collecting financial indicators such as asset-liability ratios, current ratios, and net profit margins; S2. Entity recognition: Use natural language processing technology to identify key entities in the relevant data, and extract relationships between the key entities, such as supply chain relationships and market competition relationships; S3. Graph construction: Construct a graphical knowledge graph according to the identified key entities and the relationships between the key entities. Create nodes in the graph according to the identified entities, create edges in the graph according to the identified relationships, and assign attributes to each node and edge, such as name, relationship type, and weight; S4. Feature extraction: Extract key feature information of dynamic interaction from the time series data included in the existing knowledge graph based on time-frequency analysis, and add the new features as additional information to the corresponding entities or relationships to update the knowledge graph; S5. Risk model training: Using the updated knowledge graph data, evaluate the importance of the dynamic interaction key feature information to obtain the dynamic interaction importance score. Based on the neural network learning algorithm, construct a risk prediction model, and extract the data set for training the model from the updated knowledge graph; It should be noted that the steps for establishing the risk prediction model are as follows: Data preparation: Key entities, key entity relationships, extracted dynamic interaction key features, importance scores, and memory cell feature data; Feature engineering setup: According to the given importance scores (the importance score of the change in purchase volume is 0.6, the importance score of the change in market share is 0.2, the importance scores of the changes in loan amount and repayment date are 0.2, and the importance scores of the changes in asset - liability ratio and current ratio are 0.1); Use the long - short - term memory network to analyze the changing trend of continuous input variable data over time, and use the binary cross - entropy loss function to predict the difference between the probability distribution and the true label; The mathematical expression of the binary cross - entropy loss function is as follows: Among them, N represents the number of samples, is the true label (0 or 1) of the th sample, and is the predicted probability that the model assigns the sample to the positive class; Classification output: The activation function of the output layer uses Sigmoid. If the asset - liability ratio of an enterprise exceeds 20% of the industry average level or historical average, it is output that the enterprise has financial health risks; if the current ratio is lower than the safety threshold of 1.5, it is output that the enterprise has insufficient short - term solvency and liquidity risks; when the purchase volume of raw materials provided by a certain supplier increases month by month in the past year and the growth rate for three consecutive months exceeds 10%, it is output that there are potential supply chain pressures or demand forecasting errors; if the market share of an enterprise in the main market continues to decline compared with its competitors, it is output that the enterprise faces fierce market competition pressure risks; if the total loan amount of an enterprise is increasing continuously, but there are multiple cases of failure to repay on time, it is output that there are high credit default risks; Risk prediction results: Classified into high - risk, medium - risk, and low - risk according to the classification output results. Among them, the high - risk standard is the occurrence of 5 risk classifications; the medium - risk standard is the occurrence of 3 - 5 risk classifications; the low - risk standard is the occurrence of 1 - 2 risk classifications; S6. Risk prediction: Apply the trained risk prediction model to new relevant data, analyze the output of the model, integrate the analysis results to form an overall risk assessment report, and estimate potential risk points; S7. Set a threshold for the error rate of the risk prediction model, continuously monitor and dynamically update it. When the error rate exceeds the set threshold, it is confirmed that concept drift has occurred, and the knowledge graph and risk prediction model need to be updated again.
[0016] It should be noted that a reasonable error rate threshold is set according to business requirements. For example, in financial credit risk assessment, a very low error rate, such as 0.5%, needs to be maintained; while in supply chain risk management, considering the changes in the external environment, a slightly higher error rate, such as 3%, is required. In order to capture potential concept drift in a timely manner, an online automatic testing method is adopted, that is, a small part of the samples are randomly selected for immediate verification each time a new prediction is made. In this way, the latest feedback information can be obtained to help quickly respond to any abnormal situations. In addition to directly observing the error of a single prediction, the cumulative error distribution graph over a period of time can also be summarized regularly to find out whether there are systematic biases or signs of gradual deterioration.
[0017] It should be noted that the specific steps for natural language processing technology to identify key entities are as follows: Perform word segmentation on the text data of supply chain data, market competition data, enterprise credit history data, and enterprise financial health data, and split it into independent words or phrases; Extract the features useful for identifying key entities from the independent words or phrases.
[0018] It should be noted that the features useful for identifying key entities extracted include the following: for supply chain data, extract the supplier name, product name, and purchase volume features; for market competition data, extract the competitor name and market share features; for enterprise credit history data, extract the loan amount, repayment date, and credit rating features; for enterprise financial health data, extract the financial indicators and cash flow features.
[0019] As a further method, refer to Figure 2 As shown, before feature extraction, cross-modal shared features are also included, and the specific steps are as follows: Input modalities: text modality, numerical modality, and graph modality; Vector transformation: Use word embedding to convert text into vectors, use linear transformation to convert numerical data into vectors, and use graph embedding method to convert graph data into vectors; Shared fusion: Concatenate the embedding vectors of different modalities to form multi-modal features.
[0020] As a further method, refer to Figure 3 As shown, the specific steps for extracting dynamic interaction key feature information based on time-frequency analysis are as follows: The long-time signal is segmented into multiple short-time frames, and a window function is applied to each frame to make the signal exhibit pseudo-stationary characteristics within each time frame, where the long-time signal is a time signal based on the relevant data collected from various data modalities; Within each time frame, the windowed signal is subjected to Fourier transform to obtain the frequency components of the signal within that time frame; The Fourier transform results of each time frame are arranged into a matrix to form a time-frequency diagram, where the horizontal axis in the time-frequency diagram represents time, the vertical axis represents frequency, and the intensity of each point represents the amplitude of the corresponding frequency at the corresponding moment, so as to observe the changes of the dynamic interaction characteristic signal in time and frequency.
[0021] In a specific implementation process, the model input data in step S5 includes the following specific parameters: The model input data in step S5 includes the following specific parameters: "The purchase volume of new energy raw material X by Company A in the past year (actual data: monthly purchase volume, e.g., 100 tons in January, 110 tons in February, 115 tons in March, 120 tons in April, 125 tons in May, 130 tons in June, 135 tons in July, 140 tons in August, 145 tons in September, 150 tons in October, 155 tons in November, 160 tons in December)", "The market share of Company A has been decreasing quarter by quarter in the past year (actual data: quarterly market share, e.g., 20% in Q1, 18% in Q2, 18% in Q3, 15% in Q4)", "The loan amount of Company A has been increasing month by month in the past year, but the repayment date has been gradually delayed (actual data: monthly loan amount and repayment date, e.g., loan of 1 million in January, repayment date 30 days; loan of 1.1 million in February, repayment date 35 days, loan of 1.2 million in March, repayment date 40 days, loan of 1.3 million in April, repayment date 50 days, loan of 1.4 million in May, repayment date 55 days, loan of 1.6 million in June, repayment date 60 days, loan of 1.7 million in July, repayment date 65 days, loan of 1.8 million in August, repayment date 70 days, loan of 1.9 million in September, repayment date 75 days, loan of 2 million in October, repayment date 80 days, loan of 2.1 million in November, repayment date 85 days, loan of 2.2 million in December, repayment date 90 days)", "The asset-liability ratio of Company A has increased and the current ratio has decreased in the past year (actual data: monthly asset-liability ratio and current ratio, e.g., asset-liability ratio 50% in January, current ratio 1.5; asset-liability ratio 52% in February, current ratio 1.4, asset-liability ratio 54% in March, current ratio 1.3, asset-liability ratio 56% in April, current ratio 1.2, asset-liability ratio 58% in May, current ratio 1.0, asset-liability ratio 60% in June, current ratio 0.9, asset-liability ratio 62% in July, current ratio 0.8, asset-liability ratio 64% in August, current ratio 0.7, asset-liability ratio 66% in September, current ratio 0.6, asset-liability ratio 68% in October, current ratio 0.5, asset-liability ratio 70% in November, current ratio 0.4, asset-liability ratio 72% in December, current ratio 0.3)"; "The importance score of the change in purchase volume is 0.6"; "The importance score of the change in market share is 0.2"; "The importance score of the change in loan amount and repayment date is 0.2"; "The importance score of the change in asset-liability ratio and current ratio is 0.1"; Input the above input data into the trained risk prediction model. The risk prediction results are divided into high risk, medium risk, and low risk according to the classification output results. Among them, the high-risk standard is the occurrence of 5 risk classifications; the medium-risk standard is the occurrence of 3 - 5 risk classifications; the low-risk standard is the occurrence of 1 - 2 risk classifications. The output results are as follows: Number Key entity Key entity relationship Dynamic interaction feature Importance score Risk prediction result 1 Supplier A, Product X Supply relationship Feature value 1 (high) 0.6 High risk 2 Competitor B, Market share 20% Competition relationship Feature value 2 (medium) 0.2 Medium risk 3 Loan amount 1 million, Repayment date 30 days Credit relationship Feature value 3 (low) 0.2 Low risk In addition, the model outputs the following risk assessment report: Supply chain risk: As Company A's procurement volume of raw material X increases month by month, there is a risk of supply chain disruption or cost increase.
[0022] Market competition risk: Company A's market share decreases quarter by quarter, so there is a risk that its market share will be encroached upon by competitors.
[0023] Credit risk: The loan amount of Company A increases month by month, but the repayment date is gradually delayed, so there is a risk of credit default.
[0024] Financial risk: The asset-liability ratio of Company A rises and the current ratio drops, so there is a risk of exacerbated financial risk.
[0025] Through the description of the above implementation manners, those skilled in the art can clearly understand that each implementation manner can be realized by means of software plus a general hardware platform, and of course, it can also be realized by hardware. Based on such an understanding, the essence of the above technical solutions, or rather the part that makes contributions to the related technologies, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disc, etc., and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.
[0026] The above has shown and described the basic principles, main features and advantages of the present invention. For those skilled in the art, it is obvious that the present invention is not limited to the details of the above exemplary embodiments, and without departing from the spirit or basic features of the present invention, the present invention can be implemented in other specific forms. Therefore, from any point of view, the embodiments should be regarded as exemplary and non-limiting. The scope of the present invention is defined by the appended claims rather than the above description. Therefore, it is intended to embrace all changes falling within the meaning and scope of the equivalent elements of the claims in the present invention, and any reference signs in the claims should not be regarded as limiting the claims involved.
[0027] In addition, it should be understood that although this specification is described according to implementation manners, not every implementation manner only contains an independent technical solution. This narrative manner of the specification is only for clarity. Those skilled in the art should regard the specification as a whole, and the technical solutions in each embodiment can also be appropriately combined to form other implementation manners understandable to those skilled in the art.
Claims
1. A risk estimation method based on knowledge graph, characterized in that: The steps include: S1. Data collection: collecting data on supply chain, market competition, corporate credit history, and corporate financial health in multiple dimensions; S2, entity recognition, using natural language processing technology to identify key entities in relevant data and extract relationships between key entities; S3, graph construction, constructing a knowledge graph with a graphical structure based on the identified key entities and the relationships between them; S4, feature extraction, extracting dynamic interactive key feature information from the time series data contained in the existing knowledge graph based on time-frequency analysis, and adding new features as additional information to the corresponding entities or relationships to update the knowledge graph; S5. Risk model training: using the updated knowledge graph data, evaluate the importance of key feature information of dynamic interactions to obtain dynamic interaction importance scores, build a risk prediction model based on a neural network learning algorithm, and extract a data set for training the model from the updated knowledge graph; S6. Risk prediction: Apply the trained risk prediction model to new relevant data, analyze the output of the model, integrate the analysis results, form an overall risk assessment report, and estimate potential risk points; S7. Set a threshold for the error rate of the risk prediction model, continuously monitor and dynamically update it. When the error rate exceeds the set threshold, it is confirmed that concept drift has occurred and the knowledge graph and risk prediction model need to be updated.
2. The risk prediction method based on knowledge graph according to claim 1 is characterized in that: The dataset includes key entities, key entity relationships, extracted dynamic interaction key features, and importance scoring data.
3. The risk prediction method based on knowledge graph according to claim 1 is characterized in that: Before the feature extraction step, cross-modal feature sharing is also included, which specifically includes the following steps: Input mode: text mode, numerical mode and graph mode; Vector conversion: Use word embedding to convert text into vectors, use linear transformation to convert numerical data into vectors, and use graph embedding methods to convert graph data into vectors; Shared fusion: concatenate the embedding vectors of different modalities together to form multimodal features.
4. The risk prediction method based on knowledge graph according to claim 1, characterized in that: The specific steps of extracting dynamic interaction key feature information based on the time-frequency analysis are as follows: Split the long-time signal into multiple short-time frames and apply a window function to each frame, so that the signal exhibits pseudo-stationary characteristics in each time frame; In each time frame, the windowed signal is Fourier transformed to obtain the frequency component of the signal in the time frame; The Fourier transform results of each time frame are arranged into a matrix to form a time-frequency diagram, where the horizontal axis represents time, the vertical axis represents frequency, and the intensity of each point represents the amplitude of the corresponding frequency at the corresponding moment, observing the changes in time and frequency of the dynamic interaction characteristic signal.
5. The risk prediction method based on knowledge graph according to claim 1 is characterized in that: After collecting relevant data of various data modalities in the S1 step, the data is cleaned to remove noise and irrelevant information in the text.
6. The risk prediction method based on knowledge graph according to claim 1 is characterized in that: The specific steps of natural language processing technology to identify key entities are as follows: Perform word segmentation on text data of supply chain data, market competition data, corporate credit history data, and corporate financial health data, splitting them into independent words or phrases; Extract useful features from independent words or phrases to identify key entities.
7. The risk prediction method based on knowledge graph according to claim 6 is characterized in that: The features extracted to identify key entities include: for supply chain data, extracting supplier name, product name, and purchase volume features; for market competition data, extracting competitor name and market share features; for corporate credit history data, extracting loan amount, repayment date, and credit rating features; For corporate financial health data, extract financial indicators and cash flow characteristics.
Citation Information
Cited By
Industrial chain multi-relation modeling method based on dynamic relation graph
CN120911833A
An industry chain multi-relationship modeling method based on a dynamic relationship graph
CN120911833B
Multi-platform interactive public opinion intelligent monitoring analysis method and device and computer equipment
CN121117489A
Method and device for monitoring and analyzing public opinion of multi-platform interaction and computer equipment
CN121117489B