Method, device, equipment and medium for processing associated data based on knowledge graph
By constructing a knowledge graph-based associated data processing method, the problem of low accuracy in associated data mining in complex networks is solved, and more accurate risk assessment and management are achieved.
Patent Information
- Application Number
- CN202411417300.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-10-11
- Publication Date
- 2025-09-16
- Estimated Expiration
- 2044-10-11
AI Technical Summary
Existing technologies have difficulty effectively expressing and predicting associated data in complex networks, resulting in low accuracy in associated data mining. This is especially true in the financial and transportation fields, where it is difficult to accurately assess risks and traffic conditions.
Use knowledge graphs to build complex networks, obtain the attribute information of target nodes, calculate the score values and weights of associated nodes, build differential equations for associated data calculations, and generate associated data mining reports.
The accuracy of associated data mining has been improved, which enables more accurate identification of risk transmission paths and potential risks, and enhances the effectiveness of risk management.
Smart Images

Figure CN119377311B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the fields of data request processing and financial technology, and in particular to a method, device, electronic device and storage medium for processing associated data based on a knowledge graph. Background Art
[0002] Data mining is a highly automated process for analyzing large amounts of data to identify potential connections in decision support, based on artificial intelligence (such as machine learning) and pattern recognition techniques. Currently, data mining is widely used in business operations, transportation, healthcare, market analysis, and engineering design.
[0003] For example, in the enterprise sector (such as financial institutions), structured data tables are often used to represent the relationships between related data in order to mine and assess business risks (i.e., perform risk prediction). However, in many cases, complex related data that is difficult to clearly express using structured data tables may affect business risks.
[0004] For example, in the field of transportation, data mining can be used for traffic flow forecasting, route optimization, and accident prevention. However, the complex interactions between various dynamic factors or risk events in the transportation network (such as weather conditions, emergencies, and driver behavior) are difficult to represent using simple table-structured data, making it more difficult to accurately predict traffic conditions.
[0005] Therefore, how to improve the accuracy of association data mining is a technical problem that needs to be solved urgently. Summary of the Invention
[0006] In view of the above content, it is necessary to provide a method for processing associated data based on knowledge graphs. Its purpose is to utilize the capabilities of knowledge graphs in information expression, complex network construction and intelligent reasoning to calculate the association relationships between nodes in the knowledge graph to improve the accuracy of associated data mining.
[0007] In a first aspect, a method for processing associated data based on a knowledge graph is provided, comprising:
[0008] After receiving a user's request to calculate associated data for a target node in a preset knowledge graph, obtain first attribute information corresponding to the target node from a preset database, and calculate a first information score value of the target node based on the first attribute information, where the first attribute information is basic data associated with the target node;
[0009] Obtaining second attribute information corresponding to the target node from the preset database, determining a plurality of associated nodes associated with the target node in the preset knowledge graph based on the second attribute information, and calculating an association weight between the target node and each associated node based on the second attribute information, where the second attribute information is relationship data between the target node and each associated node;
[0010] Obtaining third attribute information corresponding to each associated node from the preset database, calculating a second information score value corresponding to each associated node and a score level of each associated node based on the third attribute information of each associated node, and determining a node state corresponding to each associated node in the preset knowledge graph based on the second information score value corresponding to each associated node, wherein the node state includes a susceptible node, an infected node, and an immune node, and the third attribute information is basic data associated with each associated node;
[0011] Obtain an association data calculation differential equation constructed in advance based on the susceptible node, the infected node and the immune node, substitute the association weight and the scoring level of each association node into the association data calculation differential equation for calculation, and obtain the third information scoring value of the target node; generate an association data mining report of the target node according to the first information scoring value and the third information scoring value and feed it back to the user.
[0012] In a second aspect, a device for processing associated data based on a knowledge graph is provided, comprising:
[0013] A first calculation module is configured to, upon receiving a user's request to calculate associated data for a target node in a preset knowledge graph, obtain first attribute information corresponding to the target node from a preset database, and calculate a first information score value of the target node based on the first attribute information, where the first attribute information is basic data associated with the target node;
[0014] a second calculation module, configured to obtain second attribute information corresponding to the target node from the preset database, determine a plurality of associated nodes associated with the target node in the preset knowledge graph based on the second attribute information, and calculate an association weight between the target node and each associated node based on the second attribute information, where the second attribute information is relationship data between the target node and each associated node;
[0015] a determination module, configured to obtain third attribute information corresponding to each associated node from the preset database, calculate a second information score value corresponding to each associated node and a score level of each associated node based on the third attribute information of each associated node, and determine a node state corresponding to each associated node in the preset knowledge graph based on the second information score value corresponding to each associated node, wherein the node state includes a susceptible node, an infected node, and an immune node, and the third attribute information is basic data associated with each associated node;
[0016] A generation module is used to obtain a correlation data calculation differential equation constructed in advance based on the susceptible node, the infected node and the immune node, substitute the correlation weight and the scoring level of each correlation node into the correlation data calculation differential equation for calculation, obtain the third information scoring value of the target node, generate a correlation data mining report of the target node based on the first information scoring value and the third information scoring value, and feed it back to the user.
[0017] In a third aspect, a computer device is provided, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the steps of the above-mentioned method for processing associated data based on the knowledge graph are implemented.
[0018] In a fourth aspect, a computer-readable storage medium is provided, which stores a computer program. When the computer program is executed by a processor, the steps of the above-mentioned method for processing associated data based on the knowledge graph are implemented.
[0019] Compared with the existing technology, the present invention extracts the first attribute information of the target node from the database to score the target node and obtains the first information score value. According to the second attribute information, the association weight between other nodes associated with the target node is determined. For each associated node, its third attribute information is used to calculate the second information score value of each associated node, and the associated nodes are classified into susceptible, infected or immune states. The association weight and the score level of each associated node are substituted into the associated data calculation differential equation to obtain the first information score value of the target node. According to the first information score value and the third information score value, an associated data mining report of the target node is generated. The present invention calculates the association relationship between each node in the knowledge graph through the knowledge graph's capabilities in information expression, complex network construction and intelligent reasoning and the associated data calculation differential equation constructed with "susceptible-infectious-immune", thereby improving the accuracy of associated data mining. BRIEF DESCRIPTION OF THE DRAWINGS
[0020] Figure 1 1 is a schematic diagram of an application environment of a method for processing associated data based on a knowledge graph in an embodiment of the present invention;
[0021] Figure 2 A schematic diagram of a process for processing associated data based on a knowledge graph according to an embodiment of the present invention;
[0022] Figure 3 A schematic diagram of a module of a knowledge graph-based associated data processing device provided by one embodiment of the present invention;
[0023] Figure 4 A schematic diagram of the structure of an electronic device for implementing a method for processing associated data based on a knowledge graph according to an embodiment of the present invention;
[0024] The purpose, features and advantages of the present invention will be further described with reference to the accompanying drawings and in conjunction with the embodiments. DETAILED DESCRIPTION
[0025] In order to make the purpose, technical solutions and advantages of the present invention more clearly understood, the present invention is further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not intended to limit the present invention. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of the present invention.
[0026] It should be noted that the descriptions of "first", "second", etc. in the present invention are for descriptive purposes only and should not be understood as indicating or implying their relative importance or implicitly indicating the number of the indicated technical features. Therefore, the features defined as "first" or "second" may explicitly or implicitly include at least one of such features. In addition, the technical solutions between the various embodiments can be combined with each other, but this must be based on the fact that they can be implemented by ordinary technicians in this field. When the combination of technical solutions is contradictory or cannot be implemented, it should be deemed that such a combination of technical solutions does not exist and is not within the scope of protection required by the present invention.
[0027] The method for processing associated data of basic knowledge graph provided by the embodiment of the present invention can be applied in Figure 1In an application environment, the client communicates with the server through a network. After the server receives a user's request for calculating associated data for a target node in a preset knowledge graph through the client, it can obtain the first attribute information corresponding to the target node from the preset database, calculate the first information score value of the target node based on the first attribute information, and the first attribute information is the basic data associated with the target node; obtain the second attribute information corresponding to the target node from the preset database, determine multiple associated nodes associated with the target node in the preset knowledge graph based on the second attribute information, calculate the association weight between the target node and each associated node based on the second attribute information, and the second attribute information is the relationship data between the target node and each associated node; obtain the corresponding values of each associated node from the preset database The third attribute information is used to calculate the second information score and the score level of each associated node based on the third attribute information of each associated node. Based on the second information score of each associated node, the node status of each associated node in the preset knowledge graph is determined, and the node status includes susceptible nodes, infected nodes, and immune nodes. The association data calculation differential equation constructed in advance based on the susceptible nodes, infected nodes, and immune nodes is obtained, and the association weight and the score level of each associated node are substituted into the association data calculation differential equation for calculation to obtain the third information score of the target node. Based on the first information score and the third information score, a association data mining report for the target node is generated in response to the request. The present invention is targeted at fields such as financial enterprises, transportation, medical health, market analysis, and engineering design. It utilizes the capabilities of knowledge graphs in information expression, complex network construction, and intelligent reasoning to calculate the association relationships between each node in the knowledge graph to improve the accuracy of association data mining. The client can be, but is not limited to, various personal computers, laptops, smartphones, tablet computers, and portable wearable devices. The server can be implemented as an independent server or a server cluster consisting of multiple servers. The present invention is described in detail below through specific embodiments.
[0028] Reference Figure 2 FIG2 is a flow chart of a method for processing associated data based on a knowledge graph according to an embodiment of the present invention. The method is executed by an electronic device.
[0029] In this embodiment, the method for processing associated data based on the knowledge graph includes:
[0030] S1. After receiving a user's request to calculate associated data for a target node in a preset knowledge graph, obtain the first attribute information corresponding to the target node from a preset database, and calculate the first information score value of the target node based on the first attribute information. The first attribute information is the basic data associated with the target node.
[0031] In this embodiment, after receiving a request from a user or a development program to calculate associated data for a target node, the identification information of the target node contained in the analysis request is analyzed, and based on the identification information of the target node, the first attribute information of the target node is extracted from a preset database or other information source platform. The first attribute information is the basic data associated with the target node.
[0032] If the target node is a target node corresponding to a target enterprise in the enterprise field, the basic data associated with the target node includes news and public opinion, financial statements, operating information, and litigation records of the target enterprise.
[0033] If the target node is a target node corresponding to a target road in the transportation field, the basic data associated with the target node include the daily traffic volume, design capacity, and accident frequency of the target road.
[0034] Correlation data calculation refers to the process of estimating the possibility of a risk event occurring in the future through quantitative analysis.
[0035] According to each keyword of the first attribute information, features corresponding to each basic data are extracted from the basic data associated with the target node, the extracted features are input into a preset score prediction model, and each feature is calculated separately using multiple tree functions of the preset score prediction model to generate multiple leaf function values, and the first information score value of the target node is generated based on the multiple leaf function values.
[0036] For example:
[0037] Suppose a commercial bank's risk management department receives a request to perform a correlation calculation on a target node (target enterprise) named "Technology Company A." The department retrieves the primary attribute information of "Technology Company A" from a pre-set database. For example, the department obtains both positive and negative news about "Technology Company A" from news media, social media, industry reports, and other sources. Positive news: Technology Company A recently received funding for a major scientific research project from the government. Negative news: A product of Technology Company A was recalled due to quality issues.
[0038] For example, extracted from the annual reports, quarterly reports and other financial statements submitted by A Technology Company to regulatory agencies, such as total assets: 800 million yuan, total liabilities: 300 million yuan, current ratio: 1.8, quick ratio: 1.2, net profit margin: 12%.
[0039] The first attribute information extracts features of news and public opinion, financial statements, operating information, and litigation records. All extracted features are input into a pre-trained scoring prediction model for calculation. The model calculates the first information score for "Technology Company A" to be 0.35 (assuming the model output ranges from 0 to 1, with 0 indicating no risk and 1 indicating extremely high risk). Based on the calculation results, the first information score for "Technology Company A" is 0.35, which is a medium-low score. This means that "Technology Company A" currently faces certain risks, but is generally controllable.
[0040] In one embodiment, calculating the first information score of the target node according to the first attribute information includes:
[0041] Extracting features corresponding to each basic data from the basic data associated with the target node according to each keyword of the first attribute information;
[0042] Inputting the extracted features into a preset score prediction model, wherein the preset score prediction model is generated by an extreme gradient boosting decision tree model;
[0043] Utilizing the plurality of tree functions of the preset scoring prediction model to calculate each feature respectively, and generating a plurality of leaf function values;
[0044] A first information score value of the target node is generated based on a plurality of leaf function values.
[0045] If the target node is a target node corresponding to a target road in the field of transportation, then according to each keyword of the first attribute information (daily traffic flow, design capacity, accident frequency), the characteristics corresponding to each basic data are extracted from the first attribute information, such as the characteristics of daily traffic flow, the characteristics of design capacity, and the characteristics of accident frequency. Specifically, the average traffic flow, peak traffic flow, and minimum traffic flow of the target road within a preset time period are obtained as the characteristics of daily traffic flow, the number of lanes, width, and speed limit included in the design standard of the target road are obtained as the characteristics of design capacity, and the number and severity of accidents on the target road within a preset time period are obtained as the characteristics of accident frequency.
[0046] If the target node is the target node corresponding to the target enterprise in the enterprise field, then according to the various keywords of the first attribute information (news and public opinion, financial statements, business information, litigation records), the characteristics corresponding to each basic data are extracted from the first attribute information, such as the characteristics of news and public opinion, the characteristics of financial statements, the characteristics of business information, and the characteristics of litigation records.
[0047] In one embodiment, extracting the features of the news and public opinion, the features of the financial statements, the features of the business information, and the features of the litigation records includes:
[0048] Performing sentiment analysis on the news public opinion, and using the number and influence of positive news and negative news of the target node obtained from the analysis as the characteristics of the news public opinion;
[0049] Identifying key financial indicators in the financial statements and using the key financial indicators as features of the financial statements, the key financial indicators including current ratio, quick ratio, debt-to-asset ratio, and net profit margin;
[0050] extracting the product sales, product market share, and product competitiveness of the target node from the business information as features of the business information;
[0051] The number of litigation cases and the ratio of winning to losing of the target node within a preset time period are obtained from the litigation records as features of the litigation records.
[0052] All the feature data extracted above are input into a pre-trained score prediction model. The score prediction model is trained through machine learning methods, such as the extreme gradient boosting decision tree model (XGBoost). Within the score prediction model, multiple tree functions are used to calculate the input features. Each tree will split the features and ultimately generate a function value for a leaf node. The function values of each leaf node are integrated to generate the final first information score value of the target node.
[0053] The Extreme Gradient Boosting Decision Tree (XGBoost) model is an improvement on the GBDT algorithm. It builds a multi-tree fusion model and is a type of ensemble method. It uses distributed data loading and training, performs a second-order Taylor expansion on the loss function, and adds a regularization term to the objective function to find the optimal solution overall. This balances the objective function's reduction with model complexity to avoid overfitting.
[0054] S2. Obtain second attribute information corresponding to the target node from the preset database, determine multiple associated nodes associated with the target node in the preset knowledge graph based on the second attribute information, and calculate the association weight between the target node and each associated node based on the second attribute information, where the second attribute information is the relationship data between the target node and each associated node.
[0055] In this embodiment, it is necessary to obtain the second attribute information of the target node. If the target node is the target node corresponding to the target enterprise in the enterprise field, the relationship data between the target node and each associated node includes the asset relationship and transaction relationship between the target node and its associated nodes. The asset relationship includes but is not limited to: equity relationship (for example, the equity holdings of the target node and its associated nodes), personnel relationship (for example, the relationship between the target node and key figures such as executives and directors of the associated nodes), debt relationship (for example, the loan relationship between the target node and the bank or associated node), guarantee relationship (for example, the situation where the target node provides a guarantee for the associated node). Transaction relationship includes but is not limited to: related transactions (for example, business dealings such as purchase, sale, leasing, and services between the target node and the associated nodes), bidding (for example, the bidding relationship between the target node and the associated nodes), and supply chain (for example, the supply chain relationship between the target node and suppliers, distributors, etc.).
[0056] If the target node corresponds to a target road in the transportation sector, the relationship data between the target node and its associated nodes includes the geographic location relationship and traffic flow relationship between the target node and its associated nodes. The geographic location relationship includes, but is not limited to: the intersection or junction between the target road and other roads, the geographic coordinates of the starting and ending points of the target road, and the distance between the target road and important landmarks (such as airports, train stations, and commercial centers). The traffic flow relationship includes, but is not limited to: the comparison of the average daily traffic volume of the target road and adjacent roads, the traffic exchange between the target road and other roads during peak hours, and the impact of the target road as a main road on surrounding branches.
[0057] According to the second attribute information, multiple associated nodes associated with the target node are determined in the preset knowledge graph, and the association weight between the target node and each associated node is calculated according to the second attribute information.
[0058] In one embodiment, before receiving a user's request to calculate associated data for a target node in a preset knowledge graph, the method further includes:
[0059] Obtain a preset number of target objects from a preset database, and determine each node of the preset knowledge graph according to the preset number of target objects, wherein the target objects are entity objects corresponding to each node in the preset knowledge graph;
[0060] Edges between each node are generated based on the relationship data between each target object, and the preset knowledge graph is constructed based on each node and the edges between each node. The relationship data includes the relationship closeness, relationship depth and relationship type between each target object.
[0061] In the enterprise domain, a knowledge graph is pre-built based on the asset and transaction relationships between the target node and each associated node. In the knowledge graph, each node represents a company, and the lines between nodes represent the asset and transaction relationships between companies. The knowledge graph can be stored and queried using graph database technology to ensure efficient data processing capabilities. By constructing a knowledge graph, it is possible to clearly identify which companies have associated relationships with the target node, as well as the nature of these relationships (asset relationships, transaction relationships, etc.). This helps to accurately identify the source and path of risk transmission.
[0062] Obtain the mapping between asset relationships, transaction relationships, and edge information in the pre-set knowledge graph. This step aims to map real-world business relationships to node connections in the graph, ensuring the accuracy and completeness of the knowledge graph. Based on the mapping relationships, traverse all nodes in the pre-set knowledge graph to identify business nodes directly connected to the target node. These nodes represent multiple associated nodes associated with the target node. For each associated node, calculate the association weight between the target node and the node. This association weight is calculated based on factors such as relationship closeness, relationship depth, and relationship type.
[0063] For example, the second attribute information for "Technology Company A" is obtained from a bank's database. Based on the asset and transaction relationships between Technology Company A and each associated node, multiple associated nodes associated with Technology Company A are identified in the pre-set knowledge graph, and the association weights between Technology Company A and each associated node are calculated. The knowledge graph is traversed based on the asset and transaction relationships to identify enterprise nodes directly connected to Technology Company A. These nodes represent the multiple associated nodes associated with Technology Company A. For each associated node, the closeness, depth, and type of the relationship between Technology Company A and each associated node are obtained from the knowledge graph, and the association weights between the target node and each associated node are calculated.
[0064] In one embodiment, determining, based on the second attribute information, a plurality of associated nodes associated with the target node in the preset knowledge graph includes:
[0065] Obtaining the mapping relationship between the relationship data between the target node and each associated node and the edges of the preset knowledge graph;
[0066] All nodes of the preset knowledge graph are traversed according to the mapping relationship to obtain multiple associated nodes associated with the target node.
[0067] In the enterprise field, the asset relationship and transaction relationship between the target node and the associated nodes are obtained, and mapped with the edge information in the preset knowledge graph. These asset relationships and transaction relationships are mapped with the edge information in the preset knowledge graph to ensure that each asset relationship or transaction relationship can find the corresponding edge in the preset knowledge graph. All nodes of the preset knowledge graph are traversed according to the mapping relationship to determine the associated nodes related to the target node.
[0068] Specific steps: Set the starting node as the target node and ensure that you have a unique identifier for the target node (such as a company ID). Use a graph traversal algorithm (such as depth-first search (DFS) or breadth-first search (BFS)) to traverse the graph and find nodes directly connected to the target node. For example, choose an appropriate traversal algorithm based on the scale and characteristics of the pre-set knowledge graph. If the pre-set knowledge graph is small and not highly connected, DFS can be used; if the pre-set knowledge graph is large and highly connected, BFS is recommended.
[0069] Starting from the target node, its adjacent nodes are accessed. For each adjacent node, the mapping relationship is checked to see if there is an asset relationship or transaction relationship. For each node that has an asset relationship or transaction relationship with the target node, its identifier is recorded and considered as an associated node of the target node. Finally, based on the recorded associated node identifiers, multiple associated nodes associated with the target node are obtained. In other embodiments, an identifier list can be created to record the identifiers of all enterprises that have a direct relationship with the target node.
[0070] In one embodiment, calculating the association weight between the target node and each associated node according to the second attribute information includes:
[0071] According to the relationship data between the target node and each associated node, the relationship closeness, relationship depth and relationship type between the target node and each associated node are obtained from the preset knowledge graph;
[0072] The relationship closeness, the relationship depth and the relationship type are substituted into a preset association risk factor formula for calculation to obtain the association weight between the target node and each associated node.
[0073] In the corporate sector, relationship closeness measures the strength of the connection between two companies, quantified by the following aspects (equity ownership, senior management positions, transaction frequency, and sponsorship relationships):
[0074] Equity holding ratio: If the target company holds shares in another company, the higher the holding ratio, the closer the relationship. For example, if the target company holds more than 50% of the shares of another company, the relationship is considered to be very close. Executive position: If the executives of the target company also hold positions in the associated node, then the relationship is also relatively close. For example, the chairman of the target company is also a member of the board of directors of another company. Transaction frequency: If there are frequent transaction records between the target company and the other company, then the relationship is also relatively close. For example, the target company purchases raw materials from the other company every month, then a high transaction frequency indicates a close relationship. Guarantee relationship: If the target company has provided multiple guarantees for the other company, this is also a sign of a close relationship. For example, the target company has provided loan guarantees for the other company many times.
[0075] Relationship Depth: This measures the number of indirect relationships between two companies, including those formed through other companies. It is quantified using the following criteria: Direct Relationship: If a direct equity, transaction, or guarantee relationship exists between the target company and another company, the relationship depth is 1. For example, the target company directly holds shares in another company. Indirect Relationship: If the relationship between the target company and another company is established through a third company, the relationship depth is 2. For example, the target company has established a relationship with another company through an intermediary company. Multi-Level Indirect Relationship: If a relationship is established through multiple companies, the relationship depth increases accordingly. For example, if the target company has established a relationship with the ultimate company through an intermediary company and then through another intermediary company, the relationship depth is 3.
[0076] The relationship type is used to identify the type of relationship between two companies, including but not limited to the following (competitive relationship, upstream and downstream supply chain, partnership, equity relationship, transaction relationship, guarantee relationship): Competitive relationship: The target company and another company have a competitive relationship in the same industry, such as competition for market share. The closeness of this competitive relationship can be measured by market share comparison and competitor analysis. Upstream and downstream supply chain relationship: There is a supply chain relationship between the target company and another company, such as the target company is the supplier or customer of the other company. The closeness of this upstream and downstream supply chain can be measured by the stability and degree of dependence of the supply chain. Partnership: There is a strategic partnership between the target company and another company, such as joint research and development, joint marketing, etc. The closeness of this partnership can be measured by the closeness of the cooperation agreement.
[0077] Calculating the association weights between the target node and each associated node can quantify the degree of influence of each associated node on the target enterprise during the risk transmission process. This is crucial for subsequent associated data calculation and mining, and can help determine which associated nodes have the most significant impact on the target node.
[0078] The calculation results of the correlation weight can be used to further analyze the transmission effect of risks between enterprises. By calculating the influencing factors of different relationship types (such as equity relationships and debt relationships), the possibility and intensity of risk transmission can be more accurately evaluated.
[0079] According to the relationship closeness, relationship depth and relationship type, the preset association risk factor formula is substituted for calculation to obtain the association weight between the target node and each associated node.
[0080] As time goes by, the relationships between enterprises will change. By regularly updating the knowledge graph and recalculating the association weights, dynamic monitoring of enterprise risk status can be achieved, and new risk transmission paths and risk sources can be discovered in a timely manner.
[0081] In one embodiment, the calculation of the association weight between the target node and each associated node by substituting the relationship closeness, the relationship depth, and the relationship type into a preset association risk factor formula includes:
[0082] Multiplying the relationship closeness, the relationship depth, and the relationship type between the target node and any associated node to obtain an association weight of the any associated node to the target node;
[0083] Repeat the step of calculating the association weight of any associated node to the target node to obtain the association weight between the target node and each associated node.
[0084] The preset correlation risk factor formula is:
[0085] k(t,i)=∑w(i,j)*d(i,j)*p
[0086] Wherein, k(t,i) is the association weight of node i corresponding to the target node at time t when affected by the associated node j, w(i,j) is the closeness of the relationship between the node i and the node j, d(i,j) is the depth of the relationship between the node i and the node j, and p is the type of relationship between the node i and the node j.
[0087] In the enterprise sector, the closeness of a relationship can be quantified based on actual circumstances. For example, equity holdings: If the target node holds more than 50% of the shares of the associated node, the closeness of the relationship is high. Transaction frequency: If the target node and the associated node have dozens of transactions per year, the closeness of the relationship is high.
[0088] The depth of a relationship can be expressed as the number of direct or indirect relationships. For example, a direct relationship has d(i, j) = 1, while an indirect relationship through a third party has d(i, j) = 2.
[0089] Different types of relationships have different impacts on risk transmission, and therefore require different influencing factors. For example, equity relationships (p_equity), transaction relationships (p_trade), and guarantee relationships (p_guarantee) can be determined based on historical data analysis or expert experience.
[0090] For example:
[0091] Assume that the relationship information between target node A and associated nodes B and C is as follows:
[0092] A and B: relationship closeness w(A, B) = 0.8 (high shareholding ratio); relationship depth d(A, B) = 1 (direct relationship); relationship type is equity relationship, impact factor p_equity = 1.5.
[0093] A and C: Relationship closeness w(A, C) = 0.5 (average transaction frequency); relationship depth d(A, C) = 2 (indirect relationship through a third party); relationship type is transactional, impact factor p_trade = 1.0.
[0094] Based on the above information, the association weights are calculated as follows: k(A,B)=0.8×1×1.5=1.2, k(A,C)=0.5×2×1.0=1.0. Through the above steps, the association weights between the target node and each associated node can be quantified, thereby providing accurate data support for subsequent association data calculation and mining.
[0095] S3. Obtain the third attribute information corresponding to each associated node from the preset database, calculate the second information score value corresponding to each associated node and the score level of each associated node based on the third attribute information of each associated node, and determine the node status corresponding to each associated node in the preset knowledge graph based on the second information score value corresponding to each associated node, wherein the node status includes susceptible nodes, infected nodes, and immune nodes, and the third attribute information is basic data associated with each associated node.
[0096] In this embodiment, the method of obtaining the third attribute information of each associated node is the same as that of obtaining the second attribute information of the target node. Each feature is extracted from the third attribute information of each associated node and input into a pre-trained scoring prediction model for calculation to obtain the second information scoring value of each associated node.
[0097] The enterprise's rating is determined based on pre-set rating standards. Specifically, different rating standards are set based on the enterprise's second information score. For example, a low risk rating indicates a low inherent risk. A medium risk rating indicates a moderate inherent risk. A high risk rating indicates a high inherent risk. Based on the calculated second information score, the rating of each associated node is determined.
[0098] Based on the second information score of each associated node, its corresponding node status in the preset knowledge graph is determined. Node status includes susceptible nodes, infected nodes, and immune nodes. A susceptible node is defined as one with a high second information score, an ongoing risk event, or a source of infection for a risk event, which may transmit the risk to other companies. An infected node is defined as one with a high second information score, which is likely to be infected by a risk event and has close ties with companies in susceptible nodes. An immune node is defined as one with a low second information score, which is not easily affected by external risks and has fewer connections with companies in susceptible or infected nodes.
[0099] Through the above steps, we can accurately evaluate the second information score of each associated node and classify it into different node states. We can better understand its position in the entire knowledge graph, identify potential risk transmission paths, and take corresponding risk management measures.
[0100] In the enterprise field, if an enterprise finds itself in the state of an infected node or an immune node, it can strengthen monitoring of risky enterprises and take measures to reduce contact with high-risk enterprises in susceptible nodes to reduce the risk of infection.
[0101] In the field of transportation, if the target road is found to be in the state of an infected node or an immune node, the risk of the target road is low and it is not easily affected by external factors, so the existing management and maintenance level can be maintained. If the target road is found to be in a susceptible node (such as severe traffic congestion), immediate action is taken to resolve known problems, such as repairing damaged road surfaces, clearing obstacles, or strengthening law enforcement on traffic violations in the target road area.
[0102] In one embodiment, determining the node status corresponding to each associated node in the preset knowledge graph according to the second information score of each associated node includes:
[0103] Obtaining a predefined classification standard for node status, wherein the classification standard is used to classify node status into three types, and each type of node status has a corresponding different threshold;
[0104] The second information score value of each associated node is compared with a set threshold value of the classification standard, and the corresponding associated node state is defined according to the comparison result.
[0105] The node status classification criteria are predefined, and node status is divided into three types: susceptible nodes, infected nodes, and immune nodes. Different thresholds are set for each type of node status.
[0106] In the enterprise field, different thresholds are set to distinguish different types of node status based on the second information score of each enterprise. For example:
[0107] Immune nodes: The second information score is low, the enterprise is relatively stable and not easily affected by risks. The threshold range is: [0, 0.3] (low risk range). Infected nodes: The second information score is moderate. Although the current risk is low, it is easily affected by external risks. The threshold range is: [0.3, 0.6] (medium risk range). Susceptible nodes: The second information score is high, the enterprise faces a high risk and may spread the risk to other enterprises. The threshold range is: [0.6, 1] (high risk range).
[0108] The second information score of each associated node is compared with the threshold set above, and the node status is determined based on the comparison result. If the second information score falls within the interval [0, 0.3], the enterprise is defined as an immune node. If the second information score falls within the interval (0.3, 0.6), the enterprise is defined as an infected node. If the second information score falls within the interval (0.6, 1], the enterprise is defined as a susceptible node.
[0109] Through the above steps, the inherent risks of each associated node can be accurately assessed and divided into different node states, which can help enterprises better understand their position in the entire knowledge graph, identify potential risk transmission paths, and take corresponding risk management measures.
[0110] S4. Obtain an association data calculation differential equation constructed in advance based on the susceptible nodes, the infected nodes, and the immune nodes, substitute the association weights and the scoring levels of each association node into the association data calculation differential equation for calculation, and obtain a third information scoring value of the target node. Generate an association data mining report of the target node based on the first information scoring value and the third information scoring value and feed it back to the user.
[0111] In this embodiment, a pre-built set of correlation data calculation equations is obtained. The set of correlation data calculation equations is constructed based on the states of susceptible nodes, infected nodes, and immune nodes, and takes into account the correlation weights and ratings between the nodes.
[0112] In the enterprise field, there are many situations in the process of calculating related data:
[0113] a) With the target node as the center, radiate the corresponding risk transmission probability, where the risk is the score type of the target node itself;
[0114] b) Centered around the infected associated node, there may be multiple source risk radiations, requiring further optimization of the superposition and offsetting effects between enterprise risks;
[0115] c) Risk contagion caused by source risks that are not caused by corporate risks, such as macroeconomic downturns or industry cyclical risks, needs to be considered.
[0116] Based on the various situations existing in the above risk transmission probability process, the enterprise nodes are divided into the states of susceptible nodes, infected nodes and immune nodes to construct a specific set of correlation data calculation equations. The specific assumptions are as follows:
[0117]
[0118] The associated data calculation equations are:
[0119]
[0120] Among them, S(t) is the proportion of susceptible nodes, I(t) is the proportion of infected nodes, R(t) is the proportion of infected nodes, t is time t, and it expresses the proportion of associated nodes in the three types of states at time t. Then S(t)+I(t)+R(t)=1, and λ is the risk transmission probability of all associated nodes transmitting risk events to the target node.
[0121] Calculate the probability of risk transmission from associated node j to the target node: λ(i) = k(t,i)T(j), where k(t,i) is the association weight of node i, corresponding to the target node, under the influence of associated node j at time t, and T(j) is the rating of associated node j. k(t,i) is influenced by factors such as the closeness of the relationship, the degree of association, and the type of relationship. The closer the relationship, the more direct the association, and the more deeply the relationship type is influenced by the risk type, the larger k(t,i), and thus the greater the probability of risk transmission.
[0122] Select enterprises that have a direct relationship with the target node to form node pairs, and calculate the unilateral risk transmission probability between the enterprise node pairs. Considering that in reality, the target node will be affected by risk events of multiple associated nodes, constructing a group of associated data calculation equations can perform multilateral probability calculations in a one-time and comprehensive manner, thereby obtaining the comprehensive probability of the target node being infected, which can improve the efficiency of associated data calculation.
[0123] Substitute the association weights between the target node and each associated node, as well as the ratings of each associated node, into the risk transmission calculation formula to obtain the risk transmission probability. For example, calculate the risk transmission probability of associated node j from a risk event to the target node, λ(i) = k(t,i)T(j).
[0124] The calculated risk transmission probability is substituted into the associated data calculation equation group for calculation to obtain the third information score value of the target node, and an associated data mining report (for example, a data analysis view) of the target node is generated based on the first information score value and the third information score value.
[0125] The correlation data calculation equations used in the present invention are based on the classic SIR epidemic model and can play the following roles:
[0126] 1. The associated data calculation equations can simultaneously consider the primary information score of the target node and the tertiary information score provided by associated nodes, making the associated data calculation more comprehensive. This approach not only focuses on the health of the target node itself, but also considers the impact of the external environment on the target node.
[0127] 2. The linked data computational equations can simulate the dynamic propagation of risk events between nodes. Because the relationships and risk profiles between nodes are constantly changing, this allows nodes to monitor the evolution of risks in real time and adjust risk management strategies in a timely manner.
[0128] 3. By constructing a set of equations for calculating associated data, the impact of multiple associated nodes on the risk of the target node can be calculated at one time, avoiding the need to calculate the probability of risk transmission between each node one by one, thereby improving calculation efficiency.
[0129] To sum up, in the above steps S1-S4, the present invention extracts the first attribute information of the target node from the database to score the target node to obtain a first information score value, and determines the association weights between other nodes associated with the target node based on the second attribute information. For each associated node, its third attribute information is used to calculate the second information score value of each associated node, and the associated nodes are classified into susceptible, infected or immune states. The association weights and the score levels of each associated node are substituted into the associated data calculation differential equation to obtain the first information score value of the target node, and an associated data mining report of the target node is generated based on the first information score value and the third information score value. The present invention calculates the association relationship between each node in the knowledge graph through the knowledge graph's capabilities in information expression, complex network construction and intelligent reasoning and the associated data calculation differential equation constructed with "susceptible-infectious-immune", thereby improving the accuracy of associated data mining.
[0130] like Figure 3 , which is a module diagram of an associated data processing device based on a knowledge graph provided by one embodiment of the present invention.
[0131] The knowledge graph-based associated data processing device 100 described in the present invention can be installed in an electronic device. Depending on the functionality implemented, the knowledge graph-based associated data processing device 100 can include a first calculation module 110, a second calculation module 120, a determination module 130, and a generation module 140. A module, also referred to as a unit, is a series of computer program segments that can be executed by an electronic device processor and perform a fixed function, and is stored in the electronic device's memory.
[0132] In this embodiment, the functions of each module / unit are as follows:
[0133] A first calculation module 110 is configured to, upon receiving a user's request to calculate associated data for a target node in a preset knowledge graph, obtain first attribute information corresponding to the target node from a preset database, and calculate a first information score value for the target node based on the first attribute information, where the first attribute information is basic data associated with the target node;
[0134] A second calculation module 120 is configured to obtain second attribute information corresponding to the target node from the preset database, determine a plurality of associated nodes associated with the target node in the preset knowledge graph based on the second attribute information, and calculate an association weight between the target node and each associated node based on the second attribute information, where the second attribute information is relationship data between the target node and each associated node;
[0135] Determination module 130, configured to obtain third attribute information corresponding to each associated node from the preset database, calculate a second information score value and a score level corresponding to each associated node based on the third attribute information of each associated node, and determine a node state corresponding to each associated node in the preset knowledge graph based on the second information score value corresponding to each associated node, wherein the node state includes a susceptible node, an infected node, and an immune node, and the third attribute information is basic data associated with each associated node;
[0136] The generation module 140 is used to obtain the association data calculation differential equation constructed in advance based on the susceptible node, the infected node and the immune node, substitute the association weight and the scoring level of each association node into the association data calculation differential equation for calculation, obtain the third information scoring value of the target node, generate the association data mining report of the target node according to the first information scoring value and the third information scoring value, and feed it back to the user.
[0137] In one embodiment, the first calculation module 110 is specifically configured to:
[0138] Extracting features corresponding to each basic data from the basic data associated with the target node according to each keyword of the first attribute information;
[0139] Inputting the extracted features into a preset score prediction model, wherein the preset score prediction model is generated by an extreme gradient boosting decision tree model;
[0140] Utilizing the plurality of tree functions of the preset scoring prediction model to calculate each feature respectively, and generating a plurality of leaf function values;
[0141] A first information score value of the target node is generated based on a plurality of leaf function values.
[0142] In one embodiment, the first calculation module 110 is specifically configured to:
[0143] Obtain a preset number of target objects from a preset database, and determine each node of the preset knowledge graph according to the preset number of target objects, wherein the target objects are entity objects corresponding to each node in the preset knowledge graph;
[0144] Edges between each node are generated based on the relationship data between each target object, and the preset knowledge graph is constructed based on each node and the edges between each node. The relationship data includes the relationship closeness, relationship depth and relationship type between each target object.
[0145] In one embodiment, the second calculation module 120 is specifically configured to:
[0146] Obtaining the mapping relationship between the relationship data between the target node and each associated node and the edges of the preset knowledge graph;
[0147] All nodes of the preset knowledge graph are traversed according to the mapping relationship to obtain multiple associated nodes associated with the target node.
[0148] In one embodiment, the second calculation module 120 is specifically configured to:
[0149] According to the relationship data between the target node and each associated node, the relationship closeness, relationship depth and relationship type between the target node and each associated node are obtained from the preset knowledge graph;
[0150] The relationship closeness, the relationship depth and the relationship type are substituted into a preset association risk factor formula for calculation to obtain the association weight between the target node and each associated node.
[0151] In one embodiment, the second calculation module 120 is specifically configured to:
[0152] Multiplying the relationship closeness, the relationship depth, and the relationship type between the target node and any associated node to obtain an association weight of the any associated node to the target node;
[0153] Repeat the step of calculating the association weight of any associated node to the target node to obtain the association weight between the target node and each associated node.
[0154] In one embodiment, the determination module 130 is specifically configured to:
[0155] Obtaining a predefined classification standard for node status, wherein the classification standard is used to classify node status into three types, and each type of node status has a corresponding different threshold;
[0156] The first information score value of each associated node is compared with a set threshold value of the classification standard, and the node state of the corresponding associated node is defined according to the comparison result.
[0157] like Figure 4 , which is a schematic diagram of the structure of an electronic device for implementing a method for processing associated data based on a knowledge graph provided by an embodiment of the present invention.
[0158] In this embodiment, the electronic device 1 includes, but is not limited to, a memory 11, a processor 12, and a network interface 13, which can be interconnected through a system bus. The memory 11 stores an associated data processing program 10 based on a knowledge graph, and the associated data processing program 10 based on a knowledge graph can be executed by the processor 12. Figure 4 Only the electronic device 1 having components 11-13 and the associated data processing program 10 based on the knowledge graph is shown. It can be understood by those skilled in the art that Figure 4 The structure shown does not constitute a limitation on the electronic device 1 , and the electronic device 1 may include fewer or more components than shown in the figure, or combine certain components, or arrange the components differently.
[0159] The memory 11 includes internal memory and at least one type of readable storage medium. The memory provides a cache for the operation of the electronic device 1; the readable storage medium can be a non-volatile storage medium such as flash memory, a hard disk, a multimedia card, a card-type memory (e.g., SD or DX memory), random access memory (RAM), static random access memory (SRAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), programmable read-only memory (PROM), magnetic storage, a magnetic disk, an optical disk, etc. In some embodiments, the readable storage medium can be an internal storage unit of the electronic device 1; in other embodiments, the non-volatile storage medium can also be an external storage device of the electronic device 1, such as a plug-in hard disk equipped on the electronic device 1, a Smart Media Card (SMC), a Secure Digital (SD) card, a Flash Card, etc. In this embodiment, the readable storage medium of the memory 11 is generally used to store the operating system and various application software installed on the electronic device 1, such as the code of the knowledge graph-based associated data processing program 10 in one embodiment of the present invention. In addition, the memory 11 can also be used to temporarily store various types of data that have been output or are to be output.
[0160] In some embodiments, the processor 12 may be a central processing unit (CPU), a controller, a microcontroller, a microprocessor, or other data processing chip. The processor 12 is generally used to control the overall operation of the electronic device 1, such as performing control and processing related to data interaction or communication with other devices. In this embodiment, the processor 12 is used to execute program code stored in the memory 11 or process data, such as executing the knowledge graph-based associated data processing program 10.
[0161] The network interface 13 may include a wireless network interface or a wired network interface, and the network interface 13 is used to establish a communication connection between the electronic device 1 and a terminal (not shown in the figure).
[0162] Optionally, the electronic device 1 may further include a user interface, which may include a display and an input unit such as a keyboard. The optional user interface may also include a standard wired interface and a wireless interface. Optionally, in some embodiments, the display may be an LED display, a liquid crystal display, a touch-sensitive liquid crystal display, and an OLED (Organic Light-Emitting Diode) touch device. The display may also be appropriately referred to as a display screen or a display unit, which is used to display information processed in the electronic device 1 and to display a visual user interface.
[0163] It should be understood that the embodiment is for illustration only and the scope of the patent application is not limited to this structure.
[0164] The knowledge graph-based associated data processing program 10 stored in the memory 11 of the electronic device 1 is a combination of multiple instructions. When running in the processor 12, it can achieve the following:
[0165] After receiving a user's request to calculate associated data for a target node in a preset knowledge graph, obtain first attribute information corresponding to the target node from a preset database, and calculate a first information score value of the target node based on the first attribute information, where the first attribute information is basic data associated with the target node;
[0166] Obtaining second attribute information corresponding to the target node from the preset database, determining a plurality of associated nodes associated with the target node in the preset knowledge graph based on the second attribute information, and calculating an association weight between the target node and each associated node based on the second attribute information, where the second attribute information is relationship data between the target node and each associated node;
[0167] Obtaining third attribute information corresponding to each associated node from the preset database, calculating a second information score value corresponding to each associated node and a score level of each associated node based on the third attribute information of each associated node, and determining a node state corresponding to each associated node in the preset knowledge graph based on the second information score value corresponding to each associated node, wherein the node state includes a susceptible node, an infected node, and an immune node, and the third attribute information is basic data associated with each associated node;
[0168] Obtain an association data calculation differential equation constructed in advance based on the susceptible node, the infected node and the immune node, substitute the association weight and the scoring level of each association node into the association data calculation differential equation for calculation, and obtain the third information scoring value of the target node; generate an association data mining report of the target node according to the first information scoring value and the third information scoring value and feed it back to the user.
[0169] Specifically, the processor 12 can refer to the specific implementation method of the above-mentioned knowledge graph-based associated data processing program 10. Figure 2 The description of the relevant steps in the corresponding embodiments will not be repeated here.
[0170] Furthermore, if the modules / units integrated in the electronic device 1 are implemented as software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. The computer-readable medium can be non-volatile or non-volatile. The computer-readable medium can include: any entity or device capable of carrying the computer program code, a recording medium, a USB flash drive, a mobile hard drive, a magnetic disk, an optical disk, a computer memory, or a read-only memory (ROM).
[0171] The computer-readable storage medium stores a knowledge graph-based associated data processing program 10, which can be executed by one or more processors. The specific implementation of the computer-readable storage medium of the present invention is basically the same as the above-mentioned embodiments of the knowledge graph-based associated data processing method, and will not be repeated here.
[0172] In the several embodiments provided by the present invention, it should be understood that the disclosed devices, apparatuses, and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the module division is merely a logical function division, and other division methods may be used in actual implementation.
[0173] The modules described as separate components may or may not be physically separate, and the components shown as modules may or may not be physical units, that is, they may be located in one place or distributed across multiple network elements. Some or all of the modules may be selected to achieve the purpose of the solution of this embodiment according to actual needs.
[0174] In addition, the functional modules in various embodiments of the present invention may be integrated into a single processing unit, each unit may exist physically separately, or two or more units may be integrated into a single unit. The aforementioned integrated units may be implemented in the form of hardware or hardware plus software functional modules.
[0175] It will be apparent to those skilled in the art that the present invention is not limited to the details of the exemplary embodiments described above, and that the present invention can be implemented in other specific forms without departing from the spirit or essential characteristics of the present invention.
[0176] Therefore, the embodiments should be considered in all respects as illustrative and non-restrictive, and the scope of the invention is defined by the appended claims rather than the foregoing description, and all changes that come within the meaning and range of equivalents of the claims are intended to be embraced therein. Any reference to a figure in a claim should not be construed as limiting the claim to which it relates.
[0177] Furthermore, it is clear that the word "comprising" does not exclude other units or steps, and the singular does not exclude the plural. Multiple units or devices recited in a system claim may also be implemented by a single unit or device through software or hardware. Second-order terms are used to indicate names and do not imply any particular order.
[0178] Finally, it should be noted that the above embodiments are intended only to illustrate the technical solutions of the present invention and are not intended to limit the present invention. Although the present invention has been described in detail with reference to the preferred embodiments, it should be understood by those skilled in the art that the technical solutions of the present invention may be modified or replaced by equivalents without departing from the spirit and scope of the technical solutions of the present invention. If software tools or components other than those of our company appear in the embodiments of this application, they are merely for illustration and do not represent actual use.
Claims
1. A method for processing associated data based on knowledge graph, characterized in that: The method comprises: After receiving a user's request to calculate associated data for a target node in a preset knowledge graph, obtain first attribute information corresponding to the target node from a preset database, and calculate a first information score value of the target node based on the first attribute information, where the first attribute information is basic data associated with the target node; Obtaining second attribute information corresponding to the target node from the preset database, determining a plurality of associated nodes associated with the target node in the preset knowledge graph based on the second attribute information, and calculating an association weight between the target node and each associated node based on the second attribute information, where the second attribute information is relationship data between the target node and each associated node; Obtaining third attribute information corresponding to each associated node from the preset database, calculating a second information score value corresponding to each associated node and a score level of each associated node based on the third attribute information of each associated node, and determining a node state corresponding to each associated node in the preset knowledge graph based on the second information score value corresponding to each associated node, wherein the node state includes a susceptible node, an infected node, and an immune node, and the third attribute information is basic data associated with each associated node; Obtain an association data calculation differential equation constructed in advance based on the susceptible node, the infected node and the immune node, substitute the association weight and the scoring level of each association node into the association data calculation differential equation for calculation, and obtain the third information scoring value of the target node; generate an association data mining report of the target node according to the first information scoring value and the third information scoring value and feed it back to the user.
2. The method for processing associated data based on a knowledge graph according to claim 1, wherein: The calculating the first information score value of the target node according to the first attribute information includes: Extracting features corresponding to each basic data from the basic data associated with the target node according to each keyword of the first attribute information; Inputting the extracted features into a preset score prediction model, wherein the preset score prediction model is generated by an extreme gradient boosting decision tree model; Utilizing the plurality of tree functions of the preset scoring prediction model to calculate each feature respectively, and generating a plurality of leaf function values; A first information score value of the target node is generated based on a plurality of leaf function values.
3. The method for processing associated data based on knowledge graph according to claim 1, wherein: Before receiving a user's request to calculate associated data for a target node in a preset knowledge graph, the method further includes: Obtain a preset number of target objects from a preset database, and determine each node of the preset knowledge graph according to the preset number of target objects, wherein the target objects are entity objects corresponding to each node in the preset knowledge graph; Edges between each node are generated based on the relationship data between each target object, and the preset knowledge graph is constructed based on each node and the edges between each node. The relationship data includes the relationship closeness, relationship depth and relationship type between each target object.
4. The method for processing associated data based on knowledge graph according to claim 1, wherein: Determining, according to the second attribute information, a plurality of associated nodes associated with the target node in the preset knowledge graph, including: Obtaining the mapping relationship between the relationship data between the target node and each associated node and the edges of the preset knowledge graph; All nodes of the preset knowledge graph are traversed according to the mapping relationship to obtain multiple associated nodes associated with the target node.
5. The method for processing associated data based on a knowledge graph according to claim 1 or 3, wherein: The calculating the association weight between the target node and each associated node according to the second attribute information includes: According to the relationship data between the target node and each associated node, the relationship closeness, relationship depth and relationship type between the target node and each associated node are obtained from the preset knowledge graph; The relationship closeness, the relationship depth and the relationship type are substituted into a preset association risk factor formula for calculation to obtain the association weight between the target node and each associated node.
6. The method for processing associated data based on knowledge graph according to claim 5, characterized in that: Substituting the relationship closeness, the relationship depth, and the relationship type into a preset association risk factor formula for calculation to obtain an association weight between the target node and each associated node includes: Multiplying the relationship closeness, the relationship depth, and the relationship type between the target node and any associated node to obtain an association weight of the any associated node to the target node; Repeat the step of calculating the association weight of any associated node to the target node to obtain the association weight between the target node and each associated node.
7. The method for processing associated data based on knowledge graph according to claim 1, wherein: The determining, based on the second information score of each associated node, the corresponding node state of each associated node in the preset knowledge graph includes: Obtaining a predefined classification standard for node status, wherein the classification standard is used to classify node status into three types, and each type of node status has a corresponding different threshold; The first information score value of each associated node is compared with a set threshold value of the classification standard, and the node state of the corresponding associated node is defined according to the comparison result.
8. A device for processing associated data based on knowledge graph, characterized in that: The device comprises: A first calculation module is configured to, upon receiving a user's request to calculate associated data for a target node in a preset knowledge graph, obtain first attribute information corresponding to the target node from a preset database, and calculate a first information score value of the target node based on the first attribute information, where the first attribute information is basic data associated with the target node; a second calculation module, configured to obtain second attribute information corresponding to the target node from the preset database, determine a plurality of associated nodes associated with the target node in the preset knowledge graph based on the second attribute information, and calculate an association weight between the target node and each associated node based on the second attribute information, where the second attribute information is relationship data between the target node and each associated node; a determination module, configured to obtain third attribute information corresponding to each associated node from the preset database, calculate a second information score value corresponding to each associated node and a score level of each associated node based on the third attribute information of each associated node, and determine a node state corresponding to each associated node in the preset knowledge graph based on the second information score value corresponding to each associated node, wherein the node state includes a susceptible node, an infected node, and an immune node, and the third attribute information is basic data associated with each associated node; A generation module is used to obtain a correlation data calculation differential equation constructed in advance based on the susceptible node, the infected node and the immune node, substitute the correlation weight and the scoring level of each correlation node into the correlation data calculation differential equation for calculation, obtain the third information scoring value of the target node, generate a correlation data mining report of the target node based on the first information scoring value and the third information scoring value, and feed it back to the user.
9. An electronic device, characterized in that: The electronic device comprises: at least one processor; and, a memory communicatively connected to the at least one processor; wherein, The memory stores a knowledge graph-based associated data processing program that can be executed by the at least one processor, and the knowledge graph-based associated data processing program is executed by the at least one processor so that the at least one processor can execute the knowledge graph-based associated data processing method as described in any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that The computer-readable storage medium stores an associated data processing program based on a knowledge graph, and the associated data processing program based on a knowledge graph can be executed by one or more processors to implement the associated data processing method based on a knowledge graph as described in any one of claims 1 to 7.
Citation Information
Patent Citations
Optimal control method for malicious software propagation in smart ship heterogeneous network
CN114945175A
Telecom social network analysis driven fraud prediction and credit scoring
US20140129420A1