A data-driven method, device, and medium for improving event knowledge graphs
The data-driven method enhances event knowledge graphs by integrating numerical data through Markov boundary discovery and causal strength analysis, addressing errors and sparsity, and improving their accuracy and usability for event reasoning and decision support.
Patent Information
- Application Number
- CN202510330570.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-20
- Publication Date
- 2025-07-15
- Estimated Expiration
- 2045-03-20
AI Technical Summary
It is difficult for the prior art to effectively integrate numerical data with event knowledge graphs to achieve automatic verification, error detection and relationship completion, resulting in insufficient accuracy and completeness of event knowledge graphs, limiting their application in complex event scenarios.
The Markov boundary discovery algorithm is used to learn the direct causal relationship between events from numerical data, combined with conditional independence test and causal intensity calculation, and automatically verifies and completes the event triplets through the correlation test method, and integrates them into the event knowledge graph.
It improves the accuracy and completeness of the event knowledge graph, reduces dependence on manual verification, reduces operating costs, and shows strong robustness in sparse and high noise scenarios, supporting applications such as event reasoning and risk prediction.
Smart Images

Figure CN119849615B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of data mining and knowledge engineering, and specifically relates to an improved method, device, and medium for an event knowledge graph based on data driving. Background Art
[0002] With the rapid development of big data and artificial intelligence technologies, event knowledge graphs have become important infrastructures in fields such as information retrieval, risk prediction, and emergency response. Event knowledge graphs express event entities and their temporal and causal relationships in a structured manner, helping computers better understand and reason about complex event logics. However, in the process of constructing and applying event knowledge graphs, problems such as incorrect event knowledge, missing relationships, and data sparsity are often faced. These problems directly affect the accuracy, integrity, and usability of event knowledge graphs, thus limiting their performance in actual application scenarios. Therefore, how to effectively improve the quality of event knowledge graphs has become an important research direction.
[0003] Existing methods for improving event knowledge graphs mainly include methods based on statistical features, methods based on machine learning embeddings, and methods relying on manual verification. Statistical feature methods identify incorrect relationships and complement missing events by analyzing temporal paths, node types, and event frequency features in the event graph. However, such methods have limited effectiveness in sparse data and high-noise scenarios. Machine learning-based embedding methods attempt to map event knowledge graphs to vector spaces and quantitatively evaluate the relationships between event entities through scoring functions. However, such methods experience performance degradation in sparse data, and the models often lack interpretability. Although manual verification can better ensure the accuracy of event relationships, it requires the participation of a large number of domain experts, which is not only costly but also difficult to scale up and promote.
[0004] In the prior art, for example, patent (CN202411249513.6) verifies the quality of news by calculating the scores of title text and associated text; patent publication (CN119089995A) proposes a knowledge graph embedding method and system based on a multi-relational knowledge enhanced graph convolutional network, which can efficiently construct knowledge graph embeddings to serve as representations for evaluating the relationships between entities; patent publication (CN119150966A) proposes a quality assessment method for a vertical domain knowledge graph and designs an overall idea for verifying the quality of a knowledge graph in a specific domain. These methods all belong to the three types of traditional knowledge graph improvement methods described above and face many challenges such as high operating costs, noisy data, and sparse features.
[0005] To solve the above problems, data-driven event knowledge graph improvement methods have gradually attracted attention. Numerical data, as a quantitative record of real-world events, is widely used in emergency management, medical safety, financial risks and other fields, and contains a large amount of implicit causal relationships and statistical correlation information. Data-driven methods attempt to perform causal analysis and correlation verification through numerical data to detect erroneous event knowledge and complete missing event relationships. This method not only reduces the dependence on the internal structure of the event knowledge graph, but also shows strong adaptability in sparse and high-noise scenarios. However, there are significant differences in the representation of numerical data and event knowledge graphs - numerical data is usually presented in the form of a time series matrix, while the event knowledge graph consists of discrete event-relationship-event triplets. This difference increases the difficulty of expression fusion and joint reasoning between the two.
[0006] Therefore, how to effectively integrate numerical data with event knowledge graphs to achieve automatic verification, error detection and relationship completion of event knowledge has become a key issue that needs to be solved urgently. Summary of the invention
[0007] The present invention provides a data-driven event knowledge graph improvement method, device and medium, the purpose of which is to improve the accuracy, completeness and availability of the event knowledge graph, so as to better support application scenarios such as event reasoning, risk prediction and decision assistance.
[0008] To achieve the above object, the first aspect of the present invention provides a data-driven event knowledge graph improvement method, comprising the following steps:
[0009] Collecting numerical data related to the event knowledge graph and preprocessing the numerical data;
[0010] Using the Markov boundary discovery algorithm, we learn the direct causal relationship between events from the preprocessed numerical data and construct a set of event causal relationships;
[0011] The event pairs in the event causal relationship set are screened by the conditional independence test method, and the events with direct causal relationship with the target event are retained;
[0012] The causal strength calculation formula is used to quantitatively evaluate the relationship between each pair of screened events, and the causal strength value between each pair of events is calculated;
[0013] Preliminary screening of event triplets in the event knowledge graph is performed based on the causal strength value, and event triplets whose causal strength values do not reach the preset threshold are added to the set to be verified;
[0014] For the event triples in the set to be verified, calculate the significance level based on the correlation test method. If the significance level meets the preset threshold condition, retain the event triple; otherwise, delete it from the event knowledge graph.
[0015] Based on the calculation of causal strength, complete the event pairs that have a significant causal relationship in the preprocessed numerical data but are missing in the event knowledge graph, and generate the completed event triples.
[0016] Conduct two-way causal strength verification on the completed event triples to make the completed relationships bidirectionally valid in the event knowledge graph.
[0017] Integrate the verified and completed event triples into the original event knowledge graph to update the event knowledge graph.
[0018] Furthermore, collect numerical data related to the event knowledge graph, including the following steps:
[0019] Collect event-related numerical data from different data sources. The data sources include sensor records, time series data, log files, statistical databases, or experimental observation results. The collected numerical data includes event attributes, event occurrence times, event frequencies, and the correlations between events.
[0020] Integrate the collected event-related numerical data into a data matrix, where each column represents an event attribute and each row represents an event record, to obtain the numerical data related to the event knowledge graph.
[0021] Furthermore, the preprocessing of the numerical data includes: performing standardization, normalization, and noise reduction on the collected numerical data.
[0022] Furthermore, the standardization and normalization processing are carried out through the maximum-minimum normalization formula, and the formula is:
[0023]
[0024] Where, is the normalized event attribute value, is the original event attribute value, and are the maximum and minimum values of the original event attribute respectively.
[0025] Furthermore, when the conditional independence test meets the following conditions, the event pair is considered to have a direct causal relationship:
[0026]
[0027] Where, For the probability of the simultaneous occurrence of event and the directly causally related event under a given set of conditioning variables ; For the probability of the occurrence of event under a given set of conditioning variables ; For the probability of the occurrence of the directly causally related event under a given set of conditioning variables ; is the set of conditioning variables used in the conditional independence test, is the event, is the directly causally related event of event ; is the attribute or variable value of event ; is the attribute or variable value of the directly causally related event ;
[0028] Furthermore, the causal strength value between events is quantified using a causal strength function, where:
[0029]
[0030] where is the causal strength value between event and the directly causally related event ; represents the degree of influence of the directly causally related event on event ; represents the degree of influence of event on the directly causally related event ;
[0031] Furthermore, the correlation test uses a significance level calculation formula:
[0032]
[0033] where is the statistic for the correlation significance test between event and the directly causally related event ; is a variable value in event ; is a variable value in the directly causally related event ; is the set of all variable values in event ; is the directly causally related event The set of all variable values in is the observed value, and
[0034]
[0034]
[0035]
[0036] where represents an event and there is a direct causal relationship with the directly causally related event is the set of Markov boundaries, is the causal strength threshold, which determines whether to complete this event triple.
[0037] To achieve the above object, a second aspect of the present invention provides an electronic device, including a processor and a memory. When the processor executes the computer program stored in the memory, the steps of the data-driven event knowledge graph improvement method are implemented.
[0038] To achieve the above object, a third aspect of the present invention provides a computer-readable storage medium, on which a computer program is stored. When the computer program is run by a processor, the steps of the data-driven event knowledge graph improvement method are executed.
[0039] Advantages of the present invention:
[0040] Compared with the prior art, a data-driven event knowledge graph improvement method, device, and medium provided by the present invention verify, detect errors, and complete missing parts of the event knowledge graph in an automated manner by using causal relationships and statistical correlation information in numerical data. Specifically, the method uses the Markov boundary discovery algorithm to extract direct causal relationships between events from numerical data, and at the same time identifies other events directly related to the target event through conditional independence testing and causal strength calculation to eliminate incorrect knowledge and complete missing relationships. In addition, correlation analysis is also combined to ensure that event relationships with insufficient causality but significant correlation can be effectively identified and retained. By introducing the deep integration of causal and correlation information, this method significantly reduces the dependence on manual verification, reduces operating costs, and improves the robustness of the system in complex event scenarios. BRIEF DESCRIPTION OF THE DRAWINGS
[0041] To more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments.
[0042] Figure 1It is a flowchart of an improved method for event knowledge graph based on data driving disclosed in an embodiment of the present invention.
[0043] Figure 2 It is a flowchart of the main steps of an improved task of event knowledge graph disclosed in an embodiment of the present invention. Detailed implementation manners
[0044] In order to enable those skilled in the art to better understand the solution of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0045] According to the embodiments of the present invention, it should be noted that the steps shown in the flowchart of the drawings can be executed in a computer system such as a set of computer executable instructions, and although the logical order is shown in the following method, in some cases, the steps shown or described can be executed in a different order than here.
[0046] As Figure 1 、 Figure 2 shown, the present invention provides an improved method for event knowledge graph based on data driving, including the following steps:
[0047] Step S100: Collect numerical data related to the event knowledge graph and preprocess the numerical data;
[0048] First, it is necessary to collect numerical data related to the event knowledge graph. These data can come from different fields and data sources, such as sensor records, time series data, log files, statistical databases, or experimental observation results. Specifically, the collected numerical data should include information such as the timestamp of the event occurrence, event attributes (such as event type, loss degree, etc.), the frequency of event occurrence, and the correlation between events. These data are stored in the form of a data matrix (Data Matrix, DM), where each column represents an event attribute and each row represents an event record, obtaining the numerical data related to the event knowledge graph.
[0049] Next, preprocess the collected numerical data to ensure that the data quality is suitable for subsequent causal analysis and correlation verification. The steps of preprocessing include data standardization, normalization, and noise reduction. The purpose of standardization is to eliminate the influence of different dimensions and ensure that each event attribute is compared on the same scale; normalization is to map the numerical values of each event attribute to a unified range (such as the interval [0,1]), so that data of different attributes can be processed under the same standard. Noise reduction mainly removes irrelevant interferences and outliers in the data, which can be carried out through methods such as median filtering, smoothing, or outlier detection, to ensure that the data quality meets the requirements of subsequent causal analysis and correlation verification. In addition, for time series data, it is necessary to divide time windows, and statistically aggregate the time series data according to appropriate time granularities to better reveal the causal relationships between events. During this process, if there are missing values in the data, they should also be filled by reasonable methods to ensure the integrity and effectiveness of the data.
[0050] In this embodiment, for the event data matrix , each event attribute column is subjected to maximum-minimum normalization, and the formula is as follows:
[0051]
[0052] where, is the normalized event attribute value, is the original event attribute value, and are the maximum and minimum values of the original event attribute respectively.
[0053] Step S200: Use the Markov boundary discovery algorithm to learn the direct causal relationships between events from the preprocessed numerical data and construct an event causal relationship set;
[0054] Regard each event in the event data matrix as a node, and construct an event causal relationship network by analyzing event attributes and their correlations. In this network, edges represent causal relationships between events, and the weights of the edges represent the strength of the causal relationships. By using the Markov boundary discovery algorithm, a set of other events directly related to the target event (such as a specific event) can be automatically identified from the data, excluding those events that have no direct causal relationship with the target event. Based on the principle of conditional independence, this algorithm gradually eliminates redundant and indirect event relationships, and finally obtains an event causal relationship set that only contains direct causal relationships, ensuring that the extracted causal structure has high accuracy and representativeness.
[0055] Step S300: screening event pairs in the event causal relationship set by a conditional independence test method, and retaining events that have a direct causal relationship with the target event;
[0056] For each pair of events and directly causally related events , the conditional independence test formula is used to determine whether there is a direct causal relationship between the two:
[0057]
[0058] in, For a given set of condition variables Under the conditions, the event and directly causally related events The probability of simultaneous occurrence, For a given set of condition variables Under the conditions, the event occurs The probability of For a given set of condition variables Under the condition of direct causal relationship The probability of occurrence, is the set of conditional variables used in the conditional independence test, For events, It is an event directly causally related events, For events , Direct causal related events The value of a property or variable.
[0059] If the test results show that the causal relationship between the two events is direct, that is, the conditional independence test is not established, then the event pair is retained; if the causal relationship is not direct, it is eliminated. Through this screening step, those indirect or irrelevant event relationships can be effectively removed, ensuring that the causal relationship that is finally retained has a high degree of relevance and practical significance.
[0060] Step S400: using a causal strength calculation formula to quantitatively evaluate each pair of event relationships screened out, and calculating the causal strength value between each pair of events;
[0061] By calculating each pair of events and directly causally related events The causal strength value between:
[0062]
[0063] in, For events and the causal strength value between directly causally related events , indicating the degree of influence of the directly causally related event on the event ; indicating the degree of influence of the event on the directly causally related event .
[0064] This causal strength function can reflect the strength of the causal relationship between events and help quantify the direct causal impact between each pair of events. Through this calculation step, a causal strength value can be assigned to each pair of event relationships, thereby clarifying which pairs of events have a strong causal relationship and which may be indirect or weak relationships. According to the set causal strength threshold , pairs of events with a causal strength lower than the threshold are excluded, and only event relationships with significant causal strength are retained to ensure higher accuracy and effectiveness of the final causal relationship network.
[0065] Step S500: Preliminarily screen the event triples in the event knowledge graph according to the causal strength value, and add the event triples whose causal strength value does not reach the preset threshold to the set to be verified;
[0066] Retrieve each event triple in the event knowledge graph , and check whether this event relationship is verified in the causal strength calculation. For the event triple , if its event pair is not within the Markov boundary (MB) of the target event, that is:
[0067]
[0068] then this event relationship may be invalid and may need to be deleted from the event knowledge graph, and it is added to the set to be verified.
[0069] It should be noted that the event pair consists of two events that form a triple (the causal relationship has not been confirmed), and the event triple is a semantic unit composed of two events and their causal relationship edge. Deleting an event triple is equivalent to deleting the causal relationship between events (retaining the event nodes and only removing the relationship edge).
[0070] Step S600: For the event triples in the set to be verified, calculate the significance level based on the correlation test method. If the significance level meets the preset threshold condition, then retain this event triple, otherwise delete it from the event knowledge graph;
[0071] Identify those event pairs with insufficient causal strength but statistically significant correlation. Then, use the correlation test method to calculate the event And the direct causal-related events The significance level of the correlation , is evaluated through the statistic , and the formula is:
[0072]
[0073] Wherein, is the statistic of the significance test of the correlation between event and the direct causal-related event , is a variable value in event , is a variable value in the direct causal-related event , is the set of all variable values in event , is the set of all variable values in the direct causal-related event , is the observed value, is the expected value.
[0074] If this significance level meets the preset threshold , then retain this pair of events. Through this step, it is possible to verify those event pairs that may lack direct causal relationships but still have significant statistical correlations, thereby ensuring the integrity of the event knowledge graph, avoiding missing potential important event relationships, and achieving dynamic optimization of the event knowledge graph under the dual verification of causality and correlation.
[0075] Step S700, based on the calculation of causal strength, complete the event pairs with significant causal relationships that are missing in the preprocessed numerical data, and generate the completed event triples;
[0076] Through the Markov boundary discovery algorithm, identify those event pairs with significant causal relationships in the data matrix , but these event pairs have not yet appeared in the existing event knowledge graph. For each pair of event pairs that meet the completion conditions, if this event pair has a significant causal relationship, that is: and the causal strength , then use this event pair as a new triple ( , causes, ) and supplement it to the event knowledge graph. This process can effectively fill in the key event relationships missed due to data scarcity or noise, thereby improving the integrity and accuracy of the event knowledge graph. Among the conditions for completing event triples, represents event and the direct causal-related event There is a direct causal relationship, is the Markov boundary set, is the causal intensity threshold, which determines whether to complete this triple.
[0077] Step S800: Perform two-way causal intensity verification on the completed event triples to make the completion relationship bidirectionally valid in the event knowledge graph;
[0078] For each pair of completed event triples ( , causes, ), first calculate the causal intensity between event and the directly causally related event and , and respectively evaluate the strength of the causal relationship between the two. If both of these two causal intensity values are greater than the preset causal intensity threshold , then it is considered that the event relationship is valid in both directions, and the causal relationship of this event triple is confirmed to be stable and has reliable two-way support. If the causal intensity value is insufficient, then it is considered that the causal relationship of this event triple does not hold, and further adjustment or deletion is required. This step ensures that the completed event relationship is not only valid from the perspective of unidirectional causality, but also has a solid foundation in terms of two-way causal support, improving the reliability and consistency of the event knowledge graph.
[0079] Step S900: Integrate the verified and completed event triples into the original event knowledge graph to update the event knowledge graph.
[0080] Add all the valid triples that have passed causal verification and completion to the event knowledge graph to ensure that the causal relationship and statistical correlation of each event triple are fully verified and supplemented. Then, store the updated event knowledge graph for subsequent query, analysis, and application. The updated knowledge graph can be applied to multiple practical scenarios, such as event prediction, risk analysis, fault diagnosis, etc., providing more accurate and comprehensive decision support. In addition, during the storage process, the event knowledge graph can also be interfaced with other systems to achieve data sharing and integration, providing strong support for tasks such as event response and risk assessment in other systems.
[0081] In this embodiment, for different event knowledge graphs and data scenarios, optimize the causal intensity threshold and the correlation significance threshold , and perform performance evaluation to ensure that the improvement effect reaches the expected goal.
[0082] In the above method, the present invention has a number of significant advantages and positive effects. First, by combining causal analysis and correlation verification, it can efficiently identify error event knowledge and complete missing event relationships, thereby significantly improving the accuracy and integrity of the event knowledge graph. Second, compared with traditional methods that rely on manual verification by experts, the present invention reduces the cost and time overhead of manual intervention through a data-driven automated verification and completion process. In addition, the present invention innovatively integrates numerical data with the event knowledge graph deeply, and realizes two-way support and optimization of data and knowledge by using the causal and correlative information of the data. This method shows strong robustness in dealing with sparse knowledge and high-noise data, and can adapt to the improvement requirements of event knowledge graphs in different scenarios. Finally, the improved event knowledge graph has broad application prospects and can be used in multiple fields such as event prediction, risk analysis, and fault diagnosis, providing reliable support for decision-making in complex event scenarios.
[0083] To further illustrate the above solution, a specific embodiment of an event knowledge graph improvement method is given below (for the convenience of description, example parameters are substituted, but the specific implementation is not limited to the specific parameters of this example):
[0084] I. Data collection and preprocessing
[0085] Step 1.1: Event data collection
[0086] Collect numerical data related to events from a city emergency management system. The data includes attributes such as accident type, occurrence time, location, loss degree, response time, etc. The data sources include sensor logs, historical accident record databases, and real-time monitoring systems. These data are integrated into a data matrix, where each column represents an event attribute and each row represents an event record.
[0087] Step 1.2: Event data preprocessing
[0088] Perform standardization and normalization on the data matrix. For example, for the event attribute matrix , each attribute column is subjected to max-min normalization:
[0089]
[0090] where and are the maximum and minimum values of the original event attribute , respectively. In addition, the time series data is aggregated according to a 24-hour time window, and missing values are filled to ensure that the data quality meets the requirements of subsequent analysis.
[0091] II. Markov boundary discovery
[0092] Step 2.1: Event Causal Structure Learning
[0093] Use the Markov boundary discovery algorithm to extract the direct causal relationships between different events. For example, identify the direct causal relationship between the "heavy rain event" and the "road collapse".
[0094] Step 2.2: Conditional Independence Test
[0095] For each pair of event relationships, conduct a conditional independence test to determine whether there is a direct causal connection between the two:
[0096]
[0097] If the two satisfy the conditional independence test, it is considered that there is no direct causal connection.
[0098] Step 2.3: Event Causal Strength Calculation
[0099] Quantify the strength of the causal relationship between the "heavy rain event" and the "road collapse" and calculate its causal strength:
[0100]
[0101] When the causal strength exceeds the preset threshold , it is considered that there is a significant causal relationship between the two.
[0102] III. Event Knowledge Error Detection
[0103] Step 3.1: Verification of Error Event Knowledge
[0104] If there is a triple (heavy rain, causes, power outage) in the existing event knowledge graph, but data analysis shows that there is no causal connection between the two, then add this triple to the set to be verified:
[0105]
[0106] Step 3.2: Supplementary Verification of Event Relevance
[0107] For some event pairs with insufficient causality but significant relevance, such as "heavy rain" and "traffic congestion", conduct a significance test of relevance:
[0108]
[0109] If the significance level meets the requirements, then retain this triple.
[0110] IV. Completion of Missing Event Knowledge
[0111] Step 4.1: Completion of Event Causal Knowledge
[0112] Through Markov boundary discovery, a missing event relationship is identified:
[0113] (Heavy rain, causes, road waterlogging)
[0114] According to the calculation of causal strength, this relationship meets the completion condition:
[0115]
[0116] Therefore, this triple is supplemented into the event knowledge graph.
[0117] Step 4.2: Bidirectional event causality verification
[0118] Perform bidirectional verification on the newly added triple to ensure that the causal relationship between events is stable and has two-way support.
[0119] V. Parameter optimization
[0120] Step 5.1: Causal threshold optimization
[0121] Select better parameters from the empirical values, and finally set the causal strength threshold and the correlation significance threshold , and obtain the optimal improvement effect in this scenario.
[0122] Step 5.2: Performance evaluation
[0123] Use metrics such as Precision, Recall, and F1-Score to evaluate the improved event knowledge graph, and the results are as follows:
[0124] Precision: 90.5%, Recall: 87.3%, F1-Score: 88.8%.
[0125] VI. Event knowledge graph output
[0126] Step 6.1: Event knowledge graph update
[0127] Merge the verified and completed event triples into the original event knowledge graph to generate a new event knowledge graph.
[0128] Step 6.2: Event knowledge graph storage
[0129] Store the improved event knowledge graph in the event management database for other systems to call and query.
[0130] Step 6.3: Event knowledge graph application
[0131] In the urban risk monitoring platform, an improved event knowledge graph is used for real-time event prediction, fault diagnosis, and emergency response.
[0132] According to another aspect of the embodiments of the present application, an electronic device is further provided, including a processor and a memory. When the processor executes the computer program stored in the memory, the steps of the method are implemented.
[0133] In the above embodiments of the present invention, the descriptions of the respective embodiments have their own emphases. For the parts not detailed in a certain embodiment, reference may be made to the relevant descriptions of other embodiments.
[0134] In several embodiments provided by the present application, it should be understood that the disclosed technical content can be implemented in other ways. Among them, the device embodiments described above are only illustrative. For example, the division of the units can be a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication connection to each other can be through some interfaces. The indirect coupling or communication connection of units or modules can be in an electrical or other form.
[0135] In addition, each functional unit in the various embodiments of the present invention can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The above-mentioned integrated unit can be implemented in the form of hardware or in the form of a software functional unit.
[0136] If the above-mentioned integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or all or part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention. The aforementioned storage medium includes: USB flash drives, read-only memories (ROMs), random access memories (RAMs), mobile hard disks, magnetic disks, or optical discs, etc., which can store program codes.
[0137] The above are only the preferred embodiments of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the principle of the present invention, several improvements and refinements can be made, and these improvements and refinements should also be regarded as the protection scope of the present invention.
Claims
1. An improved method for event knowledge graph based on data-driven, characterized in that, It includes the following steps: Collect numerical data related to the event knowledge graph. The numerical data is event-related numerical data collected from a city emergency management system, including accident type, occurrence time, location, loss degree, and response time. The sources of the numerical data include sensor logs, historical accident record databases, and real-time monitoring systems, and preprocess the numerical data; Use the Markov boundary discovery algorithm to learn the direct causal relationships between events from the preprocessed numerical data and construct an event causal relationship set; Screen the event pairs in the event causal relationship set through the conditional independence test method, and retain the events that have a direct causal relationship with the target event; Adopt a causal strength calculation formula to quantitatively evaluate each pair of screened event relationships, and calculate the causal strength value between each pair of events; Preliminarily screen the event triples in the event knowledge graph according to the causal strength value, and add the event triples whose causal strength value does not reach the preset threshold to the set to be verified; For the event triples in the set to be verified, calculate the significance level based on the correlation test method. If the significance level meets the preset threshold condition, retain the event triple, otherwise delete it from the event knowledge graph; On the basis of the causal strength calculation, complete the event pairs that have significant causal relationships but are missing in the preprocessed numerical data in the event knowledge graph, and generate completed event triples; Conduct two-way causal strength verification on the completed event triples to make the completed relationships bidirectionally valid in the event knowledge graph; Integrate the verified and completed event triples into the original event knowledge graph, update the event knowledge graph, and use the improved event knowledge graph in the city risk monitoring platform for real-time event prediction, fault diagnosis, and emergency response; The causal strength value between events is quantified using a causal strength function, where: Among them, is the event and the direct causal related event between the causal strength value, indicating the direct causal related event on the event the degree of influence, indicating the event on the direct causal related event the degree of influence; The conditions for completing event triples are: Among them, represents an event and a directly causally related event have a direct causal relationship, is a Markov boundary set, is a causal intensity threshold, which determines whether to complete this event triple.
2. The method for improving an event knowledge graph based on data driving according to claim 1, characterized in that Collect numerical data related to the event knowledge graph, including the following steps: The collected numerical data includes event attributes, event occurrence time, event frequency, and the relevance between events; Integrate the collected event-related numerical data into a data matrix, where each column represents an event attribute and each row represents an event record, to obtain numerical data related to the event knowledge graph.
3. The method for improving an event knowledge graph based on data driving according to claim 2, wherein The preprocessing of the numerical data includes: standardizing, normalizing, and denoising the collected numerical data.
4. The method for improving an event knowledge graph based on data driving according to claim 3, wherein The normalization process is carried out through the maximum-minimum normalization formula, and the formula is: Among them, is the normalized event attribute value, is the original event attribute value, and are respectively the maximum value and the minimum value of the original event attribute .
5. The method for improving an event knowledge graph based on data driving according to claim 1, wherein When the following conditions are met in the conditional independence test, the event pair is considered to have a direct causal relationship: wherein, is the probability that event and the directly causally related event occur simultaneously under the condition of a given set of conditional variables , is the probability that event occurs under the condition of a given set of conditional variables , is the probability that the directly causally related event occurs under the condition of a given set of conditional variables , is the set of conditional variables used in the conditional independence test, is an event, is the directly causally related event of event , is the attribute or variable value of event , is the attribute or variable value of the directly causally related event .
6. The method for improving an event knowledge graph based on data driving according to claim 1, wherein, The correlation test adopts the significance level calculation formula: Among them, is the event and the directly causally related event The statistic for the significance test of the correlation between them, is a variable value in the event one of is a variable value in the directly causally related event one of is the event the set of all variable values in is the directly causally related event the set of all variable values in is the observed value, is the expected value.
7. An electronic device, characterized in that, It includes a processor and a memory. When the processor executes the computer program stored in the memory, it implements the steps of the data-driven event knowledge graph improvement method according to any one of claims 1 to 6.
8. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is run by the processor, it executes the steps of the data-driven event knowledge graph improvement method according to any one of claims 1 to 6.
Citation Information
Patent Citations
Multimodal false news detection method, device, equipment and medium
CN118762377B
Knowledge graph embedding method and system based on multi-relation knowledge enhancement graph convolutional network
CN119089995A
Business model driven vertical domain knowledge graph quality evaluation method
CN119150966A
Fusion map construction method and device
CN113157931A
Knowledge graph correction method based on data correlation and causal mode
CN117709449A