Fault root cause positioning method and system driven by dynamic knowledge graph
By collecting multi-source fault data to construct a dynamic knowledge graph, and combining entity association and temporal features, pattern matching and similarity analysis are performed, which solves the problem of inaccurate root cause localization in traditional fault diagnosis and improves the accuracy and efficiency of fault diagnosis.
Patent Information
- Application Number
- CN202511041863.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-28
- Publication Date
- 2025-11-14
AI Technical Summary
Traditional root cause localization methods rely on human experience, resulting in inaccurate localization and low efficiency. The knowledge graph is also slow to update and cannot reflect the latest fault situation in a timely manner, which affects the system stability and operation and maintenance efficiency.
Collect multi-source fault association datasets, extract entity associations and identify relationships, construct a dynamic knowledge graph by combining entity temporal features, determine the root cause of the fault through pattern matching reasoning and similarity matching, and conduct controlled injection testing and iterative feedback optimization.
It enables precise location of the root cause of the fault and dynamic improvement of the knowledge graph, thereby enhancing the accuracy and efficiency of fault diagnosis.
Smart Images

Figure CN120950284A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of fault root cause localization technology, specifically to a fault root cause localization method and system driven by dynamic knowledge graph. Background Technology
[0002] In complex information systems, faults often exhibit characteristics such as concealment, correlation, and dynamism. Traditional root cause analysis methods rely heavily on manual experience or simple rule matching, resulting in inaccurate localization, low efficiency, and difficulty in quickly addressing large-scale and diverse fault scenarios. While existing knowledge graph technologies can integrate fault knowledge to some extent, the lack of effective processing of multi-source data and dynamic update mechanisms means that knowledge graphs cannot reflect the latest fault conditions in a timely manner. Consequently, in actual fault diagnosis, it is impossible to accurately locate the root cause of faults, severely impacting system stability and operational efficiency.
[0003] Existing technologies suffer from inaccurate root cause localization and lagging knowledge graph updates, resulting in low efficiency in fault diagnosis. Summary of the Invention
[0004] This application provides a dynamic knowledge graph-driven fault root cause localization method and system to address the technical problems of inaccurate fault root cause localization, lagging knowledge graph updates, and low fault diagnosis efficiency in the prior art.
[0005] In view of the above problems, this application provides a method and system for fault root cause localization driven by dynamic knowledge graph.
[0006] The first aspect of this application provides a dynamic knowledge graph-driven method for fault root cause localization, the method comprising:
[0007] A multi-source fault association dataset is collected, and entity associations are extracted from the dataset to obtain a set of fault entities and a set of entity relationships. The fault entity set and entity relationship set are then updated by concatenating graph nodes and incremental learning, combining entity temporal features, to construct a fault update knowledge graph. Target fault data is monitored and acquired, and the fault update knowledge graph is used to perform pattern matching inference on the target fault data to generate a set of candidate root causes of fault patterns. Similarity matching is performed on the candidate root causes of fault patterns using a historical fault case database to determine the target fault root cause location result. Based on the target fault root cause location result, the fault update knowledge graph is subjected to controlled injection testing and iterative feedback optimization.
[0008] A second aspect of this application provides a dynamic knowledge graph-driven fault root cause localization system, the system comprising:
[0009] The system comprises the following modules: an entity association extraction module, which collects a multi-source fault association dataset and extracts entity associations from it to obtain a set of fault entities and a set of entity relationships; a fault update knowledge graph construction module, which combines entity temporal features to perform graph node concatenation and incremental learning updates on the fault entity set and entity relationship set to construct a fault update knowledge graph; a pattern matching and reasoning module, which monitors and acquires target fault data and uses the fault update knowledge graph to perform pattern matching and reasoning on the target fault data to generate a set of candidate root causes of fault patterns; and a similarity matching module, which combines a historical fault case library to perform similarity matching on the set of candidate root causes of fault patterns, determines the target fault root cause location result, and performs controlled injection testing and iterative feedback optimization on the fault update knowledge graph based on the target fault root cause location result.
[0010] One or more technical solutions provided in this application have at least the following technical effects or advantages:
[0011] A multi-source fault association dataset is collected, and entity associations are extracted from the dataset to obtain a set of fault entities and a set of entity relationships. Graph node concatenation and incremental learning updates are performed to construct a fault update knowledge graph. Target fault data is monitored and acquired, and pattern matching inference is performed to generate a set of candidate root causes of fault patterns. Similarity matching is performed on the candidate root causes of fault patterns using a historical fault case database to determine the target fault root cause location. Based on the target fault root cause location result, the fault update knowledge graph is subjected to controlled injection testing and iterative feedback optimization. This achieves the technical effect of accurately locating fault root causes and dynamically improving the knowledge graph, thereby enhancing the efficiency and accuracy of fault diagnosis. Attached Figure Description
[0012] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0013] Figure 1 A schematic flowchart of a dynamic knowledge graph-driven fault root cause localization method provided in an embodiment of this application;
[0014] Figure 2 A schematic diagram of the structure of a dynamic knowledge graph-driven fault root cause localization system provided in an embodiment of this application.
[0015] Figure labeling: Entity association extraction module 10, fault update knowledge graph construction module 20, pattern matching reasoning module 30, similarity matching module 40. Detailed Implementation
[0016] This application provides a dynamic knowledge graph-driven fault root cause localization method and system to address the technical problems of inaccurate fault root cause localization, lagging knowledge graph updates, and low fault diagnosis efficiency in the prior art.
[0017] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of this application, and not all of them. All other embodiments obtained by those skilled in the art based on the embodiments of this application without creative effort are within the scope of protection of this application.
[0018] Example 1, as Figure 1 As shown, this application provides a dynamic knowledge graph-driven method for fault root cause localization, the method comprising:
[0019] Step S100: Collect and obtain a multi-source fault association dataset, and extract entity associations from the multi-source fault association dataset to obtain a fault entity set and an entity relationship set.
[0020] Specifically, the process involves collecting multi-source fault association datasets, including system logs, performance metrics, and topology relationships. First, based on the attribute type information of the dataset, the criteria for judging multi-source noise data are determined. A cleaning program is then set up to identify and clean the dataset for noise, resulting in a usable multi-source fault association dataset. Next, historical multi-source fault datasets are mined, and fault entities are defined based on domain knowledge and experience to obtain recognition rules. After identifying entity data using these rules, a domain knowledge entity recognizer is trained and optimized using a bidirectional long short-term memory network. This recognizer then performs entity recognition on the usable dataset to obtain a set of fault entities. Finally, semantic relationship recognition and business logic association are performed on each entity in the fault entity set to determine the entity relationship set.
[0021] Step S200: Combine the temporal features of entities to perform graph node concatenation and incremental learning updates on the fault entity set and entity relationship set to construct a fault update knowledge graph.
[0022] Specifically, the process begins by combining the temporal characteristics of entities to perform time-granularity labeling and data alignment integration on the fault entity set and entity relationship set, resulting in a standard fault entity set and a standard entity relationship set. Then, each fault entity in the standard fault entity set is treated as a knowledge graph node and assigned an identifier. Nodes are then categorized and labeled with knowledge attributes. Simultaneously, each entity relationship in the standard entity relationship set is labeled with graph edge attributes. Logical concatenation analysis is performed on the graph edge set and node knowledge attribute set, combined with business logic topology relationships, to generate a graph node business concatenation network. An initial fault knowledge graph is constructed based on this network. Finally, an online streaming data processing module receives multi-source fault update data streams in real time. Fault event identification and entity parsing are performed on the update data streams to obtain newly added entity and relationship sets. The initial graph is then used for conflict detection and knowledge correction. Based on the corrected data, the initial graph is incrementally learned and updated, thereby constructing a fault update knowledge graph.
[0023] Step S300: Monitor and acquire target fault data, and use the fault update knowledge graph to perform pattern matching reasoning on the target fault data to generate a set of candidate root causes of fault modes.
[0024] Specifically, after acquiring target fault data through real-time monitoring, pattern matching and inference analysis are performed using a pre-constructed fault update knowledge graph. This knowledge graph integrates multi-source fault entities, relationships, and time-series features. Based on the cascaded network of graph nodes and edge attribute logic, it can identify entity association patterns in the target fault data, match historical fault pattern rules, and then infer possible fault causes, generating a candidate root cause set of fault patterns containing multiple potential root causes, thus providing a candidate range for subsequent precise localization.
[0025] Step S400: Combine the historical fault case library to perform similarity matching on the candidate root cause set of the fault mode, determine the target fault root cause location result, and perform controllable injection testing and iterative feedback optimization on the fault update knowledge graph based on the target fault root cause location result.
[0026] Specifically, a similarity calculation is performed between the historical fault case database and the target fault data to obtain a fault case similarity set. This set is then used to optimize the historical fault case database, selecting target-fit fault cases with high matching degree with the current fault features. The intersection matching of the candidate root cause set of fault modes is then performed based on the target-fit fault cases to determine the target fault root cause localization result. Subsequently, a controllable injection test environment is constructed, and injection tests are performed on the target fault root cause localization result in this environment to obtain the localization accuracy verification result. Based on the result, the fault update knowledge graph is iteratively updated and optimized to improve the accuracy and adaptability of the knowledge graph for fault root cause localization.
[0027] In one possible implementation, step S100 further includes:
[0028] Step S110: Determine the multi-source noise data discrimination criteria based on the attribute type information of the multi-source fault association dataset, and set up a multi-source noise data cleaning program based on the multi-source noise data discrimination criteria.
[0029] Step S120: Perform noise identification and cleaning on the multi-source fault association dataset according to the multi-source noise data discrimination criteria and the multi-source noise data cleaning procedure to obtain a usable multi-source fault association dataset.
[0030] Step S130: Construct a domain knowledge entity recognizer, and use the domain knowledge entity recognizer to perform entity recognition on the available multi-source fault association dataset to obtain a fault entity set.
[0031] Step S140: Perform semantic relationship identification and business logic association between the entities in the fault entity set to determine the entity relationship set.
[0032] Specifically, based on the attribute type information of the multi-source fault association dataset, such as data source type (system logs, performance indicators, topology relationships, etc.), numerical characteristics (normal range of parameters such as voltage, current, and frequency), timestamp integrity, and data format standardization, the discrimination criteria for multi-source noise data are clarified. For example, data that exceeds the normal numerical range by ±20%, data with missing timestamps or logical conflicts, and data with unreliable sources or incorrect formats are judged as noise data. Then, based on the above discrimination criteria, a multi-source noise data cleaning program is set up. This program covers the identification rules of noise data, the processing flow, and the replacement or removal strategies of abnormal data, so as to achieve systematic processing of noise data.
[0033] Based on established criteria for identifying multi-source noise data (such as values exceeding the normal range by ±20%, timestamp conflicts, and unreliable data sources) and a set cleaning procedure, a systematic noise identification and cleaning process is performed on the multi-source fault association dataset. The program automatically scans each record in the dataset, compares it against the criteria to identify noisy data, performs threshold correction on data exceeding the numerical range, completes or calibrates missing or conflicting timestamp data, and removes data from unreliable sources or with incorrect formats. The result is a usable multi-source fault association dataset that meets quality requirements, providing a reliable data foundation for subsequent entity identification and relation extraction.
[0034] This paper explores a multi-source fault history dataset. Based on domain knowledge and experience, fault entities such as hardware components, software modules, and performance indicators are defined to form multi-source fault entity recognition rules. Then, entity data recognition is performed on the multi-source fault history dataset using these rules, resulting in a multi-source fault entity dataset with entity annotations. Next, a Bidirectional Long Short-Term Memory (Bi-LSTM) network is used to train the dataset with proportional annotations. Through forward and backward propagation, the temporal dependencies of entity features are learned. After testing, verification, and parameter tuning, a domain knowledge entity recognizer with accurate entity recognition capabilities is constructed. Finally, this recognizer is used to extract and classify entities from the obtained usable multi-source fault association dataset, resulting in a fault entity set containing various types of fault entities.
[0035] Natural Language Processing (NLP) techniques are used to semantically parse the entity description text in the fault entity set, extracting relational keywords such as "cause," "depends on," and "associate." These keywords are then combined with a pre-defined domain relation dictionary (e.g., the mapping relationship between hardware faults and performance indicators, and the call chain rules of software modules) to identify semantic relationships between entities. Simultaneously, a business logic knowledge base is constructed based on the system topology diagram and business process documents, mapping entities to nodes in the business logic. By analyzing data flow, control flow dependencies (e.g., the interface dependency of service A calling service B in a microservice architecture), and causal propagation paths (e.g., the three-level association of database crash - application service timeout - front-end page error), business logic associations between entities are established. Finally, through cross-validation of semantic relationships and business logic, ambiguous associations are eliminated, generating an entity relation set containing entity-relationship-entity triples.
[0036] In one possible implementation, step S130 further includes:
[0037] Step S131: Mining and obtaining a multi-source fault history dataset, defining fault entities in the multi-source fault history dataset according to domain knowledge and experience, and obtaining multi-source fault entity recognition rules.
[0038] Step S132: Based on the multi-source fault entity identification rules, perform entity data identification on the multi-source fault historical dataset to obtain a multi-source fault entity dataset.
[0039] Step S133: Use a bidirectional long short-term memory network to perform proportional labeling training on the multi-source fault entity dataset to obtain an initial knowledge entity recognizer.
[0040] Step S134: Test, verify, train, and optimize the initial knowledge entity recognizer to construct the domain knowledge entity recognizer.
[0041] Specifically, historical fault records are mined and obtained from multiple sources such as system logs, hardware performance monitoring data, and network topology configuration to form a multi-source fault historical dataset. Domain experts, based on industry knowledge and operational experience, abstract and define the fault objects involved in the data. For example, specific fault phenomena or abnormal states of components such as server memory overflow, database deadlock, and excessive packet loss rate of switches are identified as independent fault entity types. For each entity type, identification rules are formulated, including feature keywords (such as memory overflow, deadlock), numerical thresholds (such as packet loss rate > 5%), and contextual semantic patterns (such as CPU utilization > 90% for 10 consecutive minutes). Finally, a set of multi-source fault entity identification rules covering the characteristics of multi-source data is formed.
[0042] Natural language processing (NLP) techniques are used to parse the text of a multi-source fault history dataset. Regular expression matching is used to identify feature keywords (such as server crashes and network latency) in the rules. A numerical calculation module is used to verify whether data indicators exceed preset thresholds (such as CPU utilization > 90%). A time series analysis algorithm is used to detect whether the contextual semantics conform to the time series pattern of the fault entities (such as three consecutive lost heartbeat packets). Entities that meet the rules are automatically extracted and labeled with entity types (hardware failure, software anomaly, etc.) using a data annotation tool. A data filtering module is used to remove records that do not meet the rules. Finally, an ETL (Extract-Transform-Load) tool is used to structure and store the processed entity data as a multi-source fault entity dataset.
[0043] A multi-source fault entity dataset is divided into training and validation sets according to a preset ratio, and trained using a bidirectional long short-term memory (Bi-LSTM) network. Through the network's bidirectional neuron structure, the network learns the positive and negative feature dependencies of fault entity data in time series, capturing temporal correlation patterns such as "equipment failure after continuous temperature rise." A proportional labeling mechanism is used to weight different types of fault entities during training, adjusting the weight ratio of each category in the loss function to address the data imbalance problem. During training, the model's entity recognition accuracy is evaluated in real time using the validation set, and network parameters (such as the number of hidden layer nodes and the learning rate) are iteratively optimized, ultimately generating an initial knowledge entity recognizer with preliminary fault entity recognition capabilities.
[0044] The initial knowledge entity recognizer was tested and validated using a test dataset independent of the training set. Its performance in fault entity recognition was evaluated by calculating metrics such as accuracy, recall, and F1 score. For recognition errors found in the test (such as missed recognition of rare fault entities and misclassification of similar entities), transfer learning was used to adjust network parameters and optimize the hidden layer structure and learning rate of the bidirectional long short-term memory network. At the same time, the model was iteratively trained with difficult examples labeled by domain experts to enhance its ability to recognize complex semantics and long-tailed fault entities. After multiple rounds of testing and optimization, a high-precision domain knowledge entity recognizer adapted to domain knowledge was constructed.
[0045] In one possible implementation, step S200 further includes:
[0046] Step S210: Combine the entity temporal characteristics to perform time granular identification and data alignment integration on the fault entity set and entity relationship set to obtain the standard fault entity set and standard entity relationship set.
[0047] Step S220: Based on the standard fault entity set and the standard entity relationship set, perform graph node concatenation to generate an initial fault knowledge graph.
[0048] Step S230: Receive multi-source fault update data streams in real time through the online streaming data processing module.
[0049] Step S240: Incrementally learn and update the initial fault knowledge graph based on the multi-source fault update data stream to obtain a fault update knowledge graph.
[0050] Specifically, for the set of faulty entities and the set of entity relationships, based on the time series characteristics of entity generation and change, a corresponding time granularity identifier is added to each entity and relationship, such as millisecond-level timestamps or minute-level statistical periods. At the same time, time series data from multiple sources and in multiple formats are aligned and integrated. Through time window sliding, interpolation completion and other methods, data from different sources are unified to the same time coordinate system to eliminate time deviations and inconsistencies, thereby forming a standardized set of faulty entities and set of entity relationships, providing a time-consistent data foundation for subsequent map construction.
[0051] Each fault entity in the standard fault entity set is treated as a node in the knowledge graph and assigned a unique identifier to form a fault knowledge graph node set. Then, each graph node is categorized and labeled with knowledge attributes, such as hardware type, software module, performance indicators, etc., to obtain a node knowledge attribute set. At the same time, each entity relationship in the standard entity relationship set is labeled with graph edge attributes to determine the edge type (such as causal relationship, dependency relationship) and weight. Combined with the business logic topology relationship, a logical concatenation analysis is performed on the graph edge set and the node knowledge attribute set to generate a graph node business concatenation network. Finally, based on this network, the fault knowledge graph node set is constructed by graph concatenation, thereby generating the initial fault knowledge graph.
[0052] By leveraging online streaming data processing modules (such as Apache Flink and Kafka Streams), a real-time data access channel is established to continuously monitor multiple data sources, including system logs, network monitoring, and hardware sensors. This allows for the capture of fault update data streams with millisecond-level response times. Simultaneously, the data streams undergo preliminary format standardization processing (such as unifying the JSON format) to ensure that fault data from different sources and with different structures (such as device alarms and sudden changes in performance metrics) can be collected in real time and transmitted to the data buffer queue, providing real-time data support for subsequent incremental updates to the knowledge graph.
[0053] The multi-source fault update data stream is parsed in real time using streaming computing frameworks (such as Apache Flink). Fault events and entities are identified through regular expression matching and semantic analysis, and new entity sets and relation sets are extracted. The conflict detection API of graph databases (such as Neo4j) is used to verify the consistency of new data with nodes (such as hardware components and performance indicators) and edges (such as dependencies and causal relationships) in the initial knowledge graph. Duplicate or contradictory data is corrected in weight or merged in version. The graph node vector representation is updated through graph embedding algorithms (such as Node2Vec), and the edge weights are dynamically adjusted using incremental learning algorithms (such as online gradient descent). Finally, the corrected entity data is synchronized to the knowledge graph storage layer through a batch import tool to complete the incremental update of the graph.
[0054] In one possible implementation, step S220 further includes:
[0055] Step S221: Treat each fault entity in the standard fault entity set as a knowledge graph node and assign an identifier to it to obtain a fault knowledge graph node set.
[0056] Step S222: Classify and label each graph node in the fault knowledge graph node set to obtain a node knowledge attribute set.
[0057] Step S223: Logically concatenate the fault knowledge graph node set based on the standard entity relationship set and the node knowledge attribute set to obtain an initial fault knowledge graph.
[0058] Specifically, each fault entity in the standard fault entity set obtained after time-granularity identification and data alignment is used as a node in the knowledge graph, and a unique identifier is assigned to each node, thus forming a fault knowledge graph node set. In this process, the identifier of each node is unique, so as to accurately identify and distinguish different fault entity nodes in the knowledge graph, laying the foundation for the subsequent construction of a complete knowledge graph.
[0059] For each node in the fault knowledge graph node set, knowledge attributes are categorized and labeled according to the type and characteristics of the fault entity it represents. For example, if a node represents a hardware device fault entity, it is labeled with attributes such as hardware type, model, and manufacturer; if it is a software module fault entity, it is labeled with attributes such as software version, functional module, and dependencies; if it belongs to a performance indicator fault entity, it is labeled with attributes such as indicator type, normal threshold, and current value. In this way, each node is given a clear attribute classification, ultimately forming a node knowledge attribute set containing various attribute information. This gives the nodes in the knowledge graph rich semantic information, providing a basis for subsequent attribute-based logical cascading.
[0060] Using the inter-entity relationships (such as causal, dependency, and influence relationships) defined in the standard entity relationship set as connecting links, and combining the types, attributes, and characteristics of each node in the node knowledge attribute set, the nodes in the fault knowledge graph node set are logically cascaded. For example, if there is a causal relationship of "server CPU overload - service response timeout" in the standard entity relationship set, and the attributes of the "server CPU overload" node and the "service response timeout" node in the corresponding node knowledge attribute set meet the triggering condition of this relationship, then a logical connection is established between the two nodes. By traversing all nodes and entity relationships, nodes with logical connections are connected according to the relationship type, ultimately forming an initial fault knowledge graph containing node attributes and logical relationships. This graph intuitively presents the relationships between fault entities in a graph structure form, providing a knowledge reasoning basis for fault root cause localization.
[0061] In one possible implementation, step S223 further includes:
[0062] Step S2231: Mark each entity relationship in the standard entity relationship set with graph edge attributes to determine the fault knowledge graph edge set.
[0063] Step S2232: Combine the business logic topology relationship to perform logical concatenation analysis on the fault knowledge graph edge set and the node knowledge attribute set to generate a graph node business concatenation network.
[0064] Step S2233: Construct the fault knowledge graph node set by graph concatenation based on the graph node service concatenation network to obtain the initial fault knowledge graph.
[0065] Specifically, for each entity relationship in the standard entity relationship set, the attributes of the graph edges are labeled to clarify the relationship type (such as causal relationship, dependency relationship, influence relationship, etc.), weight value (reflecting the strength of the relationship), direction (one-way or two-way), and timeliness (the effective time range of the relationship) represented by each edge. By standardizing and labeling the attributes of all entity relationships, the fault knowledge graph edge set is finally determined, laying the foundation for the subsequent construction of a knowledge graph with clear semantic relationships.
[0066] By combining the relational attributes (such as causal relationships, dependencies, and their weights) defined in the edge set of the fault knowledge graph with the business attributes (such as the business module and service level of the device) of each node in the node knowledge attribute set, and referring to the actual business logic topology (such as network topology and system architecture), the propagation path and scope of influence of the relationships between nodes in the business process are analyzed. For example, if there is a relationship of "database server failure - application service response timeout" in the edge set, and the node attributes show that the database server belongs to a core business module, then this relationship is included in the business cascading analysis. By traversing all the business attribute associations of the edges and nodes, a node cascading network reflecting the business logic is constructed, the propagation link of the fault in the business system is clarified, and a graph node business cascading network is generated, providing a connection basis at the business logic level for subsequent graph construction.
[0067] Using the business logic connections defined in the business cascading network of the graph nodes as a framework, the nodes in the fault knowledge graph node set are cascaded and constructed according to the fault propagation paths and logical associations determined by the business cascading network. For example, based on the cascading relationship of "core switch failure - server cluster communication interruption - application service unavailability" in the business cascading network, the corresponding fault entity nodes are connected sequentially, and the edges are assigned corresponding business relationship attributes. By traversing all logical connections in the business cascading network, the nodes and edges are structurally combined to finally form an initial fault knowledge graph containing node attributes, business relationships, and cascading logic. This graph presents the associations between fault entities in the business system in an intuitive graph structure, providing a basic architecture for knowledge reasoning in fault root cause analysis.
[0068] In one possible implementation, step S240 further includes:
[0069] Step S241: Perform fault event identification and fault entity parsing on the multi-source fault update data stream to obtain the set of newly added fault entities and the set of newly added entity relationships.
[0070] Step S242: Use the initial fault knowledge graph to perform conflict detection and knowledge correction on the newly added fault entity set and the newly added entity relationship set to obtain fault correction entity data.
[0071] Step S243: Based on the fault correction entity data, incrementally learn and update the initial fault knowledge graph to obtain the fault update knowledge graph.
[0072] Specifically, the fault event identification engine deployed in the streaming data processing module performs real-time parsing of multi-source fault update data streams (such as system log streams, performance indicator streams, and network topology streams). It uses regular expressions to match preset fault feature patterns (such as CPU utilization > 90% and connection timeout) to identify fault events. At the same time, it uses named entity recognition (NER) technology to extract fault entities (such as servers, switches, and databases) from unstructured log texts and uses dependency parsing to mine semantic relationships between entities (such as causing, affecting, and depending on). Finally, it filters and structures the newly added fault entity set and the newly added entity relationship set from the data stream, providing incremental data for subsequent knowledge graph updates.
[0073] The newly added entity set and the newly added entity relationship set are input into the initial fault knowledge graph. Utilizing the graph's structured data model, the new data is comprehensively compared with existing nodes and edges. For the newly added entity set, attributes such as node identifiers and entity types are checked. If duplicates are found with existing nodes in the graph, the attribute information of identical nodes is merged, prioritizing the retention of the attribute value with the latest timestamp or the highest confidence level. For the newly added entity relationship set, the relationship type, association direction, and weight are verified. If contradictions exist with existing relationships in the graph, the relationship weights are recalculated or the relationship direction is adjusted based on historical fault data statistics and actual business logic. After identifying and processing conflicting data, accurate and consistent fault correction entity data is finally formed.
[0074] The fault correction entity data, after conflict detection and knowledge correction, is used as the base data for incremental updates to the initial fault knowledge graph. For newly added fault entities within the fault correction entity data, they are added as new nodes to the initial knowledge graph, assigned a unique identifier, and tagged with corresponding knowledge attributes. New entity relationships are added as new edges connecting the corresponding fault entity nodes, with the relationship type and weight determined based on business logic and historical data. If the fault correction entity data contains correction information related to existing nodes or edges in the initial knowledge graph, the corresponding node attributes or edge relationship attributes are updated to ensure the information in the knowledge graph accurately reflects the latest fault data. Through this incremental learning and updating method, the fault correction entity data is integrated into the initial fault knowledge graph, ultimately resulting in a fault-updated knowledge graph that reflects the latest fault conditions and relationships between entities.
[0075] In one possible implementation, step S400 further includes:
[0076] Step S410: Calculate the similarity between the historical fault case library and the target fault data to obtain a fault case similarity set.
[0077] Step S420: Use the fault case similarity set to optimize the historical fault case library to obtain target compatible fault cases.
[0078] Step S430: Based on the target adapted fault cases, perform intersection matching on the candidate root cause set of the fault mode to determine the target fault root cause location result.
[0079] Specifically, the similarity between the multi-dimensional feature vectors in the target fault data (such as fault phenomenon text vectors, abnormal patterns in time-series indicators, and topology-related device ID sequences) and the corresponding feature vectors of each case in the historical fault case database is calculated. For textual features, a BERT pre-trained model is used to extract semantic vectors and then cosine similarity is calculated. For time-series indicators, the Dynamic Time Warping (DTW) algorithm is used to match waveform similarity. For topology, graph edit distance is used to measure the similarity of device relationships. Each feature is assigned a weight coefficient (e.g., 0.4 for phenomenon description, 0.3 for time-series indicators, and 0.3 for topology). Finally, a weighted summation is used to obtain the comprehensive similarity value between the target fault and each historical case, forming a fault case similarity set containing the similarity scores of all historical cases, providing a quantitative basis for subsequent case selection.
[0080] Based on the similarity scores between historical cases and target fault data in the fault case similarity set, the historical fault case database is filtered and sorted. A similarity threshold (e.g., 0.7) is set, prioritizing historical cases with similarity scores higher than the threshold. If there are many cases higher than the threshold, they are sorted in descending order of similarity score, and the top-scoring cases are selected as target-fit fault cases. These target-fit fault cases are most similar to the current target fault in key characteristics such as fault phenomena, involved equipment, and scope of impact, providing more targeted and referential historical experience and solutions for subsequent fault diagnosis and handling.
[0081] The root cause information recorded in the target adaptation failure cases is compiled into a list, and a corresponding list of candidate root causes for the failure mode is also formed. These two lists are then cross-compared to identify root cause items that exist in both lists. If a common root cause exists, it is determined as the root cause localization result for the target failure. If no common root cause exists, the historical frequency of each root cause in the target adaptation cases is further analyzed, and the root cause with the highest frequency is selected. Then, combined with the dependencies between devices and the impact path of the business process, logical reasoning is used to verify whether this root cause conforms to the current failure propagation pattern, ultimately determining the root cause localization result for the target failure.
[0082] In one possible implementation, step S400 further includes:
[0083] Step S440: Construct a controllable injection test environment, and perform injection test verification on the target fault root cause localization result based on the controllable injection test environment to obtain the localization accuracy verification result.
[0084] Step S450: Based on the location accuracy verification results, perform iterative feedback optimization on the fault update knowledge graph.
[0085] Specifically, virtualization technologies such as VMware or KVM are used to build a virtual testing environment that closely resembles the actual production environment, completely replicating elements such as network topology, server configuration, software versions, and business applications. Simultaneously, fault injection tools, such as the Chaos Toolkit framework, are deployed to precisely control the type, timing, and intensity of fault injection. Next, the fault type corresponding to the root cause localization result, such as hardware failure, software vulnerability, or network interruption, is translated into specific injection commands. For example, if the localization result is a server disk failure, a fault injection tool simulates disk I / O errors or a full disk; if it's a network interruption, a network simulation tool (such as Mininet) disconnects the specified network link. After fault injection, various metrics in the testing environment are monitored in real time, including system logs, service response time, and resource utilization. By comparing the changes in metrics before and after fault injection, and whether they match the actual fault phenomena, the accuracy of the root cause localization result is determined. Finally, based on the monitoring and comparative analysis results, a verification result of the localization accuracy, including whether the localization was correct and related evidence, is output.
[0086] Based on the accuracy, confidence, and reproducibility metrics in the location accuracy verification results, the fault update knowledge graph undergoes a three-level iterative optimization. Structural layer optimization: If the verification results show that the root cause location accuracy for a certain type of fault is below a threshold (e.g., 70%), the similarity between nodes is recalculated using a graph embedding algorithm (e.g., Node2Vec). This adds or adjusts the association edges between fault entities and root causes in the graph, strengthening high-confidence relationships and weakening low-confidence relationships. Parameter layer optimization: For fault scenarios reproduced during verification, relationship weights are dynamically updated (e.g., increasing the association weight of CPU overload-service unavailability from 0.6 to 0.8), and time window parameters are adjusted (e.g., correcting the fault propagation delay from 5 minutes to 3 minutes). Knowledge layer optimization: If a new fault mode not covered by the graph is discovered (e.g., unexpected cascading faults appearing during injection testing), the knowledge extraction module is triggered to extract new entities and relationships from the verification logs, which are then added to the graph after expert rule review. After each optimization, consistency checks are performed through a knowledge reasoning engine (such as Drools) to ensure the logical consistency of the graph, ultimately forming a more accurate fault update knowledge graph.
[0087] Example 2, based on the same inventive concept as the dynamic knowledge graph-driven fault root cause localization method in the foregoing examples, such as... Figure 2 As shown, this application provides a dynamic knowledge graph-driven fault root cause localization system. The system and method embodiments in this application are based on the same inventive concept. The system includes:
[0088] The entity association extraction module 10 is used to collect and obtain a multi-source fault association dataset, and to extract entity associations from the multi-source fault association dataset to obtain a fault entity set and an entity relationship set.
[0089] The fault update knowledge graph construction module 20 is used to combine the temporal features of entities to perform graph node concatenation and incremental learning updates on the fault entity set and entity relationship set, thereby constructing a fault update knowledge graph.
[0090] The pattern matching reasoning module 30 is used to monitor and acquire target fault data, and use the fault update knowledge graph to perform pattern matching reasoning on the target fault data to generate a set of candidate root causes of fault modes.
[0091] The similarity matching module 40 is used to perform similarity matching on the candidate root cause set of the fault mode in combination with the historical fault case library, determine the target fault root cause location result, and perform controllable injection testing and iterative feedback optimization on the fault update knowledge graph based on the target fault root cause location result.
[0092] Furthermore, the system is also used to implement the following functions:
[0093] Based on the attribute type information of the multi-source fault association dataset, a multi-source noise data discrimination standard is determined, and a multi-source noise data cleaning procedure is set up based on the multi-source noise data discrimination standard. Noise identification and cleaning are performed on the multi-source fault association dataset according to the multi-source noise data discrimination standard and the multi-source noise data cleaning procedure to obtain a usable multi-source fault association dataset. A domain knowledge entity recognizer is constructed, and the domain knowledge entity recognizer is used to perform entity recognition on the usable multi-source fault association dataset to obtain a fault entity set. Semantic relationship recognition and business logic association are performed between the entities in the fault entity set to determine the entity relationship set.
[0094] Furthermore, the system is also used to implement the following functions:
[0095] A multi-source fault history dataset is obtained by mining and defining fault entities in the dataset according to domain knowledge and experience, thereby obtaining multi-source fault entity recognition rules. Based on these rules, entity data recognition is performed on the dataset to obtain a multi-source fault entity dataset. A bidirectional long short-term memory network is used to train the dataset with proportional annotation to obtain an initial knowledge entity recognizer. The initial knowledge entity recognizer is then tested, validated, and optimized to construct the domain knowledge entity recognizer.
[0096] Furthermore, the system is also used to implement the following functions:
[0097] By combining the temporal characteristics of entities, the set of faulty entities and the set of entity relationships are identified and aligned with time granularity to obtain a standard set of faulty entities and a standard set of entity relationships. Based on the standard set of faulty entities and the standard set of entity relationships, graph nodes are cascaded to generate an initial fault knowledge graph. Multi-source fault update data streams are received in real time through an online streaming data processing module. Based on the multi-source fault update data streams, the initial fault knowledge graph is incrementally learned and updated to obtain a fault update knowledge graph.
[0098] Furthermore, the system is also used to implement the following functions:
[0099] Each fault entity in the standard fault entity set is treated as a knowledge graph node and assigned an identifier to obtain a fault knowledge graph node set; each graph node in the fault knowledge graph node set is classified and labeled with knowledge attributes to obtain a node knowledge attribute set; the fault knowledge graph node set is logically concatenated based on the standard entity relationship set and the node knowledge attribute set to obtain an initial fault knowledge graph.
[0100] Furthermore, the system is also used to implement the following functions:
[0101] Each entity relationship in the standard entity relationship set is identified by graph edge attributes to determine the fault knowledge graph edge set; the fault knowledge graph edge set and the node knowledge attribute set are logically concatenated by combining business logic topology relationships to generate a graph node business concatenation network; the fault knowledge graph node set is constructed by graph concatenation based on the graph node business concatenation network to obtain the initial fault knowledge graph.
[0102] Furthermore, the system is also used to implement the following functions:
[0103] The multi-source fault update data stream is subjected to fault event identification and fault entity parsing to obtain a set of newly added fault entities and a set of newly added entity relationships. The initial fault knowledge graph is used to perform conflict detection and knowledge correction on the set of newly added fault entities and the set of newly added entity relationships to obtain fault correction entity data. Based on the fault correction entity data, the initial fault knowledge graph is incrementally learned and updated to obtain the fault update knowledge graph.
[0104] Furthermore, the system is also used to implement the following functions:
[0105] Similarity calculation is performed between the historical fault case database and the target fault data to obtain a fault case similarity set; the historical fault case database is optimized using the fault case similarity set to obtain target-fit fault cases; the candidate root cause set of the fault mode is intersected and matched based on the target-fit fault cases to determine the target fault root cause location result.
[0106] Furthermore, the system is also used to implement the following functions:
[0107] A controllable injection test environment is constructed, and an injection test is performed on the target fault root cause localization result based on the controllable injection test environment to obtain the localization accuracy verification result; the fault update knowledge graph is iteratively optimized based on the localization accuracy verification result.
[0108] It should be noted that the order of the embodiments described above is merely for descriptive purposes and does not represent the superiority or inferiority of the embodiments. Furthermore, the above description focuses on specific embodiments of this specification. Additionally, the processes depicted in the accompanying drawings do not necessarily require a specific or sequential order to achieve the desired results. In some implementations, multitasking and parallel processing are possible or may be advantageous.
[0109] The above description is only a preferred embodiment of this application and is not intended to limit this application. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the protection scope of this application.
[0110] This specification and accompanying drawings are merely illustrative examples of this application and are intended to cover any and all modifications, variations, combinations, or equivalents within the scope of this application. Clearly, those skilled in the art can make various alterations and modifications to this application without departing from its scope. Therefore, if such modifications and variations fall within the scope of this application and its equivalents, this application intends to include such modifications and variations.
Claims
1. A dynamic knowledge graph-driven method for fault root cause localization, characterized in that, The method includes: Collect a multi-source fault association dataset, and extract entity associations from the multi-source fault association dataset to obtain a fault entity set and an entity relationship set; By combining the temporal characteristics of entities, the fault entity set and entity relationship set are cascaded and incrementally learned and updated to construct a fault update knowledge graph. The target fault data is monitored and acquired. The fault update knowledge graph is used to perform pattern matching reasoning on the target fault data to generate a set of candidate root causes of fault patterns. By combining the historical fault case library with the candidate root cause set of the fault mode, similarity matching is performed to determine the target fault root cause location result, and the fault update knowledge graph is subjected to controllable injection testing and iterative feedback optimization based on the target fault root cause location result.
2. The dynamic knowledge graph-driven fault root cause localization method as described in claim 1, characterized in that, The obtained set of faulty entities and set of entity relationships include: Based on the attribute type information of the multi-source fault association dataset, determine the multi-source noise data discrimination criteria, and set up a multi-source noise data cleaning procedure based on the multi-source noise data discrimination criteria. The multi-source noise data discrimination criteria and the multi-source noise data cleaning procedure are used to identify and clean the noise in the multi-source fault association dataset to obtain a usable multi-source fault association dataset. Construct a domain knowledge entity recognizer, and use the domain knowledge entity recognizer to perform entity recognition on the available multi-source fault association dataset to obtain a fault entity set; Semantic relationships and business logic associations are performed on the entities in the fault entity set to determine the entity relationship set.
3. The dynamic knowledge graph-driven fault root cause localization method as described in claim 2, characterized in that, The constructed domain knowledge entity recognizer includes: A multi-source fault history dataset is mined and obtained. Fault entities are defined in the multi-source fault history dataset according to domain knowledge and experience, and multi-source fault entity recognition rules are obtained. Based on the multi-source fault entity identification rules, entity data identification is performed on the multi-source fault historical dataset to obtain a multi-source fault entity dataset. A bidirectional long short-term memory network was used to train the multi-source fault entity dataset with proportional annotation to obtain an initial knowledge entity recognizer. The initial knowledge entity recognizer is tested, verified, trained, and optimized to construct the domain knowledge entity recognizer.
4. The dynamic knowledge graph-driven fault root cause localization method as described in claim 1, characterized in that, The construction of the fault update knowledge graph includes: By combining the temporal characteristics of entities, the set of faulty entities and the set of entity relationships are identified by time granularity and integrated with data alignment to obtain a standard set of faulty entities and a standard set of entity relationships. Based on the standard fault entity set and the standard entity relationship set, the graph nodes are concatenated to generate an initial fault knowledge graph; The online streaming data processing module receives multi-source fault update data streams in real time. Based on the multi-source fault update data stream, the initial fault knowledge graph is incrementally learned and updated to obtain a fault update knowledge graph.
5. The dynamic knowledge graph-driven fault root cause localization method as described in claim 4, characterized in that, The generation of the initial fault knowledge graph includes: Each fault entity in the standard fault entity set is used as a knowledge graph node and an identifier is assigned to it to obtain a fault knowledge graph node set. Each node in the fault knowledge graph node set is classified and labeled with knowledge attributes to obtain a node knowledge attribute set. The fault knowledge graph node set is logically concatenated based on the standard entity relationship set and the node knowledge attribute set to obtain an initial fault knowledge graph.
6. The dynamic knowledge graph-driven fault root cause localization method as described in claim 5, characterized in that, The process of obtaining the initial fault knowledge graph includes: Each entity relationship in the standard entity relationship set is labeled with a graph edge attribute to determine the fault knowledge graph edge set; By combining the business logic topology relationship, a logical concatenation analysis is performed on the edge set of the fault knowledge graph and the node knowledge attribute set to generate a business concatenation network of graph nodes. Based on the service concatenation network of the graph nodes, the fault knowledge graph node set is concatenated to construct the initial fault knowledge graph.
7. The dynamic knowledge graph-driven fault root cause localization method as described in claim 4, characterized in that, The process of obtaining the fault update knowledge graph includes: The multi-source fault update data stream is subjected to fault event identification and fault entity parsing to obtain the set of newly added fault entities and the set of newly added entity relationships. The initial fault knowledge graph is used to perform conflict detection and knowledge correction on the newly added fault entity set and the newly added entity relationship set to obtain fault correction entity data. Based on the fault correction entity data, the initial fault knowledge graph is incrementally learned and updated to obtain the fault update knowledge graph.
8. The dynamic knowledge graph-driven fault root cause localization method as described in claim 1, characterized in that, The determination of the root cause location of the target fault includes: A similarity set of fault cases is obtained by calculating the similarity between the historical fault case database and the target fault data. The historical fault case library is optimized using the fault case similarity set to obtain target-fit fault cases; Based on the target-adapted fault cases, the candidate root cause set of the fault mode is cross-matched to determine the target fault root cause location result.
9. The dynamic knowledge graph-driven fault root cause localization method as described in claim 1, characterized in that, The controlled injection testing and iterative feedback optimization of the fault update knowledge graph based on the target fault root cause localization results includes: A controllable injection test environment is constructed, and an injection test is performed on the target fault root cause localization result based on the controllable injection test environment to obtain the localization accuracy verification result; Based on the location accuracy verification results, the fault update knowledge graph is iteratively optimized.
10. A dynamic knowledge graph-driven fault root cause localization system, characterized in that, The system is used to implement the dynamic knowledge graph-driven fault root cause localization method according to any one of claims 1-9, and the system comprises: The entity association extraction module is used to collect and obtain a multi-source fault association dataset, and to extract entity associations from the multi-source fault association dataset to obtain a fault entity set and an entity relationship set. The fault update knowledge graph construction module is used to combine entity temporal features to perform graph node concatenation and incremental learning updates on the fault entity set and entity relationship set to construct a fault update knowledge graph. The pattern matching and reasoning module is used to monitor and acquire target fault data, and to perform pattern matching and reasoning on the target fault data using the fault update knowledge graph to generate a set of candidate root causes of fault patterns. The similarity matching module is used to perform similarity matching on the candidate root cause set of the fault mode in combination with the historical fault case library, determine the target fault root cause location result, and perform controllable injection testing and iterative feedback optimization on the fault update knowledge graph based on the target fault root cause location result.
Citation Information
Patent Citations
Fault diagnosis method and device based on knowledge graph and computing equipment
CN115619383A
Equipment fault information sharing system and method based on Internet of Things
CN119477277A
Container fault tracing method and device based on cross-community knowledge fusion and medium
CN120045726A
Equipment fault diagnosis recommendation method based on knowledge graph and multi-dimensional relevance indexes
CN120234376A
Operation and maintenance system and method
US20210271582A1
Cited By
Security event root cause traceability analysis method and system based on interactive questions and answers
CN121418210A
A method and system for root cause analysis of security incidents based on interactive question-and-answer
CN121418210B
Visual operation and maintenance fault diagnosis and problem attribution analysis method for information system
CN121478537A
A method for visualizing operation and maintenance fault diagnosis and problem attribution analysis of an information system
CN121478537B
Fault root cause positioning method and device, electronic equipment and storage medium
CN121567537A