Network Fault Diagnosis Using Topology Graphs and Change Rollback
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Identifying and rectifying faulty components in networked computer systems is challenging due to the complexity of interconnectivity, which multiplies fault messages and makes it difficult to pinpoint the root cause, especially with periodic system changes complicating the issue.
Innovation Solution
A method and system that generates a network topology graph, associates alerts with nodes, identifies common nodes affecting multiple components, retrieves operational changes, and determines the faulty component, allowing for rollback or reconfiguration to resolve issues.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If the system monitors all components in a networked computer system, then the reliability of fault detection is improved, but the device complexity increases due to the interconnectivity of various software and hardware components
Solution Approach 1:
The patent segments the networked computer system into distinct components (hardware components, software components, and their interconnections) represented as nodes and edges in a graph structure. This segmentation allows the system to process and analyze faults in isolated units rather than treating the entire complex system as a single monolithic entity, thereby reducing the effective complexity while maintaining comprehensive monitoring capability.
Solution Approach 2:
The patent introduces a fault analysis system that acts as an intermediary between the monitored components and the fault detection process. This intermediary system receives fault messages from various components, analyzes them through the graph structure, and identifies root causes without requiring direct monitoring of all interconnections, thus reducing system complexity while maintaining detection reliability.
2Measurement precision
If the system tracks all operational changes in the networked computer system, then the ability to identify root causes is improved, but the loss of time increases due to the volume of changes and fault messages
Solution Approach 1:
The patent implements preliminary action by pre-establishing the graph structure of the networked computer system before faults occur. The graph is built in advance with nodes representing components and edges representing interconnections, so that when faults happen, the analysis can immediately utilize this pre-configured structure without time-consuming on-the-fly mapping, thereby reducing fault identification time while maintaining precision.
Solution Approach 2:
The patent extracts only the essential information needed for fault analysis from the complete operational change logs. Instead of processing all changes uniformly, the system extracts and focuses on changes related to specific nodes in the graph that are connected to fault messages, filtering out irrelevant information and reducing the time required to identify root causes.
3Loss of information
If the system processes all fault messages from interconnected components, then the completeness of fault information is improved, but the loss of energy increases due to the multiplication of fault messages
Solution Approach 1:
The patent merges multiple fault messages that stem from the same root cause into a single analysis unit. By using the graph structure to trace connections between components, the system identifies that multiple fault messages are related and consolidates their processing, thereby maintaining information completeness while reducing the total computational resources required compared to processing each message independently.
Solution Approach 2:
The patent converts the harmful effect of message multiplication (which wastes computational resources) into a beneficial diagnostic signal. The increased number of fault messages from interconnected components is used as input to the graph-based analysis system, which then identifies patterns and root causes that would be invisible in single-message processing, turning the resource waste into valuable diagnostic information.
Data Source
AI summary
Described are a system, method, and computer program product for diagnosing faulty components in networked computer systems. The method includes generating a graph of a network topology of a networked computer system and determining a set of nodes of the graph affected by a fault in the networked computer system based on an alert associated with the set of nodes. The method also includes determining a faulty component of the networked computer system based on a common node having a plurality of edges connected to nodes in the set of nodes affected by the fault. The method further includes retrieving a set of records of operational changes to the networked computer system and determining an operational change that caused the fault. The method further includes resetting the networked computer system to a prior state before the operational change.


