Programmable Diagnosis Model for Dynamic Network Root Cause Analysis
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current root cause analysis (RCA) technologies in computer networks struggle to adapt to dynamic network changes, handle complex fault propagation scenarios, and correlate issues across multiple layers and services, especially in heterogeneous networks with frequent changes and concurrent faults.
Innovation Solution
A programmable network diagnosis model that uses a resource definition graph to model interdependencies between network devices and services, enabling forward chaining-based RCA, temporal relation consideration, and scalable integration of new services, allowing for dynamic adaptation and reliable error resilience.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If traditional RCA technologies are used to analyze network faults, then the analysis process is simple, but the system cannot adapt to dynamic network changes and complex fault propagation scenarios
Solution Approach 1:
The patent implements a dynamic resource definition graph that automatically updates when network resources change. The system monitors network state changes and dynamically adjusts the graph structure, node attributes, and interdependencies without requiring manual reconfiguration. This enables the RCA system to adapt to dynamic network changes while maintaining a manageable complexity level through automated updates.
Solution Approach 2:
The resource definition graph serves multiple functions: it models network resources, defines interdependencies, tracks temporal relationships, and supports fault propagation analysis. This universal structure handles diverse network scenarios (concurrent faults, cascading failures, temporal correlations) within a single framework, improving adaptability without proportionally increasing system complexity.
2Measurement precision
If a detailed resource definition graph is created to model all network interdependencies, then the root cause analysis precision is improved, but the system complexity and computational overhead increase
Solution Approach 1:
The patent segments the network model into discrete resources with specific attributes and interdependencies. Each network component (device, service, resource) is represented as an independent node with defined characteristics. This segmentation allows precise tracking of fault propagation paths while managing complexity through modular, hierarchical organization of the resource definition graph.
Solution Approach 2:
The system dynamically changes parameters of resources and interdependencies based on detected faults and temporal relationships. When faults occur, the system adjusts the state parameters of affected resources and updates interdependency relationships, enabling precise root cause identification. This parameter-based approach maintains precision while managing complexity through focused updates rather than complete model reconstruction.
3Reliability
If the system monitors all network events and resources continuously, then the fault detection capability is improved, but the computational load and processing time increase
Solution Approach 1:
The patent establishes telemetry rules and interdependency relationships in advance before faults occur. The resource definition graph pre-defines monitoring parameters, correlation rules, and propagation paths. When faults occur, the system executes pre-configured analysis logic rather than creating analysis frameworks from scratch, improving detection capability while reducing real-time computational load.
Solution Approach 2:
The system skips unnecessary processing steps by directly querying pre-defined interdependencies and resource attributes in the graph. Instead of performing comprehensive network-wide analysis for every event, the system rapidly traverses relevant paths in the resource definition graph based on event type and location, improving processing efficiency while maintaining reliable fault detection through targeted analysis.
Data Source
AI summary
Network management techniques are described. A controller device of this disclosure manages a device group of a network. The controller device includes processing circuitry in communication with the memory, the processing circuitry being configured to receive, using a programmable diagnosis service executed by the processing circuitry, a programming input, to form, using the programmable diagnosis service, based on the programming input, a resource definition graph that models interdependencies between a plurality of resources supported by the device group, to detect, using the programmable diagnosis service, an event affecting a first resource of the plurality of resources, and to identify, using the programmable diagnosis service, based on the interdependencies modeled in the resource definition graph formed based on the programming input, a root cause event that caused the event affecting the first resource, the root cause event occurring at a second resource of the plurality of resources.


