Cross-Device Packet and State Extraction for Network Error Diagnosis
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional methods for collecting network device state information during unexpected events or errors are slow and inefficient, often leaving networks in error states for extended periods, and lack dynamic domain-wide coordination, leading to excessive resource consumption and difficulty in determining root causes.
Innovation Solution
A mechanism for coordinating packet and state extraction across network devices using control/data plane signaling, employing a Serviceability Analytics Engine (SAE) to determine trigger events and actions, utilizing proprietary, in-band control-plane, and multicast domain probes to capture device states on-demand, allowing for rapid analysis without leaving the network in an error state.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Speed
If SNMP is used to collect telemetry data from network devices, then data collection can be performed, but the process is slow and has high latency making it unsuitable for diagnosing intermittent issues at the right time
Solution Approach 1:
The system pre-configures network devices with event triggers and corresponding actions before events occur. When an event is detected, the pre-configured actions are executed immediately, allowing rapid data collection without the latency of conventional SNMP polling or streaming telemetry setup.
Solution Approach 2:
The patent introduces an event trigger mechanism as an intermediary between network events and data collection. This trigger system coordinates packet capture and state extraction across multiple devices simultaneously, enabling synchronized rapid data collection that conventional methods cannot achieve.
2Productivity
If streaming telemetry is used to update operational states, then continuous monitoring is possible, but the update process is slow and telemetry data across devices will not be useful for triaging specific intermittent problems
Solution Approach 1:
The system configures different event triggers and actions for different network devices based on their specific roles and the particular issue being investigated. Each device captures only the relevant state information needed for the specific problem, rather than uniformly collecting all telemetry data, making the data more useful for targeted triage.
Solution Approach 2:
The event trigger system dynamically configures data collection based on detected events. When an event occurs, the system activates specific triggers and actions tailored to that event type, enabling adaptive data collection that captures exactly the information needed for the specific issue without the overhead of continuous comprehensive telemetry.
3Ease of repair
If conventional data collection methods are used to determine root cause, then analysis can be performed, but network enterprises consume excessive resources in terms of time and money
Solution Approach 1:
The system segments the data collection process into event-specific triggers and actions. Instead of collecting all possible data from all devices continuously, the system divides data collection into targeted segments based on the specific event and the devices involved, reducing resource consumption while maintaining root cause determination capability.
Solution Approach 2:
The event trigger mechanism performs exactly the right amount of data collection needed for each specific event - no more, no less. This partial action approach collects sufficient data for root cause analysis without the excessive resource consumption of comprehensive continuous telemetry, optimizing the balance between diagnostic capability and resource usage.
Data Source
AI summary
Techniques for capturing device state information from multiple network devices when a network error occurs are described. A network controller determines a trigger event and a trigger action. The trigger action is an action to be taken by a network device in response to the trigger event occurring, and the trigger event is an unexpected network event or network error. The network controller determines a method usable to configure the network devices to perform the trigger action in response to detecting the trigger event. The network controller configures, according to the method, the network devices to perform the trigger action in response to detecting the trigger event.


