Probe-Triggered Full Device State Capture for Network Diagnostics
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Network monitoring and troubleshooting face challenges in identifying the root causes of network impairments due to intermittent issues and the lack of clear data on which network nodes to gather information from, as existing synthetic network probes only reveal issues along a network path without providing underlying causes.
Innovation Solution
The implementation of probe-triggered full device state capture and correlation, where a trigger probe generates a unique correlation identifier to capture and export a 'frozen-in-time' full device state from designated network nodes, including control and data plane states, allowing for the correlation of reports to diagnose performance problems.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If synthetic network probes are used to monitor network performance, then network issues can be detected along the network path, but the underlying causes of network impairments remain unclear and difficult to diagnose
Solution Approach 1:
The system performs preliminary actions by capturing full device states at network nodes before issues fully manifest or become intermittent. When a probe detects a performance problem, the system proactively captures the complete device state (memory, registers, packet buffers) at that moment, preserving evidence before it may be lost or changed. This preliminary capture enables later analysis of root causes even for intermittent issues.
Solution Approach 2:
The invention segments the network troubleshooting process into distinct phases: (1) probe-based detection of performance problems, (2) triggered capture of full device states at specific nodes, (3) export and storage of captured states, and (4) correlation analysis. This segmentation allows each phase to be optimized independently and makes the complex troubleshooting process manageable and systematic.
2Loss of information
If full device state capture is performed at all network nodes continuously, then complete data for root cause analysis would be available, but network device resources and bandwidth would be excessively consumed
Solution Approach 1:
Instead of uniformly capturing data at all network nodes, the system applies local quality by selectively capturing full device states only at specific nodes where performance problems are detected by probes. The capture is triggered locally at the affected node based on probe results, ensuring complete diagnostic data is obtained only where needed, thus avoiding unnecessary resource consumption at unaffected nodes.
Solution Approach 2:
The system performs partial action by capturing only the necessary full device state information (memory, registers, packet buffers) at the specific moment when a problem is detected, rather than continuously capturing all possible data from all nodes. This partial capture approach provides sufficient diagnostic information while minimizing resource overhead.
3Measurement precision
If detailed full device state information is captured and stored, then root cause diagnosis becomes possible, but the complexity of data management and correlation increases
Solution Approach 1:
The system implements feedback by using probe results to trigger capture actions and using correlation identifiers to link captured device states back to the original performance problems. The correlation ID embedded in captured data provides feedback that enables automatic association of device states with specific issues, reducing manual data management complexity while maintaining high diagnostic accuracy.
4Loss of time
If network administrators manually determine which data to gather from network nodes, then data collection can be targeted, but time is lost in identifying the right data sources for intermittent issues
Solution Approach 1:
The system enables self-service by automatically triggering full device state captures at network nodes based on probe-detected performance problems, without requiring administrator intervention to identify which nodes to monitor. The system autonomously determines data collection needs and executes captures, significantly reducing troubleshooting time while maintaining operational simplicity for administrators.
Data Source
AI summary
A method comprising: at a management entity configured to communicate with a network: upon detecting a performance problem on a network path in the network, generating a trigger probe having a correlation identifier, the trigger probe configured to transit the network path and, on one or more designated network nodes of the network path, trigger (i) capturing a full device state, including a control plane state and a data plane state, and (ii) exporting a report of the full device state with the correlation identifier; sending the trigger probe along the network path; receiving, from each of the one or more designated network nodes, the report that includes the correlation identifier and the full device state; and correlating each report to the performance problem based on the correlation identifier in each report, to diagnose a root cause of the performance problem using the full device state in each report.


