Network Diagnostic Sampling in Distributed Computing
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Large and complex computer networks face inefficiencies in identifying and addressing system issues due to the high processing power and memory requirements for analyzing data from all machines, which can lead to further problems when responding to issues, such as data loss or system failure.
Innovation Solution
A central networking system employs diagnostic sampling methods and network monitoring rules that include sampling rules and safety instructions to select a subset of nodes for analysis, perform remedial actions, and ensure safe data collection without straining the network, using a subset of nodes to address issues and return the network to a stable state.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If data from all machines on the network is analyzed to identify system issues, then comprehensive network monitoring is achieved, but processing power and memory space requirements increase significantly
Solution Approach 1:
The patent divides the network monitoring task into segments by selecting a representative subset of nodes rather than analyzing all nodes. The system segments the large network into manageable samples that can be analyzed independently, reducing the overall processing burden while maintaining monitoring effectiveness.
Solution Approach 2:
The patent applies partial action by analyzing data from only a subset of nodes rather than all nodes. This selective sampling approach performs sufficient monitoring to identify network issues without the excessive processing power and memory space required for complete network analysis.
2Productivity
If remedial actions are executed immediately on machines experiencing system issues, then rapid problem resolution is achieved, but the machine may lose log data or additional functions may fail
Solution Approach 1:
The patent applies preliminary action by stabilizing the affected machine before executing remedial actions. The system first ensures the machine is in a stable state and log data is secured, then performs the necessary remediation. This prevents data loss and additional failures that would occur with immediate intervention.
Solution Approach 2:
The patent provides beforehand cushioning by preparing the machine state prior to remedial actions. The system buffers the machine stability and log data integrity before applying fixes, creating a protective measure that prevents harmful effects during the remediation process.
3Productivity
If a subset of nodes is sampled for analysis instead of all nodes, then processing efficiency is improved, but measurement precision of network-wide issues may be reduced
Solution Approach 1:
The patent uses feedback by analyzing sampled node data to detect network issues, then applying remedial actions based on those findings. The system continuously monitors the effects of remedial actions on the sampled nodes and uses this feedback to determine whether additional actions are needed, improving detection accuracy while maintaining sampling efficiency.
Data Source
AI summary
A central networking system supports efficient identification and analysis of problems that occur at associated nodes on the network. Using network monitoring rules, the central networking system samples data from a subset of nodes in response to an indication that an error or problem has occurred on the network. If the collected sample data is determined to satisfy certain network conditions, the central networking system proceeds to perform network operations on nodes of the entire network, as appropriate. Thus, the system does not need to collect data from every node in a large network to address potential network threats. The central networking system also defines rules for detecting when a node experiencing a problem violates safety conditions such that it is impossible or inadvisable to pull analytical data from the node. The system performs appropriate remedial actions to address the node problems prior to requesting data for analysis.


