Distributed Network Problem Detection via Hierarchical Zone Servers
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
In large networks, existing systems face challenges in communicating and processing information quickly, leading to delayed problem reporting and weak cause determination due to limited communication bandwidth and computing power.
Innovation Solution
A method where machines in the network identify and evaluate problems, generate messages with probabilistic activity to determine prevalence and severity, and propagate these messages through a self-organizing system to ensure only significant issues are reported to the receiver/servers, thereby reducing the load on the network.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Loss of information
If all machines report all detected problems to the receiver/server, then complete problem information is collected, but communication bandwidth is overwhelmed and processing is delayed
Solution Approach 1:
The patent segments the network into multiple zones with zone servers that aggregate problem data locally before reporting to the central receiver/server. This hierarchical segmentation reduces the immediate bandwidth burden on the central system while maintaining comprehensive problem information collection across the entire network.
Solution Approach 2:
Machines perform preliminary evaluation of detected problems against predefined criteria before reporting to the receiver/server. This preliminary filtering action reduces the volume of reports transmitted while ensuring that only significant problems are communicated, maintaining both speed and information quality.
2Measurement precision
If machines continuously monitor and report all network conditions, then problem detection accuracy is high, but communication bandwidth and computing resources are consumed excessively
Solution Approach 1:
The system implements partial monitoring by having machines report only on problems that meet specific severity thresholds or criteria. This partial action approach maintains sufficient detection accuracy for critical issues while significantly reducing unnecessary communication overhead and resource consumption for minor or transient conditions.
Solution Approach 2:
The patent employs dynamic threshold parameters that can be adjusted based on network conditions and problem severity. By changing these parameters adaptively, the system optimizes the balance between detection accuracy and resource consumption, reporting detailed information when needed but using lighter monitoring modes during normal operation.
3Loss of time
If the receiver/server processes all incoming problem reports immediately, then response time is minimized, but processing capacity is overwhelmed
Solution Approach 1:
The patent introduces zone servers as intermediate processing nodes that handle problem reports locally before forwarding to the central receiver/server. This segmentation distributes the processing load across multiple nodes, enabling timely local responses while preventing the central system from being overwhelmed by volume alone.
Solution Approach 2:
Zone servers perform preliminary processing and triage of problem reports, filtering and pre-sorting data before it reaches the central receiver/server. This preliminary action reduces the processing burden on the central system while maintaining fast response times through local pre-processing of critical issues.
Data Source
AI summary
In a network, a set of machines communicate pairwise, each conditionally adjusting messages in response to its own local state, and each in response to statistical methods conditionally propagating those messages, with the effect that problems with that network, or with a subset of its machines, are reported to a receiver/server. Only a substantially constant number of reports are made to the receiver/server, even when there are a substantial number of such machines able to detect that problem. When a problem is reported, a similar technique causes the machines to collectively evaluate and report suggested causes for that problem. Messages are propagated from each machine to another using locally random global locality. The machines in the network, in response to statistical techniques, organize hierarchically in O(log n) time, where n is the number of machines in the network, substantially without any requirement for nonlocal message exchange.


