Self-Organizing Event-Action Management for Large-Scale Networks
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Managing large-scale computing networks is challenging due to the complexity of analyzing vast volumes of metric and event data in a timely manner to prevent or control undesirable events, such as correlated failures, which existing technologies struggle to address effectively.
Innovation Solution
Implementing a self-organizing and evolving resource management system using a hierarchy of rule processing units combined with machine learning-based analytics, where rule processing units rapidly respond to events and metrics using propagated rules, and an analytics subsystem analyzes records to identify enhancements and modify configurations, allowing the system to evolve and improve over time.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If traditional monitoring and analysis methods are used for large-scale computing networks, then system stability is maintained, but the ability to analyze vast volumes of metric and event data in a timely manner deteriorates
Solution Approach 1:
The patent segments the monitoring system into multiple rule processing units distributed across the network infrastructure. Each unit independently processes specific rules and metrics, enabling parallel analysis of vast volumes of data without creating a single bottleneck, thus maintaining both high analysis capability and fast response time
Solution Approach 2:
The patent introduces rule processing units as intermediary components between data sources and analysis systems. These units pre-process and filter metrics and events locally before forwarding them upstream, reducing the volume of data requiring centralized analysis and enabling faster local responses to critical events
2Productivity
If virtualization technologies are deployed to share computing resources, then resource utilization improves, but the total number of logical computing elements to be managed increases
Solution Approach 1:
The patent implements universal rule processing units that can handle multiple types of metrics and events across different virtualization layers and computing elements. These units apply standardized rule sets to diverse resources, enabling efficient management of large numbers of virtual machines, containers, and physical hosts through a unified approach
Solution Approach 2:
The patent enables the monitoring system to automatically adapt to changing network configurations and resource allocations. Rule processing units dynamically discover new computing elements and begin monitoring them without manual intervention, allowing the system to self-manage as the infrastructure evolves and scales
3Reliability
If manual monitoring and response methods are used, then system simplicity is maintained, but the ability to prevent or rapidly control correlated failures deteriorates
Solution Approach 1:
The patent implements rule-based systems that proactively identify and respond to potential failure conditions before they escalate into correlated failures. Rules are pre-configured to detect patterns and trigger automated remediation actions, preventing failures rather than merely reacting to them after occurrence
Solution Approach 2:
The patent establishes closed-loop feedback mechanisms where rule processing units continuously monitor system state, evaluate metrics against defined rules, and automatically adjust configurations or trigger remediation actions. This automated feedback loop enables rapid response to changing conditions and prevents manual response delays from allowing failures to propagate
Data Source
AI summary
At a rule processing unit of an evolving, self-organized machine learning-based resource management service, a rule of a first rule set is applied to a value of a first collected metric, resulting in the initiation of a first corrective action. A set of metadata indicating the metric value and the corrective action is transmitted to a repository, and is used as part of an input data set for a machine learning model trained to generate rule modification recommendations. In response to determining that the corrective actions did not meet a success criterion, an escalation message is transmitted to another rule processing unit.


