Self-Organizing Event-Action Management for Large-Scale Networks

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Managing large-scale computing networks is challenging due to the complexity of analyzing vast volumes of metric and event data in a timely manner to prevent or control undesirable events, such as correlated failures, which existing technologies struggle to address effectively.

Innovation Solution

Implementing a self-organizing and evolving resource management system using a hierarchy of rule processing units combined with machine learning-based analytics, where rule processing units rapidly respond to events and metrics using propagated rules, and an analytics subsystem analyzes records to identify enhancements and modify configurations, allowing the system to evolve and improve over time.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If traditional monitoring and analysis methods are used for large-scale computing networks, then system stability is maintained, but the ability to analyze vast volumes of metric and event data in a timely manner deteriorates

Engineering Contradiction:
Improvedata analysis capabilityVSAvoidresponse time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The patent segments the monitoring system into multiple rule processing units distributed across the network infrastructure. Each unit independently processes specific rules and metrics, enabling parallel analysis of vast volumes of data without creating a single bottleneck, thus maintaining both high analysis capability and fast response time

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent introduces rule processing units as intermediary components between data sources and analysis systems. These units pre-process and filter metrics and events locally before forwarding them upstream, reducing the volume of data requiring centralized analysis and enabling faster local responses to critical events

Inventive Principle:
Principle #24Intermediary (Mediator)

2Productivity

If virtualization technologies are deployed to share computing resources, then resource utilization improves, but the total number of logical computing elements to be managed increases

Engineering Contradiction:
Improveresource utilizationVSAvoidnumber of managed elements
Core Design Contradiction:
ProductivityVSDevice complexity

Solution Approach 1:

The patent implements universal rule processing units that can handle multiple types of metrics and events across different virtualization layers and computing elements. These units apply standardized rule sets to diverse resources, enabling efficient management of large numbers of virtual machines, containers, and physical hosts through a unified approach

Inventive Principle:
Principle #6Universality (Multi-functionality)

Solution Approach 2:

The patent enables the monitoring system to automatically adapt to changing network configurations and resource allocations. Rule processing units dynamically discover new computing elements and begin monitoring them without manual intervention, allowing the system to self-manage as the infrastructure evolves and scales

Inventive Principle:
Principle #25Self-service

3Reliability

If manual monitoring and response methods are used, then system simplicity is maintained, but the ability to prevent or rapidly control correlated failures deteriorates

Engineering Contradiction:
Improvefailure prevention capabilityVSAvoidautomated response level
Core Design Contradiction:
ReliabilityVSExtent of automation

Solution Approach 1:

The patent implements rule-based systems that proactively identify and respond to potential failure conditions before they escalate into correlated failures. Rules are pre-configured to detect patterns and trigger automated remediation actions, preventing failures rather than merely reacting to them after occurrence

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent establishes closed-loop feedback mechanisms where rule processing units continuously monitor system state, evaluate metrics against defined rules, and automatically adjust configurations or trigger remediation actions. This automated feedback loop enables rapid response to changing conditions and prevents manual response delays from allowing failures to propagate

Inventive Principle:
Principle #23Feedback

Data Source

PatentUS11855849B1Artificial intelligence based self-organizing event-action management system for large-scale networks
Publication Date: 2023.12.26 AMAZON TECH INC
  • US11855849B1 patent drawing
  • US11855849B1 patent drawing
  • US11855849B1 patent drawing

AI summary

At a rule processing unit of an evolving, self-organized machine learning-based resource management service, a rule of a first rule set is applied to a value of a first collected metric, resulting in the initiation of a first corrective action. A set of metadata indicating the metric value and the corrective action is transmitted to a repository, and is used as part of an input data set for a machine learning model trained to generate rule modification recommendations. In response to determining that the corrective actions did not meet a success criterion, an escalation message is transmitted to another rule processing unit.