Event Management in Distributed Systems via Alert Segmentation
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
In distributed processing systems, the overwhelming number of error and status reports from millions of devices makes it difficult for systems administrators to identify and manage issues effectively, as these reports become unhelpful and irrelevant due to their sheer volume.
Innovation Solution
Implementing an event and alert analysis module that receives, analyzes, and manages events and alerts by assigning them to pools based on time and attributes, suppressing unnecessary alerts, and executing actions to prioritize meaningful alerts for administrators.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If error and status reports are collected from millions of devices, then comprehensive monitoring coverage is improved, but the volume of reports becomes overwhelming and unmanageable
Solution Approach 1:
The patent segments the overwhelming volume of error and status reports by organizing them into hierarchical groups based on resource type, location, and error category. This segmentation allows administrators to navigate and manage reports in manageable segments rather than facing a monolithic deluge of information.
Solution Approach 2:
The patent extracts and highlights only the most critical and relevant error reports from the massive volume of incoming data. By filtering out routine or less significant status reports and focusing on high-priority errors, the system extracts the essential information that administrators need to act upon.
2Loss of information
If all error reports are presented to administrators, then complete information availability is improved, but administrator workload and decision-making difficulty increase
Solution Approach 1:
The patent applies local quality by providing different levels of information detail to different users or contexts. Critical errors receive prominent display with full details, while routine status reports are summarized or aggregated. This ensures that each piece of information is presented with the appropriate level of detail for its importance.
Solution Approach 2:
The patent implements partial action by selectively processing and presenting only the most relevant error reports to administrators rather than displaying all reports equally. This partial presentation of information reduces cognitive load while maintaining awareness of system health through strategic sampling and prioritization.
3Reliability
If automated error recovery is implemented across all resources, then system reliability is improved, but system complexity increases
Solution Approach 1:
The patent implements self-service by enabling automated error recovery mechanisms that allow resources to diagnose and correct their own errors without human intervention. Error detection, analysis, and recovery actions are performed automatically by the system itself, reducing the need for complex manual intervention procedures.
Solution Approach 2:
The patent applies preliminary action by pre-configuring error recovery policies and procedures for common error scenarios. Recovery actions are prepared and validated in advance, allowing the system to respond to errors automatically using pre-planned recovery strategies rather than requiring complex real-time decision-making.
Data Source
AI summary
Methods, systems, and computer program products for event management in a distributed processing system are provided. Embodiments include receiving, by the incident analyzer, one or more events from one or more resources, each event identifying a location of the resource producing the event; identifying, by the incident analyzer, an action in dependence upon the one or more events and the location of the one or more resources producing the one or more events; identifying, by the incident analyzer, a location scope for the action in dependence upon the one or more events; and executing, by the incident analyzer, the identified action.


