Service Event Correlation Graphs for Network Alarm Noise Reduction
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Large-scale networked systems face challenges in managing vast numbers of metrics and alarms, making it difficult to identify the most critical issues and understand dependencies, leading to overwhelming and ineffective monitoring and diagnosis of network health.
Innovation Solution
A computer-implemented method using a multiple-layer relational graph to correlate service events, comprising a configuration, observation, and learned layer to determine relationships between services and events, facilitating improved service event analysis and root cause prediction.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If engineers monitor all service events and alarms in large-scale networked systems, then complete visibility of network health is achieved, but the complexity of managing and analyzing the vast number of metrics becomes overwhelming
Solution Approach 1:
The patent segments the monitoring system into multiple layers: data collection layer, data processing layer, and presentation layer. Service events are collected from multiple sources, processed through correlation engines that apply filtering and aggregation rules, and then presented through unified dashboards. This segmentation allows complete monitoring coverage while managing complexity through structured processing stages.
Solution Approach 2:
The patent introduces intermediary components including event correlation engines, aggregation services, and normalization layers that sit between raw service events and the user interface. These intermediaries filter, correlate, and consolidate events before presentation, reducing the complexity burden on both the system and users while maintaining comprehensive monitoring capability.
2Reliability
If engineers respond to all service alarms, then potential issues are thoroughly addressed, but the time required to respond to each alarm increases due to alert noise
Solution Approach 1:
The patent applies partial action by selectively responding to only the most critical and relevant service events. Correlation engines identify patterns and group related events, allowing engineers to respond to consolidated alerts representing multiple underlying issues rather than individual events. This approach maintains thorough issue resolution while significantly reducing response time by focusing attention on priority events.
Solution Approach 2:
The patent implements feedback mechanisms where the system learns from engineer responses to alarms, adjusting correlation rules and alert priorities based on historical data. Frequently triggered alerts are consolidated more aggressively, and the system provides feedback loops that refine event correlation over time, reducing unnecessary alarm responses while ensuring critical issues are addressed.
3Measurement precision
If engineers analyze individual service events separately, then detailed examination of each event is possible, but the ability to understand relationships and dependencies between events is lost
Solution Approach 1:
The patent merges individual service event analysis with relationship context through correlation engines that identify and group related events. The system combines detailed event examination with contextual information about dependencies and relationships, presenting both granular event data and holistic relationship views simultaneously. This allows engineers to maintain precise event analysis while preserving understanding of event interconnections.
Data Source
AI summary
Examples described herein generally relate to receiving a query context for service events occurring on one or more networks, determining, based on the query context, a set of service events occurring on the one or more networks, querying multiple layers of a multiple-layer relational graph to determine one or more other service events having a defined relationship with the set of service events at one or more of the multiple layers, where the multiple layers include a configuration layer, an observation layer, and learned layer, defining relationships between services or service events, and indicating, via a user interface and in response to the query context, the one or more other service events.


