Adaptive IT Event Clustering for Incident Resolution Prediction
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing technologies struggle to effectively manage and predict IT incidents in complex, distributed, heterogeneous, and dynamically changing environments, leading to a flood of disparate event messages and overwhelming IT teams tasked with managing IT systems, resulting in a flood of disparate event messages and overwhelming IT teams tasked with managing IT systems, which are difficult to understand.
Innovation Solution
A real-time adaptive performance management system that provides contextual awareness to inform rapid response, reduces system downtime, accurately diagnoses and predicts issues with the most significant impact before incidents become bigger problems, and accurately diagnoses the system and predicts issues with the most significant impact before incidents become bigger problems.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If manual techniques and pre-programmed rules are used for event management, then IT teams can manage complex systems, but the labor intensity and cost increase significantly
Solution Approach 1:
The system enables automated self-service through machine learning models that automatically analyze events, identify patterns, and generate insights without human intervention. The adaptive model continuously learns from historical data and autonomously manages event correlation and prediction, reducing dependency on manual IT operations.
Solution Approach 2:
Manual mechanical processes of event analysis and pattern recognition are replaced with computational intelligence. Machine learning algorithms automatically process and correlate events, substituting human cognitive efforts with automated computational systems that scale efficiently.
2Loss of information
If monitoring systems are arrayed to provide visibility to operational metrics, then system visibility improves, but the sheer size and complexity result in a flooding of disparate event messages
Solution Approach 1:
The system merges disparate event messages from multiple monitoring sources into unified event clusters. By correlating events across different subsystems and time periods, the system combines scattered information into coherent patterns, reducing the apparent volume of individual events while preserving comprehensive operational visibility.
Solution Approach 2:
The machine learning model acts as an intermediary layer between raw monitoring data and operational insights. It mediates the complex flow of events by automatically filtering, correlating, and prioritizing messages, transforming the flooding of disparate events into manageable, actionable information.
3Extent of automation
If pre-programmed rules are used for event management, then some automation is achieved, but the ability to scale and evolve for future advances is limited
Solution Approach 1:
The system transitions from static pre-programmed rules to dynamic machine learning models that continuously adapt. The adaptive model evolves automatically by learning from new data patterns, enabling the system to scale and respond to emerging event types without requiring manual rule updates or reprogramming.
Solution Approach 2:
The system implements continuous feedback loops where model predictions and outcomes are fed back into the learning process. This feedback mechanism enables automatic refinement and adaptation of event management strategies, allowing the system to evolve with changing operational conditions and scale effectively.
Data Source
AI summary
Operations events received from different monitoring systems are transformed into a common event format having standardized fields including a source origin identifier, a source component identifier, a creation time, and event data. Frequency-time analysis is performed on the transformed Operations events by segmenting time-domain data into time bins and determining source origin counts for each time bin. Event clusters are identified based on time-frequency-space envelopes determined from the frequency-time analysis. Each event cluster is associated with resolution metrics including a time-to-resolve value and a number of responders required for resolution. A predictive model is trained using the event clusters and their associated resolution metrics. New incoming Operations events are grouped into a new event cluster. The trained predictive model is applied to the new event cluster to predict resolution requirements, and a remediation action is initiated based on the predicted resolution requirements.


