Incident Prediction System Using Ticket Sequence Analysis
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing systems for predicting catastrophic incidents in IT service management lack mechanisms to identify catastrophic incidents and their probability of occurrence based on individual service patterns and do not consider the sequence of events for root cause identification across multiple services.
Innovation Solution
A method and system that utilize a machine learning-based approach to predict catastrophic incidents by analyzing historical and real-time tickets, identifying sequence rules, and generating risk scores to proactively alert IT management systems, thereby preventing disruptive incidents.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If existing techniques use log filtering, OPTICS, LSTM, and similarity score methods to predict system failure, then system failure prediction capability is improved, but the ability to identify catastrophic incidents and their probability of occurrence for multiple services based on individual patterns is insufficient
Solution Approach 1:
The patent segments the prediction system into multiple service-specific prediction models, each trained on individual service patterns. This allows the system to identify catastrophic incidents for multiple services based on their unique patterns rather than using a generic prediction approach, thereby improving measurement precision for catastrophic incident identification.
Solution Approach 2:
The patent applies local quality by creating service-specific prediction capabilities where each service has its own pattern analysis. This enables the system to tailor the prediction accuracy to the specific characteristics of each service, improving the overall precision of catastrophic incident identification across different service types.
2Ease of operation
If existing systems use reactive mechanisms to respond to high severity incidents after occurrence, then response to critical events is achieved, but system downtime increases and incurred costs increase
Solution Approach 1:
The patent implements preliminary action by predicting catastrophic incidents before they occur using pattern recognition and risk scoring. The system identifies sequences of events that precede high severity incidents and generates early warnings, enabling proactive measures to be taken before system failure occurs, thereby reducing system downtime and associated costs.
Solution Approach 2:
The patent uses feedback mechanisms by continuously monitoring ticket data and updating prediction models with new information. The system learns from historical patterns and adjusts risk scores based on evolving service behaviors, improving the accuracy of predictions and enabling more effective preventive actions over time.
3Measurement precision
If existing techniques analyze cluster time analysis and console logs to predict system failure, then basic prediction warnings are provided, but sequence of events for root cause identification is not considered
Solution Approach 1:
The patent adds the dimension of temporal sequencing by analyzing the sequence of events leading up to catastrophic incidents. Instead of only analyzing static cluster data, the system examines the chronological order of tickets and events, enabling root cause identification through pattern recognition in event sequences while maintaining prediction accuracy.
Data Source
AI summary
Disclosed herein an incident prediction system and a method for predicting catastrophic incidents in a service system. The system receives one or more tickets for one or more services along with information of the one or more tickets. Further, the system identifies a set of tickets for each of a predefined event window and a predefined non-event window from the one or more tickets and identifies a set of words for each of the set of tickets using a prediction model. Furthermore, the system identifies cluster data for each set of tickets from pre-defined clusters using a ticket clustering model and identifies presence of one or more sequence rules in cluster data. Finally, the system generates a risk score for each of the one or more services and root-causes for the set of tickets based on the one or more sequence rules, one or more parameters and creation time.


