AI Incident Clustering for IT System Resolution
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Large IT organizations face inefficiencies in resolving recurring incidents across complex IT landscapes due to decentralized personnel and systems, leading to significant burdens and time wastage in identifying and addressing system performance changes, which can result in costly outages and reputation risks.
Innovation Solution
A machine-learning based method that clusters incident records using a rolling time window and Hierarchical Density-Based Spatial Clustering of Applications with Noise (HDBSCAN), creating a problem record, linking incident records to it, and updating the model to learn from resolutions, enabling automated identification and resolution of recurring incidents.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Productivity
If decentralized personnel and systems are used to resolve recurring incidents, then individual flexibility is maintained, but significant inefficiencies and time wastage occur
Solution Approach 1:
The patent merges multiple incident records into unified cluster tickets using machine learning algorithms. Incident clustering technology combines similar incidents that occur across different teams and systems, transforming decentralized incident handling into centralized cluster management. This reduces the number of separate tickets personnel must handle while maintaining the ability to address system-specific nuances.
Solution Approach 2:
The patent introduces an intermediary AI-powered clustering system between decentralized incident reporters and human resolution teams. This intermediary automatically groups related incidents, identifies patterns, and creates cluster tickets, reducing the burden on individual teams while preserving their ability to resolve specific incidents. The clustering system acts as a mediator that processes raw incident data before presenting it to human operators.
2Loss of time
If manual incident analysis is performed to determine cause of performance changes, then detailed investigation is possible, but time is lost and incidents cost money
Solution Approach 1:
The patent implements preliminary automated analysis of incident records using machine learning models before human investigation. The system pre-processes incident data, identifies potential causes, and groups similar incidents into clusters, providing investigators with pre-analyzed information. This preliminary action reduces the time required for manual analysis while maintaining accuracy through the combination of automated pattern recognition and human expertise.
Solution Approach 2:
The patent replaces manual mechanical analysis of incident records with automated machine learning-based analysis. The system uses algorithms to automatically identify patterns, correlations, and root causes across large volumes of incident data, substituting human manual investigation with computational analysis. This substitution dramatically reduces time loss while maintaining or improving measurement precision through consistent, scalable analysis.
3Productivity
If individual tickets are handled separately, then specific incident details are maintained, but significant inefficiencies occur in large IT organizations
Solution Approach 1:
The patent merges individual incident tickets into cluster tickets while preserving essential information from each incident. The clustering system maintains detailed records of constituent incidents within each cluster, allowing teams to see both the aggregated pattern and specific incident details. This merging approach improves productivity by reducing ticket volume while preventing information loss through structured data retention.
Solution Approach 2:
The patent segments incident information into hierarchical levels: cluster-level summaries that show overall patterns and individual incident-level details that preserve specific information. This segmentation allows the system to present aggregated views for productivity improvement while maintaining access to granular details when needed. The segmented structure enables efficient cluster management without sacrificing incident-specific information.
Data Source
AI summary
A method for identifying and handling related incidents using a machine-learning based model, includes, performing by one or more processors, operations including: clustering a sub-group of incident records from among a plurality of incident records using the machine-learning based model based on a rolling time window and a number of records in the sub-group of incident records; creating a problem record based on the clustered incident records; populating the problem record with information related to the clustered incident records; linking the clustered incident records to the problem record; providing a notification that the problem record has been created; receiving a resolution for the problem record; and updating, based on the resolution, the machine-learning based model to learn an association between extracted features of the resolution and extracted features of the clustered incident records.


