IT Incident Time Prediction Using Correlated Performance Indicators
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Complex and distributed IT infrastructures face challenges in accurately detecting anomalies and predicting failures due to the variability of performance indicators and the complexity of interdependent hardware and software components, leading to delayed detection and overwhelming maintenance alerts.
Innovation Solution
A method and device that analyze performance indicator values to identify abnormal and risky indicators, using a correlation base to calculate an estimated duration before a technical incident, allowing for proactive maintenance and reducing false alerts by processing data in real-time and continuously.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If performance indicator monitoring is implemented to detect anomalies in IT infrastructure, then system reliability is improved, but the complexity of the monitoring system increases due to the large number of interconnected components
Solution Approach 1:
The monitoring system is segmented into multiple specialized modules: a collection module for gathering performance indicators, a data processing module for analyzing the indicators, and a correlation base for storing relationships between indicators. This segmentation allows each module to handle specific tasks independently, reducing overall system complexity while maintaining comprehensive monitoring capability.
Solution Approach 2:
A correlation base is introduced as an intermediary structure that stores pre-established relationships between performance indicators. This intermediary allows the system to leverage known correlations without requiring complex real-time analysis of all indicator relationships, simplifying the monitoring process while improving reliability through informed anomaly detection.
2Measurement precision
If direct use of performance indicator values is applied for anomaly detection, then measurement precision is maintained, but detection accuracy deteriorates due to great variability of observable values over time
Solution Approach 1:
The system performs preliminary actions by establishing a correlation base that stores relationships between performance indicators before actual anomaly detection occurs. This pre-processing of correlation data allows the system to contextualize raw performance indicator values, transforming precise but variable measurements into meaningful anomaly detections that account for temporal variations and inter-indicator relationships.
3Reliability
If comprehensive anomaly detection is implemented across all performance indicators, then reliability is improved, but loss of time increases due to the overwhelming number of maintenance alerts
Solution Approach 1:
The system implements feedback mechanisms where the correlation base is continuously updated with new correlation data derived from historical performance indicator relationships. This feedback loop allows the system to learn from past anomalies and improve its detection accuracy over time, reducing false alerts and enabling maintenance teams to respond more efficiently to genuine issues while maintaining high reliability.
Data Source
Figure 1
Figure 2
Figure 3
AI summary
The invention relates to a method and a device for determining an estimated time before a technical incident, said method comprising: - a step (120) of receiving values of performance indicators, - a step (140) of identifying performance indicators in anomaly, - a step (150) of identifying first indicators at risk, - a step (160) of identifying other indicators at risk, and - a step (170) of determining an estimated time before a technical incident comprising a calculation, from the indicators in anomaly and the identified indicators at risk, of a shortest path leading to a risk of a technical incident, and a calculation of an estimated time before a technical incident, said estimated time before a technical incident being calculated from the values of time before switching to anomaly between correlated performance indicators for each of the performance indicators constituting the shortest calculated path.