Performance Indicator Correlation for Incident Lead-Time Prediction
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Complex and distributed computing infrastructures face challenges in detecting anomalies and predicting failures due to high variability in performance metrics, leading to late detection and overwhelming maintenance teams with false alarms, necessitating a method to determine an estimated duration before technical incidents from weak signals.
Innovation Solution
A method executed by a computing device with a data processing module and correlation base, which receives performance indicator values, identifies anomalous and at-risk indicators, and calculates an estimated duration before a technical incident by analyzing correlations and historical durations, enabling predictive maintenance and reducing response time.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If existing monitoring devices report anomalies directly, then anomaly detection is achieved, but maintenance teams are overwhelmed with false alarms and warnings
Solution Approach 1:
The patent introduces an intermediary processing layer between anomaly detection and maintenance teams. This layer analyzes multiple performance indicators, identifies correlations, and filters warnings to distinguish true anomalies from false alarms. The intermediary processes raw anomaly data through correlation analysis and patterns recognition before presenting filtered, prioritized information to maintenance teams, thereby reducing alarm fatigue while maintaining detection sensitivity.
2Measurement precision
If monitoring thresholds are made sensitive to detect early anomalies, then detection precision improves, but false alarms increase due to normal performance variability
Solution Approach 1:
The patent merges multiple performance indicator measurements and correlates them across different system components. Instead of relying on single-threshold alerts, the system combines data from multiple sources and analyzes their interrelationships. This merging approach allows the system to distinguish true anomalies (which affect multiple correlated indicators) from normal variability (which affects single indicators independently), thereby improving detection precision while reducing false alarms.
Solution Approach 2:
The system implements feedback loops where anomaly detection results are continuously refined based on correlation analysis and historical patterns. When potential anomalies are detected, the system feeds this information back through correlation validation and patterns matching to confirm or reject the anomaly. This feedback mechanism allows sensitive detection thresholds to be maintained while filtering out false alarms through iterative validation.
3Reliability
If maintenance teams respond to all anomalies immediately, then system reliability improves, but response time efficiency decreases due to prioritization challenges
Solution Approach 1:
The patent applies partial action by prioritizing maintenance responses based on anomaly severity and correlation strength. Instead of treating all anomalies equally, the system identifies and addresses only the most critical anomalies that have high correlation with system failures. This selective approach allows maintenance teams to focus on partial but high-impact actions first, improving system reliability for critical issues while avoiding waste of time on low-priority false alarms.
Data Source
AI summary
A method for determining an estimated duration before a technical incident in a computing infrastructure, executed by a computing device. The computing device including a data processing module, a storage module that stores in memory at least one correlation base between performance indicators, wherein the correlation base includes values of duration before becoming anomalous between correlated performance indicators, and a collection module. The method includes receiving performance indicator values, identifying anomalous performance indicators, identifying first and other at-risk indicators. The method includes determining an estimated duration before a technical incident including a calculation, from the anomalous indicators and at-risk indicators, a shorter path leading to a risk of technical incident, and a calculation of an estimated duration before a technical incident. The estimated duration is calculated from the values of duration before becoming anomalous between correlated performance indicators for each of the performance indicators constituting the shortest path calculated.


