Anomaly Detection in Computing Infrastructure
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Complex and distributed computing infrastructures face challenges in detecting anomalies and predicting technical incidents due to high variability in performance indicators, leading to delayed or false alerts, overwhelming maintenance services and increasing the risk of system failures.
Innovation Solution
A method and device that determine a technical incident risk value by analyzing performance indicators, identifying anomalous and at-risk indicators, creating an augmented anomalies vector, and comparing it to reference data to anticipate potential failures, enabling proactive maintenance.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Ease of operation
If performance indicator values are used directly for anomaly detection, then the monitoring process is simple, but the detection threshold values become ineffective due to high variability over time
Solution Approach 1:
The patent applies preliminary action by pre-calculating and storing correlation data between performance indicators before actual anomaly detection occurs. This correlation base is built in advance using historical data, allowing the system to anticipate relationships between indicators and use them for more accurate anomaly detection without requiring complex real-time calculations.
Solution Approach 2:
The patent introduces an intermediary approach by using correlation data as a mediator between performance indicator values and anomaly detection. Instead of directly analyzing raw performance values against fixed thresholds, the system uses pre-computed correlation information to transform and contextualize these values, enabling more effective anomaly detection despite variability.
2Quantity of substance
If existing monitoring devices report only anomalies, then the monitoring output is concise, but maintenance services become overwhelmed with anomaly or failure warnings
Solution Approach 1:
The system performs preliminary analysis by calculating and storing correlation data in advance, which enables it to predict which performance indicators are likely to become anomalous before they actually do. This allows the system to proactively identify at-risk indicators and provide maintenance services with advance warning, reducing the volume of reactive alerts while maintaining high productivity.
Solution Approach 2:
The patent implements feedback mechanisms by continuously monitoring performance indicators, comparing them against correlation-based expectations, and adjusting anomaly detection thresholds dynamically. This feedback loop allows the system to learn from historical patterns and improve its prediction accuracy over time, reducing false alarms and optimizing maintenance service response.
3Measurement precision
If the system detects anomalies from weak signals, then the detection sensitivity increases, but the risk of false alarms increases
Solution Approach 1:
The system performs preliminary computation of correlation data using extensive historical performance indicator data, building a robust baseline of normal relationships between indicators. This pre-computed correlation information serves as a reference that enables the system to distinguish between genuine weak signals and normal variations, reducing false alarms while maintaining high sensitivity for actual anomalies.
Solution Approach 2:
The patent applies parameter changes by dynamically adjusting detection thresholds and correlation weights based on historical data patterns and system conditions. Instead of using fixed thresholds, the system adapts its sensitivity parameters over time, learning from past performance to optimize the balance between detecting weak signals and avoiding false alarms in changing operational environments.
Data Source
AI summary
The invention relates to a device and a method (100) for determining a technical incident risk value in an infrastructure (5), said method comprising:a step of receiving (120) performance indicator values,a step of identifying (140) anomalous performance indicators, so as to identify abnormal values, and identifying performance indicators associated with these abnormal values,a step of determining (150) at-risk indicators, comprising an identification of performance indicators of the computing infrastructure that are correlated with the identified anomalous indicators,a step of creating (160) an augmented anomalies vector, comprising the identifiers of the identified anomalous indicators and the identifiers of the determined at-risk indicators,a determination step (170), comprising the comparison of the augmented anomalies vector with predetermined technical incident reference data.


