Supercomputer Anomaly Detection via Predictive Sensor Analysis
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Supercomputers face challenges in real-time anomaly detection due to the complexity of their infrastructure and high-speed networks, leading to potential service interruptions and reduced reliability, which existing methods struggle to address effectively.
Innovation Solution
A method and system that utilize sensors to send statistical data to a maintenance assistance system, where a prediction algorithm forecasts future variations and a detection algorithm identifies anomalies in real-time, employing filtering and aggregation algorithms to optimize supercomputer performance and reliability.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If human analysis is used for monitoring supercomputer infrastructure, then operational simplicity is maintained, but response time to errors is too long resulting in service interruptions
Solution Approach 1:
The system implements self-service through automated anomaly detection algorithms that continuously monitor infrastructure parameters and automatically identify deviations from normal operation. The detection algorithm compares real-time sensor data against predicted values generated by the prediction algorithm, enabling the system to autonomously detect anomalies without human intervention, thus reducing response time while maintaining reliability
Solution Approach 2:
The patent replaces manual human analysis with computational algorithms. The prediction algorithm uses historical data to forecast future parameter values, and the detection algorithm automatically compares actual sensor readings against these predictions, substituting mechanical human monitoring with automated computational systems that operate continuously without fatigue or delay
2Measurement precision
If arbitrary parameters are included in anomaly detection, then comprehensive monitoring is achieved, but false predictions or anomaly detections increase
Solution Approach 1:
The system dynamically adjusts which parameters are monitored based on their relevance to actual anomalies. The prediction algorithm learns from historical data which parameters are meaningful predictors of future states, and the detection algorithm focuses comparison on these relevant parameters rather than all available sensors, thereby maintaining detection accuracy while reducing processing complexity through selective parameter usage
3Reliability
If all sensors are used for monitoring, then complete coverage is achieved, but processing time and computational load increase
Solution Approach 1:
The system extracts and processes only the essential subset of sensor data needed for effective anomaly detection. The prediction algorithm identifies which sensor parameters are most predictive of future system states, and the detection algorithm focuses computational resources on comparing these key parameters against predictions, thereby maintaining comprehensive monitoring coverage while significantly reducing processing time and computational load through selective data extraction
Data Source
Figure 1
Figure 2
Figure 3
AI summary
The invention relates to a method for aiding maintenance and optimization of a supercomputer which comprises the dispatching to a system for aiding maintenance by at least one sensor of a signal representative of statistical data of at least one calculation node of the supercomputer, prediction at regular intervals of the future variations of the statistical data on the basis of signals representative of the statistical data, dispatched by the sensor or sensors, the detection of anomalies of variations of the signals representative of the statistical data, dispatched by the sensor or sensors, with respect to the future variations predicted in the prediction step. The invention also relates to a system for aiding maintenance and optimization.