Supercomputer Anomaly Detection via Predictive Sensor Analysis

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Supercomputers face challenges in real-time anomaly detection due to the complexity of their infrastructure and high-speed networks, leading to potential service interruptions and reduced reliability, which existing methods struggle to address effectively.

Innovation Solution

A method and system that utilize sensors to send statistical data to a maintenance assistance system, where a prediction algorithm forecasts future variations and a detection algorithm identifies anomalies in real-time, employing filtering and aggregation algorithms to optimize supercomputer performance and reliability.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If human analysis is used for monitoring supercomputer infrastructure, then operational simplicity is maintained, but response time to errors is too long resulting in service interruptions

Engineering Contradiction:
Improveservice continuityVSAvoidresponse time to errors
Core Design Contradiction:
ReliabilityVSLoss of time

Solution Approach 1:

The system implements self-service through automated anomaly detection algorithms that continuously monitor infrastructure parameters and automatically identify deviations from normal operation. The detection algorithm compares real-time sensor data against predicted values generated by the prediction algorithm, enabling the system to autonomously detect anomalies without human intervention, thus reducing response time while maintaining reliability

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The patent replaces manual human analysis with computational algorithms. The prediction algorithm uses historical data to forecast future parameter values, and the detection algorithm automatically compares actual sensor readings against these predictions, substituting mechanical human monitoring with automated computational systems that operate continuously without fatigue or delay

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

2Measurement precision

If arbitrary parameters are included in anomaly detection, then comprehensive monitoring is achieved, but false predictions or anomaly detections increase

Engineering Contradiction:
Improveanomaly detection accuracyVSAvoiddata processing complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The system dynamically adjusts which parameters are monitored based on their relevance to actual anomalies. The prediction algorithm learns from historical data which parameters are meaningful predictors of future states, and the detection algorithm focuses comparison on these relevant parameters rather than all available sensors, thereby maintaining detection accuracy while reducing processing complexity through selective parameter usage

Inventive Principle:
Principle #35Parameter changes

3Reliability

If all sensors are used for monitoring, then complete coverage is achieved, but processing time and computational load increase

Engineering Contradiction:
Improvemonitoring coverageVSAvoidprocessing speed
Core Design Contradiction:
ReliabilityVSProductivity

Solution Approach 1:

The system extracts and processes only the essential subset of sensor data needed for effective anomaly detection. The prediction algorithm identifies which sensor parameters are most predictive of future system states, and the detection algorithm focuses computational resources on comparing these key parameters against predictions, thereby maintaining comprehensive monitoring coverage while significantly reducing processing time and computational load through selective data extraction

Inventive Principle:
Principle #2Taking out (Extraction)

Data Source

PatentEP3380942B1Method and system for aiding maintenance and optimization of a supercomputer
Publication Date: 2023.02.15 BULL SA
  • EP3380942B1 patent drawingFigure 1
  • EP3380942B1 patent drawingFigure 2
  • EP3380942B1 patent drawingFigure 3

AI summary

The invention relates to a method for aiding maintenance and optimization of a supercomputer which comprises the dispatching to a system for aiding maintenance by at least one sensor of a signal representative of statistical data of at least one calculation node of the supercomputer, prediction at regular intervals of the future variations of the statistical data on the basis of signals representative of the statistical data, dispatched by the sensor or sensors, the detection of anomalies of variations of the signals representative of the statistical data, dispatched by the sensor or sensors, with respect to the future variations predicted in the prediction step. The invention also relates to a system for aiding maintenance and optimization.