Server Cluster Failure Prediction via Segmented Predictive Models
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current failure detection systems in computer clusters are inadequate as they treat clusters as monoliths, failing to analyze individual server states and resource consumption, leading to undetected failures and insufficient lead-time for prevention.
Innovation Solution
An automated predictive system that uses data mining and machine learning to detect operating conditions preceding device failures, providing proactive measures such as resource allocation modifications, load balancing, and maintenance schedules to prevent failures, and communicating potential crash profiles to monitor devices and distribute predictive models across local and remote systems.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Device complexity
If current failure detection systems treat clusters as monoliths, then system complexity is reduced, but measurement precision of individual server states deteriorates
Solution Approach 1:
The patent divides the cluster monitoring system into individual server-level monitoring units. Each server's operating state and resource consumption are analyzed separately through dedicated predictive models, enabling precise measurement of individual server conditions while maintaining overall system manageability through modular architecture.
2Measurement precision
If detailed individual server analysis is implemented, then measurement precision of failure conditions improves, but device complexity increases
Solution Approach 1:
The patent employs multiple copies of the same predictive model architecture deployed across different servers. Each server receives an instance of the predictive model that analyzes its specific operating state. This copying approach enables detailed individual analysis without increasing overall system complexity, as the same proven model structure is replicated and adapted to each server's context.
3Reliability
If predictive models are distributed across local and remote systems, then reliability of failure prediction improves, but device complexity increases
Solution Approach 1:
The predictive modeling function is segmented and distributed across multiple locations - local systems at each server and remote centralized systems. This segmentation allows the system to leverage both local real-time data processing and remote computational resources, improving prediction reliability through distributed intelligence while maintaining manageable complexity through clear functional separation.
Solution Approach 2:
The predictive model architecture is designed to be universal and multi-functional, operating effectively whether deployed locally at individual servers or remotely at centralized systems. This universal design enables the same core predictive logic to serve multiple deployment contexts, improving reliability through consistent prediction capabilities across distributed locations without proportionally increasing complexity.
Data Source
AI summary
A system and method predicts events in a computer system. The system and method includes a controller that receives a crash profile. The controller generates granular information that identifies data indicative of a potential server cluster failure in an enterprise system without needing to identify an originating cause of the potential server cluster failure. The system and method trains a model by sampling portions of a profile that may include directives and data indicative of a normal operating state and a conditioned preamble operating state. The system and method provides a trained model to a prediction engine. The system and method modifies an allocation of computing resources in response to the prediction of the potential server cluster failure by the prediction engine monitoring one or more servers of a server cluster.


