Server Cluster Failure Prediction via Segmented Predictive Models

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current failure detection systems in computer clusters are inadequate as they treat clusters as monoliths, failing to analyze individual server states and resource consumption, leading to undetected failures and insufficient lead-time for prevention.

Innovation Solution

An automated predictive system that uses data mining and machine learning to detect operating conditions preceding device failures, providing proactive measures such as resource allocation modifications, load balancing, and maintenance schedules to prevent failures, and communicating potential crash profiles to monitor devices and distribute predictive models across local and remote systems.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Device complexity

If current failure detection systems treat clusters as monoliths, then system complexity is reduced, but measurement precision of individual server states deteriorates

Engineering Contradiction:
Improvesystem complexityVSAvoidmeasurement precision of individual server states
Core Design Contradiction:
Device complexityVSMeasurement precision

Solution Approach 1:

The patent divides the cluster monitoring system into individual server-level monitoring units. Each server's operating state and resource consumption are analyzed separately through dedicated predictive models, enabling precise measurement of individual server conditions while maintaining overall system manageability through modular architecture.

Inventive Principle:
Principle #1Segmentation

2Measurement precision

If detailed individual server analysis is implemented, then measurement precision of failure conditions improves, but device complexity increases

Engineering Contradiction:
Improvemeasurement precision of failure conditionsVSAvoiddevice complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent employs multiple copies of the same predictive model architecture deployed across different servers. Each server receives an instance of the predictive model that analyzes its specific operating state. This copying approach enables detailed individual analysis without increasing overall system complexity, as the same proven model structure is replicated and adapted to each server's context.

Inventive Principle:
Principle #26Copying

3Reliability

If predictive models are distributed across local and remote systems, then reliability of failure prediction improves, but device complexity increases

Engineering Contradiction:
Improvereliability of failure predictionVSAvoiddevice complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The predictive modeling function is segmented and distributed across multiple locations - local systems at each server and remote centralized systems. This segmentation allows the system to leverage both local real-time data processing and remote computational resources, improving prediction reliability through distributed intelligence while maintaining manageable complexity through clear functional separation.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The predictive model architecture is designed to be universal and multi-functional, operating effectively whether deployed locally at individual servers or remotely at centralized systems. This universal design enables the same core predictive logic to serve multiple deployment contexts, improving reliability through consistent prediction capabilities across distributed locations without proportionally increasing complexity.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Data Source

PatentUS10838791B1Robust event prediction
Publication Date: 2020.11.17 PROGRESSIVE CASUALTY INSURANCE CO
  • US10838791B1 patent drawing
  • US10838791B1 patent drawing
  • US10838791B1 patent drawing

AI summary

A system and method predicts events in a computer system. The system and method includes a controller that receives a crash profile. The controller generates granular information that identifies data indicative of a potential server cluster failure in an enterprise system without needing to identify an originating cause of the potential server cluster failure. The system and method trains a model by sampling portions of a profile that may include directives and data indicative of a normal operating state and a conditioned preamble operating state. The system and method provides a trained model to a prediction engine. The system and method modifies an allocation of computing resources in response to the prediction of the potential server cluster failure by the prediction engine monitoring one or more servers of a server cluster.