Server Failure Prediction via ML MTBF Analysis

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Maintaining scalability, availability, durability, and performance in distributed computing environments with heterogeneous computing resources is complex, especially as the number of devices increases, due to occasional failures and differences in devices such as disk drives, which complicates management and requires effective predictive maintenance.

Innovation Solution

A failure prediction subsystem uses machine learning algorithms and predictive models to determine the Mean Time Between Failures (MTBF) for computing resources by analyzing hardware and software metrics, allowing for proactive decommissioning and repair of failing components before they impact users.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Productivity

If the number of devices in distributed computing environments increases to improve scalability and performance, then computing power and storage capacity increase, but system complexity and difficulty of management increase

Engineering Contradiction:
Improvecomputing powerVSAvoidsystem complexity
Core Design Contradiction:
ProductivityVSDevice complexity

Solution Approach 1:

The patent segments the distributed computing system into multiple independent computing resources (virtual machines, storage devices, network devices) that can be individually monitored and managed. Each resource has its own health metrics and failure predictions, allowing the complex system to be broken down into manageable units while maintaining overall scalability.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The system performs preliminary actions by predicting failures before they occur using machine learning models. By analyzing historical data and current metrics, the system proactively identifies resources at risk of failure and schedules maintenance before actual failures impact system performance, thus managing complexity through advance planning.

Inventive Principle:
Principle #10Preliminary action

2Adaptability or versatility

If heterogeneous computing resources are used to improve adaptability and versatility, then system flexibility increases, but reliability and predictability of performance decrease

Engineering Contradiction:
Improvesystem flexibilityVSAvoidperformance predictability
Core Design Contradiction:
Adaptability or versatilityVSReliability

Solution Approach 1:

The patent monitors and analyzes multiple parameters (metrics) from heterogeneous computing resources including CPU utilization, memory usage, storage I/O, network traffic, and temperature. By tracking these varying parameters across different device types, the system maintains reliability through comprehensive monitoring while preserving the adaptability to use diverse hardware configurations.

Inventive Principle:
Principle #35Parameter changes

Solution Approach 2:

The system implements continuous feedback loops where metrics from heterogeneous resources are collected, analyzed by machine learning models, and used to adjust predictions and maintenance schedules. This feedback mechanism ensures reliable performance prediction across diverse hardware by constantly learning from actual resource behavior patterns.

Inventive Principle:
Principle #23Feedback

3Reliability

If proactive failure prediction and maintenance are implemented to improve reliability and reduce downtime, then system availability increases, but computational resources are consumed for analysis and model training

Engineering Contradiction:
Improvesystem availabilityVSAvoidcomputational resource consumption
Core Design Contradiction:
ReliabilityVSUse of energy by moving object

Solution Approach 1:

The patent applies partial action by focusing monitoring and prediction efforts on the most critical computing resources and failure modes. Rather than analyzing every possible metric for every resource equally, the system identifies and prioritizes key metrics and resources that have the greatest impact on system availability, thus reducing overall computational consumption while maintaining reliability.

Inventive Principle:
Principle #16Partial or excessive action

Solution Approach 2:

The system uses machine learning models that are trained on historical data and then copied/deployed to make predictions on current resources. Instead of performing complex real-time analysis on all resources simultaneously, pre-trained models are used to efficiently predict failures, reducing the computational burden during operation while maintaining high reliability through accurate predictions.

Inventive Principle:
Principle #26Copying

Data Source

PatentUS10613962B1Server failure predictive model
Publication Date: 2020.04.07 AMAZON TECH INC
  • US10613962B1 patent drawing
  • US10613962B1 patent drawing
  • US10613962B1 patent drawing

AI summary

A failure prediction subsystem may obtain metrics information from a set of servers, the metrics information including measurements of the operation of the set of servers. The failure prediction subsystem may then determine a mean time between failures for at least one server of the set of servers by at least providing a portion of the metrics information as an input to a machine learning algorithm. The machine learning algorithm may output a mean time between failure for at least one server of the set of servers and the failure prediction subsystem may determine if mitigating action should be taken based at least in part on the mean time between failure.