Server Failure Prediction via Performance Metrics and Remedial Actions

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing approaches to maintaining a fleet of servers across multi-site/region datacenters are costly and lead to reduced capacity and service interruptions due to the time-consuming nature of maintenance actions, resulting in financial losses and loss of goodwill.

Innovation Solution

A method and system for predicting system component failures within a computer system using performance metrics, AI/ML models, and remedial actions to mitigate failures proactively, thereby improving system integrity and robustness.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If traditional maintenance approaches are used to monitor and repair server components, then system reliability is maintained, but maintenance time and service interruptions increase

Engineering Contradiction:
Improvesystem reliabilityVSAvoidmaintenance time
Core Design Contradiction:
ReliabilityVSLoss of time

Solution Approach 1:

The system performs preliminary actions by predicting component failures before they actually occur. The failure prediction module analyzes performance metrics and generates failure probabilities in advance, allowing maintenance to be scheduled proactively rather than reactively, thus reducing unplanned service interruptions and maintenance time.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The system implements continuous feedback loops where performance metrics are constantly monitored, analyzed, and used to update failure predictions. The remedial action module receives feedback from failure predictions and automatically executes appropriate maintenance actions, creating a closed-loop system that continuously improves reliability while minimizing maintenance interruptions.

Inventive Principle:
Principle #23Feedback

2Reliability

If frequent system checks and maintenance actions are performed, then component failures are detected earlier, but productivity and service capacity are reduced

Engineering Contradiction:
Improvecomponent failure detectionVSAvoidservice capacity
Core Design Contradiction:
ReliabilityVSProductivity

Solution Approach 1:

The system applies partial action by performing maintenance only on components that show elevated failure probabilities. Instead of checking and maintaining all components uniformly, the remedial action module targets specific components based on predicted failure risks, thus maintaining high detection capability while preserving service capacity for non-critical components.

Inventive Principle:
Principle #16Partial or excessive action

Solution Approach 2:

The system changes the parameter of maintenance frequency from a uniform schedule to a dynamic, risk-based schedule. Components with higher failure probabilities undergo more frequent monitoring and earlier maintenance, while components with low risk continue operating with minimal interruption, thus optimizing the balance between failure detection and productivity.

Inventive Principle:
Principle #35Parameter changes

3Ease of manufacture

If reactive maintenance is performed after failures occur, then maintenance costs are reduced, but service interruptions and financial losses increase

Engineering Contradiction:
Improvemaintenance costVSAvoidservice availability
Core Design Contradiction:
Ease of manufactureVSReliability

Solution Approach 1:

The system performs preliminary maintenance actions by predicting failures before they occur and executing remedial actions proactively. This prevents catastrophic failures and associated high costs of emergency repairs, while maintaining service availability through planned, non-disruptive maintenance execution.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The system skips the traditional reactive maintenance cycle by directly transitioning from prediction to preventive action. The automated remedial action module executes maintenance tasks before failures occur, eliminating the need for costly emergency responses and service interruptions associated with reactive maintenance.

Inventive Principle:
Principle #21Skipping (Rushing through)

Data Source

PatentUS20250199894A1Method and system for predicting server hardware and server hardware component failures
Publication Date: 2025.06.19 JPMORGAN CHASE BANK NA
  • US20250199894A1 patent drawing
  • US20250199894A1 patent drawing
  • US20250199894A1 patent drawing

AI summary

A system for predicting system component failures within a computer system that comprises a plurality of hardware components and a plurality of software components. The system may comprise memory storing instructions that, when executed, cause a processor to: obtain performance metrics by monitoring a network interface of the computer system; generate component failure probabilities by processing the performance metrics; determine that a first component failure probability among the component failure probabilities exceeds a risk threshold; determine remedial actions that mitigate a first component failure probability; and mitigate the first component failure probability by initiating an execution of the remedial actions.