Automated Server Fault Analysis and Recovery System

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional data center management tools rely on manual methods, making it difficult to predict and proactively manage fault events, especially new and unpredictable hardware faults, which can lead to significant downtime and impact customer services.

Innovation Solution

An automated system and method for managing fault events in data centers, involving hardware fault event analysis, reporting, statistical data processing, and predictive analytics to identify fault sources, assess risks, and schedule recovery mechanisms, including instant or delayed repair based on recovery policies.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If manual fault management methods are used in conventional data center tools, then operational simplicity is maintained, but fault prediction capability and proactive management are insufficient

Engineering Contradiction:
Improvefault prediction capabilityVSAvoidmanagement system complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The system performs preliminary actions by continuously collecting hardware fault event data and performing statistical analysis to predict potential faults before they occur. The automated analytics engine processes historical fault data to identify patterns and generate predictions, enabling proactive fault management rather than reactive manual intervention.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The system implements self-service through automated fault detection, analysis, and prediction capabilities. The analytics engine autonomously processes hardware fault event data from multiple sources, performs statistical analysis, and generates predictions without requiring manual intervention, thereby improving reliability while managing complexity through automation.

Inventive Principle:
Principle #25Self-service

2Loss of time

If automated fault analysis and prediction systems are implemented, then fault prediction and proactive management improve, but system complexity and data processing requirements increase

Engineering Contradiction:
Improveserver downtimeVSAvoidautomated management system complexity
Core Design Contradiction:
Loss of timeVSDevice complexity

Solution Approach 1:

The system performs preliminary analysis of hardware fault events and predicts potential failures before they cause server downtime. By continuously monitoring and analyzing fault data, the system identifies patterns that indicate impending failures, allowing administrators to take preventive actions and reduce unexpected downtime.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The system implements feedback mechanisms by continuously collecting hardware fault event data, analyzing it through statistical methods, and using the results to improve future predictions. The automated analytics engine processes feedback from actual fault occurrences to refine its prediction models, thereby reducing downtime while managing complexity through iterative improvement.

Inventive Principle:
Principle #23Feedback

3Measurement precision

If comprehensive hardware fault event data is collected and analyzed, then prediction accuracy improves, but data processing time and computational resources increase

Engineering Contradiction:
Improvefault analysis accuracyVSAvoiddata processing time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The system extracts only the most relevant features and patterns from comprehensive hardware fault event data for analysis. The automated analytics engine identifies and extracts key indicators from large volumes of raw data, focusing computational resources on the most predictive elements rather than processing every detail, thereby maintaining accuracy while reducing processing time.

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The system changes parameters by transforming raw hardware fault event data into standardized statistical metrics and patterns. The analytics engine applies statistical analysis methods that convert complex raw data into meaningful predictive indicators, improving measurement precision while optimizing processing efficiency through parameter transformation.

Inventive Principle:
Principle #35Parameter changes

Data Source

PatentUS10761926B2Server hardware fault analysis and recovery
Publication Date: 2020.09.01 QUANTA COMPUTER INC
  • US10761926B2 patent drawing
  • US10761926B2 patent drawing
  • US10761926B2 patent drawing

AI summary

A method and system for automatically managing a fault event occurring in a datacenter system provided. The method includes collecting hardware fault event analysis corresponding with the hardware fault event. The hardware fault event analysis is organized into a report for a server device suffering from the hardware fault event. The method also includes processing statistical data received from the report for the server device. The method also includes performing hardware recovery based on the evaluated statistical data.