Self-Learning Analytics for Datacenter Fault Management

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

In data centers, existing fault management solutions struggle to accurately diagnose hardware/firmware failures in newer systems due to the lack of historic failure data, often leading to ineffective error analysis and incorrect identification of faulty Field Replaceable Units (FRUs).

Innovation Solution

A self-learning analytics system that includes a self-learning manager and error analysis engine, which generates service events with prioritized cause and action recommendations, validates support engineer actions, and dynamically updates recommendations based on field data and failure rates, enabling closed-loop solutions for hardware/firmware fault management.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Extent of automation

If automated hardware/firmware analysis is performed using existing fault management solutions, then service events are generated with recommendation actions, but the accuracy of fault diagnosis deteriorates due to lack of historic failure data for new hardware/firmware

Engineering Contradiction:
Improveautomated hardware/firmware analysisVSAvoidfault diagnosis accuracy
Core Design Contradiction:
Extent of automationVSMeasurement precision

Solution Approach 1:

The system implements feedback by collecting actual support engineer actions and field failure data, then using this information to dynamically update and refine the error analysis algorithms. The self-learning manager continuously learns from real-world outcomes to improve future diagnostic accuracy, transforming static algorithms into adaptive systems that evolve with accumulated experience.

Inventive Principle:
Principle #23Feedback

Solution Approach 2:

The system enables self-service through self-learning algorithms that automatically improve their own performance without external intervention. The error analysis engine autonomously updates its diagnostic models by processing field data and support actions, allowing the system to self-optimize and adapt to new hardware/firmware configurations without requiring manual algorithm updates.

Inventive Principle:
Principle #25Self-service

2Ease of manufacture

If error analysis algorithms are developed based on past experience and data, then diagnostic recommendations can be generated, but the effectiveness deteriorates when applied to newer hardware/firmware products due to technological changes

Engineering Contradiction:
Improvealgorithm development based on past dataVSAvoideffectiveness on newer hardware/firmware
Core Design Contradiction:
Ease of manufactureVSAdaptability or versatility

Solution Approach 1:

The system transitions from static, historically-based algorithms to dynamic, self-learning algorithms that continuously adapt to new hardware/firmware configurations. The error analysis engine automatically updates its diagnostic models by learning from field data and support actions, enabling it to remain effective across evolving technology generations without requiring complete algorithm redesign.

Inventive Principle:
Principle #15Dynamics

Solution Approach 2:

The system uses feedback loops where actual field failure data and support engineer actions are continuously collected and used to refine the error analysis algorithms. This feedback mechanism allows the system to adapt to new hardware/firmware products by learning from real-world performance data, maintaining effectiveness despite technological changes.

Inventive Principle:
Principle #23Feedback

3Measurement precision

If manual support engineer intervention is used for complex error analysis spanning multiple subsystems, then detailed analysis can be performed, but the time and resource consumption increases

Engineering Contradiction:
Improvedetailed error analysis capabilityVSAvoidsupport engineer time consumption
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The system introduces a self-learning error analysis engine as an intermediary between the fault detection system and support engineers. This intermediary automatically performs preliminary analysis of complex errors spanning multiple subsystems, filtering and prioritizing cases that truly require human intervention. The system generates service events with prioritized recommendation actions, reducing the time support engineers need to spend on routine diagnostic tasks.

Inventive Principle:
Principle #24Intermediary (Mediator)

4Productivity

If field replaceable unit identification is performed without considering common system bus faults, then quick FRU replacement can be recommended, but incorrect identification occurs when the fault is actually in shared resources

Engineering Contradiction:
ImproveFRU replacement speedVSAvoidFRU identification accuracy
Core Design Contradiction:
ProductivityVSMeasurement precision

Solution Approach 1:

The system uses the self-learning error analysis engine as an intermediary that performs intelligent fault isolation before generating service events. The engine analyzes error patterns and contextual information to distinguish between FRU-specific faults and shared resource faults on common system buses. This intermediary analysis layer prevents incorrect FRU identification by considering the broader system context, reducing unnecessary FRU replacements while maintaining quick response times.

Inventive Principle:
Principle #24Intermediary (Mediator)

Data Source

PatentUS10157100B2Support action based self learning and analytics for datacenter device hardware/firmare fault management
Publication Date: 2018.12.18 HEWLETT PACKARD ENTERPRISE DEV LP
  • US10157100B2 patent drawing
  • US10157100B2 patent drawing
  • US10157100B2 patent drawing

AI summary

Techniques for support activity based self learning and analytics for datacenter device hardware/firmware fault management are described. In one example, a service event including a unique service event ID, a set of prioritized cause and support engineer actions/recommendations and associated unique support engineer action codes are generated for the support engineer upon detecting a hardware/firmware failure event. Support engineer actions taken by the support engineer upon completing the service event are then received. The support engineer actions are then analyzed using the set of prioritized cause and support engineer actions/recommendations and FRU configuration information before and after servicing the hardware/firmware failure. Any potential errors resulting from the support engineer actions are then determined and notified at real-time to the support engineer based on the outcome of the analysis. Any needed updates to the set of prioritized cause and support engineer actions/recommendations are then recommended.