Self-Learning Analytics for Datacenter Fault Management
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
In data centers, existing fault management solutions struggle to accurately diagnose hardware/firmware failures in newer systems due to the lack of historic failure data, often leading to ineffective error analysis and incorrect identification of faulty Field Replaceable Units (FRUs).
Innovation Solution
A self-learning analytics system that includes a self-learning manager and error analysis engine, which generates service events with prioritized cause and action recommendations, validates support engineer actions, and dynamically updates recommendations based on field data and failure rates, enabling closed-loop solutions for hardware/firmware fault management.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Extent of automation
If automated hardware/firmware analysis is performed using existing fault management solutions, then service events are generated with recommendation actions, but the accuracy of fault diagnosis deteriorates due to lack of historic failure data for new hardware/firmware
Solution Approach 1:
The system implements feedback by collecting actual support engineer actions and field failure data, then using this information to dynamically update and refine the error analysis algorithms. The self-learning manager continuously learns from real-world outcomes to improve future diagnostic accuracy, transforming static algorithms into adaptive systems that evolve with accumulated experience.
Solution Approach 2:
The system enables self-service through self-learning algorithms that automatically improve their own performance without external intervention. The error analysis engine autonomously updates its diagnostic models by processing field data and support actions, allowing the system to self-optimize and adapt to new hardware/firmware configurations without requiring manual algorithm updates.
2Ease of manufacture
If error analysis algorithms are developed based on past experience and data, then diagnostic recommendations can be generated, but the effectiveness deteriorates when applied to newer hardware/firmware products due to technological changes
Solution Approach 1:
The system transitions from static, historically-based algorithms to dynamic, self-learning algorithms that continuously adapt to new hardware/firmware configurations. The error analysis engine automatically updates its diagnostic models by learning from field data and support actions, enabling it to remain effective across evolving technology generations without requiring complete algorithm redesign.
Solution Approach 2:
The system uses feedback loops where actual field failure data and support engineer actions are continuously collected and used to refine the error analysis algorithms. This feedback mechanism allows the system to adapt to new hardware/firmware products by learning from real-world performance data, maintaining effectiveness despite technological changes.
3Measurement precision
If manual support engineer intervention is used for complex error analysis spanning multiple subsystems, then detailed analysis can be performed, but the time and resource consumption increases
Solution Approach 1:
The system introduces a self-learning error analysis engine as an intermediary between the fault detection system and support engineers. This intermediary automatically performs preliminary analysis of complex errors spanning multiple subsystems, filtering and prioritizing cases that truly require human intervention. The system generates service events with prioritized recommendation actions, reducing the time support engineers need to spend on routine diagnostic tasks.
4Productivity
If field replaceable unit identification is performed without considering common system bus faults, then quick FRU replacement can be recommended, but incorrect identification occurs when the fault is actually in shared resources
Solution Approach 1:
The system uses the self-learning error analysis engine as an intermediary that performs intelligent fault isolation before generating service events. The engine analyzes error patterns and contextual information to distinguish between FRU-specific faults and shared resource faults on common system buses. This intermediary analysis layer prevents incorrect FRU identification by considering the broader system context, reducing unnecessary FRU replacements while maintaining quick response times.
Data Source
AI summary
Techniques for support activity based self learning and analytics for datacenter device hardware/firmware fault management are described. In one example, a service event including a unique service event ID, a set of prioritized cause and support engineer actions/recommendations and associated unique support engineer action codes are generated for the support engineer upon detecting a hardware/firmware failure event. Support engineer actions taken by the support engineer upon completing the service event are then received. The support engineer actions are then analyzed using the set of prioritized cause and support engineer actions/recommendations and FRU configuration information before and after servicing the hardware/firmware failure. Any potential errors resulting from the support engineer actions are then determined and notified at real-time to the support engineer based on the outcome of the analysis. Any needed updates to the set of prioritized cause and support engineer actions/recommendations are then recommended.


