Disk Array Failure Prediction and Prevention

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current storage systems face challenges in preventing data unavailable (DU) and data lost (DL) events, which can lead to user dissatisfaction and increased support pressure, as existing methods are inadequate in detecting potential issues proactively and providing effective solutions, especially for hardware-related problems and large-capacity drives.

Innovation Solution

A method and device that collect data from disk arrays, analyze it for potential failure events using a knowledge database and matching algorithm, and generate reports to alert users of impending issues, enabling timely action to prevent DU and DL events by determining appropriate actions based on collected data and historical failure events.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If two storage pools are configured to work in active mode for high reliability, then system availability is improved, but the risk of simultaneous failure causing data unavailability increases

Engineering Contradiction:
Improvesystem availabilityVSAvoiddata unavailability risk
Core Design Contradiction:
ReliabilityVSObject-affected harmful factors

Solution Approach 1:

The system performs preliminary analysis of collected data to determine potential failure events before they occur. By proactively identifying disks at risk of failure and taking preventive actions (such as data migration or replacement), the system avoids the harmful effect of simultaneous failures that would cause data unavailability, while maintaining the high availability architecture.

Inventive Principle:
Principle #10Preliminary action

2Device complexity

If conventional RAID configuration is used without proactive failure detection, then system complexity is reduced, but data unavailable and data lost rates increase

Engineering Contradiction:
Improveconfiguration simplicityVSAvoiddata unavailable rate
Core Design Contradiction:
Device complexityVSReliability

Solution Approach 1:

The system implements a feedback mechanism by continuously collecting data from storage devices, analyzing this data to determine potential failure events, and providing alerts or automated responses. This feedback loop enables proactive detection of failing disks and triggers preventive actions, significantly reducing data unavailable and data lost rates while maintaining relatively simple conventional RAID configurations.

Inventive Principle:
Principle #23Feedback

3Reliability

If proactive failure detection and prevention mechanisms are implemented, then data unavailable and data lost rates are reduced, but system complexity and data processing requirements increase

Engineering Contradiction:
Improvedata lost rateVSAvoiddetection system complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The system employs self-service mechanisms where storage devices automatically provide diagnostic information and health status data without requiring external intervention. The analysis system processes this self-provided data to identify potential failures and trigger preventive actions, reducing the need for complex manual monitoring and intervention systems while achieving low data lost rates.

Inventive Principle:
Principle #25Self-service

Data Source

PatentUS11126501B2Method, device and program product for avoiding a fault event of a disk array
Publication Date: 2021.09.21 EMC IP HLDG CO LLC
  • US11126501B2 patent drawing
  • US11126501B2 patent drawing
  • US11126501B2 patent drawing

AI summary

Techniques involve avoiding a potential failure event on a disk array. Along these lines, data collected for a disk array are obtained. It is determined, based on the collected data, whether a potential failure event is to occur on the disk array. In response to determining that the potential failure event is to occur on the disk array, an action to be taken for the disk array is determined, to avoid occurrence of the potential failure event.