Storage Extent Failure Prediction and Data Rebuild

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

In storage systems with multiple devices in a resource pool, predicting and preventing failures is challenging due to varying wear degrees and service states, leading to potential data loss as existing solutions often require replacing entire devices even if only a part fails.

Innovation Solution

A method and apparatus that monitor service states and features of storage device extents, identify potential failure points using association relations, and rebuild data to free extents, allowing for finer granularity in failure management and reducing data loss.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If entire storage devices are replaced when failure occurs, then system reliability is maintained, but resource waste increases and cost rises

Engineering Contradiction:
Improvesystem reliabilityVSAvoidstorage capacity loss
Core Design Contradiction:
ReliabilityVSLoss of substance

Solution Approach 1:

The patent divides the storage device into multiple extents (logical blocks), allowing failure isolation at the extent level rather than replacing the entire device. When an extent fails, only that specific extent is marked as failed and excluded from service, while other extents remain operational. This segmentation enables fine-grained failure management and maximizes usable storage capacity.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent applies different quality treatment to different extents within a storage device based on their service states. Extents are monitored individually for service time, service state, and failure characteristics, allowing selective replacement or maintenance only of failed extents rather than uniform replacement of the entire device. This local quality approach optimizes resource utilization.

Inventive Principle:
Principle #3Local quality

2Ease of operation

If storage devices with different wear degrees are managed uniformly, then management simplicity is maintained, but failure prediction accuracy decreases

Engineering Contradiction:
Improvemanagement simplicityVSAvoidfailure prediction accuracy
Core Design Contradiction:
Ease of operationVSMeasurement precision

Solution Approach 1:

The patent implements dynamic monitoring and evaluation of storage device extents based on real-time service states. The system continuously collects data on service time, service state, and failure characteristics, then dynamically adjusts failure predictions and management strategies. This dynamic approach allows accurate failure prediction for extents with different wear degrees while maintaining manageable operations through automated evaluation.

Inventive Principle:
Principle #15Dynamics

Solution Approach 2:

The patent establishes a feedback mechanism where failure characteristics of extents are continuously monitored and fed back into the evaluation model. This feedback enables the system to refine its failure prediction accuracy by learning from actual failure patterns and adjusting its assessment of extents with different wear degrees, thereby improving prediction precision without complicating management.

Inventive Principle:
Principle #23Feedback

3Device complexity

If partial failure of storage device is ignored, then system complexity is reduced, but data loss risk increases

Engineering Contradiction:
Improvefailure management complexityVSAvoiddata safety
Core Design Contradiction:
Device complexityVSReliability

Solution Approach 1:

The patent performs preliminary identification and evaluation of potential failure extents before actual data loss occurs. By monitoring service states and failure characteristics in advance, the system can predict which extents are likely to fail and take preventive actions such as data migration or isolation. This preliminary action prevents data loss while managing complexity through automated prediction algorithms.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent introduces an intermediary evaluation mechanism that acts as a mediator between storage device extents and the failure management system. This intermediary layer (the evaluation model) processes raw service state data and failure characteristics, transforming them into actionable failure predictions. It simplifies the overall system by providing a standardized approach to handling partial failures while maintaining high data safety through proactive identification and management.

Inventive Principle:
Principle #24Intermediary (Mediator)

Data Source

PatentUS11036587B2Method, apparatus, and computer program product for managing storage system using partial drive failure prediction
Publication Date: 2021.06.15 EMC IP HLDG CO LLC
  • US11036587B2 patent drawing
  • US11036587B2 patent drawing
  • US11036587B2 patent drawing

AI summary

According to implementations of the present disclosure, there is provided a method for managing a storage system, extents in the storage system being from multiple storage devices in a resource pool associated with the storage system. In the method, regarding multiple extents comprised in a storage device among the multiple storage devices, respective service states of the multiple extents are obtained. Respective features of respective extents among the multiple extents are determined on the basis of respective service states of the multiple extents. An association relation between a failure in an extent in a storage device in the resource pool and a feature of the extent is obtained. A failure extent in which a failure is to be occurred is identified from the multiple extents on the basis of respective features of the multiple extents and the association relation.