Storage Extent Failure Prediction and Data Rebuild
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
In storage systems with multiple devices in a resource pool, predicting and preventing failures is challenging due to varying wear degrees and service states, leading to potential data loss as existing solutions often require replacing entire devices even if only a part fails.
Innovation Solution
A method and apparatus that monitor service states and features of storage device extents, identify potential failure points using association relations, and rebuild data to free extents, allowing for finer granularity in failure management and reducing data loss.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If entire storage devices are replaced when failure occurs, then system reliability is maintained, but resource waste increases and cost rises
Solution Approach 1:
The patent divides the storage device into multiple extents (logical blocks), allowing failure isolation at the extent level rather than replacing the entire device. When an extent fails, only that specific extent is marked as failed and excluded from service, while other extents remain operational. This segmentation enables fine-grained failure management and maximizes usable storage capacity.
Solution Approach 2:
The patent applies different quality treatment to different extents within a storage device based on their service states. Extents are monitored individually for service time, service state, and failure characteristics, allowing selective replacement or maintenance only of failed extents rather than uniform replacement of the entire device. This local quality approach optimizes resource utilization.
2Ease of operation
If storage devices with different wear degrees are managed uniformly, then management simplicity is maintained, but failure prediction accuracy decreases
Solution Approach 1:
The patent implements dynamic monitoring and evaluation of storage device extents based on real-time service states. The system continuously collects data on service time, service state, and failure characteristics, then dynamically adjusts failure predictions and management strategies. This dynamic approach allows accurate failure prediction for extents with different wear degrees while maintaining manageable operations through automated evaluation.
Solution Approach 2:
The patent establishes a feedback mechanism where failure characteristics of extents are continuously monitored and fed back into the evaluation model. This feedback enables the system to refine its failure prediction accuracy by learning from actual failure patterns and adjusting its assessment of extents with different wear degrees, thereby improving prediction precision without complicating management.
3Device complexity
If partial failure of storage device is ignored, then system complexity is reduced, but data loss risk increases
Solution Approach 1:
The patent performs preliminary identification and evaluation of potential failure extents before actual data loss occurs. By monitoring service states and failure characteristics in advance, the system can predict which extents are likely to fail and take preventive actions such as data migration or isolation. This preliminary action prevents data loss while managing complexity through automated prediction algorithms.
Solution Approach 2:
The patent introduces an intermediary evaluation mechanism that acts as a mediator between storage device extents and the failure management system. This intermediary layer (the evaluation model) processes raw service state data and failure characteristics, transforming them into actionable failure predictions. It simplifies the overall system by providing a standardized approach to handling partial failures while maintaining high data safety through proactive identification and management.
Data Source
AI summary
According to implementations of the present disclosure, there is provided a method for managing a storage system, extents in the storage system being from multiple storage devices in a resource pool associated with the storage system. In the method, regarding multiple extents comprised in a storage device among the multiple storage devices, respective service states of the multiple extents are obtained. Respective features of respective extents among the multiple extents are determined on the basis of respective service states of the multiple extents. An association relation between a failure in an extent in a storage device in the resource pool and a feature of the extent is obtained. A failure extent in which a failure is to be occurred is identified from the multiple extents on the basis of respective features of the multiple extents and the association relation.


