Storage System Reliability Estimation via Retrieval Point Modeling
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
The complexity of data protection techniques and their configuration parameters makes it difficult for system administrators to design storage systems that meet dependability goals, with unclear reliability and potentially excessive costs.
Innovation Solution
A method to estimate storage system reliability by modeling the design under a workload, determining retrieval points for failure scenarios, and calculating data loss time periods and recovery times to quantify dependability and costs.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If multiple data protection techniques are combined to improve reliability, then storage system reliability is improved, but device complexity increases
Solution Approach 1:
The patent segments the storage system into distinct protection layers (RAID for hardware failure, mirroring for site failure, backup for logical errors) and evaluates each independently. This segmentation allows administrators to understand and manage each protection technique's contribution to overall reliability without being overwhelmed by the combined complexity.
Solution Approach 2:
The patent introduces an intermediary evaluation framework that acts as a mediator between the complex multi-layer protection system and system administrators. This framework translates the complex interactions into understandable metrics (data loss time period, recovery time, availability) that administrators can use to make informed decisions.
2Reliability
If more retrieval points are maintained to reduce data loss time period, then storage system reliability is improved, but storage capacity requirements increase
Solution Approach 1:
The patent applies partial action by evaluating whether to maintain full copies or partial copies of retrieval points at each secondary storage node. This allows optimization of storage capacity while still achieving acceptable data loss time periods by retaining only the necessary portion of data at each location.
Solution Approach 2:
The patent enables dynamic adjustment of retrieval point retention parameters (retention count, retention window) based on evaluated reliability metrics and storage capacity constraints. This allows the system to optimize the balance between data loss time period and storage capacity requirements.
3Reliability
If synchronous mirroring is used to minimize recovery time, then availability is improved, but bandwidth consumption increases
Solution Approach 1:
The patent enables dynamic selection between synchronous and asynchronous mirroring modes based on evaluated recovery time requirements and available bandwidth. This dynamic approach allows the system to use synchronous mirroring only when absolutely necessary (minimal recovery time required) and asynchronous mirroring when bandwidth conservation is more important.
Solution Approach 2:
The patent evaluates mirroring operations as periodic actions with different characteristics (synchronous updates vs. background propagation). This allows optimization of bandwidth usage by scheduling asynchronous updates during periods of lower network utilization while still meeting recovery time objectives.
4Reliability
If comprehensive evaluation of all failure scenarios is performed to ensure dependability goals are met, then storage system reliability is improved, but computational complexity increases
Solution Approach 1:
The patent segments the evaluation process into independent components (modeling retrieval points, evaluating each failure scenario separately, calculating individual metrics). This segmentation reduces computational complexity by avoiding the need to evaluate all failure scenarios simultaneously while still ensuring comprehensive coverage of dependability goals.
Data Source
AI summary
An embodiment of a method of estimating storage system reliability begins with a first step of modeling a storage system design in operation under a workload to determine location of retrieval points. The retrieval points provide sources for primary storage recovery for a plurality of failure scenarios. The method continues with a second step of finding a most recent retrieval point relative to a target recovery time that is available for recovery for a particular failure scenario. In a third step, a difference between the target recovery time and a retrieval point creation time for the most recent retrieval point is determined. The difference indicates a data loss time period.


