Resiliency Groups for Scalable Storage Data Protection
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
As storage systems expand, the probability of data loss due to multiple device failures grows exponentially, leading to poor scalability in data recovery and loss survivability, especially in multi-chassis clusters, where traditional redundancy methods fail to maintain effective data protection.
Innovation Solution
The formation of resiliency groups within storage systems, where data is written across a subset of blades, and the system dynamically reconfigures these groups in response to changes in geometry, such as adding or removing blades, to maintain stable failure probabilities and ensure data recovery capabilities.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If traditional redundancy methods are used in storage systems, then data protection is provided for small-scale systems, but data loss probability grows exponentially as the number of blades increases
Solution Approach 1:
The patent divides the storage system into multiple resiliency groups, where each group independently manages a subset of blades and data. This segmentation allows the system to scale by adding more resiliency groups without increasing the failure probability within each group, thereby resolving the contradiction between maintaining data protection reliability and achieving system scalability.
2Productivity
If the number of blades is increased to expand storage capacity, then storage system scalability is improved, but the probability of multiple device failures increases
Solution Approach 1:
By segmenting blades into resiliency groups where each group maintains independent redundancy, the system can add more blades (increasing storage capacity) without proportionally increasing the failure probability. Each resiliency group's failure probability remains bounded, allowing scalable expansion while maintaining reliability.
Solution Approach 2:
The patent changes the organizational parameter from a monolithic structure to a modular resiliency group structure. This parameter change allows the system to maintain constant failure probability characteristics within each group while scaling the total number of blades, thus resolving the contradiction between storage capacity and failure probability.
3Productivity
If data is distributed across all blades in the storage system, then storage efficiency is improved, but data recovery becomes complex and non-scalable
Solution Approach 1:
The patent segments data distribution into resiliency groups, where data is distributed across blades within each group rather than across the entire system. This segmentation simplifies data recovery by localizing it to individual resiliency groups, making the recovery process less complex and more scalable while maintaining storage efficiency.
Data Source
AI summary
A method of operating a plurality of blades of a storage system, performed by the storage system, is provided. The method includes writing data stripes across one or more sets of blades of the plurality of blades within resiliency groups, the plurality of blades having computing resources and storage memory, each resiliency group supporting data recovery in case of loss of a specified number of blades of the resiliency group. The method includes transferring data from a first resiliency group to a second resiliency group, responsive to a change in geometry of the storage system.


