Resiliency Groups Stabilize Data Loss in Scalable Storage Systems
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
The probability of data loss due to multiple blade failures in storage systems increases exponentially with the number of blades, leading to poor scalability of data recovery and loss survivability as storage systems expand, particularly in multi-chassis clusters.
Innovation Solution
Forming resiliency groups within storage systems, where each group has a specified subset of blades, and reorganizing these groups when changes occur in the cluster geometry to maintain a stable probability of failure, using non-volatile random-access memory (NVRAM) for efficient data management and erasure coding to ensure data integrity across multiple storage devices.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Quantity of substance
If the number of blades in a storage system is increased to improve storage capacity, then the storage capacity is improved, but the probability of data loss due to multiple blade failures increases exponentially
Solution Approach 1:
The patent divides the storage system into multiple resiliency groups, where each group contains a subset of blades configured with erasure coding (e.g., N+2 redundancy). This segmentation isolates failure domains, ensuring that failures in one group do not propagate to other groups, thereby maintaining system-wide reliability even as total blade count increases.
Solution Approach 2:
The patent changes the redundancy parameter from a global N+2 configuration to localized resiliency groups with specified subset configurations. By adjusting the scope and composition of redundancy groups dynamically, the system maintains stable failure probability parameters while scaling storage capacity through increased blade count.
2Adaptability or versatility
If more blades are added to expand the storage system, then the storage system scalability is improved, but the data recovery effectiveness deteriorates
Solution Approach 1:
By organizing blades into discrete resiliency groups with defined boundaries, the system enables independent data recovery operations within each group. This segmentation allows scalable expansion where new groups can be added without impacting recovery complexity in existing groups, maintaining effective data recovery despite increased system scale.
Solution Approach 2:
The system pre-configures resiliency groups with erasure coding redundancy before failures occur. This preliminary organization of data across N+2 blades within each group ensures that recovery operations can proceed efficiently using pre-established redundancy relationships, regardless of the total number of blades in the expanded system.
3Reliability
If the storage system is configured with N+2 redundancy to survive two blade failures, then the loss survivability is improved, but the device complexity increases
Solution Approach 1:
The patent segments the complex N+2 redundancy configuration into manageable resiliency groups, where each group independently implements erasure coding. This segmentation simplifies the overall system complexity by breaking down the global redundancy management into localized group configurations, making the system more tractable while maintaining dual-blade-failure survivability.
4Stability of the object's composition
If resiliency groups are reorganized when cluster geometry changes, then the stability of failure probability is improved, but the system operation complexity increases
Solution Approach 1:
The system implements dynamic resiliency group reorganization that automatically adapts to cluster geometry changes while maintaining stable failure probability characteristics. The dynamic adjustment of group compositions in response to blade additions, removals, or failures enables the system to preserve reliability statistics without manual intervention, balancing operational simplicity with statistical stability.
Data Source
AI summary
A method of operating a storage system, and related storage system, are provided. The storage system establishes resiliency groups, each having a defined level of redundancy of resources of the storage system. The resiliency groups include at least one compute resources resiliency group and at least one storage resources resiliency group. The storage system supports capability of configurations that have multiples of each of the resiliency groups. Blades of the storage system perform distributed data and metadata storage across modular storage devices, in accordance with the resiliency groups.


