Resiliency Groups Stabilize Data Loss in Scalable Storage Systems

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

The probability of data loss due to multiple blade failures in storage systems increases exponentially with the number of blades, leading to poor scalability of data recovery and loss survivability as storage systems expand, particularly in multi-chassis clusters.

Innovation Solution

Forming resiliency groups within storage systems, where each group has a specified subset of blades, and reorganizing these groups when changes occur in the cluster geometry to maintain a stable probability of failure, using non-volatile random-access memory (NVRAM) for efficient data management and erasure coding to ensure data integrity across multiple storage devices.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Quantity of substance

If the number of blades in a storage system is increased to improve storage capacity, then the storage capacity is improved, but the probability of data loss due to multiple blade failures increases exponentially

Engineering Contradiction:
Improvestorage capacityVSAvoiddata loss probability
Core Design Contradiction:
Quantity of substanceVSReliability

Solution Approach 1:

The patent divides the storage system into multiple resiliency groups, where each group contains a subset of blades configured with erasure coding (e.g., N+2 redundancy). This segmentation isolates failure domains, ensuring that failures in one group do not propagate to other groups, thereby maintaining system-wide reliability even as total blade count increases.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent changes the redundancy parameter from a global N+2 configuration to localized resiliency groups with specified subset configurations. By adjusting the scope and composition of redundancy groups dynamically, the system maintains stable failure probability parameters while scaling storage capacity through increased blade count.

Inventive Principle:
Principle #35Parameter changes

2Adaptability or versatility

If more blades are added to expand the storage system, then the storage system scalability is improved, but the data recovery effectiveness deteriorates

Engineering Contradiction:
Improvesystem scalabilityVSAvoiddata recovery effectiveness
Core Design Contradiction:
Adaptability or versatilityVSReliability

Solution Approach 1:

By organizing blades into discrete resiliency groups with defined boundaries, the system enables independent data recovery operations within each group. This segmentation allows scalable expansion where new groups can be added without impacting recovery complexity in existing groups, maintaining effective data recovery despite increased system scale.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The system pre-configures resiliency groups with erasure coding redundancy before failures occur. This preliminary organization of data across N+2 blades within each group ensures that recovery operations can proceed efficiently using pre-established redundancy relationships, regardless of the total number of blades in the expanded system.

Inventive Principle:
Principle #10Preliminary action

3Reliability

If the storage system is configured with N+2 redundancy to survive two blade failures, then the loss survivability is improved, but the device complexity increases

Engineering Contradiction:
Improveloss survivabilityVSAvoidredundancy configuration complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The patent segments the complex N+2 redundancy configuration into manageable resiliency groups, where each group independently implements erasure coding. This segmentation simplifies the overall system complexity by breaking down the global redundancy management into localized group configurations, making the system more tractable while maintaining dual-blade-failure survivability.

Inventive Principle:
Principle #1Segmentation

4Stability of the object's composition

If resiliency groups are reorganized when cluster geometry changes, then the stability of failure probability is improved, but the system operation complexity increases

Engineering Contradiction:
Improvefailure probability stabilityVSAvoidsystem reorganization complexity
Core Design Contradiction:
Stability of the object's compositionVSEase of operation

Solution Approach 1:

The system implements dynamic resiliency group reorganization that automatically adapts to cluster geometry changes while maintaining stable failure probability characteristics. The dynamic adjustment of group compositions in response to blade additions, removals, or failures enables the system to preserve reliability statistics without manual intervention, balancing operational simplicity with statistical stability.

Inventive Principle:
Principle #15Dynamics

Data Source

PatentUS11782625B2Heterogeneity supportive resiliency groups
Publication Date: 2023.10.10 PURE STORAGE INC
  • US11782625B2 patent drawing
  • US11782625B2 patent drawing
  • US11782625B2 patent drawing

AI summary

A method of operating a storage system, and related storage system, are provided. The storage system establishes resiliency groups, each having a defined level of redundancy of resources of the storage system. The resiliency groups include at least one compute resources resiliency group and at least one storage resources resiliency group. The storage system supports capability of configurations that have multiples of each of the resiliency groups. Blades of the storage system perform distributed data and metadata storage across modular storage devices, in accordance with the resiliency groups.