Resiliency Groups for Scalable Storage Data Protection

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

As storage systems expand, the probability of data loss due to multiple device failures grows exponentially, leading to poor scalability in data recovery and loss survivability, especially in multi-chassis clusters, where traditional redundancy methods fail to maintain effective data protection.

Innovation Solution

The formation of resiliency groups within storage systems, where data is written across a subset of blades, and the system dynamically reconfigures these groups in response to changes in geometry, such as adding or removing blades, to maintain stable failure probabilities and ensure data recovery capabilities.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If traditional redundancy methods are used in storage systems, then data protection is provided for small-scale systems, but data loss probability grows exponentially as the number of blades increases

Engineering Contradiction:
Improvedata protectionVSAvoidscalability
Core Design Contradiction:
ReliabilityVSProductivity

Solution Approach 1:

The patent divides the storage system into multiple resiliency groups, where each group independently manages a subset of blades and data. This segmentation allows the system to scale by adding more resiliency groups without increasing the failure probability within each group, thereby resolving the contradiction between maintaining data protection reliability and achieving system scalability.

Inventive Principle:
Principle #1Segmentation

2Productivity

If the number of blades is increased to expand storage capacity, then storage system scalability is improved, but the probability of multiple device failures increases

Engineering Contradiction:
Improvestorage capacityVSAvoidfailure probability
Core Design Contradiction:
ProductivityVSReliability

Solution Approach 1:

By segmenting blades into resiliency groups where each group maintains independent redundancy, the system can add more blades (increasing storage capacity) without proportionally increasing the failure probability. Each resiliency group's failure probability remains bounded, allowing scalable expansion while maintaining reliability.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent changes the organizational parameter from a monolithic structure to a modular resiliency group structure. This parameter change allows the system to maintain constant failure probability characteristics within each group while scaling the total number of blades, thus resolving the contradiction between storage capacity and failure probability.

Inventive Principle:
Principle #35Parameter changes

3Productivity

If data is distributed across all blades in the storage system, then storage efficiency is improved, but data recovery becomes complex and non-scalable

Engineering Contradiction:
Improvestorage efficiencyVSAvoiddata recovery complexity
Core Design Contradiction:
ProductivityVSDevice complexity

Solution Approach 1:

The patent segments data distribution into resiliency groups, where data is distributed across blades within each group rather than across the entire system. This segmentation simplifies data recovery by localizing it to individual resiliency groups, making the recovery process less complex and more scalable while maintaining storage efficiency.

Inventive Principle:
Principle #1Segmentation

Data Source

PatentUS11138103B1Resiliency groups
Publication Date: 2021.10.05 PURE STORAGE INC
  • US11138103B1 patent drawing
  • US11138103B1 patent drawing
  • US11138103B1 patent drawing

AI summary

A method of operating a plurality of blades of a storage system, performed by the storage system, is provided. The method includes writing data stripes across one or more sets of blades of the plurality of blades within resiliency groups, the plurality of blades having computing resources and storage memory, each resiliency group supporting data recovery in case of loss of a specified number of blades of the resiliency group. The method includes transferring data from a first resiliency group to a second resiliency group, responsive to a change in geometry of the storage system.