Resiliency Group Configurations for Storage Scalability

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

As storage systems scale, the mean time to failure (MTF) and mean time to data loss (MTTDL) worsen due to increasing storage memory and component failure rates, leading to challenges in maintaining data recoverability and survivability.

Innovation Solution

The solution involves dynamically forming failure domains in a storage system with multiple blades, using a failure domain formation policy to identify suitable configurations that balance reliability, data survivability, and resource efficiency.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Quantity of substance

If storage systems scale larger with more storage memory and components, then storage capacity increases, but mean time to failure and mean time to data loss worsen due to increasing component failure rates

Engineering Contradiction:
Improvestorage capacityVSAvoidmean time to failure
Core Design Contradiction:
Quantity of substanceVSReliability

Solution Approach 1:

The storage system is divided into multiple failure domains, where each domain is a group of storage devices that can fail together due to common causes. By segmenting the storage system into isolated failure domains, the patent prevents a single failure event from affecting the entire system, thus improving reliability while maintaining large storage capacity.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent introduces failure domains as intermediary groups between individual storage devices and the overall storage system. These failure domains act as buffers that contain and isolate failures, allowing the system to maintain data survivability even as storage capacity scales by adding more devices.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Quantity of substance

If storage systems scale larger with more components, then storage capacity increases, but data recoverability becomes more challenging due to multiple component failures

Engineering Contradiction:
Improvestorage capacityVSAvoiddata recoverability
Core Design Contradiction:
Quantity of substanceVSLoss of information

Solution Approach 1:

By segmenting the storage system into multiple failure domains with limited device groupings, the patent ensures that data within each domain can be recovered even if multiple devices fail. This segmentation prevents cascading failures across the entire system, maintaining data recoverability as storage capacity scales.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent implements failure domains as a preemptive protective measure, creating boundaries that cushion the system against the propagation of failures. This prior cushioning allows the system to withstand multiple component failures within a domain while preserving data recoverability across the entire storage system.

Inventive Principle:
Principle #11Beforehand cushioning (Prior cushioning)

3Reliability

If traditional RAID stripes with error correction coding are used, then data recoverability is maintained for single failures, but data is lost when multiple component failures exceed recovery capability

Engineering Contradiction:
Improvedata recoverabilityVSAvoiddata loss in multiple failures
Core Design Contradiction:
ReliabilityVSLoss of information

Solution Approach 1:

The patent divides the storage system into multiple failure domains, each with its own error correction coding and data recovery capabilities. This segmentation allows each domain to independently handle multiple device failures without affecting other domains, thereby preventing data loss that would occur in traditional unified RAID systems when failure thresholds are exceeded.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

Each failure domain is configured with local error correction coding and recovery mechanisms tailored to its specific devices. This local quality approach ensures that data protection and recovery capabilities are optimized for each domain's failure characteristics, preventing data loss even when multiple components fail within a domain.

Inventive Principle:
Principle #3Local quality

Data Source

PatentUS20250156286A1Resiliency group configurations
Publication Date: 2025.05.15 PURE STORAGE INC
  • US20250156286A1 patent drawing
  • US20250156286A1 patent drawing
  • US20250156286A1 patent drawing

AI summary

A storage system with storage drives and a processing device establishes resiliency groups of storage system resources. The storage system determines an explicit trade-off between data survivability over resource failures and data capacity efficiency, for the resiliency groups. Responsive to adding at least one storage drive, the storage system establishes re-formed resiliency groups according to the explicit trade-off, without decreasing data survivability. The storage system may bias to have more and narrower resiliency groups to increase mean time to data loss.