Failure-decoupled volume-level redundancy coding for data storage

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current network computing and storage systems face challenges in optimizing data performance and integrity, particularly in ensuring availability and durability of data across multiple volumes, especially when dealing with failures and the need for efficient storage and retrieval of large datasets.

Innovation Solution

The implementation of failure-decoupled volume-level redundancy coding techniques, which involve sorting data into failure-decorrelated subsets, using redundancy codes like erasure codes to distribute data across multiple volumes, and generating indices for efficient storage and retrieval, allowing for the regeneration of data even if some volumes or shards become unavailable.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If traditional redundancy coding is used to ensure data availability across multiple volumes, then data durability is improved, but the number of volumes required increases, reducing storage efficiency

Engineering Contradiction:
Improvedata availabilityVSAvoidnumber of volumes
Core Design Contradiction:
ReliabilityVSQuantity of substance

Solution Approach 1:

The patent segments data into failure-decorrelated subsets that can be independently stored on different volumes. By dividing the data storage system into independent failure domains, the system ensures that a failure in one volume does not compromise the entire dataset, thereby maintaining data availability while optimizing the number of volumes required.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent applies failure-decoupled volume-level redundancy coding techniques in advance to sort data into failure-decorrelated subsets before storage. This preliminary organization ensures that data is pre-positioned across volumes in a way that maximizes availability while minimizing the total number of volumes needed, resolving the contradiction between reliability and storage efficiency.

Inventive Principle:
Principle #10Preliminary action

2Reliability

If data is distributed across multiple volumes for fault tolerance, then data durability is improved, but storage efficiency deteriorates due to increased overhead

Engineering Contradiction:
Improvedata durabilityVSAvoidstorage overhead
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

By segmenting data into failure-decorrelated subsets and storing them across volumes with independent failure domains, the patent reduces the overhead associated with traditional redundancy coding. Each subset can be independently managed and recovered, reducing the complexity and overhead of managing redundant data across the entire system.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent changes the parameters of redundancy coding by applying failure-decoupled volume-level techniques that optimize the balance between data durability and storage overhead. By adjusting how data is encoded and distributed across volumes based on failure correlations, the system achieves improved durability without proportionally increasing storage overhead.

Inventive Principle:
Principle #35Parameter changes

3Reliability

If more volumes are used to ensure data availability, then data integrity is improved, but retrieval time increases due to distributed access

Engineering Contradiction:
Improvedata integrityVSAvoidretrieval time
Core Design Contradiction:
ReliabilityVSLoss of time

Solution Approach 1:

The patent segments data into failure-decorrelated subsets that can be retrieved independently. When data needs to be accessed, the system can retrieve only the specific subset needed rather than accessing all distributed volumes, significantly reducing retrieval time while maintaining data integrity through the redundant subset structure.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

By pre-organizing data into failure-decorrelated subsets with clear indexing and location information, the patent enables rapid identification and retrieval of specific data portions. This preliminary organization allows the system to access data quickly from the minimum necessary volumes without scanning or accessing the entire distributed storage system.

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS10089179B2Failure-decoupled volume-level redundancy coding techniques
Publication Date: 2018.10.02 AMAZON TECH INC
  • US10089179B2 patent drawing
  • US10089179B2 patent drawing
  • US10089179B2 patent drawing

AI summary

Techniques described and suggested herein include systems and methods for storing, indexing, and retrieving original data of data archives on data storage systems using redundancy coding techniques. For example, redundancy codes, such as erasure codes, may be applied to archives (such as those received from a customer of a computing resource service provider) so as allow the storage of original data of the individual archives available on a minimum of volumes, such as those of a data storage system, while retaining availability, durability, and other guarantees imparted by the application of the redundancy code. Sparse indexing techniques may be implemented so as to reduce the footprint of indexes used to locate the original data, once stored. The volumes may be apportioned into failure-decorrelated subsets, and archives stored thereto may be apportioned to such subsets.