Locality-Sensitive Data Retrieval for Redundancy Coded Storage

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Network computing and storage systems face challenges in optimizing data performance and integrity, particularly in efficiently storing and retrieving redundancy coded data across multiple volumes while ensuring availability and durability.

Innovation Solution

The implementation of redundancy coding techniques, such as erasure codes, that distribute original data across multiple volumes as shards, allowing for efficient storage and retrieval by using identity shards for direct access and additional shards for regeneration, while also optimizing storage across heterogeneous components.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If redundancy coded data is distributed across multiple volumes, then data durability and availability are improved, but data retrieval performance deteriorates due to increased access complexity

Engineering Contradiction:
Improvedata durabilityVSAvoiddata retrieval performance
Core Design Contradiction:
ReliabilityVSSpeed

Solution Approach 1:

The patent segments data into identity shards and encoded shards stored across multiple volumes. Identity shards contain original data portions that can be directly retrieved, while encoded shards contain redundancy information. This segmentation allows the system to maintain durability through distributed storage while enabling fast retrieval by prioritizing direct access to identity shards.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent introduces an optimization engine as an intermediary component that mediates between data retrieval requests and the distributed shard storage. The optimization engine determines whether to retrieve data directly from identity shards or to regenerate data from encoded shards, thereby optimizing retrieval performance while maintaining the benefits of distributed redundancy storage.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Quantity of substance

If redundancy coding is implemented across heterogeneous storage components, then storage capacity and cost efficiency are improved, but system complexity increases

Engineering Contradiction:
Improvestorage capacityVSAvoidsystem complexity
Core Design Contradiction:
Quantity of substanceVSDevice complexity

Solution Approach 1:

The patent implements a universal shard interface that allows heterogeneous storage components to be accessed through a common protocol. Both identity shards and encoded shards can be stored across different storage media types, and the optimization engine provides a unified interface for retrieving data regardless of the underlying storage heterogeneity, thereby managing system complexity while maximizing storage capacity utilization.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Solution Approach 2:

The patent dynamically adjusts retrieval parameters based on storage conditions and data characteristics. The optimization engine evaluates factors such as storage media performance, data access patterns, and shard availability to determine optimal retrieval strategies, allowing the system to adapt to heterogeneous storage environments without requiring complex manual configuration.

Inventive Principle:
Principle #35Parameter changes

3Speed

If direct access to original data is prioritized, then data retrieval speed is improved, but data integrity verification becomes more difficult

Engineering Contradiction:
Improvedata retrieval speedVSAvoiddata integrity verification
Core Design Contradiction:
SpeedVSMeasurement precision

Solution Approach 1:

The patent performs preliminary actions by storing both identity shards (original data portions) and encoded shards (redundancy information) in advance across multiple volumes. This preliminary distribution of data in multiple forms enables the system to quickly retrieve data directly from identity shards when available, while also maintaining the capability to verify data integrity through encoded shards without requiring additional real-time processing.

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS10311020B1Locality-sensitive data retrieval for redundancy coded data storage systems
Publication Date: 2019.06.04 AMAZON TECH INC
  • US10311020B1 patent drawing
  • US10311020B1 patent drawing
  • US10311020B1 patent drawing

AI summary

Techniques described and suggested herein include systems and methods for optimizing retrieval, based on localities associated with a requestor and that of various components of a data storage system, of data archives stored on data storage systems using redundancy coding techniques. For example, redundancy coded shards, which may include identity shards that contain unencoded original data of archives, may be configured such that a variable number of the shards can be leveraged to meet performance requirements or time-to-retrieval limitations for retrieval requests associated with the archives stored and/or encoded therein. Under some circumstances, implementing systems may monitor relative geographic locations, among other performance-related metrics, so as to retrieve data such that fewer hosting data storage facilities are used for a given retrieval.