Locality-Sensitive Data Retrieval for Redundancy Coded Storage
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Network computing and storage systems face challenges in optimizing data performance and integrity, particularly in efficiently storing and retrieving redundancy coded data across multiple volumes while ensuring availability and durability.
Innovation Solution
The implementation of redundancy coding techniques, such as erasure codes, that distribute original data across multiple volumes as shards, allowing for efficient storage and retrieval by using identity shards for direct access and additional shards for regeneration, while also optimizing storage across heterogeneous components.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If redundancy coded data is distributed across multiple volumes, then data durability and availability are improved, but data retrieval performance deteriorates due to increased access complexity
Solution Approach 1:
The patent segments data into identity shards and encoded shards stored across multiple volumes. Identity shards contain original data portions that can be directly retrieved, while encoded shards contain redundancy information. This segmentation allows the system to maintain durability through distributed storage while enabling fast retrieval by prioritizing direct access to identity shards.
Solution Approach 2:
The patent introduces an optimization engine as an intermediary component that mediates between data retrieval requests and the distributed shard storage. The optimization engine determines whether to retrieve data directly from identity shards or to regenerate data from encoded shards, thereby optimizing retrieval performance while maintaining the benefits of distributed redundancy storage.
2Quantity of substance
If redundancy coding is implemented across heterogeneous storage components, then storage capacity and cost efficiency are improved, but system complexity increases
Solution Approach 1:
The patent implements a universal shard interface that allows heterogeneous storage components to be accessed through a common protocol. Both identity shards and encoded shards can be stored across different storage media types, and the optimization engine provides a unified interface for retrieving data regardless of the underlying storage heterogeneity, thereby managing system complexity while maximizing storage capacity utilization.
Solution Approach 2:
The patent dynamically adjusts retrieval parameters based on storage conditions and data characteristics. The optimization engine evaluates factors such as storage media performance, data access patterns, and shard availability to determine optimal retrieval strategies, allowing the system to adapt to heterogeneous storage environments without requiring complex manual configuration.
3Speed
If direct access to original data is prioritized, then data retrieval speed is improved, but data integrity verification becomes more difficult
Solution Approach 1:
The patent performs preliminary actions by storing both identity shards (original data portions) and encoded shards (redundancy information) in advance across multiple volumes. This preliminary distribution of data in multiple forms enables the system to quickly retrieve data directly from identity shards when available, while also maintaining the capability to verify data integrity through encoded shards without requiring additional real-time processing.
Data Source
AI summary
Techniques described and suggested herein include systems and methods for optimizing retrieval, based on localities associated with a requestor and that of various components of a data storage system, of data archives stored on data storage systems using redundancy coding techniques. For example, redundancy coded shards, which may include identity shards that contain unencoded original data of archives, may be configured such that a variable number of the shards can be leveraged to meet performance requirements or time-to-retrieval limitations for retrieval requests associated with the archives stored and/or encoded therein. Under some circumstances, implementing systems may monitor relative geographic locations, among other performance-related metrics, so as to retrieve data such that fewer hosting data storage facilities are used for a given retrieval.


