Multi-Durability Storage Architecture for Low-Latency Recovery
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing block-based storage systems face challenges in maintaining data durability and low latency due to single-point failures in server nodes or control planes, leading to significant storage capacity unavailability and high recovery latencies across multiple locations.
Innovation Solution
A data storage system architecture with head nodes and data storage sleds, where data is replicated across multiple sleds and nodes, allowing for independent operation without relying on a zonal control plane, and utilizing redundant networks and power distribution to ensure high reliability and durability.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If data is stored across multiple devices in multiple locations to increase durability, then data durability is improved, but latency for data recovery increases due to data being located across multiple locations
Solution Approach 1:
The system segments data into multiple parts and distributes them across different storage devices and locations. Each part can be independently accessed, allowing parallel data recovery operations that reduce overall recovery latency while maintaining durability through distributed storage.
Solution Approach 2:
The patent introduces a hierarchical storage architecture with multiple durability levels (hot, warm, cold storage) organized in different dimensions of the storage system. This allows data to be retrieved from the appropriate level based on access patterns, optimizing the balance between durability and recovery latency by serving frequently accessed data from lower-latency locations.
2Ease of operation
If a common control plane is used to manage data in multiple locations, then system management is simplified, but a failure of a control plane component impacts a large quantity of storage capacity
Solution Approach 1:
The control plane is segmented into distributed control nodes that operate independently across different storage locations. Each control node manages a subset of storage devices, eliminating the single point of failure in centralized control while maintaining manageable system operation through modular control architecture.
Solution Approach 2:
Storage devices and control nodes are designed to autonomously perform failover and data redistribution operations without requiring centralized control plane intervention. This self-service capability ensures continued operation and data availability even when control plane components fail, while maintaining ease of operation through automated management protocols.
3Adaptability or versatility
If extensive networks are used to move data between multiple locations, then data distribution capability is improved, but network complexity and cost increase
Solution Approach 1:
The patent combines multiple network communication paths into a unified data distribution architecture that leverages existing network infrastructure. By merging control and data plane communications and utilizing standard network protocols, the system achieves versatile data distribution capability while reducing network complexity and avoiding the need for extensive specialized networking infrastructure.
Data Source
AI summary
A data storage system includes multiple head nodes and multiple data storage sleds mounted in a rack. For a particular volume or volume partition one of the head nodes is designated as a primary head node for the volume or volume partition. The primary head node is configured to store data for the volume in a data storage of the primary head node and cause the data to be replicated to a secondary head node. The primary head node is also configured to cause the data for the volume to be stored in a plurality of respective mass storage devices each in different ones of the plurality of data storage sleds of the data storage system.


