Redundant Rack Storage Networks for Low-Latency Data Durability

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing block-based storage systems face challenges in maintaining data durability and low latency due to single-point failures, such as server node failures or common control plane failures, which can lead to significant storage capacity unavailability and high recovery latencies across multiple locations.

Innovation Solution

A data storage system with a rack-mounted configuration of head nodes and data storage sleds, where data is replicated across multiple head nodes and mass storage devices, allowing for independent operation without relying on a zonal control plane, and utilizing redundant networks and power systems to ensure high reliability and durability.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If data is stored across multiple devices in multiple locations using a common control plane, then data durability is improved, but system complexity increases and single-point failures can still impact large quantities of storage capacity

Engineering Contradiction:
Improvedata durabilityVSAvoidsystem complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The patent divides the storage system into independent rack units, each with its own control plane and networking. This segmentation isolates failures to individual racks rather than affecting the entire storage system, while maintaining data durability through replication across multiple racks. Each rack operates autonomously with head nodes that can independently manage data operations.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent introduces rack-mounted network switches as intermediaries that facilitate communication within each rack while maintaining independence from other racks. These switches enable local data operations without requiring communication with a centralized control plane, thus reducing system complexity and eliminating single-point failures while preserving data durability through distributed architecture.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Ease of operation

If storage systems use server nodes to service multiple storage nodes, then ease of operation is improved, but reliability deteriorates due to single-point failures rendering large storage capacity unusable

Engineering Contradiction:
Improveease of operationVSAvoidstorage availability
Core Design Contradiction:
Ease of operationVSReliability

Solution Approach 1:

The patent segments the storage system into independent rack units where each rack can operate autonomously. Head nodes are distributed across racks rather than centralized, eliminating single-point failures. Each rack maintains its own control plane and networking infrastructure, ensuring that failures in one rack do not impact other racks or render large storage capacity unusable.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

Each rack is designed to be self-sufficient with its own control plane, networking infrastructure, and head nodes. This self-service capability allows racks to continue operating independently even when other racks experience failures, maintaining storage availability without requiring centralized server node management.

Inventive Principle:
Principle #25Self-service

3Reliability

If data is located across multiple geographic locations, then data durability is improved, but latency increases due to extensive networks required for data movement

Engineering Contradiction:
Improvedata durabilityVSAvoiddata recovery latency
Core Design Contradiction:
ReliabilityVSLoss of time

Solution Approach 1:

The patent transitions from geographic distribution to physical rack-level distribution within data centers. By organizing storage units in racks with dedicated networking infrastructure, the system achieves data durability through spatial distribution across multiple racks while minimizing network latency through localized high-speed connections within each rack, eliminating the need for extensive geographic network infrastructure.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

Data Source

PatentUS10771550B2Data storage system with redundant internal networks
Publication Date: 2020.09.08 AMAZON TECH INC
  • US10771550B2 patent drawing
  • US10771550B2 patent drawing
  • US10771550B2 patent drawing

AI summary

A data storage system includes a rack, multiple head nodes, multiple data storage sleds, and at least two networking devices. The at least two network devices are configured to implement at least two redundant networks within the data storage system. Also, each of the head nodes is assigned at least two network addresses for communication with the data storage sleds of the data storage system via the at least two networking devices. The data storage sleds each include multiple mass storage devices and a sled controller that is configured to couple with the at least two network switches. In some embodiments, the data storage system further includes redundant power systems within a rack in which the head nodes, the data storage sleds, and the at least two networking devices are mounted.