Partition Switch for Distributed Storage Latency Reduction
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
In large-scale distributed storage systems, internal data messages often traverse multiple network hops, leading to increased latency and high-bandwidth interconnect link costs, which can be costly and inefficient.
Innovation Solution
Implementing a partition switch that directly connects storage drives and associated computational processes, reducing the number of network hops to as few as two, thereby minimizing latency and allowing for the use of lower-bandwidth, lower-cost links.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Speed
If traditional multi-hop network architecture is used for connecting storage drives, then system coverage and connectivity are improved, but latency increases and link bandwidth requirements increase
Solution Approach 1:
The network architecture is segmented into multiple hierarchical levels: edge switches at the storage drive level, aggregation switches at the rack level, and core switches at the data center level. This segmentation allows messages to traverse fewer hops within local segments while maintaining overall system connectivity, thereby reducing latency without sacrificing coverage.
Solution Approach 2:
The patent introduces a hierarchical dimension to the network architecture, organizing switches and storage drives across multiple levels (edge, aggregation, core). This dimensional organization enables localized fast communication at the edge level while providing broader connectivity through higher levels, resolving the contradiction between speed and coverage.
2Area of stationary object
If more network hops are used to connect storage drives across racks, then system coverage is improved, but latency and link costs increase
Solution Approach 1:
The data center is segmented into multiple racks, each with its own aggregation switch, and storage drives are grouped with edge switches at the rack level. This segmentation enables wide system coverage across many racks while keeping message hops minimal within each segment, thus reducing latency despite expanded coverage area.
Solution Approach 2:
Aggregation switches serve as intermediaries between edge switches at the rack level and core switches at the data center level. These intermediary devices enable messages to traverse the expanded network coverage area efficiently by providing optimized routing paths, thereby reducing overall transmission latency.
3Speed
If high-bandwidth interconnect links are used throughout the network, then message transmission speed is improved, but hardware costs increase
Solution Approach 1:
The network implements local quality by providing high-bandwidth links only where most critical - at the edge switch level where storage drives directly communicate. Lower-bandwidth links are used for aggregation and core routing where full bandwidth is less critical, thereby maintaining high transmission speed for critical operations while reducing overall hardware costs.
Solution Approach 2:
The network architecture provides dynamic bandwidth allocation through its hierarchical structure, where bandwidth requirements are met locally at the edge level for time-critical storage operations, while aggregate traffic at higher levels can tolerate lower bandwidth. This dynamic approach optimizes the balance between transmission speed and hardware cost.
Data Source
AI summary
Devices, computer-readable media, and methods for reducing the number of “hops” that internal messages must traverse in data center switching architectures are disclosed. In one example, a data center includes a first rack housing a first server, a first computational process associated to a first storage drive hosted on the first server and residing within a first level of the distributed storage system, a second rack housing a second server, a second computational process associated to a second storage drive hosted on the second server and residing within the first level of the distributed storage system, and a first switch communicatively coupled to the first level to receive messages directly from the first computational process and the second computational process.


