Hashing Address Space to Storage Servers for Load Balancing
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
In distributed storage systems, high-speed network coupling leads to significant disk access latency, which can cause delays in data retrieval, and existing techniques for replication and erasure coding do not effectively balance loads across storage servers, resulting in inefficient data access.
Innovation Solution
A method of hashing an address space to a plurality of storage servers, where the address space is divided into data segments, assigned to servers based on a sequence, and load measurements are used to adjust data shares to balance server loads while maintaining base addresses, thereby improving response times and cache hit ratios.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Speed
If data is stored in a distributed storage system with high-speed network coupling, then network transmission speed is improved, but disk access latency causes significant delays in data retrieval
Solution Approach 1:
The address space is divided into multiple data segments, each assigned to different storage servers. This segmentation allows parallel access to different portions of data, reducing overall access latency by eliminating the need to sequence through entire datasets on single disks.
Solution Approach 2:
The patent introduces a new dimension of data organization by hashing the address space to distribute data across multiple storage servers rather than relying on single-disk sequential access. This dimensional shift from vertical (single disk depth) to horizontal (multiple server breadth) access patterns mitigates disk access latency.
2Reliability
If replication and erasure coding are used to enhance reliability, then data redundancy is improved, but load balancing across storage servers deteriorates
Solution Approach 1:
The system performs preliminary hashing of the address space to determine optimal data segment assignments before actual data storage. This preliminary action ensures that replication and erasure coding operations are distributed evenly across storage servers from the outset, preventing load imbalance issues.
Solution Approach 2:
The patent implements a feedback mechanism that monitors load distribution across storage servers and dynamically adjusts data segment assignments. This feedback loop ensures that as data is added or removed, the system maintains balanced loads while preserving the reliability benefits of replication and erasure coding.
3Ease of manufacture
If data segments are assigned to storage servers based on a fixed sequence, then assignment simplicity is improved, but load balance across servers deteriorates
Solution Approach 1:
The patent changes the assignment parameter from a fixed sequential index to a hash-based address space mapping. This parameter change maintains simplicity by using a deterministic function (hashing) while achieving superior load balance by distributing data segments based on address patterns rather than arbitrary sequence numbers.
Data Source
AI summary
An embodiment of a method of hashing an address space to a plurality of storage servers begins with a first step of dividing the address space by a number of the storage servers to form data segments. Each data segment comprises a base address. A second step assigns the data segments to the storage servers according to a sequence. The method continues with a third step of measuring a load on each of the storage servers. According to an embodiment, the method concludes with a fourth step of adjusting data shares assigned to the storage servers according to the sequence to approximately balances the loads on the storage servers while maintaining the base address for each data segment on an originally assigned storage server. According to another embodiment, the method periodically performs the third and fourth steps to maintain an approximately balanced load on the storage servers.


