Network-Aware Data Storage for Lower Cross-Cluster Traffic
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Decentralized storage systems generate excessive data traffic across network boundaries, causing congestion and requiring additional network equipment, without considering operator network boundaries or bottlenecks.
Innovation Solution
A network-aware coordinator node optimizes data distribution by selecting storage nodes based on their network infrastructure positions relative to the data source, ensuring a minimum number of data blocks are stored in the same cluster as the data source, with additional blocks distributed across other clusters to maintain resiliency and minimize cross-cluster traffic.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If data blocks are distributed across multiple clusters without considering network boundaries, then storage resiliency is improved, but data traffic volume increases causing network congestion
Solution Approach 1:
The patent applies local quality by making the data distribution strategy location-dependent. Data blocks are preferentially stored in the same network cluster as the data source, with only necessary blocks placed in remote clusters. This creates different storage densities and traffic patterns for different network locations, optimizing both resiliency and traffic reduction simultaneously.
2Speed
If data blocks are stored in the same cluster as the data source, then read/write access times are improved, but storage resiliency may be reduced
Solution Approach 1:
The patent applies partial action by storing only the minimum necessary number of data blocks in remote clusters rather than distributing all blocks universally. This partial distribution across clusters maintains resiliency while keeping the majority of blocks locally accessible, thus improving access times without sacrificing reliability.
3Device complexity
If decentralized storage systems operate without network awareness, then system simplicity is maintained, but additional network equipment is required causing capital expenditure increase
Solution Approach 1:
The patent applies self-service by enabling the storage system to automatically identify and utilize existing network cluster structures without requiring external network equipment or complex configuration. The system self-organizes data blocks according to network boundaries, eliminating the need for additional capital expenditure on network infrastructure.
Data Source
AI summary
Methods and apparatus are provided. In an example aspect, a method of storing data in a communication system is provided. The data comprises n data blocks, and k data blocks of the n data blocks are required to recover the data. The communication system includes a plurality of clusters of data storage nodes including a first cluster that is a closest cluster of the plurality of clusters to a source of the data. The method comprises receiving, from a node associated with a respective network operator of each of the clusters of storage nodes, information identifying the data storage nodes that are in each of the clusters. The method also comprises causing each of at least k data blocks of the n data blocks to be stored in a different storage node in the first cluster.


