Distributed Storage Node Recovery With Minimal Repair Bandwidth
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current distributed storage systems face inefficiencies in network bandwidth usage and recovery time due to the lack of effective methods for concurrent recovery of multiple failed nodes, particularly with existing MSR codes that require complex coordination and high overhead.
Innovation Solution
A method using minimum storage generating codes based on interference alignment, where storage nodes are divided into systematical and parity nodes, with assistant nodes sending encoded helper data to regenerate missing blocks, allowing for concurrent recovery of multiple failed nodes with reduced bandwidth and time requirements.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Quantity of substance
If erasure coding is used to maximize storage efficiency, then space utilization is improved, but network bandwidth consumption increases during failed node repair
Solution Approach 1:
The patent segments the repair process into two distinct phases: a first phase where the new node downloads data from surviving nodes, and a second phase where multiple new nodes exchange data among themselves. This segmentation allows optimization of network bandwidth by reducing redundant data transmission across the network.
Solution Approach 2:
The patent applies preliminary action by having the new node download necessary data from surviving nodes before the actual repair process begins. This preliminary data collection enables subsequent new nodes to perform repairs more efficiently by exchanging data locally rather than all downloading from the same surviving nodes, thus reducing overall network bandwidth consumption.
2Loss of energy
If cooperative concurrent recovery strategies are implemented to reduce total repair bandwidth, then network bandwidth efficiency is improved, but implementation complexity increases
Solution Approach 1:
The patent divides the concurrent recovery process into sequential phases with clear responsibilities: first phase for initial data download from surviving nodes, second phase for inter-node data exchange. This segmentation simplifies coordination by establishing a clear execution order and reducing the need for complex real-time synchronization among nodes.
Solution Approach 2:
The patent introduces dynamic node roles where nodes can transition from being data sources in the first phase to data consumers and sharers in the second phase. This dynamic approach allows flexible coordination without requiring pre-assigned fixed roles, reducing implementation complexity while maintaining bandwidth efficiency.
3Reliability
If complete data segments are collected to repair failed nodes using traditional erasure coding, then data security is ensured, but recovery time increases
Solution Approach 1:
The patent applies partial action by having new nodes download only the necessary portions of data from surviving nodes during the first phase, rather than complete data segments. During the second phase, nodes exchange only the specific data portions needed for repair. This partial data collection approach maintains data security through erasure coding properties while significantly reducing recovery time by avoiding unnecessary data transmission.
4Manufacturing precision
If exact regenerating codes are used to recover identical data blocks, then data integrity is improved, but recovery bandwidth increases compared to functional regenerating codes
Solution Approach 1:
The patent changes the approach from exact regeneration to functional regeneration, where the focus shifts from recovering identical data blocks to recovering data blocks that satisfy MDS and bandwidth efficiency properties. This parameter change in the regeneration objective allows for reduced recovery bandwidth while maintaining sufficient data integrity for storage system operation.
Data Source
AI summary
Disclosed is a method of data concurrent recovery for a distributed storage system, that is, a method for synchronous repair of multiple failed nodes with a minimum recovery bandwidth when a node in a distributed storage system fails. First an assistant node is selected to get helper data sub-block, then the repair matrix related to the data block stored in the node to be repaired is constructed, and finally the lost data block is reconstructed by multiplying the repair matrix and the helper data helper data sub-block; the missing data block is reconstructed by decoding, wherein the node to be recovered includes all failed systematical nodes, or all or partly failed parity nodes. The method is applicable to concurrently recover multiple failed nodes at minimal recovery bandwidth, and the nodes to be recovered are selected according to the demand to reduce the recovery bandwidth as much as possible.


