Distributed Storage Node Recovery With Minimal Repair Bandwidth

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current distributed storage systems face inefficiencies in network bandwidth usage and recovery time due to the lack of effective methods for concurrent recovery of multiple failed nodes, particularly with existing MSR codes that require complex coordination and high overhead.

Innovation Solution

A method using minimum storage generating codes based on interference alignment, where storage nodes are divided into systematical and parity nodes, with assistant nodes sending encoded helper data to regenerate missing blocks, allowing for concurrent recovery of multiple failed nodes with reduced bandwidth and time requirements.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Quantity of substance

If erasure coding is used to maximize storage efficiency, then space utilization is improved, but network bandwidth consumption increases during failed node repair

Engineering Contradiction:
Improvestorage efficiencyVSAvoidnetwork bandwidth consumption
Core Design Contradiction:
Quantity of substanceVSLoss of energy

Solution Approach 1:

The patent segments the repair process into two distinct phases: a first phase where the new node downloads data from surviving nodes, and a second phase where multiple new nodes exchange data among themselves. This segmentation allows optimization of network bandwidth by reducing redundant data transmission across the network.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent applies preliminary action by having the new node download necessary data from surviving nodes before the actual repair process begins. This preliminary data collection enables subsequent new nodes to perform repairs more efficiently by exchanging data locally rather than all downloading from the same surviving nodes, thus reducing overall network bandwidth consumption.

Inventive Principle:
Principle #10Preliminary action

2Loss of energy

If cooperative concurrent recovery strategies are implemented to reduce total repair bandwidth, then network bandwidth efficiency is improved, but implementation complexity increases

Engineering Contradiction:
Improvetotal repair bandwidthVSAvoidcoordination complexity
Core Design Contradiction:
Loss of energyVSDevice complexity

Solution Approach 1:

The patent divides the concurrent recovery process into sequential phases with clear responsibilities: first phase for initial data download from surviving nodes, second phase for inter-node data exchange. This segmentation simplifies coordination by establishing a clear execution order and reducing the need for complex real-time synchronization among nodes.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent introduces dynamic node roles where nodes can transition from being data sources in the first phase to data consumers and sharers in the second phase. This dynamic approach allows flexible coordination without requiring pre-assigned fixed roles, reducing implementation complexity while maintaining bandwidth efficiency.

Inventive Principle:
Principle #15Dynamics

3Reliability

If complete data segments are collected to repair failed nodes using traditional erasure coding, then data security is ensured, but recovery time increases

Engineering Contradiction:
Improvedata securityVSAvoidrecovery time
Core Design Contradiction:
ReliabilityVSLoss of time

Solution Approach 1:

The patent applies partial action by having new nodes download only the necessary portions of data from surviving nodes during the first phase, rather than complete data segments. During the second phase, nodes exchange only the specific data portions needed for repair. This partial data collection approach maintains data security through erasure coding properties while significantly reducing recovery time by avoiding unnecessary data transmission.

Inventive Principle:
Principle #16Partial or excessive action

4Manufacturing precision

If exact regenerating codes are used to recover identical data blocks, then data integrity is improved, but recovery bandwidth increases compared to functional regenerating codes

Engineering Contradiction:
Improvedata block accuracyVSAvoidrecovery bandwidth
Core Design Contradiction:
Manufacturing precisionVSLoss of energy

Solution Approach 1:

The patent changes the approach from exact regeneration to functional regeneration, where the focus shifts from recovering identical data blocks to recovering data blocks that satisfy MDS and bandwidth efficiency properties. This parameter change in the regeneration objective allows for reduced recovery bandwidth while maintaining sufficient data integrity for storage system operation.

Inventive Principle:
Principle #35Parameter changes

Data Source

PatentUS11188404B2Methods of data concurrent recovery for a distributed storage system and storage medium thereof
Publication Date: 2021.11.30 HERE DATA TECH
  • US11188404B2 patent drawing
  • US11188404B2 patent drawing
  • US11188404B2 patent drawing

AI summary

Disclosed is a method of data concurrent recovery for a distributed storage system, that is, a method for synchronous repair of multiple failed nodes with a minimum recovery bandwidth when a node in a distributed storage system fails. First an assistant node is selected to get helper data sub-block, then the repair matrix related to the data block stored in the node to be repaired is constructed, and finally the lost data block is reconstructed by multiplying the repair matrix and the helper data helper data sub-block; the missing data block is reconstructed by decoding, wherein the node to be recovered includes all failed systematical nodes, or all or partly failed parity nodes. The method is applicable to concurrently recover multiple failed nodes at minimal recovery bandwidth, and the nodes to be recovered are selected according to the demand to reduce the recovery bandwidth as much as possible.