Distributed Storage Recovery Matrix for Concurrent Node Failures
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current distributed storage systems face challenges in efficiently recovering data from multiple node failures due to high recovery bandwidth requirements and complex implementation of concurrent recovery strategies, particularly with minimum storage regenerating (MSR) codes constructed via product matrix (PM) methods, which are limited to single-node failures and lack practical solutions for general coding parameters and multiple node failures.
Innovation Solution
A method for data recovery in distributed storage systems using MSR codes constructed via PM, which includes centralized and distributed modes, where the number of assistant nodes is chosen based on the number of failed nodes, and a repair matrix is calculated to reconstruct missing data blocks, allowing for concurrent recovery of multiple nodes without requiring data exchange between new nodes, thus reducing recovery bandwidth and computing overhead.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If original PM MSR code is used for single node failure recovery, then the storage efficiency is maximized and recovery bandwidth is minimized, but it cannot handle multiple concurrent node failures
Solution Approach 1:
The patent segments the recovery process into two distinct phases: (1) assistant nodes download helper data from surviving nodes, and (2) substitute nodes exchange and process this data to reconstruct failed node data. This segmentation allows the system to handle multiple failures concurrently while maintaining the efficiency benefits of the original PM MSR code
Solution Approach 2:
Assistant nodes perform preliminary actions by downloading and storing helper data from surviving nodes before the actual reconstruction process. This preliminary data collection enables substitute nodes to efficiently reconstruct failed node data through local computation and data exchange, rather than requiring complex real-time coordination during failure recovery
2Ease of manufacture
If failed nodes are recovered one by one, then the recovery process is straightforward to implement, but the total bandwidth consumption increases significantly
Solution Approach 1:
The patent merges the recovery processes of multiple failed nodes into a single concurrent operation. Assistant nodes collect helper data for multiple failed nodes simultaneously, and substitute nodes exchange data to reconstruct all failed nodes in parallel. This merging reduces total bandwidth consumption compared to sequential recovery while maintaining implementation feasibility through clear phase separation
3Loss of energy
If cooperative recovery with data exchange between substitute nodes is implemented, then the recovery bandwidth is minimized, but the completion time increases and requires complex protocols
Solution Approach 1:
The patent introduces dynamic elements to the recovery process by allowing substitute nodes to flexibly exchange data based on their specific needs and the data already possessed by other substitute nodes. This dynamic data exchange enables optimized bandwidth utilization while the parallel processing architecture maintains relatively short completion times by avoiding sequential dependencies
Data Source
AI summary
A method of data recovery for a distributed storage system is a method of recovering multiple failed nodes concurrently with the minimum feasible bandwidth when failed nodes exist in a distributed storage system. By means of selecting assistant nodes, obtaining helper data sub-blocks through computing the selected assistant nodes, then computing a repair matrix and finally multiple the repair matrix and the helper data sub-blocks, the missing data blocks are reconstructed; or the missing data blocks are reconstructed by decoding. The method is applicable to data recovery in the case of any number of failed nodes and any reasonable combinations of coding parameters. The data recovery herein can reach the theoretical lower limit of the minimum recovery bandwidth.


