Erasure Coding Scale-Out With Partial Re-Encoding
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional erasure coding technologies incur excessive processing overhead and bandwidth usage when scaling out storage systems, as they require complete re-encoding of data and storage of new coding fragments, which is inefficient and resource-intensive.
Innovation Solution
Implementing a partial encoding method for data storage clusters, where an initial protection scheme is determined based on the initial number of storage devices, and modified as additional devices are added, allowing for incremental updates to the coding fragments, reducing processing overhead and bandwidth usage.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If complete re-encoding is performed when storage devices are added to a group, then data protection reliability is improved, but processing overhead and bandwidth usage increase excessively
Solution Approach 1:
The patent applies partial encoding by only encoding the data portions that need to be redistributed to new storage devices, rather than performing complete re-encoding of all data. This partial action reduces processing overhead and bandwidth consumption while still achieving the necessary data protection reliability for the expanded storage group.
Solution Approach 2:
The patent segments the encoding process into distinct portions: identifying which data fragments need re-encoding, separating the encoding of those specific fragments, and selectively transferring only the necessary encoded fragments to new storage devices. This segmentation enables efficient resource utilization during scale-out operations.
2Reliability
If complete re-encoding is performed when storage devices are added to a group, then data protection reliability is improved, but bandwidth usage increases excessively
Solution Approach 1:
The patent performs partial encoding that generates only the specific coding fragments needed for the expanded storage configuration, rather than regenerating all coding fragments. This reduces network bandwidth consumption during the scale-out operation while maintaining adequate data protection reliability.
Solution Approach 2:
The patent extracts and identifies only the specific coding fragments that need to be transferred to new storage devices, separating them from the complete set of coding fragments. This extraction approach minimizes bandwidth usage by transmitting only the necessary data portions.
3Reliability
If conventional erasure coding technologies are used during scale out, then data protection is maintained, but processing time and resource consumption increase
Solution Approach 1:
The patent implements partial encoding that processes only the specific data fragments requiring re-encoding during scale-out operations, significantly reducing processing time compared to complete re-encoding while maintaining data protection integrity.
Solution Approach 2:
The patent performs preliminary identification of which data fragments require re-encoding before initiating the encoding process. This preliminary action enables the system to prepare only the necessary computations, reducing overall processing time during scale-out operations.
Data Source
AI summary
Scale out data protection with erasure coding is presented herein. Based on an initial number of storage devices determined to have been included in an initial stage of a data storage cluster, an initial protection scheme for the initial stage can determine first coding fragment(s) for data stored within the data storage cluster to facilitate a first recovery, from the initial stage, of the data using the first coding fragment(s). Further, in response to a defined number of additional storage devices being determined to have been added to the data storage cluster to generate a modified data storage cluster, the initial protection scheme can be modified to obtain a modified protection scheme that can determine, for the modified data storage cluster, second coding fragment(s) for the data to facilitate a second recovery of the data using the first coding fragment(s) and the second coding fragment(s).


