Parity Code Regeneration With Skipper Parity for Low Repair Traffic
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional parity-checking algorithms face significant computational and storage requirements, and are unable to scale effectively, particularly in large storage applications, and traditional Reed Solomon codes do not provide sufficient traffic savings during data reconstruction, especially in cases of storage device failures.
Innovation Solution
A data storage system that generates a skipper parity using invertible operations, combining data elements to produce horizontal and skipper parity entries, allowing for efficient data recreation with reduced repair traffic, where only half of the remaining data is needed to recreate a failed storage node, and updates to parities require minimal changes.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If conventional parity-checking algorithms are used, then data integrity can be verified, but computational and storage requirements increase significantly and scalability deteriorates
Solution Approach 1:
The patent segments the parity-checking process into two distinct components: horizontal parity (combining data elements across storage nodes) and skipper parity (combining data elements within a storage node after transformation). This segmentation allows each parity type to handle specific aspects of error detection and data recovery independently, reducing the overall computational complexity compared to conventional unified parity-checking algorithms.
Solution Approach 2:
The patent introduces a transformation dimension by applying invertible transformations to data elements before combining them to form skipper parity. This dimensional change enables more efficient data recovery operations, as the transformation creates additional mathematical relationships that can be exploited during reconstruction, thereby reducing computational requirements while maintaining reliability.
2Reliability
If traditional Reed Solomon code is used for error correction, then up to 2 node failures can be tolerated, but repair traffic increases significantly during data reconstruction
Solution Approach 1:
The patent extracts and utilizes specific mathematical properties of the parity structure during data recovery. By leveraging the invertible transformation relationships embedded in the skipper parity, the system can recover lost data elements using only the necessary parity information and surviving data elements, rather than requiring all surviving nodes to participate in repair operations as in traditional Reed Solomon codes.
Solution Approach 2:
The patent changes the mathematical parameters of the parity code by using invertible transformations (such as XOR operations or other bijective mappings) on data elements before parity combination. This parameter transformation enables more efficient repair traffic calculations, where the recovery process requires fewer data transmissions compared to conventional Reed Solomon reconstruction.
3Reliability
If all surviving nodes are contacted for data recreation, then complete data recovery is possible, but traffic requirements increase
Solution Approach 1:
The patent applies partial action by enabling data recovery using only a subset of available parity information and surviving data elements, rather than requiring all surviving nodes. The invertible transformation relationships in the skipper parity allow the system to perform sufficient recovery operations with partial data, reducing traffic requirements while maintaining complete data recovery capability.
Data Source
AI summary
The disclosed technology can advantageously provide an efficient data recovery system including a plurality of storage nodes including a first storage node and a second storage node, and a storage logic that is coupled to the storage nodes and that manages storage of data on the storage nodes. The storage logic is executable to: receive a data set including data elements including a first set of data elements associated with the first storage node and a second set of data elements associated with the second storage node; generate a first parity of the data set, the first parity including a horizontal parity including a set of horizontal parity entries; and combine the data elements from the data set to produce a skipper parity including a set of skipper parity entries. Combining the data elements includes transforming a subset of the data elements from the data set using an invertible operation, the set of horizontal parity entries being different from the set of skipper parity entries.


