Interleaved Erasure Coding Across Tape Cartridges for Correlated Errors
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional tape data storage systems face challenges in achieving high durability due to non-random and correlated errors, such as mechanical and environmental disturbances, which conventional erasure codes and internal error correction algorithms struggle to handle efficiently, leading to data loss and increased costs from the need for multiple copies and complex tape management.
Innovation Solution
The implementation of fountain erasure coding, which distributes data across multiple tape cartridges in an interleaved manner, allowing for double protection against errors and enabling efficient handling of localized damage and system-level errors, while reducing the overhead required to achieve eleven nines of durability.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If conventional erasure codes and internal error correction algorithms are used, then data protection against random errors is achieved, but non-random and correlated errors (mechanical and environmental disturbances) cause data loss
Solution Approach 1:
The patent segments data into multiple code words and further divides each code word into multiple chunks, which are then distributed across multiple tape cartridges. This segmentation allows the system to handle non-random errors by ensuring that no single error event can corrupt all chunks of a code word, thereby improving adaptability to correlated errors while maintaining data protection
Solution Approach 2:
The patent introduces a new dimension of distribution by spreading data chunks across multiple tape cartridges rather than concentrating them on single cartridges or within a single code structure. This multi-dimensional distribution strategy enables the system to withstand localized damage and non-random errors that would otherwise affect concentrated data storage
2Reliability
If multiple copies of tape cartridges are created to achieve required durability, then data protection against hard errors is improved, but overhead and tape management complexity increase significantly
Solution Approach 1:
The patent introduces erasure codes as an intermediary layer between the data and the tape cartridges. Instead of creating multiple physical copies of data, the system uses encoded chunks distributed across cartridges, with the erasure code structure serving as a mediator that enables data recovery from any sufficient subset of chunks, thereby reducing tape management complexity while maintaining durability
Solution Approach 2:
The patent creates logical copies of encoded data chunks across multiple tape cartridges rather than physical copies of the entire data set. This selective copying of encoded segments, combined with erasure coding mathematics, allows reconstruction of original data from any sufficient combination of chunks, significantly reducing the number of tapes needed compared to full replication while achieving the same durability goals
3Reliability
If conventional replication is used to handle hard errors, then data durability is achieved, but the number of tapes required increases (e.g., 24 tapes for 4 tapes distributed over 6 copies)
Solution Approach 1:
The patent changes the fundamental parameter of data representation by applying erasure coding transformations to data before distribution. Instead of storing raw data copies, the system transforms data into encoded chunks where mathematical relationships between chunks enable recovery. This parameter change in data representation allows achieving the same durability with fewer tapes by exploiting the redundancy properties of erasure codes rather than requiring full physical copies
4Reliability
If tape drives retry reads with repositioning when errors are encountered, then attempt to recover data is made, but tape deterioration and damage worsen
Solution Approach 1:
The patent performs preliminary distribution of encoded data chunks across multiple tape cartridges before any read operation. This preliminary encoding and distribution creates a resilient data structure where loss or damage of individual chunks does not result in total data loss. Consequently, the system can tolerate tape errors without needing aggressive retry mechanisms that would further damage the tape, as the distributed encoding already provides built-in error tolerance
Data Source
AI summary
Methods, apparatus, and other embodiments associated with doubly distributing erasure encoded data in a data storage system are described. One example apparatus includes a set of data storage devices and a set of logics that includes an encoding logic that generates an erasure encoded object that includes code-words, and chunks the code-words into code-word chunks, and a distribution logic that interleaves members of the set of code-word chunks into a plurality of records, and distributes the records across the data storage devices and within individual data storage devices. Example apparatus may include a read logic that reads the plurality of stored records from the data storage devices, and ignores read errors, and a repair logic that monitors the set of data storage devices, replaces or repairs failing data storage devices, generates replacement records, and stores the replacement records on a replacement data storage device.


