MDS Data Chunk Storage Across Multiple Data Centers
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional data storage systems face challenges in ensuring data accessibility and redundancy across multiple data centers, particularly when failures occur, as they require extensive storage capacity and time-consuming reconstruction processes.
Innovation Solution
An encoding system that strategically stores data chunks and error-correcting code chunks across multiple data centers, using a systematic maximal-distance separable (MDS) code, allowing data to be accessed from two centers without reconstruction and reconstructed when f data centers are inaccessible, with each data chunk stored at a home and distinct secondary data center, and each code chunk stored at a distinct third data center.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If data is stored redundantly across multiple data centers, then data reliability is improved, but storage capacity requirements increase
Solution Approach 1:
The patent changes the parameter of data distribution by using systematic MDS codes to encode data into (n-f) data chunks and (f-1) code chunks, allowing flexible control over redundancy level while minimizing storage requirements. This mathematical encoding approach optimizes the balance between reliability and storage capacity.
Solution Approach 2:
Instead of storing complete redundant copies of data across multiple data centers, the patent creates encoded copies (chunks) that collectively represent the original data. Any (n-f) data chunks or combination of data and code chunks can reconstruct the original data, reducing total storage capacity while maintaining reliability.
2Reliability
If data is stored with extensive redundancy, then data accessibility is improved, but reconstruction time increases
Solution Approach 1:
The patent segments data into (n-f) data chunks and generates (f-1) separate code chunks using MDS encoding. This segmentation allows parallel access to multiple chunks from different data centers simultaneously, reducing reconstruction time while maintaining accessibility even when f data centers are unavailable.
Solution Approach 2:
The patent performs preliminary encoding of data into chunks and code chunks before storage, organizing them across different data centers according to a block design. This preliminary arrangement ensures that when data needs reconstruction, the required chunks are already positioned and ready for rapid assembly without time-consuming processing.
3Reliability
If data chunks are stored at multiple data centers, then data availability is improved, but system complexity increases
Solution Approach 1:
The patent creates a universal storage scheme where the same MDS encoding and block design principles apply regardless of the number of data centers or failure scenarios. The systematic approach using (n-f) data chunks and (f-1) code chunks provides a unified method for handling data distribution, storage, and reconstruction across any number of data centers, reducing operational complexity.
Data Source
AI summary
Methods, systems, and apparatus, including computer programs encoded on a computer storage medium, for storing data reliably across groups of storage nodes. In one aspect, a method includes receiving (n−f) data chunks for storage across n groups of storage nodes and generating (f−1) error-correcting code chunks using an error-correcting code and the (n−f) data chunks. The (n−f) data chunks are stored at a first group of storage nodes. Each data chunk of the (n−f) data chunks is stored at a respective second group of storage nodes. Each code chunk of the (f−1) code chunks is stored at a respective third group of storage nodes. Each second group of storage nodes and each third group of storage nodes is distinct from each other and from the first group of storage nodes.


