Geographically Diverse Erasure Coding With Protection Set Consolidation
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Geographically diverse data storage systems face challenges in efficiently recovering data due to high computing resource burdens and storage overhead, particularly when using erasure coding, which can be resource-intensive and inefficient in managing data access and storage across distant zones.
Innovation Solution
The system employs erasure coding with a 4+2 scheme, distributing data chunks across six zones, allowing recovery from any two inaccessible zones, and adapts by redistributing chunks during scale-out events to combine complementary protection sets, reducing storage overhead from 50% to 25% without re-encoding, using matrix operations and coding matrixes to maintain data protection and resilience.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If erasure coding is used to protect data across geographically diverse zones, then data resilience and protection are improved, but computing resource burden and storage overhead increase
Solution Approach 1:
The system segments data into multiple chunks and distributes them across different geographically diverse zones. Each zone stores a portion of the erasure-coded data, allowing the system to protect data while reducing the computing burden on any single node by dividing the overall encoding and recovery workload across multiple locations
Solution Approach 2:
The patent introduces geographic location as an additional dimension for data distribution beyond traditional storage tiers. By organizing data chunks across geographically diverse zones rather than just across storage devices, the system achieves data protection while distributing computing resources more effectively across distributed locations
2Reliability
If erasure coding with 4+2 scheme is used across six zones, then data recovery capability is improved, but storage overhead increases to 50%
Solution Approach 1:
The system merges multiple erasure-coded protection sets by consolidating data chunks from different zones. When zones are combined, the system can reuse existing encoded chunks across protection sets, reducing redundant storage requirements and lowering overall storage overhead while maintaining the same data recovery capability
Solution Approach 2:
The patent makes storage chunks universal by allowing the same encoded chunk to serve multiple protection sets simultaneously. A single chunk can contribute to the protection of multiple different data sets across zones, increasing the utility of stored data and reducing the total quantity of storage resources needed
3Adaptability or versatility
If data chunks are redistributed during scale-out events, then system adaptability is improved, but data loss risk increases during transitions
Solution Approach 1:
The system performs preliminary validation and verification actions before completing the redistribution of data chunks during scale-out events. By checking data integrity and verifying protection set validity before and during transitions, the system enables adaptability to scaling events while minimizing data loss risk through proactive error detection and correction
Solution Approach 2:
The patent implements feedback mechanisms that monitor the state of protection sets and data chunks during scale-out events. This feedback allows the system to detect potential data loss risks during redistribution and trigger corrective actions, enabling safe adaptation to changing system configurations
Data Source
AI summary
Erasure coding for scaling-out of a geographically diverse data storage system is disclosed. Chunks can be stored according to a first erasure coding scheme in zones of a geographically diverse data storage system. In response to scaling-out the geographically diverse data storage system, chunks can be moved to store data in a more diverse manner. The more diverse chunk storage can facilitate changing storage from the first erasure coding scheme to a second erasure coding scheme. The second erasure coding scheme can have a lower storage overhead than the first erasure coding scheme. In an aspect, the erasure coding scheme change can occur by combining erasure coding code chunks having complementary coding matrixes. Combining erasure coding code chunks having complementary coding matrixes can consume fewer computing resources than re-encoding data chunks for the second erasure coding scheme in a conventional manner.


