Hierarchical Data Protection for Distributed Storage Overhead
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current data storage systems face inefficiencies due to excessive storage overhead and inability to handle multiple failure types when using a single data protection scheme, leading to unacceptable repair loads and vulnerabilities.
Innovation Solution
Implementing a policy-based hierarchical data protection method that determines data protection schemes at different levels of a storage system using Information Lifecycle Management (ILM) policies, allowing for the combination of multiple protection schemes at various layers to tailor data protection to performance, reliability, and storage overhead requirements.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If replication is used for data protection, then data availability is improved, but storage overhead increases significantly
Solution Approach 1:
The patent segments data protection into multiple hierarchical levels: object level (within storage nodes), storage node level (across nodes), and site level (across geographic locations). Each level applies appropriate protection schemes (RAID, erasure coding, replication) tailored to specific failure modes, avoiding unnecessary overhead at each层级.
Solution Approach 2:
Different data protection schemes are applied locally at different hierarchical levels based on specific requirements. For example, RAID 5/6 is applied at the object level for disk failures, erasure coding at the storage node level for node failures, and selective replication at the site level for geographic failures. This localized approach optimizes the balance between reliability and storage overhead.
2Device complexity
If a single data protection scheme is used across all levels, then system complexity is reduced, but ability to support multiple failure types deteriorates
Solution Approach 1:
The patent implements a universal hierarchical data protection architecture that can adapt to multiple failure types through a single unified framework. The system universally applies protection at multiple levels (object, node, site) and can selectively activate different schemes (RAID, erasure coding, replication) based on failure scenarios, achieving multi-functionality without proportionally increasing complexity.
Solution Approach 2:
The data protection system is dynamic and adaptive, automatically selecting appropriate protection schemes at each hierarchical level based on configured policies and detected failure conditions. The system can dynamically adjust the level and type of protection applied, transitioning between different protection modes as needed to handle various failure types effectively.
3Reliability
If full object copies are stored for node and site failure protection, then reliability against catastrophic failures is improved, but storage overhead and repair load increase
Solution Approach 1:
The patent applies copying selectively at the site level rather than universally. Full object copies are created and distributed to remote sites only when configured for geographic redundancy, while local storage nodes use more space-efficient schemes like erasure coding. This selective copying approach provides catastrophic failure protection where needed while minimizing overall storage overhead.
Solution Approach 2:
The patent implements nested data protection where multiple protection schemes are layered hierarchically. Inner layers (object-level RAID) protect against disk failures, middle layers (node-level erasure coding) protect against node failures, and outer layers (site-level replication) protect against catastrophic failures. Each layer provides protection for specific failure modes, creating a nested structure that achieves comprehensive protection with optimized storage overhead at each level.
Data Source
AI summary
A storage management computing device obtains an information lifecycle management (ILM) policy. A data protection scheme to be applied at a storage node computing device level is determined and a plurality of storage node computing devices are identified based on an application of the ILM policy to metadata received from one of the storage node computing devices and associated with an object ingested by the one of the storage node computing devices. The one of the storage node computing devices is instructed to generate one or more copies of the object or fragments of the object according to the data protection scheme and to distribute the object copies or one of the object fragments to one or more other of the storage node computing devices to be stored by at least the one or more other storage node computing devices on one or more disk storage devices.


