Policy-Based Hierarchical Data Protection for Rebuild Resilience
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current data storage systems face challenges with excess storage overhead, high repair loads, and inability to support multiple failure types due to reliance on single data protection schemes, which can leave systems vulnerable to additional failures during rebuild processes.
Innovation Solution
Implementing a policy-based data protection system that uses a storage management computing device to evaluate information lifecycle management policies and dynamically select appropriate data protection schemes, such as replication or erasure coding, at both storage node and disk levels, based on metadata and application requirements, to ensure high availability and failure protection with optimized resource utilization.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Ease of manufacture
If replication is used for data protection, then simplicity of implementation is improved, but storage overhead increases due to multiple copies
Solution Approach 1:
The patent implements dynamic data protection by allowing different data protection schemes (replication, RAID, erasure coding) to be applied at different hierarchical levels (storage node, disk, storage pool) rather than using a single static scheme throughout the system. This dynamic approach optimizes the balance between storage overhead and protection effectiveness.
Solution Approach 2:
The patent segments the data protection function into multiple hierarchical levels: storage node level, disk level, and storage pool level. Each level can independently apply appropriate data protection schemes, allowing the system to reduce overall storage overhead while maintaining comprehensive protection through layered defense.
2Quantity of substance
If RAID protection schemes are used, then storage overhead is reduced, but repair load and system vulnerability during rebuild increases
Solution Approach 1:
The system dynamically adjusts data protection strategies during failure recovery events. When a failure is detected, the patent can temporarily enhance protection at higher hierarchical levels or use erasure coding with higher redundancy ratios during the rebuild process, then return to normal operation after recovery, thus managing vulnerability during critical rebuild periods.
Solution Approach 2:
The patent implements prior cushioning by maintaining spare capacity and pre-configured recovery resources at the storage pool level. Erasure coding schemes are designed with built-in redundancy that provides a cushion against multiple failures, and the system pre-allocates resources for rapid rebuild operations to minimize the vulnerable window during recovery.
3Quantity of substance
If erasure coding is used, then storage overhead is reduced compared to replication, but repair bandwidth requirements increase
Solution Approach 1:
The patent segments erasure coding operations across multiple hierarchical levels and storage pools. By dividing the erasure coding process into smaller, localized operations at the disk level and storage node level, the system reduces the bandwidth requirements for any single repair operation while maintaining overall data protection through the hierarchical structure.
4Reliability
If hierarchical data protection with replication at node level and RAID/DDP at disk level is implemented, then protection against node and site failures is improved, but storage overhead increases due to full object copies
Solution Approach 1:
The patent changes the fundamental parameter of data protection by replacing full object replication with erasure coding at the storage pool level. This parameter change allows the system to achieve the same level of protection against node and site failures while significantly reducing storage overhead, as erasure coding distributes redundant information more efficiently than complete object copies.
Data Source
AI summary
A storage management computing device obtains an information lifecycle management (ILM) policy based on a query from a storage node ingesting an object into a distributed storage system. A hierarchical data protection plan comprising protection schemes at different layers of the distributed storage system is determined. A data protection scheme to be applied at a storage node computing device level is determined and a plurality of storage node computing devices are identified based on the hierarchical data protection plan. The storage management computing device instructs the ingesting storage node to store the object into the distributed storage system according to the hierarchical data protection plan.


