Snapshot-Based Hydration of Cloud Storage via DAG Replication
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Traditional storage systems face inefficiencies in data replication and management, particularly in handling snapshots and data distribution across multiple storage arrays and cloud services, leading to complexities in data integrity and availability.
Innovation Solution
A system that employs a directed acyclic graph (DAG) of mediums and medium mapping tables to efficiently replicate data, utilizing deduplication and compression, and enabling seamless data distribution across storage arrays and cloud services, ensuring data integrity and availability.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If traditional data replication methods are used across multiple storage arrays and cloud services, then data distribution is achieved, but data redundancy increases and management complexity increases
Solution Approach 1:
The patent combines multiple storage systems (local storage arrays and cloud services) into a unified storage pool managed by a single controller. This merging eliminates data redundancy by creating a single source of truth while maintaining high availability through the distributed architecture, directly resolving the contradiction between reliability and data redundancy
Solution Approach 2:
The storage controller provides universal management capabilities across heterogeneous storage systems, enabling a single system to manage both local and cloud-based storage resources. This multi-functionality allows the system to achieve data availability without duplicating data across multiple independent systems, thereby reducing redundancy while maintaining reliability
2Reliability
If traditional data replication methods are used across multiple storage arrays and cloud services, then data distribution is achieved, but system complexity increases
Solution Approach 1:
The patent introduces a central storage controller as an intermediary that manages all data operations between clients and the distributed storage system. This mediator abstracts the complexity of managing multiple storage arrays and cloud services, providing simplified data management while ensuring data availability across the distributed infrastructure
Solution Approach 2:
The system segments storage resources into a modular distributed architecture where data is divided into blocks stored across multiple storage units. This segmentation, combined with centralized control, allows the system to achieve high availability through distribution while maintaining manageable complexity through the controller's coordination of all storage operations
3Reliability
If snapshots are replicated to cloud services, then data availability is improved, but data distribution complexity increases
Solution Approach 1:
The system creates snapshots locally at the storage array before replicating them to cloud services. This preliminary action ensures that snapshot availability is maintained locally for immediate recovery operations, while cloud replication provides backup availability without requiring complex real-time synchronization mechanisms
Data Source
AI summary
Systems, methods, and computer readable storage mediums for snapshot-based hydration of a cloud-based storage system, including: storing, in a cloud computing environment, a snapshot of a dataset that is stored on a separate storage system, wherein the snapshot includes a self-described copy of the dataset such that the dataset can be reconstructed without accessing the separate storage system; creating, in a cloud computing environment, at least a portion of a cloud-based storage system; and populating, from the snapshot that is stored in the cloud computing environment, at least a portion of a storage layer within the cloud-based storage system, wherein the cloud-based storage system can service I/O operations to the dataset after the storage layer has been populated.


