Distributed Deduplication Across Storage Clusters
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Data backup systems face challenges in efficiently using storage space due to limited deduplication of data across objects and in meeting recovery time objectives (RTOs) due to I/O bandwidth limitations, leading to potential storage cluster inefficiencies and extended restore times.
Innovation Solution
A system dynamically selects a storage distribution mode between deduplication and restore modes based on optimization criteria, such as deduplication efficiency and RTO, to balance data storage across multiple deduplication domains, allowing for efficient deduplication and rapid data restoration by distributing data across multiple storage clusters for parallel restoration.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Quantity of substance
If data is stored in a single deduplication domain to maximize deduplication efficiency, then storage space utilization is improved, but restore time increases due to I/O bandwidth limitations
Solution Approach 1:
The patent divides data into multiple data strips and distributes them across multiple deduplication domains. Each data strip can be restored in parallel from different storage clusters, thereby reducing overall restore time while maintaining deduplication efficiency through distributed deduplication operations.
Solution Approach 2:
The patent introduces a new dimension of data distribution by organizing data across multiple deduplication domains rather than confining it to a single domain. This multi-dimensional storage architecture enables parallel restore operations while preserving deduplication benefits through coordinated deduplication across domains.
2Loss of time
If data is distributed across multiple deduplication domains to enable parallel restoration, then restore time is reduced, but storage efficiency decreases due to limited deduplication across domains
Solution Approach 1:
The patent merges deduplication operations across multiple deduplication domains by coordinating deduplication processes and sharing deduplication metadata. This allows the system to maintain high deduplication efficiency even when data is distributed across multiple domains, as deduplication operations can identify and eliminate redundant data across the entire distributed system.
Solution Approach 2:
The patent implements feedback mechanisms where deduplication information from one deduplication domain is used to optimize storage and retrieval in other domains. This cross-domain deduplication feedback enables the system to maintain storage efficiency while benefiting from parallel restore capabilities across multiple domains.
3Productivity
If more storage clusters are added to increase restore throughput, then restore speed is improved, but system complexity and cost increase
Solution Approach 1:
The patent designs the distributed deduplication system to serve multiple functions: it enables parallel restore operations, maintains deduplication efficiency, and provides scalable storage capacity. The same distributed architecture that enables fast restore also provides flexible storage expansion without proportionally increasing system complexity, as the system can dynamically allocate resources across multiple deduplication domains.
Data Source
Figure 1
Figure 2
Figure 3
AI summary
A plurality of objects sharing one or more common attributes are identified. A storage distribution mode for the identified objects sharing the one or more common attributes is determined based at least in part on one or more optimization criteria. The storage distribution mode is caused to be implemented by one or more of a plurality of storage clusters.