Clustered Storage Deduplication via Data Source Identifiers
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Clustered storage systems face inefficiencies in deduplication due to uneven distribution of data across nodes, leading to a loss of deduplication efficiency, as multiple generations of related data are stored on different nodes, complicating data management in a global namespace.
Innovation Solution
The system employs data source identification to intelligently distribute backup files by associating them with a common data protection set, ensuring that data from the same source is stored on the same deduplication domain, thereby increasing deduplication efficiency by grouping related data together.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Ease of operation
If data chunks are evenly distributed across cluster nodes, then storage capacity and resource utilization are balanced, but deduplication efficiency deteriorates
Solution Approach 1:
The patent applies local quality by creating deduplication domains with specialized characteristics for handling related data. Instead of uniform distribution, data from the same source is routed to specific deduplication domains where deduplication operations can be performed efficiently, while maintaining overall system balance through multiple domains working in parallel.
Solution Approach 2:
The patent segments the storage system into multiple deduplication domains, each capable of independent deduplication operations. This segmentation allows the system to maintain both balanced resource utilization across domains and high deduplication efficiency within each domain by keeping related data together.
2Productivity
If multiple generations of related data are stored on different nodes, then distribution evenness improves, but deduplication efficiency is lost
Solution Approach 1:
The patent implements preliminary action by assigning data to deduplication domains based on data source identification before actual storage. This pre-assignment ensures that related data generations are directed to the same domain in advance, enabling deduplication operations to find and eliminate duplicates efficiently without requiring post-storage redistribution.
Solution Approach 2:
The patent introduces deduplication domains as intermediary layers between the global namespace and physical storage nodes. These domains act as mediators that receive data with source identifiers and route related data to appropriate domains, maintaining both distribution efficiency and deduplication capability.
3Ease of operation
If a global namespace is used to represent the storage system, then external access is simplified, but data source tracking becomes difficult
Solution Approach 1:
The patent implements a nested structure where deduplication domains are embedded within the global namespace. The global namespace provides simplified external access, while nested deduplication domains maintain data source tracking information. This nested architecture allows both simplicity for external systems and complexity for internal deduplication operations to coexist.
Data Source
AI summary
Described is a system (and method) that intelligently distributes data within a clustered storage environment. To provide such a capability, the system may distribute backup files by considering a source of the data to be backed-up. In particular, the system may leverage the ability of front-end components such as a backup application to perform a granular data source identification of data. Such information may be propagated to back-end components such as a storage filesystem in the form of a data source identifier (e.g. placement tag). The data source identifiers may then be accessed by the clustered storage system to intelligently distribute backup files amongst a set of storage nodes forming a cluster. For example, backup files from the same data source may be stored on the same storage node to obtain the same deduplication efficiency as a single storage system.


