Integrated Deduplication Index for Source-Target Data Storage
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing deduplication solutions do not allow sharing of data chunks generated by deduplication operations on either the source or target, leading to inefficient storage and processing, as they typically require separate appliances or incomplete deduplication processes.
Innovation Solution
An integrated approach that enables seamless switching between source and target deduplication activities, using a shared deduplication index and process to identify and store unique data chunks, allowing reuse across multiple nodes and files, with policies to determine the deduplication location based on factors like time, system load, and file type.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If separate deduplicating appliances are deployed for source and target, then deduplication can be performed at both locations, but device complexity and cost increase
Solution Approach 1:
The patent merges source-side and target-side deduplication functionalities into a single integrated system. The storage system performs both source deduplication (on incoming data streams) and target deduplication (on stored data) using unified hardware resources, eliminating the need for separate appliances while maintaining comprehensive deduplication coverage.
Solution Approach 2:
The storage system is designed with multi-functional capability to perform multiple deduplication operations simultaneously. The same storage array and processing units can switch between acting as a deduplication source for external systems and as a deduplication target for backup operations, maximizing resource utilization and reducing overall system complexity.
2Loss of energy
If data is deduplicated at source only, then network bandwidth is reduced, but storage capacity at target is not optimized
Solution Approach 1:
The system implements continuous deduplication at both source and target locations. Source deduplication continuously reduces network traffic by eliminating duplicates before transmission, while target deduplication continuously optimizes storage capacity by eliminating duplicates after reception, ensuring both network and storage resources are efficiently utilized at all times.
3Quantity of substance
If data is deduplicated at target only, then storage capacity is optimized, but network bandwidth is not reduced
Solution Approach 1:
The system performs preliminary deduplication at the source before data enters the network, reducing the amount of data that needs to be transmitted. This preliminary action decreases network bandwidth consumption while the target subsequently performs additional deduplication to optimize storage capacity, addressing both concerns in sequence.
4Ease of operation
If chunks are not shared between source and target, then deduplication processes are independent, but storage efficiency is reduced
Solution Approach 1:
The patent introduces a shared index as an intermediary data structure that enables coordination between source and target deduplication processes. The index stores metadata about deduplicated chunks and is accessible by both source and target operations, allowing them to share deduplication knowledge and achieve higher storage efficiency while maintaining operational simplicity through a unified management interface.
Data Source
AI summary
One aspect of the present invention includes a configuration of a storage management system that enables the performance of deduplication activities at both the client (source) and at the server (target) locations. The location of deduplication operations can then be optimized based on system conditions or predefined policies. In one embodiment, seamless switching of deduplication activities between the client and the server is enabled by utilizing uniform deduplication process algorithms and accessing the same deduplication index (containing information on the hashed data chunks). Additionally, any data transformations on the chunks are performed subsequent to identification of the data chunks. Accordingly, with use of this storage configuration, the storage system can find and utilize matching chunks generated with either client- or server-side deduplication.


