Seed Data Selection for Deduplicated Storage Replication
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
In data replication between remote data storage systems, using slower, less expensive communication links can lead to prolonged data recovery times when one system becomes unavailable, especially when deduplicated data stores are empty or initially replicating to a new target, as only un-replicated data is efficiently communicated over slower links.
Innovation Solution
The implementation of a method that uses removable storage to provide seed data, where a deduplication engine identifies preferred manifests or combinations of manifests based on chunk identifier duplication levels, allowing efficient replication by transferring seed data over high-bandwidth connections and populating the deduplicated data store, thereby reducing recovery times.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Loss of energy
If data is replicated over slower, less expensive communication links, then cost is reduced, but data recovery time increases
Solution Approach 1:
The patent applies preliminary action by transferring seed data to the replacement data storage system before the actual replication process begins. The seed data is selected from the source data storage system and transferred over a fast communication link to populate the deduplicated data store of the replacement system, enabling subsequent efficient replication over slower links.
2Productivity
If deduplicated data store is empty during replication, then data can be replicated, but replication time increases significantly
Solution Approach 1:
The patent implements preliminary action by pre-transferring seed data to populate the deduplicated data store before the main replication process. This ensures that when replication begins, the replacement data storage system already has a foundation of data, dramatically reducing the time required to replicate subsequent data over slower communication links.
Solution Approach 2:
The patent uses an intermediary approach by introducing a seed data transfer mechanism that acts as a bridge between the source and replacement systems. The seed data serves as an intermediary foundation that enables the replacement system to efficiently receive and process subsequent replication data without waiting for the entire data set to be transferred.
3Productivity
If only un-replicated data is communicated over slow link, then replication is efficient, but initial replication to empty store is slow
Solution Approach 1:
The patent applies preliminary action by performing the seed data transfer before the main replication process. This preliminary transfer populates the deduplicated data store with essential data, allowing subsequent replication operations to proceed efficiently over slower communication links without the penalty of transferring entire data sets from an empty store.
Data Source
AI summary
There is disclosed a computer system operable to process a plurality of logical storage unit manifests the manifests comprising respective pluralities of chunk identifiers identifying data chunks in a deduplicated data chunk store The computer system can determine at least one preferred manifest or preferred combination of manifests according to levels of duplication of the chunk identifiers within respective said manifests, and/or within respective combinations of said manifests. The computer system can provide preferred seed data corresponding to data chunks identified by the at least one preferred manifest or preferred combination of manifests. A method and computer readable medium are also disclosed. At least some embodiments facilitate timely and convenient transfer and storage of relevant data chunks to a receiving deduplicated data chunk store of a data storage system.


