Data Replication Using Synthetic File Relationships
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing data replication methods across namespaces fail to effectively utilize synthetic information, leading to inefficient replication processes due to the inability to track and propagate synthetic relationships between files, resulting in unnecessary file copying and loss of deduplication benefits.
Innovation Solution
A method is introduced to track and correct base file information using a dictionary that records source and destination file relationships during cloning operations, ensuring synthetic information is updated to point to correct base files within the same namespace, enabling efficient replication by leveraging deduplication advantages.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If traditional file copying is used for replication across namespaces, then file replication can be performed, but synthetic information is lost and deduplication benefits cannot be utilized
Solution Approach 1:
The system performs preliminary actions by tracking and recording file cloning operations in a dictionary before replication occurs. This allows the replication process to later reference the dictionary and correctly identify base files, ensuring synthetic information is preserved and deduplication can be applied during replication.
Solution Approach 2:
A dictionary data structure is introduced as an intermediary between the file cloning operation and the replication process. The dictionary stores mappings between cloned files and their base files, enabling the replication system to retrieve synthetic information and maintain deduplication relationships across namespaces.
2Ease of manufacture
If file cloning is performed without tracking, then cloning operation is simple, but base file relationships are lost and synthetic replication cannot occur
Solution Approach 1:
The file cloning process automatically updates the dictionary with base file relationships without requiring manual intervention. The system self-services by tracking its own cloning operations, maintaining the information needed for synthetic replication while keeping the cloning operation itself simple and transparent to users.
3Productivity
If synthetic information is not tracked during cloning, then replication process is straightforward, but redundant data transfer occurs and deduplication benefits are lost
Solution Approach 1:
The system implements feedback by having the replication process query the dictionary to determine whether files have been cloned and what their base file relationships are. This feedback mechanism enables the replication system to avoid redundant data transfer by leveraging existing synthetic relationships and applying deduplication where applicable.
Data Source
AI summary
Methods of cloning data backup across namespaces are disclosed. One or more source files are cloned from a first namespace to a second namespace, as one or more destination files. When the cloning of the source file(s) is performed, a data structure including source file information and destination file information is generated. A source synthetic file is cloned from the first namespace to the second namespace, as a destination synthetic file, where the source synthetic file uses the source file(s) as one or more base files on the first namespace. When the cloning of the source synthetic file is performed, the data structure is looked up to obtain the source file information and the destination file information. Based on the source file information and the destination file information, synthetic information of the destination synthetic file is updated to use the destination file(s) as one or more base files on the second namespace.


