Unified Replication Metadata Model for Data Recovery
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current data replication strategies complicate disaster recovery due to scattered metadata across multiple replication tools, leading to increased recovery time and resource-intensive indexing processes, which impact system availability and production efficiency.
Innovation Solution
A unified replication metadata model is generated to correlate metadata from different replication tools, allowing for the selection of a proper subset of replicas to index, thereby creating a unified content index that facilitates faster data recovery by identifying candidate replicas and their confidence scores.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If multiple replication-based strategies are deployed for data protection, then data resiliency and protection capability are improved, but metadata management complexity increases and recovery time extends
Solution Approach 1:
The patent combines multiple separate metadata sources from different replication strategies into a single unified metadata store. This unified store correlates metadata across synchronous replication, asynchronous replication, and backup operations, eliminating the need to query multiple separate metadata repositories and significantly reducing recovery complexity.
Solution Approach 2:
The patent introduces a unified metadata store as an intermediary layer between various replication strategies and the recovery process. This intermediary correlates and standardizes metadata from different sources, providing a single point of access for recovery operations without requiring changes to the underlying replication mechanisms.
2Manufacturing precision
If complete dataset crawling or mining is performed to build indexes, then indexing coverage is improved, but system resource usage increases and production is impacted
Solution Approach 1:
The patent performs indexing selectively on only those replicas that are relevant to recovery operations, rather than crawling or mining the complete dataset. By using the unified metadata store to identify which replicas contain needed data, the system indexes only necessary portions, reducing resource consumption while maintaining adequate indexing coverage for recovery purposes.
3Reliability
If all replicas are indexed to ensure complete data retrieval, then data recovery completeness is improved, but recovery time and resource usage increase
Solution Approach 1:
The patent pre-correlates metadata from multiple replication strategies in the unified metadata store before recovery operations are needed. This preliminary organization of metadata allows the system to quickly identify which specific replicas contain the required data during recovery, eliminating the need to examine or index all replicas and significantly reducing recovery time while ensuring completeness.
4Reliability
If aggressive replication strategies are used to meet resiliency requirements, then data protection capability is improved, but the rate of replica generation increases impacting system resources
Solution Approach 1:
The patent extracts and separates metadata management from the high-speed replication process. By pulling metadata out into a unified store that is independently managed and correlated, the system can maintain aggressive replication rates for data protection without proportionally increasing the overhead of metadata processing and indexing, thus improving resource efficiency.
Data Source
AI summary
An approach for managing replicated data is presented. A current usage of resources in a system and a threshold usage of the resources are determined. Based on inter-replica correlation(s) and inter-data correlation(s) specified by a unified replication metadata model, a proper subset of replicas included in a plurality of replicas is indexed by (i) if the current usage is less than the threshold usage, determining an expected additional resource usage due to performing an indexing task online and based on the expected additional resource usage, determining a resource affinity score for performing the indexing task online, or (ii) if the current usage is greater than or equal to the threshold usage, determining an expected resource usage due to performing the indexing task offline and based on the expected resource usage, determining a resource affinity score for performing the indexing task offline.


