Hypervisor Recovery Time Minimization via Volume Assignment
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current data protection systems face challenges in minimizing downtime during recovery and keeping pace with data transactions at production sites, leading to potential system shutdowns due to backlog of un-logged data transactions.
Innovation Solution
A computer-implemented method and system that determines recovery time for hypervisors, assigns volumes to replication appliances, and manages IOs to minimize recovery time, using a splitter and replication appliances to replicate data and maintain data consistency across sites.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If journaling is used to enable continuous data protection and rollback capability, then the ability to recover to any specified point in time is improved, but the overhead of multiple data transactions at the backup site causes the backup site to fall behind when data transaction rates are high
Solution Approach 1:
The patent segments the backup site into multiple independent backup sites, allowing parallel processing of data transactions. Each backup site can handle a portion of the journaling overhead independently, distributing the load and preventing any single site from becoming a bottleneck when data transaction rates are high.
Solution Approach 2:
The patent introduces a new dimension to the backup architecture by implementing multi-site replication instead of single-site journaling. This dimensional expansion allows the system to handle high data transaction rates by distributing journaling operations across multiple geographic or logical sites, thereby maintaining both reliability and productivity.
2Stability of the object's composition
If the backup site processes all data transactions to maintain continuous protection, then data consistency is improved, but the production site must slow down to prevent backlog buildup
Solution Approach 1:
The patent segments the journaling workload across multiple backup sites, allowing the production site to maintain high data transaction rates without overwhelming a single backup site. Each backup site processes a portion of the transactions independently, maintaining data consistency while preserving production speed.
Solution Approach 2:
The patent implements partial replication where not all data transactions need to be fully processed by every backup site simultaneously. Instead, the system uses a combination of full replication for critical data and partial/asynchronous replication for less critical data, maintaining consistency while allowing the production site to operate at full speed.
3Device complexity
If a single backup site is used for continuous data protection, then system complexity is reduced, but the backup site cannot keep pace with high data transaction rates
Solution Approach 1:
The patent merges multiple backup sites into a unified continuous data protection system, where each site functions independently but contributes to the overall protection capability. This merging approach increases total data processing capacity while maintaining relatively simple individual site structures, allowing the system to handle high transaction rates without excessive complexity.
Solution Approach 2:
The patent creates backup sites that serve multiple functions: they can independently handle journaling operations, provide rollback capabilities, and serve as disaster recovery sites. This multi-functionality allows the system to achieve high productivity with a modular architecture that doesn't significantly increase overall system complexity.
Data Source
AI summary
A computer implemented method, system, and computer program product for recovering from a crash of a system being replicated, the method comprising determining the amount of recovery time due to the crash of each of a set of hypervisors; wherein each of the hypervisors runs one or more data replication elements selected from the group consisting of a splitter and a replication appliance; wherein each of the splitters and replication appliances replicates one or more volumes, creating an assignment of the one or more volumes to the set of replication appliances and creating an assignment of one or more replication appliances to a set of hypervisors to minimize the amount of recovery time.


