Pre-booted Replica VMs for Zero RTO Data Protection
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current data protection systems face challenges in reducing Recovery Time Objective (RTO) during data recovery operations, which can lead to significant damage and increased costs due to the time required to restore applications and systems.
Innovation Solution
Implementing a data protection system that replicates production virtual machines to replica virtual machines, allowing for pre-configured and pre-booted replica virtual machines to reduce RTO by sharing resources and optimizing replication strategies based on Quality of Service, failure domains, and application relationships, thereby minimizing the need for extensive startup and configuration processes.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If traditional data protection systems perform full recovery operations from backup media, then data can be restored, but the Recovery Time Objective (RTO) increases significantly due to extensive startup and configuration processes
Solution Approach 1:
The system pre-boots replica virtual machines before actual recovery is needed, performing startup and configuration processes in advance. This preliminary action ensures that when recovery is required, the replica VMs are already in a ready state, dramatically reducing RTO from minutes to seconds while maintaining full data restoration capability
Solution Approach 2:
The system creates and maintains replica virtual machines as copies of production virtual machines. These replicas are kept in a pre-booted state with synchronized data, allowing rapid failover without the need to perform full recovery operations from backup media, thus reducing RTO while ensuring data restoration capability
2Reliability
If multiple applications are replicated to separate dedicated virtual machines, then each application can be recovered independently, but the cost and resource overhead increase significantly
Solution Approach 1:
The system merges multiple replica virtual machines into a single shared infrastructure, allowing multiple applications to be replicated to the same physical or virtual host. This consolidation reduces resource overhead and cost while maintaining the ability to recover applications independently through virtualization isolation and selective activation
Solution Approach 2:
The shared virtual machine infrastructure is designed to serve multiple applications simultaneously. A single host can host multiple replica VMs that can be independently activated for different applications, providing universal recovery capability that reduces overall resource requirements while maintaining independent recovery for each application
3Loss of time
If replica virtual machines are kept in a pre-booted state to reduce RTO, then recovery time decreases, but the cost of maintaining ready-state replicas increases
Solution Approach 1:
The system dynamically manages the state of replica virtual machines, transitioning them between pre-booted and suspended states based on recovery priorities and resource availability. Critical applications maintain pre-booted replicas for immediate recovery, while less critical applications use suspended replicas that can be activated quickly, optimizing the balance between RTO and maintenance cost
Solution Approach 2:
The system applies different quality levels to different replica VMs based on application criticality. High-priority applications receive full pre-booted replicas with all resources allocated, while lower-priority applications use streamlined replicas with minimal resources. This local differentiation reduces overall cost while maintaining fast recovery for critical systems
Data Source
AI summary
Data protection operations including replication operations from a production site to a replica site are disclosed. An example method assessing applications operating on a production virtual machine based on a replication strategy. The replication strategy is configured to identify related applications and ensure that the related applications are replicated to different replica virtual machines. The applications are then replicated from the production virtual machines to the replica virtual machines according to the replication strategy. The replication strategy can improve performance of the recovery operation.


