Predictive VM Replication for Rapid Disaster Recovery
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current data protection systems face challenges in reducing Recovery Time Objective (RTO) and managing costs effectively, particularly in restoring applications and failing over to minimize data loss during disasters, as existing methods often require significant time to configure and boot virtual machines, leading to increased damage and costs.
Innovation Solution
Implementing a data protection system that replicates Input/Output (IO) from production virtual machines to replica virtual machines, where the replica machines are pre-configured and powered on demand, allowing for rapid boot and application start-up, and optimizing resource allocation by sharing replica virtual machines among multiple applications, using dynamic replication strategies and machine learning for predictive failure management.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If virtual machines are configured and booted during recovery operations, then data protection and application restoration are achieved, but Recovery Time Objective (RTO) increases leading to greater damage and costs
Solution Approach 1:
The system pre-configures replica virtual machines with application binaries and dependencies before failures occur. During normal operation, replica VMs are prepared with all necessary software components, so when a failure is detected, the recovery process can immediately activate the pre-configured replica without needing to perform configuration or installation steps, thereby dramatically reducing RTO while maintaining data protection
Solution Approach 2:
The system dynamically adjusts the state of replica virtual machines based on operational needs and failure conditions. Replica VMs can be transitioned between different states (prepared, active, standby) depending on the situation, allowing the system to optimize both protection and recovery speed by having replicas ready to activate immediately when needed
2Loss of time
If replica virtual machines are pre-configured and powered on demand, then Recovery Time Objective is reduced, but resource allocation and system complexity increase
Solution Approach 1:
The system creates a universal replica virtual machine template that can serve multiple applications and workloads. Instead of maintaining separate complex configurations for each application, a single replicated VM image can be dynamically instantiated and configured for different purposes, reducing overall system complexity while enabling rapid recovery across multiple scenarios
Solution Approach 2:
The system uses virtual machine replication to create copies of the replica VM that can be rapidly deployed. Rather than manually configuring each recovery instance, the system copies pre-configured VM images and instantiates them as needed, which simplifies the deployment process and reduces the operational complexity of managing multiple recovery instances
3Use of energy by moving object
If replica virtual machines are shared among multiple applications, then resource utilization is optimized, but the ability to handle simultaneous failures of different applications is reduced
Solution Approach 1:
The system dynamically allocates replica virtual machines based on real-time failure conditions and resource availability. When applications are running normally, replicas are shared to optimize resource utilization. When simultaneous failures occur, the system dynamically provisions additional replica instances or allocates dedicated resources, ensuring that availability requirements are met during critical failure scenarios while maintaining efficiency during normal operation
Data Source
AI summary
Data protection operations including replication operations that dynamically adapt a topology of replica virtual machines. A data protection system may implement a machine model that is trained using, as input, characteristics of virtual machines. When a failure is predicted, a topology of the replica virtual machines is changed. The topology may also change when changes in the environment are detected. The changes may include redistributing the protected applications to the replica virtual machines and/or scaling the replica virtual machines.


