Application Stack Recovery Time Estimation Across Cloud Regions
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
In the event of a disaster in a primary cloud region, identifying and orchestrating the recovery of enterprise applications to a standby region is difficult and time-consuming, with existing solutions lacking efficiency and accuracy.
Innovation Solution
A recovery protection group (RPG) is defined to manage cloud resources, enabling automatic generation of recovery plans and simulated drills, with deep introspection to identify resource characteristics, and providing accurate recovery time estimates.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Productivity
If manual identification and orchestration of cloud resources is used for recovery, then flexibility and control are maintained, but the process becomes time-consuming and error-prone
Solution Approach 1:
The system enables self-service by automatically identifying cloud resources, generating recovery plans, and executing recovery operations without requiring manual intervention. The recovery protection group autonomously monitors its members, detects outages, and orchestrates the recovery process, transforming a complex manual task into an automated self-service operation that speeds up recovery while reducing errors.
Solution Approach 2:
The system performs preliminary actions by pre-identifying cloud resources, pre-generating recovery plans, and pre-configuring recovery parameters before an outage occurs. The recovery protection group establishes recovery procedures in advance, so when an outage happens, the recovery process can be executed immediately without time-consuming on-site analysis and decision-making.
2Reliability
If comprehensive identification of all cloud resources is performed, then recovery completeness is improved, but the time required for identification and orchestration increases
Solution Approach 1:
The system performs preliminary identification and cataloging of all cloud resources within the recovery protection group before an outage occurs. By maintaining an up-to-date inventory of resources, their locations, and inter-dependencies in advance, the system ensures complete recovery identification without time-consuming analysis during the actual recovery operation.
Solution Approach 2:
The system continuously monitors the status of cloud resources and updates the recovery protection group's knowledge base in real-time. This feedback mechanism ensures that the system always has current information about resource locations, statuses, and relationships, enabling complete and accurate recovery identification without repeated time-consuming surveys during outages.
Data Source
AI summary
Techniques for estimating a time to recover an application stack from a primary region to a standby region. In one technique, a recovery plan that comprises a plurality of actions to perform relative to a plurality of cloud resources in a recovery protection group is selected. Historical data that indicates actual times to perform one or more actions pertaining to recovering cloud resources is stored. Based on the historical data, a total time to execute the recovery plan is estimated. The total time is stored in association with the recovery plan.


