Disaggregated Datacenter Disaster Recovery via Dynamic Workload Replication
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Traditional disaster recovery systems in cloud computing are expensive and inflexible, requiring hot standby datacenters and dedicated servers, which waste resources since disasters are rare, and cannot provide immediate operation continuation without interruption.
Innovation Solution
Implementing a disaggregated computing system that continuously replicates workload and state data from a primary site to a secondary site using a direct point-to-point connection, allowing compute resources to be connected instantly in case of a disaster, and utilizing different service level agreements (SLAs) to efficiently allocate resources during failover.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If hot standby datacenters and dedicated servers are used for disaster recovery, then reliability is improved, but device complexity and cost increase
Solution Approach 1:
The patent creates a virtual copy of the primary datacenter environment at the secondary site through virtualization. Instead of maintaining physical hot standby datacenters, the system uses virtual machine images and state data replication to create a functional duplicate that can be activated immediately upon disaster occurrence, thereby maintaining reliability while reducing device complexity
Solution Approach 2:
The secondary datacenter infrastructure is designed to serve multiple purposes: it can function as a disaster recovery site, a test environment, or a development platform. This multi-functionality eliminates the need for dedicated hot standby infrastructure, reducing overall system complexity while maintaining disaster recovery capability
2Reliability
If hot standby datacenters are maintained continuously, then immediate operation continuation is achieved, but resource waste increases
Solution Approach 1:
The system dynamically adjusts the state of the secondary datacenter based on disaster recovery requirements and resource availability. Virtual machine images and state data are replicated continuously, but full computational resources are only activated when needed for failover, allowing the system to maintain operation continuity capability while minimizing resource consumption during normal operations
Solution Approach 2:
Critical data replication and virtual machine image preparation are performed in advance during normal operations. When disaster occurs, the pre-prepared state data and images enable immediate restoration without requiring real-time resource allocation, thus achieving operation continuity while avoiding continuous resource waste
3Reliability
If dedicated standby servers are used, then disaster recovery reliability is improved, but cost increases
Solution Approach 1:
The patent merges the disaster recovery function with other datacenter functions by allowing the secondary site to simultaneously serve as a test environment, development platform, or additional capacity resource. This consolidation eliminates the need for separate dedicated standby servers, reducing resource allocation while maintaining disaster recovery reliability
Solution Approach 2:
The secondary datacenter infrastructure is designed to serve multiple purposes: disaster recovery, testing, development, and additional computational capacity. This multi-functionality allows the same physical resources to fulfill multiple roles, reducing the total quantity of resources needed compared to dedicated standby servers
Data Source
AI summary
Embodiments for disaster recovery in a disaggregated computing system. A memory is allocated at a secondary, disaster recovery site for data received from a primary site. A degree of resiliency is defined for respective workloads associated with the data at the primary site to specify how critical each respective workload is to execute in case of disaster, and the data is replicated to the allocated memory at the disaster recovery site according to the degree of resiliency.


