Disaggregated Datacenter Memory Replication for Failover
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Traditional disaster recovery systems in cloud computing are costly and inefficient, requiring hot standby datacenters and dedicated servers for continuous data mirroring, which wastes resources and is not cost-effective for rare disaster scenarios.
Innovation Solution
Implementing a disaggregated computing system that continuously replicates workload and state data from a primary site to a secondary site using a point-to-point connection between memory pools, without requiring compute resources at the secondary site, allowing instant allocation of compute resources during failover and efficient resource utilization.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If hot standby datacenters and dedicated servers are used for continuous data mirroring, then data replication reliability is improved, but resource utilization deteriorates and costs increase
Solution Approach 1:
The system dynamically allocates compute resources at the secondary site based on disaster occurrence. During normal operations, no compute resources are allocated, but memory resources are pre-allocated for potential data reception. Upon disaster detection, compute resources are immediately assigned to the memory resources to handle failover workloads, eliminating the need for continuous resource dedication while maintaining reliability
Solution Approach 2:
The invention extracts the compute resource requirement from the data replication process. Memory resources are continuously allocated and data is continuously replicated to the secondary site without requiring compute resources to be attached. This separation allows memory to remain in a low-power state during normal operations while still maintaining replication capability
2Reliability
If hot standby datacenters are maintained for disaster recovery, then operational continuity is improved, but device complexity and costs increase
Solution Approach 1:
The system segments computing resources into separate pools: memory resources and compute resources. This segmentation allows memory to be continuously allocated and pre-positioned at the secondary site without requiring associated compute resources. The separation simplifies the disaster recovery architecture by eliminating the need for complete standby server systems while maintaining operational continuity capability
3Speed
If dedicated servers are used for continuous data mirroring, then data replication speed is improved, but resource utilization deteriorates
Solution Approach 1:
Memory resources are preliminarily allocated at the secondary site before any disaster occurs. This preliminary allocation of memory (without compute resources) enables immediate data replication capability to be established, allowing fast replication speed when needed while avoiding the continuous resource consumption that would occur with pre-configured dedicated servers
Data Source
AI summary
Embodiments for disaster recovery in a disaggregated computing system. Memory resources are allocated at a secondary, disaster recovery site for data received from a primary site. The data from the primary site is continuously replicated to the allocated memory resources at the disaster recovery site without requiring any compute resources to be attached to the allocated memory resources. Responsive to determining a disaster recovery failover is in progress, the compute resources are assigned to the allocated memory resources for performing a failover workload, and the failover workload is executed at the disaster recovery site.


