Disaggregated Datacenter Memory Replication for Failover

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Traditional disaster recovery systems in cloud computing are costly and inefficient, requiring hot standby datacenters and dedicated servers for continuous data mirroring, which wastes resources and is not cost-effective for rare disaster scenarios.

Innovation Solution

Implementing a disaggregated computing system that continuously replicates workload and state data from a primary site to a secondary site using a point-to-point connection between memory pools, without requiring compute resources at the secondary site, allowing instant allocation of compute resources during failover and efficient resource utilization.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If hot standby datacenters and dedicated servers are used for continuous data mirroring, then data replication reliability is improved, but resource utilization deteriorates and costs increase

Engineering Contradiction:
Improvedata replication reliabilityVSAvoidresource waste
Core Design Contradiction:
ReliabilityVSLoss of energy

Solution Approach 1:

The system dynamically allocates compute resources at the secondary site based on disaster occurrence. During normal operations, no compute resources are allocated, but memory resources are pre-allocated for potential data reception. Upon disaster detection, compute resources are immediately assigned to the memory resources to handle failover workloads, eliminating the need for continuous resource dedication while maintaining reliability

Inventive Principle:
Principle #15Dynamics

Solution Approach 2:

The invention extracts the compute resource requirement from the data replication process. Memory resources are continuously allocated and data is continuously replicated to the secondary site without requiring compute resources to be attached. This separation allows memory to remain in a low-power state during normal operations while still maintaining replication capability

Inventive Principle:
Principle #2Taking out (Extraction)

2Reliability

If hot standby datacenters are maintained for disaster recovery, then operational continuity is improved, but device complexity and costs increase

Engineering Contradiction:
Improveoperational continuityVSAvoidsystem complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The system segments computing resources into separate pools: memory resources and compute resources. This segmentation allows memory to be continuously allocated and pre-positioned at the secondary site without requiring associated compute resources. The separation simplifies the disaster recovery architecture by eliminating the need for complete standby server systems while maintaining operational continuity capability

Inventive Principle:
Principle #1Segmentation

3Speed

If dedicated servers are used for continuous data mirroring, then data replication speed is improved, but resource utilization deteriorates

Engineering Contradiction:
Improvedata replication speedVSAvoidresource utilization
Core Design Contradiction:
SpeedVSProductivity

Solution Approach 1:

Memory resources are preliminarily allocated at the secondary site before any disaster occurs. This preliminary allocation of memory (without compute resources) enables immediate data replication capability to be established, allowing fast replication speed when needed while avoiding the continuous resource consumption that would occur with pre-configured dedicated servers

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS10983881B2Disaster recovery and replication in disaggregated datacenters
Publication Date: 2021.04.20 INTERNATIONAL BUSINESS MACHINE CORPORATION
  • US10983881B2 patent drawing
  • US10983881B2 patent drawing
  • US10983881B2 patent drawing

AI summary

Embodiments for disaster recovery in a disaggregated computing system. Memory resources are allocated at a secondary, disaster recovery site for data received from a primary site. The data from the primary site is continuously replicated to the allocated memory resources at the disaster recovery site without requiring any compute resources to be attached to the allocated memory resources. Responsive to determining a disaster recovery failover is in progress, the compute resources are assigned to the allocated memory resources for performing a failover workload, and the failover workload is executed at the disaster recovery site.