Stretched Cluster Migration for Zero-Downtime VMs

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current active/active storage systems are limited to metropolitan-scale distances, making it impossible to implement seamless data migration between data centers located over 100 kilometers apart without reverting to active/passive architectures, which results in downtime during migration.

Innovation Solution

Implementing a distributed virtual volume system with a stretched cluster mechanism and relay stations to enable synchronous data migration across long distances, using protocols like VPLEX Metro and VMware vMotion, allowing for zero-downtime migration of active virtual machines by encapsulating their state and transferring it over high-speed networks.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If active/active storage systems are used for data migration, then data availability and zero-downtime migration are achieved, but the geographic distance is limited to metropolitan-scale (100 kilometers or less)

Engineering Contradiction:
Improvedata availabilityVSAvoidgeographic distance
Core Design Contradiction:
ReliabilityVSLength of stationary object

Solution Approach 1:

The system segments the long-distance data migration path into multiple metropolitan-scale segments, each handled by a relay station. Virtual machine state is transferred through intermediate relay stations that maintain active/active storage relationships, breaking down the impossible long-distance direct connection into manageable short-distance segments that can each support synchronous replication

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

Relay stations serve as intermediary systems between source and destination data centers. These relay stations maintain active storage relationships with both previous and next stations in the chain, acting as mediators that enable state transfer across geographic distances that would otherwise exceed the capabilities of active/active storage protocols

Inventive Principle:
Principle #24Intermediary (Mediator)

2Length of stationary object

If active/passive storage systems are used for long-distance data migration, then geographic distance limitations are overcome, but data downtime occurs during migration

Engineering Contradiction:
Improvegeographic distanceVSAvoiddata availability
Core Design Contradiction:
Length of stationary objectVSReliability

Solution Approach 1:

The system performs preliminary actions by establishing active storage relationships and pre-configuring relay stations before migration begins. Virtual machine state is prepared and staged at relay stations in advance, allowing the actual migration to occur without downtime when the virtual machine is actively moved through the relay chain

Inventive Principle:
Principle #10Preliminary action

3Manufacturing precision

If synchronous replication is used for active/active storage, then data consistency is maintained, but the maximum transmission distance is limited to metropolitan scales

Engineering Contradiction:
Improvedata consistencyVSAvoidtransmission distance
Core Design Contradiction:
Manufacturing precisionVSLength of stationary object

Solution Approach 1:

The long-distance transmission path is segmented into multiple short-distance synchronous replication links between adjacent relay stations. Each segment maintains data consistency through synchronous replication, while the chain of segments collectively spans much larger geographic distances than a single synchronous link could achieve

Inventive Principle:
Principle #1Segmentation

Data Source

PatentUS10083057B1Migration of active virtual machines across multiple data centers
Publication Date: 2018.09.25 EMC IP HLDG CO LLC
  • US10083057B1 patent drawing
  • US10083057B1 patent drawing
  • US10083057B1 patent drawing

AI summary

A method of providing migration of active virtual machines by performing data migrations between data centers using distributed volume and stretched cluster mechanisms to migrate the data synchronously within the distance and time latency limits defined by the distributed volume protocol, then performing data migrations within data centers using local live migration, and for long distance data migrations on a scale or distance that may exceed synchronous limits of the distributed volume protocol, combining appropriate inter- and intra-site data migrations so that data migrations can be performed exclusively using synchronous transmission.