Virtual Machine Migration for Storage Failure Recovery

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Hardware failures, particularly storage and network outages, significantly impact virtualized infrastructure, leading to application downtime and reduced consolidation ratios, which undermines the capital expenditure benefits of virtualization and customer confidence.

Innovation Solution

A system comprising a master and slave host, where the slave host monitors virtual machines for Permanent Device Loss (PDL) and All Paths Down (APD) failures, and implements specific remedies to terminate or restart virtual machines, ensuring service continuity by migrating them to healthy hosts.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Productivity

If consolidation ratios are increased to maximize virtualization benefits, then capital expenditure efficiency improves, but the impact of hardware failures increases and reliability deteriorates

Engineering Contradiction:
Improvecapital expenditure efficiencyVSAvoidsystem reliability
Core Design Contradiction:
ProductivityVSReliability

Solution Approach 1:

The system performs preliminary actions by proactively detecting storage connectivity failures (PDL and APD conditions) before they cause complete virtual machine downtime. The host system identifies failure conditions and initiates remediation actions in advance, migrating virtual machines to healthy hosts before service is completely lost, thus maintaining reliability while allowing high consolidation ratios

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The system implements continuous feedback monitoring of storage connectivity status for each virtual machine. When storage paths are monitored and failure conditions are detected, the system responds by triggering migration workflows that move affected virtual machines to alternative healthy hosts, thereby maintaining service availability despite hardware failures in highly consolidated environments

Inventive Principle:
Principle #23Feedback

2Reliability

If hardware failure protection is enhanced to maintain service availability, then reliability improves, but system complexity and operational overhead increase

Engineering Contradiction:
Improveservice availabilityVSAvoidsystem complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The system enables self-service automation where the virtualization infrastructure automatically detects storage connectivity failures and performs remediation by migrating virtual machines to healthy hosts without manual intervention. This automation reduces operational complexity while maintaining high service availability, as the system handles failure response autonomously

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The system prepares remediation actions in advance by continuously monitoring storage connectivity and having migration workflows ready to execute. When failures are detected, pre-configured remediation actions are immediately triggered, reducing the complexity of real-time failure response while ensuring service continuity

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS9268642B2Protecting paired virtual machines
Publication Date: 2016.02.23 VMWARE INC
  • US9268642B2 patent drawing
  • US9268642B2 patent drawing
  • US9268642B2 patent drawing

AI summary

A system for monitoring virtual machines includes a master host and a slave host. The slave host includes a primary virtual machine and a secondary virtual machine. The slave host is configured to identify a failure that impacts an ability of at least one of the primary virtual machine and the secondary virtual machine to provide service. If the failure is a Permanent Device Loss failure, the slave host is configured to terminate each impacted virtual machine. If the failure is an All Paths Down failure, the master host is configured to apply one of the following: a first remedy if the primary virtual machine is impacted and the secondary virtual machine is not impacted; a second remedy if the secondary virtual machine is impacted and the primary virtual machine is not impacted; or a third remedy if both the primary virtual machine and the secondary virtual machine are impacted.