Virtual Machine Migration for Storage Failure Recovery
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Hardware failures, particularly storage and network outages, significantly impact virtualized infrastructure, leading to application downtime and reduced consolidation ratios, which undermines the capital expenditure benefits of virtualization and customer confidence.
Innovation Solution
A system comprising a master and slave host, where the slave host monitors virtual machines for Permanent Device Loss (PDL) and All Paths Down (APD) failures, and implements specific remedies to terminate or restart virtual machines, ensuring service continuity by migrating them to healthy hosts.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Productivity
If consolidation ratios are increased to maximize virtualization benefits, then capital expenditure efficiency improves, but the impact of hardware failures increases and reliability deteriorates
Solution Approach 1:
The system performs preliminary actions by proactively detecting storage connectivity failures (PDL and APD conditions) before they cause complete virtual machine downtime. The host system identifies failure conditions and initiates remediation actions in advance, migrating virtual machines to healthy hosts before service is completely lost, thus maintaining reliability while allowing high consolidation ratios
Solution Approach 2:
The system implements continuous feedback monitoring of storage connectivity status for each virtual machine. When storage paths are monitored and failure conditions are detected, the system responds by triggering migration workflows that move affected virtual machines to alternative healthy hosts, thereby maintaining service availability despite hardware failures in highly consolidated environments
2Reliability
If hardware failure protection is enhanced to maintain service availability, then reliability improves, but system complexity and operational overhead increase
Solution Approach 1:
The system enables self-service automation where the virtualization infrastructure automatically detects storage connectivity failures and performs remediation by migrating virtual machines to healthy hosts without manual intervention. This automation reduces operational complexity while maintaining high service availability, as the system handles failure response autonomously
Solution Approach 2:
The system prepares remediation actions in advance by continuously monitoring storage connectivity and having migration workflows ready to execute. When failures are detected, pre-configured remediation actions are immediately triggered, reducing the complexity of real-time failure response while ensuring service continuity
Data Source
AI summary
A system for monitoring virtual machines includes a master host and a slave host. The slave host includes a primary virtual machine and a secondary virtual machine. The slave host is configured to identify a failure that impacts an ability of at least one of the primary virtual machine and the secondary virtual machine to provide service. If the failure is a Permanent Device Loss failure, the slave host is configured to terminate each impacted virtual machine. If the failure is an All Paths Down failure, the master host is configured to apply one of the following: a first remedy if the primary virtual machine is impacted and the secondary virtual machine is not impacted; a second remedy if the secondary virtual machine is impacted and the primary virtual machine is not impacted; or a third remedy if both the primary virtual machine and the secondary virtual machine are impacted.


