Proactive Resource Reservation for Virtual Machine Failover
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Hardware failures, particularly storage and network outages, significantly impact virtualized infrastructure, leading to application downtime and reduced consolidation ratios, which undermines the capital expenditure benefits of virtualization and customer confidence.
Innovation Solution
A proactive resource reservation system is implemented, comprising a cluster of hosts with a master and slave configuration, where the system detects storage access failures and automatically recovers impacted virtual machines by reserving resources on healthy hosts, ensuring continuous service.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Loss of energy
If hardware consolidation ratios are increased to reduce capital expenditure, then cost efficiency is improved, but the impact of hardware failures increases leading to more virtual machine outages
Solution Approach 1:
The system performs preliminary actions by proactively detecting hardware failures before they cause virtual machine outages. The failure detector continuously monitors hardware status and triggers resource reservations in advance, ensuring that backup resources are ready before failures impact service delivery.
Solution Approach 2:
The system implements beforehand cushioning by reserving backup computational resources on healthy hosts before failures occur. When hardware failures are detected, the system has pre-positioned resources ready to immediately接纳 and run failed virtual machines, cushioning the impact of failures and maintaining service availability.
2Reliability
If proactive resource reservation is implemented to protect against hardware failures, then virtual machine availability is improved, but system complexity increases
Solution Approach 1:
The system implements self-service by enabling virtual machines to automatically detect their own hardware failures and trigger resource reservations. The virtual machines monitor their execution environment, detect failures independently, and initiate failover procedures without requiring complex external management intervention.
Solution Approach 2:
The system merges the failure detection and resource reservation functions into an integrated automated workflow. The failure detector, resource manager, and virtual machine monitoring are combined into a unified system that automatically coordinates failure detection, resource allocation, and failover execution, reducing operational complexity.
3Loss of time
If automated failure detection and resource reservation is implemented, then downtime is reduced, but computational overhead increases
Solution Approach 1:
The system implements skipping by rapidly executing failover procedures once failures are detected. The automated resource reservation and virtual machine migration processes are optimized to complete quickly, rushing through the failover sequence to minimize the time window where services are degraded or unavailable.
Solution Approach 2:
The system uses feedback mechanisms where the failure detector continuously monitors hardware status and provides real-time information to the resource manager. This feedback loop enables the system to dynamically adjust resource allocations and trigger failover only when actually needed, avoiding unnecessary computational overhead during normal operation.
Data Source
AI summary
A system for proactive resource reservation for protecting virtual machines. The system includes a cluster of hosts, wherein the cluster of hosts includes a master host, a first slave host, and one or more other slave hosts, and wherein the first slave host executes one or more virtual machines thereon. The first slave host is configured to identify a failure that impacts an ability of the one or more virtual machines to provide service, and calculate a list of impacted virtual machines. The master host is configured to receive a request to reserve resources on another host in the cluster of hosts to enable the impacted one or more virtual machines to failover, calculate a resource capacity among the cluster of hosts, determine whether the calculated resource capacity is sufficient to reserve the resources, and send an indication as to whether the resources are reserved.


