Virtual Machine Recovery Using Priority-Based Standby Server Allocation
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Data centers face challenges in providing reliable and efficient recovery of virtual machines due to hardware failures, especially with varying importance scores of customers and workloads, making it impractical to have a corresponding number of standby servers readily available.
Innovation Solution
A method and system for managing servers that involve obtaining a configuration of server groups, including standby servers, and automatically allocating and remotely restarting virtual machines based on priority values of servers and virtual machines, prioritizing important workloads and users.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If standby servers are maintained for all possible failures, then reliability is improved, but device complexity and cost increase significantly
Solution Approach 1:
The system performs preliminary actions by pre-configuring virtual machine images and priorities before failures occur. When a server fails, the system can rapidly deploy new virtual machines using pre-configured images from a repository, eliminating the need for complex manual recovery procedures and reducing the requirement for extensive standby server pools.
Solution Approach 2:
The system creates copies of virtual machine images from a centralized repository and deploys them to available servers. Instead of maintaining exact standby servers for every possible failure scenario, the system uses image copying and restoration techniques to rapidly recreate failed virtual machines on any available hardware, significantly reducing the complexity of the standby infrastructure.
2Productivity
If priority-based allocation is implemented, then resource utilization efficiency is improved, but system complexity increases
Solution Approach 1:
The system changes parameters by assigning priority values to different virtual machines and servers. This simple parameter-based approach enables automated decision-making for resource allocation and failure recovery without requiring complex algorithms. The priority parameters guide the system in selecting which virtual machines to restore first and which servers to use, improving resource utilization through straightforward parameter-driven logic.
3Productivity
If automatic recovery systems are implemented, then productivity is improved, but ease of operation decreases
Solution Approach 1:
The system implements self-service capabilities by automatically detecting server failures, selecting appropriate virtual machine images from the repository, allocating resources based on priority parameters, and deploying recovered virtual machines without human intervention. This automation dramatically improves recovery speed while the standardized self-service process actually simplifies operations by eliminating manual recovery procedures.
Data Source
AI summary
A method and system for managing servers by obtaining a configuration of a plurality of server groups, the configuration including information about virtual machines on servers of the server groups, and the plurality of server groups including at least one pool of standby servers that is at least operatively distinct from a rest of the plurality of server groups. The method further includes detecting a failure event of a failed server, automatically allocating, a standby server of the at least one pool of standby servers to a server group of the failed server based on a server priority value of the failed server and/or a virtual machine priority value of one or more virtual machines of the failed server, and remotely restarting the one or more virtual machines of the failed server on the standby server.


