Virtual Machine Recovery Using Priority-Based Standby Server Allocation

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Data centers face challenges in providing reliable and efficient recovery of virtual machines due to hardware failures, especially with varying importance scores of customers and workloads, making it impractical to have a corresponding number of standby servers readily available.

Innovation Solution

A method and system for managing servers that involve obtaining a configuration of server groups, including standby servers, and automatically allocating and remotely restarting virtual machines based on priority values of servers and virtual machines, prioritizing important workloads and users.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If standby servers are maintained for all possible failures, then reliability is improved, but device complexity and cost increase significantly

Engineering Contradiction:
Improvevirtual machine recovery reliabilityVSAvoidstandby server pool complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The system performs preliminary actions by pre-configuring virtual machine images and priorities before failures occur. When a server fails, the system can rapidly deploy new virtual machines using pre-configured images from a repository, eliminating the need for complex manual recovery procedures and reducing the requirement for extensive standby server pools.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The system creates copies of virtual machine images from a centralized repository and deploys them to available servers. Instead of maintaining exact standby servers for every possible failure scenario, the system uses image copying and restoration techniques to rapidly recreate failed virtual machines on any available hardware, significantly reducing the complexity of the standby infrastructure.

Inventive Principle:
Principle #26Copying

2Productivity

If priority-based allocation is implemented, then resource utilization efficiency is improved, but system complexity increases

Engineering Contradiction:
Improveresource utilization efficiencyVSAvoidpriority management system complexity
Core Design Contradiction:
ProductivityVSDevice complexity

Solution Approach 1:

The system changes parameters by assigning priority values to different virtual machines and servers. This simple parameter-based approach enables automated decision-making for resource allocation and failure recovery without requiring complex algorithms. The priority parameters guide the system in selecting which virtual machines to restore first and which servers to use, improving resource utilization through straightforward parameter-driven logic.

Inventive Principle:
Principle #35Parameter changes

3Productivity

If automatic recovery systems are implemented, then productivity is improved, but ease of operation decreases

Engineering Contradiction:
Improverecovery speedVSAvoidsystem management complexity
Core Design Contradiction:
ProductivityVSEase of operation

Solution Approach 1:

The system implements self-service capabilities by automatically detecting server failures, selecting appropriate virtual machine images from the repository, allocating resources based on priority parameters, and deploying recovered virtual machines without human intervention. This automation dramatically improves recovery speed while the standardized self-service process actually simplifies operations by eliminating manual recovery procedures.

Inventive Principle:
Principle #25Self-service

Data Source

PatentUS20260044361A1Automatic recovery of virtual machines
Publication Date: 2026.02.12 INTERNATIONAL BUSINESS MACHINE CORPORATION
  • US20260044361A1 patent drawing
  • US20260044361A1 patent drawing
  • US20260044361A1 patent drawing

AI summary

A method and system for managing servers by obtaining a configuration of a plurality of server groups, the configuration including information about virtual machines on servers of the server groups, and the plurality of server groups including at least one pool of standby servers that is at least operatively distinct from a rest of the plurality of server groups. The method further includes detecting a failure event of a failed server, automatically allocating, a standby server of the at least one pool of standby servers to a server group of the failed server based on a server priority value of the failed server and/or a virtual machine priority value of one or more virtual machines of the failed server, and remotely restarting the one or more virtual machines of the failed server on the standby server.