Cluster Virtual Machine High Availability Recovery Strategy
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
OpenStack lacks a comprehensive high availability solution for virtual machines, leading to unavailability of applications and disruptions in cloud services when virtual machines fail, as they often require human intervention for recovery.
Innovation Solution
A method and apparatus for implementing high availability in cluster virtual machines, which involves detecting abnormal states, acquiring specific recovery strategies (such as restart, rebuild, or ignore strategies), and performing recovery operations, including monitoring host and physical volume states, and rebuilding virtual machines on alternative hosts if necessary.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If OpenStack is used for virtual machine management, then cloud computing capability is improved, but high availability of virtual machines deteriorates due to lack of complete recovery solutions
Solution Approach 1:
The system enables virtual machines to automatically detect their own abnormal states and execute recovery operations without human intervention. The high availability mechanism allows the virtual machine to monitor its runtime state, detect failures, and autonomously restart or rebuild itself, transforming the traditional manual recovery model into an automated self-service model that significantly improves reliability while maintaining cloud computing versatility
Solution Approach 2:
The system pre-configures multiple recovery strategies (restart strategy, rebuild strategy, ignore strategy) for different virtual machine failure scenarios before failures occur. When an abnormal state is detected, the system selects and executes the pre-planned recovery strategy, eliminating the need for real-time human decision-making and ensuring rapid recovery, thus improving high availability while preserving cloud computing capabilities
2Ease of operation
If manual intervention is required for virtual machine recovery, then operational control is maintained, but service availability deteriorates due to recovery time loss
Solution Approach 1:
The virtual machine monitors its own runtime state and automatically executes recovery operations when abnormalities are detected. The system retrieves recovery strategies from storage, selects appropriate actions (restart, rebuild, or ignore), and performs recovery without requiring operator intervention, thereby maintaining service availability while simplifying operational control to automated self-management
Solution Approach 2:
The system continuously monitors the runtime state of virtual machines and provides feedback when abnormal conditions occur. Upon detecting failures, the system automatically triggers the appropriate recovery strategy and monitors the recovery process, creating a closed-loop feedback mechanism that ensures rapid restoration of service availability without manual intervention while maintaining operational transparency
3Reliability
If recovery strategies are customized for each virtual machine, then recovery accuracy is improved, but system complexity increases
Solution Approach 1:
The system divides recovery strategies into distinct, independent types (restart strategy, rebuild strategy, ignore strategy) that can be selectively applied to different virtual machines based on their specific needs and failure modes. Each strategy is a separate module that handles specific recovery scenarios, allowing customized recovery accuracy for each virtual machine while keeping the overall system structure modular and manageable rather than monolithic and complex
Data Source
Figure 1
Figure 2~3
Figure 4
AI summary
Provided is a method and an apparatus (800) of implementing a high availability of a cluster virtual machine. The method includes: detecting (S 110) a state of at least one virtual machine in a cluster; acquiring (S 120), in response to the at least one virtual machine being in an abnormal state, a high availability recovery strategy specific to the at least one virtual machine; and performing (S130) a recovery operation on the at least one virtual machine according to the acquired high availability recovery strategy specific to the at least one virtual machine.