Virtual Machine Manager Hardware Fault Detection
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current virtual machine management systems are inefficient in managing large-scale data centers due to limitations in detecting hardware problems accurately and reliably, and they cannot customize backup and error recovery solutions for each virtual machine, leading to potential downtime and management complexities.
Innovation Solution
A system and method that integrates a virtual machine manager with a blade server management module, enabling direct reception of hardware problem information and sending processing commands to virtual machine hypervisors, along with resource management and predefined strategy handling, to address hardware issues promptly and accurately.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If virtual machine manager uses heartbeat detection to monitor virtual machine host status, then the system can detect abnormal conditions, but the detection accuracy and reliability are insufficient leading to delayed response
Solution Approach 1:
The patent introduces a blade server management module as an intermediary between the virtual machine manager and the virtual machine host. This module directly monitors hardware status through system management bus, providing reliable and timely detection of hardware problems without the delays inherent in network-based heartbeat methods.
Solution Approach 2:
The patent replaces the software-based network heartbeat detection mechanism with a hardware-level monitoring system using system management bus and blade server management module. This substitution enables direct, real-time hardware status monitoring with higher reliability and faster response time.
2Quantity of substance
If virtual machine manager manages large number of virtual machine hosts (2000+ servers), then the system can handle large-scale data centers, but the management complexity and resource requirements increase significantly
Solution Approach 1:
The patent segments the management function by introducing a blade server management module that handles hardware-level monitoring and basic operations. This allows the virtual machine manager to focus on higher-level virtual machine management, reducing its complexity while enabling support for larger numbers of hosts through hierarchical decomposition of management responsibilities.
Solution Approach 2:
The blade server management module serves multiple functions including hardware monitoring, event detection, and coordination with the virtual machine manager. This multi-functionality reduces the overall system complexity by consolidating management tasks in a single module rather than requiring the virtual machine manager to handle all aspects directly.
3Adaptability or versatility
If virtual machine manager cannot customize backup and error recovery solutions, then the system is simpler to operate, but the ability to handle specific hardware failures is insufficient
Solution Approach 1:
The patent implements preliminary action by pre-configuring backup and error recovery solutions for different hardware failure scenarios. When hardware problems are detected, the system can immediately execute predetermined recovery procedures, providing both customization capability and operational simplicity through automation of the recovery process.
Data Source
AI summary
A system and method are provided for virtual machine management. The system comprises a virtual machine manager, a blade server management module, at least one blade server, and a virtual machine manager. The virtual machine manager comprises an abnormal event receiving module for receiving information about a blade server having a hardware problem directly from the blade server management module and additionally a virtual machine management module for sending a processing command to a virtual machine hypervisor on the blade server having the hardware problem. The virtual machine management module receives the information about the hardware problem from the abnormal event receiving module. The processing command is determined in accordance with the information about the hardware problem and strategies for handling predefined hardware problems.


