Distributed Hypervisor State Management for Virtual Machine Fault Tolerance
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
In networking systems, consolidating server services into a single physical machine improves efficiency but reduces fault tolerance, as a failure in one service can affect others, and using multiple operating systems enhances protection but may not optimize resource usage effectively.
Innovation Solution
A method and system that distribute hypervisor functions across multiple network devices to control virtual machines, store and retrieve state information for fault tolerance and resilience, and allocate resources based on network device parameters, enabling load balancing and proactive migration of virtual machines.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Productivity
If server services are consolidated into a single physical machine, then server efficiency is improved, but fault tolerance deteriorates
Solution Approach 1:
The patent segments the hypervisor control functions across multiple network devices rather than consolidating them in a single device. Each network device maintains a portion of the hypervisor functions and state information, creating distributed virtual machine management. This segmentation allows the system to maintain high resource utilization (efficiency) while ensuring that failure of one device does not affect the entire system (fault tolerance).
Solution Approach 2:
The patent implements local quality by storing state information locally at each network device while maintaining distributed coordination. Each network device maintains its own local state information for the virtual machines it manages, allowing independent operation and recovery. This local quality enables each device to continue serving its virtual machines even when other devices fail, thus improving fault tolerance while maintaining overall system efficiency.
2Reliability
If multiple operating systems are used to maintain protection, then fault tolerance is improved, but resource utilization efficiency deteriorates
Solution Approach 1:
The patent implements universality by creating a distributed hypervisor system where multiple network devices collectively provide virtual machine management services. Instead of requiring separate dedicated operating systems for each service, the distributed hypervisor allows any network device to potentially manage any virtual machine, providing multi-functionality across the system. This maintains fault tolerance through distribution while improving resource utilization through flexible allocation.
Solution Approach 2:
The patent transitions from a single-dimension approach (single hypervisor on one device) to a multi-dimensional distributed architecture where hypervisor functions are spread across multiple network devices. This dimensional change allows the system to simultaneously achieve the protection benefits of multiple systems while maintaining the resource efficiency of shared virtualization, as resources can be dynamically allocated across the distributed network devices.
Data Source
AI summary
Hypervisor functions, which may control operations of one or more virtual machines, may be distributed across a plurality of network devices. State information may be stored for the virtual machines on network devices for fault tolerance and resilience. The virtual machines may retrieve stored state information to recover from a fault. The hypervisor may control the storage of the state information. Resources of the network devices may be allocated for fault tolerance and resilience of the virtual machines based on network device parameters, which may include storage capacity, processor usage, access time, communication bandwidth, and/or latency. The state information may include program content, cache content, memory content, and/or register content information, may be stored on a continuous, periodic, or an aperiodic basis, and may be shared among the network devices to enable the processing of data by the virtual machines when a fault occurs, and may be incrementally updated.


