Virtual Machine Remediation System for Cluster State Changes
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
In complex virtual machine systems, ensuring performance guarantees after user-initiated or automatic changes, such as adding or removing VMs and hosts, is challenging due to the need for continuous resource management and the inability to react to changes like host failures without violating resource reservations.
Innovation Solution
A method that adjusts host computer configurations by identifying reasons for failed state changes, associating these reasons with remediation actions, assigning costs to these actions, and determining a set of actions to perform based on their costs, allowing the state change to pass both present and future checks, thereby ensuring resource availability and performance guarantees.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If resource reservations are strictly enforced for each VM, then performance guarantees are maintained, but the system's ability to react to changes (such as host failures) is hindered
Solution Approach 1:
The patent implements dynamic resource management by introducing a remediation system that can adjust resource allocations in response to changing conditions. Instead of static reservations, the system continuously evaluates the need for resource guarantees and adapts allocations based on current system state, allowing both performance guarantees and flexibility to coexist through time-varying resource management
Solution Approach 2:
The system changes resource allocation parameters dynamically by modifying reservation levels based on detected conditions. When changes occur in the virtual machine environment, the remediation system adjusts resource parameters (CPU, memory, storage allocations) to maintain performance guarantees while adapting to new configurations, effectively using parameter transformation to resolve the contradiction
2Reliability
If fixed resources are reserved for each VM, then performance predictability is ensured, but resource utilization efficiency decreases
Solution Approach 1:
The patent applies partial resource reservation rather than full fixed allocation. Instead of reserving all potentially needed resources for each VM, the system reserves only the minimum necessary to guarantee performance, allowing excess resources to be dynamically allocated to other VMs when not needed, thus improving overall utilization while maintaining predictability through controlled partial reservation
Solution Approach 2:
The remediation system creates a universal resource pool that serves multiple VMs simultaneously. Rather than dedicating resources exclusively to individual VMs, the system manages a shared pool that can be allocated to any VM needing resources, enabling both performance guarantees through minimum reservations and high utilization through flexible shared access
3Reliability
If continuous resource management is implemented, then performance guarantees are maintained, but system complexity increases
Solution Approach 1:
The patent implements self-service mechanisms where the remediation system automatically detects performance issues and executes corrective actions without human intervention. The system monitors resource allocations, identifies when performance guarantees are at risk, and autonomously applies remediation strategies, reducing the perceived complexity for users while maintaining continuous performance management
Solution Approach 2:
The system establishes continuous feedback loops that monitor resource usage and performance metrics. This feedback mechanism enables the remediation system to make informed decisions about resource allocations, automatically adjusting allocations based on real-time conditions while maintaining performance guarantees, thereby managing complexity through structured information flow
Data Source
AI summary
A method for adjusting the configuration of host computers in a cluster on which virtual machines are running in response to a failed change in state is disclosed. The method involves receiving at least one reason a change in state failed the present check or the future check, associating the at least one reason with at least one remediation action, wherein the remediation action would allow the change in state to pass both a present check and a future check, assigning the at least one remediation action a cost, and determining a set of remediation actions to perform based on the cost assigned to each remediation action. In an embodiment, the steps of this method may be implemented in a non-transitory computer-readable storage medium having instructions that, when executed in a computing device, causes the computing device to carry out the steps.


