VM Placement Optimization for Cluster Failure Resilience

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

In high-availability virtualized computing clusters, existing techniques face challenges in optimizing virtual machine placements to ensure resilience and performance, particularly when node failures occur, due to the complexity of permutation scenarios and the need for frequent reconfigurations, which can lead to suboptimal configurations and increased resource demands.

Innovation Solution

The implementation of advanced techniques that proactively monitor and manage computing cluster configurations, performing node-to-node VM migrations and optimizations before failure events to achieve optimal high-availability placements, utilizing data organization, communication paths, and module interrelationships to enhance flexibility and reduce resource demands.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If VM placements are optimized in advance using advanced techniques, then high-availability compliance and resilience are improved, but device complexity and processing overhead increase

Engineering Contradiction:
Improvehigh-availability complianceVSAvoidplacement optimization complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The system performs VM placement optimization in advance of failure events by proactively analyzing current VM placements, identifying feasible placement scenarios, and selecting optimal placements before failures occur. This preliminary action ensures that when a node fails, the pre-computed optimal placement can be executed immediately without complex real-time calculations, thus improving reliability while managing complexity.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The placement optimization process is divided into distinct phases: analyzing current VM placements, identifying feasible placement scenarios, scoring scenarios based on optimization criteria, and selecting the optimal placement. This segmentation allows each phase to be handled independently with appropriate algorithms, reducing overall system complexity while achieving comprehensive optimization.

Inventive Principle:
Principle #1Segmentation

2Reliability

If VM migrations are performed frequently to maintain optimal placements, then resilience and performance are improved, but loss of time and productivity decrease

Engineering Contradiction:
ImproveresilienceVSAvoidmigration time
Core Design Contradiction:
ReliabilityVSLoss of time

Solution Approach 1:

The system identifies and prepares optimal VM placement scenarios in advance of failures by analyzing current placements and computing optimal target configurations. This preliminary computation ensures that when a failure occurs, the system can execute pre-determined migrations rather than performing complex real-time optimization, significantly reducing migration time while maintaining resilience.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The optimization process focuses on local adjustments to VM placements rather than comprehensive reconfigurations. By identifying specific VMs that need migration and their optimal target nodes, the system minimizes unnecessary migrations and reduces overall migration time while still achieving optimal resilience.

Inventive Principle:
Principle #3Local quality

3Manufacturing precision

If comprehensive placement scenario analysis is performed, then manufacturing precision and measurement precision are improved, but device complexity and resource demands increase

Engineering Contradiction:
Improveplacement optimization precisionVSAvoidanalysis complexity
Core Design Contradiction:
Manufacturing precisionVSDevice complexity

Solution Approach 1:

The system employs scoring functions that evaluate placement scenarios based on multiple parameters including resource utilization, load balancing, and failure recovery capabilities. By transforming the complex multi-dimensional optimization problem into a scored ranking system, the patent achieves high placement precision while managing analysis complexity through parameter-based evaluation rather than exhaustive scenario analysis.

Inventive Principle:
Principle #35Parameter changes

Solution Approach 2:

The system analyzes a sufficient subset of placement scenarios rather than all possible permutations. By identifying feasible placements and scoring the most promising candidates, the system achieves high optimization precision without the computational complexity of exhaustive analysis, performing only the necessary level of analysis required for optimal decisions.

Inventive Principle:
Principle #16Partial or excessive action

4Adaptability or versatility

If cluster reconfiguration activities are performed after failure events, then adaptability is improved, but productivity and time efficiency worsen due to degraded cluster performance

Engineering Contradiction:
Improvefailure recovery capabilityVSAvoidreconfiguration efficiency
Core Design Contradiction:
Adaptability or versatilityVSProductivity

Solution Approach 1:

The system computes optimal VM placement scenarios and prepares migration plans before failures occur. When a node fails, the pre-computed placement can be executed immediately without requiring complex real-time reconfiguration decisions, thus maintaining high productivity during recovery while preserving full adaptability to handle various failure scenarios.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The system maintains pre-computed optimal placement scenarios that act as a buffer or cushion against failure events. These pre-prepared scenarios ensure that when failures occur, the cluster can rapidly reconfigure using predetermined optimal configurations rather than performing time-consuming analysis during the critical recovery window, thereby maintaining productivity while ensuring adaptability.

Inventive Principle:
Principle #11Beforehand cushioning (Prior cushioning)

Data Source

PatentUS12093153B2Optimizing high-availability virtual machine placements in advance of a computing cluster failure event
Publication Date: 2024.09.17 NUTANIX INC
  • US12093153B2 patent drawing
  • US12093153B2 patent drawing
  • US12093153B2 patent drawing

AI summary

Placement scenario optimization mechanisms for automatic placement of computing entities onto nodes of a running multi-node computing cluster. A set of failure mode parameters define a high-availability requirement of the multi-node computing cluster. In advance of a failure event, and responsive to a determination that a then-current computing entity placement does not satisfy the high-availability requirement, the cluster is analyzed and a plurality of feasible placement scenarios are generated. Optimization criteria are applied to the feasible placement scenarios such that a best choice from among the feasible placement scenarios is identified and applied to the virtual machine placements over the cluster. A change monitoring and detection facility continually observes the multi-node computing cluster to detect a change of a failure mode parameter or to detect a change to the configuration of the virtual machines. Certain of such changes cause feasible placement scenarios to be generated, evaluated, selected, and applied.