Admission Control for Accelerator Failover Capacity

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current failover techniques in data centers do not account for accelerator devices, leading to potential outages during host failures, as they do not reserve or manage these devices for failover scenarios, resulting in VMs being unable to restart on other hosts without available accelerators.

Innovation Solution

Implementing high-availability admission control policies that reserve accelerator devices as failover capacity, ensuring they are available for restarting computing entities in case of host failures, by determining the necessary number of devices to reserve based on tolerance policies and enforcing these policies within the network.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If failover techniques are implemented without reserving accelerator devices, then device complexity is reduced, but reliability deteriorates as VMs cannot restart on other hosts without available accelerators

Engineering Contradiction:
Improveavailability of accelerator devices during failoverVSAvoidcomplexity of failover management
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The patent reserves accelerator devices in advance for failover scenarios before actual host failures occur. The system pre-allocates accelerator devices to standby hosts, ensuring that when a host failure occurs, there are already available accelerators ready for immediate assignment to failing VMs, eliminating the need for complex real-time allocation during failover events.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent introduces an intermediary admission control mechanism that manages the allocation and reservation of accelerator devices between hosts and VMs. This intermediary system tracks accelerator availability, enforces reservation policies, and coordinates failover operations, simplifying the overall failover management while ensuring reliability.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Reliability

If accelerator devices are reserved as failover capacity, then reliability is improved, but productivity deteriorates due to reduced available capacity for new computing entities

Engineering Contradiction:
Improveavailability of accelerator devices for failoverVSAvoidthroughput of computing entities
Core Design Contradiction:
ReliabilityVSProductivity

Solution Approach 1:

The patent implements partial reservation of accelerator devices for failover capacity rather than reserving all accelerators. The system reserves only the minimum necessary number of accelerator devices to handle expected failover scenarios, allowing the majority of accelerator capacity to remain available for productive workloads. This partial action approach ensures adequate failover protection while minimizing impact on system throughput.

Inventive Principle:
Principle #16Partial or excessive action

Solution Approach 2:

The patent employs dynamic admission control that adjusts accelerator reservations based on current system conditions, workload demands, and failure rates. The reservation policy is not static but adapts to changing conditions, increasing reservations when failure risk is high and decreasing them when system conditions are stable, thereby optimizing both reliability and productivity dynamically.

Inventive Principle:
Principle #15Dynamics

3Reliability

If admission control policies are enforced to reserve accelerator devices, then reliability is improved, but device complexity increases due to policy management overhead

Engineering Contradiction:
Improveavailability of accelerator devicesVSAvoidcomplexity of policy enforcement
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The patent implements self-service mechanisms where the admission control system automatically monitors accelerator device availability, enforces reservation policies, and manages failover operations without requiring manual intervention. The system self-adjusts reservations based on observed failure patterns and current resource utilization, reducing the operational complexity of policy management while maintaining high reliability.

Inventive Principle:
Principle #25Self-service

Data Source

PatentUS11748142B2High-availability admission control for accelerator devices
Publication Date: 2023.09.05 VMWARE INC
  • US11748142B2 patent drawing
  • US11748142B2 patent drawing
  • US11748142B2 patent drawing

AI summary

The disclosure provides an approach for high-availability admission control. Embodiments include determining a number of slots present on the cluster of hosts. Embodiments include receiving an indication of a number of host failures to tolerate. Embodiments include determining a number of slots that are assigned to existing computing instances on the cluster of hosts. Embodiments include determining an available cluster capacity based on the number of slots present on the cluster of hosts, the number of host failures to tolerate, and the number of slots that are assigned to existing computing instances on the cluster of hosts. Embodiments include determining whether to admit a given computing instance to the cluster of hosts based on the available cluster capacity.