Admission Control for Accelerator Failover Capacity
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current failover techniques in data centers do not account for accelerator devices, leading to potential outages during host failures, as they do not reserve or manage these devices for failover scenarios, resulting in VMs being unable to restart on other hosts without available accelerators.
Innovation Solution
Implementing high-availability admission control policies that reserve accelerator devices as failover capacity, ensuring they are available for restarting computing entities in case of host failures, by determining the necessary number of devices to reserve based on tolerance policies and enforcing these policies within the network.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If failover techniques are implemented without reserving accelerator devices, then device complexity is reduced, but reliability deteriorates as VMs cannot restart on other hosts without available accelerators
Solution Approach 1:
The patent reserves accelerator devices in advance for failover scenarios before actual host failures occur. The system pre-allocates accelerator devices to standby hosts, ensuring that when a host failure occurs, there are already available accelerators ready for immediate assignment to failing VMs, eliminating the need for complex real-time allocation during failover events.
Solution Approach 2:
The patent introduces an intermediary admission control mechanism that manages the allocation and reservation of accelerator devices between hosts and VMs. This intermediary system tracks accelerator availability, enforces reservation policies, and coordinates failover operations, simplifying the overall failover management while ensuring reliability.
2Reliability
If accelerator devices are reserved as failover capacity, then reliability is improved, but productivity deteriorates due to reduced available capacity for new computing entities
Solution Approach 1:
The patent implements partial reservation of accelerator devices for failover capacity rather than reserving all accelerators. The system reserves only the minimum necessary number of accelerator devices to handle expected failover scenarios, allowing the majority of accelerator capacity to remain available for productive workloads. This partial action approach ensures adequate failover protection while minimizing impact on system throughput.
Solution Approach 2:
The patent employs dynamic admission control that adjusts accelerator reservations based on current system conditions, workload demands, and failure rates. The reservation policy is not static but adapts to changing conditions, increasing reservations when failure risk is high and decreasing them when system conditions are stable, thereby optimizing both reliability and productivity dynamically.
3Reliability
If admission control policies are enforced to reserve accelerator devices, then reliability is improved, but device complexity increases due to policy management overhead
Solution Approach 1:
The patent implements self-service mechanisms where the admission control system automatically monitors accelerator device availability, enforces reservation policies, and manages failover operations without requiring manual intervention. The system self-adjusts reservations based on observed failure patterns and current resource utilization, reducing the operational complexity of policy management while maintaining high reliability.
Data Source
AI summary
The disclosure provides an approach for high-availability admission control. Embodiments include determining a number of slots present on the cluster of hosts. Embodiments include receiving an indication of a number of host failures to tolerate. Embodiments include determining a number of slots that are assigned to existing computing instances on the cluster of hosts. Embodiments include determining an available cluster capacity based on the number of slots present on the cluster of hosts, the number of host failures to tolerate, and the number of slots that are assigned to existing computing instances on the cluster of hosts. Embodiments include determining whether to admit a given computing instance to the cluster of hosts based on the available cluster capacity.


