Accelerator Pooling Over Fabric for Compute Resource Allocation
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Hardware accelerators in compute devices often experience inefficient resource allocation due to varying levels of usage, leading to underutilization as they may be idle for extended periods.
Innovation Solution
A system for pooling accelerators over a fabric, allowing compute devices to transparently access either local or remote accelerator devices based on usage and availability, managed by an accelerator manager that selects the most appropriate device through a network connection, enabling efficient resource allocation and utilization.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Speed
If hardware accelerators are incorporated into compute devices, then computing tasks can be performed more quickly, but the accelerators may be unused much of the time leading to inefficient resource allocation
Solution Approach 1:
The patent merges multiple accelerator devices into a pooled resource that can be dynamically allocated to different compute devices. The accelerator manager consolidates control over multiple accelerators and distributes them based on demand, transforming isolated accelerator-unit systems into an integrated resource-sharing system that improves both speed and resource efficiency.
Solution Approach 2:
The accelerator manager enables accelerators to serve multiple functions by dynamically assigning them to different compute devices based on their availability and usage patterns. This universal access model allows a single accelerator to support multiple compute workloads over time, eliminating the inefficiency of dedicated single-purpose accelerator assignments.
2Ease of operation
If hardware accelerators are dedicated to specific compute devices, then access is straightforward and fast, but utilization is inefficient when the accelerator is idle
Solution Approach 1:
The accelerator manager implements feedback mechanisms to monitor accelerator availability and compute device requests in real-time. Based on this feedback, the system dynamically adjusts accelerator assignments, moving accelerators between compute devices to match actual demand patterns and maximize utilization while maintaining simple access for applications.
Solution Approach 2:
The system transitions from static dedicated accelerator assignments to dynamic resource allocation. The accelerator manager continuously adapts accelerator assignments based on real-time usage patterns, availability status, and compute device needs, enabling the system to respond flexibly to changing workloads while maintaining operational simplicity.
3Productivity
If multiple accelerators are pooled and dynamically allocated, then resource efficiency improves, but system complexity increases due to the need for an accelerator manager
Solution Approach 1:
The accelerator manager serves as an intermediary layer between compute devices and accelerator hardware. This mediator handles the complexity of resource pooling, availability tracking, and dynamic assignment, shielding individual compute devices from the underlying system complexity while enabling efficient resource utilization across the distributed system.
Data Source
AI summary
Technologies for pooling accelerators over fabric are disclosed. In the illustrative embodiment, an application may access an accelerator device over an application programming interface (API) and the API can access an accelerator device that is either local or a remote accelerator device that is located on a remote accelerator sled over a network fabric. The API may employ a send queue and a receive queue to send and receive command capsules to and from the accelerator sled.


