GPU Fabric Port Throttling for Isolated VM Bandwidth
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
The challenge of managing multi-tenancy in datacenter graphics processors is exacerbated by the unsustainable scaling of silicon architectures, leading to inefficiencies in measuring and throttling per-VM performance, which impacts the total cost of ownership (TCO) and overall silicon footprint.
Innovation Solution
Implementing configurable fabric bandwidth throttling for GPU virtualized workloads, allowing for dynamic allocation and isolation of compute resources among virtual machines (VMs) to ensure quality of service (QoS) and optimize resource utilization.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If duplication method of scaling silicon architectures is applied, then multi-tenancy capability is improved, but per-VM performance measurement and throttling becomes unsustainable
Solution Approach 1:
The GPU fabric is segmented into multiple virtual channels, each dedicated to a specific virtual machine. This segmentation allows independent bandwidth management and performance measurement for each VM without requiring duplication of the entire silicon architecture, thereby maintaining multi-tenancy capability while making per-VM performance throttling sustainable.
Solution Approach 2:
The patent changes the bandwidth allocation parameter from a static, architecture-level setting to a dynamic, per-VM virtual channel configuration. By adjusting the number of virtual channels and their associated bandwidth thresholds, the system can efficiently manage multi-tenancy without the computational overhead of full architecture duplication.
2Adaptability or versatility
If silicon architecture duplication is used to scale, then multi-tenancy is improved, but total cost of ownership increases
Solution Approach 1:
Instead of duplicating the entire silicon architecture for each VM, the system segments the fabric bandwidth into virtual channels. This allows multiple VMs to share the same physical silicon while maintaining isolation and performance guarantees, thereby reducing the silicon footprint and associated TCO.
Solution Approach 2:
The physical GPU fabric is designed to serve multiple VMs simultaneously through virtual channelization. The same physical infrastructure performs the function of multiple virtual systems, eliminating the need for dedicated silicon duplication for each tenant and reducing overall silicon footprint.
3Reliability
If fabric bandwidth throttling is implemented, then resource management and QoS are improved, but device complexity increases
Solution Approach 1:
The fabric is divided into virtual channels that are automatically associated with specific VMs. This segmentation enables QoS guarantees through simple bandwidth threshold enforcement at each virtual channel level, avoiding the need for complex centralized bandwidth management while maintaining reliable resource allocation.
Solution Approach 2:
The virtual channel mechanism enables self-service bandwidth management where each VM's bandwidth is automatically enforced through the virtual channel infrastructure. The system self-regulates resource allocation based on configured thresholds without requiring complex external control, thereby achieving QoS with manageable complexity.
Data Source
AI summary
One embodiment provides a graphics processor comprising a memory interface, a graphics core cluster including a plurality of graphics cores, and an interconnect fabric to interconnect a plurality of hardware clients including the plurality of graphics cores. The interconnect fabric include a plurality of fabric ports coupled with the plurality of graphics cores. A fabric port is configured to limit bandwidth available to an associated graphics core via a bandwidth throttler circuit coupled with the fabric port.


