GPU Fabric Port Throttling for Isolated VM Bandwidth

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

The challenge of managing multi-tenancy in datacenter graphics processors is exacerbated by the unsustainable scaling of silicon architectures, leading to inefficiencies in measuring and throttling per-VM performance, which impacts the total cost of ownership (TCO) and overall silicon footprint.

Innovation Solution

Implementing configurable fabric bandwidth throttling for GPU virtualized workloads, allowing for dynamic allocation and isolation of compute resources among virtual machines (VMs) to ensure quality of service (QoS) and optimize resource utilization.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If duplication method of scaling silicon architectures is applied, then multi-tenancy capability is improved, but per-VM performance measurement and throttling becomes unsustainable

Engineering Contradiction:
Improvemulti-tenancy capabilityVSAvoidper-VM performance measurement and throttling efficiency
Core Design Contradiction:
Adaptability or versatilityVSProductivity

Solution Approach 1:

The GPU fabric is segmented into multiple virtual channels, each dedicated to a specific virtual machine. This segmentation allows independent bandwidth management and performance measurement for each VM without requiring duplication of the entire silicon architecture, thereby maintaining multi-tenancy capability while making per-VM performance throttling sustainable.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent changes the bandwidth allocation parameter from a static, architecture-level setting to a dynamic, per-VM virtual channel configuration. By adjusting the number of virtual channels and their associated bandwidth thresholds, the system can efficiently manage multi-tenancy without the computational overhead of full architecture duplication.

Inventive Principle:
Principle #35Parameter changes

2Adaptability or versatility

If silicon architecture duplication is used to scale, then multi-tenancy is improved, but total cost of ownership increases

Engineering Contradiction:
Improvemulti-tenancy capabilityVSAvoidsilicon footprint
Core Design Contradiction:
Adaptability or versatilityVSQuantity of substance

Solution Approach 1:

Instead of duplicating the entire silicon architecture for each VM, the system segments the fabric bandwidth into virtual channels. This allows multiple VMs to share the same physical silicon while maintaining isolation and performance guarantees, thereby reducing the silicon footprint and associated TCO.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The physical GPU fabric is designed to serve multiple VMs simultaneously through virtual channelization. The same physical infrastructure performs the function of multiple virtual systems, eliminating the need for dedicated silicon duplication for each tenant and reducing overall silicon footprint.

Inventive Principle:
Principle #6Universality (Multi-functionality)

3Reliability

If fabric bandwidth throttling is implemented, then resource management and QoS are improved, but device complexity increases

Engineering Contradiction:
Improvequality of service guaranteeVSAvoidfabric bandwidth management complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The fabric is divided into virtual channels that are automatically associated with specific VMs. This segmentation enables QoS guarantees through simple bandwidth threshold enforcement at each virtual channel level, avoiding the need for complex centralized bandwidth management while maintaining reliable resource allocation.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The virtual channel mechanism enables self-service bandwidth management where each VM's bandwidth is automatically enforced through the virtual channel infrastructure. The system self-regulates resource allocation based on configured thresholds without requiring complex external control, thereby achieving QoS with manageable complexity.

Inventive Principle:
Principle #25Self-service

Data Source

PatentUS20250292351A1Configurable fabric bandwidth throttling for GPU virtualized workloads
Publication Date: 2025.09.18 INTEL CORP
  • US20250292351A1 patent drawing
  • US20250292351A1 patent drawing
  • US20250292351A1 patent drawing

AI summary

One embodiment provides a graphics processor comprising a memory interface, a graphics core cluster including a plurality of graphics cores, and an interconnect fabric to interconnect a plurality of hardware clients including the plurality of graphics cores. The interconnect fabric include a plurality of fabric ports coupled with the plurality of graphics cores. A fabric port is configured to limit bandwidth available to an associated graphics core via a bandwidth throttler circuit coupled with the fabric port.