Adaptive GPU Core Virtualization for Multi-VM Resource Utilization

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing data center GPU virtualization techniques result in inefficient utilization of compute resources due to workloads not fully utilizing the entire GPU during time-sliced virtualization, where a single VM owns all resources but may not utilize them fully.

Innovation Solution

Implementing adaptive virtualization of GPU cores and engine-based virtualization, allowing dynamic allocation of resources across multiple virtual machines (VMs) to optimize resource utilization.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Ease of operation

If time-sliced virtualization is used where a single VM owns all GPU resources for a period of time, then resource isolation and simplicity are improved, but resource utilization efficiency deteriorates because the VM may not utilize all execution resources fully

Engineering Contradiction:
Improveresource isolationVSAvoidresource utilization efficiency
Core Design Contradiction:
Ease of operationVSProductivity

Solution Approach 1:

The patent segments the GPU into multiple virtual GPU cores, each of which can be independently allocated to different VMs. This segmentation allows fine-grained resource distribution while maintaining isolation, resolving the contradiction between simple isolation and efficient utilization.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent implements dynamic allocation where GPU cores can be reassigned between VMs based on real-time workload demands. This dynamic approach allows the system to optimize resource utilization efficiency while maintaining isolation through virtualization layers.

Inventive Principle:
Principle #15Dynamics

2Reliability

If the entire GPU is dedicated to a single VM during a time slice, then isolation and security are improved, but resource utilization efficiency deteriorates due to idle execution resources

Engineering Contradiction:
ImproveisolationVSAvoididle execution resources
Core Design Contradiction:
ReliabilityVSLoss of energy

Solution Approach 1:

By dividing the GPU into multiple virtual cores, the system can allocate only the necessary portion of resources to each VM based on its actual needs, reducing idle resources while maintaining isolation through the virtualization layer.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent changes the allocation parameter from whole-GPU time-slicing to fine-grained core-level allocation, allowing the system to adjust the degree of resource dedication based on workload requirements and optimize energy utilization.

Inventive Principle:
Principle #35Parameter changes

3Device complexity

If GPU resources are statically allocated to VMs, then system simplicity is improved, but adaptability to varying workloads deteriorates

Engineering Contradiction:
Improveallocation managementVSAvoidworkload adaptability
Core Design Contradiction:
Device complexityVSAdaptability or versatility

Solution Approach 1:

The patent implements dynamic allocation mechanisms where GPU core assignments can change based on real-time workload detection, allowing the system to adapt to varying demands while managing complexity through automated scheduling.

Inventive Principle:
Principle #15Dynamics

Solution Approach 2:

The system incorporates feedback loops that monitor workload characteristics and adjust GPU core allocations accordingly, enabling adaptability to varying workloads while keeping the management system simple through closed-loop control.

Inventive Principle:
Principle #23Feedback

Data Source

PatentUS20250291620A1Adaptive virtualization of GPU cores and engine based virtualization
Publication Date: 2025.09.18 INTEL CORP
  • US20250291620A1 patent drawing
  • US20250291620A1 patent drawing
  • US20250291620A1 patent drawing

AI summary

One embodiment provides a graphics processor comprising a memory interface, a plurality of interfaces to a plurality of compute engines, a processing resource cluster including a plurality of processing resources, the plurality of processing resources configured to execute instructions on behalf of the plurality of compute engines, and virtualization circuitry configured to enable time-sliced virtualization of the plurality of processing resources via the plurality of compute engines, wherein the virtualization circuitry to concurrently process workloads from a plurality of guest software environments during a time-slice via dynamic assignment of the workloads to the plurality of interfaces to the plurality of compute engines.