Adaptive GPU Core Virtualization for Multi-VM Resource Utilization
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing data center GPU virtualization techniques result in inefficient utilization of compute resources due to workloads not fully utilizing the entire GPU during time-sliced virtualization, where a single VM owns all resources but may not utilize them fully.
Innovation Solution
Implementing adaptive virtualization of GPU cores and engine-based virtualization, allowing dynamic allocation of resources across multiple virtual machines (VMs) to optimize resource utilization.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Ease of operation
If time-sliced virtualization is used where a single VM owns all GPU resources for a period of time, then resource isolation and simplicity are improved, but resource utilization efficiency deteriorates because the VM may not utilize all execution resources fully
Solution Approach 1:
The patent segments the GPU into multiple virtual GPU cores, each of which can be independently allocated to different VMs. This segmentation allows fine-grained resource distribution while maintaining isolation, resolving the contradiction between simple isolation and efficient utilization.
Solution Approach 2:
The patent implements dynamic allocation where GPU cores can be reassigned between VMs based on real-time workload demands. This dynamic approach allows the system to optimize resource utilization efficiency while maintaining isolation through virtualization layers.
2Reliability
If the entire GPU is dedicated to a single VM during a time slice, then isolation and security are improved, but resource utilization efficiency deteriorates due to idle execution resources
Solution Approach 1:
By dividing the GPU into multiple virtual cores, the system can allocate only the necessary portion of resources to each VM based on its actual needs, reducing idle resources while maintaining isolation through the virtualization layer.
Solution Approach 2:
The patent changes the allocation parameter from whole-GPU time-slicing to fine-grained core-level allocation, allowing the system to adjust the degree of resource dedication based on workload requirements and optimize energy utilization.
3Device complexity
If GPU resources are statically allocated to VMs, then system simplicity is improved, but adaptability to varying workloads deteriorates
Solution Approach 1:
The patent implements dynamic allocation mechanisms where GPU core assignments can change based on real-time workload detection, allowing the system to adapt to varying demands while managing complexity through automated scheduling.
Solution Approach 2:
The system incorporates feedback loops that monitor workload characteristics and adjust GPU core allocations accordingly, enabling adaptability to varying workloads while keeping the management system simple through closed-loop control.
Data Source
AI summary
One embodiment provides a graphics processor comprising a memory interface, a plurality of interfaces to a plurality of compute engines, a processing resource cluster including a plurality of processing resources, the plurality of processing resources configured to execute instructions on behalf of the plurality of compute engines, and virtualization circuitry configured to enable time-sliced virtualization of the plurality of processing resources via the plurality of compute engines, wherein the virtualization circuitry to concurrently process workloads from a plurality of guest software environments during a time-slice via dynamic assignment of the workloads to the plurality of interfaces to the plurality of compute engines.


