GPU Virtualization with Abstract uGPU Layer for Scalable Over-provisioning
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current GPU virtualization methods in client/server architectures lack scalable over-provisioning capabilities, leading to underutilization of resources as allocated vGPUs become idle and cannot be easily migrated or shared among users.
Innovation Solution
The introduction of an abstract layer, referred to as a unique GPU (uGPU), which decouples vGPU allocation from specific GPU servers, allowing for over-provisioning by mapping multiple uGPUs to a single vGPU and enabling dynamic migration based on load balancing policies.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Productivity
If vGPUs are allocated to specific GPU servers, then allocation decisions can be made, but resource utilization decreases when vGPUs become idle and cannot be easily migrated
Solution Approach 1:
The patent introduces an intermediate abstraction layer between the client and physical GPU servers. This abstraction layer decouples the allocation decision from the actual vGPU instance, allowing the system to manage and migrate virtualized processing units without directly binding clients to specific servers. The intermediary enables flexible resource allocation while maintaining simple client-side interfaces.
Solution Approach 2:
The patent segments the allocation process into two independent components: the allocation decision (which client gets which type of processing unit) and the actual vGPU instance (which physical server hosts it). This segmentation allows the system to independently optimize resource utilization by migrating vGPU instances without affecting allocation decisions, thereby resolving the contradiction between productivity and device complexity.
2Adaptability or versatility
If vGPUs are bound to specific servers, then stable allocation is achieved, but flexibility and adaptability decrease when load balancing is needed
Solution Approach 1:
The patent implements dynamic resource allocation where the mapping between abstract processing units and physical vGPU instances can change over time. The system can migrate vGPU instances between servers based on load conditions, resource availability, and performance requirements. This dynamic approach maintains allocation stability for clients while enabling adaptability for system-wide load balancing and resource optimization.
Data Source
AI summary
Techniques are disclosed for processing unit virtualization with scalable over-provisioning in an information processing system. For example, the method accesses a data structure that maps a correspondence between a plurality of virtualized processing units and a plurality of abstracted processing units, wherein the plurality of abstracted processing units are configured to decouple an allocation decision from the plurality of virtualized processing units, and further wherein at least one of the virtualized processing units is mapped to multiple ones of the abstracted processing units. The method allocates one or more virtualized processing units to execute a given application by allocating one or more abstracted processing units identified from the data structure. The method also enables migration of one or more virtualized processing units across the system. Examples of processing units with which scalable over-provisioning functionality can be applied include, but are not limited to, accelerators such as GPUs.


