Resource Quota Management for Shared GPU Processing Units
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current systems lack effective Quality of Service (QoS) control over shared dedicated processors, such as GPUs, in public clouds and data centers, leading to inefficient resource utilization and impact on applications with varying priorities.
Innovation Solution
A method and system that set predetermined resource quotas based on application priority, allocating dedicated processing unit resources only when within the quota and shifting to general-purpose processing units when limits are approached, ensuring higher-priority applications receive necessary resources while minimizing impact on lower-priority ones.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Productivity
If shared GPU instances are deployed to provide computing-intensive services, then service capability is improved, but resource allocation fairness and quality of service control deteriorate
Solution Approach 1:
The patent changes the resource allocation parameter from unlimited sharing to quota-based allocation. By introducing predetermined resource quotas associated with different priority levels, the system transforms the GPU resource distribution mechanism to ensure quality of service control while maintaining shared deployment capabilities.
Solution Approach 2:
The patent segments the GPU resource allocation into different priority levels with corresponding quotas. By dividing resources into segments based on priority (e.g., high, medium, low priority levels with different quota allocations), the system achieves both shared service capability and controlled quality of service.
2Reliability
If resource quotas are imposed on applications, then quality of service control is improved, but system flexibility and adaptability deteriorate
Solution Approach 1:
The patent implements dynamic resource allocation where the system can adaptively assign GPU resources based on application priority levels and current availability. The controller dynamically adjusts resource distribution within the quota framework, allowing the system to respond to varying workload demands while maintaining quality of service control.
Solution Approach 2:
The patent creates a universal resource management framework that handles multiple application types and priority levels through a single quota-based system. This multi-functional approach allows the same mechanism to serve diverse computing workloads while maintaining flexibility and adaptability across different scenarios.
3Productivity
If dedicated processors are shared among multiple applications, then resource utilization efficiency is improved, but application performance consistency deteriorates
Solution Approach 1:
The patent changes the performance parameter by introducing priority-based quota mechanisms that guarantee minimum resource allocations. By adjusting the allocation parameters based on priority levels, the system maintains both efficient resource utilization and consistent application performance for high-priority workloads.
Solution Approach 2:
The patent applies preliminary action by pre-establishing resource quotas and priority levels before applications start executing. This preliminary configuration ensures that performance requirements are met from the outset while maintaining efficient shared resource utilization, preventing performance degradation during execution.
Data Source
AI summary
Embodiments of the present disclosure provide a method, a server system and a computer program product of managing resources. The method may comprise receiving a request for a first amount of resources of the dedicated processing unit from an application with an assigned priority. The method may further comprise determining a total amount of resources of the dedicated processing unit to be occupied by the application based on the request. The method may also comprise in response to the total amount approximating or exceeding a predetermined quota associated with the priority, allocating the first amount of resources of the general-purpose processing unit to the application.


