Dynamic GPU Provisioning for Data Analytics Workloads

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current GPU-based resource provisioning in cloud environments is static, leading to low GPU utilization due to workload requirements variability, necessitating enhanced techniques for dynamic provisioning and scaling of GPUs for data analytic workloads.

Innovation Solution

A method and apparatus for dynamically allocating and scaling GPUs based on workload requirements, utilizing a GPU configuration determination and scaling component that adjusts GPU resources in real-time according to monitored resource consumption, allowing for fine-grained allocation and de-allocation of GPUs within a pool, enabling efficient use of GPU resources for data analytic workloads.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Ease of operation

If static GPU provisioning is used, then device complexity is reduced and ease of operation is improved, but GPU utilization deteriorates due to workload variability

Engineering Contradiction:
Improveease of GPU provisioningVSAvoidGPU utilization
Core Design Contradiction:
Ease of operationVSProductivity

Solution Approach 1:

The patent implements dynamic GPU provisioning that automatically adjusts resource allocation based on real-time workload monitoring. The system transitions from static pre-allocation to dynamic on-demand allocation, where GPUs are provisioned or de-provisioned based on actual utilization needs, thereby maintaining ease of operation while significantly improving GPU utilization efficiency

Inventive Principle:
Principle #15Dynamics

Solution Approach 2:

The system incorporates continuous workload monitoring and feedback mechanisms that track GPU utilization metrics. Based on this feedback, the provisioning system automatically adjusts resource allocation decisions, scaling GPU resources up or down to match actual demand patterns, thus resolving the contradiction between operational simplicity and resource utilization

Inventive Principle:
Principle #23Feedback

2Reliability

If workloads are assigned to particular GPUs, then reliability is improved through dedicated resources, but adaptability deteriorates when workload requirements vary

Engineering Contradiction:
Improveworkload processing reliabilityVSAvoidworkload requirement adaptability
Core Design Contradiction:
ReliabilityVSAdaptability or versatility

Solution Approach 1:

The patent creates a universal GPU pool where resources can serve multiple workload types dynamically. Instead of dedicating specific GPUs to specific workloads, the system enables any GPU in the pool to handle any workload based on current requirements, thereby maintaining reliability through consistent resource availability while achieving adaptability to varying workload demands

Inventive Principle:
Principle #6Universality (Multi-functionality)

Solution Approach 2:

The system dynamically reassigns GPUs from completed or low-priority workloads to new or high-priority workloads based on real-time conditions. This dynamic reallocation maintains system reliability by ensuring workloads always have available resources while adapting to changing requirements without rigid workload-GPU bindings

Inventive Principle:
Principle #15Dynamics

3Productivity

If more GPUs are provisioned to handle peak workloads, then productivity is improved, but loss of substance increases due to idle resources during low-demand periods

Engineering Contradiction:
Improveworkload processing capacityVSAvoididle GPU resources
Core Design Contradiction:
ProductivityVSLoss of substance

Solution Approach 1:

The patent implements elastic GPU provisioning that scales resource capacity dynamically according to demand fluctuations. During peak periods, the system provisions additional GPUs to maintain productivity, while during low-demand periods, it de-provisions or reassigns those same resources, thereby eliminating idle resource waste while preserving processing capacity when needed

Inventive Principle:
Principle #15Dynamics

Solution Approach 2:

The system temporarily discards (de-provisions) GPU resources during low-utilization periods and recovers them when demand increases. This cyclical provisioning strategy ensures that resources are not permanently allocated and left idle, instead being continuously reused across different workloads based on actual needs, thus reducing substance loss while maintaining productivity

Inventive Principle:
Principle #34Discarding and recovering

Data Source

PatentUS9916636B2Dynamically provisioning and scaling graphic processing units for data analytic workloads in a hardware cloud
Publication Date: 2018.03.13 BLUE HERON DEVELOPMENT LLC
  • US9916636B2 patent drawing
  • US9916636B2 patent drawing
  • US9916636B2 patent drawing

AI summary

Server resources in a data center are disaggregated into shared server resource pools, including a graphics processing unit (GPU) pool. Servers are constructed dynamically, on-demand and based on workload requirements, by allocating from these resource pools. According to this disclosure, GPU utilization in the data center is managed proactively by assigning GPUs to workloads in a fine granularity and agile way, and de-provisioning them when no longer needed. In this manner, the approach is especially advantageous to automatically provision GPUs for data analytic workloads. The approach thus provides for a “micro-service” enabling data analytic workloads to automatically and transparently use GPU resources without providing (e.g., to the data center customer) the underlying provisioning details. Preferably, the approach dynamically determines the number and the type of GPUs to use, and then during runtime auto-scales the GPUs based on workload.