Dynamic FPGA Allocation for Data Analytics Workloads

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current accelerator-based resource provisioning in cloud environments is static, leading to low utilization and requires expertise, making it inefficient for dynamic data analytic workloads.

Innovation Solution

A method and system for dynamically allocating and scaling accelerators based on workload requirements, using a pool of resources that adjusts the number and type of accelerators during runtime based on monitored resource consumption, allowing for fine-grained allocation and de-allocation.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If static accelerator provisioning is used, then hardware resources are allocated in advance, but accelerator utilization is low and requires expertise

Engineering Contradiction:
Improveaccelerator utilizationVSAvoidprovisioning complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The patent implements dynamic accelerator provisioning by allowing the system to automatically adjust accelerator allocation based on real-time workload characteristics. The accelerator manager monitors workload progress and dynamically adds or removes accelerators from workload instances, transitioning from static pre-provisioning to dynamic runtime adjustment, thereby improving utilization without requiring manual expertise

Inventive Principle:
Principle #15Dynamics

Solution Approach 2:

The system employs self-service mechanisms where the accelerator manager automatically provisions, scales, and manages accelerator resources without human intervention. Workload instances self-report their accelerator needs, and the manager autonomously allocates appropriate accelerator counts, eliminating the need for user expertise while maintaining high utilization efficiency

Inventive Principle:
Principle #25Self-service

2Productivity

If dynamic accelerator scaling is implemented, then accelerator utilization is improved, but system complexity increases

Engineering Contradiction:
Improveworkload processing performanceVSAvoidsystem architecture complexity
Core Design Contradiction:
ProductivityVSDevice complexity

Solution Approach 1:

The patent segments the accelerator management function into a dedicated accelerator manager component that operates independently from workload instances. This segmentation allows dynamic scaling functionality to be centralized and managed separately, reducing the complexity burden on individual workload instances while maintaining overall system productivity

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The accelerator manager acts as an intermediary between workload instances and the accelerator hardware pool. It mediates the complex interactions by abstracting accelerator management details from workloads, providing a simplified interface for dynamic scaling while handling the underlying complexity of resource allocation and hardware communication

Inventive Principle:
Principle #24Intermediary (Mediator)

Data Source

PatentUS11275622B2Utilizing accelerators to accelerate data analytic workloads in disaggregated systems
Publication Date: 2022.03.15 INTERNATIONAL BUSINESS MACHINE CORPORATION
  • US11275622B2 patent drawing
  • US11275622B2 patent drawing
  • US11275622B2 patent drawing

AI summary

Server resources in a data center are disaggregated into shared server resource pools, including an accelerator (e.g., FPGA) pool. Servers are constructed dynamically, on-demand and based on workload requirements, by allocating from these resource pools. According to this disclosure, accelerator utilization in the data center is managed proactively by assigning accelerators to workloads in a fine granularity and agile way, and de-provisioning them when no longer needed. In this manner, the approach is especially advantageous to automatically provision accelerators for data analytic workloads. The approach thus provides for a “micro-service” enabling data analytic workloads to automatically and transparently use FPGA resources without providing (e.g., to the data center customer) the underlying provisioning details. Preferably, the approach dynamically determines the number and the type of FPGAs to use, and then during runtime auto-scales the FPGAs based on workload.