Hybrid Computing Workload Manager for Throughput SLA Satisfaction

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional hybrid computing environments lack high-throughput computing capabilities due to inadequate unified specification and management of computing requirements across heterogeneous systems, often scheduling less compute-intensive workloads on servers and more compute-intensive workloads on accelerators, which limits their overall performance and energy efficiency.

Innovation Solution

A workload manager dynamically schedules data-parallel workload tasks across server and accelerator resources based on high-throughput computing service level agreements (SLAs), utilizing surrogate execution and resource guarantees to ensure sustained throughput, and enables workload mobility and redistribution to optimize performance and energy usage.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Productivity

If compute-intensive workloads are scheduled on accelerators and less compute-intensive workloads on servers, then workload distribution is optimized, but high-throughput computing capability is not achieved

Engineering Contradiction:
Improvehigh-throughput computing capabilityVSAvoidunified specification and management of computing requirements
Core Design Contradiction:
ProductivityVSDevice complexity

Solution Approach 1:

The workload manager provides unified specification and management of computing requirements across heterogeneous systems (servers and accelerators). It schedules both compute-intensive and throughput-oriented workloads across the hybrid environment, enabling the system to achieve high-throughput computing capability while maintaining manageable complexity through a single management framework.

Inventive Principle:
Principle #6Universality (Multi-functionality)

2Reliability

If dynamic scheduling of data-parallel workload tasks is implemented, then throughput SLAs are satisfied, but system complexity increases

Engineering Contradiction:
ImproveSLA satisfactionVSAvoiddynamic scheduling mechanism
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The workload manager continuously monitors resource availability and workload status in the hybrid computing environment, using this feedback to dynamically schedule data-parallel workload tasks. This feedback mechanism enables the system to satisfy throughput SLAs by adapting to changing conditions while managing complexity through automated decision-making based on real-time information.

Inventive Principle:
Principle #23Feedback

3Productivity

If resource allocation is optimized for throughput, then computing performance improves, but energy consumption increases

Engineering Contradiction:
Improvethroughput performanceVSAvoidenergy consumption
Core Design Contradiction:
ProductivityVSUse of energy by moving object

Solution Approach 1:

The workload manager dynamically allocates resources based on actual workload characteristics and system state. It can shift workloads between servers and accelerators in real-time, enabling the system to achieve high throughput performance when needed while reducing energy consumption during lower-demand periods through adaptive resource management.

Inventive Principle:
Principle #15Dynamics

Data Source

PatentUS8739171B2High-throughput-computing in a hybrid computing environment
Publication Date: 2014.05.27 INTERNATIONAL BUSINESS MACHINE CORPORATION
  • US8739171B2 patent drawing
  • US8739171B2 patent drawing
  • US8739171B2 patent drawing

AI summary

Embodiments of the present invention provide high-throughput computing in a hybrid processing system. A set of high-throughput computing service level agreements (SLAs) is analyzed. The set of high-throughput computing SLAs are associated with a hybrid processing system. The hybrid processing system includes at least one server system that includes a first computing architecture and a set of accelerator systems each including a second computing architecture that is different from the first computing architecture. A first set of resources at the server system and a second set of resources at the set of accelerator systems are monitored. A set of data-parallel workload tasks is dynamically scheduled across at least one resource in the first set of resources and at least one resource in the second set of resources. The dynamic scheduling of the set of data-parallel workload tasks substantially satisfies the set of high-throughput computing SLAs.