Latency-Aware Instruction Scheduling for GPU-CPU Alignment

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing techniques for allocating processing resources in computing devices result in inefficient scheduling due to variations in performance across different computing clusters, leading to unpredictability in workload completion times.

Innovation Solution

A scheduling system that allocates CPU and GPU resources within a shared cluster by using a GPU Affinity aware fitness algorithm, which enforces constraints on job sizes and placement, ensuring optimal resource allocation based on latency and communication speeds between nodes.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Productivity

If existing scheduling techniques are used to allocate processing resources across computing clusters, then resource allocation can be performed, but scheduling efficiency deteriorates due to performance variations across clusters

Engineering Contradiction:
Improvescheduling efficiencyVSAvoidpredictability of scheduling
Core Design Contradiction:
ProductivityVSReliability

Solution Approach 1:

The patent changes the scheduling parameters from simple resource allocation to latency-aware placement. The fitness function incorporates latency values of interconnects as a key parameter, transforming the scheduling decision-making process to consider communication costs between processors, thereby resolving the contradiction between scheduling efficiency and predictability.

Inventive Principle:
Principle #35Parameter changes

Solution Approach 2:

The patent replaces traditional mechanical/resource-based scheduling mechanisms with a latency-aware scheduling model. Instead of solely based on resource availability and simple task assignment, the system uses latency measurements of interconnects to determine optimal processor placement, substituting the scheduling mechanism with a more intelligent latency-driven approach.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

2Power

If computing clusters are used for workload scheduling, then processing capacity is increased, but performance variation across clusters worsens scheduling predictability

Engineering Contradiction:
Improveprocessing capacityVSAvoidscheduling predictability
Core Design Contradiction:
PowerVSReliability

Solution Approach 1:

The patent implements feedback mechanisms where latency values of interconnects are measured and fed back into the scheduling process. The fitness function uses these latency feedback values to adjust processor placement decisions, creating a closed-loop system that maintains scheduling predictability despite increased processing capacity across heterogeneous clusters.

Inventive Principle:
Principle #23Feedback

3Speed

If resources are allocated without considering interconnect latency, then allocation speed is improved, but communication efficiency deteriorates

Engineering Contradiction:
Improveallocation speedVSAvoidcommunication efficiency
Core Design Contradiction:
SpeedVSLoss of energy

Solution Approach 1:

The patent applies preliminary action by measuring and pre-calculating latency values of interconnects before actual workload scheduling occurs. These pre-measured latency values are stored and used in the fitness function during scheduling, allowing the system to make optimal placement decisions without real-time communication testing, thus maintaining allocation speed while improving communication efficiency.

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS20240069964A1Scheduling instructions using latency of interconnects of processors
Publication Date: 2024.02.29 NVIDIA CORP
  • US20240069964A1 patent drawing
  • US20240069964A1 patent drawing
  • US20240069964A1 patent drawing

AI summary

Apparatuses, systems, and techniques for scheduling instructions in a cluster to guarantee GPU-CPU alignment for these instructions. In at least one embodiment, jobs are scheduled based on constraints on job sizes and job placement. In at least one embodiment, a processor comprises circuits to schedule instructions to be performed by processors based on latency of interconnects coupled to these processors.