GPU-remoting Latency Aware VM Migration

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Resource schedulers in datacenters often fail to adequately consider Graphics Processing Units (GPUs) and GPU-based workloads for load balancing and migration decisions, leading to suboptimal performance in environments with remoting-enabled GPUs.

Innovation Solution

Implementing GPU-remoting latency aware virtual machine migration mechanisms within resource schedulers to optimize the placement and migration of GPU-based workloads, taking into account GPU-remoting latency, host memory, and compute capability for balanced resource allocation.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Productivity

If resource schedulers use traditional load balancing methods without GPU-specific considerations, then the scheduling system remains simple and easy to implement, but GPU-based workloads experience suboptimal performance and resource allocation

Engineering Contradiction:
ImproveGPU workload performanceVSAvoidscheduling system complexity
Core Design Contradiction:
ProductivityVSDevice complexity

Solution Approach 1:

The patent modifies traditional scheduling parameters by introducing GPU-specific metrics including GPU-remoting latency, host memory capacity, and compute capability. The scheduler transitions from generic resource scheduling to GPU-aware scheduling by incorporating these additional parameters, enabling optimized placement and migration decisions that improve GPU workload performance while maintaining manageable system complexity through structured parameter integration

Inventive Principle:
Principle #35Parameter changes

2Loss of time

If virtual machines are migrated without considering GPU-remoting latency, then migration decisions are simplified and faster, but latency increases and workload performance deteriorates

Engineering Contradiction:
ImproveGPU-remoting latencyVSAvoidworkload performance
Core Design Contradiction:
Loss of timeVSProductivity

Solution Approach 1:

The patent implements feedback mechanisms where the scheduler continuously monitors GPU-remoting latency and uses this information to make informed migration decisions. The system measures actual latency values, compares them against thresholds and historical data, and adjusts migration strategies accordingly, creating a closed-loop control system that minimizes latency while maintaining workload performance

Inventive Principle:
Principle #23Feedback

Solution Approach 2:

The system performs preliminary latency measurements and host capability assessments before migration decisions are made. By pre-evaluating GPU-remoting latency and host resources in advance, the scheduler can predict migration outcomes and select optimal destination hosts beforehand, avoiding post-migration latency issues and improving overall workload performance

Inventive Principle:
Principle #10Preliminary action

3Productivity

If GPUs are virtualized and shared across multiple workloads, then resource utilization increases and hardware costs decrease, but resource contention and latency increase

Engineering Contradiction:
Improveresource utilizationVSAvoidresource contention
Core Design Contradiction:
ProductivityVSObject-affected harmful factors

Solution Approach 1:

The patent applies local quality principles by assigning different weights and priorities to GPU resources based on individual workload requirements and host characteristics. The scheduler considers factors such as workload sensitivity to latency, host GPU-remoting capabilities, and local resource availability to make differentiated scheduling decisions, thereby reducing resource contention while maintaining high utilization through targeted resource allocation

Inventive Principle:
Principle #3Local quality

4Productivity

If traditional load balancing is used without GPU awareness, then the scheduling system remains simple, but placement decisions fail to optimize for GPU-specific resources

Engineering Contradiction:
Improveload balancing efficiencyVSAvoidscheduling mechanism complexity
Core Design Contradiction:
ProductivityVSDevice complexity

Solution Approach 1:

The patent extends the scheduling system to serve multiple functions simultaneously: traditional CPU-based load balancing, GPU resource allocation, latency optimization, and migration management. By integrating these functions into a unified GPU-aware scheduler, the system achieves efficient multi-objective optimization without requiring separate independent systems, thereby improving load balancing efficiency while controlling overall complexity through functional consolidation

Inventive Principle:
Principle #6Universality (Multi-functionality)

Data Source

PatentUS11886898B2GPU-remoting latency aware virtual machine migration
Publication Date: 2024.01.30 VMWARE INC
  • US11886898B2 patent drawing
  • US11886898B2 patent drawing
  • US11886898B2 patent drawing

AI summary

Various aspects are disclosed for graphics processing unit (GPU)-remoting latency aware migration. In some aspects, a host executes a GPU-remoting client that includes a GPU workload. GPU-remoting latencies are identified for hosts of a cluster. A destination host is identified based on having a lower GPU-remoting latency than the host currently executing the GPU-remoting client. The GPU-remoting client is migrated from its current host to the destination host.