Live GPU State Migration for Cloud Latency Reduction

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Virtual machines deployed in cloud settings offer quick startup but slow runtime performance due to network distance, while micro datacenters provide good performance but with longer startup times, and GPUs or software renderers are often underutilized due to dedicated user allocation and changing application utilization.

Innovation Solution

Implement a system for live migration of GPU states by recording and replaying GPU commands from a source GPU to a destination GPU, predicting downtime, and connecting clients to the destination GPU during low usage periods to optimize resource utilization and reduce latency.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Loss of time

If virtual machines are deployed in cloud settings, then startup time is reduced, but runtime performance deteriorates due to network distance

Engineering Contradiction:
Improvestartup timeVSAvoidruntime performance
Core Design Contradiction:
Loss of timeVSSpeed

Solution Approach 1:

The patent implements dynamic migration of GPU virtual machines between different physical locations (cloud data centers and micro datacenters) based on runtime conditions. The system monitors utilization metrics and automatically relocates VMs to optimize the trade-off between startup time and runtime performance, allowing the deployment architecture to adapt dynamically rather than being static

Inventive Principle:
Principle #15Dynamics

2Speed

If micro datacenters are used, then runtime performance is improved, but startup time increases

Engineering Contradiction:
Improveruntime performanceVSAvoidstartup time
Core Design Contradiction:
SpeedVSLoss of time

Solution Approach 1:

The system enables dynamic migration of GPU VMs between cloud data centers and micro datacenters based on real-time utilization monitoring. When micro datacenters become underutilized, VMs are automatically migrated back to cloud data centers, reducing startup time while maintaining performance benefits during high-utilization periods

Inventive Principle:
Principle #15Dynamics

3Reliability

If GPUs are allocated to dedicated users, then reliability is improved, but device utilization deteriorates due to changing application needs

Engineering Contradiction:
Improveservice reliabilityVSAvoidGPU utilization
Core Design Contradiction:
ReliabilityVSProductivity

Solution Approach 1:

The patent implements GPU virtualization that allows a single physical GPU to serve multiple virtual machines and users simultaneously. The hypervisor manages resource allocation dynamically, enabling the same GPU hardware to be shared across multiple applications and users while maintaining isolation and reliability for each user's workload

Inventive Principle:
Principle #6Universality (Multi-functionality)

Solution Approach 2:

The system dynamically migrates GPU VMs between different physical GPUs and locations based on utilization metrics. When a GPU becomes underutilized, its VMs are migrated to other GPUs, ensuring high overall utilization while maintaining dedicated service reliability for each user through virtualized isolation

Inventive Principle:
Principle #15Dynamics

Data Source

PatentUS9965823B2Migration of graphics processing unit (GPU) states
Publication Date: 2018.05.08 MICROSOFT TECHNOLOGY LICENSING LLC
  • US9965823B2 patent drawing
  • US9965823B2 patent drawing
  • US9965823B2 patent drawing

AI summary

The claimed subject matter includes techniques for live migration of a graphics processing unit (GPU) state. An example method includes receiving recorded GPU commands from a relay at a destination GPU. The method also includes replaying the recorded GPU commands at the destination GPU. The method also includes detecting a downtime for the GPU commands. The method further includes establishing a connection between the destination GPU and the client during the detected downtime.