Remote GPU Access via DPU Protocol Offloading

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

The existing GPU remote virtualization technologies face inefficiencies due to cumbersome and inefficient I/O operations, which increase communication delay and waste GPU clock cycles, thereby reducing overall computing throughput.

Innovation Solution

A method and device for remotely accessing graphics processing units, where a virtual machine generates target information including an application request, and a processor determines a target service module and transmission path at a second node connected to a target graphics processing unit, allowing the application request to be sent and processing results to be received efficiently.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If data is transmitted through multiple VMM layers and software protocol stacks, then remote GPU access is enabled, but communication delay increases and computing throughput decreases

Engineering Contradiction:
Improveremote GPU access capabilityVSAvoidcommunication delay
Core Design Contradiction:
Adaptability or versatilityVSLoss of time

Solution Approach 1:

The patent extracts the network protocol stack processing from the software layer (VMM and host OS) and relocates it to a dedicated hardware device (DPU). This extraction eliminates the need for data to pass through multiple software layers, directly reducing communication delay while preserving remote GPU access capability.

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The patent introduces a DPU as an intermediary device between the virtual machine and the remote GPU. This intermediary handles network protocol processing in hardware, acting as a mediator that reduces the communication path and eliminates software layer overhead, thereby reducing communication delay.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Ease of operation

If complex network protocol stack processing is performed in software, then network communication is achieved, but GPU clock cycles are wasted on I/O operations

Engineering Contradiction:
Improvenetwork communication capabilityVSAvoidcomputing throughput
Core Design Contradiction:
Ease of operationVSProductivity

Solution Approach 1:

The patent replaces the mechanical software-based network protocol processing with a hardware-based system (DPU). This substitution moves protocol stack processing from the GPU's software environment to dedicated hardware, preventing GPU clock cycles from being wasted on I/O operations and improving computing throughput.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Solution Approach 2:

The DPU acts as an intermediary that handles all network protocol processing, freeing the GPU from I/O operations. This mediator absorbs the computational overhead of network communication, allowing the GPU to focus on computing tasks and thereby improving overall productivity.

Inventive Principle:
Principle #24Intermediary (Mediator)

3Adaptability or versatility

If data is copied multiple times through software layers, then remote transmission is achieved, but I/O operation efficiency decreases

Engineering Contradiction:
Improveremote transmission capabilityVSAvoidI/O operation complexity
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The patent extracts the data copying and protocol processing operations from the software layers and consolidates them in the hardware DPU. This extraction reduces the number of times data needs to be copied through multiple software layers, simplifying I/O operations while maintaining remote transmission capability.

Inventive Principle:
Principle #2Taking out (Extraction)

Data Source

PatentUS20250191108A1Method and device for remotely accessing graphics processing units
Publication Date: 2025.06.12 LENOVO (BEIJING) LTD
  • US20250191108A1 patent drawing
  • US20250191108A1 patent drawing
  • US20250191108A1 patent drawing

AI summary

A method applied to a first node including a first processor and a virtual machine includes the virtual machine generating target information including an application request and satisfying a preset transmission information rule for transmission between a virtual graphics processing unit driver module of the virtual machine and the first processor, and the first processor determining, based on the target information, a target service module at a second node and connected to a target graphics processing unit and a transmission path connecting the virtual graphics processing unit driver, the first processor, a second processor at the second node, and the target service module, sending the application request to the target service module based on the transmission path, and receiving a processing result fed back by the second processor through the transmission path.