Local PCIe Switch for NVM to GPU Data Transfer

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

The existing methods for data transfer between non-volatile memory (NVM) and graphics processing unit (GPU) local memory require involvement of the host computing system's root complex, leading to increased traffic and congestion.

Innovation Solution

Incorporating a local PCIe switch within the solid state graphics (SSG) card to enable direct peer-to-peer data transfer between NVM and local memory, bypassing the PCIe root complex, allowing the GPU to initiate data transfers without processor intervention.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If data transfer from NVM to GPU local memory goes through host memory and root complex, then data transfer can be performed using existing memory architectures, but system latency increases and performance decreases

Engineering Contradiction:
Improvedata transfer capabilityVSAvoidsystem latency
Core Design Contradiction:
ReliabilityVSLoss of time

Solution Approach 1:

The patent segments the data transfer path by introducing a local PCIe switch that creates a dedicated transfer channel between NVM and GPU local memory, separating this traffic from the host memory path through the root complex. This segmentation allows simultaneous independent operations and reduces contention on shared resources.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The local PCIe switch acts as an intermediary device that enables direct peer-to-peer data transfer between NVM and GPU local memory. This intermediary provides the necessary protocol handling and routing functions locally, eliminating the need for data to traverse the host memory and root complex, thereby reducing latency while maintaining transfer capability.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Adaptability or versatility

If data transfer involves root complex of host computing system, then data can be transferred between different memory architectures, but traffic and congestion increase

Engineering Contradiction:
Improvememory architecture compatibilityVSAvoidsystem throughput
Core Design Contradiction:
Adaptability or versatilityVSProductivity

Solution Approach 1:

The patent segments system traffic into separate pathways: local peer-to-peer transfers between NVM and GPU local memory through the local PCIe switch, and host memory access through the root complex. This segmentation prevents NVM-GPU transfer traffic from congesting the host memory path, maintaining high throughput for both operations simultaneously.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The local PCIe switch serves as an intermediary that handles protocol conversion and data routing between NVM and GPU local memory locally, eliminating the need for these transfers to traverse the root complex. This maintains adaptability between different memory architectures while preserving system throughput by keeping traffic local.

Inventive Principle:
Principle #24Intermediary (Mediator)

3Ease of operation

If processor intervention is required for data transfer, then data transfer can be initiated and controlled, but resource usage increases and efficiency decreases

Engineering Contradiction:
Improvetransfer control capabilityVSAvoidprocessor resource usage
Core Design Contradiction:
Ease of operationVSUse of energy by moving object

Solution Approach 1:

The patent implements self-service by enabling the GPU to autonomously initiate and control data transfers from NVM to its local memory through the local PCIe switch without processor intervention. The GPU can directly issue memory read commands to the NVM controller, and the local switch handles the transfer automatically, freeing the processor from involvement in routine data transfer operations while maintaining control capability.

Inventive Principle:
Principle #25Self-service

Data Source

PatentUS10678733B2Apparatus for connecting non-volatile memory locally to a GPU through a local switch
Publication Date: 2020.06.09 ATI TECHNOLOGIES ULC
  • US10678733B2 patent drawing
  • US10678733B2 patent drawing
  • US10678733B2 patent drawing

AI summary

Described herein are a method and device for transferring data in a computer system. The device includes a host processor, a plurality of first memory architectures, a switch, a redundant array of independent drives (RAID) assist unit; and a second memory architecture. The host processor is configured to send a data transfer command to the RAID assist unit via the switch. The RAID assist unit is configured to create a set of parallel memory transactions between the plurality of first memory architectures and the second memory architecture, execute the set of parallel memory transactions via the local switch and absent interaction with the host processor; and notify the host processor upon completion of data transfer. In an implementation, the plurality of first memory architectures is non-volatile memories (NVMs) and the second memory architecture is local memory.