GPU Acceleration via PCIe Switch Topology

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

The existing GPU scheduling methods create a bandwidth bottleneck in the PCIe bus, limiting the acceleration potential of graphics processing units (GPUs) due to inherent limitations, which restricts their maximum utilization in deep learning and engineering applications.

Innovation Solution

A GPU accelerating device that includes a communication unit, processor, and storage device, allowing multiple GPUs to interact with CPUs and each other through switches, with a method to dynamically arrange GPU resources based on user requests, optimizing data transmission by adjusting the number and arrangement of GPUs and switches to maximize bandwidth and processing efficiency.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Ease of operation

If a job scheduler (SLURM/LSF/BPS) is used to allocate tasks to GPUs, then task scheduling is achieved, but a bandwidth bottleneck is created in the PCIe bus

Engineering Contradiction:
Improvetask schedulingVSAvoidGPU acceleration performance
Core Design Contradiction:
Ease of operationVSProductivity

Solution Approach 1:

The system segments the GPU pool into multiple groups, with each group having a dedicated PCIe switch. This segmentation allows parallel data transmission paths, preventing the single-bus bottleneck while maintaining centralized scheduling control through the management module.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent introduces PCIe switches as intermediary devices between the management module and GPU groups. These switches act as mediators that enable high-speed data transmission without congesting the main PCIe bus, thus resolving the bandwidth bottleneck while preserving scheduling functionality.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Quantity of substance

If multiple GPUs are connected through PCIe bus, then GPU resources are available for processing, but bandwidth limitations restrict maximum utilization

Engineering Contradiction:
Improvenumber of GPUsVSAvoiddata transmission bandwidth
Core Design Contradiction:
Quantity of substanceVSProductivity

Solution Approach 1:

The patent transitions from a single-dimension PCIe bus architecture to a multi-dimensional network topology by introducing PCIe switches that create multiple transmission paths. This dimensional change allows data to flow through parallel channels, exponentially increasing bandwidth capacity while supporting larger numbers of GPUs.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

Solution Approach 2:

The system merges multiple PCIe switches into a coordinated network under management module control. This merging creates a unified high-bandwidth infrastructure that supports multiple GPUs simultaneously without the limitations of individual bus connections, achieving scalable bandwidth expansion.

Inventive Principle:
Principle #5Merging (Combining)

Data Source

PatentUS10867363B2Device and method for accelerating graphics processor units, and computer readable storage medium
Publication Date: 2020.12.15 SHENZHEN FULIAN FUGUI PRECISION INDUSTRY CO LTD
  • US10867363B2 patent drawing
  • US10867363B2 patent drawing
  • US10867363B2 patent drawing

AI summary

A method for accelerating graphics processing units (GPUs) receives a request for usage of GPU resource sent by a user, calculates a quantity of GPUs which are necessary, and arrange the GPUs in several ways to maximize data transmission from and between the GPUs, and between the GPUs and one or more central processing units (CPUs) connected by switches between the GPUs and the CPUs. A device for accelerating GPUs is also provided.