GPU Acceleration via PCIe Switch Topology
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
The existing GPU scheduling methods create a bandwidth bottleneck in the PCIe bus, limiting the acceleration potential of graphics processing units (GPUs) due to inherent limitations, which restricts their maximum utilization in deep learning and engineering applications.
Innovation Solution
A GPU accelerating device that includes a communication unit, processor, and storage device, allowing multiple GPUs to interact with CPUs and each other through switches, with a method to dynamically arrange GPU resources based on user requests, optimizing data transmission by adjusting the number and arrangement of GPUs and switches to maximize bandwidth and processing efficiency.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Ease of operation
If a job scheduler (SLURM/LSF/BPS) is used to allocate tasks to GPUs, then task scheduling is achieved, but a bandwidth bottleneck is created in the PCIe bus
Solution Approach 1:
The system segments the GPU pool into multiple groups, with each group having a dedicated PCIe switch. This segmentation allows parallel data transmission paths, preventing the single-bus bottleneck while maintaining centralized scheduling control through the management module.
Solution Approach 2:
The patent introduces PCIe switches as intermediary devices between the management module and GPU groups. These switches act as mediators that enable high-speed data transmission without congesting the main PCIe bus, thus resolving the bandwidth bottleneck while preserving scheduling functionality.
2Quantity of substance
If multiple GPUs are connected through PCIe bus, then GPU resources are available for processing, but bandwidth limitations restrict maximum utilization
Solution Approach 1:
The patent transitions from a single-dimension PCIe bus architecture to a multi-dimensional network topology by introducing PCIe switches that create multiple transmission paths. This dimensional change allows data to flow through parallel channels, exponentially increasing bandwidth capacity while supporting larger numbers of GPUs.
Solution Approach 2:
The system merges multiple PCIe switches into a coordinated network under management module control. This merging creates a unified high-bandwidth infrastructure that supports multiple GPUs simultaneously without the limitations of individual bus connections, achieving scalable bandwidth expansion.
Data Source
AI summary
A method for accelerating graphics processing units (GPUs) receives a request for usage of GPU resource sent by a user, calculates a quantity of GPUs which are necessary, and arrange the GPUs in several ways to maximize data transmission from and between the GPUs, and between the GPUs and one or more central processing units (CPUs) connected by switches between the GPUs and the CPUs. A device for accelerating GPUs is also provided.


