Multi-GPU Sled Architecture for High Density and Serviceability
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing server and GPU platforms are inefficient in terms of space utilization and require full board replacement if one server fails, limiting server density and performance within a given form factor.
Innovation Solution
A modular chassis configuration with multi-server and multi-GPU sleds that allow for easy serviceability and efficient use of space, featuring a cubby chassis with vertically oriented side-plane PCBs, edge guides, and air duct covers to enhance airflow and signal integrity, enabling the housing of multiple server or GPU cards in a compact form factor.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Quantity of substance
If conventional GPU cards are used, then they provide adequate processing power, but they are very large which limits their density for a given form factor
Solution Approach 1:
The invention segments the traditional single-GPU card architecture by placing multiple GPUs on a single PCB. The PCB is divided into multiple regions, each hosting a GPU die and associated components. This segmentation allows multiple processing units to coexist in a compact form factor, directly increasing GPU density while reducing the volume occupied by each individual GPU card.
Solution Approach 2:
The invention transitions from a vertical stacking arrangement (single GPU per card) to a horizontal plane arrangement (multiple GPUs per card). By utilizing the two-dimensional surface area of the PCB more efficiently and arranging GPUs in a planar configuration rather than vertical stacking, the design maximizes the number of GPUs that can fit within the constrained form factor volume.
2Area of stationary object
If a single PCB with multiple servers is used, then space utilization is improved, but if one server fails, the entire PCB must be replaced
Solution Approach 1:
The invention segments the server functionality into independent modular units that can be individually replaced. Each server component is designed as a separate module on the PCB, allowing failure isolation to specific modules rather than the entire board. This enables targeted replacement of only the failed component, simplifying repair operations while maintaining high space utilization through the multi-server PCB architecture.
3Productivity
If server density is increased within a given form factor, then performance is improved, but serviceability becomes more difficult
Solution Approach 1:
The invention implements dynamic serviceability features including hot-swappable server modules and retractable PCIe brackets that allow easy insertion and removal of server components. The design incorporates movable components and flexible connections that maintain electrical connectivity during module replacement, enabling service personnel to quickly swap out failed or upgraded servers without shutting down the entire system or dealing with complex wiring.
Data Source
AI summary
A multiple graphics processing unit (multi-GPU) platform including a cubby chassis and at least one multi-GPU sled. The cubby chassis includes partitions defining a plurality of sled positions. The multi-GPU sled includes a chassis having a vertical sidewall and a horizontal bottom wall with an open top and an open side. A side-plane PCB is mounted to the vertical sidewall and a plurality of dividers are attached to the bottom wall and oriented perpendicular to the side-plane PCB. One or more GPU cards are connected to the side-plane PCB and are supported on the plurality of dividers. The GPU cards include a GPU PCB having a first side facing the bottom wall and an outward facing second side. A cover is coupled to the horizontal bottom wall to enclose the open side of the sled chassis and help direct airflow across the GPU cards.


