Internal GPU-NIC Parallel Switch Fabric for AI Scalability
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional Ethernet AI fabrics face scalability issues, limiting the number of server devices and GPUs that can be interconnected, which may not satisfy future demands.
Innovation Solution
An internal parallel switch fabric is introduced to connect Graphics Processing Units (GPUs) directly to Network Interface Controllers (NICs), allowing each GPU access to multiple external fabrics, thereby increasing scalability.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If conventional PCIe switch devices are used to connect GPUs to NICs, then the system can be manufactured with existing hardware, but the scalability of the Ethernet AI fabric is limited
Solution Approach 1:
The internal switch fabric is segmented into multiple PCIe switch devices, each handling specific GPU-NIC connections. This segmentation allows the system to scale by adding more switch devices rather than requiring a single complex switch, thereby improving scalability while managing device complexity through modular architecture
Solution Approach 2:
The patent introduces a new dimensional approach by creating multiple parallel paths between GPUs and NICs through the internal switch fabric. Instead of a single connection dimension, GPUs can now access multiple NICs through different switch fabric paths, enabling exponential scalability as more GPUs and NICs are added to the system
2Quantity of substance
If each GPU is mapped to a single NIC, then the configuration is simple, but the number of interconnected GPUs is limited
Solution Approach 1:
The internal switch fabric provides universal connectivity where each GPU can access multiple NICs and each NIC can serve multiple GPUs. This multi-functional capability allows any GPU to be mapped to any NIC dynamically, increasing the number of interconnected GPUs while the switch fabric manages the complexity of mappings through its switching logic
Solution Approach 2:
The GPU-NIC mapping becomes dynamic rather than static. The internal switch fabric can dynamically route connections between GPUs and NICs based on workload requirements, allowing the system to adaptively increase the number of interconnected GPUs without requiring complex manual configuration of fixed mappings
3Adaptability or versatility
If multiple NICs are connected to each GPU via internal switch fabric, then access to external fabrics is improved, but the internal switch fabric complexity increases
Solution Approach 1:
The internal switch fabric acts as an intermediary layer between GPUs and NICs. Rather than directly connecting each GPU to multiple NICs (which would create complex point-to-point wiring), the switch fabric mediates all connections through standardized switch ports, improving access to external fabrics while the switch fabric itself manages the complexity of multiple connections
Data Source
AI summary
An internal Graphics Processing Unit (GPU)/Network Interface Controller (NIC) parallel switch fabric system includes a computing device that is coupled to the plurality of external fabrics. The computing device includes a NIC set that provides access to each of the plurality of external fabrics, and a plurality of GPUs. The computing device also includes an internal switch fabric that is configured to couple each of the plurality of GPUs to the NIC set to provide each of the plurality of GPUs access to each of the plurality of external fabrics.


