Dynamic AI Accelerator Topology via Flexible Cable Interconnects
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing AI accelerator chip topologies are either clumsy or require substantial hardware overhead, making them inefficient for distributed AI training, which necessitates the development of a more dynamic and flexible connectivity solution.
Innovation Solution
The implementation of a data processing system with dynamically activatable and deactivatable inter-card and inter-chip connections using cable connections, such as CCIX, allows for the creation of AI chip topologies of varying sizes by activating or deactivating connections between base boards, enabling efficient AI model training with reduced hardware overhead.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If PCB wires on base board are used to connect AI accelerator chips, then small topology can be built, but the approach becomes clumsy and requires substantial hardware overhead
Solution Approach 1:
The patent implements dynamically reconfigurable interconnects that can be activated or deactivated based on training requirements. Cable connections between base boards can be dynamically enabled or disabled to create topologies of varying sizes, allowing the system to adapt from small to large configurations without permanent hardwired connections for all possible configurations.
Solution Approach 2:
The system divides the AI accelerator cluster into multiple base boards, each containing multiple chips. These base boards can be independently connected or disconnected via cable interconnects, allowing modular assembly of different topology sizes. Each base board segment can function semi-independently and be combined in various configurations.
2Adaptability or versatility
If Ethernet is used to connect different base boards for large topology, then scalability is improved, but hardware overhead and complexity increase substantially
Solution Approach 1:
The cable interconnects serve multiple functions: they provide high-speed communication between base boards for large topology configurations, can be individually activated or deactivated for dynamic reconfiguration, and replace the need for extensive PCB wiring infrastructure. The same physical infrastructure supports both small and large topologies through selective activation.
Solution Approach 2:
Instead of creating unique hardwired connections for each possible topology configuration, the system uses reusable cable connections that can be plugged and unplugged to create different topologies. The same set of cables can be configured multiple times for different training needs without requiring dedicated wiring for each configuration.
3Device complexity
If fixed topology configurations are used, then hardware overhead is reduced, but adaptability to different training needs is limited
Solution Approach 1:
The system employs dynamically controllable interconnects that can be activated or deactivated through control logic. This allows the hardware overhead to remain minimal (cables connected but not necessarily active) while providing full adaptability to switch between different topology configurations based on training requirements.
Solution Approach 2:
The system changes the operational state of interconnects (activated/deactivated) rather than physically reconfiguring the hardware structure. By changing the activation parameter of existing cable connections, the system achieves different topology configurations without adding or removing physical hardware, maintaining low overhead while providing high adaptability.
Data Source
Figure 1
Figure 2A~2C
Figure 2D~2F
AI summary
A data processing system includes a central processing unit (CPU) (107,109) and ac-celerator cards coupled to the CPU (107,109) over a bus, each of the accelerator card-s having a plurality of data processing (DP) accelerators to receive DP tasks from the CPU (107,109) and to perform the received DP tasks. At least two of the accelerator cards are coupled to each other via an inter-card connection, and at least two of the DP accelerators are coupled to each other via an inter-chip connection. Each of the inter-card connection and the inter-chip connection is capable of being dynamically activated or deactivated, such that in response to a request received from the CPU (107,109), any one of the accelerator cards or any one of the DP accelerators within any one of the accelerator cards can be enabled or disabled to process any one of the DP tasks received from the CPU (107,109).