AI Accelerator Switch Linking for Transformer Workloads
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Transformer-based neural network models require significant computational resources and time to process, leading to inefficiencies and high energy consumption, especially as model sizes increase to billions or trillions of parameters.
Innovation Solution
The development of AI accelerator apparatuses and chiplet devices with in-memory compute (IMC) capabilities, which include digital IMC (DIMC) devices for matrix computations and SIMD devices for non-matrix computations, along with a modular architecture that allows for scalable and efficient processing of transformer workloads.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If transformer model size increases to handle more complex NLP tasks, then model capability improves, but computational resource requirements and processing time increase significantly
Solution Approach 1:
The system segments the transformer model processing into multiple AI processing systems, each handling specific portions of the computational workload. The model is divided across multiple chiplet devices within each AI processing system, allowing parallel processing of different model layers or operations, thereby reducing overall processing time while maintaining full model capability.
Solution Approach 2:
The patent introduces a multi-dimensional architecture by stacking multiple AI processing systems vertically and connecting them via switches in a network fabric. This adds a spatial dimension to the computational architecture, enabling massive parallel processing across distributed systems while maintaining the integrity of the large transformer model through coordinated communication.
2Adaptability or versatility
If transformer model size increases to handle more complex NLP tasks, then model capability improves, but energy consumption increases significantly
Solution Approach 1:
By segmenting the computational workload across multiple AI processing systems and chiplet devices, the system distributes energy consumption across many smaller units rather than concentrating it in a single high-power processor. This segmentation allows for more efficient energy utilization through parallel processing at lower power levels.
Solution Approach 2:
The switch fabric acts as an intermediary that optimizes data flow between AI processing systems, reducing unnecessary data transmission and associated energy consumption. The hierarchical memory structure also serves as an intermediary, caching frequently accessed model parameters to minimize energy-intensive memory accesses.
3Device complexity
If traditional CPU-based processing is used for transformer workloads, then system simplicity is maintained, but computational efficiency and throughput are insufficient
Solution Approach 1:
The system segments the monolithic CPU into multiple specialized AI processing systems, each with dedicated hardware accelerators for transformer operations. This segmentation enables highly efficient parallel processing of attention mechanisms and matrix operations while maintaining manageable complexity through modular design.
Solution Approach 2:
The patent replaces traditional general-purpose CPU mechanics with specialized hardware accelerators and neural network processing units optimized for transformer workloads. This substitution provides dedicated circuitry for matrix multiplications and activation functions, dramatically improving computational efficiency for AI inference tasks.
4Adaptability or versatility
If 1024 GPUs are used to train large models like GPT-3, then model training capability is achieved, but training time remains extremely long (about 4 months)
Solution Approach 1:
The system segments the training workload across multiple AI processing systems working in parallel, with each system handling specific model layers or parameter subsets. This segmentation enables faster convergence by distributing the computational burden, reducing training time from months to shorter durations while maintaining the ability to train trillion-parameter models.
Data Source
AI summary
A server system using switch linking between groups of AI processing systems. The system includes at least a first central processing unit (CPU) coupled to a first group of AI processing systems via a first switch and a second CPU coupled to a second group of AI processing systems via a second switch. Each of the CPUs is also coupled to a separate group of memory devices and a communication link is configured between the first switch and the second switch to communicate information between the first group of AI processing systems and the second group of AI processing systems. Each of these AI processing system groups include a plurality of AI processing modules, and each of these modules include a plurality of chiplet devices. Each chiplet device is configured with a plurality of in-memory compute (IMC) devices for processing neural network model workloads.


