In-Memory Compute Chiplets for Transformer Attention Acceleration
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Transformer-based neural network models require significant computational resources and time to process large datasets, making it challenging to efficiently serve NLP models at scale.
Innovation Solution
The use of AI accelerator apparatuses with in-memory compute chiplet devices, which include digital in-memory compute (DIMC) devices and single input multiple data (SIMD) devices, to accelerate transformer computations by integrating computational functions and memory fabric.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Productivity
If traditional computing architectures are used to process transformer workloads, then model training and inference can be performed, but computational performance is insufficient and time consumption is excessive
Solution Approach 1:
The system is divided into multiple chiplet devices, each containing specialized digital in-memory compute units. These chiplets can be independently configured and scaled to match specific transformer workload requirements, enabling parallel processing of different model layers or batches, thereby significantly improving computational throughput and reducing training/inference time
Solution Approach 2:
A high-speed interconnect fabric acts as an intermediary between the digital in-memory compute units and external memory systems. This intermediary enables efficient data movement and communication, reducing bottlenecks and ensuring that computational units can operate at full speed without waiting for data, thus improving overall productivity while maintaining scalable architecture
2Measurement precision
If model size and compute requirements are increased to improve NLP model performance, then inference accuracy is improved, but device complexity and resource requirements increase
Solution Approach 1:
The system segments the computational workload across multiple chiplet devices, each handling specific portions of the transformer model. This segmentation allows independent optimization of each chiplet for specific model layers or functions, managing device complexity through modular design while supporting large-scale models that require high inference accuracy
Solution Approach 2:
The patent transitions from traditional von Neumann architecture to an in-memory compute architecture, adding a new dimension of computation within the memory fabric. This dimensional change enables parallel processing operations that were not feasible in traditional architectures, allowing the system to handle larger model sizes and more complex computations without proportionally increasing device complexity
3Speed
If more computational resources are allocated to accelerate transformer workloads, then processing speed is improved, but power consumption increases
Solution Approach 1:
The patent merges computational logic and memory storage into a unified in-memory compute architecture. By combining these functions within the same physical substrate, the system eliminates data movement between separate CPU and memory components, reducing energy consumption while maintaining high processing speed through localized computation within the memory fabric
Solution Approach 2:
The system replaces traditional mechanical/electrical data movement mechanisms with field-based in-memory computation. Instead of physically moving data between memory and processing units, computations are performed directly within the memory array using electrical field interactions, significantly reducing energy consumption while maintaining or improving processing speed
4Adaptability or versatility
If NLP models are scaled up to handle larger datasets and more complex tasks, then model capability is improved, but training time and computational resource requirements increase
Solution Approach 1:
The system segments large-scale NLP model training workloads across multiple chiplet devices, each capable of independently processing specific model layers or data batches. This segmentation enables parallel training operations that significantly reduce training time while maintaining the ability to handle large, complex models with high adaptability and versatility
Solution Approach 2:
The in-memory compute chiplet architecture provides a universal computing platform that can be configured to handle various NLP model architectures and workloads. The same hardware infrastructure can efficiently train different model sizes and types, providing multi-functionality that reduces training time across diverse applications without requiring specialized hardware for each model type
Data Source
AI summary
An AI accelerator apparatus using in-memory compute chiplet devices. The apparatus includes one or more chiplets, each of which includes a plurality of tiles. Each tile includes a plurality of slices, a central processing unit (CPU), and a hardware dispatch device. Each slice can include a digital in-memory compute (DIMC) device configured to perform high throughput computations. In particular, the DIMC device can be configured to accelerate the computations of attention functions for transformer-based models (a.k.a. transformers) applied to machine learning applications. A single input multiple data (SIMD) device configured to further process the DIMC output and compute softmax functions for the attention functions. The chiplet can also include die-to-die (D2D) interconnects, a peripheral component interconnect express (PCIe) bus, a dynamic random access memory (DRAM) interface, and a global CPU interface to facilitate communication between the chiplets, memory and a server or host system.


