In-Memory Compute Chiplet for Transformer Workloads
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Transformer-based neural network models, particularly in natural language processing, face significant challenges due to high computational intensity and memory requirements, making it inefficient to serve large-scale NLP models effectively.
Innovation Solution
The implementation of AI accelerator apparatuses using chiplet devices with digital in-memory compute (DIMC) functionality, which integrates computational functions and memory fabric, and includes SIMD devices to accelerate attention functions and softmax computations, along with die-to-die interconnects and PCIe buses for efficient communication, enabling scalable and efficient processing of transformer workloads.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Productivity
If traditional computing architectures are used to process transformer workloads, then general-purpose computing capability is maintained, but computational performance is insufficient and power consumption is high
Solution Approach 1:
The system is divided into multiple chiplet devices, each containing specialized DIMC engines and memory fabric. These chiplets can be independently configured and scaled to match specific transformer workload requirements, enabling efficient parallel processing while maintaining modular power management
Solution Approach 2:
The patent replaces traditional von Neumann architecture with an in-memory compute architecture where computation is performed directly within the memory fabric. This eliminates the need for continuous data movement between separate CPU and memory units, dramatically reducing energy consumption while increasing computational throughput for transformer operations
2Measurement precision
If model size and compute requirements are increased to improve NLP performance, then inference accuracy is improved, but training time and computational resource requirements increase significantly
Solution Approach 1:
The system performs preliminary configuration of chiplet devices to match specific transformer model architectures. The modular chiplet design allows pre-optimization of compute and memory resources for particular model sizes and types, enabling rapid deployment of large-scale models without requiring linear increases in training time
Solution Approach 2:
Multiple chiplet devices are combined and interconnected through high-speed interfaces to form a unified computing system. This parallel architecture enables large transformer models to be processed across multiple devices simultaneously, reducing training time while maintaining the ability to handle trillion-parameter models
3Productivity
If computational resources are increased to serve large-scale NLP models, then model serving capability is improved, but device complexity and scalability challenges increase
Solution Approach 1:
The system uses segmented chiplet devices that can be independently configured for specific transformer workload types. Each chiplet contains self-contained DIMC engines and memory fabric, allowing the system to scale by simply adding identical or heterogeneous chiplets rather than redesigning the entire system architecture
Solution Approach 2:
The chiplet devices are designed with universal interfaces and configurable DIMC engines that can handle various transformer operations including attention mechanisms, feedforward networks, and normalization layers. This multi-functionality allows a single chiplet design to serve multiple model types and sizes, reducing overall system complexity
Data Source
AI summary
An AI accelerator apparatus using in-memory compute chiplet devices. The apparatus includes one or more chiplets, each of which includes a plurality of tiles. Each tile includes a plurality of slices, a central processing unit (CPU), and a hardware dispatch device. Each slice can include a digital in-memory compute (DIMC) device configured to perform high throughput computations. In particular, the DIMC device can be configured to accelerate the computations of attention functions for transformer-based models (a.k.a. transformers) applied to machine learning applications, including generative AI. A single input multiple data (SIMD) device configured to further process the DIMC output and compute softmax functions for the attention functions. The chiplet can also include die-to-die (D2D) interconnects, a peripheral component interconnect express (PCIe) bus, a dynamic random access memory (DRAM) interface, and a global CPU interface to facilitate communication between the chiplets, memory and a server or host system.


