In-DRAM Transformer Accelerator Using Stochastic Attention Computation
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Transformer neural networks face challenges in computational efficiency and energy consumption due to high data movement and complex computations, with existing PIM architectures optimized for traditional DNNs like CNNs, not adequately addressing transformer-specific characteristics.
Innovation Solution
A PIM system with DRAM tiles and MOMCAPs for stochastic multiplication and analog accumulation, employing a token-based dataflow to compute attention scores efficiently, minimizing data movement and energy consumption.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Productivity
If traditional PIM architectures are used for transformer neural networks, then computational operations can be performed in-memory, but data movement bottlenecks and high energy consumption persist due to inadequate optimization for transformer-specific characteristics
Solution Approach 1:
The patent changes the computational approach from deterministic to stochastic computing, using probabilistic bit streams to represent data and perform multiplications. This parameter change enables more efficient exploitation of memory parallelism and reduces energy consumption while maintaining computational accuracy for transformer operations
Solution Approach 2:
The patent segments the transformer computation into distinct stochastic processing stages (stochastic multiplication, stochastic accumulation, stochastic softmax) that can be efficiently mapped to memory subsystem operations. This segmentation allows each stage to be optimized independently for energy efficiency and speed
2Measurement precision
If complex computations are performed to achieve high accuracy in transformer networks, then computational precision is improved, but data movement and computational complexity increase
Solution Approach 1:
The patent replaces complex deterministic computational mechanics with simpler stochastic processes. Instead of performing complex sequential arithmetic operations, the system uses random bit stream generation and probabilistic counting, which can be executed with simpler hardware logic while achieving the same computational precision
3Productivity
If more computational resources are allocated to handle transformer workloads, then processing capability is improved, but hardware complexity and resource requirements increase
Solution Approach 1:
The patent makes the memory subsystem universally capable of performing multiple functions: data storage, stochastic multiplication, stochastic accumulation, and stochastic softmax computation. This multi-functionality eliminates the need for separate dedicated hardware units for each operation, reducing overall hardware complexity while improving processing capability
Solution Approach 2:
The memory subsystem serves itself by performing computational operations directly on stored data without requiring external computational units. The memory cells and interconnects are utilized for both data movement and computation, making the hardware work harder but not more complex
Applied Scientific Principles
This section explains which scientific principles are used to turn an abstract innovation direction into a practical engineering solution.
Function Achieved in This Case
The system accelerates transformer neural networks by reducing data movement and energy consumption, achieving faster and more precise computations with lower hardware complexity, eliminating the need for external high-bandwidth memory.
Implementation Method 1
a first metal-oxide-metal capacitor (MOMCAP) for accumulating analog values
Data Source
AI summary
The present disclosure provides a processing-in-memory (PIM) system and method for accelerating transformer neural networks. The system comprises a plurality of subarrays, each subarray of the plurality of subarrays including a plurality of DRAM tiles. The subarrays are equipped with bitlines for performing stochastic multiplication operations, metal-oxide-metal capacitors (MOMCAPs) for accumulating analog values, and stochastic-to-analog (S_to_A) circuits for converting stochastic data into analog charge. The system employs a token-based dataflow scheme to efficiently compute attention scores in transformer layers.


