Transformer Mapping to Analog Compute-in-Memory via MLP Approximation
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Analog compute-in-memory (ACIM) hardware faces challenges in supporting complex neural network architectures like transformer models due to non-native operations such as layer normalization, softmax functions, and GELU activation functions, leading to computational bottlenecks and reduced efficiency.
Innovation Solution
Implementing multi-layer perceptrons to approximate non-vector-matrix multiplication operations, using crossbar arrays to store weight values as analog quantities, and integrating shift, shift-scale, and dense neural networks to decompose complex functions into linear transformations suitable for ACIM processing.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Use of energy by moving object
If ACIM hardware uses traditional digital processors (CPU, GPU, TPU) to execute transformer models, then computational flexibility and support for non-native operations are maintained, but energy consumption increases and computational efficiency decreases
Solution Approach 1:
The patent transforms non-native transformer operations into native ACIM operations by changing the computational parameters. Multi-layer perceptrons are trained to approximate layer normalization, softmax, and GELU operations, converting them into forms compatible with analog compute-in-memory hardware's native vector-matrix multiplication capabilities, thereby reducing energy consumption while maintaining functional equivalence
Solution Approach 2:
The patent replaces digital processing mechanisms with analog computing mechanisms. By substituting digital processor executions with analog crossbar array computations, the system eliminates the energy-intensive data movement between memory and processing units inherent in digital architectures, achieving lower energy consumption for the same computational tasks
2Productivity
If ACIM hardware implements heterogeneous architecture with specialized digital processing units, then non-native operations can be executed, but computational bottlenecks increase and overall efficiency decreases
Solution Approach 1:
The patent makes the ACIM hardware universally capable of executing both native and non-native operations through a single unified architecture. By training multi-layer perceptrons to approximate various transformer operations (layer normalization, softmax, GELU), the system eliminates the need for separate specialized digital processing units, achieving high computational throughput without heterogeneous complexity
Solution Approach 2:
The patent merges the functions of multiple specialized processing units into a single analog compute-in-memory array. By combining layer normalization, softmax, and activation function operations into unified multi-layer perceptron approximations that execute on the same crossbar array, the system reduces device complexity while maintaining computational throughput
3Measurement precision
If ACIM hardware uses complex analog implementations for non-native operations, then operational accuracy is maintained, but hardware complexity and design difficulty increase
Solution Approach 1:
The patent creates simplified analog copies of complex digital operations. By training multi-layer perceptrons to approximate transformer operations and implementing them through relatively simple analog circuits on crossbar arrays, the system maintains operational accuracy without requiring complex analog implementations, thereby reducing hardware complexity and design difficulty
Applied Scientific Principles
This section explains which scientific principles are used to turn an abstract innovation direction into a practical engineering solution.
Function Achieved in This Case
Enables efficient execution of transformer models in ACIM hardware by approximating non-native operations, enhancing computational capacity and maintaining compatibility with analog compute-in-memory constraints.
Implementation Method 1
ACIM systems leverage the physical properties of memory devices, such as resistive random-access memory (RRAM) or non-volatile capacitors, to store synaptic weights and perform vector-matrix multiplications through analog operations
Implementation Method 2
The analog compute-in-memory architecture may comprise crossbar arrays of memory elements that store weight values as analog quantities using conductance or capacitance properties
Data Source
AI summary
The present disclosure provides a method for implementing transformer models in analog compute-in-memory hardware. The method comprises training a target neural network using one or more operators on one or more graphics processing units, generating one or more datasets from full network traces to capture input-output relationships of non-vector-matrix multiplication operations, training one or more multi-layer perceptrons to approximate the non-vector-matrix multiplication operations using the one or more datasets, replacing the original non-vector-matrix multiplication operations with the trained one or more multi-layer perceptrons, and mapping the resulting multi-layer perceptron-only neural network to an analog compute-in-memory architecture. The non-vector-matrix multiplication operations comprise layer normalization operations, softmax operations, and GELU activation operations. The analog compute-in-memory architecture comprises crossbar arrays of memory elements that store weight values as analog quantities using conductance or capacitance properties.


