Transformer Mapping to Analog Compute-in-Memory via MLP Approximation

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Analog compute-in-memory (ACIM) hardware faces challenges in supporting complex neural network architectures like transformer models due to non-native operations such as layer normalization, softmax functions, and GELU activation functions, leading to computational bottlenecks and reduced efficiency.

Innovation Solution

Implementing multi-layer perceptrons to approximate non-vector-matrix multiplication operations, using crossbar arrays to store weight values as analog quantities, and integrating shift, shift-scale, and dense neural networks to decompose complex functions into linear transformations suitable for ACIM processing.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Use of energy by moving object

If ACIM hardware uses traditional digital processors (CPU, GPU, TPU) to execute transformer models, then computational flexibility and support for non-native operations are maintained, but energy consumption increases and computational efficiency decreases

Engineering Contradiction:
Improveenergy consumptionVSAvoidsupport for non-native operations
Core Design Contradiction:
Use of energy by moving objectVSAdaptability or versatility

Solution Approach 1:

The patent transforms non-native transformer operations into native ACIM operations by changing the computational parameters. Multi-layer perceptrons are trained to approximate layer normalization, softmax, and GELU operations, converting them into forms compatible with analog compute-in-memory hardware's native vector-matrix multiplication capabilities, thereby reducing energy consumption while maintaining functional equivalence

Inventive Principle:
Principle #35Parameter changes

Solution Approach 2:

The patent replaces digital processing mechanisms with analog computing mechanisms. By substituting digital processor executions with analog crossbar array computations, the system eliminates the energy-intensive data movement between memory and processing units inherent in digital architectures, achieving lower energy consumption for the same computational tasks

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

2Productivity

If ACIM hardware implements heterogeneous architecture with specialized digital processing units, then non-native operations can be executed, but computational bottlenecks increase and overall efficiency decreases

Engineering Contradiction:
Improvecomputational throughputVSAvoidheterogeneous architecture
Core Design Contradiction:
ProductivityVSDevice complexity

Solution Approach 1:

The patent makes the ACIM hardware universally capable of executing both native and non-native operations through a single unified architecture. By training multi-layer perceptrons to approximate various transformer operations (layer normalization, softmax, GELU), the system eliminates the need for separate specialized digital processing units, achieving high computational throughput without heterogeneous complexity

Inventive Principle:
Principle #6Universality (Multi-functionality)

Solution Approach 2:

The patent merges the functions of multiple specialized processing units into a single analog compute-in-memory array. By combining layer normalization, softmax, and activation function operations into unified multi-layer perceptron approximations that execute on the same crossbar array, the system reduces device complexity while maintaining computational throughput

Inventive Principle:
Principle #5Merging (Combining)

3Measurement precision

If ACIM hardware uses complex analog implementations for non-native operations, then operational accuracy is maintained, but hardware complexity and design difficulty increase

Engineering Contradiction:
Improveoperational accuracyVSAvoidanalog implementation complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent creates simplified analog copies of complex digital operations. By training multi-layer perceptrons to approximate transformer operations and implementing them through relatively simple analog circuits on crossbar arrays, the system maintains operational accuracy without requiring complex analog implementations, thereby reducing hardware complexity and design difficulty

Inventive Principle:
Principle #26Copying

Applied Scientific Principles

This section explains which scientific principles are used to turn an abstract innovation direction into a practical engineering solution.

Function Achieved in This Case

Enables efficient execution of transformer models in ACIM hardware by approximating non-native operations, enhancing computational capacity and maintaining compatibility with analog compute-in-memory constraints.

Implementation Method 1

ACIM systems leverage the physical properties of memory devices, such as resistive random-access memory (RRAM) or non-volatile capacitors, to store synaptic weights and perform vector-matrix multiplications through analog operations

Methodology Applied
Scientific EffectConductance: Conduction (electrical)

Implementation Method 2

The analog compute-in-memory architecture may comprise crossbar arrays of memory elements that store weight values as analog quantities using conductance or capacitance properties

Methodology Applied
Scientific EffectCapacitance: Capacitance

Data Source

PatentUS20260065046A1Techniques to support transformer models in analog compute-in-memory hardware
Publication Date: 2026.03.05 GEORGIA TECH RES CORP
  • US20260065046A1 patent drawing
  • US20260065046A1 patent drawing
  • US20260065046A1 patent drawing

AI summary

The present disclosure provides a method for implementing transformer models in analog compute-in-memory hardware. The method comprises training a target neural network using one or more operators on one or more graphics processing units, generating one or more datasets from full network traces to capture input-output relationships of non-vector-matrix multiplication operations, training one or more multi-layer perceptrons to approximate the non-vector-matrix multiplication operations using the one or more datasets, replacing the original non-vector-matrix multiplication operations with the trained one or more multi-layer perceptrons, and mapping the resulting multi-layer perceptron-only neural network to an analog compute-in-memory architecture. The non-vector-matrix multiplication operations comprise layer normalization operations, softmax operations, and GELU activation operations. The analog compute-in-memory architecture comprises crossbar arrays of memory elements that store weight values as analog quantities using conductance or capacitance properties.