Memory Crossbar MAC Using 3-Phase Analog Bit Partitioning

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing matrix-vector multiplication (MVM) operations in high-performance computing face challenges due to their recurrence, universality, matrix size, and memory requirements, leading to inefficiencies in traditional von Neumann architectures and high energy consumption.

Innovation Solution

A crossbar array structure with interleaved switched-capacitor analogue multipliers and adders operates using a 3-phase clocking scheme and a specific bit partition for partial multiplications and additions in the analogue domain, reducing analogue compute signal-to-noise ratio requirements while maintaining pipeline behavior.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Ease of manufacture

If traditional von Neumann architecture is used for MVM operations, then data storage and processing are separated, but data transfer congestion and power consumption increase

Engineering Contradiction:
Improvearchitectural simplicityVSAvoidpower consumption
Core Design Contradiction:
Ease of manufactureVSUse of energy by moving object

Solution Approach 1:

The patent merges memory storage and arithmetic processing into a single integrated crossbar array structure. Each memory cell contains both the weight storage and the multiply-accumulate unit, eliminating the need for separate data transfer paths between memory and processing units. This integration directly reduces power consumption by eliminating continuous data transfer operations while maintaining architectural functionality.

Inventive Principle:
Principle #5Merging (Combining)

2Productivity

If dedicated hardware acceleration devices with crossbar arrays are used, then MVM operations are accelerated, but computational precision requirements increase

Engineering Contradiction:
ImproveMVM operation speedVSAvoidcomputational precision
Core Design Contradiction:
ProductivityVSMeasurement precision

Solution Approach 1:

The patent segments the computational precision into multiple stages through bit-partitioning of the input vector and weight matrix. Instead of requiring high precision throughout the entire computation, the system divides the multiplication into multiple lower-precision stages that are accumulated progressively. This segmentation allows the crossbar array to operate at high speed while maintaining acceptable precision through the accumulation of partial products.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent dynamically adjusts the bit-partition parameters to optimize between speed and precision. By changing the partitioning granularity and the number of accumulation stages, the system can adapt to different computational requirements. This parameter adjustment allows the same hardware architecture to serve both high-speed inference applications and higher-precision training applications.

Inventive Principle:
Principle #35Parameter changes

3Productivity

If multi-bit analogue operations are performed in a single step, then computational throughput increases, but analogue compute signal-to-noise ratio requirements increase

Engineering Contradiction:
Improvecomputational throughputVSAvoidsignal-to-noise ratio
Core Design Contradiction:
ProductivityVSMeasurement precision

Solution Approach 1:

The patent segments multi-bit analogue operations into multiple single-bit or low-bit operations performed sequentially through the 3-phase clocking scheme. Each phase processes a subset of bits, and the results are accumulated in subsequent phases. This segmentation reduces the signal-to-noise ratio requirements for each individual operation while maintaining high overall throughput through pipelined execution across multiple phases.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent implements continuous useful action through the pipelined 3-phase clocking scheme, where operations are performed continuously across multiple phases without idle time. The accumulation register continuously integrates the results from each phase, ensuring that the useful computational action never stops. This continuous operation maintains high throughput while distributing the precision requirements across multiple smaller operations rather than requiring high precision in a single step.

Inventive Principle:
Principle #20Continuity of useful action

Data Source

PatentUS20260099298A9Multi-bit analog multiply-accumulate operations with memory crossbar arrays
Publication Date: 2026.04.09 AXELERA AI BV
  • US20260099298A9 patent drawing
  • US20260099298A9 patent drawing
  • US20260099298A9 patent drawing

AI summary

The invention is notably directed to a method of processing data. The method relies on a memory device having a crossbar array structure. The latter includes K×L cells, which interconnect K rows and Z columns. The cells include respective memory systems, which store respective A-bit weights. The memory systems are connected to respective compute units, which are configured as interleaved switched-capacitor analogue multipliers and adders. According to the proposed method, input signals encoding respective M-bit input words are synchronously applied to respective ones of the K rows. The compute units are operated according to a 3-phase clocking scheme, with a view to obtaining MAC results for each of the L columns, where K≥2, L>2, N≥2, and M≥2. Remarkably, the 3-phase clocking scheme is here set to perform n×m partial multiplications, in the analogue domain, according to a specific bit partition, so as to obtain n×m partial output signals in output of each of the compute units. This partition decomposes each of the N-bit weights into n groups of bits and each of the M-bit input words into m groups of bits. Each of the n groups and the m groups includes at least one bit. However, at least one of the n groups and/or the m groups includes at least two bits, whereby N+M>n+m≥3. Moreover, the MAC results are obtained by summing the partial output signals obtained by the compute units for each of the Z columns. The summed output signals are converted into digital signals encoding partial values. The partial values are shifted according to corresponding bit positions, which are set in accordance with the bit partition, and the shifted values are finally added, so as to recompose the desired output vector components. The invention is further directed to related apparatuses and systems.