Analog MAC Architecture With Charge Transfer Capacitors for Low Power

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing multiplier-accumulator architectures for machine learning applications face challenges in power consumption due to synchronous clocked stages and increased gate complexity, particularly in forming dot products for large matrices, which results in high power dissipation and inefficiency.

Innovation Solution

A scalable asynchronous multiplier-accumulator architecture with a common charge transfer bus for MAC, Bias, and ADC unit elements, utilizing NAND-groups and binary weighted charge transfer capacitors to minimize displacement currents and power consumption, and incorporating a Successive Approximation Register (SAR) controller for efficient analog-to-digital conversion.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Ease of operation

If synchronous clocked stages are used in multiplier-accumulator architecture, then operation control is simplified, but power dissipation increases

Engineering Contradiction:
Improveoperation controlVSAvoidpower dissipation
Core Design Contradiction:
Ease of operationVSUse of energy by moving object

Solution Approach 1:

The patent employs periodic charge transfer operations where capacitors transfer charges at specific phases of computation. The computation proceeds through periodic phases: accumulation phase where charges are transferred to summing nodes, and reset phase where nodes are prepared for next computation. This periodic action eliminates continuous clocking while maintaining controlled operation.

Inventive Principle:
Principle #19Periodic action

Solution Approach 2:

The patent replaces the mechanical clocked switching system with an analog charge-based computation system. Instead of using clocked digital switches to control data flow, the system uses charge transfer through capacitors and current mirrors, where the computation is driven by voltage levels and charge accumulation rather than clock edges.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

2Productivity

If large n×n multipliers are implemented for machine learning applications, then computational capability increases, but gate complexity increases as n2

Engineering Contradiction:
Improvecomputational capabilityVSAvoidgate complexity
Core Design Contradiction:
ProductivityVSDevice complexity

Solution Approach 1:

The patent extracts the multiplication and accumulation functions into separate dedicated circuit blocks. The multiplier section uses capacitor arrays to compute partial products, while the accumulator section separately sums these products. This extraction allows each section to be optimized independently and enables modular scaling for different matrix sizes without requiring complete redesign.

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The patent segments the large n×n multiplication into multiple smaller computational units. Each unit handles a subset of the computation using binary-weighted capacitor arrays, and multiple such units are combined to achieve the full n×n functionality. This segmentation reduces the gate complexity of individual units while maintaining overall computational capability.

Inventive Principle:
Principle #1Segmentation

3Adaptability or versatility

If additional adders are added for multiply-accumulate operations, then computational functionality improves, but device complexity increases

Engineering Contradiction:
Improvecomputational functionalityVSAvoiddevice complexity
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The patent merges the multiplication and accumulation operations into a unified charge transfer mechanism. The same capacitor arrays that perform multiplication also perform accumulation by transferring charges to shared summing nodes. This merging eliminates the need for separate adder circuits for each multiply-accumulate operation, reducing device complexity while maintaining full MAC functionality.

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The patent designs universal computational units that can perform multiple functions. The same hardware block can function as a multiplier for different matrix dimensions, an accumulator for different accumulation depths, and can be configured for different precision requirements. This universality allows the system to handle various ML workloads without requiring dedicated hardware for each specific operation.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Applied Scientific Principles

This section explains which scientific principles are used to turn an abstract innovation direction into a practical engineering solution.

Function Achieved in This Case

The architecture achieves reduced power consumption by minimizing displacement currents and enabling efficient dot product operations, while allowing for scalability and flexible configuration, thereby improving the performance and efficiency of machine learning operations.

Implementation Method 1

each NAND gate having a positive output coupled through a binary weighted positive charge transfer capacitor to a positive charge transfer line and a negative output coupled through a binary weighted negative charge transfer capacitor to a negative charge transfer line

Methodology Applied
Scientific EffectCharge transfer: Electrostatics

Implementation Method 2

This power savings can be realized by an architecture which minimizes displacement currents when the kernel (coefficient matrix W) is mostly static as is commonly the case in ML applications

Methodology Applied
Scientific EffectDisplacement current minimization: Electrical Resistance

Implementation Method 3

a third plurality of ADC unit elements (ADC UEs) operative to convert a charge present on the differential charge transfer lines into a digital output value

Methodology Applied
Scientific EffectAnalog to digital conversion:

Data Source

PatentUS11689213B2Architecture for analog multiplier-accumulator with binary weighted charge transfer capacitors
Publication Date: 2023.06.27 CEREMORPHIC INC
  • US11689213B2 patent drawing
  • US11689213B2 patent drawing
  • US11689213B2 patent drawing

AI summary

An architecture for a multiplier-accumulator (MAC) uses a common Unit Element (UE) for each aspect of operation, the MAC formed as a plurality of MAC UEs, a plurality of Bias UEs, and a plurality of Analog to Digital Conversion (ADC) UEs which collectively perform a scalable MAC operation and generate a binary result. Each MAC UE, BIAS UE and ADC UE comprises groups of NAND gates with complementary outputs arranged in NAND-groups, each NAND gate coupled to a differential charge transfer bus through a binary weighted charge transfer capacitor to form an analog multiplication product as a charge applied to the differential charge transfer bus. The analog charge transfer bus is coupled to groups of ADC UEs with an ADC controller which enables and disables the ADC UEs using successive approximation to determine the accumulated multiplication result.