Shared Charge-Bus Multiplier-Accumulator for Low-Power Scaling

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing multiplier-accumulator architectures for machine learning applications face challenges in power consumption due to synchronous clocked operations and increased gate complexity, particularly in scalable hardware designs for multiply-accumulate operations.

Innovation Solution

An asynchronous multiplier-accumulator architecture utilizing AND gates and charge transfer capacitors to minimize internal state changes, with a charge summing unit and analog-to-digital converter to generate digital outputs, allowing for cascaded operations and reduced power consumption.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If synchronous clocked stages are used for multiplication operations, then operational reliability is improved, but power consumption increases

Engineering Contradiction:
Improveoperational reliabilityVSAvoidpower consumption
Core Design Contradiction:
ReliabilityVSUse of energy by moving object

Solution Approach 1:

The patent inverts the conventional synchronous clocked approach by implementing an asynchronous multiplication architecture. Instead of using clock signals to coordinate operations, the system uses hand-shaking signals and ready/valid flags to indicate when operations are complete, thereby eliminating continuous clocking and reducing power consumption while maintaining operational reliability

Inventive Principle:
Principle #13The other way round (Inversion)

Solution Approach 2:

The patent employs periodic action through the use of reset signals that periodically clear the accumulation register and through the periodic hand-shaking protocol between multiplier and accumulator stages. This periodic resetting and signal exchange ensures reliable operation without requiring continuous clocking, thus reducing power consumption

Inventive Principle:
Principle #19Periodic action

2Productivity

If n×n multiplier size is increased for machine learning applications, then computational capability is improved, but gate complexity increases as n2

Engineering Contradiction:
Improvecomputational capabilityVSAvoidgate complexity
Core Design Contradiction:
ProductivityVSDevice complexity

Solution Approach 1:

The patent segments large n×n multiplication operations into smaller m×m unit element blocks. Each block performs local multiplication and accumulation, and the results are combined to produce the final output. This segmentation reduces the gate complexity of individual units while maintaining the overall computational capability through parallel block operations

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent implements a nested structure where multiple m×m unit element blocks are nested within a larger n×n multiplier-accumulator array. The unit elements are arranged in a grid pattern where blocks are interconnected through shared buses and accumulation paths, allowing hierarchical computation that scales efficiently with increasing n

Inventive Principle:
Principle #7Nested doll (Nesting)

3Adaptability or versatility

If multiple adders are added for multiply-accumulate operations, then computational functionality is improved, but device complexity increases

Engineering Contradiction:
Improvecomputational functionalityVSAvoiddevice complexity
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The patent merges the multiplication and accumulation functions into a single integrated unit element block. The multiply-accumulate operation is performed by combining the AND gate multiplication stage with an accumulation register and addition logic within the same block, eliminating the need for separate adder circuits and reducing overall device complexity while maintaining full computational functionality

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The unit element block is designed as a universal module that can perform multiple functions: multiplication via AND gates, accumulation via the accumulation register, and output via the ready/valid flag system. This multi-functional design eliminates the need for dedicated separate circuits for each operation, reducing device complexity while maintaining adaptability

Inventive Principle:
Principle #6Universality (Multi-functionality)

4Use of energy by moving object

If internal state changes are minimized for static weighting matrices, then power consumption is reduced, but operational flexibility is constrained

Engineering Contradiction:
Improvepower consumptionVSAvoidoperational flexibility
Core Design Contradiction:
Use of energy by moving objectVSAdaptability or versatility

Solution Approach 1:

The patent implements a dynamic architecture where the multiplication and accumulation operations are triggered by hand-shaking signals and ready/valid flags rather than continuous clocking. The system dynamically activates computation only when input data is ready and the accumulator is prepared, minimizing unnecessary state changes and power consumption while maintaining operational flexibility through signal-based control

Inventive Principle:
Principle #15Dynamics

Applied Scientific Principles

This section explains which scientific principles are used to turn an abstract innovation direction into a practical engineering solution.

Function Achieved in This Case

The solution enables low-power, scalable, and efficient multiply-accumulate operations by minimizing internal state changes and power dissipation, suitable for machine learning applications with static weighting matrices, while maintaining flexibility and accuracy in digital output generation.

Implementation Method 1

each AND gate output coupled through a capacitor of value Cu to a particular analog charge line of an analog charge bus according to bit order

Methodology Applied
Scientific EffectCharge transfer: Capacitance

Implementation Method 2

the charge summing capacitors having a second terminal which are coupled together and also coupled to the input of an analog to digital converter

Methodology Applied
Scientific EffectCharge summing: Capacitance

Data Source

PatentUS12014151B2Scaleable analog multiplier-accumulator with shared result bus
Publication Date: 2024.06.18 CEREMORPHIC INC
  • US12014151B2 patent drawing
  • US12014151B2 patent drawing
  • US12014151B2 patent drawing

AI summary

A plurality of unit elements share a charge transfer bus, each unit element accepts A and B digital inputs and generates a product P as an analog charge transferred to the charge transfer bus, each unit element comprised of groups of AND gates coupled to charge transfer lines through a capacitor Cu. Each unit element receives one bit of the B input applied to all of the AND gates of the unit element, and each unit element having the bits of A applied to each associated AND gate input of each unit element. The AND gates of each unit element are coupled to charge transfer lines through a capacitor Cu, and the charge transfer lines couple to binary weighted charge summing capacitors which sum and scale the charges contributed by all unit elements to the charge transfer lines according to a bit weight and converted to a digital value output.