Differential MAC Unit on a Shared Charge Bus for Low-Power Scaling

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing multiplier-accumulator architectures for machine learning applications face challenges in scalability and power consumption due to synchronous operation and increased gate complexity, particularly in performing multiply-accumulate operations for large matrices, which results in high power dissipation and inefficiency.

Innovation Solution

A scalable asynchronous multiplier-accumulator architecture utilizing a common charge transfer bus for multiplier-accumulator, bias, and analog-to-digital converter unit elements, with NAND-groups and charge transfer capacitors to minimize displacement currents and power consumption by enabling asynchronous operation and shared binary weighted charge transfer lines.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Power

If synchronous clocked stages are used for multiplier operation, then operation timing is controlled, but power dissipation increases

Engineering Contradiction:
Improvepower dissipationVSAvoidclocked stage complexity
Core Design Contradiction:
PowerVSDevice complexity

Solution Approach 1:

The patent transitions from static synchronous clocked stages to dynamic asynchronous operation where the system adapts its timing based on actual computation completion. The multiplier-accumulator uses dynamic timing signals that are generated based on when computations are actually complete, rather than being driven by a fixed global clock, thereby reducing unnecessary switching activity and power dissipation.

Inventive Principle:
Principle #15Dynamics

Solution Approach 2:

The patent extracts and removes the global clocking mechanism from the multiplier-accumulator architecture. By eliminating the synchronous clocked stages and their associated clock distribution networks, the design reduces power dissipation while maintaining proper operation timing through alternative asynchronous signaling methods.

Inventive Principle:
Principle #2Taking out (Extraction)

2Productivity

If large n×n multipliers are used for machine learning operations, then computational capability increases, but gate complexity increases as n2

Engineering Contradiction:
Improvecomputational capabilityVSAvoidgate complexity
Core Design Contradiction:
ProductivityVSDevice complexity

Solution Approach 1:

The patent segments the large n×n multiplication task into multiple smaller multiplier-accumulator units that operate in parallel. Each unit handles a portion of the computation, and the results are accumulated through shared charge transfer buses. This segmentation reduces the gate complexity of individual units while maintaining the overall computational capability through parallel processing.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent designs universal multiplier-accumulator units that can be configured and cascaded to handle different matrix sizes and operations. These multi-functional units can perform multiplication, accumulation, and can be arranged in various configurations to support different computational requirements, thereby reducing the need for dedicated complex circuits for each specific operation size.

Inventive Principle:
Principle #6Universality (Multi-functionality)

3Adaptability or versatility

If multiple adders are added for multiply-accumulate operations, then computational functionality is enhanced, but power dissipation and complexity increase

Engineering Contradiction:
Improvecomputational functionalityVSAvoidpower dissipation
Core Design Contradiction:
Adaptability or versatilityVSLoss of energy

Solution Approach 1:

The patent merges the multiplication and accumulation functions into unified multiplier-accumulator units that share common hardware resources. The accumulation function is integrated with the multiplication logic, allowing the same circuit elements to perform both operations sequentially or in parallel, thereby reducing the total number of separate adders and associated power dissipation.

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The patent replaces traditional ripple-carry adder chains with charge-based accumulation mechanisms. Instead of using multiple sequential adders that require extensive carry propagation logic, the design uses charge transfer and summation at node points, which reduces the logical complexity and power consumption of the accumulation function while maintaining the same computational capability.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

4Loss of energy

If asynchronous operation is implemented to reduce power consumption, then power dissipation decreases, but timing control becomes more challenging

Engineering Contradiction:
Improvepower dissipationVSAvoidtiming control
Core Design Contradiction:
Loss of energyVSEase of operation

Solution Approach 1:

The patent implements feedback mechanisms where completion signals from computational units are fed back to control the timing of subsequent operations. The system uses carry-propagation detection and accumulation completion signals to dynamically control the timing of data transfer and processing stages, ensuring proper synchronization without requiring a global clock. This feedback-based timing control maintains ease of operation while enabling asynchronous low-power operation.

Inventive Principle:
Principle #23Feedback

Applied Scientific Principles

This section explains which scientific principles are used to turn an abstract innovation direction into a practical engineering solution.

Function Achieved in This Case

This architecture reduces power consumption by minimizing displacement currents and allows for scalable and efficient multiply-accumulate operations, particularly when the kernel is static, while maintaining accuracy and flexibility in handling various matrix sizes.

Implementation Method 1

each NAND gate having a positive output coupled through a binary weighted positive charge transfer capacitor to a unique positive charge transfer line and a binary weighted negative output coupled through a negative charge transfer capacitor to a unique negative charge transfer line

Methodology Applied
Scientific EffectCharge transfer: Capacitance

Implementation Method 2

each MAC UE providing a result as a charge transferred to a differential charge transfer bus

Methodology Applied
Scientific EffectDifferential charge transfer: Capacitance

Implementation Method 3

a second plurality of Bias unit elements (Bias UEs) performing a bias operation and placing a bias value as a charge onto the differential charge transfer bus

Methodology Applied
Scientific EffectCharge transfer: Capacitance

Implementation Method 4

a third plurality of ADC unit elements (ADC UEs) operative to convert a charge present on the differential charge transfer bus into a digital output value

Methodology Applied
Scientific EffectCharge-to-digital conversion: Capacitance

Data Source

PatentUS12026479B2Differential unit element for multiply-accumulate operations on a shared charge transfer bus
Publication Date: 2024.07.02 CEREMORPHIC INC
  • US12026479B2 patent drawing
  • US12026479B2 patent drawing
  • US12026479B2 patent drawing

AI summary

A Unit Element (UE) has a digital X input and a digital W input, and comprises groups of NAND gates generating complementary outputs which are coupled to differential charge transfer lines through respective charge transfer capacitor Cu. The number of bits in the X input determines the number of NAND gates in a NAND-group and the number of bits in the W input determines the number of NAND groups. Each NAND-group receives one bit of the W input applied to all of the NAND gates of the NAND-group, and each unit element having the bits of X applied to each associated NAND gate input of each unit element. The NAND gate outputs are coupled through a charge transfer capacitor Cu to charge transfer lines. Multiple Unit Elements may be placed in parallel to sum and scale the charges from the charge transfer lines, the charges coupled to an analog to digital converter which forms the dot product output.