Analog MAC Architecture With Charge Transfer Capacitors for Low Power
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing multiplier-accumulator architectures for machine learning applications face challenges in power consumption due to synchronous clocked stages and increased gate complexity, particularly in forming dot products for large matrices, which results in high power dissipation and inefficiency.
Innovation Solution
A scalable asynchronous multiplier-accumulator architecture with a common charge transfer bus for MAC, Bias, and ADC unit elements, utilizing NAND-groups and binary weighted charge transfer capacitors to minimize displacement currents and power consumption, and incorporating a Successive Approximation Register (SAR) controller for efficient analog-to-digital conversion.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Ease of operation
If synchronous clocked stages are used in multiplier-accumulator architecture, then operation control is simplified, but power dissipation increases
Solution Approach 1:
The patent employs periodic charge transfer operations where capacitors transfer charges at specific phases of computation. The computation proceeds through periodic phases: accumulation phase where charges are transferred to summing nodes, and reset phase where nodes are prepared for next computation. This periodic action eliminates continuous clocking while maintaining controlled operation.
Solution Approach 2:
The patent replaces the mechanical clocked switching system with an analog charge-based computation system. Instead of using clocked digital switches to control data flow, the system uses charge transfer through capacitors and current mirrors, where the computation is driven by voltage levels and charge accumulation rather than clock edges.
2Productivity
If large n×n multipliers are implemented for machine learning applications, then computational capability increases, but gate complexity increases as n2
Solution Approach 1:
The patent extracts the multiplication and accumulation functions into separate dedicated circuit blocks. The multiplier section uses capacitor arrays to compute partial products, while the accumulator section separately sums these products. This extraction allows each section to be optimized independently and enables modular scaling for different matrix sizes without requiring complete redesign.
Solution Approach 2:
The patent segments the large n×n multiplication into multiple smaller computational units. Each unit handles a subset of the computation using binary-weighted capacitor arrays, and multiple such units are combined to achieve the full n×n functionality. This segmentation reduces the gate complexity of individual units while maintaining overall computational capability.
3Adaptability or versatility
If additional adders are added for multiply-accumulate operations, then computational functionality improves, but device complexity increases
Solution Approach 1:
The patent merges the multiplication and accumulation operations into a unified charge transfer mechanism. The same capacitor arrays that perform multiplication also perform accumulation by transferring charges to shared summing nodes. This merging eliminates the need for separate adder circuits for each multiply-accumulate operation, reducing device complexity while maintaining full MAC functionality.
Solution Approach 2:
The patent designs universal computational units that can perform multiple functions. The same hardware block can function as a multiplier for different matrix dimensions, an accumulator for different accumulation depths, and can be configured for different precision requirements. This universality allows the system to handle various ML workloads without requiring dedicated hardware for each specific operation.
Applied Scientific Principles
This section explains which scientific principles are used to turn an abstract innovation direction into a practical engineering solution.
Function Achieved in This Case
The architecture achieves reduced power consumption by minimizing displacement currents and enabling efficient dot product operations, while allowing for scalability and flexible configuration, thereby improving the performance and efficiency of machine learning operations.
Implementation Method 1
each NAND gate having a positive output coupled through a binary weighted positive charge transfer capacitor to a positive charge transfer line and a negative output coupled through a binary weighted negative charge transfer capacitor to a negative charge transfer line
Implementation Method 2
This power savings can be realized by an architecture which minimizes displacement currents when the kernel (coefficient matrix W) is mostly static as is commonly the case in ML applications
Implementation Method 3
a third plurality of ADC unit elements (ADC UEs) operative to convert a charge present on the differential charge transfer lines into a digital output value
Data Source
AI summary
An architecture for a multiplier-accumulator (MAC) uses a common Unit Element (UE) for each aspect of operation, the MAC formed as a plurality of MAC UEs, a plurality of Bias UEs, and a plurality of Analog to Digital Conversion (ADC) UEs which collectively perform a scalable MAC operation and generate a binary result. Each MAC UE, BIAS UE and ADC UE comprises groups of NAND gates with complementary outputs arranged in NAND-groups, each NAND gate coupled to a differential charge transfer bus through a binary weighted charge transfer capacitor to form an analog multiplication product as a charge applied to the differential charge transfer bus. The analog charge transfer bus is coupled to groups of ADC UEs with an ADC controller which enables and disables the ADC UEs using successive approximation to determine the accumulated multiplication result.


