Charge-Transfer Bus MAC Architecture for Low-Power Matrix Multiplication
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing hardware multiplier-accumulators for machine learning applications face challenges in scalability and power consumption due to synchronous operation and increased gate complexity, particularly in performing multiply-accumulate operations for large matrices, which results in high power dissipation and inefficiency.
Innovation Solution
A scalable asynchronous multiplier-accumulator architecture utilizing a common charge transfer bus for multiplier-accumulator, bias, and analog-to-digital converter unit elements, with NAND-groups and charge transfer capacitors to minimize displacement currents and power consumption by enabling asynchronous operation and shared binary weighted charge transfer lines.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If synchronous clocked operation is used in multiplier-accumulator, then operation timing is controlled and synchronized, but power dissipation increases due to displacement currents
Solution Approach 1:
The patent uses periodic clock signals to control the switching of unit elements in the multiplier-accumulator. The clock signals are applied periodically to enable multiply, accumulate, and clear operations at specific intervals, replacing continuous displacement currents with periodic charge transfers that occur only when needed.
Solution Approach 2:
The patent changes the operational mode from synchronous continuous clocking to asynchronous event-driven operation. The unit elements transition between held state, multiply state, and accumulate state based on control signals, changing the temporal parameters of operation to reduce unnecessary charge transfers and minimize displacement currents.
2Productivity
If large n×n multiplier is used for machine learning applications, then computational capability increases, but gate complexity increases as n2
Solution Approach 1:
The patent segments the multiplier-accumulator into multiple unit elements, each handling a portion of the computational task. Instead of implementing a single large n×n multiplier with n2 gates, the system uses multiple smaller unit elements that can be cascaded or operated in parallel, with each unit element having significantly fewer gates.
Solution Approach 2:
The patent introduces a shared charge transfer bus as an intermediary between unit elements. This bus enables multiple unit elements to share common charge transfer pathways, reducing the need for dedicated charge transfer lines between each pair of unit elements and thereby reducing overall gate complexity while maintaining computational capability.
3Productivity
If multiple adders are added for multiply-accumulate operations, then accumulation capability improves, but device complexity increases
Solution Approach 1:
The patent merges multiple accumulation functions into a single shared charge transfer bus. Instead of using separate adders for each accumulation operation, multiple unit elements transfer their partial products as charges to the same bus, where charges are summed through parallel charge transfer. This combining approach maintains accumulation capability while reducing the number of discrete adder circuits.
Solution Approach 2:
The shared charge transfer bus serves multiple functions: it acts as the output bus for multiply operations, the input bus for accumulate operations, and the interconnection medium between multiple unit elements. This multi-functionality eliminates the need for separate dedicated adders for each operation type.
4Device complexity
If common charge transfer bus is used for MAC, Bias, and ADC unit elements, then device complexity reduces, but charge interference may increase
Solution Approach 1:
The patent uses periodic clocking and controlled enabling of different unit element types (MAC, Bias, ADC) on the shared charge transfer bus. By activating only one type of unit element at a time through periodic control signals, the system prevents charge interference between different functional units while maintaining architectural simplicity through bus sharing.
Solution Approach 2:
The patent extracts and separates the control logic for different unit element types, using specific control signals to enable only the appropriate unit elements at each time interval. This extraction of control functions allows multiple unit element types to share the common bus without interference, as each type is activated only when its control signal is asserted.
Applied Scientific Principles
This section explains which scientific principles are used to turn an abstract innovation direction into a practical engineering solution.
Function Achieved in This Case
This architecture reduces power consumption by allowing asynchronous operation and minimizing displacement currents, while providing a scalable and flexible solution for machine learning applications by efficiently performing multiply-accumulate operations with reduced gate complexity and common mode imbalances.
Implementation Method 1
a scalable asynchronous multiplier-accumulator architecture utilizing a common charge transfer bus for multiplier-accumulator, bias, and analog-to-digital converter unit elements
Data Source
AI summary
An architecture for a multiplier-accumulator (MAC) uses a common Unit Element (UE) for each aspects of operation, the MAC formed as a plurality of MAC UEs, a plurality of Bias UEs, and a plurality of Analog to Digital Conversion (ADC) UEs which collectively perform a scalable MAC operation and generate a binary result. Each MAC UE, BIAS UE and ADC UE comprises groups of NAND gates with complementary outputs arranged in AND-groups, each AND gate coupled to a charge transfer bus through a charge transfer capacitor Cu to form an analog multiplication product. Each UE transfers differential charge to the charge transfer bus. The analog charge transfer bus is coupled to groups of ADC UEs with an ADC controller which enables and disables the ADC UEs using successive approximation to determine the accumulated multiplication result.


