Shared Analog Bus Layout for Asynchronous MAC Unit Elements
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing multiplier-accumulator architectures for machine learning applications face challenges in scalability and power consumption due to synchronous operation and increased gate complexity, particularly in forming dot products for large matrices, which results in high power dissipation and inefficiency.
Innovation Solution
A scalable asynchronous multiplier-accumulator architecture utilizing a common charge transfer bus for multiplier-accumulator, bias, and analog-to-digital converter unit elements, with NAND-groups and charge transfer capacitors to minimize displacement currents and power consumption by enabling asynchronous operation and shared binary weighted charge transfer lines.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Ease of operation
If synchronous clocked stages are used for multiplier operation, then timing control is simplified, but power dissipation increases due to continuous clocking
Solution Approach 1:
The patent replaces continuous synchronous clocking with periodic pulse generation only when data changes occur. The flip-flop outputs generate enable pulses that trigger charge transfer only when input data changes, eliminating continuous periodic clocking and reducing power dissipation while maintaining timing control.
Solution Approach 2:
The system uses its own data change events to trigger operation. When input data changes, the flip-flops automatically generate enable pulses that initiate the charge transfer sequence, making the system self-regulating without external clock control, thus reducing power consumption.
2Productivity
If large n×n multipliers are implemented, then computational capability increases, but gate complexity increases as n2
Solution Approach 1:
The patent divides the multiplication process into segmented unit elements, each handling a specific bit position. Each unit element processes one bit of the multiplicand against all bits of the multiplier, with results accumulated through charge transfer. This segmentation reduces the complexity of individual units while maintaining overall computational capability.
Solution Approach 2:
The patent introduces charge transfer capacitors and differential charge transfer buses as intermediaries between computational units. These intermediaries carry multiplication results without requiring complex digital logic interconnections, simplifying the overall system architecture while enabling large-scale multiplication operations.
3Adaptability or versatility
If more adders are added for multiply-accumulate operations, then MAC functionality is enhanced, but power consumption and circuit complexity increase
Solution Approach 1:
The patent combines multiplication and accumulation functions into integrated unit elements. The charge transfer mechanism serves both to convey multiplication results and to perform accumulation by summing charges on the differential bus, eliminating the need for separate adder circuits and reducing power consumption.
Solution Approach 2:
The patent replaces traditional digital adder logic with an analog charge summation mechanism. Charges representing multiplication results are transferred and summed on the differential charge transfer bus, converting complex digital addition operations into simpler analog charge accumulation, thereby reducing power consumption and circuit complexity.
4Loss of energy
If static kernel values are utilized, then power savings can be realized, but the architecture must minimize displacement currents
Solution Approach 1:
The patent pre-charges the differential charge transfer bus to match the static kernel values before multiplication operations begin. This preliminary action eliminates displacement currents during subsequent operations with static kernels, as no charge transfer is needed when kernel values do not change, thereby reducing power consumption.
Solution Approach 2:
The patent changes the operating parameters of the charge transfer bus dynamically. When kernel values are static, the bus remains in a stable charged state without active transfer. When kernel values change, the bus is recharged to the new values. This parameter change approach minimizes displacement currents during static periods while maintaining functionality during transitions.
Applied Scientific Principles
This section explains which scientific principles are used to turn an abstract innovation direction into a practical engineering solution.
Function Achieved in This Case
This architecture reduces power consumption by minimizing displacement currents and allows for flexible, scalable, and efficient multiply-accumulate operations, particularly when the kernel is static, thereby enhancing the performance of machine learning applications.
Implementation Method 1
each MAC UE providing a result as a charge transferred to a differential charge transfer bus
Implementation Method 2
each NAND gate having a positive output coupled through a binary weighted positive charge transfer capacitor to a unique positive charge transfer line and a binary weighted negative output coupled through a negative charge transfer capacitor to a unique negative charge transfer line
Data Source
AI summary
A planar fabrication charge transfer capacitor for coupling charge from a Unit Element (UE) generates a positive charge first output V_PP and a positive charge second output V_NP, the first output coupled to a positive charge line comprising a continuous first planar conductor, a continuous second planar conductor parallel to the first planar conductor, and a continuous third planar conductor parallel to the first planar conductor and second planar conductor, the charge transfer capacitor comprising, in sequence: a first co-planar conductor segment, the first planar conductor, a second co-planar conductor segment, the second planar conductor, a third co-planar conductor segment, the third planar conductor, and a fourth coplanar conductor segment, the first and third coplanar conductor segments capacitively edge coupled to the UE first output V_PP, the second and fourth coplanar conductor segments capacitively edge coupled to the UE second output V_NP.


