Differential MAC Unit on a Shared Charge Bus for Low-Power Scaling
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing multiplier-accumulator architectures for machine learning applications face challenges in scalability and power consumption due to synchronous operation and increased gate complexity, particularly in performing multiply-accumulate operations for large matrices, which results in high power dissipation and inefficiency.
Innovation Solution
A scalable asynchronous multiplier-accumulator architecture utilizing a common charge transfer bus for multiplier-accumulator, bias, and analog-to-digital converter unit elements, with NAND-groups and charge transfer capacitors to minimize displacement currents and power consumption by enabling asynchronous operation and shared binary weighted charge transfer lines.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Power
If synchronous clocked stages are used for multiplier operation, then operation timing is controlled, but power dissipation increases
Solution Approach 1:
The patent transitions from static synchronous clocked stages to dynamic asynchronous operation where the system adapts its timing based on actual computation completion. The multiplier-accumulator uses dynamic timing signals that are generated based on when computations are actually complete, rather than being driven by a fixed global clock, thereby reducing unnecessary switching activity and power dissipation.
Solution Approach 2:
The patent extracts and removes the global clocking mechanism from the multiplier-accumulator architecture. By eliminating the synchronous clocked stages and their associated clock distribution networks, the design reduces power dissipation while maintaining proper operation timing through alternative asynchronous signaling methods.
2Productivity
If large n×n multipliers are used for machine learning operations, then computational capability increases, but gate complexity increases as n2
Solution Approach 1:
The patent segments the large n×n multiplication task into multiple smaller multiplier-accumulator units that operate in parallel. Each unit handles a portion of the computation, and the results are accumulated through shared charge transfer buses. This segmentation reduces the gate complexity of individual units while maintaining the overall computational capability through parallel processing.
Solution Approach 2:
The patent designs universal multiplier-accumulator units that can be configured and cascaded to handle different matrix sizes and operations. These multi-functional units can perform multiplication, accumulation, and can be arranged in various configurations to support different computational requirements, thereby reducing the need for dedicated complex circuits for each specific operation size.
3Adaptability or versatility
If multiple adders are added for multiply-accumulate operations, then computational functionality is enhanced, but power dissipation and complexity increase
Solution Approach 1:
The patent merges the multiplication and accumulation functions into unified multiplier-accumulator units that share common hardware resources. The accumulation function is integrated with the multiplication logic, allowing the same circuit elements to perform both operations sequentially or in parallel, thereby reducing the total number of separate adders and associated power dissipation.
Solution Approach 2:
The patent replaces traditional ripple-carry adder chains with charge-based accumulation mechanisms. Instead of using multiple sequential adders that require extensive carry propagation logic, the design uses charge transfer and summation at node points, which reduces the logical complexity and power consumption of the accumulation function while maintaining the same computational capability.
4Loss of energy
If asynchronous operation is implemented to reduce power consumption, then power dissipation decreases, but timing control becomes more challenging
Solution Approach 1:
The patent implements feedback mechanisms where completion signals from computational units are fed back to control the timing of subsequent operations. The system uses carry-propagation detection and accumulation completion signals to dynamically control the timing of data transfer and processing stages, ensuring proper synchronization without requiring a global clock. This feedback-based timing control maintains ease of operation while enabling asynchronous low-power operation.
Applied Scientific Principles
This section explains which scientific principles are used to turn an abstract innovation direction into a practical engineering solution.
Function Achieved in This Case
This architecture reduces power consumption by minimizing displacement currents and allows for scalable and efficient multiply-accumulate operations, particularly when the kernel is static, while maintaining accuracy and flexibility in handling various matrix sizes.
Implementation Method 1
each NAND gate having a positive output coupled through a binary weighted positive charge transfer capacitor to a unique positive charge transfer line and a binary weighted negative output coupled through a negative charge transfer capacitor to a unique negative charge transfer line
Implementation Method 2
each MAC UE providing a result as a charge transferred to a differential charge transfer bus
Implementation Method 3
a second plurality of Bias unit elements (Bias UEs) performing a bias operation and placing a bias value as a charge onto the differential charge transfer bus
Implementation Method 4
a third plurality of ADC unit elements (ADC UEs) operative to convert a charge present on the differential charge transfer bus into a digital output value
Data Source
AI summary
A Unit Element (UE) has a digital X input and a digital W input, and comprises groups of NAND gates generating complementary outputs which are coupled to differential charge transfer lines through respective charge transfer capacitor Cu. The number of bits in the X input determines the number of NAND gates in a NAND-group and the number of bits in the W input determines the number of NAND groups. Each NAND-group receives one bit of the W input applied to all of the NAND gates of the NAND-group, and each unit element having the bits of X applied to each associated NAND gate input of each unit element. The NAND gate outputs are coupled through a charge transfer capacitor Cu to charge transfer lines. Multiple Unit Elements may be placed in parallel to sum and scale the charges from the charge transfer lines, the charges coupled to an analog to digital converter which forms the dot product output.


