Hybrid MAC Unit Architecture Using Charge Reuse and Analog Summation
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing hardware accelerators for applications like Convolution Neural Networks (CNN) are power-hungry due to inefficient multiplication and accumulation (MAC) operations in digital circuits, which require additional power and slow down the circuit operation when resetting capacitors or using complex resistor strings.
Innovation Solution
A hybrid architecture using a 16x16 matrix of processing elements (PEs) connected via digital-to-analog converters (DACs) and analog-to-digital converters (ADCs), employing capacitors and resistors to perform 256 multiplications and additions within a single clock cycle with low power consumption, without the need for extra clock cycles to reset capacitors and without using buffers or additional current sources.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Productivity
If capacitors are reset to zero between successive multiplication operations, then the circuit can perform multiple independent multiplications, but the operation speed decreases and power consumption increases due to extra clock cycles and wasted stored charge
Solution Approach 1:
The patent maintains continuous useful action by preserving the charge stored in capacitors between multiplication operations. Instead of resetting capacitors to zero, the system keeps the stored charge intact, allowing the same capacitors to be reused for subsequent operations without requiring discharge and recharge cycles, thereby eliminating idle time and reducing power consumption.
Solution Approach 2:
The patent recovers the charge stored in capacitors that would otherwise be discarded during reset operations. By maintaining the stored charge between operations, the system recovers the energy investment made during charging, converting what would be wasted energy into useful computational resource for future operations.
2Productivity
If 28 resistors and 256 switches are used to implement an 8 bit multiplier unit, then multiplication operations can be performed, but the circuit occupies larger area similar to single resistive string DACs
Solution Approach 1:
The patent segments the large-scale resistor network into multiple smaller resistive strings. Instead of using a single large array of 28 resistors and 256 switches, the system divides the computation into multiple independent resistive strings, each handling a portion of the multiplication task. This segmentation reduces the area required for each individual string while maintaining overall computational capability.
Solution Approach 2:
The patent transitions from a two-dimensional array layout to a three-dimensional stacked architecture by implementing multiple resistive strings in vertical layers. This dimensional change allows the circuit to achieve the required computation capability with reduced planar area, as the resistive elements are distributed across multiple vertical levels rather than spread out in a single large plane.
3Reliability
If buffers are used to connect two resistive strings, then loading effects are avoided, but additional power is dissipated and extra die area is occupied
Solution Approach 1:
The patent extracts and removes the buffer components from the circuit architecture. Instead of using buffers to connect resistive strings, the system directly couples the strings through carefully designed switch networks that prevent loading effects without requiring additional active buffering elements, thereby eliminating the associated power dissipation and area overhead.
Solution Approach 2:
The patent introduces switch networks as intermediary elements between resistive strings instead of buffers. These switches act as mediators that control the interaction between strings, preventing direct loading effects while requiring minimal power and occupying small area compared to buffer amplifiers.
4Productivity
If additional current sources are used to drive sub-resistive strings, then multiplication operations can be performed, but the system complexity increases
Solution Approach 1:
The patent makes the existing current sources universal by designing them to serve multiple functions. Instead of adding dedicated current sources for each sub-resistive string, the system configures the existing current sources to drive multiple strings through selective switching, thereby maintaining computation capability while avoiding the complexity increase that would result from proliferating dedicated current sources.
Solution Approach 2:
The patent merges the current driving function across multiple sub-resistive strings into shared current sources. By combining the current supply function and using switch networks to distribute current to different strings as needed, the system achieves the required computation capability with fewer current sources, thereby reducing overall system complexity.
Applied Scientific Principles
This section explains which scientific principles are used to turn an abstract innovation direction into a practical engineering solution.
Function Achieved in This Case
This approach enables faster and more energy-efficient MAC operations by reusing capacitor charge and eliminating the need for extra power to reset capacitors, while maintaining the accuracy of output voltage without the loading effects of sub-resistive strings.
Implementation Method 1
In a conventional circuit using capacitors, the multiplication operation is realized by charging the analog input into one capacitor value which is proportional to the digital input and, sharing the charge into another fixed value capacitor.
Implementation Method 2
To implement a conventional 8 bit multiplier unit with resistors, it requires 28 (=256) resistors and 256 switches to generate 256 distinct voltage samples from the input voltage
Implementation Method 3
The digital inputs (110) are converted to analog signal (114) using digital to analog converters (DAC) (102)
Implementation Method 4
Analog to digital converters (ADC) (104) are used to convert the analog MAC output (116) to digital form (118)
Data Source
AI summary
The present invention provides an analog-digital hybrid architecture, which performs 256 multiplications and additions at a time. The system comprises 256 Processing Elements (PE) (108), which are arranged in a matrix form (16 rows and 16 columns). The digital inputs (110) are converted to analog signal (114) using digital to analog converters (DAC) (102). One PE (108) produces one analog output (115) which is nothing but the multiplication of the analog input (114) and the digital weight input (112). The implementation of PE is done by using i) capacitors and switches and ii) resistor and switches. The outputs from multiple PEs (108) in a column are connected together to produce one analog MAC output (116). In the similar manner, the system produces 16 MAC outputs (118) corresponding to 16 columns. Analog to digital converters (ADC) (104) are used to convert the analog MAC output (116) to digital form (118).


