Unary-Weighted MAC Unit for Low Area AI Computing
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional binary weighted current steering DACs for MAC units in AI applications are inefficient due to large area consumption, high power usage, and scalability issues, requiring extensive calibration and being unsuitable for deep learning computations.
Innovation Solution
A high resolution, low area, low power MAC unit utilizing a unary-weighted current steering architecture with fewer current sources and hierarchical current division, employing smaller switching transistors to achieve tight current matching and efficient current scaling, compatible with MRAM memory devices.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Speed
If binary weighted current steering DAC architecture is used for MAC units, then high speed and low latency are achieved, but large area consumption and high power usage occur
Solution Approach 1:
The patent segments the current steering DAC into multiple smaller current source groups (e.g., 16 groups of 4-current-source arrays for an 8-bit DAC) rather than using a single large binary-weighted current source array. Each group handles a portion of the bit resolution, and the groups are combined through current summation nodes. This segmentation reduces the area of individual current sources while maintaining the overall high-speed performance through parallel operation of multiple smaller groups.
2Speed
If binary weighted current steering DAC architecture is used for MAC units, then high speed and low latency are achieved, but high power usage occurs
Solution Approach 1:
The patent divides the high-power binary-weighted current sources into multiple lower-power segmented current source groups. Each segment uses fewer current sources (e.g., 4-current-source arrays) that consume less power individually. The segments operate in parallel and their outputs are summed, achieving the same computational function with reduced total power consumption compared to a single large binary-weighted array.
3Measurement precision
If higher number of bits is used in current steering DAC, then higher resolution is achieved, but extensive calibration is needed due to mismatch
Solution Approach 1:
The patent achieves high resolution (e.g., 8-bit) by segmenting the DAC into multiple lower-resolution current source groups (e.g., 16 groups with 0.5-bit resolution each). Each segment uses a small number of current sources that are easier to match, reducing the calibration complexity. The high overall resolution is achieved through the combination of multiple segments rather than through a single large array of precisely matched current sources.
4Measurement precision
If higher number of bits is used in current steering DAC, then higher resolution is achieved, but area increases as number of current sources increases
Solution Approach 1:
The patent implements high-resolution current steering by dividing the DAC into multiple segments, where each segment contains a small number of current sources (e.g., 4-current-source arrays). For an 8-bit DAC, 16 such segments are used, each contributing to the overall resolution. This segmented approach achieves high resolution without requiring a single large array of 256 current sources, thereby significantly reducing the total area while maintaining the required bit resolution.
5Productivity
If conventional MAC unit architecture is used, then basic multiply-accumulate operations are performed, but scalability issues arise for deep learning computations
Solution Approach 1:
The patent creates a universal MAC unit architecture that can perform various computational functions required for deep learning by configuring the segmented current source groups and switching networks in different patterns. The same physical hardware can be reconfigured to perform different types of computations (e.g., different MAC operations, vector operations, matrix operations) by changing the switching control signals, providing the adaptability and versatility needed for deep learning applications while maintaining basic MAC operation capability.
Data Source
AI summary
An improved electronic mixed mode multiplier and accumulate circuit for artificial intelligence and computing system applications that perform vector-vector, vector-matrix and other multiply-accumulate computations. The circuit is provided is a high resolution, high linearity, low area, low power multiply—accumulate (MAC) unit to interface with a memory device for storing computation output results. The MAC unit uses a less number of current carrying elements resulting in much lower integrated circuit area, and provides a tight matching between the current elements thus preserving inherent linearity requirements due to current mode operation. Further the MAC performs current scaling using switches and current division where the current switches occupy minimum size transistors requiring a small area to implement that renders it compatible with MRAM such as a magnetic tunnel junction device. The MAC is hierarchically extended for increased number of bits to provide a delay implementation using orthogonal vector and current addition.


