Unary-Weighted MAC Unit for Low Area AI Computing

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional binary weighted current steering DACs for MAC units in AI applications are inefficient due to large area consumption, high power usage, and scalability issues, requiring extensive calibration and being unsuitable for deep learning computations.

Innovation Solution

A high resolution, low area, low power MAC unit utilizing a unary-weighted current steering architecture with fewer current sources and hierarchical current division, employing smaller switching transistors to achieve tight current matching and efficient current scaling, compatible with MRAM memory devices.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Speed

If binary weighted current steering DAC architecture is used for MAC units, then high speed and low latency are achieved, but large area consumption and high power usage occur

Engineering Contradiction:
Improvecomputation speedVSAvoidDAC area
Core Design Contradiction:
SpeedVSArea of stationary object

Solution Approach 1:

The patent segments the current steering DAC into multiple smaller current source groups (e.g., 16 groups of 4-current-source arrays for an 8-bit DAC) rather than using a single large binary-weighted current source array. Each group handles a portion of the bit resolution, and the groups are combined through current summation nodes. This segmentation reduces the area of individual current sources while maintaining the overall high-speed performance through parallel operation of multiple smaller groups.

Inventive Principle:
Principle #1Segmentation

2Speed

If binary weighted current steering DAC architecture is used for MAC units, then high speed and low latency are achieved, but high power usage occurs

Engineering Contradiction:
Improvecomputation speedVSAvoidpower consumption
Core Design Contradiction:
SpeedVSUse of energy by stationary object

Solution Approach 1:

The patent divides the high-power binary-weighted current sources into multiple lower-power segmented current source groups. Each segment uses fewer current sources (e.g., 4-current-source arrays) that consume less power individually. The segments operate in parallel and their outputs are summed, achieving the same computational function with reduced total power consumption compared to a single large binary-weighted array.

Inventive Principle:
Principle #1Segmentation

3Measurement precision

If higher number of bits is used in current steering DAC, then higher resolution is achieved, but extensive calibration is needed due to mismatch

Engineering Contradiction:
ImproveDAC resolutionVSAvoidcalibration complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent achieves high resolution (e.g., 8-bit) by segmenting the DAC into multiple lower-resolution current source groups (e.g., 16 groups with 0.5-bit resolution each). Each segment uses a small number of current sources that are easier to match, reducing the calibration complexity. The high overall resolution is achieved through the combination of multiple segments rather than through a single large array of precisely matched current sources.

Inventive Principle:
Principle #1Segmentation

4Measurement precision

If higher number of bits is used in current steering DAC, then higher resolution is achieved, but area increases as number of current sources increases

Engineering Contradiction:
ImproveDAC resolutionVSAvoidDAC area
Core Design Contradiction:
Measurement precisionVSArea of stationary object

Solution Approach 1:

The patent implements high-resolution current steering by dividing the DAC into multiple segments, where each segment contains a small number of current sources (e.g., 4-current-source arrays). For an 8-bit DAC, 16 such segments are used, each contributing to the overall resolution. This segmented approach achieves high resolution without requiring a single large array of 256 current sources, thereby significantly reducing the total area while maintaining the required bit resolution.

Inventive Principle:
Principle #1Segmentation

5Productivity

If conventional MAC unit architecture is used, then basic multiply-accumulate operations are performed, but scalability issues arise for deep learning computations

Engineering Contradiction:
ImproveMAC operation capabilityVSAvoidscalability for deep learning
Core Design Contradiction:
ProductivityVSAdaptability or versatility

Solution Approach 1:

The patent creates a universal MAC unit architecture that can perform various computational functions required for deep learning by configuring the segmented current source groups and switching networks in different patterns. The same physical hardware can be reconfigured to perform different types of computations (e.g., different MAC operations, vector operations, matrix operations) by changing the switching control signals, providing the adaptability and versatility needed for deep learning applications while maintaining basic MAC operation capability.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Data Source

PatentUS11544037B2Low area multiply and accumulate unit
Publication Date: 2023.01.03 INTERNATIONAL BUSINESS MACHINE CORPORATION
  • US11544037B2 patent drawing
  • US11544037B2 patent drawing
  • US11544037B2 patent drawing

AI summary

An improved electronic mixed mode multiplier and accumulate circuit for artificial intelligence and computing system applications that perform vector-vector, vector-matrix and other multiply-accumulate computations. The circuit is provided is a high resolution, high linearity, low area, low power multiply—accumulate (MAC) unit to interface with a memory device for storing computation output results. The MAC unit uses a less number of current carrying elements resulting in much lower integrated circuit area, and provides a tight matching between the current elements thus preserving inherent linearity requirements due to current mode operation. Further the MAC performs current scaling using switches and current division where the current switches occupy minimum size transistors requiring a small area to implement that renders it compatible with MRAM such as a magnetic tunnel junction device. The MAC is hierarchically extended for increased number of bits to provide a delay implementation using orthogonal vector and current addition.