Digital IMC Macro Using Approximate Arithmetic for Area Efficiency

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current in-memory computing (IMC) SRAMs using analog-mixed-signal (AMS) hardware face significant precision and accuracy limitations due to process, voltage, and temperature variations, while digital IMC SRAMs are less energy-efficient and have worse area efficiency due to requiring more transistors.

Innovation Solution

A digital IMC macro circuit utilizing approximate arithmetic hardware with custom full adder circuits and pass gate logic in a ripple carry adder tree, reducing the number of transistors and devices while maintaining reduced variability, and incorporating an approximation-aware training algorithm and multi-bit XNOR number format to minimize inference accuracy degradation.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If digital logic is used in IMC SRAM to eliminate variability, then reliability is improved, but area efficiency deteriorates due to requiring more transistors

Engineering Contradiction:
Improvecomputing precisionVSAvoidarea efficiency
Core Design Contradiction:
ReliabilityVSArea of stationary object

Solution Approach 1:

The patent changes the parameter of arithmetic precision from exact to approximate by using approximate compressors instead of exact compressors. This parameter change allows the circuit to achieve sufficient computational accuracy while using fewer transistors, thereby improving area efficiency while maintaining acceptable reliability for AI workloads

Inventive Principle:
Principle #35Parameter changes

Solution Approach 2:

The patent employs approximate arithmetic hardware that is simpler and requires fewer transistors compared to exact digital logic. This approach accepts a certain level of computational approximation in exchange for significantly reduced hardware resources, effectively trading precision for area efficiency

Inventive Principle:
Principle #27Cheap short-living objects (Disposable)

2Area of stationary object

If approximate arithmetic hardware is used, then area efficiency is improved, but manufacturing precision deteriorates due to PVT variations

Engineering Contradiction:
Improvearea efficiencyVSAvoidcomputing accuracy
Core Design Contradiction:
Area of stationary objectVSManufacturing precision

Solution Approach 1:

The patent implements calibration mechanisms that use feedback to adjust and compensate for PVT variations. By measuring actual circuit behavior under different conditions and applying correction factors, the system maintains computing accuracy despite using approximate arithmetic hardware that is more susceptible to process, voltage, and temperature variations

Inventive Principle:
Principle #23Feedback

3Area of stationary object

If full adder circuits with pass gate logic are used, then area efficiency is improved, but device complexity increases due to custom circuit design

Engineering Contradiction:
Improvearea efficiencyVSAvoidcircuit complexity
Core Design Contradiction:
Area of stationary objectVSDevice complexity

Solution Approach 1:

The patent segments the adder tree into multiple levels with approximate compressors at each level. This segmentation allows the use of simplified full adder circuits with pass gate logic in each segment, reducing the transistor count per unit area while managing complexity through modular hierarchical structure

Inventive Principle:
Principle #1Segmentation

Data Source

PatentUS20230266943A1Digital in-memory computing macro based on approximate arithmetic hardware
Publication Date: 2023.08.24 THE TRUSTEES OF COLUMBIA UNIV IN THE CITY OF NEW YORK
  • US20230266943A1 patent drawing
  • US20230266943A1 patent drawing
  • US20230266943A1 patent drawing

AI summary

Various embodiments described herein provide for a digital In-Memory Computing (IMC) macro circuit that utilizes approximate arithmetic hardware to reduce the number of transistors and devices in the circuit relative to a convention digital IMC, thereby improving the area-efficiency of the digital IMC, but while retaining the benefits of reduced variability relative to an analog-mixed-signal (AMS) circuit. The proposed digital IMC macro circuit also includes custom full adder (FA) circuits with pass gate logic in a ripple carry adder (RCA) tree. The disclosed digital IMC macro circuit can also perform a vector-matrix dot product in one cycle while achieving high energy and area efficiency.