3D Neural Network Circuit In-Memory MAC Operations

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional deep neural networks (DNNs) mapped on Von Neumann computing architectures face an energy and time bottleneck due to the extensive movement of synaptic weights between memory and processing units, which is inefficient and consumes more energy than computation itself, particularly in resource-constrained devices like IoT and mobile devices.

Innovation Solution

A hardware implementation of a dense and energy-efficient neural network architecture where synaptic weights and neuron functionalities are integrated in a 3D stacked memory array, utilizing transistors operating in the subthreshold regime to perform MAC operations and non-linearity functions without the need for DACs or OPAMPS, by comparing competing current components generated by p-channel and n-channel MOSFETs.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Use of energy by moving object

If DNNs are mapped on conventional Von Neumann computing architectures, then computation can be performed, but energy consumption increases significantly due to extensive data movement between memory and processing units

Engineering Contradiction:
Improveenergy consumptionVSAvoidenergy wasted in data movement
Core Design Contradiction:
Use of energy by moving objectVSLoss of energy

Solution Approach 1:

The patent merges the memory unit and processing unit into a single integrated neuromorphic chip, eliminating the separation between storage and computation. Synaptic weights are stored directly within the processing units, allowing computations to occur in-place without transferring data between memory and processor, thus resolving the energy waste associated with data movement in conventional Von Neumann architectures

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The patent transitions from the traditional sequential data movement model to a parallel in-memory computation model by introducing conductance-mode neurons that perform MAC operations directly within the memory array. This dimensional shift from von Neumann bottleneck to in-memory parallel processing fundamentally changes how data is handled, enabling energy-efficient computation without repeated data transfers

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

2Use of energy by moving object

If data transfer precision is reduced from floating point to integer or 1bit precision, then energy efficiency improves and data transfer is reduced, but computational accuracy may be compromised

Engineering Contradiction:
Improveenergy efficiencyVSAvoidcomputational accuracy
Core Design Contradiction:
Use of energy by moving objectVSMeasurement precision

Solution Approach 1:

The patent replaces conventional digital arithmetic operations with analog conductance-based computations. By using conductance-mode neurons that perform multiplication through physical conductance values rather than digital arithmetic, the system achieves energy-efficient computation while maintaining accuracy through analog signal processing. This substitution of mechanical/digital operations with physical analog operations enables low-power computing without sacrificing precision

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

3Productivity

If a dense neural network architecture is implemented, then computational capability increases, but hardware complexity and resource requirements increase

Engineering Contradiction:
Improvecomputational capabilityVSAvoidhardware complexity
Core Design Contradiction:
ProductivityVSDevice complexity

Solution Approach 1:

The patent implements multi-functionality by integrating multiple operations within single neuromorphic units. Each neuron performs both storage (through conductance modulation) and computation (through current-based MAC operations) simultaneously. This universal design allows dense network architectures to be implemented without proportionally increasing hardware complexity, as the same basic unit handles multiple functions required by dense connectivity patterns

Inventive Principle:
Principle #6Universality (Multi-functionality)

Applied Scientific Principles

This section explains which scientific principles are used to turn an abstract innovation direction into a practical engineering solution.

Function Achieved in This Case

This solution significantly reduces energy consumption by performing computations in a compact, high-density 3D architecture, allowing for efficient mapping of DNNs on a standalone chip, thereby enhancing battery life in resource-constrained devices.

Implementation Method 1

utilizing transistors operating in the subthreshold regime to perform MAC operations and non-linearity functions

Methodology Applied
Scientific EffectSubthreshold regime operation:

Implementation Method 2

by comparing competing current components generated by p-channel and n-channel MOSFETs

Methodology Applied
Scientific EffectMOSFET current generation:

Data Source

PatentEP3654250B1Machine learning accelerator
Publication Date: 2023.04.12 INTERUNIVERSITAIR MICRO ELECTRONICS CENT (IMEC VZW)
  • EP3654250B1 patent drawingFigure 1~2
  • EP3654250B1 patent drawingFigure 3
  • EP3654250B1 patent drawingFigure 4

AI summary

A neural network circuit for providing a thresholded weighted sum of input signals comprises - at least two arrays of transistors with programmable threshold voltage, each transistor storing a synaptic weight as a threshold voltage and having a control electrode for receiving an activation input signal, each transistor of the at least two arrays providing an output current for either a positive weighted current component in an array of a set of first arrays or a negative weighted current component in an array of a set of second arrays, - for each array of transistors, a reference network associated therewith, for providing a reference signal to be combined with the positive or negative weight current components of the transistors of the associated array, the reference signal having opposite sign compared to the weight current components of the associated array, thereby providing the thresholding of the weighted sums of the currents, - at least one bit line for receiving the combined positive and/or negative current components, each combined with their associated reference signals.