Compute Near Memory Binary Neural Network Circuits

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Deep learning neural networks face significant energy inefficiencies due to high energy consumption in data transfer between memory and processors, particularly in von Neumann architectures, which creates bottlenecks in machine-learning applications and limits the energy efficiency of binary neural networks.

Innovation Solution

Implementing a Compute Near Memory (CNM) binary neural network accelerator with digital circuits, utilizing a two-level hierarchy of interleaved memory and compute units, including latch-based memory arrays and wide vector inner product execution units, to reduce data movement and enhance energy efficiency through high parallelism and near threshold voltage operation.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Productivity

If data transfer between memory and processor is performed in von Neumann architecture, then computational tasks can be executed, but energy consumption increases significantly and data transfer bottleneck occurs

Engineering Contradiction:
Improvecomputational throughputVSAvoidenergy consumption
Core Design Contradiction:
ProductivityVSUse of energy by moving object

Solution Approach 1:

The patent merges memory and compute units into a single integrated architecture where memory is placed adjacent to computational logic. This eliminates the separation between storage and processing, allowing data to be processed locally without frequent transfers back to the processor, thereby reducing energy consumption while maintaining computational throughput

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The patent segments the neural network processing into distributed units where each unit contains both memory and computational logic. This segmentation allows independent processing of data blocks, reducing the need for centralized data transfer and minimizing energy consumption associated with moving data between separate memory and processor components

Inventive Principle:
Principle #1Segmentation

2Productivity

If data transfer between memory and processor is performed in von Neumann architecture, then computational tasks can be executed, but data transfer bottleneck limits processing speed

Engineering Contradiction:
Improvecomputational throughputVSAvoiddata transfer speed
Core Design Contradiction:
ProductivityVSSpeed

Solution Approach 1:

By merging memory and compute units, the patent eliminates the data transfer bottleneck inherent in von Neumann architecture. Data can be processed locally within each unit without waiting for transfers from remote memory, significantly improving processing speed and throughput

Inventive Principle:
Principle #5Merging (Combining)

3Use of energy by moving object

If binary neural network operations are implemented, then energy efficiency improves, but circuit noise and process variations affect reliability

Engineering Contradiction:
Improveenergy efficiencyVSAvoidcircuit noise immunity
Core Design Contradiction:
Use of energy by moving objectVSReliability

Solution Approach 1:

The patent implements local quality by designing compute units with dedicated memory storage adjacent to each other, creating localized processing domains. This localization reduces the impact of circuit noise and process variations by confining operations within controlled, isolated units rather than across the entire chip, thereby improving reliability while maintaining energy efficiency

Inventive Principle:
Principle #3Local quality

Data Source

PatentEP3828775A1Energy efficient compute near memory binary neural network circuits
Publication Date: 2021.06.02 INTEL CORP
  • EP3828775A1 patent drawingFigure 1
  • EP3828775A1 patent drawingFigure 2
  • EP3828775A1 patent drawingFigure 3

AI summary

A compute near memory binary neural network accelerator with digital circuits that achieves energy efficiencies comparable to or surpassing a compute near memory binary neural network accelerator with analog circuits is provided. The compute near memory binary neural network accelerator with digital circuits is more process scalable, robust to process, voltage and temperature variations, and immune to circuit noise.