Compute Near Memory Binary Neural Network Circuits
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Deep learning neural networks face significant energy inefficiencies due to high energy consumption in data transfer between memory and processors, particularly in von Neumann architectures, which creates bottlenecks in machine-learning applications and limits the energy efficiency of binary neural networks.
Innovation Solution
Implementing a Compute Near Memory (CNM) binary neural network accelerator with digital circuits, utilizing a two-level hierarchy of interleaved memory and compute units, including latch-based memory arrays and wide vector inner product execution units, to reduce data movement and enhance energy efficiency through high parallelism and near threshold voltage operation.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Productivity
If data transfer between memory and processor is performed in von Neumann architecture, then computational tasks can be executed, but energy consumption increases significantly and data transfer bottleneck occurs
Solution Approach 1:
The patent merges memory and compute units into a single integrated architecture where memory is placed adjacent to computational logic. This eliminates the separation between storage and processing, allowing data to be processed locally without frequent transfers back to the processor, thereby reducing energy consumption while maintaining computational throughput
Solution Approach 2:
The patent segments the neural network processing into distributed units where each unit contains both memory and computational logic. This segmentation allows independent processing of data blocks, reducing the need for centralized data transfer and minimizing energy consumption associated with moving data between separate memory and processor components
2Productivity
If data transfer between memory and processor is performed in von Neumann architecture, then computational tasks can be executed, but data transfer bottleneck limits processing speed
Solution Approach 1:
By merging memory and compute units, the patent eliminates the data transfer bottleneck inherent in von Neumann architecture. Data can be processed locally within each unit without waiting for transfers from remote memory, significantly improving processing speed and throughput
3Use of energy by moving object
If binary neural network operations are implemented, then energy efficiency improves, but circuit noise and process variations affect reliability
Solution Approach 1:
The patent implements local quality by designing compute units with dedicated memory storage adjacent to each other, creating localized processing domains. This localization reduces the impact of circuit noise and process variations by confining operations within controlled, isolated units rather than across the entire chip, thereby improving reliability while maintaining energy efficiency
Data Source
Figure 1
Figure 2
Figure 3
AI summary
A compute near memory binary neural network accelerator with digital circuits that achieves energy efficiencies comparable to or surpassing a compute near memory binary neural network accelerator with analog circuits is provided. The compute near memory binary neural network accelerator with digital circuits is more process scalable, robust to process, voltage and temperature variations, and immune to circuit noise.