Computational Memory Banks with SIMD Controllers
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing computational memory architectures for deep learning and neural networks face inefficiencies in power consumption, complexity, and chip area requirements due to high power usage in moving data between memory and processing elements, and are inadequate for handling large two-dimensional images and complex operations like permutations and table lookups.
Innovation Solution
The design incorporates a plurality of computational memory banks with arrays of memory units and processing elements connected by row and column buses, featuring single instruction, multiple data (SIMD) controllers within each bank for efficient instruction distribution and flexible low-precision arithmetic, allowing for local storage and decoding of instructions and coefficients, thereby reducing power consumption and increasing processing efficiency.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Use of energy by moving object
If data is physically moved between memory and processing elements, then data access is enabled, but power consumption increases significantly
Solution Approach 1:
The patent combines memory units and processing elements into a single computational memory bank, eliminating the need for physical data movement between separate memory and processing components. The processing elements are directly coupled to memory units within the same bank, enabling computations to be performed on data without transferring it externally, thus reducing power consumption while maintaining data accessibility.
Solution Approach 2:
The computational memory is divided into multiple independent banks, each containing its own processing elements and memory units. This segmentation allows parallel processing across multiple banks while keeping the distance between memory and processing elements minimal within each bank, thereby reducing the power required for data access operations.
2Productivity
If more processing elements are added to handle deep learning computations, then computational capability increases, but chip area requirements increase
Solution Approach 1:
By merging memory units and processing elements into integrated computational memory banks, the patent achieves high computational capability without proportionally increasing chip area. The co-location of processing elements with memory units eliminates the need for separate memory arrays and data buses, thereby reducing overall chip area while maintaining or enhancing computational throughput for deep learning workloads.
3Quantity of substance
If data is moved over longer distances between memory and processing elements, then larger memory capacity is achieved, but power consumption due to capacitance charging increases
Solution Approach 1:
The patent segments the computational memory into multiple smaller banks, each with its own processing elements. This segmentation ensures that the distance between memory units and processing elements remains short within each bank, reducing the capacitance charging requirements and associated power consumption, while the overall system achieves large effective memory capacity through the aggregation of multiple banks.
Data Source
AI summary
An example device includes a plurality of computational memory banks. Each computational memory bank of the plurality of computational memory banks includes an array of memory units and a plurality of processing elements connected to the array of memory units. The device further includes a plurality of single instruction, multiple data (SIMD) controllers. Each SIMD controller of the plurality of SIMD controllers is contained within at least one computational memory bank of the plurality of computational memory banks. Each SIMD controller is to provide instructions to the at least one computational memory bank.


