Dual-Mode In-Memory Memory Array for Wide Vector Neural Computing
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing memory architectures struggle to support high processing speed and low power consumption in neural computing applications, particularly in deep neural networks, due to the need for wide vector access and local processing, while also requiring error correction capabilities.
Innovation Solution
A memory array architecture that supports both conventional memory access and digital in-memory computation modes, utilizing a matrix of sub-arrays with word line and bit line configurations, enabling parallel processing and error correction through column multiplexing and error correction coding, allowing for wide vector access and efficient computation.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Productivity
If conventional memory access mode is used with single word line actuation, then error correction capabilities are maintained, but processing speed and computational efficiency are limited
Solution Approach 1:
The row decoder circuit is designed to dynamically switch between two operational modes: conventional mode (actuating only one word line) for error-corrected memory access, and computational mode (simultaneously actuating one word line per sub-array) for high-speed in-memory computation. This dynamic reconfiguration allows the same hardware to serve dual purposes without permanent structural complexity
Solution Approach 2:
The memory array architecture is designed to perform multiple functions using the same physical structure: it can operate as a conventional error-corrected memory system or as an in-memory computation engine for neural networks. The universal design eliminates the need for separate dedicated hardware for each function, resolving the complexity issue
2Productivity
If wide vector access is implemented for neural computing, then computational efficiency improves, but power consumption increases
Solution Approach 1:
The patent merges memory access operations with computational operations by enabling simultaneous actuation of multiple word lines (one per sub-array) during in-memory computation. This combining of memory and compute functions in the same hardware structure eliminates the need for separate data movement operations, achieving wide vector access efficiency while reducing overall power consumption compared to conventional von Neumann architecture
Solution Approach 2:
The memory array performs computational operations directly within its structure through in-memory computation, where the memory cells themselves participate in the computation process. This self-service approach eliminates the need for external processing units to fetch and process data, reducing power consumption while maintaining high computational efficiency
3Productivity
If simultaneous word line actuation is used for in-memory computation, then processing throughput increases, but error correction capabilities are reduced
Solution Approach 1:
The memory array is segmented into multiple independent sub-arrays, each with its own local bit lines and output circuits. During in-memory computation, one word line is actuated per sub-array simultaneously, allowing parallel processing throughput increase. The segmentation isolates potential errors to individual sub-arrays while maintaining overall system reliability through the modular structure
Data Source
AI summary
A memory array includes sub-arrays with memory cells arranged in a row-column matrix where each row includes a word line and each sub-array column includes a local bit line. A control circuit supports: a first mode where only one word line in the memory array is actuated during a column multiplexed memory access operation; and a second mode where one word line per sub-array is simultaneously actuated during an in-memory computation operation. An input/output circuit for each column includes inputs to the local bit lines of the sub-arrays, a column data output coupled to the bit line inputs to provide data read from the array in the first mode, and a sub-array data output coupled to each bit line input to provide weight data read from the array in the second mode. A computational circuit executes the in-memory computation as a function of feature data and the read weight data.


