In-Memory SRAM and Page-Buffer Computing for AI Memory Bottlenecks
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
The memory wall in AI systems, caused by the lag in memory access speed compared to processor computation speed, hinders high-performance computing, particularly in large transformer models that require significant data and computation, leading to high power consumption and performance constraints.
Innovation Solution
A memory device with a peripheral circuit containing process units that perform calculations under control logic, distributing AI system tasks within the memory device, especially those requiring large data-width, thereby reducing the need for data transfer to the processor.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Speed
If memory access speed is improved to match processor computation speed, then the memory wall constraint is resolved and high-performance computing is enabled, but device complexity increases due to the need for integrated process units and control logic
Solution Approach 1:
The patent merges the memory device with processing capabilities by integrating process units and control logic directly into the memory device structure. This combination allows computation to occur at the memory location, eliminating the need for frequent data transfers between separate memory and processor components, thereby resolving the memory wall constraint while managing complexity through integration.
Solution Approach 2:
The memory device is designed to perform multiple functions: it serves as both a storage medium and a processing unit. The integrated process units can perform calculations directly on data stored in the memory cells, making the device universal in its capability to handle both storage and computation tasks, thus improving effective access speed without proportionally increasing overall system complexity.
2Reliability
If large amounts of data are transferred to the processor for computation, then accurate AI calculations are performed, but power consumption increases and performance is constrained due to the memory wall
Solution Approach 1:
The control logic within the memory device prepares and processes data before it is transferred to the processor. By performing preliminary calculations and data preparation operations within the memory device itself, the system reduces the volume of data that needs to be transferred to the processor, thereby lowering power consumption while ensuring that the data ready for processing is accurately prepared for AI calculations.
3Measurement precision
If data is processed outside the memory device, then calculation accuracy is maintained, but the memory wall constraint persists and calculation speed is reduced
Solution Approach 1:
The integrated process units within the memory device act as intermediaries between the stored data and the external processor. These intermediaries can perform initial processing, filtering, and preparation of data, reducing the burden on external processors and enabling faster overall calculation speed while maintaining accuracy through the coordinated operation of the control logic and process units.
Data Source
AI summary
A memory device, a memory system, and a method for data calculation with the memory device are provided. The memory device includes an array of memory cells and a peripheral circuit coupled to the memory cells is provided. The peripheral circuit includes a static random-access memory (SRAM) configured to obtain first data transmitted from a data interface of the memory device, page buffers configured to sense second data from the array of memory cells, and at least one process unit coupled to the SRAM and the page buffers via a data-path bus of the peripheral circuit. At least one process unit is configured to perform calculation based on the first data and the second data. The peripheral circuit further includes a control logic configured to program the second data into the array of memory cells.


