Memory Device Page-Buffer Processing for AI Memory Bottlenecks
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
The memory wall constraint in high-performance computing systems, where the memory access speed lags behind the processor computation speed, limits the performance of AI systems, particularly those using large transformer models, due to high power consumption and memory bottlenecks.
Innovation Solution
A memory device with a peripheral circuit that includes page buffers, process units, and control logic to perform calculations within the memory device, distributing computational tasks and breaking the memory wall by processing data within the memory system.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Quantity of substance
If data is stored in external memory and processed by separate processor, then data storage capacity is improved, but memory access speed lags behind processor computation speed causing memory bottleneck
Solution Approach 1:
The patent merges the processing unit and memory into a single integrated memory device. The processing unit is directly coupled to the memory cell array through a data bus, enabling data processing to be performed within the memory device itself. This eliminates the need for separate external memory and processor, resolving the speed mismatch between memory access and processor computation by making them co-located and synchronized.
Solution Approach 2:
The patent introduces a data bus as an intermediary component that directly connects the processing unit to the memory cell array. This data bus serves as a high-speed communication channel that mediates data transfer between storage and processing functions within the same device, enabling faster data access compared to external memory interfaces.
2Adaptability or versatility
If large transformer models are used for AI computation, then AI reasoning capability is improved, but power consumption increases and memory bottlenecks form
Solution Approach 1:
The patent combines multiple functions (storage, processing, and computation) into a single memory device. By integrating the processing unit within the memory device, data can be processed locally without requiring frequent transfers to external processors, reducing overall power consumption while supporting large transformer models for AI reasoning.
3Productivity
If data is transferred between memory and processor, then computation is improved, but data transfer time increases creating memory wall
Solution Approach 1:
The patent extracts the processing function from the external processor and places it directly within the memory device. This extraction eliminates the need for data to be transferred between separate memory and processor components, removing the data transfer time bottleneck while maintaining high computation performance through local processing.
Solution Approach 2:
The patent merges storage and processing functions into a single integrated device. The processing unit is directly coupled to the memory cell array, enabling data to be processed in-place without physical transfer between components. This eliminates data transfer time entirely while maintaining high computation productivity through immediate local processing.
Data Source
AI summary
A memory device, a memory system, and a method for data calculation with the memory device are provided. The memory device includes an array of memory cells and a peripheral circuit coupled to the memory cells is provided. The peripheral circuit includes page buffers configured to store first data transmitted from a data interface of the memory device and to sense second data from the array of memory cells. The peripheral circuit further includes at least one process unit coupled to the page buffers via a data-path bus of the peripheral circuit and configured to perform calculation based on the first data and the second data. The peripheral circuit further includes a control logic configured to program the second data into the array of memory cells.


