Integrated Processing Device in Volatile Memory for Parallel Access
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current machine learning systems face limitations due to limited memory bandwidth, particularly in deep neural networks, where the interface between processor chips and DRAM introduces latency and power consumption overhead, making it inefficient to store and process large amounts of data.
Innovation Solution
Integrating a processing device within a volatile memory device, such as DRAM, to enable parallel access to memory regions, reducing reliance on external memory interfaces and minimizing power consumption by allowing simultaneous read/write operations across multiple banks.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Quantity of substance
If processor chips and DRAM are used as separate external devices, then high density memory storage is achieved, but memory bandwidth is limited and power consumption increases due to interface overhead
Solution Approach 1:
The patent merges the processor and memory into a single integrated device, eliminating the external interface between them. The processing device is built within the volatile memory device, allowing direct access to memory regions without going through external interfaces, thus resolving the bandwidth limitation while maintaining high storage capacity
2Quantity of substance
If processor chips and DRAM are used as separate external devices, then high density memory storage is achieved, but power consumption increases due to interface overhead
Solution Approach 1:
By integrating the processing device within the volatile memory device, the patent eliminates power-consuming external interfaces and data transfers. The processing device accesses memory regions directly within the same device, significantly reducing power consumption while maintaining high storage capacity
3Quantity of substance
If memory is accessed through external interfaces, then large amounts of memory are available, but latency increases and bandwidth is limited
Solution Approach 1:
The integration of processing and memory into a single device eliminates external interface delays. The processing device can directly access any memory region within the volatile memory device without latency introduced by external interfaces, achieving both large capacity and high access speed
Solution Approach 2:
The volatile memory device is divided into multiple memory regions that can be accessed in parallel by the processing device. This segmentation allows simultaneous read/write operations across different regions, increasing overall memory access speed while maintaining large total capacity
4Productivity
If parallel access to memory regions is enabled, then processing efficiency increases, but device complexity increases
Solution Approach 1:
The volatile memory device is segmented into multiple memory regions, each independently accessible by the processing device. This segmentation enables parallel read/write operations across different regions, improving processing efficiency while the integrated design keeps the overall device complexity manageable
Data Source
AI summary
A memory system having a processing device (e.g., CPU) and memory regions (e.g., in a DRAM device) on the same chip or die. The memory regions store data used by the processing device during machine learning processing (e.g., using a neural network). One or more controllers are coupled to the memory regions and configured to: read data from a first memory region (e.g., a first bank), including reading first data from the first memory region, where the first data is for use by the processing device in processing associated with machine learning; and write data to a second memory region (e.g., a second bank), including writing second data to the second memory region. The reading of the first data and writing of the second data are performed in parallel.


