Multichannel Memory with Off-Chip Scratchpads for Capacity Limits
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing memory-intensive systems, such as those in neural network designs, face bandwidth and capacity constraints due to the limitations of on-chip scratchpads and global memory access, leading to inefficiencies and increased latency.
Innovation Solution
Implementing a multichannel DRAM architecture with auxiliary scratchpads configured as a unitary device, positioned off-chip and coupled to processing cores, to augment the capacity of on-chip scratchpads, allowing for efficient data orchestration through software management and reducing energy consumption.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Productivity
If the scratchpad capacity is increased to reduce data replenishment frequency, then the productivity is improved, but the area of the on-chip memory is worsened
Solution Approach 1:
The patent transitions from a single on-chip scratchpad to a hierarchical memory architecture that includes off-chip auxiliary scratchpads, effectively adding a spatial dimension to the memory system. This allows the system to achieve larger effective scratchpad capacity without proportionally increasing the on-chip area, as the auxiliary scratchpads are positioned off-chip and accessed through a memory interface.
2Loss of time
If the scratchpad capacity is increased to reduce data replenishment frequency, then the loss of time is improved, but the area of the on-chip memory is worsened
Solution Approach 1:
By introducing off-chip auxiliary scratchpads and a memory interface, the system extends the scratchpad functionality into an additional spatial dimension. This hierarchical approach allows larger effective capacity that reduces data replenishment frequency and time, without requiring proportional increases in on-chip area.
3Speed
If the bandwidth is increased to improve data access speed, then the speed is improved, but the device complexity is worsened
Solution Approach 1:
The memory system is segmented into distinct components: on-chip main scratchpads for high-speed access, off-chip auxiliary scratchpads for additional capacity, and a memory interface for coordination. This segmentation allows each component to be optimized for its specific function, achieving high bandwidth where needed while managing overall system complexity through modular design.
Solution Approach 2:
The memory interface acts as an intermediary between the processing elements and the auxiliary scratchpads. It manages data transfer, address translation, and coordination, thereby simplifying the interaction complexity while enabling high-speed access to the extended memory capacity.
4Use of energy by moving object
If the on-chip scratchpad capacity is increased to reduce global memory access frequency, then the use of energy is improved, but the area of the on-chip memory is worsened
Solution Approach 1:
The system implements local quality by providing high-speed on-chip scratchpads for frequently accessed data near the processing elements, while using off-chip auxiliary scratchpads for less frequently accessed data. This localized high-performance memory approach minimizes energy-consuming global memory accesses while avoiding the need to expand the entire scratchpad system uniformly across the chip.
Data Source
Figure 1A
Figure 1B
Figure 2A
AI summary
A memory system, a method of assembling the memory system, and a computer system. The memory system includes a global memory device coupled to a plurality of processing elements. The global memory device is positioned external to a chip on which the plurality of processing devices reside. The memory system also includes at least one main scratchpad coupled to the at least one processing element of the plurality of processing devices and the global memory device. The memory system further includes a plurality of auxiliary scratchpads coupled to the plurality of processing elements and the global memory device. The one or more auxiliary scratchpads are configured to store static tensors. At least a portion of the plurality of auxiliary scratchpads are configured as a unitary multichannel device.