Interleaved Memory Processing Unit for Low Latency AI Compute
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional computing systems face challenges in reducing processing latency and power consumption due to the time-consuming transfer of large data sets between memory and processing units.
Innovation Solution
The design of a memory processing unit (MPU) with interleaved memory and processing regions, where compute cores are strategically coupled between adjacent memory blocks to optimize dataflow based on neural network models, core utilization, and memory reuse, enhancing bandwidth balance and reducing latency.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Productivity
If data is transferred from memory to processing units and back, then processing can be performed, but processing latency and power consumption increase
Solution Approach 1:
The patent merges memory and processing units into a single integrated chip architecture. Memory blocks are directly coupled to processing regions through interconnects, eliminating the need for separate data transfer paths. This integration allows data to be processed in-place within the memory structure itself, significantly reducing transfer latency and improving overall processing throughput.
Solution Approach 2:
The integrated chip is divided into multiple processing regions, each containing compute cores and local memory blocks. This segmentation allows different parts of the chip to process data independently and simultaneously, improving throughput while maintaining low latency through local data access within each region.
2Productivity
If data is transferred from memory to processing units and back, then processing can be performed, but power consumption increases
Solution Approach 1:
By merging memory and processing units on the same chip with direct coupling through interconnects, the patent eliminates repeated data transfers between memory and processing units. This integration reduces the energy consumption associated with data transfer operations while maintaining high processing throughput through efficient in-place computation.
3Speed
If core groups are coupled between adjacent memory blocks, then memory bandwidth is improved, but device complexity increases
Solution Approach 1:
The chip architecture is segmented into multiple processing regions with compute cores and memory blocks organized in a systematic pattern. Each processing region has local memory blocks that are directly coupled to the compute cores through dedicated interconnects, enabling high bandwidth memory access while keeping the overall architecture modular and manageable.
Solution Approach 2:
The patent implements local coupling between core groups and adjacent memory blocks within each processing region. This local quality approach ensures that each compute core has direct access to nearby memory blocks with high bandwidth, while the global architecture maintains a regular, repeating pattern that simplifies design and fabrication.
Data Source
AI summary
A memory processing unit (MPU) can include a first memory, a second memory, a plurality of processing regions and control logic. The first memory can include a plurality of memory regions. The plurality of memory regions can be organized in a plurality of memory blocks. The plurality of processing regions can be interleaved between the plurality of processing regions of the first memory. The plurality of processing regions can be organized in a plurality of core groups include a plurality of compute cores. The compute groups in the processing regions can be coupled to a plurality of adjacent memory blocks in the adjacent memory regions. The second memory can be coupled to the plurality of processing regions. The memory processing architectures can advantageously enable the design selection of the number of processing regions, the number of groups of compute cores, the number of compute cores, the number of compute clusters, the types of compute cores, the organization of memories, the types of memories, the size of the memories, the number of memory regions, the number of memory blocks and the like.


