Interleaved Memory Processing Unit Architecture for Lower Data Latency
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional computing systems experience high processing latency and power consumption due to the time-consuming transfer of large data volumes between memory and processing units.
Innovation Solution
A memory processing unit (MPU) architecture with interleaved processing regions and memories, including near memory compute cores and arithmetic compute cores, optimized for dataflow and synchronization, reduces data transfer latency and power consumption by integrating compute functions closer to memory.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Productivity
If data is transferred from memory to processing units and back, then computations can be performed, but processing latency and data latency increase
Solution Approach 1:
The patent merges memory and processing units into a unified memory processing unit (MPU) architecture where processing regions are interleaved between memory regions. This integration allows compute cores to access memory directly without external data transfer, eliminating the time-consuming data movement between separate memory and processing components.
Solution Approach 2:
The patent introduces near-memory compute cores as intermediary processing elements positioned between memory and traditional processing units. These compute cores perform preliminary computations close to the data source, reducing the volume of data that needs to be transferred to main processing units and thereby reducing overall processing latency.
2Productivity
If data is transferred from memory to processing units, then computations can be performed, but power consumption increases
Solution Approach 1:
The patent merges memory and processing units into a unified memory processing unit (MPU) architecture where processing regions are interleaved between memory regions. This integration allows compute cores to access memory directly without external data transfer, eliminating the power-consuming data movement between separate memory and processing components.
Solution Approach 2:
The patent introduces near-memory compute cores as intermediary processing elements positioned between memory and traditional processing units. These compute cores perform preliminary computations close to the data source, reducing the volume of data that needs to be transferred to main processing units and thereby reducing overall power consumption associated with data transfer.
3Loss of time
If processing regions are interleaved between memory regions, then data transfer latency is reduced, but device complexity increases
Solution Approach 1:
The patent segments the MPU into multiple processing regions interleaved with memory regions, each containing specialized compute cores. This segmentation allows independent operation of different processing regions, reducing inter-dependency and simplifying control logic while achieving low latency through parallel access to nearby memory regions.
Solution Approach 2:
The patent implements local quality by providing each processing region with its own near-memory compute cores and associated resources. This localized architecture allows each region to operate semi-independently with dedicated resources, reducing the complexity of global resource management while maintaining low data transfer latency through local access.
Data Source
AI summary
A memory processing unit (MPU) can include a first memory, a second memory, a plurality of processing regions and control logic. The first memory can include a plurality of regions. The plurality of processing regions can be interleaved between the plurality of regions of the first memory. The processing regions can include a plurality of compute cores. The second memory can be coupled to the plurality of processing regions. The control logic can configure data flow between compute cores of one or more of the processing regions and corresponding adjacent regions of the first memory. The control logic can also configure data flow between the second memory and the compute cores of one or more of the processing regions. The control logic can also configure data flow between compute cores within one or more respective ones of the processing regions. The control logic can also configure array data for storage memory of the MPU.


