Interleaved MPU Regions with Inter-Layer Communication for Low Latency
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional computing systems face challenges in reducing processing latency and power consumption due to the time-consuming transfer of large data volumes between memory and processing units.
Innovation Solution
A memory processing unit (MPU) architecture with interleaved processing regions and inter-layer communication (ILC) modules that facilitate efficient data flow and synchronization between compute cores and memory regions, reducing the need for off-chip data communications.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Productivity
If data is transferred from memory to processing units and back, then computations can be performed, but processing latency and power consumption increase
Solution Approach 1:
The patent merges memory and processing units into a unified memory processing unit (MPU) architecture where processing regions are interleaved with memory regions. This integration eliminates separate data transfer stages, allowing computations to be performed directly within the memory structure and reducing the time required to move data between components.
Solution Approach 2:
The patent introduces a new spatial dimension by interleaving processing regions between memory regions in a three-dimensional stacked architecture. This vertical integration creates direct proximity connections between memory and processing units, transforming the traditional sequential data flow into a parallel, in-situ computation model that reduces latency.
2Productivity
If data is transferred from memory to processing units and back, then computations can be performed, but power consumption increases
Solution Approach 1:
The patent merges memory and processing units into a unified memory processing unit (MPU) architecture where processing regions are interleaved with memory regions. This integration eliminates separate data transfer stages, allowing computations to be performed directly within the memory structure and reducing the time required to move data between components.
Solution Approach 2:
The patent extracts the data transfer function from the traditional von Neumann architecture by implementing in-memory computing capabilities. Processing operations are taken out of separate processing units and embedded directly within memory regions, eliminating the energy-intensive data movement between distinct memory and processing components.
3Productivity
If compute cores access shared buffers in memory regions, then data reuse is enhanced, but synchronization complexity increases
Solution Approach 1:
The patent introduces inter-layer communication (ILC) modules as intermediary components that manage communication between compute cores accessing shared buffers. These ILC modules handle synchronization commands and coordinate access to shared memory regions, reducing the complexity burden on individual compute cores while enabling efficient data reuse across multiple processing operations.
Solution Approach 2:
The patent implements synchronization mechanisms where compute cores issue synchronization commands that are tracked and managed by the system. This feedback loop allows the architecture to coordinate access to shared buffers, ensuring data consistency while enabling multiple compute cores to efficiently reuse data from shared memory regions.
Data Source
AI summary
A memory processing unit (MPU) can include a first memory, a second memory, a plurality of processing regions and control logic. The first memory can include a plurality of regions. The plurality of processing regions can be interleaved between the plurality of regions of the first memory. The processing regions can include a plurality of compute cores. The second memory can be coupled to the plurality of processing regions. The control logic can configure data flow between compute cores of one or more of the processing regions and corresponding adjacent regions of the first memory. The control logic can also configure data flow between the second memory and the compute cores of one or more of the processing regions. The control logic can also configure data flow between compute cores within one or more respective ones of the processing regions.


