Integrated Memory Coprocessor Bypassing Cache Latency
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current network processing units (NPUs) face latency issues due to the need for data transfer between on-die caches and off-chip memory, which complicates packet processing and increases power consumption, especially when handling high-bandwidth network traffic.
Innovation Solution
An integrated main memory and coprocessor chip (MMCC) design that eliminates the need for caching data from off-chip resources, allowing for direct on-chip processing and storage, reducing latency and power consumption by bypassing external memory access and cache coherence overhead.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Loss of time
If data is stored in off-chip memory and accessed through cache, then memory capacity is sufficient, but latency is increased due to multiple access steps
Solution Approach 1:
The patent merges the main memory and coprocessor onto a single chip, eliminating the need for separate off-chip memory and cache structures. This integration allows direct access between the coprocessor and memory without requiring multiple access steps through cache, thereby reducing latency while maintaining sufficient memory capacity.
Solution Approach 2:
The patent extracts the cache layer from the memory access path by providing direct coprocessor access to main memory. This removes the intermediate caching step that contributes to latency, allowing the coprocessor to access memory directly without the complexity of cache management.
2Use of energy by moving object
If multiple processors share external memory, then data sharing is enabled, but power consumption increases due to repeated external access
Solution Approach 1:
By integrating multiple coprocessors and memory on the same chip, the patent enables them to share the integrated memory structure without requiring repeated external memory access. This sharing mechanism reduces power consumption while maintaining high data processing throughput through on-chip communication pathways.
3Speed
If cache is used to store frequently accessed data, then access speed improves, but cache coherence overhead increases
Solution Approach 1:
The patent removes the cache coherence management overhead by eliminating the cache layer entirely. Instead of managing cache coherence between multiple processors, the system provides direct access to the integrated memory structure, achieving fast data access without the complexity of cache coherence protocols.
4Loss of time
If data is transferred between chip interfaces for processing, then processing capability is sufficient, but latency is incurred due to framing and transmission
Solution Approach 1:
The patent merges the coprocessor and memory onto the same chip, eliminating the need for data transfer between separate chip interfaces. This integration removes the framing and transmission steps that contribute to latency, enabling high packet processing rates through direct on-chip data access and processing.
Data Source
AI summary
System, method, and apparatus for integrated main memory (MM) and configurable coprocessor (CP) chip for processing subset of network functions. Chip supports external accesses to MM without additional latency from on-chip CP. On-chip memory scheduler resolves all bank conflicts and configurably load balances MM accesses. Instruction set and data on which the CP executes instructions are all disposed on-chip with no on-chip cache memory, thereby avoiding latency and coherency issues. Multiple independent and orthogonal threading domains used: a FIFO-based scheduling domain (SD) for the I/O; a multi-threaded processing domain for the CP. The CP is an array of independent, autonomous, unsequenced processing engines processing on-chip data tracked by SD of external CMD and reordered per FIFO CMD sequence before transmission. Paired I/O ports tied to unique global on-chip SD allow multiple external processors to slave chip and its resources independently and autonomously without scheduling between the external processors.


