Smart Memory Tiles for Packet Processing Bandwidth
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Network systems face performance bottlenecks due to internal memory bandwidth and chip/memory I/O bandwidth limitations, especially in packet processing applications that require large memory access and high random-access bandwidth, leading to latency issues and power consumption concerns.
Innovation Solution
A smart memory architecture comprising memory tiles with integrated processing elements and an interconnection network that enables local computation and efficient data streaming, reducing latency and bandwidth demands by creating static communication paths between tiles for sequential data processing.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Quantity of substance
If large memory structures are used to support high packet rates, then memory capacity increases, but memory bandwidth and I/O bandwidth become bottlenecks
Solution Approach 1:
The memory system is divided into multiple memory tiles, each with its own processing element. This segmentation allows parallel access to different memory regions, effectively increasing the overall memory bandwidth without requiring a single large memory structure. Each tile can be accessed independently, enabling simultaneous data retrieval across multiple tiles.
Solution Approach 2:
The patent introduces a new dimension of processing by integrating processing elements directly within memory tiles. This transforms the traditional single-dimension memory access model into a multi-dimensional architecture where computation can occur at the memory location itself, reducing the need for high-speed memory bandwidth while maintaining processing capability.
2Adaptability or versatility
If random access to memory locations is performed repeatedly, then data processing flexibility increases, but latency increases
Solution Approach 1:
Each memory tile includes an integrated processing element that can perform computations directly on data stored in that tile without requiring retrieval to external processors. This self-service capability eliminates the latency associated with random access to external memory while maintaining the flexibility to process data in any order.
Solution Approach 2:
The patent merges storage and processing functions into a unified memory tile structure. By combining memory cells with processing elements in the same tile, the system eliminates the separation between data retrieval and computation, thereby reducing latency while preserving processing flexibility.
3Productivity
If data is moved repeatedly between memory and processing units, then data processing capability increases, but power consumption increases
Solution Approach 1:
Processing elements within each memory tile perform computations directly on stored data without requiring data to be moved to external processing units. This eliminates the energy-consuming data transfer operations while maintaining full data processing capability, as each tile is self-sufficient for both storage and computation.
Solution Approach 2:
By merging memory and processing functions into the same tile structure, the patent eliminates the need for repeated data movement between separate memory and processing units. This integration significantly reduces power consumption associated with data transfer while preserving enhanced data processing capability.
Data Source
Figure 1
Figure 2
Figure 3
AI summary
An apparatus comprising a storage device comprising a plurality of memory tiles each comprising a memory block and a processing element, and an interconnection network coupled to the storage device and configured to interconnect the memory tiles, wherein the processing elements are configured to perform at least one packet processing feature, and wherein the interconnection network is configured to promote communication between the memory tiles. Also disclosed is a network component comprising a receiver configured to receive network data, a logic unit configured to convert the network data for suitable deterministic memory caching and processing, a serial input/output (I/O) interface configured to forward the converted network data in a serialized manner, a memory comprising a plurality of memory tiles configured to store and process the converted network data from the serial I/O interface, and a transmitter configured to forward the processed network data from the serial I/O interface.