Distributed Neuromorphic Accelerator Bandwidth Reduction
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current neuromorphic architectures for artificial neural networks (ANNs) face significant bandwidth requirements, which are not feasible in low-power system-on-chip (SoC) architectures, leading to inefficient energy usage and limited implementation in power-constrained devices like mobile devices.
Innovation Solution
A distributed neuromorphic accelerator architecture with multiple processing units that maximizes data reuse, shares data among units, and optimizes bandwidth by operating at lower frequencies, reducing external memory accesses and interconnect width, thereby improving energy efficiency and area savings.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Productivity
If a monolithic neuromorphic accelerator architecture is used, then processing capability is improved, but bandwidth requirements and energy consumption increase significantly
Solution Approach 1:
The patent divides the neuromorphic accelerator into multiple distributed processing units (clusters) that operate independently but cooperatively. Each cluster processes a portion of the neural network computation, allowing data to be reused locally without requiring high-bandwidth interconnects between centralized components. This segmentation reduces the total bandwidth requirement while maintaining overall processing capability.
Solution Approach 2:
The patent transitions from a centralized monolithic architecture to a distributed spatial arrangement of processing units. By organizing processors and memory across multiple dimensions (spatial distribution rather than centralized hierarchy), the system reduces interconnect bandwidth requirements while preserving computational throughput through parallel operation of distributed units.
2Productivity
If a monolithic neuromorphic accelerator architecture is used, then processing capability is improved, but area usage increases
Solution Approach 1:
The patent segments the large monolithic accelerator into smaller distributed processing clusters that can be spatially distributed across the chip. Each cluster contains its own processing units and local memory, reducing the need for large centralized interconnect structures and allowing more efficient utilization of chip area through parallel distributed computation.
3Productivity
If data is shared among multiple processing units, then data reuse is maximized, but interconnect complexity increases
Solution Approach 1:
The patent merges the interconnect resources across multiple processing clusters into a shared communication fabric. Instead of providing dedicated point-to-point connections between each processing unit and memory, the system uses a unified interconnect structure that all clusters access, reducing overall interconnect complexity while enabling efficient data sharing and reuse across the distributed system.
Data Source
AI summary
A storage device and method are described for performing convolution operations. For example, one embodiment of an apparatus to perform convolution operations comprises a plurality of processing units to execute convolution operations on input data and partial results; a unified scratchpad memory comprising a plurality of memory banks communicatively coupled to the plurality of processing units through a plurality of read/write ports, each of the plurality of memory banks partitioned to store both the input data and partial results; a control unit to allocate the input data and partial results to the memory banks to ensure a minimum quality of service in accordance with the specified number of read/write ports and the specified convolution operation to be performed.


