Distributed Neuromorphic Accelerator Bandwidth Reduction

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current neuromorphic architectures for artificial neural networks (ANNs) face significant bandwidth requirements, which are not feasible in low-power system-on-chip (SoC) architectures, leading to inefficient energy usage and limited implementation in power-constrained devices like mobile devices.

Innovation Solution

A distributed neuromorphic accelerator architecture with multiple processing units that maximizes data reuse, shares data among units, and optimizes bandwidth by operating at lower frequencies, reducing external memory accesses and interconnect width, thereby improving energy efficiency and area savings.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Productivity

If a monolithic neuromorphic accelerator architecture is used, then processing capability is improved, but bandwidth requirements and energy consumption increase significantly

Engineering Contradiction:
Improveprocessing capabilityVSAvoidenergy consumption
Core Design Contradiction:
ProductivityVSUse of energy by moving object

Solution Approach 1:

The patent divides the neuromorphic accelerator into multiple distributed processing units (clusters) that operate independently but cooperatively. Each cluster processes a portion of the neural network computation, allowing data to be reused locally without requiring high-bandwidth interconnects between centralized components. This segmentation reduces the total bandwidth requirement while maintaining overall processing capability.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent transitions from a centralized monolithic architecture to a distributed spatial arrangement of processing units. By organizing processors and memory across multiple dimensions (spatial distribution rather than centralized hierarchy), the system reduces interconnect bandwidth requirements while preserving computational throughput through parallel operation of distributed units.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

2Productivity

If a monolithic neuromorphic accelerator architecture is used, then processing capability is improved, but area usage increases

Engineering Contradiction:
Improveprocessing capabilityVSAvoidarea usage
Core Design Contradiction:
ProductivityVSArea of stationary object

Solution Approach 1:

The patent segments the large monolithic accelerator into smaller distributed processing clusters that can be spatially distributed across the chip. Each cluster contains its own processing units and local memory, reducing the need for large centralized interconnect structures and allowing more efficient utilization of chip area through parallel distributed computation.

Inventive Principle:
Principle #1Segmentation

3Productivity

If data is shared among multiple processing units, then data reuse is maximized, but interconnect complexity increases

Engineering Contradiction:
Improvedata reuse efficiencyVSAvoidinterconnect complexity
Core Design Contradiction:
ProductivityVSDevice complexity

Solution Approach 1:

The patent merges the interconnect resources across multiple processing clusters into a shared communication fabric. Instead of providing dedicated point-to-point connections between each processing unit and memory, the system uses a unified interconnect structure that all clusters access, reducing overall interconnect complexity while enabling efficient data sharing and reuse across the distributed system.

Inventive Principle:
Principle #5Merging (Combining)

Data Source

PatentUS9971540B2Storage device and method for performing convolution operations
Publication Date: 2018.05.15 INTEL CORP
  • US9971540B2 patent drawing
  • US9971540B2 patent drawing
  • US9971540B2 patent drawing

AI summary

A storage device and method are described for performing convolution operations. For example, one embodiment of an apparatus to perform convolution operations comprises a plurality of processing units to execute convolution operations on input data and partial results; a unified scratchpad memory comprising a plurality of memory banks communicatively coupled to the plurality of processing units through a plurality of read/write ports, each of the plurality of memory banks partitioned to store both the input data and partial results; a control unit to allocate the input data and partial results to the memory banks to ensure a minimum quality of service in accordance with the specified number of read/write ports and the specified convolution operation to be performed.