Memory-Centric Neural Network Accelerator Architecture
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current neural network hardware accelerators face challenges in achieving optimal performance and energy efficiency, particularly in mobile devices, due to von Neumann architecture limitations and memory bandwidth constraints, which hinder their scalability with increasing dataset sizes and task complexities.
Innovation Solution
A memory-centric neural network hardware accelerator architecture that includes a processing unit, semiconductor memory devices, a weight matrix constructed with rows and columns of memory cells, timestamp registers, and a lookup table for updating weights based on adjusting values, enabling efficient data processing and online learning capabilities.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Productivity
If von Neumann-based GPU architecture is used, then computation performance is high, but memory bandwidth cannot scale with increasing dataset size and task complexity
Solution Approach 1:
The system segments memory operations into separate functional units: weight memory, input memory, output memory, and dedicated memory controllers for each type. This segmentation allows parallel access paths and eliminates the bottleneck of unified memory architecture, enabling memory bandwidth to scale independently with dataset size and task complexity.
Solution Approach 2:
The patent introduces intermediary buffers and memory controllers that mediate between the processing units and memory arrays. These intermediaries facilitate efficient data movement and caching strategies, allowing the system to handle large datasets without proportionally increasing memory bandwidth requirements.
2Speed
If traditional neural network hardware accelerator is used, then inference speed is improved, but power consumption increases
Solution Approach 1:
The system implements local weight storage in dedicated weight memory units close to the processing elements, eliminating the need for repeated global memory accesses. This local caching strategy maintains fast inference speed while significantly reducing power consumption by minimizing energy-intensive memory transactions.
Solution Approach 2:
The patent employs periodic weight updates and batched inference operations that allow the system to enter low-power states between computational tasks. By processing data in batches and updating weights periodically rather than continuously, the system maintains inference performance while reducing average power consumption.
3Adaptability or versatility
If CPU is used for neural network computations, then flexibility is maintained, but performance and efficiency decrease
Solution Approach 1:
The patent designs a universal neural network processing architecture that can handle multiple network types (CNN, RNN, Transformer) and operations (inference, training, fine-tuning) through a single hardware platform. The programmable processing units and configurable memory architecture provide both the flexibility of software and the efficiency of dedicated hardware.
4Loss of time
If memory-centric architecture is implemented, then real-time processing is enabled, but device complexity increases
Solution Approach 1:
The system employs a nested memory hierarchy with multiple levels of caching: fast on-chip weight memory nested within processing units, intermediate buffers nested within memory controllers, and larger off-chip memory for bulk storage. This nested structure enables real-time processing by keeping frequently accessed data in inner layers while maintaining manageable complexity through hierarchical organization.
Data Source
AI summary
A memory-centric neural network system and operating method thereof includes: a processing unit; semiconductor memory devices coupled to the processing unit, the semiconductor memory devices containing instructions executed by the processing unit; a weight matrix constructed with rows and columns of memory cells, inputs of the memory cells of a same row being connected to one of axons, outputs of the memory cells of a same column being connected to one of neurons; timestamp registers registering timestamps of the axons and the neurons; and a lookup table containing adjusting values indexed in accordance with the timestamps, wherein the processing unit updates the weight matrix in accordance with the adjusting values.


