Memory-Centric Neural Network Accelerator Architecture

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current neural network hardware accelerators face challenges in achieving optimal performance and energy efficiency, particularly in mobile devices, due to von Neumann architecture limitations and memory bandwidth constraints, which hinder their scalability with increasing dataset sizes and task complexities.

Innovation Solution

A memory-centric neural network hardware accelerator architecture that includes a processing unit, semiconductor memory devices, a weight matrix constructed with rows and columns of memory cells, timestamp registers, and a lookup table for updating weights based on adjusting values, enabling efficient data processing and online learning capabilities.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Productivity

If von Neumann-based GPU architecture is used, then computation performance is high, but memory bandwidth cannot scale with increasing dataset size and task complexity

Engineering Contradiction:
Improvecomputation performanceVSAvoidmemory bandwidth scalability
Core Design Contradiction:
ProductivityVSQuantity of substance

Solution Approach 1:

The system segments memory operations into separate functional units: weight memory, input memory, output memory, and dedicated memory controllers for each type. This segmentation allows parallel access paths and eliminates the bottleneck of unified memory architecture, enabling memory bandwidth to scale independently with dataset size and task complexity.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent introduces intermediary buffers and memory controllers that mediate between the processing units and memory arrays. These intermediaries facilitate efficient data movement and caching strategies, allowing the system to handle large datasets without proportionally increasing memory bandwidth requirements.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Speed

If traditional neural network hardware accelerator is used, then inference speed is improved, but power consumption increases

Engineering Contradiction:
Improveinference speedVSAvoidpower consumption
Core Design Contradiction:
SpeedVSUse of energy by moving object

Solution Approach 1:

The system implements local weight storage in dedicated weight memory units close to the processing elements, eliminating the need for repeated global memory accesses. This local caching strategy maintains fast inference speed while significantly reducing power consumption by minimizing energy-intensive memory transactions.

Inventive Principle:
Principle #3Local quality

Solution Approach 2:

The patent employs periodic weight updates and batched inference operations that allow the system to enter low-power states between computational tasks. By processing data in batches and updating weights periodically rather than continuously, the system maintains inference performance while reducing average power consumption.

Inventive Principle:
Principle #19Periodic action

3Adaptability or versatility

If CPU is used for neural network computations, then flexibility is maintained, but performance and efficiency decrease

Engineering Contradiction:
ImproveflexibilityVSAvoidcomputation efficiency
Core Design Contradiction:
Adaptability or versatilityVSProductivity

Solution Approach 1:

The patent designs a universal neural network processing architecture that can handle multiple network types (CNN, RNN, Transformer) and operations (inference, training, fine-tuning) through a single hardware platform. The programmable processing units and configurable memory architecture provide both the flexibility of software and the efficiency of dedicated hardware.

Inventive Principle:
Principle #6Universality (Multi-functionality)

4Loss of time

If memory-centric architecture is implemented, then real-time processing is enabled, but device complexity increases

Engineering Contradiction:
Improveprocessing latencyVSAvoidarchitecture complexity
Core Design Contradiction:
Loss of timeVSDevice complexity

Solution Approach 1:

The system employs a nested memory hierarchy with multiple levels of caching: fast on-chip weight memory nested within processing units, intermediate buffers nested within memory controllers, and larger off-chip memory for bulk storage. This nested structure enables real-time processing by keeping frequently accessed data in inner layers while maintaining manageable complexity through hierarchical organization.

Inventive Principle:
Principle #7Nested doll (Nesting)

Data Source

PatentUS11501131B2Neural network hardware accelerator architectures and operating method thereof
Publication Date: 2022.11.15 SK HYNIX INC
  • US11501131B2 patent drawing
  • US11501131B2 patent drawing
  • US11501131B2 patent drawing

AI summary

A memory-centric neural network system and operating method thereof includes: a processing unit; semiconductor memory devices coupled to the processing unit, the semiconductor memory devices containing instructions executed by the processing unit; a weight matrix constructed with rows and columns of memory cells, inputs of the memory cells of a same row being connected to one of axons, outputs of the memory cells of a same column being connected to one of neurons; timestamp registers registering timestamps of the axons and the neurons; and a lookup table containing adjusting values indexed in accordance with the timestamps, wherein the processing unit updates the weight matrix in accordance with the adjusting values.