Runtime Configurable Memory Hierarchy for Neural Network Accelerators

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

In distributed edge computing, particularly in cognitive IoT domains, neural network accelerators face increased computational loads due to large memory requirements of complex neural networks, making it impractical to design chips with sufficient on-chip SRAM and necessitating a memory hierarchy that trades off between reliability, accuracy, and energy efficiency.

Innovation Solution

A closed-loop, runtime configurable technique that dynamically adjusts operational parameters of SRAM and DRAM modules in a neural network accelerator to optimize energy consumption while maintaining target accuracy by monitoring bit error rates and adjusting settings to ensure inference accuracy above a predetermined threshold.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Quantity of substance

If on-chip SRAM capacity is increased to meet large memory requirements of complex neural networks, then memory capacity is improved, but chip area and manufacturing complexity worsen

Engineering Contradiction:
Improvememory capacityVSAvoidchip design complexity
Core Design Contradiction:
Quantity of substanceVSDevice complexity

Solution Approach 1:

The patent divides the memory system into a hierarchy of multiple memory types (SRAM, DRAM, and potentially other memory levels) with different capacities and characteristics. Each memory level serves specific neural network layers, allowing the system to achieve large total memory capacity without requiring a single large SRAM block, thus reducing chip area complexity while meeting memory requirements.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent transitions from a single-dimensional memory capacity increase (larger SRAM) to a multi-dimensional memory hierarchy approach, utilizing different memory technologies at different levels. This dimensional shift allows the system to achieve equivalent or superior memory capacity with reduced SRAM area requirements and manufacturing complexity.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

2Use of energy by moving object

If memory operational parameters are adjusted to reduce energy consumption, then energy efficiency is improved, but bit error rate increases and accuracy deteriorates

Engineering Contradiction:
Improveenergy consumptionVSAvoidprocessing accuracy
Core Design Contradiction:
Use of energy by moving objectVSReliability

Solution Approach 1:

The patent implements dynamic adjustment of memory operational parameters based on runtime conditions. The system can adaptively change parameters such as refresh rates, voltage levels, and memory access patterns to optimize energy consumption while maintaining accuracy thresholds, rather than using static conservative settings.

Inventive Principle:
Principle #15Dynamics

Solution Approach 2:

The patent systematically varies memory operational parameters (frequency, voltage, refresh intervals) to identify optimal settings that balance energy consumption and accuracy. By exploring the parameter space and selecting configurations that meet minimum accuracy requirements, the system achieves energy efficiency without compromising processing quality.

Inventive Principle:
Principle #35Parameter changes

3Reliability

If memory operational parameters are adjusted to increase accuracy, then reliability is improved, but energy consumption increases

Engineering Contradiction:
Improveprocessing accuracyVSAvoidenergy consumption
Core Design Contradiction:
ReliabilityVSUse of energy by moving object

Solution Approach 1:

The patent applies partial action by adjusting memory parameters only to the extent necessary to meet minimum accuracy thresholds for different neural network layers. Rather than uniformly maximizing accuracy across all memory operations, the system applies accuracy-enhancing parameter adjustments only where and when needed, reducing overall energy consumption while maintaining sufficient reliability.

Inventive Principle:
Principle #16Partial or excessive action

4Use of energy by moving object

If a memory hierarchy with multiple memory types is implemented, then energy efficiency is improved through parameter tuning, but device complexity increases

Engineering Contradiction:
Improveenergy efficiencyVSAvoidmemory hierarchy complexity
Core Design Contradiction:
Use of energy by moving objectVSDevice complexity

Solution Approach 1:

The patent implements a memory controller that universally manages multiple memory types (SRAM, DRAM, and potentially other memory levels) through a unified interface and control logic. This multi-functional controller handles parameter adjustment, error monitoring, and memory allocation across different memory technologies, reducing the need for separate control mechanisms and managing complexity through consolidation.

Inventive Principle:
Principle #6Universality (Multi-functionality)

5Reliability

If bit error rate monitoring and dynamic parameter adjustment are implemented, then accuracy is maintained under varying conditions, but control system complexity increases

Engineering Contradiction:
Improveaccuracy maintenanceVSAvoidcontrol system complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The patent implements a feedback mechanism where bit error rates are continuously monitored and used to dynamically adjust memory operational parameters. The controller receives error rate information and automatically modifies parameters such as refresh rates or voltage levels to maintain accuracy thresholds, creating a self-regulating system that adapts to changing conditions without manual intervention.

Inventive Principle:
Principle #23Feedback

Data Source

PatentUS11334786B2System and method for an error-aware runtime configurable memory hierarchy for improved energy efficiency
Publication Date: 2022.05.17 INTERNATIONAL BUSINESS MACHINE CORPORATION
  • US11334786B2 patent drawing
  • US11334786B2 patent drawing
  • US11334786B2 patent drawing

AI summary

A method (and structure and computer product) to optimize an operation in a Neural Network Accelerator (NNAccel) that includes a hierarchy of neural network layers as computational stages for the NNAccel and a configurable hierarchy of memory modules including one or more on-chip Static Random-Access Memory (SRAM) modules and one or more Dynamic Random-Access Memory (DRAM) modules, where each memory module is controlled by a plurality of operational parameters that are adjustable by a controller of the NNAcc. The method includes detecting bit error rates of memory modules currently being used by the NNAccel and determining, by the controller, whether the detected bit error rates are sufficient for a predetermined threshold value for an accuracy of a processing of the NNAccel. One or more operational parameters of one or more memory modules are dynamically changed by the controller to move to a higher accuracy state when the accuracy is below the predetermined threshold value.