Runtime Configurable Memory Hierarchy for Neural Network Accelerators
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
In distributed edge computing, particularly in cognitive IoT domains, neural network accelerators face increased computational loads due to large memory requirements of complex neural networks, making it impractical to design chips with sufficient on-chip SRAM and necessitating a memory hierarchy that trades off between reliability, accuracy, and energy efficiency.
Innovation Solution
A closed-loop, runtime configurable technique that dynamically adjusts operational parameters of SRAM and DRAM modules in a neural network accelerator to optimize energy consumption while maintaining target accuracy by monitoring bit error rates and adjusting settings to ensure inference accuracy above a predetermined threshold.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Quantity of substance
If on-chip SRAM capacity is increased to meet large memory requirements of complex neural networks, then memory capacity is improved, but chip area and manufacturing complexity worsen
Solution Approach 1:
The patent divides the memory system into a hierarchy of multiple memory types (SRAM, DRAM, and potentially other memory levels) with different capacities and characteristics. Each memory level serves specific neural network layers, allowing the system to achieve large total memory capacity without requiring a single large SRAM block, thus reducing chip area complexity while meeting memory requirements.
Solution Approach 2:
The patent transitions from a single-dimensional memory capacity increase (larger SRAM) to a multi-dimensional memory hierarchy approach, utilizing different memory technologies at different levels. This dimensional shift allows the system to achieve equivalent or superior memory capacity with reduced SRAM area requirements and manufacturing complexity.
2Use of energy by moving object
If memory operational parameters are adjusted to reduce energy consumption, then energy efficiency is improved, but bit error rate increases and accuracy deteriorates
Solution Approach 1:
The patent implements dynamic adjustment of memory operational parameters based on runtime conditions. The system can adaptively change parameters such as refresh rates, voltage levels, and memory access patterns to optimize energy consumption while maintaining accuracy thresholds, rather than using static conservative settings.
Solution Approach 2:
The patent systematically varies memory operational parameters (frequency, voltage, refresh intervals) to identify optimal settings that balance energy consumption and accuracy. By exploring the parameter space and selecting configurations that meet minimum accuracy requirements, the system achieves energy efficiency without compromising processing quality.
3Reliability
If memory operational parameters are adjusted to increase accuracy, then reliability is improved, but energy consumption increases
Solution Approach 1:
The patent applies partial action by adjusting memory parameters only to the extent necessary to meet minimum accuracy thresholds for different neural network layers. Rather than uniformly maximizing accuracy across all memory operations, the system applies accuracy-enhancing parameter adjustments only where and when needed, reducing overall energy consumption while maintaining sufficient reliability.
4Use of energy by moving object
If a memory hierarchy with multiple memory types is implemented, then energy efficiency is improved through parameter tuning, but device complexity increases
Solution Approach 1:
The patent implements a memory controller that universally manages multiple memory types (SRAM, DRAM, and potentially other memory levels) through a unified interface and control logic. This multi-functional controller handles parameter adjustment, error monitoring, and memory allocation across different memory technologies, reducing the need for separate control mechanisms and managing complexity through consolidation.
5Reliability
If bit error rate monitoring and dynamic parameter adjustment are implemented, then accuracy is maintained under varying conditions, but control system complexity increases
Solution Approach 1:
The patent implements a feedback mechanism where bit error rates are continuously monitored and used to dynamically adjust memory operational parameters. The controller receives error rate information and automatically modifies parameters such as refresh rates or voltage levels to maintain accuracy thresholds, creating a self-regulating system that adapts to changing conditions without manual intervention.
Data Source
AI summary
A method (and structure and computer product) to optimize an operation in a Neural Network Accelerator (NNAccel) that includes a hierarchy of neural network layers as computational stages for the NNAccel and a configurable hierarchy of memory modules including one or more on-chip Static Random-Access Memory (SRAM) modules and one or more Dynamic Random-Access Memory (DRAM) modules, where each memory module is controlled by a plurality of operational parameters that are adjustable by a controller of the NNAcc. The method includes detecting bit error rates of memory modules currently being used by the NNAccel and determining, by the controller, whether the detected bit error rates are sufficient for a predetermined threshold value for an accuracy of a processing of the NNAccel. One or more operational parameters of one or more memory modules are dynamically changed by the controller to move to a higher accuracy state when the accuracy is below the predetermined threshold value.


