Heterogeneous Multiply-Accumulate Unit Array for DNN Accelerator
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional deep neural network (DNN) accelerators face limitations in supply voltage scaling due to timing errors in multiply-accumulate units, which restrict power consumption reduction without compromising DNN accuracy.
Innovation Solution
A deep neural network accelerator design featuring a heterogeneous unit array with operational units of varying sizes, proportional to their cumulative importance values, along with an activator, weighting unit, and accumulator, which allows for different weight mappings and activation propagation paths to optimize DNN operations.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Loss of energy
If supply voltage scaling is applied to reduce power consumption, then power consumption decreases, but timing errors occur in multiply-accumulate units limiting further voltage reduction
Solution Approach 1:
The patent applies local quality by creating heterogeneous multiply-accumulate units with different hardware sizes within the same array. Each unit's size is optimized based on its specific computational importance, allowing critical units to operate reliably at lower voltages while less critical units can tolerate higher voltages or have reduced size, thus resolving the contradiction between power consumption and timing error probability.
Solution Approach 2:
The patent changes the parameter of multiply-accumulate unit size to address the timing error issue. By varying the hardware size of individual units based on cumulative importance values, the system can adjust the critical path delay of each unit, thereby controlling timing error probability while enabling overall power reduction through voltage scaling.
2Device complexity
If single-sized multiply-accumulate units are used, then device structure is simple, but supply voltage scaling is limited due to timing errors
Solution Approach 1:
The patent transitions from uniform-sized units to heterogeneous units with locally optimized sizes. Each multiply-accumulate unit's hardware size is determined by its cumulative importance value, creating local variations in structure that enable better power efficiency while maintaining overall system functionality.
Solution Approach 2:
The patent introduces dynamic weight mapping that can reconfigure which weights are assigned to which multiply-accumulate units based on their sizes and computational importance. This dynamic allocation allows the system to optimize performance and power consumption adaptively, overcoming the limitations of static single-sized unit designs.
3Reliability
If larger multiply-accumulate units are used, then timing error probability decreases due to shorter critical path delay, but device complexity and power consumption increase
Solution Approach 1:
The patent applies local quality by assigning different sizes to multiply-accumulate units based on their specific computational importance. Rather than uniformly increasing all unit sizes to reduce timing errors, only the necessary units are enlarged, minimizing the overall power consumption increase while achieving the required timing reliability for critical operations.
Solution Approach 2:
The patent uses partial action by applying size optimization only to the extent necessary for each unit's computational importance. Not all units are enlarged to the maximum size; instead, each unit receives just enough hardware resources to achieve acceptable timing performance, avoiding excessive power consumption in units that don't require full optimization.
Data Source
AI summary
A deep neural network accelerator includes a unit array including a first sub-array including a first operational unit and a second sub-array including a second operational unit. The first and second operational units have different sizes from each other, the sizes of the first and second operational units are in proportion to each cumulative importance value accumulated in each operational unit of the unit array while performing a deep neural network operation, and the each cumulative importance value is obtained by accumulating an importance for each weight mapped to the each operational unit of the unit array.


