Heterogeneous Multiply-Accumulate Unit Array for DNN Accelerator

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional deep neural network (DNN) accelerators face limitations in supply voltage scaling due to timing errors in multiply-accumulate units, which restrict power consumption reduction without compromising DNN accuracy.

Innovation Solution

A deep neural network accelerator design featuring a heterogeneous unit array with operational units of varying sizes, proportional to their cumulative importance values, along with an activator, weighting unit, and accumulator, which allows for different weight mappings and activation propagation paths to optimize DNN operations.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Loss of energy

If supply voltage scaling is applied to reduce power consumption, then power consumption decreases, but timing errors occur in multiply-accumulate units limiting further voltage reduction

Engineering Contradiction:
Improvepower consumptionVSAvoidtiming error probability
Core Design Contradiction:
Loss of energyVSReliability

Solution Approach 1:

The patent applies local quality by creating heterogeneous multiply-accumulate units with different hardware sizes within the same array. Each unit's size is optimized based on its specific computational importance, allowing critical units to operate reliably at lower voltages while less critical units can tolerate higher voltages or have reduced size, thus resolving the contradiction between power consumption and timing error probability.

Inventive Principle:
Principle #3Local quality

Solution Approach 2:

The patent changes the parameter of multiply-accumulate unit size to address the timing error issue. By varying the hardware size of individual units based on cumulative importance values, the system can adjust the critical path delay of each unit, thereby controlling timing error probability while enabling overall power reduction through voltage scaling.

Inventive Principle:
Principle #35Parameter changes

2Device complexity

If single-sized multiply-accumulate units are used, then device structure is simple, but supply voltage scaling is limited due to timing errors

Engineering Contradiction:
Improvemultiply-accumulate unit structureVSAvoidpower consumption reduction capability
Core Design Contradiction:
Device complexityVSLoss of energy

Solution Approach 1:

The patent transitions from uniform-sized units to heterogeneous units with locally optimized sizes. Each multiply-accumulate unit's hardware size is determined by its cumulative importance value, creating local variations in structure that enable better power efficiency while maintaining overall system functionality.

Inventive Principle:
Principle #3Local quality

Solution Approach 2:

The patent introduces dynamic weight mapping that can reconfigure which weights are assigned to which multiply-accumulate units based on their sizes and computational importance. This dynamic allocation allows the system to optimize performance and power consumption adaptively, overcoming the limitations of static single-sized unit designs.

Inventive Principle:
Principle #15Dynamics

3Reliability

If larger multiply-accumulate units are used, then timing error probability decreases due to shorter critical path delay, but device complexity and power consumption increase

Engineering Contradiction:
Improvetiming error probabilityVSAvoidpower consumption
Core Design Contradiction:
ReliabilityVSLoss of energy

Solution Approach 1:

The patent applies local quality by assigning different sizes to multiply-accumulate units based on their specific computational importance. Rather than uniformly increasing all unit sizes to reduce timing errors, only the necessary units are enlarged, minimizing the overall power consumption increase while achieving the required timing reliability for critical operations.

Inventive Principle:
Principle #3Local quality

Solution Approach 2:

The patent uses partial action by applying size optimization only to the extent necessary for each unit's computational importance. Not all units are enlarged to the maximum size; instead, each unit receives just enough hardware resources to achieve acceptable timing performance, avoiding excessive power consumption in units that don't require full optimization.

Inventive Principle:
Principle #16Partial or excessive action

Data Source

PatentUS12182689B2Deep neural network accelerator using heterogeneous multiply-accumulate unit
Publication Date: 2024.12.31 KOREA UNIV RES & BUSINESS FOUND
  • US12182689B2 patent drawing
  • US12182689B2 patent drawing
  • US12182689B2 patent drawing

AI summary

A deep neural network accelerator includes a unit array including a first sub-array including a first operational unit and a second sub-array including a second operational unit. The first and second operational units have different sizes from each other, the sizes of the first and second operational units are in proportion to each cumulative importance value accumulated in each operational unit of the unit array while performing a deep neural network operation, and the each cumulative importance value is obtained by accumulating an importance for each weight mapped to the each operational unit of the unit array.