TULIP BNN ASIC Using Programmable Threshold Logic Cells

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current machine learning technologies, particularly deep neural networks (DNNs), face significant challenges in achieving energy efficiency and performance due to high on-chip storage requirements, leading to substantial energy and delay penalties, despite efforts like weight pruning, quantization, and binary neural networks (BNNs), which still require efficient hardware implementations.

Innovation Solution

A configurable binary neural network (BNN) application-specific integrated circuit (ASIC) using a network of programmable threshold logic standard cells, referred to as TULIP, is designed with unique processing elements that perform inner product and thresholding operations natively, allowing for optimal scheduling and energy-efficient execution of BNN operations.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Productivity

If conventional DNNs are deployed on hardware platforms, then computation performance can be achieved, but on-chip storage requirements become excessively high, leading to large energy and delay penalties

Engineering Contradiction:
Improvecomputation performanceVSAvoidenergy consumption
Core Design Contradiction:
ProductivityVSUse of energy by moving object

Solution Approach 1:

The patent applies quantization by changing the parameter representation from full-precision floating-point to reduced-precision fixed-point, and further to binary values. This parameter transformation drastically reduces storage requirements and enables energy-efficient computation while maintaining acceptable accuracy for neural network operations

Inventive Principle:
Principle #35Parameter changes

Solution Approach 2:

The patent replaces conventional arithmetic logic units with dedicated binary neural network processing elements that use simplified logic circuits. This substitution of computational mechanics with specialized hardware architecture eliminates the need for complex arithmetic operations, reducing energy consumption and increasing computation speed

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

2Productivity

If on-chip storage is increased to meet DNN requirements, then computation performance improves, but device area and manufacturing cost increase

Engineering Contradiction:
Improvecomputation performanceVSAvoiddevice area
Core Design Contradiction:
ProductivityVSArea of stationary object

Solution Approach 1:

By transforming weight and activation parameters from high-precision representations to binary values, the patent reduces the storage footprint of neural network models by one to two orders of magnitude, enabling deployment on devices with limited on-chip memory while maintaining computation performance

Inventive Principle:
Principle #35Parameter changes

Solution Approach 2:

The patent segments the neural network computation into discrete binary operations that can be processed efficiently by dedicated hardware elements. This segmentation allows for compact implementation of processing units that require minimal on-chip storage for intermediate results

Inventive Principle:
Principle #1Segmentation

3Use of energy by moving object

If binary neural networks are used to reduce storage requirements, then energy consumption decreases, but hardware implementation complexity increases

Engineering Contradiction:
Improveenergy consumptionVSAvoidhardware implementation complexity
Core Design Contradiction:
Use of energy by moving objectVSDevice complexity

Solution Approach 1:

The patent designs a universal binary neural network processing element that can execute multiple neural network operations (matrix multiplication, convolution, activation functions) using the same hardware structure. This multi-functionality reduces overall hardware complexity compared to implementing separate dedicated circuits for each operation type

Inventive Principle:
Principle #6Universality (Multi-functionality)

Solution Approach 2:

The patent implements reconfigurable processing elements that can dynamically adjust their operation mode and parameters to match the specific requirements of different neural network layers and operations, optimizing performance while maintaining hardware efficiency

Inventive Principle:
Principle #15Dynamics

Data Source

PatentUS20220121915A1Configurable BNN ASIC using a network of programmable threshold logic standard cells
Publication Date: 2022.04.21 THE ARIZONA BOARD OF REGENTS ON BEHALF OF THE UNIV OF ARIZONA
  • US20220121915A1 patent drawing
  • US20220121915A1 patent drawing
  • US20220121915A1 patent drawing

AI summary

A configurable binary neural network (BNN) application-specific integrated circuit (ASIC) using a network of programmable threshold logic standard cells is provided. A new architecture is presented for a BNN that uses an optimal schedule for executing the operations of an arbitrary BNN. This architecture, also referred to herein as TULIP, is designed with the goal of maximizing energy efficiency per classification. At the top-level, TULIP consists of a collection of unique processing elements (TULIP-PEs) that are organized in a single instruction, multiple data (SIMD) fashion. Each TULIP-PE consists of a small network of binary neurons, and a small amount of local memory per neuron. Novel algorithms are presented herein for mapping arbitrary nodes of a BNN onto the TULIP-PEs. Comparison results show that TULIP is consistently 3× more energy-efficient than conventional designs, without any penalty in performance, area, or accuracy.