Neural Network Processing Unit Hybrid Precision Computing

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Neural network computing faces challenges in balancing accuracy, power consumption, and data bandwidth, as existing technologies rely heavily on high-precision floating-point numbers that consume high power and data bandwidth.

Innovation Solution

A neural network processing unit that employs hybrid-precision and mixed-precision computing by using both floating-point and fixed-point number representations, with dedicated circuitry for conversion between these representations, allowing for configurable operations across layers to optimize accuracy, power consumption, and bandwidth.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If floating-point numbers with large bit-width are used for high accuracy, then measurement precision is improved, but use of energy and data bandwidth increase

Engineering Contradiction:
Improvecomputational accuracyVSAvoidpower consumption
Core Design Contradiction:
Measurement precisionVSUse of energy by moving object

Solution Approach 1:

The patent applies different number representations to different parts of the neural network computation process. Specifically, floating-point representation is used where high precision is critical (such as in certain layer computations), while fixed-point representation is used in other parts where lower precision is acceptable. This local differentiation allows the system to maintain necessary accuracy while reducing overall power consumption and data bandwidth requirements.

Inventive Principle:
Principle #3Local quality

Solution Approach 2:

The patent dynamically changes the precision parameter of number representations based on computational requirements. By converting between floating-point and fixed-point representations and selectively applying them to different layers or operations, the system adapts the precision level to match the actual computational needs, thereby optimizing the trade-off between accuracy and energy consumption.

Inventive Principle:
Principle #35Parameter changes

2Measurement precision

If floating-point numbers with large bit-width are used for high accuracy, then measurement precision is improved, but data bandwidth increases

Engineering Contradiction:
Improvecomputational accuracyVSAvoiddata bandwidth
Core Design Contradiction:
Measurement precisionVSQuantity of substance

Solution Approach 1:

The patent applies different number representations to different parts of the neural network computation process. Specifically, floating-point representation is used where high precision is critical (such as in certain layer computations), while fixed-point representation is used in other parts where lower precision is acceptable. This local differentiation allows the system to maintain necessary accuracy while reducing overall power consumption and data bandwidth requirements.

Inventive Principle:
Principle #3Local quality

Solution Approach 2:

The patent dynamically changes the precision parameter of number representations based on computational requirements. By converting between floating-point and fixed-point representations and selectively applying them to different layers or operations, the system adapts the precision level to match the actual computational needs, thereby optimizing the trade-off between accuracy and energy consumption.

Inventive Principle:
Principle #35Parameter changes

3Use of energy by moving object

If fixed-point number representation is used to reduce power consumption and data bandwidth, then use of energy and data bandwidth are reduced, but measurement precision deteriorates

Engineering Contradiction:
Improvepower consumptionVSAvoidcomputational accuracy
Core Design Contradiction:
Use of energy by moving objectVSMeasurement precision

Solution Approach 1:

The patent applies different number representations to different parts of the neural network computation process. Specifically, floating-point representation is used where high precision is critical (such as in certain layer computations), while fixed-point representation is used in other parts where lower precision is acceptable. This local differentiation allows the system to maintain necessary accuracy while reducing overall power consumption and data bandwidth requirements.

Inventive Principle:
Principle #3Local quality

Solution Approach 2:

The patent dynamically changes the precision parameter of number representations based on computational requirements. By converting between floating-point and fixed-point representations and selectively applying them to different layers or operations, the system adapts the precision level to match the actual computational needs, thereby optimizing the trade-off between accuracy and energy consumption.

Inventive Principle:
Principle #35Parameter changes

4Quantity of substance

If fixed-point number representation is used to reduce data bandwidth, then data bandwidth is reduced, but measurement precision deteriorates

Engineering Contradiction:
Improvedata bandwidthVSAvoidcomputational accuracy
Core Design Contradiction:
Quantity of substanceVSMeasurement precision

Solution Approach 1:

The patent applies different number representations to different parts of the neural network computation process. Specifically, floating-point representation is used where high precision is critical (such as in certain layer computations), while fixed-point representation is used in other parts where lower precision is acceptable. This local differentiation allows the system to maintain necessary accuracy while reducing overall power consumption and data bandwidth requirements.

Inventive Principle:
Principle #3Local quality

Solution Approach 2:

The patent dynamically changes the precision parameter of number representations based on computational requirements. By converting between floating-point and fixed-point representations and selectively applying them to different layers or operations, the system adapts the precision level to match the actual computational needs, thereby optimizing the trade-off between accuracy and energy consumption.

Inventive Principle:
Principle #35Parameter changes

Data Source

PatentUS20220156567A1Neural network processing unit for hybrid and mixed precision computing
Publication Date: 2022.05.19 MEDIATEK INC
  • US20220156567A1 patent drawing
  • US20220156567A1 patent drawing
  • US20220156567A1 patent drawing

AI summary

A neural network (NN) processing unit includes an operation circuit to perform tensor operations of a given layer of a neural network in one of a first number representation and a second number representation. The NN processing unit further includes a conversion circuit coupled to at least one of an input port and an output port of the operation circuit to convert between the first number representation and the second number representation. The first number representation is one of a fixed-point number representation and a floating-point number representation, and the second number representation is the other one of the fixed-point number representation and the floating-point number representation.