Neural Network Processing Unit Hybrid Precision Computing
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Neural network computing faces challenges in balancing accuracy, power consumption, and data bandwidth, as existing technologies rely heavily on high-precision floating-point numbers that consume high power and data bandwidth.
Innovation Solution
A neural network processing unit that employs hybrid-precision and mixed-precision computing by using both floating-point and fixed-point number representations, with dedicated circuitry for conversion between these representations, allowing for configurable operations across layers to optimize accuracy, power consumption, and bandwidth.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If floating-point numbers with large bit-width are used for high accuracy, then measurement precision is improved, but use of energy and data bandwidth increase
Solution Approach 1:
The patent applies different number representations to different parts of the neural network computation process. Specifically, floating-point representation is used where high precision is critical (such as in certain layer computations), while fixed-point representation is used in other parts where lower precision is acceptable. This local differentiation allows the system to maintain necessary accuracy while reducing overall power consumption and data bandwidth requirements.
Solution Approach 2:
The patent dynamically changes the precision parameter of number representations based on computational requirements. By converting between floating-point and fixed-point representations and selectively applying them to different layers or operations, the system adapts the precision level to match the actual computational needs, thereby optimizing the trade-off between accuracy and energy consumption.
2Measurement precision
If floating-point numbers with large bit-width are used for high accuracy, then measurement precision is improved, but data bandwidth increases
Solution Approach 1:
The patent applies different number representations to different parts of the neural network computation process. Specifically, floating-point representation is used where high precision is critical (such as in certain layer computations), while fixed-point representation is used in other parts where lower precision is acceptable. This local differentiation allows the system to maintain necessary accuracy while reducing overall power consumption and data bandwidth requirements.
Solution Approach 2:
The patent dynamically changes the precision parameter of number representations based on computational requirements. By converting between floating-point and fixed-point representations and selectively applying them to different layers or operations, the system adapts the precision level to match the actual computational needs, thereby optimizing the trade-off between accuracy and energy consumption.
3Use of energy by moving object
If fixed-point number representation is used to reduce power consumption and data bandwidth, then use of energy and data bandwidth are reduced, but measurement precision deteriorates
Solution Approach 1:
The patent applies different number representations to different parts of the neural network computation process. Specifically, floating-point representation is used where high precision is critical (such as in certain layer computations), while fixed-point representation is used in other parts where lower precision is acceptable. This local differentiation allows the system to maintain necessary accuracy while reducing overall power consumption and data bandwidth requirements.
Solution Approach 2:
The patent dynamically changes the precision parameter of number representations based on computational requirements. By converting between floating-point and fixed-point representations and selectively applying them to different layers or operations, the system adapts the precision level to match the actual computational needs, thereby optimizing the trade-off between accuracy and energy consumption.
4Quantity of substance
If fixed-point number representation is used to reduce data bandwidth, then data bandwidth is reduced, but measurement precision deteriorates
Solution Approach 1:
The patent applies different number representations to different parts of the neural network computation process. Specifically, floating-point representation is used where high precision is critical (such as in certain layer computations), while fixed-point representation is used in other parts where lower precision is acceptable. This local differentiation allows the system to maintain necessary accuracy while reducing overall power consumption and data bandwidth requirements.
Solution Approach 2:
The patent dynamically changes the precision parameter of number representations based on computational requirements. By converting between floating-point and fixed-point representations and selectively applying them to different layers or operations, the system adapts the precision level to match the actual computational needs, thereby optimizing the trade-off between accuracy and energy consumption.
Data Source
AI summary
A neural network (NN) processing unit includes an operation circuit to perform tensor operations of a given layer of a neural network in one of a first number representation and a second number representation. The NN processing unit further includes a conversion circuit coupled to at least one of an input port and an output port of the operation circuit to convert between the first number representation and the second number representation. The first number representation is one of a fixed-point number representation and a floating-point number representation, and the second number representation is the other one of the fixed-point number representation and the floating-point number representation.


