Floating-Point Neural Network Arithmetic With Reduced-Width Align-Shifting
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Neural network devices face challenges in efficiently processing complex input data in real-time due to limited resources, necessitating a technology that maximizes performance while reducing operations, particularly in low-power systems like smartphones.
Innovation Solution
A neural network device incorporating a floating point arithmetic circuit that performs dot-product operations by align-shifting fraction part multiplying results based on exponent part adding operations, using a reduced shiftable bit width align-shifter to minimize power consumption and hardware usage.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If a neural network device performs complex dot-product operations on floating point data pairs, then processing accuracy is improved, but power consumption and hardware resource usage increase
Solution Approach 1:
The patent changes the parameter of bit width for align-shifting operations. Instead of using a fixed large bit width, the system dynamically determines the required bit width based on the actual data characteristics and operation requirements. This parameter optimization reduces the hardware resources allocated to the align-shifter while maintaining the necessary processing accuracy for floating point dot-product operations.
Solution Approach 2:
The patent applies local quality by making the align-shifter's bit width adaptable rather than uniform. The system determines the minimum necessary bit width for each specific operation context, allocating resources only where needed. This localized optimization ensures that the align-shifter uses exactly the resources required for accurate processing without excessive power consumption from oversized hardware components.
2Device complexity
If a neural network device uses a reduced bit width align-shifter, then hardware resource usage is reduced, but processing capability may be compromised
Solution Approach 1:
The patent implements dynamics by making the align-shifter's bit width adaptive rather than static. The system dynamically determines the appropriate bit width based on the characteristics of the floating point data being processed and the specific operation requirements. This dynamic adaptation ensures that the hardware resources are optimized for each operation, maintaining full processing capability when needed while reducing complexity when possible.
Solution Approach 2:
The system changes the operational parameter of bit width based on real-time analysis of the data and operation requirements. By calculating the necessary precision for each dot-product operation and adjusting the align-shifter bit width accordingly, the system maintains processing capability while minimizing hardware resource usage. The parameter adjustment ensures that no processing capability is lost despite the reduced bit width.
3Speed
If a neural network device processes floating point data in real-time, then responsiveness is improved, but the amount of operations required increases
Solution Approach 1:
The patent optimizes the operational parameters of the align-shifter to reduce the number of operations required. By determining the minimum necessary bit width for each operation rather than using a conservative fixed value, the system reduces redundant processing steps. This parameter optimization allows real-time processing of floating point data with fewer operations, improving responsiveness while reducing computational complexity.
Data Source
AI summary
A neural network device for performing a neural network operation includes a floating point arithmetic circuit configured to perform a dot-product operation for each of a plurality of floating point data pairs, wherein the floating point arithmetic circuit is configured to, in the dot-product operation, align-shift a plurality of fraction part multiplying operation results respectively corresponding to the floating point data pairs based on a maximum value determined from a plurality of exponent part adding operation results respectively corresponding to the floating point data pairs.


