Operation Unit for Neural Network Bit Width Configuration
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing neural network operations face challenges in efficiently processing data with different bit widths, leading to increased power consumption and hardware area requirements due to the need for multiple processors and operators.
Innovation Solution
An operation unit and method that determine the bit width of operation data and either use a matching operator directly or combine lower-bit width operators to create a new operator with the necessary bit width, allowing for efficient neural network, matrix, and vector operations while reducing the number of operation units and hardware area.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If multiple processors and operators with different bit widths are used to handle operation data of different bit widths, then the neural network operations can be performed accurately, but the hardware area and device complexity increase
Solution Approach 1:
The patent applies universality by designing a single operator that can handle multiple bit widths through configurable parameters. The operator is designed to be multi-functional, accepting operation data of different bit widths (e.g., 8-bit, 16-bit, 32-bit) without requiring separate dedicated operators for each bit width, thereby reducing hardware area while maintaining accuracy
Solution Approach 2:
The patent uses parameter changes by introducing a bit width configuration parameter that dynamically adjusts the operator's behavior. The operator can change its effective bit width based on the input data characteristics, allowing it to adapt between different precision requirements without physical reconfiguration, thus resolving the contradiction between accuracy and hardware area
2Measurement precision
If multiple processors and operators with different bit widths are used to handle operation data of different bit widths, then the neural network operations can be performed accurately, but the device complexity and number of processors increase
Solution Approach 1:
The patent applies universality by designing a single operator that can handle multiple bit widths through configurable parameters. The operator is designed to be multi-functional, accepting operation data of different bit widths (e.g., 8-bit, 16-bit, 32-bit) without requiring separate dedicated operators for each bit width, thereby reducing hardware area while maintaining accuracy
Solution Approach 2:
The patent applies merging by combining multiple bit width handling capabilities into a single operator entity. Instead of having separate processors for different bit widths, the patent merges these functions into one unified operator that can be configured dynamically, thus reducing the total number of processors and simplifying device architecture
3Area of stationary object
If low-bit width operators are combined to form high-bit width operators, then the number of processors and hardware area are reduced, but the operation complexity increases
Solution Approach 1:
The patent uses parameter changes by introducing a bit width configuration parameter that dynamically adjusts the operator's behavior. The operator can change its effective bit width based on the input data characteristics, allowing it to adapt between different precision requirements without physical reconfiguration, thus resolving the contradiction between accuracy and hardware area
Data Source
Figure 1~3
Figure 4~6
Figure 7~9
AI summary
The present invention provides an operation unit, an operation method and an operation device, configures the bit width of the operation data through configuring the bit width of the configuration instruction; when the operation is performed according to the instruction, it determines whether there is an operator with the same bit width as the operation data, and if so, the operand is transmitted directly to the corresponding operator, otherwise, an operator combination strategy is generated and a plurality of operators are combined to a new operator according to the operator combination strategy so that the bit width of the new operator matches the bit width of the operand and the operand is transmitted to the new operator; then, the operator that obtains the operand is commanded to perform a neural network operation/matrix operation/vector operation. The present invention can support the operation of the operation data with different bit-widths to realize efficient neural network operations, matrix operations and vector operations, and at the same time, save the number of operation units and reduce the hardware area.