Neural Network Apparatus Zero Skipping Controller
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing neural network processors face inefficiencies in performing parallelized operations, particularly in determining shared operands between pixel values of input feature maps and weight values of kernels, which affects the speed and effectiveness of convolution operations.
Innovation Solution
The neural network apparatus configures processing units to perform parallelized operations based on determining shared operands as either pixel values or weight values of input feature maps and kernels, allowing for efficient parallel processing and zero skipping mechanisms to optimize hardware structure and reduce operation counts.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Productivity
If traditional neural network processors perform parallelized operations between pixel values and weight values, then convolution operations can be executed, but hardware complexity increases and processing efficiency decreases due to inability to skip zero operations
Solution Approach 1:
The patent extracts and separates zero value detection and skipping logic from the main parallelized operation pipeline. The controller independently identifies zero values in pixel values or weight values before they enter the processing units, allowing the processing units to skip unnecessary operations without requiring complex internal zero-detection circuitry, thus improving efficiency while maintaining hardware simplicity
Solution Approach 2:
The controller performs preliminary identification of zero values in pixel values or weight values before the parallelized operations begin. This preliminary action allows the system to pre-determine which processing units should be activated or deactivated, avoiding wasted computational resources and reducing hardware complexity by eliminating the need for runtime zero-detection in each processing unit
2Speed
If the neural network apparatus uses fixed parallel processing units for all operations, then hardware structure is simplified, but processing speed decreases when zero values are present in input data
Solution Approach 1:
The patent implements dynamic configuration of processing units based on the presence of zero values. The controller dynamically determines which processing units should be activated or deactivated based on zero value identification results, allowing the system to adapt its processing capacity to the actual data characteristics, thereby improving speed without requiring a permanently complex hardware structure
Solution Approach 2:
The processing units are designed with universal functionality to handle both zero and non-zero operations. The same processing units can be dynamically activated or deactivated based on the controller's determination, eliminating the need for separate specialized hardware paths for zero and non-zero operations, thus improving speed while maintaining hardware simplicity
3Loss of energy
If the apparatus performs operations on all pixel values and weight values without skipping zeros, then processing is simplified, but energy consumption increases and processing time is wasted
Solution Approach 1:
The controller implements a feedback mechanism where it receives information about zero values from the input data or previous processing stages, and uses this feedback to dynamically configure which processing units should be activated. This feedback loop allows the system to avoid unnecessary energy consumption on zero value operations while maintaining a relatively simple control structure
Solution Approach 2:
The system uses the inherent zero value information present in the neural network data itself to guide its own processing configuration. The zero values in the input data or weight values serve as natural signals for the controller to deactivate certain processing units, eliminating the need for external control signals or complex additional circuitry to identify optimization opportunities
Data Source
Figure 1
Figure 2A
Figure 2B
AI summary
A neural network apparatus includes one or more processors comprising: a controller configured to determine a shared operand to be shared in parallelized operations as being either one of a pixel value among pixel values of an input feature map and a weight value among weight values of a kernel, based on either one or both of a feature of the input feature map and a feature of the kernel; and one or more processing units configured to perform the parallelized operations based on the determined shared operand.