Neural Network Accelerator Feedback Paths for Dynamic Command Execution
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Dynamic neural networks are challenging to execute efficiently in hardware due to the need for host intervention to determine segment command streams, leading to inefficient use of resources.
Innovation Solution
Incorporating an embedded processor with a hardware feedback path into neural network accelerators, allowing the command decoder to dynamically determine the next command stream based on branch commands, enabling conditional or unconditional branching without host intervention.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If host intervention is used to determine segment command streams, then flexibility in executing dynamic neural networks is maintained, but resource efficiency and processing speed deteriorate
Solution Approach 1:
The neural network accelerator executes dynamic neural networks autonomously by determining the next command stream based on branch commands without requiring host intervention. The embedded processor and command decoder work together to select and execute appropriate command streams, enabling the system to serve itself rather than relying on continuous host control.
Solution Approach 2:
Command streams are pre-loaded into the neural network accelerator's memory before execution. The system determines which pre-loaded command stream to execute next based on branch conditions, eliminating the need for host intervention during the execution process and improving processing speed.
2Adaptability or versatility
If host intervention is required for branch commands, then control flexibility is maintained, but hardware resource usage increases
Solution Approach 1:
The embedded processor and command decoder enable the neural network accelerator to autonomously determine the next command stream based on branch commands, eliminating the need for host intervention and reducing hardware resource usage while maintaining control flexibility.
Solution Approach 2:
An embedded processor is introduced as an intermediary component between the command decoder and the execution units. This embedded processor handles the complex decision-making for branch commands locally within the accelerator, reducing the burden on external host hardware resources.
3Device complexity
If dynamic neural networks are executed without embedded processor, then hardware simplicity is maintained, but execution efficiency and adaptability deteriorate
Solution Approach 1:
The embedded processor serves multiple functions: it determines the next command stream based on branch conditions, coordinates command stream execution, and enables adaptive behavior. This multi-functionality improves execution efficiency without requiring separate dedicated components for each function.
Solution Approach 2:
The system changes the operational parameters by introducing an embedded processor that can dynamically select command streams based on runtime conditions. This parameter change enables efficient execution of dynamic neural networks while keeping the overall hardware architecture relatively simple.
Data Source
AI summary
Neural network accelerators with one or more neural network accelerator cores. Each neural network accelerator core has hardware accelerators configured to accelerate neural network operations, an embedded processor, a command decoder, and a hardware feedback path between the embedded processor and the command decoder. The command decoder is configured to control the hardware accelerators and the embedded processor of that core in accordance with commands of a command stream, and when the command stream comprises a set of one or more branch commands that indicate a conditional branch is to be performed, cause the embedded processor to determine a next command stream, and in response to receiving information from the embedded processor identifying the next command stream via the hardware feedback path, control the one or more hardware accelerators and the embedded processor in accordance with commands of the next command stream.


