Dynamic Routing for Neural Network Accelerators
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current deep learning technologies face challenges in achieving improvements in accuracy, performance, and energy efficiency, particularly in the context of neural network training and inference.
Innovation Solution
The implementation of a deep learning accelerator system that utilizes an array of processing elements with routers, enabling flow-based computations on wavelets. This system includes enhanced instruction set architectures, wavelet filtering, and dynamic routing techniques to optimize neural network processing.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Ease of manufacture
If static routing is used in neural network accelerators, then device complexity is reduced and ease of manufacture is improved, but adaptability to different neural network architectures and dynamic workloads deteriorates
Solution Approach 1:
The patent implements dynamic routing that can adapt to different neural network architectures and workloads by allowing data to be routed through multiple possible paths based on runtime conditions. The routing configuration can be modified dynamically to optimize performance for different computation patterns, resolving the contradiction between ease of manufacture and adaptability.
Solution Approach 2:
The routing infrastructure is designed to support multiple routing modes (static and dynamic) within the same hardware platform. This universal routing system can handle different neural network types, layer configurations, and data flow patterns, providing both manufacturing simplicity and architectural versatility simultaneously.
2Productivity
If more processing elements are added to increase computational throughput, then productivity is improved, but device complexity and energy consumption increase
Solution Approach 1:
The patent divides the neural network computation into discrete wavelets that can be processed independently by multiple processing elements. This segmentation allows the system to scale productivity by adding more PEs without proportionally increasing complexity, as each PE handles a standardized unit of work.
Solution Approach 2:
Multiple processing elements are merged into a coordinated array that shares common routing infrastructure and memory resources. This merging approach increases productivity through parallel processing while controlling device complexity by sharing resources across PEs rather than duplicating full processing stacks.
3Adaptability or versatility
If dynamic routing is implemented to optimize data flow, then adaptability and performance are improved, but device complexity and routing overhead increase
Solution Approach 1:
The patent performs preliminary routing configuration based on neural network architecture characteristics before execution. Routing paths are pre-computed and cached, allowing dynamic adaptation to different network types without requiring complex real-time routing decisions, thus reducing routing overhead and device complexity.
Solution Approach 2:
The patent introduces routing configuration structures and control logic as intermediaries between the hardware fabric and neural network workloads. These intermediaries manage the complexity of dynamic routing by providing an abstraction layer that simplifies path selection and reduces the burden on individual processing elements.
4Use of energy by moving object
If wavelet filtering is applied to reduce data transmission, then energy efficiency is improved, but processing time and computational overhead increase
Solution Approach 1:
The patent applies wavelet filtering selectively to reduce data transmission only when it provides energy efficiency benefits. The filtering intensity and application scope are adjusted based on workload characteristics, achieving energy savings without unnecessarily increasing processing time for all operations.
Solution Approach 2:
Wavelet filtering is applied locally at processing elements based on their specific needs and the characteristics of incoming data. This localized approach allows energy-efficient filtering where beneficial while minimizing processing overhead by avoiding unnecessary filtering in other parts of the system.
Data Source
AI summary
Techniques in dynamic routing for advanced deep learning provide improvements in one or more of accuracy, performance, and energy efficiency. An array of processing elements comprising a portion of a neural network accelerator performs flow-based computations on wavelets of data. Each processing element comprises a compute element enabled to execute programmed instructions using the data and a router enabled to route the wavelets via static routing, dynamic routing, or both. The routing is in accordance with a respective virtual channel specifier of each of the wavelets and controlled by routing configuration information of the router. The static techniques enable statically specifiable neuron connections. The dynamic techniques enable information from the wavelets to alter the routing configuration information during neural network processing.


