Reconfigurable Systolic Neural Network Engine for Training and Inference
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional processors are inefficient for performing neural network computations due to high processing time and power consumption, and existing hardware accelerators like GPUs are costly and limited to inference tasks only.
Innovation Solution
A special-purpose hardware accelerator, the systolic neural network engine, uses a systolic array with data processing units connected in a local region to perform computations during both training and inference, supporting reconfiguration for various neural network architectures.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Productivity
If conventional processors are used for neural network computations, then device complexity is low and versatility is high, but processing time is excessive and power consumption is high
Solution Approach 1:
The patent replaces conventional von Neumann architecture processors with a systolic array architecture where data flows continuously through processing elements in a pipelined manner, eliminating the need for repeated memory access cycles and enabling parallel computation of neural network operations
Solution Approach 2:
The neural network computation is divided into multiple processing elements arranged in a systolic array, where each element performs a specific computation on a portion of the data, enabling parallel processing and reducing overall computation time
2Productivity
If GPUs are used as hardware accelerators, then processing speed improves, but device cost increases and adaptability is limited to inference tasks only
Solution Approach 1:
The systolic array architecture is designed to be reconfigurable, allowing the same hardware to dynamically adapt between training and inference modes by changing data flow patterns and computation operations, providing both speed and versatility
Solution Approach 2:
The patent designs a universal systolic array platform that can perform both forward propagation (inference) and backward propagation (training) operations, as well as support different neural network architectures, eliminating the need for separate hardware for different tasks
3Productivity
If systolic array architecture is implemented, then processing efficiency improves, but device complexity increases
Solution Approach 1:
The patent combines multiple functions (data storage, computation, and data transfer) into integrated processing elements within the systolic array, reducing the need for separate components and interconnects while maintaining high computation efficiency
Data Source
AI summary
Some embodiments include a special-purpose hardware accelerator that can perform specialized machine learning tasks during both training and inference stages. For example, this hardware accelerator uses a systolic array having a number of data processing units (“DPUs”) that are each connected to a small number of other DPUs in a local region. Data from the many nodes of a neural network is pulsed through these DPUs with associated tags that identify where such data was originated or processed, such that each DPU has knowledge of where incoming data originated and thus is able to compute the data as specified by the architecture of the neural network. These tags enable the systolic neural network engine to perform computations during backpropagation, such that the systolic neural network engine is able to support training.


