Neural Network Backpropagation Using Discrete Data and Parallel Processing
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current methods for supporting backpropagation in multilayer artificial neural networks, such as using general-purpose processors or GPUs, face performance bottlenecks due to high power consumption and inefficient data handling, particularly with continuous data representation requiring extensive decoding and off-chip data movement.
Innovation Solution
An apparatus and method utilizing a direct memory access unit to exchange discrete data for multilayer neural network operations, with a master computation module calculating input gradient vectors and slave computation modules parallelly calculating output vectors, reducing the need for extensive decoding and minimizing power consumption by using on-chip caching.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If general-purpose processors are used to execute backpropagation algorithms, then flexibility and programmability are improved, but operational performance and power efficiency deteriorate due to sequential execution and high decoding overhead
Solution Approach 1:
The patent replaces the general-purpose processor's sequential instruction decoding and execution mechanism with a dedicated neural network processing architecture that directly executes neural network operations through specialized hardware circuits, eliminating the need for software interpretation and significantly improving operational performance
Solution Approach 2:
The patent changes the data representation parameter from continuous floating-point numbers to discrete values, which simplifies the computational operations and enables more efficient hardware implementation, thereby improving both operational performance and power efficiency
2Measurement precision
If continuous data representation is used to store floating-point numbers, then precision is improved, but storage space and computational resource consumption increase
Solution Approach 1:
The patent changes the data representation parameter from continuous floating-point format to discrete quantized values, reducing the number of bits required to represent each weight value and activation, thereby decreasing storage space requirements and computational resource consumption while maintaining acceptable precision
Solution Approach 2:
The patent uses simplified discrete data representations that require less complex hardware to process and store, trading off some precision for significant reductions in storage space and computational complexity, making the system more efficient for neural network operations
3Productivity
If GPUs are used to execute neural network operations, then parallel processing capability is improved, but power consumption increases due to repeated off-chip data movement
Solution Approach 1:
The patent implements an on-chip caching hierarchy that nests multiple levels of memory within the processing unit, allowing frequently accessed neural network data to be stored closer to the computation units, thereby reducing the need for repeated off-chip data movement and lowering power consumption
Solution Approach 2:
The patent introduces on-chip cache memory as an intermediary between the external memory and the processing units, acting as a buffer that stores weight values and activation data locally, thus reducing the frequency and volume of data transfers to and from off-chip memory and decreasing power consumption
Data Source
AI summary
Aspects for backpropagation of a multilayer neural network (MNN) in a neural network processor are described herein. The aspects may include a computation module configured to receive one or more groups of MNN data. The computation module may further include a master computation module configured to calculate an input gradient vector based on a first output gradient vector from an adjacent layer and based on a data type of each of the one or more groups of MNN data. Further still, the computation module may include one or more slave computation modules configured to parallelly calculate portions of a second output vector based on the input gradient vector calculated by the master computation module and based on the data type of each of the one or more groups of MNN data.


