Neural Network Arithmetic Processing Device Parallel MAC Units
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current neural network arithmetic processing devices face inefficiencies in multiply-accumulate operations, particularly in increasing speed and efficiency while minimizing circuit scale, especially when applying arithmetic functions across multiple layers, as they require complex temporal designs and increased circuit complexity for parallelization or pipelining.
Innovation Solution
The proposed solution involves a neural network arithmetic processing device with multiple multiply-accumulate units operating in parallel, where each unit includes memories for input variables and weight data, multipliers, adders, and output units, with parallel execution of arithmetic operations across layers, utilizing ring buffer memories to simplify design and reduce circuit scale, and activation function units for enhanced processing.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Productivity
If parallelization or pipelining is implemented to increase speed and efficiency of multiply-accumulate operations, then processing speed improves, but circuit scale and design complexity increase
Solution Approach 1:
The neural network arithmetic processing device is divided into multiple independent multiply-accumulate units (MAC units), each capable of performing arithmetic operations autonomously. This segmentation allows parallel execution of operations across multiple layers without requiring complex inter-unit communication, thereby improving processing speed while maintaining manageable circuit scale through modular design
Solution Approach 2:
The patent transitions from sequential layer-by-layer processing to simultaneous multi-layer processing by introducing a temporal dimension for parallel execution. Multiple MAC units operate on different layers at the same time, effectively adding a time-parallelism dimension that increases productivity without proportionally increasing spatial circuit complexity
2Reliability
If the number of layers of the neural network is increased to improve performance, then neural network performance improves, but arithmetic operation time and circuit scale increase
Solution Approach 1:
The neural network is segmented into multiple layers that can be processed simultaneously by dedicated MAC units. Each MAC unit handles a specific layer's computations independently, allowing deep neural networks with many layers to achieve high performance without linearly increasing total arithmetic operation time, as multiple layers progress in parallel rather than sequentially
Solution Approach 2:
The patent ensures continuous utilization of arithmetic resources by maintaining multiple active MAC units operating on different layers simultaneously. This continuity eliminates idle time between layer completions and keeps the computational pipeline full, thereby reducing total arithmetic operation time while supporting increased network depth for improved performance
Data Source
AI summary
A neural network arithmetic processing device is capable of implementing a further increase in speed and efficiency of multiply-accumulate arithmetic operation, suppressing an increase in circuit scale, and performing multiply-accumulate arithmetic operation with simple design. A neural network arithmetic processing device includes a first multiply-accumulate arithmetic unit, a register connected to the first multiply-accumulate arithmetic unit, and a second multiply-accumulate arithmetic unit connected to the register. The first multiply-accumulate arithmetic unit has a first memory, a second memory, a first multiplier, a first adder, and a first output unit. The second multiply-accumulate arithmetic unit has an input unit, a third memory, second multipliers, second adders, and second output units.


