Neural Network Layer Fusion to Cut I/O Overhead on AI Chips
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
The increasing number of layers and parameters in neural networks leads to significant on-chip and off-chip input/output accesses, consuming resources and delaying operation time, necessitating a mechanism to reduce these overheads.
Innovation Solution
An integrated circuit apparatus and method for neural network computing that includes a template fuse unit, compiler, linker, and computing apparatus to dynamically fuse multiple layers, reducing input/output overheads by loading data required for computing at a time.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Power
If the number of layers and parameters in neural networks is increased to improve computing capability, then the computing power is improved, but the on-chip and off-chip input/output accesses increase, consuming more resources and delaying operation time
Solution Approach 1:
The patent merges multiple adjacent layers of the neural network into a single fused layer, allowing multiple computing operations to be performed in one execution cycle. This reduces the number of separate input/output operations between memory and computing units, thereby decreasing the time lost to data transfer overhead while maintaining the cumulative computing power of all fused layers
2Power
If the number of layers and parameters in neural networks is increased to improve computing capability, then the computing power is improved, but the resource consumption increases due to frequent input/output accesses
Solution Approach 1:
By fusing multiple layers into one, the patent reduces the frequency of memory access operations. Each fused layer processes multiple operations internally without requiring repeated data loading from off-chip memory, thereby reducing energy consumption associated with memory bandwidth usage and I/O operations while preserving the total computing capability
3Ease of operation
If multiple layers are processed separately to maintain modularity, then the ease of operation is improved, but the input/output overheads increase and computational efficiency decreases
Solution Approach 1:
The patent introduces a dynamic layer fusion mechanism that can adaptively combine multiple layers based on computational requirements and resource constraints. This allows the system to switch between fused and non-fused modes, maintaining operational flexibility and modularity while achieving high computational efficiency when fusion is applied to reduce I/O overhead
Data Source
AI summary
The present disclosure relates to an apparatus and a method for performing neural network computing, a board card, and a readable storage medium. The computing apparatus of the present disclosure is included in an integrated circuit apparatus. The integrated circuit apparatus includes a general interconnection interface and other processing apparatus. The computing apparatus interacts with other processing apparatus to jointly complete a computing operation specified by a user. The integrated circuit apparatus further includes a storage apparatus. The storage apparatus is connected to the computing apparatus and other processing apparatus, respectively. The storage apparatus is used for data storage of the computing apparatus and other processing apparatus.


