Neural Network Layer Fusion to Cut I/O Overhead on AI Chips

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

The increasing number of layers and parameters in neural networks leads to significant on-chip and off-chip input/output accesses, consuming resources and delaying operation time, necessitating a mechanism to reduce these overheads.

Innovation Solution

An integrated circuit apparatus and method for neural network computing that includes a template fuse unit, compiler, linker, and computing apparatus to dynamically fuse multiple layers, reducing input/output overheads by loading data required for computing at a time.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Power

If the number of layers and parameters in neural networks is increased to improve computing capability, then the computing power is improved, but the on-chip and off-chip input/output accesses increase, consuming more resources and delaying operation time

Engineering Contradiction:
Improvecomputing powerVSAvoidoperation time
Core Design Contradiction:
PowerVSLoss of time

Solution Approach 1:

The patent merges multiple adjacent layers of the neural network into a single fused layer, allowing multiple computing operations to be performed in one execution cycle. This reduces the number of separate input/output operations between memory and computing units, thereby decreasing the time lost to data transfer overhead while maintaining the cumulative computing power of all fused layers

Inventive Principle:
Principle #5Merging (Combining)

2Power

If the number of layers and parameters in neural networks is increased to improve computing capability, then the computing power is improved, but the resource consumption increases due to frequent input/output accesses

Engineering Contradiction:
Improvecomputing powerVSAvoidresource consumption
Core Design Contradiction:
PowerVSUse of energy by moving object

Solution Approach 1:

By fusing multiple layers into one, the patent reduces the frequency of memory access operations. Each fused layer processes multiple operations internally without requiring repeated data loading from off-chip memory, thereby reducing energy consumption associated with memory bandwidth usage and I/O operations while preserving the total computing capability

Inventive Principle:
Principle #5Merging (Combining)

3Ease of operation

If multiple layers are processed separately to maintain modularity, then the ease of operation is improved, but the input/output overheads increase and computational efficiency decreases

Engineering Contradiction:
ImprovemodularityVSAvoidcomputational efficiency
Core Design Contradiction:
Ease of operationVSProductivity

Solution Approach 1:

The patent introduces a dynamic layer fusion mechanism that can adaptively combine multiple layers based on computational requirements and resource constraints. This allows the system to switch between fused and non-fused modes, maintaining operational flexibility and modularity while achieving high computational efficiency when fusion is applied to reduce I/O overhead

Inventive Principle:
Principle #15Dynamics

Data Source

PatentUS12547902B2Device and method for neural network computing, and board and readable storage medium
Publication Date: 2026.02.10 CAMBRICON TECH CO LTD
  • US12547902B2 patent drawing
  • US12547902B2 patent drawing
  • US12547902B2 patent drawing

AI summary

The present disclosure relates to an apparatus and a method for performing neural network computing, a board card, and a readable storage medium. The computing apparatus of the present disclosure is included in an integrated circuit apparatus. The integrated circuit apparatus includes a general interconnection interface and other processing apparatus. The computing apparatus interacts with other processing apparatus to jointly complete a computing operation specified by a user. The integrated circuit apparatus further includes a storage apparatus. The storage apparatus is connected to the computing apparatus and other processing apparatus, respectively. The storage apparatus is used for data storage of the computing apparatus and other processing apparatus.