Neural Network Forward Fusion Template Fuse Unit

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

The increasing number of layers and parameters in neural networks leads to significant on-chip and off-chip input/output accesses, consuming resources and delaying operation times, necessitating a mechanism to reduce these overheads.

Innovation Solution

An integrated circuit apparatus and method for forward fusion of a neural network, which creates a template fuse unit to perform neural network computing, reducing input/output overheads by fusing adjacent layers and optimizing data transfer between on-chip and off-chip memory.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If the number of layers and parameters in neural network is increased, then the computing capability and accuracy are improved, but the on-chip and off-chip input/output accesses increase, consuming more resources and delaying operation time

Engineering Contradiction:
Improveneural network computing capabilityVSAvoidoperation time
Core Design Contradiction:
ReliabilityVSLoss of time

Solution Approach 1:

The patent merges adjacent layers into a template fuse unit, combining multiple sequential operations into a single integrated computing unit. This reduces the number of separate input/output accesses between layers, thereby decreasing the time spent on data transfer while maintaining the computational capability of the original multi-layer structure.

Inventive Principle:
Principle #5Merging (Combining)

2Reliability

If the number of layers and parameters in neural network is increased, then the computing capability and accuracy are improved, but the on-chip and off-chip input/output accesses increase, consuming more resources

Engineering Contradiction:
Improveneural network computing capabilityVSAvoidresource consumption
Core Design Contradiction:
ReliabilityVSUse of energy by moving object

Solution Approach 1:

By fusing adjacent layers into a template fuse unit, the patent reduces the frequency of data transfers between on-chip and off-chip memory. This consolidation decreases the total number of input/output operations, thereby reducing energy consumption associated with these transfers while preserving the computational functionality.

Inventive Principle:
Principle #5Merging (Combining)

3Reliability

If multiple layers are used in neural network, then the feature extraction and classification performance are improved, but the data transfer overhead between on-chip and off-chip memory increases

Engineering Contradiction:
Improvefeature extraction performanceVSAvoiddata transfer overhead
Core Design Contradiction:
ReliabilityVSLoss of energy

Solution Approach 1:

The patent combines multiple adjacent layers into a template fuse unit that processes data in a single integrated operation. This reduces the number of intermediate data transfers between on-chip and off-chip memory, thereby reducing transfer overhead and energy consumption while maintaining the feature extraction capabilities through the fused layer structure.

Inventive Principle:
Principle #5Merging (Combining)

Data Source

PatentUS20230259746A1Device for forward fusion of neural network, board, method, and readable storage medium
Publication Date: 2023.08.17 CAMBRICON TECH CO LTD
  • US20230259746A1 patent drawing
  • US20230259746A1 patent drawing
  • US20230259746A1 patent drawing

AI summary

The present disclosure relates to an apparatus and a method for forward fusing a neural network, a board card, and a readable storage medium. The computing apparatus of the present disclosure is included in an integrated circuit apparatus. The integrated circuit apparatus includes a general interconnection interface and other processing apparatus. The computing apparatus interacts with other processing apparatus to jointly complete a computing operation specified by a user. The integrated circuit apparatus further includes a storage apparatus. The storage apparatus is connected to the computing apparatus and other processing apparatus, respectively. The storage apparatus is used for data storage of the computing apparatus and other processing apparatus.