Neural Network Processing Units for Parallel Computation Pipelining
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Deep neural network models face challenges in efficiently implementing large numbers of computations required for complex applications, leading to inefficient hardware utilization and performance issues.
Innovation Solution
A computing system comprising multiple processing units is arranged to improve performance and hardware utilization by distributing computations across units, allowing for flexible pipelining designs and even computation loads.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Productivity
If multiple processing units are used to process computations of the same layer, then computation speed and performance are improved, but hardware utilization becomes less efficient
Solution Approach 1:
The patent segments the neural network computation into different layers and assigns them to different processing units. Specifically, multiple processing units are used to parallelize computations within the same layer (e.g., different feature maps or channels), while maintaining a pipeline structure where each processing unit handles a specific layer. This segmentation enables both improved computation speed through parallelization and better hardware utilization through specialized assignment.
Solution Approach 2:
The patent introduces a temporal dimension to the hardware utilization problem by using a pipeline architecture. Different processing units operate at different stages of the computation pipeline simultaneously, allowing the system to maintain high utilization across all units over time. This dimensional approach transforms the static resource allocation problem into a dynamic temporal scheduling problem.
2Loss of energy
If a single processing unit processes multiple layers, then hardware utilization is improved, but computation load becomes uneven and performance decreases
Solution Approach 1:
The patent applies local quality by assigning different processing capabilities and specialized functions to different processing units based on the specific requirements of the layers they handle. Each processing unit is optimized for its specific layer's computation characteristics, creating local specialization rather than uniform general-purpose processing. This ensures both high utilization and balanced load distribution.
3Device complexity
If processing units are arranged in a fixed pipeline structure, then computation flow is simplified, but flexibility in handling different neural network configurations is reduced
Solution Approach 1:
The patent implements dynamics by making the processing unit arrangement and data flow configurable rather than fixed. The system can dynamically adjust which processing units handle which layers and how data is routed between them, allowing adaptation to different neural network architectures while maintaining the benefits of pipeline structure. This dynamic configurability resolves the contradiction between simplified flow and flexibility.
Data Source
AI summary
The present application discloses a computing system for implementing an artificial neural network model. The artificial neural network model has a structure of multiple layers. The computing system comprises a first processing unit, a second processing unit, and a third processing unit. The first processing unit performs computations of the first layer based on a first part of input data of the first layer to generate a first part of output data. The second processing unit performs computations of the first layer based on a second part of the input data of the first layer so as to generate a second part of the output data. The third processing unit performs computations of the second layer based on the first part and the second part of the output data. The first processing unit, the second processing unit, and the third processing unit have the same structure.


