Neural Network Data Reorganization for Bus Bandwidth Utilization
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Traditional CPU and GPU methods are inadequate for processing neural networks due to their large number of parameters and high degree of parallelism, necessitating a specialized processor for deep learning applications.
Innovation Solution
A data processing method and device for neural networks that involves receiving output feature data subsets from a processing circuit array, reorganizing them according to specific positions within the network layers, and writing the reorganized data subsets into a destination storage device, optimizing system efficiency and bus bandwidth utilization.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Productivity
If traditional CPU or GPU methods are used for neural network processing, then general-purpose computing capability is maintained, but processing efficiency and bandwidth utilization deteriorate due to the large number of parameters and high degree of parallelism in neural networks
Solution Approach 1:
The patent segments the neural network processing into distinct functional modules: processing circuit arrays for parallel computation, reorganization circuits for data arrangement, and writing circuits for memory storage. This segmentation allows each module to be optimized for its specific function, improving overall processing efficiency while managing complexity through modular design.
Solution Approach 2:
The patent introduces reorganization circuits as intermediary components between the processing circuit arrays and memory systems. These intermediary circuits buffer and reorganize data subsets, decoupling the parallel processing operations from sequential memory access requirements, thereby improving bandwidth utilization without requiring fundamental changes to the memory architecture.
2Speed
If data subsets are processed in parallel across multiple processing circuits, then processing speed improves, but data organization and memory bandwidth utilization worsen due to scattered data positions
Solution Approach 1:
The patent applies preliminary action by reorganizing data subsets in advance before they are written to memory or used in subsequent processing stages. The reorganization circuits arrange data according to their logical positions in the neural network layers before storage, preventing scattered data access patterns and improving memory bandwidth utilization in subsequent operations.
3Productivity
If specialized processing circuits are designed for neural network operations, then processing efficiency improves, but adaptability to different data formats and network architectures deteriorates
Solution Approach 1:
The patent implements universality through reorganization circuits that can handle multiple data formats and network layer types. These circuits are designed to reorganize data subsets from different processing circuit arrays regardless of the specific neural network architecture or data format, allowing the same hardware infrastructure to support various deep learning models and operations.
Data Source
AI summary
A data processing method and device for a neural network, a neural network processing device, and a storage medium. The neural network includes a plurality of network layers, the plurality of network layers include a first network layer, and the method includes: receiving at least two output feature data subsets, obtained by a processing circuit array for at least two different portions of the first network layer, from the processing circuit array; splicing and combining the at least two output feature data subsets according to positions of the at least two different portions in the first network layer to obtain a reorganized data subset; and writing the reorganized data subset into a destination storage device. The data processing method for neural network processing may be capable of improving system efficiency and improving bus bandwidth utilization.


