Convolutional Operation Device Dynamic Dimensional Conversion
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Deep convolutional neural networks face inefficiencies in parallel processing due to fixed structural limitations in convolutional operation devices, leading to underutilization of Processing Elements (PEs) and reduced computational efficiency across varying convolutional layers.
Innovation Solution
A convolutional operation device with a dynamic dimensional structure, featuring input sharing networks, MAC arrays, and output shift networks, allows for optimized dimensional configuration of Processing Elements to align with the characteristics of each convolutional layer, enhancing parallel operation efficiency.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Productivity
If a fixed structural configuration of Processing Elements is used in convolutional operation devices, then the device complexity is reduced and manufacturing is easier, but the utilization rate of Processing Elements decreases and parallel operation efficiency is reduced when handling diverse convolutional layers
Solution Approach 1:
The patent implements a dynamic dimensional conversion mechanism that allows the convolutional operation device to adjust its processing elements arrangement from a fixed structure to a configurable dynamic structure. The system can transform between different dimensional configurations (e.g., 2D to 1D or 3D arrangements) based on the specific requirements of each convolutional layer, thereby maximizing PE utilization while maintaining manageable structural complexity through algorithmic control.
Solution Approach 2:
The patent applies dimensional conversion by changing the arrangement dimension of processing elements from traditional 2D grids to 1D sequences or 3D blocks depending on the convolutional layer characteristics. This dimensional flexibility enables the system to optimize data flow patterns and memory access efficiency for different convolutional operations, directly improving parallel operation efficiency without being constrained by a single fixed structural configuration.
2Productivity
If Processing Elements are arranged in a fixed dimensional structure, then the device is simpler to implement, but the utilization rate of Processing Elements is low when processing varying convolutional layer characteristics
Solution Approach 1:
The system dynamically reconfigures the dimensional arrangement of processing elements based on the characteristics of each convolutional layer being processed. By implementing runtime dimensional conversion, the device can adapt to varying computational requirements, maximizing PE utilization for each specific convolutional operation while maintaining the ability to handle diverse layer types through a unified flexible architecture.
Solution Approach 2:
The patent changes key parameters of the processing element arrangement, such as dimensional configuration, data flow direction, and computation granularity, to match the specific requirements of different convolutional layers. This parameter adaptation enables high PE utilization across varying workloads by adjusting the computational paradigm rather than requiring hardware reconfiguration.
3Productivity
If a fixed dimensional structure is used for convolutional operations, then the hardware design is simpler, but computational efficiency is reduced due to underutilization of Processing Elements across varying convolutional layers
Solution Approach 1:
The patent introduces a dynamic dimensional conversion mechanism that enables the system to switch between different processing configurations based on computational needs. This dynamic reconfiguration capability significantly improves computational efficiency by optimizing PE utilization for each convolutional layer, while the complexity is managed through software-controlled algorithms rather than hardware complexity.
Solution Approach 2:
The patent replaces fixed mechanical hardware configurations with a flexible software-controlled dimensional conversion system. Instead of designing separate hardware circuits for different dimensional arrangements, the system uses algorithmic control to achieve various dimensional configurations, thereby improving computational efficiency while keeping the hardware design relatively simple.
Data Source
AI summary
A convolutional operation device for performing convolutional neural network processing includes an input sharing network including first and second input feature map registers configured to shift each input feature map, which is inputted in row units, in a row or column direction and output the shifted input feature map and arranged in rows and columns, a first MAC array connected to the first input feature map registers, an input feature map switching network configured to select one of the first and second input feature map registers, a second MAC array connected to one selected by the input feature map switching network among the first and second input feature map registers, and an output shift network configured to shift the output feature map from the first MAC array and the second MAC array to transmit the shifted output feature map to an output memory.


