DNN Partitioning Across Computation Areas for Edge Memory Limits
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing deep neural network (DNN) computations face inefficiencies in computational speed and power consumption due to spatial partitioning, especially in the latter half of the network, and require significant memory resources, particularly in edge devices with limited capacity.
Innovation Solution
A computing apparatus that switches between spatial and channel partitioning methods during DNN computation, utilizing a plurality of computation areas with intermediate memory to manage feature maps efficiently, allowing parallel processing and dynamic allocation of feature maps without subpartitioning.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Productivity
If spatial partitioning is used to distribute feature maps across multiple edge devices, then the computation can be parallelized, but computational efficiency decreases in the latter half of the DNN network where feature map size is reduced
Solution Approach 1:
The patent dynamically switches between spatial partitioning and channel partitioning based on the DNN layer being processed. In the first half of the network where feature maps are large, spatial partitioning is used. In the latter half where feature maps are small but channels are numerous, channel partitioning is used. This dynamic adaptation resolves the contradiction by optimizing the partitioning strategy to match the characteristics of each network stage.
Solution Approach 2:
The patent changes the partitioning parameter from spatial division to channel division based on the processing stage. By monitoring the feature map characteristics (size and channel count) at different layers, the system transitions between partitioning methods to maintain optimal computational efficiency throughout the entire DNN computation.
2Measurement precision
If large-scale DNN is used to achieve highly accurate image analysis, then accuracy improves, but memory requirements increase significantly
Solution Approach 1:
The patent segments the DNN computation into multiple partitions using both spatial and channel partitioning methods. By dividing the computation across multiple edge devices and organizing feature maps into manageable groups, the system can handle large-scale DNNs with limited memory capacity on individual devices, thereby achieving high accuracy without requiring excessive memory resources.
Solution Approach 2:
The patent introduces channel partitioning as an additional dimension for organizing and processing feature maps. Instead of only dividing feature maps spatially, the system partitions along the channel dimension, creating a two-dimensional partitioning strategy that efficiently manages memory resources while maintaining the ability to process large-scale DNNs for high-accuracy image analysis.
3Power
If multiple edge computing devices are operated to perform partitioned DNN computation, then computation capacity increases, but power consumption increases
Solution Approach 1:
The patent uses channel partitioning to process only the necessary channels in the latter half of the DNN network, avoiding redundant computations. By selectively activating computation for specific channel groups rather than processing all channels on all devices, the system reduces overall power consumption while maintaining sufficient computation capacity for accurate image analysis.
Data Source
AI summary
A computing apparatus includes at least one computing device including a computation area. The computing apparatus acquires input information, generates a plurality of input feature maps from the input information, and performs DNN computation in parallel on the generated plurality of input feature maps by at least one DNN partitioning method including channel partitioning. In the channel partitioning, by grouping the feature maps into sets grouping the feature maps without partitioning the feature maps and allocating each set grouping the feature maps to each computation area, computation for layers is performed.


