Systolic Array PE Distribution for Multi-NN Execution
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current deep learning operation devices using systolic arrays face limitations in achieving high throughput and real-time performance due to their inability to efficiently distribute processing elements (PEs) across multiple neural networks (NNs) simultaneously, leading to suboptimal execution of deep learning operations.
Innovation Solution
The proposed solution involves a processor with a systolic array that distributes PEs to perform deep learning operations across multiple NNs or sub-NNs based on their characteristics, allowing for simultaneous execution by setting optimal propagation directions for input data and output partial sums, thereby improving resource utilization and operational efficiency.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Power
If a systolic array is used to perform deep learning operations, then computational power is improved, but the ability to handle multiple neural networks simultaneously deteriorates
Solution Approach 1:
The systolic array is divided into multiple processing element groups, where each group can be independently allocated to different neural networks. This segmentation enables the system to handle multiple NNs simultaneously while maintaining high computational power within each segment.
Solution Approach 2:
The processor dynamically configures the systolic array by adjusting the number of processing elements allocated to each neural network based on real-time requirements. This dynamic reconfiguration allows the system to adapt to varying computational demands of different NNs while maximizing overall computational efficiency.
2Productivity
If processing elements are dedicated to a single neural network, then computational efficiency for that network is improved, but resource utilization deteriorates
Solution Approach 1:
Processing elements in the systolic array are designed to be universal and reconfigurable, allowing them to serve multiple neural networks sequentially or concurrently. This multi-functionality ensures high computational efficiency for each NN while maximizing overall resource utilization across the system.
Solution Approach 2:
The system dynamically allocates and reconfigures processing elements between different neural networks based on current computational demands. This dynamic resource management maintains high computational efficiency for active NNs while ensuring optimal utilization of all available processing resources.
3Speed
If the systolic array is configured for high throughput, then processing speed is improved, but flexibility in handling different neural network characteristics deteriorates
Solution Approach 1:
The systolic array configuration is dynamically adjusted based on the specific characteristics of each neural network being processed. Processing speed is optimized for each NN's requirements while maintaining the ability to adapt to different network architectures, layer configurations, and computational patterns.
Solution Approach 2:
Different regions or groups of processing elements within the systolic array can be configured with different operational characteristics tailored to specific neural network requirements. This local customization enables high processing speed for each NN's specific needs while maintaining overall system flexibility.
Data Source
AI summary
A method and a device with deep learning operations. An electronic device includes a processor configured to simultaneously perform, using a systolic array, a plurality of tasks, wherein the processor includes the systolic array having a plurality of processing elements (PEs), and a first on-chip network that performs data propagation between two or more of the plurality of PEs, where each of the plurality of tasks includes one or more deep learning operations.


