Systolic Array PE Distribution for Multi-NN Execution

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current deep learning operation devices using systolic arrays face limitations in achieving high throughput and real-time performance due to their inability to efficiently distribute processing elements (PEs) across multiple neural networks (NNs) simultaneously, leading to suboptimal execution of deep learning operations.

Innovation Solution

The proposed solution involves a processor with a systolic array that distributes PEs to perform deep learning operations across multiple NNs or sub-NNs based on their characteristics, allowing for simultaneous execution by setting optimal propagation directions for input data and output partial sums, thereby improving resource utilization and operational efficiency.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Power

If a systolic array is used to perform deep learning operations, then computational power is improved, but the ability to handle multiple neural networks simultaneously deteriorates

Engineering Contradiction:
Improvecomputational powerVSAvoidability to handle multiple neural networks
Core Design Contradiction:
PowerVSAdaptability or versatility

Solution Approach 1:

The systolic array is divided into multiple processing element groups, where each group can be independently allocated to different neural networks. This segmentation enables the system to handle multiple NNs simultaneously while maintaining high computational power within each segment.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The processor dynamically configures the systolic array by adjusting the number of processing elements allocated to each neural network based on real-time requirements. This dynamic reconfiguration allows the system to adapt to varying computational demands of different NNs while maximizing overall computational efficiency.

Inventive Principle:
Principle #15Dynamics

2Productivity

If processing elements are dedicated to a single neural network, then computational efficiency for that network is improved, but resource utilization deteriorates

Engineering Contradiction:
Improvecomputational efficiencyVSAvoidresource utilization
Core Design Contradiction:
ProductivityVSQuantity of substance

Solution Approach 1:

Processing elements in the systolic array are designed to be universal and reconfigurable, allowing them to serve multiple neural networks sequentially or concurrently. This multi-functionality ensures high computational efficiency for each NN while maximizing overall resource utilization across the system.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Solution Approach 2:

The system dynamically allocates and reconfigures processing elements between different neural networks based on current computational demands. This dynamic resource management maintains high computational efficiency for active NNs while ensuring optimal utilization of all available processing resources.

Inventive Principle:
Principle #15Dynamics

3Speed

If the systolic array is configured for high throughput, then processing speed is improved, but flexibility in handling different neural network characteristics deteriorates

Engineering Contradiction:
Improveprocessing speedVSAvoidflexibility in handling different neural network characteristics
Core Design Contradiction:
SpeedVSAdaptability or versatility

Solution Approach 1:

The systolic array configuration is dynamically adjusted based on the specific characteristics of each neural network being processed. Processing speed is optimized for each NN's requirements while maintaining the ability to adapt to different network architectures, layer configurations, and computational patterns.

Inventive Principle:
Principle #15Dynamics

Solution Approach 2:

Different regions or groups of processing elements within the systolic array can be configured with different operational characteristics tailored to specific neural network requirements. This local customization enables high processing speed for each NN's specific needs while maintaining overall system flexibility.

Inventive Principle:
Principle #3Local quality

Data Source

PatentUS20220138563A1Method and device with deep learning operations
Publication Date: 2022.05.05 SAMSUNG ELECTRONICS CO LTD
  • US20220138563A1 patent drawing
  • US20220138563A1 patent drawing
  • US20220138563A1 patent drawing

AI summary

A method and a device with deep learning operations. An electronic device includes a processor configured to simultaneously perform, using a systolic array, a plurality of tasks, wherein the processor includes the systolic array having a plurality of processing elements (PEs), and a first on-chip network that performs data propagation between two or more of the plurality of PEs, where each of the plurality of tasks includes one or more deep learning operations.