Reconfigurable Hardware Block for Neural Processing Unit Workload Distribution
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional Neural Processing Units (NPUs) face inefficiencies when handling mixed workloads of different deep neural network operations, leading to resource wastage and delayed processing due to fixed hardware structures that are not adaptable to varying workloads of CNN and RNN operations.
Innovation Solution
A neural processing unit with a reconfigurable hardware block and a work distributer that dynamically allocates workload between processing cores and the hardware block based on workload demands, allowing the hardware block to be reconfigured for different neural network operations, thereby optimizing resource utilization and reducing idle time.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If the NPU has a fixed hardware structure for one kind of DNN operation, then the hardware structure is simple, but another type of DNN operation cannot be performed or is performed with high delay
Solution Approach 1:
The patent implements a universal hardware block that can be reconfigured to perform different types of DNN operations (CNN, RNN, etc.) through software configuration rather than requiring separate dedicated hardware cores for each operation type. This allows a single hardware structure to serve multiple functions, resolving the contradiction between versatility and complexity.
Solution Approach 2:
The hardware block is designed with dynamic reconfigurability, allowing its internal structure to be changed based on the specific DNN operation being performed. The system can transition between different operational modes through configuration changes, enabling adaptability without requiring multiple fixed hardware architectures.
2Productivity
If the NPU has multiple processing cores for different DNN operations, then different DNN operations can be performed simultaneously, but resources are wasted when one core completes earlier and becomes idle
Solution Approach 1:
The system dynamically adjusts the operational state of processing cores based on workload distribution. When one core completes its task earlier, the system can redistribute remaining work to other cores or reconfigure the hardware block to handle additional operations, preventing idle time and resource wastage while maintaining high processing throughput.
Solution Approach 2:
The system monitors the completion status of each processing core and uses this feedback to dynamically redistribute workloads. This ensures that all processing cores remain actively engaged in productive work, optimizing resource utilization while maintaining high productivity across diverse DNN operation workloads.
3Adaptability or versatility
If the NPU has dedicated hardware cores for CNN and RNN operations, then both operation types can be processed, but the time for application-work is delayed when amounts of operations differ
Solution Approach 1:
The patent employs a universal hardware block that can be configured to handle any type of DNN operation, allowing flexible workload distribution. This eliminates the need for separate dedicated cores for CNN and RNN operations, enabling more efficient processing when the amounts of different operation types differ by allowing dynamic workload allocation across a single versatile hardware unit.
Data Source
AI summary
Provided is a neural processing unit that performs application-work including a first neural network operation, the neural processing unit includes a first processing core configured to execute the first neural network operation, a hardware block reconfigurable as a hardware core configured to perform hardware block-work, and at least one processor configured to execute computer-readable instructions to distribute a part of the application-work as the hardware block-work to the hardware block based on a first workload of the first processing core.


