Edge NPU Resource Configuration for Faster Custom AI Deployment
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
The manual design process for neural processing units (NPUs) is time-consuming and inefficient in configuring customized devices according to customer requests, as it does not account for individual customer needs and resource optimization.
Innovation Solution
A method for automatically adjusting basic resources of NPUs to meet customer-specific instructions by predicting performance and optimizing computational resources, including adjusting parameters such as channel numbers, memory capacity, and data bandwidth.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If manual design method is used for NPU configuration, then design flexibility and customization are achieved, but design time and development efficiency are significantly increased
Solution Approach 1:
The patent pre-configures multiple NPU architectures with different resource allocations (compute units, memory, bandwidth) before customer requests arrive. When a customer orders an edge device, the system selects from these pre-configured architectures rather than designing from scratch, dramatically reducing design time while maintaining customization capability through the variety of pre-prepared configurations.
Solution Approach 2:
The patent creates a family of NPU architectures by systematically varying key parameters such as the number of compute units, memory size, and data bandwidth. This allows the system to offer customized solutions for different customer needs (different computational performance, power consumption, chip area requirements) by selecting appropriate parameter combinations from the pre-configured family, rather than performing manual design adjustments for each request.
2Productivity
If basic resources of NPU are increased to meet customer instructions, then computational performance is improved, but resource consumption and chip area are increased
Solution Approach 1:
The patent configures NPUs with non-uniform resource distribution tailored to specific computational workloads. Instead of uniformly increasing all resources, the system adjusts local resource allocations (e.g., more compute units for compute-intensive tasks, more memory for memory-intensive tasks) based on the customer's specific instruction requirements, achieving high computational performance with minimized overall resource consumption.
Solution Approach 2:
The patent provides a family of NPU architectures with dynamically selectable resource configurations. The optimal configuration is selected based on the specific customer instruction and workload characteristics, allowing the system to match computational performance exactly to requirements without over-provisioning resources, thereby minimizing resource consumption and chip area while meeting performance targets.
3Productivity
If multiple NPU architectures are pre-configured with different resource allocations, then customer-specific requirements are met faster, but device complexity and design overhead are increased
Solution Approach 1:
The patent segments the NPU design space into discrete, pre-configured architecture options with clearly defined resource allocations. Each segment represents a complete, validated configuration that can be independently selected and deployed. This segmentation transforms the complex continuous design space into manageable discrete choices, reducing design overhead while enabling fast customer-specific configuration through simple selection from the segmented options.
4Manufacturing precision
If manual adjustment of NPU resources is performed for each customer request, then optimal configuration is achieved, but development efficiency and time-to-market are reduced
Solution Approach 1:
The patent creates a library of validated NPU architecture templates that can be copied and selected for different customer requests. Instead of performing manual optimization for each new customer, the system copies appropriate pre-optimized configurations from the library and makes minor adjustments if needed, maintaining configuration optimization quality while dramatically improving development efficiency and reducing time-to-market.
Data Source
AI summary
Disclosed is a neural network optimizing method performed by a device including performing a computation on an input value with a basic resource of a first neural network based on a customer-requested instruction for the basic resource used for computation execution of the first neural network, checking computational performance including a computational processing speed, a power consumption amount, and a chip area according to the computation execution of the first neural network, adjusting at least one of the basic resource based on the checked computational performance, and re-performing the computation on the input value based on the customer-requested instruction with a resource of the second neural network after changing to an environment of a second neural network by adjusting the at least one of the basic resource.


