Edge NPU Resource Configuration for Faster Custom AI Deployment

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

The manual design process for neural processing units (NPUs) is time-consuming and inefficient in configuring customized devices according to customer requests, as it does not account for individual customer needs and resource optimization.

Innovation Solution

A method for automatically adjusting basic resources of NPUs to meet customer-specific instructions by predicting performance and optimizing computational resources, including adjusting parameters such as channel numbers, memory capacity, and data bandwidth.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If manual design method is used for NPU configuration, then design flexibility and customization are achieved, but design time and development efficiency are significantly increased

Engineering Contradiction:
Improvecustomization capabilityVSAvoiddesign time
Core Design Contradiction:
Adaptability or versatilityVSLoss of time

Solution Approach 1:

The patent pre-configures multiple NPU architectures with different resource allocations (compute units, memory, bandwidth) before customer requests arrive. When a customer orders an edge device, the system selects from these pre-configured architectures rather than designing from scratch, dramatically reducing design time while maintaining customization capability through the variety of pre-prepared configurations.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent creates a family of NPU architectures by systematically varying key parameters such as the number of compute units, memory size, and data bandwidth. This allows the system to offer customized solutions for different customer needs (different computational performance, power consumption, chip area requirements) by selecting appropriate parameter combinations from the pre-configured family, rather than performing manual design adjustments for each request.

Inventive Principle:
Principle #35Parameter changes

2Productivity

If basic resources of NPU are increased to meet customer instructions, then computational performance is improved, but resource consumption and chip area are increased

Engineering Contradiction:
Improvecomputational performanceVSAvoidresource consumption
Core Design Contradiction:
ProductivityVSLoss of energy

Solution Approach 1:

The patent configures NPUs with non-uniform resource distribution tailored to specific computational workloads. Instead of uniformly increasing all resources, the system adjusts local resource allocations (e.g., more compute units for compute-intensive tasks, more memory for memory-intensive tasks) based on the customer's specific instruction requirements, achieving high computational performance with minimized overall resource consumption.

Inventive Principle:
Principle #3Local quality

Solution Approach 2:

The patent provides a family of NPU architectures with dynamically selectable resource configurations. The optimal configuration is selected based on the specific customer instruction and workload characteristics, allowing the system to match computational performance exactly to requirements without over-provisioning resources, thereby minimizing resource consumption and chip area while meeting performance targets.

Inventive Principle:
Principle #15Dynamics

3Productivity

If multiple NPU architectures are pre-configured with different resource allocations, then customer-specific requirements are met faster, but device complexity and design overhead are increased

Engineering Contradiction:
Improveconfiguration speedVSAvoiddesign complexity
Core Design Contradiction:
ProductivityVSDevice complexity

Solution Approach 1:

The patent segments the NPU design space into discrete, pre-configured architecture options with clearly defined resource allocations. Each segment represents a complete, validated configuration that can be independently selected and deployed. This segmentation transforms the complex continuous design space into manageable discrete choices, reducing design overhead while enabling fast customer-specific configuration through simple selection from the segmented options.

Inventive Principle:
Principle #1Segmentation

4Manufacturing precision

If manual adjustment of NPU resources is performed for each customer request, then optimal configuration is achieved, but development efficiency and time-to-market are reduced

Engineering Contradiction:
Improveconfiguration optimizationVSAvoiddevelopment efficiency
Core Design Contradiction:
Manufacturing precisionVSProductivity

Solution Approach 1:

The patent creates a library of validated NPU architecture templates that can be copied and selected for different customer requests. Instead of performing manual optimization for each new customer, the system copies appropriate pre-optimized configurations from the library and makes minor adjustments if needed, maintaining configuration optimization quality while dramatically improving development efficiency and reducing time-to-market.

Inventive Principle:
Principle #26Copying

Data Source

PatentUS12614076B2Neural network optimization device for edge device meeting on-demand instruction and method using the same
Publication Date: 2026.04.28 MOBILINT INC
  • US12614076B2 patent drawing
  • US12614076B2 patent drawing
  • US12614076B2 patent drawing

AI summary

Disclosed is a neural network optimizing method performed by a device including performing a computation on an input value with a basic resource of a first neural network based on a customer-requested instruction for the basic resource used for computation execution of the first neural network, checking computational performance including a computational processing speed, a power consumption amount, and a chip area according to the computation execution of the first neural network, adjusting at least one of the basic resource based on the checked computational performance, and re-performing the computation on the input value based on the customer-requested instruction with a resource of the second neural network after changing to an environment of a second neural network by adjusting the at least one of the basic resource.