Workload Management Engine for AI Neural Network Optimization
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional artificial intelligence systems lack comprehensive workload management for processing units, leading to inefficiencies such as increased inference latency, reduced throughput, and suboptimal inference speeds due to unoptimized neural network processing on shared resources.
Innovation Solution
A workload management engine that dynamically switches between different neural network models based on workload management factors and logic, using strategies like quantization and pruning to optimize processing unit performance.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If multiple neural networks operate simultaneously on a shared NPU, then application versatility is improved, but computational resource competition increases causing suboptimal inference speeds
Solution Approach 1:
The system dynamically switches between different neural network models based on real-time workload management factors including power mode, operational mode, NPU capacity, and physical environment conditions. This dynamic adaptation allows the system to optimize inference speed for specific applications while maintaining versatility across multiple applications.
Solution Approach 2:
The workload management engine changes system parameters by selecting different neural network models based on identified workload management factors. The engine monitors factors such as power supply mode, operational mode, and NPU capacity to determine optimal model selection, thereby resolving the contradiction between versatility and inference speed.
2Measurement precision
If full neural network models are deployed, then processing accuracy is improved, but computational resource consumption increases
Solution Approach 1:
The system employs reduced neural network models that use partial computational resources compared to full models. These reduced models are selected based on workload management factors, providing sufficient processing accuracy for specific tasks while consuming fewer computational resources and energy.
Solution Approach 2:
The workload management engine changes the model size parameter by selecting between full and reduced neural network models based on identified factors such as power mode and NPU capacity. This allows optimization of the balance between processing accuracy and computational resource consumption.
3Productivity
If neural network models are reduced using quantization and pruning, then computational efficiency is improved, but model performance may deteriorate
Solution Approach 1:
The system changes the model reduction parameters by applying different quantization levels and pruning intensities to generate multiple reduced models. The workload management engine then selects the appropriate reduced model based on workload management factors, optimizing the balance between computational efficiency and model performance.
Solution Approach 2:
The system applies partial reduction techniques (quantization and pruning) to create reduced models that maintain sufficient performance for specific tasks while improving computational efficiency. The workload management engine determines the appropriate level of reduction based on identified factors.
4Adaptability or versatility
If dynamic model switching is implemented, then workload adaptability is improved, but system complexity increases
Solution Approach 1:
The workload management engine acts as an intermediary component that handles the complexity of dynamic model switching. It monitors workload management factors and coordinates model selection, thereby improving workload adaptability while containing system complexity within a dedicated management layer.
Solution Approach 2:
The workload management engine provides multi-functional capabilities including monitoring, decision-making, and coordination of neural network model selection. This universal management approach improves workload adaptability across different applications while consolidating complexity into a single engine.
Data Source
AI summary
Methods, systems, and computer storage media for providing workload management using a workload management engine in an artificial intelligence (AI) system. In particular, workload management incorporates adaptive strategies that adjust the neural network models employed by a processing unit (e.g., NPU/GPU/TPU) based on the dynamic nature of workloads, workload management factors, and workload management logic. The workload management engine provides the workload management logic to support strategic decision-making for processor optimization. In operation, a plurality states of workload management factors are identified. A task associated with a workload processing unit is identified. Based on the task and the plurality of states of the workload processing unit, a neural network model from a plurality of neural network models is selected. The plurality of neural network models include a full neural network model and a reduced neural network model. The task is caused to be executed using the identified neural network model.


