Workload Management Engine for AI Neural Network Optimization

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional artificial intelligence systems lack comprehensive workload management for processing units, leading to inefficiencies such as increased inference latency, reduced throughput, and suboptimal inference speeds due to unoptimized neural network processing on shared resources.

Innovation Solution

A workload management engine that dynamically switches between different neural network models based on workload management factors and logic, using strategies like quantization and pruning to optimize processing unit performance.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If multiple neural networks operate simultaneously on a shared NPU, then application versatility is improved, but computational resource competition increases causing suboptimal inference speeds

Engineering Contradiction:
Improveapplication versatilityVSAvoidinference speed
Core Design Contradiction:
Adaptability or versatilityVSProductivity

Solution Approach 1:

The system dynamically switches between different neural network models based on real-time workload management factors including power mode, operational mode, NPU capacity, and physical environment conditions. This dynamic adaptation allows the system to optimize inference speed for specific applications while maintaining versatility across multiple applications.

Inventive Principle:
Principle #15Dynamics

Solution Approach 2:

The workload management engine changes system parameters by selecting different neural network models based on identified workload management factors. The engine monitors factors such as power supply mode, operational mode, and NPU capacity to determine optimal model selection, thereby resolving the contradiction between versatility and inference speed.

Inventive Principle:
Principle #35Parameter changes

2Measurement precision

If full neural network models are deployed, then processing accuracy is improved, but computational resource consumption increases

Engineering Contradiction:
Improveprocessing accuracyVSAvoidcomputational resource consumption
Core Design Contradiction:
Measurement precisionVSUse of energy by moving object

Solution Approach 1:

The system employs reduced neural network models that use partial computational resources compared to full models. These reduced models are selected based on workload management factors, providing sufficient processing accuracy for specific tasks while consuming fewer computational resources and energy.

Inventive Principle:
Principle #16Partial or excessive action

Solution Approach 2:

The workload management engine changes the model size parameter by selecting between full and reduced neural network models based on identified factors such as power mode and NPU capacity. This allows optimization of the balance between processing accuracy and computational resource consumption.

Inventive Principle:
Principle #35Parameter changes

3Productivity

If neural network models are reduced using quantization and pruning, then computational efficiency is improved, but model performance may deteriorate

Engineering Contradiction:
Improvecomputational efficiencyVSAvoidmodel performance
Core Design Contradiction:
ProductivityVSMeasurement precision

Solution Approach 1:

The system changes the model reduction parameters by applying different quantization levels and pruning intensities to generate multiple reduced models. The workload management engine then selects the appropriate reduced model based on workload management factors, optimizing the balance between computational efficiency and model performance.

Inventive Principle:
Principle #35Parameter changes

Solution Approach 2:

The system applies partial reduction techniques (quantization and pruning) to create reduced models that maintain sufficient performance for specific tasks while improving computational efficiency. The workload management engine determines the appropriate level of reduction based on identified factors.

Inventive Principle:
Principle #16Partial or excessive action

4Adaptability or versatility

If dynamic model switching is implemented, then workload adaptability is improved, but system complexity increases

Engineering Contradiction:
Improveworkload adaptabilityVSAvoidsystem complexity
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The workload management engine acts as an intermediary component that handles the complexity of dynamic model switching. It monitors workload management factors and coordinates model selection, thereby improving workload adaptability while containing system complexity within a dedicated management layer.

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The workload management engine provides multi-functional capabilities including monitoring, decision-making, and coordination of neural network model selection. This universal management approach improves workload adaptability across different applications while consolidating complexity into a single engine.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Data Source

PatentUS20250209304A1Workload management engine in an artificial intelligence system
Publication Date: 2025.06.26 MICROSOFT TECHNOLOGY LICENSING LLC
  • US20250209304A1 patent drawing
  • US20250209304A1 patent drawing
  • US20250209304A1 patent drawing

AI summary

Methods, systems, and computer storage media for providing workload management using a workload management engine in an artificial intelligence (AI) system. In particular, workload management incorporates adaptive strategies that adjust the neural network models employed by a processing unit (e.g., NPU/GPU/TPU) based on the dynamic nature of workloads, workload management factors, and workload management logic. The workload management engine provides the workload management logic to support strategic decision-making for processor optimization. In operation, a plurality states of workload management factors are identified. A task associated with a workload processing unit is identified. Based on the task and the plurality of states of the workload processing unit, a neural network model from a plurality of neural network models is selected. The plurality of neural network models include a full neural network model and a reduced neural network model. The task is caused to be executed using the identified neural network model.