Neural Network Scheduler for Dynamic Resource Allocation

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Allocating sufficient computing resources for neural networks to perform inferences is challenging due to significant resource demands and competition among machine learning models, leading to inefficient resource utilization and performance degradation.

Innovation Solution

An AI-assisted system that uses neural networks to predict performance characteristics of machine learning models, allowing for dynamic allocation of resources through load balancing and reassignment of models to optimize computing resource utilization across servers.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If computing resources are allocated to multiple machine learning models, then model diversity and functionality are improved, but resource competition and utilization efficiency deteriorate

Engineering Contradiction:
Improvemodel diversityVSAvoidresource utilization efficiency
Core Design Contradiction:
Adaptability or versatilityVSLoss of energy

Solution Approach 1:

The system dynamically adjusts resource allocation based on real-time performance monitoring and neural network predictions. Computing resources are not statically assigned but continuously reallocated according to changing workload conditions, model performance requirements, and resource availability, resolving the contradiction between supporting multiple models and maintaining efficiency

Inventive Principle:
Principle #15Dynamics

Solution Approach 2:

The system changes key parameters such as batch size, learning rate, and resource allocation ratios based on neural network predictions of performance characteristics. These parameter adjustments optimize resource utilization for each model while maintaining the ability to run multiple models simultaneously

Inventive Principle:
Principle #35Parameter changes

2Productivity

If more computing resources are allocated to neural networks, then inference performance is improved, but resource availability for other models deteriorates

Engineering Contradiction:
Improveinference performanceVSAvoidresource availability
Core Design Contradiction:
ProductivityVSAdaptability or versatility

Solution Approach 1:

The system implements continuous feedback loops where performance metrics are monitored, neural networks predict resource requirements, and resource allocation is adjusted accordingly. This feedback mechanism ensures that inference performance is optimized while maintaining adequate resource availability for other models through iterative adjustments

Inventive Principle:
Principle #23Feedback

Solution Approach 2:

The system uses neural networks to predict performance characteristics and resource requirements before actual inference workloads execute. This preliminary prediction allows proactive resource allocation that ensures sufficient resources for high-priority inference tasks while pre-reserving capacity for other models

Inventive Principle:
Principle #10Preliminary action

3Loss of energy

If computing resources are dynamically reallocated, then resource utilization efficiency is improved, but system complexity and management overhead increase

Engineering Contradiction:
Improveresource utilization efficiencyVSAvoidsystem complexity
Core Design Contradiction:
Loss of energyVSDevice complexity

Solution Approach 1:

The system employs self-service mechanisms where neural networks automatically predict resource requirements and performance characteristics without human intervention. The automated prediction and allocation system reduces manual management overhead while maintaining high resource utilization efficiency through intelligent, autonomous decision-making

Inventive Principle:
Principle #25Self-service

Data Source

PatentUS20220180178A1Neural network scheduler
Publication Date: 2022.06.09 NVIDIA CORP
  • US20220180178A1 patent drawing
  • US20220180178A1 patent drawing
  • US20220180178A1 patent drawing

AI summary

Apparatuses, systems, and techniques to allocate computing resources to perform inferences. In at least one embodiment, one or more neural networks cause computing resources to be identified based, at least in part, on performance requirements of one or more neural networks to perform inferences.