Inference Model Migration for Edge Resource Optimization
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing information handling systems face challenges in efficiently managing workload migration between client and edge devices, particularly in determining when to migrate AI/ML inference models to edge servers for improved performance.
Innovation Solution
The system employs resource detection circuitry to collect data on resource utilization, determining performance levels with and without executing AI/ML inference models. Based on performance gains, network latency, and power consumption, the system decides to migrate the inference model to an edge server for execution.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Productivity
If the inference model is executed in the information handling system, then the system can process AI/ML tasks locally, but resource utilization increases and performance may degrade
Solution Approach 1:
The system dynamically determines whether to execute the inference model locally or migrate it to an edge server based on real-time resource availability and performance requirements. This dynamic decision-making allows the system to adapt resource allocation to changing conditions, improving application performance when resources are available while reducing resource utilization when the model is migrated to the edge server
Solution Approach 2:
The system introduces an edge server as an intermediary between the information handling system and the inference model execution. By migrating the inference model to the edge server, the system offloads computational tasks to an external resource, thereby reducing local resource utilization while maintaining or improving application performance through access to enhanced computing capabilities
2Speed
If the inference model is executed locally, then processing speed may be fast, but power consumption increases
Solution Approach 1:
The system dynamically evaluates the trade-off between processing speed and power consumption by monitoring resource utilization and performance metrics. When local execution of the inference model causes excessive power consumption, the system migrates the model to an edge server, thereby reducing power consumption while maintaining acceptable processing speeds through the edge server's computational resources
3Use of energy by moving object
If the inference model is migrated to edge server, then resource utilization decreases, but network latency may increase
Solution Approach 1:
The system dynamically determines the optimal execution location by evaluating both resource utilization and network latency considerations. When migrating the inference model to an edge server reduces resource utilization, the system monitors for performance gains and adjusts the execution location accordingly, accepting network latency trade-offs only when they result in overall performance improvement
Data Source
AI summary
An information handling system includes resource detection circuitry that collects data associated with resources being utilized in the information handling system. The system determines resources for execution of an inference model, and receives the data associated with the resources from the resource detection circuitry. Based on the resources for the execution of the inference model, the system determines one performance of an application when the inference model is executed in the information handling system. The system determines another performance level of the application when the inference model is not executed in the information handling system. Based on the two performance levels, the system determines whether the application has a performance gain by the inference model not being executed in the information handling system. In response to the performance gain, the system migrates the inference model to an edge server for execution.


