Dynamic Load Balancing for ML Inference in Mobile Systems
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current cloud-based mobile systems face challenges in providing efficient machine learning due to high latency and bandwidth saturation, as they rely on a centralized structure that is geographically distant from users, leading to inefficient power usage and performance fluctuations.
Innovation Solution
Implementing dynamic load balancing of machine learning operations between edge computing devices and cloud computing systems, where inference operations are dynamically shifted based on environmental factors such as bandwidth and CPU frequency to optimize performance and power efficiency.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Power
If machine learning operations are executed on centralized cloud computing systems, then processing power and computational resources are improved, but latency and bandwidth saturation increase due to geographical distance from users
Solution Approach 1:
The patent segments machine learning operations into two categories: training operations executed on centralized cloud computing systems and inference operations executed on distributed edge computing devices. This segmentation allows the system to leverage the computational power of the cloud while reducing latency for real-time inference by processing data locally at the network edge, thus resolving the contradiction between centralized processing power and distributed low-latency requirements.
2Loss of time
If machine learning operations are executed on edge computing devices, then latency is reduced and bandwidth consumption is decreased, but power consumption increases on mobile devices
Solution Approach 1:
The patent implements dynamic load balancing that adaptively adjusts the distribution of machine learning operations between edge devices and cloud systems based on real-time environmental factors such as device power state, network conditions, and computational workload. This dynamic approach allows the system to optimize the trade-off between latency reduction and power consumption by flexibly migrating operations between execution locations rather than using a static deployment model.
3Productivity
If machine learning operations are dynamically balanced between edge devices and cloud systems, then power efficiency and performance are improved, but system complexity increases
Solution Approach 1:
The patent employs feedback mechanisms where the load balancing system continuously monitors environmental factors including device power state, network bandwidth conditions, and computational performance metrics. Based on this feedback, the system dynamically adjusts the distribution of training and inference operations between edge devices and cloud systems, enabling automated optimization of power efficiency and performance without requiring complex manual configuration or intervention.
Data Source
AI summary
Various embodiments are provided for load balancing of machine learning operations in a computing environment by a processor. One or more machine learning operations performing inference or training operations may by dynamically balanced between one or more edge computing devices in a wireless communication network and a cloud computing system for increasing performance of a selected metric.


