DRL Load Balancing for Mobile Edge Computing Networks
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current load balancing methods in mobile edge computing (MEC) networks face challenges in efficiently distributing traffic across base stations, leading to increased end-to-end delay and load variation, particularly due to the restrictive nature of cell individual offset (CIO)-based algorithms which do not adequately consider caching and computational requirements.
Innovation Solution
A deep reinforcement learning (DRL)-based load balancing algorithm that uses multi-agent reinforcement learning (MARL) to determine base station association decisions based on communication, computing, and caching (3C) load components, and joint load status, minimizing the load in the most overloaded base station and reducing end-to-end delay.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Ease of operation
If CIO-based load balancing algorithm is used, then base station association decisions are simplified, but load distribution efficiency deteriorates and end-to-end delay increases
Solution Approach 1:
The patent transforms the load balancing decision from a simple CIO offset adjustment to a multi-parameter optimization problem involving communication load, caching load, and computational load. The DRL agent learns optimal base station associations by considering multiple load components simultaneously, changing the decision parameters from single-dimension to multi-dimension optimization.
Solution Approach 2:
The patent replaces the traditional mechanical CIO-based adjustment mechanism with an intelligent DRL-based decision system. Instead of using fixed or heuristic rules for base station association, the system employs reinforcement learning agents that adaptively learn optimal associations through continuous interaction with the network environment, substituting rule-based mechanics with learning-based intelligence.
2Device complexity
If traditional load balancing methods are used, then system complexity is reduced, but load variation and end-to-end delay worsen
Solution Approach 1:
The patent introduces dynamic adaptability into the load balancing system through DRL agents that continuously learn and adjust base station associations based on real-time network conditions. The system transitions from static or periodically updated load balancing configurations to dynamic, continuously adapting decisions that respond to changing communication, caching, and computational loads across the network.
Solution Approach 2:
The DRL agents perform preliminary learning and exploration during idle periods or low-traffic conditions, building knowledge bases of optimal base station associations for various network states. This preliminary action allows the system to make rapid, informed decisions during high-traffic periods without requiring complex real-time computations, thus reducing end-to-end delay while maintaining intelligent load balancing.
3Ease of manufacture
If CIO-based algorithms are used, then implementation simplicity is maintained, but resource utilization efficiency deteriorates
Solution Approach 1:
The patent creates a universal DRL-based load balancing framework that simultaneously optimizes multiple network resources including communication bandwidth, caching capacity, and computational power. The multi-agent DRL system serves multiple functions: load balancing, resource allocation, and QoS optimization, replacing the single-function CIO adjustment mechanism with a multi-functional intelligent system that handles diverse resource management tasks.
Solution Approach 2:
The DRL agents act as intelligent intermediaries between mobile devices and base stations, making optimized routing and association decisions that balance multiple resource constraints. Instead of direct device-to-base station connections determined by simple signal strength or CIO offsets, the DRL intermediary analyzes comprehensive network state information and mediates the association process to achieve optimal resource utilization across communication, caching, and computational dimensions.
Data Source
AI summary
A method includes obtaining at least one policy parameter of a neural network corresponding to a load balancing policy, receiving trajectories for each mobile device in a plurality of mobile devices of the wireless network, each trajectory corresponding to a sequence of states of a respective mobile device, wherein the sequence of states is generated based on a continuous interaction of an existing policy of the respective mobile device with the wireless network, estimating advantage functions for each mobile device in the plurality of mobile devices based on the trajectories for each respective mobile device, and updating the at least one policy parameter based on the estimated advantage functions such that the load balancing policy is determined based on states of each mobile device in the plurality of mobile devices.


