Reinforcement Learning for Dynamic Edge Computing Job Assignment
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Infrastructure providers face challenges in accommodating dynamic client requests for intent-based computing jobs in mobile edge computing, as existing methods require manual labeling and pre-processing, and are inefficient in resource allocation and virtual topology design.
Innovation Solution
An unsupervised reinforcement learning approach using two neural networks, a policy neural network and a value neural network, with a novel reward function, action space, and batch processing technique for efficient virtual topology design and resource allocation, allowing for online decision-making and dynamic request handling.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If manual labeling and pre-processing methods are used for intent-based computing job assignment, then the system can handle static requests, but it cannot efficiently accommodate dynamic client requests in real-time
Solution Approach 1:
The patent transforms the static job assignment system into a dynamic one by implementing online reinforcement learning that continuously adapts to changing client requests in real-time. The system dynamically updates policy neural networks and value neural networks based on incoming requests, enabling flexible accommodation of dynamic workload patterns without manual reconfiguration.
Solution Approach 2:
The system employs self-service mechanisms through automated reinforcement learning algorithms that autonomously learn optimal assignment strategies from observed data. The policy neural network and value neural network automatically improve their decision-making capabilities through online learning, eliminating the need for manual labeling and pre-processing while adapting to new request patterns.
2Productivity
If traditional resource allocation methods are used, then the infrastructure can support basic computing jobs, but it cannot efficiently utilize physical resources for intent-based computing
Solution Approach 1:
The patent creates a universal resource allocation framework that handles diverse intent-based computing jobs through a single reinforcement learning system. The policy neural network and value neural network work together to universally manage various types of computing requests (data processing, training, inference) across multiple edge servers, maximizing physical resource utilization without requiring separate specialized systems for each job type.
Solution Approach 2:
The system dynamically changes key parameters including virtual topology configuration, resource allocation ratios, and job assignment decisions based on real-time system state and learned policies. The reinforcement learning algorithm continuously optimizes these parameters to improve resource utilization efficiency while adapting to changing workload characteristics and infrastructure conditions.
3Adaptability or versatility
If online training and decision-making is implemented for dynamic requests, then the system can adapt to changing conditions, but it requires complex neural networks and batch processing techniques
Solution Approach 1:
The patent segments the complex online training process into distinct functional components: a policy neural network for decision-making, a value neural network for evaluation, and a batch processing mechanism for updates. This segmentation allows each component to specialize in specific tasks, managing overall system complexity while enabling sophisticated online adaptation through coordinated interaction between segmented modules.
Data Source
AI summary
An advance in the art is made according to aspects of the present disclosure directed to a method that determines virtual topology design and resource allocation for dynamic intent-based computing jobs in a mobile edge computing infrastructure when client requests are dynamic. Our method according to aspects of the present disclosure is an unsupervised machine learning approach, so that there is no need for manual labeling or pre-processing in advance, while a training process and decision making is performed online. In sharp contrast to the prior art, our method according to aspects of the present disclosure utilizes reinforcement learning techniques to make an efficient assignment in which two neural networks—a policy neural network and a value neural network—are used interactively to achieve the assignment. A training process is performed through a batch (or group) processing style in an online manner.


