Reinforcement Learning Agent for Dynamic Bandwidth Scheduling
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Static bandwidth management strategies in polymorphic smart networks fail to adapt to dynamic changes in network modal traffic, leading to potential network overload and communication quality issues.
Innovation Solution
A reinforcement learning agent training method and apparatus that constructs a deep neural network model with a new and old execution network and action evaluation network to dynamically schedule bandwidth resources based on global network characteristics, using reinforcement learning algorithms to continuously interact with the network and adjust actions according to changing states and reward values.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If static bandwidth management strategies are used, then network resource allocation is simple to implement, but the system cannot adapt to dynamic changes in network modal traffic
Solution Approach 1:
The patent applies reinforcement learning to create a dynamic bandwidth management system where the agent continuously learns optimal bandwidth allocation strategies through interaction with the network environment. The system adapts to changing traffic patterns by updating its policy based on observed states and rewards, transforming the static management approach into a dynamic one that responds to real-time network conditions.
Solution Approach 2:
The reinforcement learning framework implements continuous feedback loops where the agent observes network states, executes bandwidth allocation actions, receives reward signals based on network performance, and updates its policy accordingly. This feedback mechanism enables the system to learn from past decisions and improve future bandwidth allocation, resolving the contradiction between simplicity and adaptability.
2Reliability
If static bandwidth limiting strategies are applied, then network overload is prevented, but communication transmission quality deteriorates under dynamic traffic conditions
Solution Approach 1:
The reinforcement learning agent dynamically adjusts bandwidth allocation to maintain reliable communication under varying traffic conditions. By continuously learning from network states and performance outcomes, the system optimizes bandwidth distribution to ensure quality of service while adapting to dynamic traffic patterns, preventing both overload and quality deterioration.
Solution Approach 2:
The system changes bandwidth allocation parameters dynamically based on learned policies. Instead of fixed bandwidth limits, the reinforcement learning agent adjusts bandwidth parameters in response to changing network states, maintaining communication quality while adapting to different traffic scenarios through parameter optimization.
3Productivity
If reinforcement learning training is performed with frequent updates, then bandwidth allocation optimality improves, but training time and computational resources increase
Solution Approach 1:
The patent employs incremental learning where the reinforcement learning agent performs partial updates based on sampled experiences rather than requiring complete retraining. The agent processes experiences in batches and performs incremental policy updates, achieving good bandwidth allocation efficiency without the excessive computational burden of frequent full-model retraining.
Solution Approach 2:
The system uses experience replay mechanisms where past experiences are stored and reused for training. By copying and reusing historical state-action-reward tuples, the system maximizes learning from available data, improving bandwidth allocation efficiency while reducing the time required for new training iterations through efficient data utilization.
Data Source
AI summary
The present disclosure discloses a reinforcement learning agent training method, modal bandwidth resource scheduling method and apparatus. The reinforcement learning agent training method utilizes a reinforcement learning agent to continuously interact with a network environment in a polymorphic smart network to obtain the latest global network characteristics and output updated actions. By adjusting the bandwidth occupied by modals, a reward value is set to determine an optimization target for the agent, the scheduling of modals is realized, and the rational use of polymorphic smart network resources is guaranteed. The trained reinforcement learning agent is applied to the modal bandwidth resource scheduling method, and can adapt to networks with different characteristics, and thus can be used for intelligent management and control of polymorphic smart networks and has good adaptability and scheduling performance.


