Reinforcement Learning for OTN Service Creation and Resource Allocation
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing methods for optimizing optical transport network (OTN) resources are inefficient in achieving global resource allocation optimization, leading to suboptimal capital expenditure (CAPEX)/operation expenditure (OPEX) and transmission performance.
Innovation Solution
A method and apparatus utilizing reinforcement learning to determine an action policy for creating OTN services, calculating timely rewards, and updating optimized objective policy parameters to rank service creations, ensuring efficient resource allocation and global optimization.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If traditional methods are used for optimizing OTN resources, then the optimization process is simpler, but the convergence and reliability of resource allocation optimization deteriorates
Solution Approach 1:
The patent implements feedback mechanisms through reward functions that evaluate service creation decisions. The reinforcement learning agent receives feedback in the form of rewards or penalties based on the quality of resource allocation decisions, enabling continuous improvement of optimization reliability through iterative learning from environmental feedback.
Solution Approach 2:
The system employs self-service through autonomous reinforcement learning agents that independently optimize OTN resource allocation without requiring external intervention. The agents learn optimal policies through self-interaction with the environment, making the optimization process self-improving and highly reliable.
2Reliability
If reinforcement learning is applied to optimize OTN resources, then the convergence and reliability improve, but the computational complexity and training time increases
Solution Approach 1:
The patent applies preliminary action by pre-training reinforcement learning agents offline before deployment. The agents undergo extensive training in simulated environments to learn optimal policies beforehand, so that during actual operation, they can quickly converge and make reliable decisions without requiring extensive real-time training.
Solution Approach 2:
The system uses copying by creating virtual copies of the OTN environment for training purposes. Multiple simulated environments allow parallel training of reinforcement learning agents, significantly reducing training time compared to single-environment training, while maintaining the reliability benefits of comprehensive learning.
3Productivity
If reinforcement learning is used for service creation ranking, then the global resource allocation optimization improves, but the system complexity increases
Solution Approach 1:
The patent applies segmentation by dividing the complex resource allocation problem into smaller sub-problems handled by specialized reinforcement learning agents. Each agent focuses on specific aspects of service creation and resource allocation, making the overall system more manageable and less complex while achieving high global optimization through coordinated agent actions.
Data Source
AI summary
The present disclosure provides a method for optimizing OTN resources, including: determining and creating, according to an action policy, a service to be created in a current service creation state, calculating a timely reward in the current service creation state, entering a next service creation state until an episode is ended, and calculating and updating, according to the timely reward in each service creation state, an optimized objective policy parameter in each service creation state; iterating a preset number of episodes to calculate and update the optimized objective policy parameter in each service creation state; determining, according to the optimized objective policy parameter in each service creation state in the preset number of episodes, a resultant optimized objective policy parameter in each service creation state; and updating the action policy according to the resultant optimized objective policy parameter in each service creation state.

