Reinforcement Learning for OTN Service Creation and Resource Allocation

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing methods for optimizing optical transport network (OTN) resources are inefficient in achieving global resource allocation optimization, leading to suboptimal capital expenditure (CAPEX)/operation expenditure (OPEX) and transmission performance.

Innovation Solution

A method and apparatus utilizing reinforcement learning to determine an action policy for creating OTN services, calculating timely rewards, and updating optimized objective policy parameters to rank service creations, ensuring efficient resource allocation and global optimization.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If traditional methods are used for optimizing OTN resources, then the optimization process is simpler, but the convergence and reliability of resource allocation optimization deteriorates

Engineering Contradiction:
Improvereliability of OTN resource optimizationVSAvoidcomplexity of optimization method
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The patent implements feedback mechanisms through reward functions that evaluate service creation decisions. The reinforcement learning agent receives feedback in the form of rewards or penalties based on the quality of resource allocation decisions, enabling continuous improvement of optimization reliability through iterative learning from environmental feedback.

Inventive Principle:
Principle #23Feedback

Solution Approach 2:

The system employs self-service through autonomous reinforcement learning agents that independently optimize OTN resource allocation without requiring external intervention. The agents learn optimal policies through self-interaction with the environment, making the optimization process self-improving and highly reliable.

Inventive Principle:
Principle #25Self-service

2Reliability

If reinforcement learning is applied to optimize OTN resources, then the convergence and reliability improve, but the computational complexity and training time increases

Engineering Contradiction:
Improveconvergence of optimizationVSAvoidtraining time
Core Design Contradiction:
ReliabilityVSLoss of time

Solution Approach 1:

The patent applies preliminary action by pre-training reinforcement learning agents offline before deployment. The agents undergo extensive training in simulated environments to learn optimal policies beforehand, so that during actual operation, they can quickly converge and make reliable decisions without requiring extensive real-time training.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The system uses copying by creating virtual copies of the OTN environment for training purposes. Multiple simulated environments allow parallel training of reinforcement learning agents, significantly reducing training time compared to single-environment training, while maintaining the reliability benefits of comprehensive learning.

Inventive Principle:
Principle #26Copying

3Productivity

If reinforcement learning is used for service creation ranking, then the global resource allocation optimization improves, but the system complexity increases

Engineering Contradiction:
Improveresource allocation efficiencyVSAvoidsystem complexity
Core Design Contradiction:
ProductivityVSDevice complexity

Solution Approach 1:

The patent applies segmentation by dividing the complex resource allocation problem into smaller sub-problems handled by specialized reinforcement learning agents. Each agent focuses on specific aspects of service creation and resource allocation, making the overall system more manageable and less complex while achieving high global optimization through coordinated agent actions.

Inventive Principle:
Principle #1Segmentation

Data Source

PatentUS12401442B2Method for optimizing OTN resources, computer device and storage medium
Publication Date: 2025.08.26 ZTE CORP
  • US12401442B2 patent drawing
  • US12401442B2 patent drawing

AI summary

The present disclosure provides a method for optimizing OTN resources, including: determining and creating, according to an action policy, a service to be created in a current service creation state, calculating a timely reward in the current service creation state, entering a next service creation state until an episode is ended, and calculating and updating, according to the timely reward in each service creation state, an optimized objective policy parameter in each service creation state; iterating a preset number of episodes to calculate and update the optimized objective policy parameter in each service creation state; determining, according to the optimized objective policy parameter in each service creation state in the preset number of episodes, a resultant optimized objective policy parameter in each service creation state; and updating the action policy according to the resultant optimized objective policy parameter in each service creation state.