OTN Service Resource Configuration Using Reinforcement Learning

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing optical transport network (OTN) resource configuration methods lack comprehensive optimization, efficiency, and reliability in configuring service resources, particularly in the optical channel layer, due to the lack of a unified algorithm for path and resource allocation.

Innovation Solution

Employing reinforcement learning technology to create a comprehensive optimization model for OTN single service resource configuration, utilizing action policies and reward mechanisms to iteratively optimize resource parameters such as route, wavelength, spectrum, and modulation format, with impairment verification analysis, to achieve optimal resource allocation.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If traditional step-by-step resource configuration methods are used, then the configuration process is simple to implement, but the optimization comprehensiveness and reliability are insufficient

Engineering Contradiction:
Improveresource configuration reliabilityVSAvoidconfiguration algorithm complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The patent implements a feedback mechanism through impairment verification (IV) analysis that evaluates the quality of configured service resources and provides feedback to the reinforcement learning model. The model uses this feedback to iteratively optimize its action policy, adjusting resource configuration decisions based on observed outcomes. This closed-loop feedback system enhances configuration reliability by continuously learning from actual network performance.

Inventive Principle:
Principle #23Feedback

Solution Approach 2:

The reinforcement learning model performs self-learning and self-optimization through iterative episodes without requiring manual intervention or complex external control systems. The model autonomously improves its resource configuration capabilities by learning from accumulated experience and reward signals, reducing the need for complex human-managed configuration algorithms while maintaining high reliability.

Inventive Principle:
Principle #25Self-service

2Manufacturing precision

If reinforcement learning iterative optimization is applied, then the optimization comprehensiveness and convergence are improved, but the calculation time and processing complexity increase

Engineering Contradiction:
Improveresource configuration precisionVSAvoidconfiguration time
Core Design Contradiction:
Manufacturing precisionVSLoss of time

Solution Approach 1:

The patent pre-defines the action policy structure, state space, and reward function before execution. The reinforcement learning model is pre-trained through multiple episodes to learn optimal configuration strategies in advance. This preliminary preparation allows the model to make rapid, precise configuration decisions during actual operation without requiring extensive real-time computation, thus reducing configuration time while maintaining high precision.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The model performs impairment verification analysis selectively based on configuration confidence and resource criticality. For high-confidence, low-risk configurations, the model may skip exhaustive IV analysis to reduce processing time. For critical or uncertain configurations, full IV analysis is performed to ensure precision. This partial application of verification balances speed and accuracy.

Inventive Principle:
Principle #16Partial or excessive action

3Measurement precision

If comprehensive impairment verification analysis is performed, then the service quality assessment accuracy is improved, but the processing overhead and complexity increase

Engineering Contradiction:
Improveservice quality measurement precisionVSAvoidverification analysis complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent segments the impairment verification analysis into distinct, modular components that evaluate different aspects of service quality independently (e.g., optical signal quality, resource allocation efficiency, configuration validity). Each segment can be processed separately and contributes to the overall quality assessment. This segmentation improves measurement precision by allowing detailed evaluation of each quality dimension while reducing overall complexity through modular processing.

Inventive Principle:
Principle #1Segmentation

4Productivity

If multiple resource parameters are optimized simultaneously, then the overall resource allocation efficiency is improved, but the control difficulty and computational complexity increase

Engineering Contradiction:
Improveresource allocation efficiencyVSAvoidparameter control complexity
Core Design Contradiction:
ProductivityVSDevice complexity

Solution Approach 1:

The patent employs a universal reinforcement learning model that handles multiple resource parameters (wavelength assignment, spectrum allocation, modulation format selection, routing decisions) through a unified action policy framework. The model learns to coordinate optimization of all parameters simultaneously by integrating them into a single decision-making process, rather than managing each parameter separately. This universal approach improves overall allocation efficiency while the model's learned policies reduce the effective control complexity compared to traditional multi-parameter optimization methods.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Data Source

PatentEP4236345B9Single service resource configuration method and apparatus, computer device and medium
Publication Date: 2026.04.15 ZTE CORP
  • EP4236345B9 patent drawingFigure 1
  • EP4236345B9 patent drawingFigure 2
  • EP4236345B9 patent drawingFigure 3~5

AI summary

The present disclosure provides a service resource configuration method. The method comprises: according to an action policy, configuring resource parameters for a service to be configured, calculating a timely reward in a current state, performing injury validation (IV) analysis according to the action policy after all the resource parameters are configured, and ending one round after the IV analysis is completed, wherein an action in the action policy enters a next state after being completed, and the action comprises an action of configuring a resource parameter or an action of performing IV analysis; calculating and updating optimization target policy parameters in each state according to the timely reward in each state; iterating a preset number of rounds to calculate and update the optimization target policy parameters in each state; respectively determining optimal target policy parameters in each state according to the optimization target policy parameters in each state in the preset number of rounds; and updating the action policy according to the optimal target policy parameters in each state. Embodiments of the present disclosure further provide a single service resource configuration apparatus, a computer device and a computer-readable medium.