OTN Service Resource Configuration Using Reinforcement Learning
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing optical transport network (OTN) resource configuration methods lack comprehensive optimization, efficiency, and reliability in configuring service resources, particularly in the optical channel layer, due to the lack of a unified algorithm for path and resource allocation.
Innovation Solution
Employing reinforcement learning technology to create a comprehensive optimization model for OTN single service resource configuration, utilizing action policies and reward mechanisms to iteratively optimize resource parameters such as route, wavelength, spectrum, and modulation format, with impairment verification analysis, to achieve optimal resource allocation.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If traditional step-by-step resource configuration methods are used, then the configuration process is simple to implement, but the optimization comprehensiveness and reliability are insufficient
Solution Approach 1:
The patent implements a feedback mechanism through impairment verification (IV) analysis that evaluates the quality of configured service resources and provides feedback to the reinforcement learning model. The model uses this feedback to iteratively optimize its action policy, adjusting resource configuration decisions based on observed outcomes. This closed-loop feedback system enhances configuration reliability by continuously learning from actual network performance.
Solution Approach 2:
The reinforcement learning model performs self-learning and self-optimization through iterative episodes without requiring manual intervention or complex external control systems. The model autonomously improves its resource configuration capabilities by learning from accumulated experience and reward signals, reducing the need for complex human-managed configuration algorithms while maintaining high reliability.
2Manufacturing precision
If reinforcement learning iterative optimization is applied, then the optimization comprehensiveness and convergence are improved, but the calculation time and processing complexity increase
Solution Approach 1:
The patent pre-defines the action policy structure, state space, and reward function before execution. The reinforcement learning model is pre-trained through multiple episodes to learn optimal configuration strategies in advance. This preliminary preparation allows the model to make rapid, precise configuration decisions during actual operation without requiring extensive real-time computation, thus reducing configuration time while maintaining high precision.
Solution Approach 2:
The model performs impairment verification analysis selectively based on configuration confidence and resource criticality. For high-confidence, low-risk configurations, the model may skip exhaustive IV analysis to reduce processing time. For critical or uncertain configurations, full IV analysis is performed to ensure precision. This partial application of verification balances speed and accuracy.
3Measurement precision
If comprehensive impairment verification analysis is performed, then the service quality assessment accuracy is improved, but the processing overhead and complexity increase
Solution Approach 1:
The patent segments the impairment verification analysis into distinct, modular components that evaluate different aspects of service quality independently (e.g., optical signal quality, resource allocation efficiency, configuration validity). Each segment can be processed separately and contributes to the overall quality assessment. This segmentation improves measurement precision by allowing detailed evaluation of each quality dimension while reducing overall complexity through modular processing.
4Productivity
If multiple resource parameters are optimized simultaneously, then the overall resource allocation efficiency is improved, but the control difficulty and computational complexity increase
Solution Approach 1:
The patent employs a universal reinforcement learning model that handles multiple resource parameters (wavelength assignment, spectrum allocation, modulation format selection, routing decisions) through a unified action policy framework. The model learns to coordinate optimization of all parameters simultaneously by integrating them into a single decision-making process, rather than managing each parameter separately. This universal approach improves overall allocation efficiency while the model's learned policies reduce the effective control complexity compared to traditional multi-parameter optimization methods.
Data Source
Figure 1
Figure 2
Figure 3~5
AI summary
The present disclosure provides a service resource configuration method. The method comprises: according to an action policy, configuring resource parameters for a service to be configured, calculating a timely reward in a current state, performing injury validation (IV) analysis according to the action policy after all the resource parameters are configured, and ending one round after the IV analysis is completed, wherein an action in the action policy enters a next state after being completed, and the action comprises an action of configuring a resource parameter or an action of performing IV analysis; calculating and updating optimization target policy parameters in each state according to the timely reward in each state; iterating a preset number of rounds to calculate and update the optimization target policy parameters in each state; respectively determining optimal target policy parameters in each state according to the optimization target policy parameters in each state in the preset number of rounds; and updating the action policy according to the optimal target policy parameters in each state. Embodiments of the present disclosure further provide a single service resource configuration apparatus, a computer device and a computer-readable medium.