Single service resource configuration method, device, computer equipment and medium

By configuring single-service resources in the OTN network through a reinforcement learning algorithm, the problem of inconsistent resource configuration in existing technologies is solved, single-service resource optimization with high convergence and reliability is achieved, and an OTN service path with comprehensive indicator optimization is provided.

CN114584865BActive Publication Date: 2025-10-03ZTE CORP
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202011293457.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2020-11-18
Publication Date
2025-10-03
Estimated Expiration
2040-11-18

AI Technical Summary

Technical Problem

The existing OTN network single-service resource allocation scheme lacks a unified algorithm, resulting in poor resource allocation results and insufficient comprehensive optimization and efficiency.

Method used

Using reinforcement learning algorithms, resource parameters are configured through action strategies, timely rewards are calculated, and damage confirmation IV analysis is performed. The target strategy parameters are iteratively optimized to determine the optimal optimization target strategy, update the action strategy, and achieve comprehensive optimization of single business resources.

Benefits of technology

It improves the convergence and reliability of single-service resource configuration in the OTN network, provides service paths optimized with comprehensive indicators, ensures the perfect combination of reinforcement learning technology and the OTN network, and realizes intelligent optimization.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114584865B_ABST
    Figure CN114584865B_ABST
Patent Text Reader

Abstract

The present disclosure provides a method for configuring business resources. The method configures resource parameters for a business to be configured according to an action strategy, calculates timely rewards in the current state, and after all resource parameters are configured, performs damage confirmation IV analysis according to the action strategy. A round ends, wherein after an action is completed, the next state is entered. The action includes configuring a resource parameter action or performing an IV analysis action; calculating and updating the optimization target strategy parameters in each state based on the timely rewards in each state; iterating a preset number of rounds to calculate and update the optimization target strategy parameters in each state; determining the optimal optimization target strategy parameters in each state based on the optimization target strategy parameters in each state in the preset number of rounds; and updating the action strategy based on the optimal optimization target strategy parameters in each state. The present disclosure also provides a single-business resource configuration device, a computer device, and a computer-readable medium.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of artificial intelligence technology, and in particular to a single-service resource configuration method, apparatus, computer equipment, and computer-readable medium. Background Art

[0002] With the development of artificial intelligence (AI), reinforcement learning (RL) is gaining increasing attention across various fields and industries. Reinforcement learning (RL), also known as re-enforcement learning or evaluation learning, is an important machine learning method with numerous applications in areas such as intelligent robotic control and network analysis and prediction. Within the connectionist school of machine learning, learning algorithms are categorized into three types: unsupervised learning, supervised learning, and reinforcement learning.

[0003] Reinforcement learning involves an agent learning through trial and error. Rewards earned through interaction with the environment guide behavior, with the goal of maximizing the reward. Reinforcement learning differs from supervised learning in connectionist learning primarily in its reinforcement signals. In reinforcement learning, the reinforcement signals provided by the environment evaluate the quality of an action (typically a scalar signal), rather than telling the reinforcement learning system (RLS) how to perform the correct action. Because the external environment provides little information, RLS must rely on its own experience to learn. In this way, RLS acquires knowledge in an action-evaluation environment and refines its action plans to adapt to the environment.

[0004] In recent years, with the application and promotion of reinforcement learning technology, how to apply the advantages of this technology to the field of intelligent management, control, and operation and maintenance of OTN (Optical Transport Network), especially the application of reinforcement learning in the optimization of service resource allocation at the optical layer of OTN network, has attracted widespread attention from OTN experts.

[0005] The OTN network local single-channel resource optimization (LSO) solution based on the SDON (Software Defined Optical Network) architecture is as follows: Figure 1As shown in the figure, in the SDON architecture, the Path Computation Element (PCE) is primarily responsible for performing route calculation and resource allocation for OTN network services, providing optimized paths that meet cost and other policy objectives. Based on this, it performs resource allocation and evaluation, including RWA (Routing and Wavelength Assignment), RSA (Routing and Spectrum Assignment), SDO (Software Defined Optics), and IV (Impairment Verification), ultimately obtaining a service resource path that meets comprehensive optimization criteria. Traditional OTN network resource allocation solutions for single services implement path resource calculation and allocation in a step-by-step manner, rather than being fully implemented within a unified algorithm. Consequently, these solutions suffer from limitations in terms of service resource optimization effectiveness, overall optimization level, optimization efficiency, and theoretical rigor of the optimization algorithm. Summary of the Invention

[0006] In view of the above-mentioned deficiencies in the prior art, the present disclosure provides a single-service resource configuration method, apparatus, computer device, and computer-readable medium.

[0007] In a first aspect, an embodiment of the present disclosure provides a single-service resource configuration method, including:

[0008] Configure resource parameters for the business to be configured according to the action strategy and calculate the timely reward in the current state. After all resource parameters are configured, perform damage confirmation IV analysis according to the action strategy. One round ends. After an action is completed, enter the next state. The action includes configuring a resource parameter action or performing IV analysis action.

[0009] Calculate and update the optimization target strategy parameters in each state based on the timely rewards in each state;

[0010] Iterate a preset number of rounds to calculate and update the optimized target strategy parameters under each state;

[0011] Determine the optimal optimization target strategy parameters in each state according to the optimization target strategy parameters in each state in the preset number of rounds;

[0012] The action strategy is updated according to the optimal optimization target strategy parameters in each state.

[0013] In some embodiments, the resource parameters include routing, wavelength, spectrum, and modulation format. In one round, the resource parameters are configured for the service to be configured in the following order: routing, wavelength, spectrum, and modulation format.

[0014] In some embodiments, the states include: a pending routing state, a pending wavelength state, a pending spectrum state, a pending modulation format state, a pending IV analysis state, and a terminated state. The timely reward R0 in the pending routing state is 0, and the timely rewards in other states satisfy one or any combination of the following:

[0015] The timely reward R1 in the state of the wavelength to be configured is a function of the working routing cost, and R1 is in a monotonically decreasing relationship with the function of the working routing cost;

[0016] The timely reward R2 in the state of the spectrum to be configured is a function of the wavelength resource utilization rate, and R2 and the function of the wavelength resource utilization rate are in a monotonically increasing relationship;

[0017] The timely reward R3 in the modulation format to be configured state is a function of the spectrum width occupied by the service, and R3 and the function of the spectrum width occupied by the service are in a monotonically decreasing relationship;

[0018] The timely reward R4 in the pending IV analysis state is a function of the service spectrum efficiency, and R4 and the function of the service spectrum efficiency are in a monotonically increasing relationship;

[0019] The timely reward R5 in the termination state is related to the IV analysis result. When the IV analysis result is qualified, R5 is a positive number, and when the IV analysis result is unqualified, R5 is a negative number.

[0020] In some embodiments, the action strategy includes a random action strategy and a deterministic action strategy, and configuring resource parameters for the service to be configured according to the action strategy includes: configuring routing, wavelength, spectrum, and modulation format for the service to be configured according to the random action strategy;

[0021] The performing damage confirmation IV analysis according to the action strategy includes: performing IV analysis according to the deterministic action strategy.

[0022] In some embodiments, when the route of the service to be configured includes multiple relay segments, configuring resource parameters for the service to be configured according to the action policy includes: configuring resource parameters for the service to be configured according to the action policy in each relay segment; performing damage confirmation IV analysis according to the action policy includes: performing damage confirmation IV analysis according to the action policy in each relay segment;

[0023] The damage confirmation IV analysis includes:

[0024] Calculating the budget value of the optical signal-to-noise ratio of each relay segment in the route of the service to be configured respectively;

[0025] In response to the budgeted values ​​of the optical signal-to-noise ratios of the respective relay segments satisfying a preset condition, determining that the IV analysis result is qualified;

[0026] In response to a budget value of an optical signal-to-noise ratio of at least one relay segment not satisfying a preset condition, the IV analysis result is determined to be unqualified.

[0027] In some embodiments, performing an injury confirmation IV analysis comprises:

[0028] Calculating a budget value of an optical signal-to-noise ratio of the service to be configured;

[0029] In response to the budget value of the optical signal-to-noise ratio meeting a preset condition, determining that the IV analysis result is qualified;

[0030] In response to the budget value of the optical signal-to-noise ratio not meeting a preset condition, the IV analysis result is determined to be unqualified.

[0031] In some embodiments, the budget value of the optical signal-to-noise ratio satisfies a preset condition, including:

[0032] OSNR 预算值 -OSNR 平坦度 ≥OSNR 传输门限

[0033] OSNR 传输门限 =OSNR B2B +OSNR 非线性 +OSNR CD +OSNR PMD +OSNR 滤波 +OSNR PDL +OSNR 波动 +OSNR 净余量

[0034] Among them, OSNR 预算值 is the budget value of optical signal-to-noise ratio, OSNR 平坦度 is the flatness of the optical signal-to-noise ratio, OSNR B2B is the back-to-back optical signal-to-noise ratio, OSNR 非线性 is the nonlinear cost of optical signal-to-noise ratio, OSNR CD is the dispersion penalty of optical signal-to-noise ratio, OSNR PMD is the optical signal-to-noise ratio (OSNR) of the polarization film dispersion penalty. 滤波 The filter cost of optical signal-to-noise ratio, OSNR PDL Polarization-dependent loss penalty for optical signal-to-noise ratio, OSNR 波动 is the fluctuation of optical signal-to-noise ratio, OSNR 净余量The net margin required for the optical signal-to-noise ratio; OSNR 平坦度 、OSNR B2B 、OSNR 非线性 、OSNR CD 、OSNR PMD 、OSNR 滤波 、OSNR PDL 、OSNR 波动 、OSNR 净余量 is the default value.

[0035] In some embodiments, the calculation and updating of the optimization target strategy parameters in each state based on the timely rewards in each state includes:

[0036] Calculate the expected return in the current state based on the immediate rewards in each state after the next state;

[0037] Calculate and update the optimization target strategy parameters in the current state based on the expected return in the current state.

[0038] In some embodiments, the expected return under the current state is calculated according to the following formula:

[0039]

[0040] Among them, G t State S t Next, perform action a t The expected return, γ is the discount coefficient, 0<γ<1; R is the immediate reward, t is the state S t The number of configured resource parameters, t = (0, ..., n-1), n-1 is the total number of resource parameters.

[0041] In some embodiments, the optimization target strategy parameters include the state behavior value Q π (s,a), or,

[0042] The optimization target strategy parameters include the state value V π (s), Among them, π(a|s) is the probability of taking action a according to the action strategy π(s,a) in state S, and A is the set of actions performed in each state.

[0043] In some embodiments, when the optimization target strategy parameter is the state behavior value Q π (s, a), the Monte Carlo algorithm, the temporal difference algorithm with different strategies, or the temporal difference algorithm with the same strategy are used to calculate and update the optimization target strategy parameters in each state;

[0044] The updating of the action strategy according to the optimal optimization target strategy parameters in each state includes: π (s,a) Update the action strategy.

[0045] In some embodiments, when the optimization target strategy parameter is the state value V π (s), a dynamic programming algorithm is used to calculate the optimization target strategy parameters;

[0046] The updating of the action strategy according to the optimal optimization target strategy parameters in each state includes: π (s) Update the action strategy.

[0047] In another aspect, the present disclosure further provides a single service resource configuration device, comprising: a first processing module, a second processing module, and an updating module.

[0048] The first processing module is used to configure resource parameters for the service to be configured according to the action strategy and calculate the timely reward in the current state. After all resource parameters are configured, damage confirmation IV analysis is performed according to the action strategy, and one round ends. After an action is completed, the next state is entered. The action includes configuring a resource parameter action or performing an IV analysis action; calculating and updating the optimization target strategy parameters in each state according to the timely reward in each state; iterating a preset number of rounds to calculate and update the optimization target strategy parameters in each state;

[0049] The second processing module is used to determine the optimal optimization target strategy parameters in each state according to the optimization target strategy parameters in each state in the preset number of rounds;

[0050] The updating module is used to update the action strategy according to the optimal optimization target strategy parameters in each state.

[0051] In another aspect, an embodiment of the present disclosure further provides a computer device, comprising:

[0052] one or more processors;

[0053] a storage device having one or more programs stored thereon;

[0054] When the one or more programs are executed by the one or more processors, the one or more processors implement the single-service resource configuration method as described above.

[0055] On the other hand, an embodiment of the present disclosure further provides a computer-readable medium having a computer program stored thereon, wherein the program implements the business resource configuration method as described above when executed.

[0056] The single-service resource configuration method and device provided by the embodiment of the present disclosure configure resource parameters for the service to be configured according to the action strategy, and calculate the timely reward in the current state. After all resource parameters are configured, damage confirmation IV analysis is performed according to the action strategy, and one round ends. After one action is completed, the next state is entered. The action includes configuring a resource parameter action or performing an IV analysis action; calculating and updating the optimization target policy parameters in each state according to the timely reward in each state; iterating a preset number of rounds to calculate and update the optimization target policy parameters in each state; determining the optimal optimization target policy parameters in each state according to the optimization target policy parameters in each state in the preset number of rounds; and updating the action policy according to the optimal optimization target policy parameters in each state. The embodiment of the present disclosure utilizes the reward and punishment mechanism of the reinforcement learning algorithm to comprehensively optimize multiple resources and performance indicators, optimize the resource configuration of a single OTN network service, and then provide users with an OTN service path optimized with comprehensive indicators. The obtained action strategy has good convergence, rigor and high reliability. BRIEF DESCRIPTION OF THE DRAWINGS

[0057] Figure 1 This is a schematic diagram of single-service resource configuration in an OTN network under the SDON architecture;

[0058] Figure 2 A schematic diagram of a single service resource configuration process provided in an embodiment of the present disclosure;

[0059] Figure 3 A schematic diagram of a process for performing IV analysis provided in an embodiment of the present disclosure;

[0060] Figure 4 A schematic diagram of a process for calculating optimization target strategy parameters provided by an embodiment of the present disclosure;

[0061] Figure 5 A schematic diagram of the structure of a single-service resource configuration device provided in an embodiment of the present disclosure. DETAILED DESCRIPTION

[0062] Example embodiments will be described more fully hereinafter with reference to the accompanying drawings, but the example embodiments may be embodied in different forms and should not be construed as limited to the embodiments set forth herein. Rather, these embodiments are provided so that this disclosure will be thorough and complete and will fully convey the scope of this disclosure to those skilled in the art.

[0063] As used herein, the term "and / or" includes any and all combinations of one or more of the associated listed items.

[0064] The terms used herein are used only to describe specific embodiments and are not intended to limit the present disclosure. As used herein, the singular forms "a," "an," and "the" are also intended to include the plural forms, unless the context clearly indicates otherwise. It will also be understood that when the terms "comprising" and / or "made of" are used in this specification, the presence of the features, wholes, steps, operations, elements, and / or components is specified, but the presence or addition of one or more other features, wholes, steps, operations, elements, components, and / or groups thereof is not excluded.

[0065] The embodiments described herein may be described with reference to plan views and / or cross-sectional views, with the aid of idealized schematic diagrams of the present disclosure. Thus, the example illustrations may be modified based on manufacturing techniques and / or tolerances. Therefore, the embodiments are not limited to the embodiments shown in the accompanying drawings, but include modifications of the configurations formed based on the manufacturing process. Therefore, the regions illustrated in the accompanying drawings are schematic in nature, and the shapes of the regions shown in the drawings illustrate specific shapes of the regions of the elements, but are not intended to be limiting.

[0066] Unless otherwise defined, all terms (including technical and scientific terms) used herein have the same meaning as commonly understood by one of ordinary skill in the art. It will also be understood that terms such as those defined in commonly used dictionaries should be interpreted as having a meaning consistent with their meaning in the context of the relevant art and the present disclosure, and will not be interpreted as having an idealized or overly formal meaning unless expressly defined as such herein.

[0067] The disclosed embodiments use reinforcement learning technology to implement the creation and comprehensive optimization of single-line services in the OTN network. The reinforcement learning algorithm model is designed to closely follow the processes and rules of the SDON management and control system for single-line service creation and resource allocation. This ensures that reinforcement learning technology comprehensively covers all processes and functional steps of the comprehensive optimization of single-line OTN services, and perfectly integrates reinforcement learning technology with the scenario of comprehensive optimization of single-line OTN services. This ensures that reinforcement learning technology solves the comprehensive optimization problem of single-line OTN services in a tailored and timely manner, and fully utilizes the intelligent optimization function of reinforcement learning for single-line OTN services.

[0068] The single OTN network service mentioned in the embodiments of the present disclosure refers to the optical layer, that is, the OCh (Optical Channel) layer, also known as the wavelength-level service of the L0 layer. In the process of optimizing and calculating a single OTN service, the network environment in which the OTN service is located (mainly including the network topology status, the number of other services in the OTN network, routing, resource allocation status, etc.) is required not to change.

[0069] The relevant parameters of the reinforcement learning algorithm model are defined as follows:

[0070] 1. Define Episode

[0071] A complete Episode is defined as the entire process of determining the route, configuring wavelengths, configuring spectrum resources, configuring the modulation format (SDO), and performing impairment verification (IV) analysis for an OTN network service using a certain action strategy.

[0072] 2. Define action a t and action strategies

[0073] The action includes configuring a resource parameter action or performing an IV analysis action. The resource parameters may include: routing, wavelength, spectrum resources, modulation format, t = (0, ..., n-1), n-1 is the total number of resource parameters. In the embodiment of the present disclosure, n-1 = 4, that is, one round includes 5 actions, namely a0-a4.

[0074] The action strategy π(s,a) includes random action strategy and deterministic action strategy. The random action strategy can be represented by π'(s,a), and the deterministic action strategy can be represented by μ(s,a) (also written as μ(s)).

[0075] The action set for configuring a single OTN network service resource is as follows:

[0076] (1) a0 action set: configures routing actions. Based on routing constraints (including must-pass and must-avoid nodes and links, etc.) and the differences in routing selection, the a0 action set includes multiple routing actions such as assigning A route, B route, C route, etc. to the service to be configured.

[0077] (2) a1 action set: Configure wavelength actions. This action sets wavelengths for the service routes to be configured, ensuring wavelength consistency and continuity. Based on wavelength values, the a1 action set includes multiple wavelength configuration actions, such as wavelength L, wavelength M, wavelength N, and so on, for the service routes to be configured. It should be noted that if a route includes relay spans, different relay spans can use different wavelengths.

[0078] (3) Action set a2: Configure spectrum actions, which configure spectrum for the service route to be configured. Based on the bandwidth value, the action set a2 includes multiple Configure Spectrum d actions, such as bandwidth x, bandwidth y, bandwidth z, and so on, for the service route to be configured. It should be noted that if the route includes relay spans, different relay spans can use different spectrum widths.

[0079] (4) Action Set a3: Configure SDO (Modulation Format) actions, which configure the modulation format for the service route to be configured. Based on the value of the modulation format attribute, the action set a3 includes multiple SDO configuration actions, such as assigning modulation format i, modulation format j, modulation format k, etc. to the service route to be configured. It should be noted that if the route includes relay spans, different relay spans can use different modulation formats.

[0080] (5) Action set a4: IV analysis action, based on the service route to be configured and the network resources configured along the route, IV analysis is performed along the route. It should be noted that if the route includes relay spans, different back-to-back optical signal-to-noise ratios (OSNRs) need to be used in different relay spans. B2B Perform IV analysis.

[0081] In a multi-relay span scenario, due to the different resource configuration and IV analysis of each span, under the premise of satisfying the resource configuration constraints and attribute setting constraints of the OTN network for each service, each action set can be further split into multiple action sets based on the relay span, such as a 11 The action set represents the action of allocating wavelengths to the first relay span of the service route to be configured. 12 The action set indicates the action of allocating wavelengths to the second relay span of the service route to be configured.

[0082] 3. Define the state S during the configuration of a single business resource t

[0083] Each state in a round is S t , t=(0,…,n), n is the total number of resource parameters+1. In the embodiment of the present disclosure, the total number of resource parameters is 4, n=5, so one round includes 6 states, namely S0-S5. S0 is the initial state, that is, the state of waiting for routing configuration. In this state, no resource parameters are configured; S1 is the state of waiting for wavelength configuration. In this state, routing has been configured (that is, routing configuration action a0 has been executed) but wavelength has not been configured; S2 is the state of waiting for spectrum configuration. In this state, wavelength has been configured (that is, wavelength configuration action a1 has been executed) but spectrum has not been configured; S3 is the state of waiting for modulation format configuration. In this state, spectrum has been configured (that is, spectrum configuration action a2 has been executed) but modulation format has not been configured; S4 is the state of IV analysis. In this state, modulation format has been configured (that is, modulation format configuration action a3 has been executed) but IV analysis has not been performed; S5 is the termination state. In this state, all resource parameters have been configured and IV analysis has been completed. Once the termination state is entered, it means that one round is over.

[0084] Taking the service request from node A to node D as an example, in state S0, the routing configuration action a0 is executed to select a working route from the alternative routes; in state S1, the wavelength configuration action a1 is executed to allocate a wavelength to the working route; in state S2, the spectrum configuration action a2 is executed to allocate a spectrum to the route with the allocated wavelength; in state S3, the SDO configuration action a3 is executed to configure the modulation format for the route with the allocated wavelength and spectrum; in state S4, the IV analysis action a4 is executed to perform IV analysis on the route with the allocated wavelength, spectrum, and SDO; in state S5, the IV analysis is completed, and this round ends.

[0085] In a multi-relay span scenario, the next states S1, S2, S3, S4, and S5 corresponding to actions a0, a1, a2, a3, and a4 can be split into multiple states as the corresponding actions are split across different relay spans. In other words, states can be divided by relay span. In a multi-relay span scenario, the round ends only when all relay spans enter the S5 terminal state.

[0086] The present disclosure provides a method for configuring single service resources. Figure 2 As shown, the method includes the following steps:

[0087] Step 11: Configure resource parameters for the business to be configured according to the action strategy, and calculate the timely reward under the current state. After all resource parameters are configured, IV analysis is performed according to the action strategy. One round ends, wherein, after an action is completed, the next state is entered. The action includes configuring a resource parameter action or performing an IV analysis action.

[0088] In this step, in one round, resource parameters are configured for the service to be configured according to the action strategy π(s, a). After each resource parameter is configured, the immediate reward for that state is calculated, the current state ends, and the next state is entered. Following the above steps, each resource parameter is configured in one round, and the immediate reward for the corresponding state is calculated. After all resource parameters are configured, IV analysis is performed, and the round ends.

[0089] Step 12: Calculate and update the optimization target strategy parameters in each state based on the timely rewards in each state.

[0090] In this step, different algorithms can be used to calculate and update the optimization target strategy parameters. It should be noted that different algorithms are used, and the optimization target strategy parameters are also different. Various algorithms will be explained in detail later.

[0091] Step 13: Iterate a preset number of rounds to calculate and update the optimized target strategy parameters in each state.

[0092] In this step, steps 11-12 are repeated for a preset number of rounds, and the optimized target strategy parameters under each state in each round are calculated and updated.

[0093] Step 14: Determine the optimal optimization target strategy parameters in each state according to the optimization target strategy parameters in each state in a preset number of rounds.

[0094] In this step, for each state, the optimal optimization target policy parameters for that state are determined from the optimization target policy parameters from different rounds. It should be noted that the method for determining the optimal optimization target policy parameters varies depending on the algorithm used. This step yields the optimal optimization target policy parameters for all states of the service to be configured.

[0095] Step 15: Update the action strategy according to the optimal optimization target strategy parameters in each state.

[0096] The optimization target strategy parameters are used to characterize states and actions. Once the optimal optimization target strategy parameters in a certain state are determined, the optimal action a in that state can be determined. t , the optimal action a t That is, the action of configuring the optimal resource parameters in this state can determine the optimal resource parameters in this state, thereby obtaining the action set of all optimal resource parameters, which is the optimized action strategy π(s,a).

[0097] The single-service resource configuration method and device provided by the embodiment of the present disclosure configure resource parameters for the service to be configured according to the action strategy, and calculate the timely reward in the current state. After all resource parameters are configured, damage confirmation IV analysis is performed according to the action strategy, and one round ends. After one action is completed, the next state is entered. The action includes configuring a resource parameter action or performing an IV analysis action; calculating and updating the optimization target policy parameters in each state according to the timely reward in each state; iterating a preset number of rounds to calculate and update the optimization target policy parameters in each state; determining the optimal optimization target policy parameters in each state according to the optimization target policy parameters in each state in the preset number of rounds; and updating the action policy according to the optimal optimization target policy parameters in each state. The embodiment of the present disclosure utilizes the reward and punishment mechanism of the reinforcement learning algorithm to comprehensively optimize multiple resources and performance indicators, optimize the resource configuration of a single OTN network service, and then provide users with an OTN service path optimized with comprehensive indicators. The obtained action strategy has good convergence, rigor and high reliability.

[0098] In some embodiments, resource parameters may include routing, wavelength, spectrum, and modulation format. In a single round, resource parameters are configured for the service to be configured in the following order: routing, wavelength, spectrum, and modulation format. It should be noted that while the disclosed embodiments illustrate the order of routing, wavelength, spectrum, and modulation format configuration, those skilled in the art will appreciate that the order of configuring the resource parameters, as well as the type and number of resource parameters, is not limited, as long as IV analysis is performed after all resource parameters are configured.

[0099] In some embodiments, the states include: a routing state to be configured S0, a wavelength state to be configured S1, a spectrum state to be configured S2, a modulation format state to be configured S3, an IV analysis state to be configured S4, and a termination state S5. t Indicates state S t The immediate reward obtained under state S t-1 Next, perform action a t-1 Then migrate to state S t The immediate reward obtained when t is the state S t The number of resource parameters configured under t = (0, ..., n-1), where n-1 is the total number of resource parameters. It should be noted that the timely reward R0 in the routing state S0 to be configured is 0, and the timely rewards in other states meet one or any combination of the following (1)-(5):

[0100] (1) The timely reward R1 in the state S1 of the wavelength to be configured is a function of the working routing cost, and R1 is in a monotonically decreasing relationship with the function of the working routing cost; that is, R1 can be a function of the working routing cost SvcCost obtained after the service to be configured performs action a0, and the two are in a monotonically decreasing relationship.

[0101] (2) The timely reward R2 in the state of the spectrum to be configured is a function of the wavelength resource utilization rate, and R2 is in a monotonically increasing relationship with the function of the wavelength resource utilization rate; that is, R2 can be the timely reward obtained by the service to be configured after obtaining the working routing wavelength resource after executing action a1. Under the conditions of meeting the wavelength consistency and continuity constraints, R2 can be the current network wavelength resource utilization rate U λ function, and there is a monotonically increasing relationship between the two.

[0102] (3) The timely reward R3 in the modulation format state to be configured is a function of the spectrum width occupied by the service, and R3 is in a monotonically decreasing relationship with the function of the spectrum width occupied by the service; that is, R3 can be the timely reward obtained by the service to be configured after executing action a2 to obtain the spectrum resources of the working route. Under the constraint condition that the minimum bandwidth usage threshold of the service to be configured is met, R3 can be the spectrum width F currently occupied by the service to be configured. w function, and there is a monotonically decreasing relationship between the two.

[0103] (4) The timely reward R4 in the pending IV analysis state is a function of the service spectrum efficiency, and R4 is in a monotonically increasing relationship with the function of the service spectrum efficiency; that is, R4 can be the timely reward obtained by the service to be configured after executing action a3 to obtain the SDO (modulation format configuration) of the working route. Under the constraint condition that the minimum bandwidth usage threshold of the service to be configured is met, R4 can be a function of the current spectrum efficiency ξ of the service to be configured, and the two are in a monotonically increasing relationship.

[0104] (5) The timely reward R5 in the termination state is related to the IV analysis result. When the IV analysis result is qualified, R5 is a positive number, and when the IV analysis result is unqualified, R5 is a negative number. In other words, R5 can be the timely reward obtained by the service to be configured after executing action a4 and completing the IV analysis. If the IV analysis result is qualified, it means that the working route meets the service transmission performance requirements, and R5 will be assigned a positive reward, the reward value of which exceeds the sum of the previous four timely rewards; if the IV analysis result is unqualified, it means that the working route does not meet the service transmission performance requirements, and R5 will be assigned a negative reward as a penalty, the absolute value of which exceeds the sum of the previous four timely rewards.

[0105] In some embodiments, the action strategy π(s,a) includes a random action strategy π'(S,a) and a deterministic action strategy μ(S,a). Configuring resource parameters for the service to be configured according to the action strategy includes configuring a route, wavelength, spectrum, and modulation format for the service to be configured according to the random action strategy π'(S,a). Performing IV analysis according to the action strategy includes performing IV analysis according to the deterministic action strategy μ(S,a).

[0106] Actions a0, a1, a2, and a3 correspond to the four operations of route selection, wavelength assignment, spectrum allocation, and SDO setup, respectively. Each operation has multiple alternative options. For example, an OTN network service can select one of multiple alternative routes as the working route. If selecting one of these routes as the working route is considered a specific action, then action a0 in state S0 actually corresponds to a set of actions. The specific action can be executed according to the action strategy to perform route selection. Therefore, the initial strategy for action a0 in state S0 can be the random strategy π'(s0, a0). Similarly, the initial strategies for actions a1, a2, and a3 in states S1, S2, and S3 also use the random action strategies π'(s1, a1), π'(s2, a2), and π'(s3, a3). Action a4 in state S4 corresponds to the IV analysis operation, and its action strategy can be the deterministic action strategy μ(s4, a4).

[0107] IV analysis is used to evaluate the impact of factors such as back-to-back OSNR, fiber nonlinearity, fiber CD, fiber PMD, optical filtering, PDL introduced by optical devices, cumulative OSNR fluctuation across multiple service spans, and OSNR flatness on system performance. Based on this, the feasibility of the performance of optical link resources through which OTN network services are transmitted is evaluated and analyzed according to customer requirements for OSNR margin configuration and equipment vendor strategies.

[0108] In some embodiments, as Figure 3 As shown, the damage confirmation IV analysis comprises the following steps:

[0109] Step 21: Calculate the budget value of the optical signal-to-noise ratio (OSNR) of the service to be configured. 预算值 .

[0110] In some embodiments, the budget value of the optical signal-to-noise ratio (OSNR) of the service to be configured is 预算值 It can be calculated using Formula 58, which will not be repeated here.

[0111] Step 22: Determine the budget value of the optical signal-to-noise ratio (OSNR). 预算值 Whether the preset conditions are met, if so, execute step 23, otherwise, execute step 24.

[0112] In some embodiments, the budget value of the optical signal-to-noise ratio OSNR 预算值 Satisfy the pre-conditions, including:

[0113] OSNR 预算值 -OSNR 平坦度 ≥OSNR 传输门限

[0114] OSNR 传输门限 =OSNRB2B +OSNR 非线性 +OSNR CD +OSNR PMD +OSNR 滤波 +OSNR PDL +OSNR 波动 +OSNR 净余量

[0115] Among them, OSNR 平坦度 is the flatness of the optical signal-to-noise ratio, which is the empirical value of OTN network statistics; OSNR B2B Back-to-back optical signal-to-noise ratio (OSNR) can be obtained by querying the optical module manual. 非线性 is the nonlinear cost of optical signal-to-noise ratio, OSNR CD is the optical signal-to-noise ratio (OSNR) chromatic dispersion (CD) penalty. PMD The cost of polarization mode dispersion (PMD) for the optical signal-to-noise ratio, OSNR 滤波 The filter cost of optical signal-to-noise ratio, OSNR PDL Polarization Dependent Loss (PDL) penalty for optical signal-to-noise ratio, OSNR 波动 is the fluctuation of optical signal-to-noise ratio, which is the empirical value of OTN network statistics; OSNR 净余量 The net margin required for optical signal-to-noise ratio (OSNR) is determined based on actual needs. 平坦度 、OSNR B2B 、OSNR 非线性 、OSNR CD 、OSNR PMD 、OSNR 滤波 、OSNR PDL 、OSNR 波动 、OSNR 净余量 is the default value.

[0116] Step 23: Determine whether the IV analysis result is qualified.

[0117] Step 24: Determine that the IV analysis result is unqualified.

[0118] In some embodiments, in a multi-relay span scenario, the route of the service to be configured includes multiple relay segments. Accordingly, configuring resource parameters for the service to be configured according to the action policy includes the following steps: configuring resource parameters for the service to be configured according to the action policy in each relay segment respectively. Performing IV analysis according to the action policy includes the following steps: performing IV analysis according to the action policy in each relay segment respectively. Performing damage confirmation IV analysis includes the following steps: calculating the budget value of the optical signal-to-noise ratio of each relay segment in the route of the service to be configured; in response to the budget value of the optical signal-to-noise ratio of each relay segment meeting the preset conditions, determining that the IV analysis result is qualified; in response to the budget value of the optical signal-to-noise ratio of at least one relay segment not meeting the preset conditions, determining that the IV analysis result is unqualified.

[0119] In some embodiments, as Figure 4 As shown, the calculation and updating of the optimization target strategy parameters in each state according to the timely rewards in each state include the following steps:

[0120] Step 31, calculate the expected reward in the current state based on the timely rewards in each state after the next state.

[0121] In some embodiments, the expected return under the current state can be calculated according to the following formula:

[0122]

[0123] Among them, G t State S t Next, perform action a t The expected return, γ is the discount coefficient, 0<γ<1; R is the immediate reward, t is the state S t The number of services created under t=(0, ..., n-1), where n-1 is the total number of resource parameters.

[0124] It should be noted that the expected return in the last state is the immediate reward in that state.

[0125] Step 32: Calculate and update the optimization target strategy parameters in the current state based on the expected return in the current state.

[0126] Through steps 31-32, the reward and punishment mechanism of the enhanced algorithm can be used to optimize the optimization target strategy parameters.

[0127] In some embodiments, the optimization target strategy parameter may be the state behavior value Q π (s,a), Indicates that the agent is in state S t The expectation of the cumulative reward after executing action a according to the action strategy π(s,a).

[0128] In some embodiments, the optimization target strategy parameter may also be the state value V π (s), Represents all state behavior values ​​Q under state S π (s,a). Among them, π(a|s) is the probability of executing action a according to the action strategy π(s,a) in state S, and A is the set of actions executed in each state. It should be noted that if the action strategy π(s,a) is a deterministic action strategy, then V π (s,a)=Q π (s,a).

[0129] In some embodiments, when the optimization target strategy parameter is the state behavior value Q π (s, a), the Monte Carlo Process (MCP) algorithm, the temporal difference of different strategies (TD-Error of different strategies) algorithm or the temporal difference of the same strategy (TD-Error of the same strategy) algorithm can be used to calculate and update the optimization target strategy parameters in each state. In some embodiments, the Q-Learning algorithm in the TD-Error algorithm of different strategies can be selected, or the SASA (State-Action-Reward-Action) algorithm in the TD-Error algorithm of the same strategy can be selected. Accordingly, the updating of the action strategy according to the optimal optimization target strategy parameters in each state (i.e., step 15) includes: according to the state behavior value Q π (s,a) Update the action strategy.

[0130] For example, if the Q-Learning algorithm or SASA algorithm is used, determining the optimal optimization target strategy parameters in each state (i.e., step 14) may include: π (s,a)), the maximum value of the optimal optimization target strategy parameter under each state is determined respectively.

[0131] In some embodiments, when the optimization target strategy parameter is the state value V π (s), a dynamic programming algorithm can be used to calculate and update the optimization target strategy parameters. Accordingly, the updating of the action strategy according to the optimal optimization target strategy parameters in each state (i.e., step 15) includes: according to the state value V π (s) Update the action policy μ(s,a).

[0132] The following describes the process of implementing single-service resource configuration in an OTN network using the Monte Carlo algorithm, Q-Learning algorithm, SASA algorithm, and dynamic programming algorithm.

[0133] (1) The process of implementing single-service resource configuration in an OTN network using the exploratory initialization Monte Carlo algorithm is as follows:

[0134] Initialize the entire network topology environment, for all s∈S,a∈A(s),

[0135] Q(s,a)←0; the initial value of the action strategy is μ(s,a);

[0136] returns(s,a)←emptylist;

[0137] Repeat the following process:

[0138] {

[0139] Select s0∈S, a0∈A(s) according to μ(s,a) and generate a new Episode;

[0140] For each pair (s, a) in this Episode:

[0141] The reward after the first appearance of G←(s,a);

[0142] Add G to returns(s,a);

[0143] Let the state behavior value Q(s,a)←average(returns(s,a)) take the average of the returns;

[0144] For each s in this Episode:

[0145] π(s)←argmax a Q(s,a);

[0146] }

[0147] (2) The process of implementing single-service resource configuration in an OTN network using the Q-Learning (i.e., TD-Error with different strategies) algorithm is as follows:

[0148] Initialize the entire network topology environment, for all s∈S,a∈A(s),

[0149] Q(s,a)←0; the action strategy is μ(s,a);

[0150] Repeat the following process for each Episode loop:

[0151] Initialize the state space S;

[0152] Repeat (repeat the following process for each step in this Episode):

[0153] According to the strategy μ(s,a), at s t State selection action a t ;

[0154] Execute action a t , and get timely reward R t+1 and next state s t+1 ;

[0155] Let Q(s t ,a t )←Q(s t ,a t )+α[R t+1 +γmax a Q(s t+1 ,a)-Q(s t ,a t )];

[0156] Among them, α is the learning rate;

[0157] s t ←s t+1 ;

[0158] Until s t is the terminated state;

[0159] Until all Q(s,a) converge;

[0160] Output final strategy: π(s)←argmax a Q(s,a);

[0161] (3) The process of implementing single-service resource allocation in an OTN network using the SARSA (TD-Error with the same strategy) algorithm is as follows:

[0162] Initialize the entire network topology environment, for all s∈S,a∈A(s), Q(s,a)←0;

[0163] Repeat the following process for each Episode loop:

[0164] Initialize the state space S;

[0165] Given the starting state s0, and according to the greedy strategy ε (taking the action that maximizes the immediate reward), select action a0;

[0166] Repeat (repeat the following process for each step in this Episode):

[0167] According to the greedy strategy ε, in s t State selection action a t , get immediate reward R t+1 and the next state s t+1 ;

[0168] According to the greedy strategy ε, action a is obtained t+1 ;

[0169] Let Q(s t ,a t )←Q(s t ,a t )+α[R t+1 +γQ(s t+1 ,a t+1 )-Q(s t ,a t )];

[0170] Among them, α is the learning rate;

[0171] s t ←s t+1 ;a t ←a t+1 ;

[0172] Until s t is the terminated state;

[0173] Until all Q(s,a) converge;

[0174] Output final strategy: π(s)←argmax a Q(s,a);

[0175] (4) The process of implementing single-service resource allocation in an OTN network using a dynamic programming algorithm based on policy iteration is as follows:

[0176] Step 1: Initialize the entire network topology environment.

[0177] For all s t ∈S,a∈A(s),V(s t )=0, let all The action strategy is initialized to μ(s);

[0178] Step 2: Strategy Evaluation

[0179] Here p(s t+1 ,R t+1 |s t ,μ(s)) and p(s t+1 ,R t+1 |s t ,a) represents the strategy μ(s) in state st The probability of executing the corresponding action a under

[0180] The Repeat loop repeats the following process:

[0181] Δ←0;

[0182] For each s t ∈S:

[0183] v←V(s t );

[0184]

[0185] Δ←max(Δ,|vV(s t )|);

[0186] Until Δ<θ (θ is a specified constant) converges;

[0187] Step 3: Strategy Improvement

[0188] For each s t ∈S:

[0189] a←μ(s);

[0190]

[0191] If a≠μ(s), then the strategy does not converge, otherwise the strategy converges;

[0192] If the strategy converges, the algorithm ends and returns V(s) and μ(s), otherwise it continues to return to step 2;

[0193] The disclosed embodiments can be applied to intelligent optical network management, control, and operation and maintenance. Reinforcement learning technology is used to comprehensively optimize the various resources and performance indicators of a single OTN service, thereby providing users with an OTN service path with optimized comprehensive indicators. Reinforcement learning algorithms enable comprehensive path optimization, and it is expected that ideal path optimization results can be intelligently achieved through iterative improvements to action strategies.

[0194] Based on the same technical concept, the embodiment of the present disclosure also provides a single service resource configuration device, such as Figure 5As shown, the single-service resource configuration device includes: a first processing module 101, a second processing module 102 and an update module 103. The first processing module 101 is used to configure resource parameters for the service to be configured according to the action strategy, and calculate the timely reward under the current state. After all resource parameters are configured, damage confirmation IV analysis is performed according to the action strategy, and one round ends. After one action is completed, the next state is entered. The action includes configuring a resource parameter action or performing an IV analysis action; calculating and updating the optimization target strategy parameters under each state according to the timely reward under each state; iterating a preset number of rounds to calculate and update the optimization target strategy parameters under each state.

[0195] The second processing module 102 is configured to determine the optimal optimization target strategy parameters in each state according to the optimization target strategy parameters in each state in the preset number of rounds.

[0196] The updating module 103 is used to update the action strategy according to the optimal optimization target strategy parameters in each state.

[0197] In some embodiments, the resource parameters include routing, wavelength, spectrum, and modulation format. In one round, the resource parameters are configured for the service to be configured in the following order: routing, wavelength, spectrum, and modulation format.

[0198] In some embodiments, the states include: a pending routing state, a pending wavelength state, a pending spectrum state, a pending modulation format state, a pending IV analysis state, and a terminated state. The timely reward R0 in the pending routing state is 0, and the timely rewards in other states satisfy one or any combination of the following:

[0199] The timely reward R1 in the state of the wavelength to be configured is a function of the working routing cost, and R1 is in a monotonically decreasing relationship with the function of the working routing cost;

[0200] The timely reward R2 in the state of the spectrum to be configured is a function of the wavelength resource utilization rate, and R2 and the function of the wavelength resource utilization rate are in a monotonically increasing relationship;

[0201] The timely reward R3 in the modulation format to be configured state is a function of the spectrum width occupied by the service, and R3 and the function of the spectrum width occupied by the service are in a monotonically decreasing relationship;

[0202] The timely reward R4 in the pending IV analysis state is a function of the service spectrum efficiency, and R4 and the function of the service spectrum efficiency are in a monotonically increasing relationship;

[0203] The timely reward R5 in the termination state is related to the IV analysis result. When the IV analysis result is qualified, R5 is a positive number, and when the IV analysis result is unqualified, R5 is a negative number.

[0204] In some embodiments, the action strategy includes a random action strategy and a deterministic action strategy. The first processing module 101 is configured to configure routing, wavelength, spectrum and modulation format for the service to be configured according to the random action strategy; and perform IV analysis according to the deterministic action strategy.

[0205] In some embodiments, when the route of the service to be configured includes multiple relay segments, the first processing module 101 is used to configure resource parameters for the service to be configured according to the action strategy in each relay segment; and perform damage confirmation IV analysis according to the action strategy in each relay segment.

[0206] The first processing module 101 is configured to respectively calculate a budget value of an optical signal-to-noise ratio for each relay segment in the route of the service to be configured; determine that an IV analysis result is qualified in response to the budget value of the optical signal-to-noise ratio of each relay segment satisfying a preset condition; and determine that the IV analysis result is unqualified in response to the budget value of the optical signal-to-noise ratio of at least one relay segment not satisfying the preset condition.

[0207] In some embodiments, the first processing module 101 is configured to calculate a budget value of an optical signal-to-noise ratio of the service to be configured; determine that an IV analysis result is qualified in response to the budget value of the optical signal-to-noise ratio satisfying a preset condition; and determine that the IV analysis result is unqualified in response to the budget value of the optical signal-to-noise ratio not satisfying the preset condition.

[0208] In some embodiments, the budget value of the optical signal-to-noise ratio satisfies a preset condition, including:

[0209] OSNR 预算值 -OSNR 平坦度 ≥OSNR 传输门限

[0210] OSNR 传输门限 =OSNR B2B +OSNR 非线性 +OSNR CD +OSNR PMD +OSNR 滤波 +OSNR PDL +OSNR 波动 +OSNR 净余量

[0211] Among them, OSNR 预算值 is the budget value of optical signal-to-noise ratio, OSNR 平坦度 is the flatness of the optical signal-to-noise ratio, OSNR B2Bis the back-to-back optical signal-to-noise ratio, OSNR 非线性 is the nonlinear cost of optical signal-to-noise ratio, OSNR CD is the dispersion penalty of optical signal-to-noise ratio, OSNR PMD is the optical signal-to-noise ratio (OSNR) of the polarization film dispersion penalty. 滤波 The filter cost of optical signal-to-noise ratio, OSNR PDL Polarization-dependent loss penalty for optical signal-to-noise ratio, OSNR 波动 is the fluctuation of optical signal-to-noise ratio, OSNR 净余量 The net margin required for the optical signal-to-noise ratio; OSNR 平坦度 、OSNR B2B 、OSNR 非线性 、OSNR CD 、OSNR PMD 、OSNR 滤波 、OSNR PDL 、OSNR 波动 、OSNR 净余量 is the default value.

[0212] In some embodiments, the first processing module 101 is used to calculate the expected return in the current state based on the timely rewards in each state after the next state; and calculate and update the optimization target strategy parameters in the current state based on the expected return in the current state.

[0213] In some embodiments, the first processing module 101 is configured to calculate the expected return under the current state according to the following formula:

[0214]

[0215] Among them, G t State S t Next, perform action a t The expected return, γ is the discount coefficient, 0<γ<1; R is the immediate reward, t is the state S t The number of configured resource parameters, t = (0, ..., n-1), n-1 is the total number of resource parameters.

[0216] In some embodiments, the optimization target strategy parameters include the state behavior value Q π (s,a), or,

[0217] The optimization target strategy parameters include the state value V π (s), Among them, π(a|s) is the probability of taking action a according to the action strategy π(s,a) in state S, and A is the set of actions performed in each state.

[0218] In some embodiments, the first processing module 101 is used to, when the optimization target strategy parameter is the state behavior value Q π When (s, a), the Monte Carlo algorithm, the temporal difference algorithm with different strategies or the temporal difference algorithm with the same strategy are used to calculate and update the optimization target strategy parameters in each state.

[0219] The updating module 103 is used to update the state behavior value Q according to the state behavior value Q π (s,a) Update the action strategy.

[0220] In some embodiments, the first processing module 101 is used to, when the optimization target strategy parameter is the state value V π (s), a dynamic programming algorithm is used to calculate the optimization target strategy parameters.

[0221] The updating module 103 is used to update the state value V π (s) Update the action strategy.

[0222] An embodiment of the present disclosure also provides a computer device, which includes: one or more processors and a storage device; wherein one or more programs are stored on the storage device, and when the one or more programs are executed by the one or more processors, the one or more processors implement the single-service resource configuration method provided in the aforementioned embodiments.

[0223] The embodiments of the present disclosure further provide a computer-readable medium having a computer program stored thereon, wherein when the computer program is executed, the single-service resource configuration method provided in the aforementioned embodiments is implemented.

[0224] It will be appreciated by those skilled in the art that all or some of the steps in the method disclosed above, and the functional modules / units in the device can be implemented as software, firmware, hardware, and appropriate combinations thereof. In a hardware implementation, the division between the functional modules / units mentioned in the above description does not necessarily correspond to the division of physical components; for example, a physical component may have multiple functions, or a function or step may be performed by several physical components in cooperation. Some or all physical components may be implemented as software executed by a processor, such as a central processing unit, a digital signal processor, or a microprocessor, or implemented as hardware, or implemented as an integrated circuit, such as an application-specific integrated circuit. Such software may be distributed on a computer-readable medium, which may include a computer storage medium (or non-transitory medium) and a communication medium (or temporary medium). As is well known to those skilled in the art, the term computer storage medium includes volatile and non-volatile, removable and non-removable media implemented in any method or technology for storing information (such as computer-readable instructions, data structures, program modules, or other data). Computer storage media include, but are not limited to, RAM, ROM, EEPROM, flash memory or other memory technology, CD-ROM, digital versatile disks (DVD) or other optical disk storage, magnetic cassettes, magnetic tape, magnetic disk storage or other magnetic storage devices, or any other medium that can be used to store the desired information and can be accessed by a computer. In addition, it is well known to those skilled in the art that communication media generally embodies computer-readable instructions, data structures, program modules, or other data in a modulated data signal such as a carrier wave or other transport mechanism, and may include any information delivery media.

[0225] Example embodiments have been disclosed herein, and although specific terms are employed, they are used and should be interpreted only in a general illustrative sense and not for purposes of limitation. In some instances, it will be apparent to those skilled in the art that, unless otherwise expressly indicated, features, characteristics, and / or elements described in conjunction with a particular embodiment may be used alone or in combination with features, characteristics, and / or elements described in conjunction with other embodiments. Therefore, it will be understood by those skilled in the art that various changes in form and detail may be made without departing from the scope of the invention as set forth in the appended claims.

Claims

1. A single service resource configuration method for an OTN optical transport network, characterized in that: include: Configure resource parameters for the business to be configured according to the action strategy and calculate the timely reward in the current state. After all resource parameters are configured, perform damage confirmation IV analysis according to the action strategy. One round ends. After an action is completed, enter the next state. The action includes configuring a resource parameter action or performing IV analysis action. Calculate and update the optimization target strategy parameters in each state based on the timely rewards in each state; Iterate a preset number of rounds to calculate and update the optimized target strategy parameters under each state; Determine the optimal optimization target strategy parameters in each state according to the optimization target strategy parameters in each state in the preset number of rounds; The action strategy is updated according to the optimal optimization target strategy parameters in each state.

2. The method according to claim 1, wherein The resource parameters include routing, wavelength, spectrum and modulation format. In one round, the resource parameters are configured for the service to be configured in the following order: configuring routing, configuring wavelength, configuring spectrum and configuring modulation format.

3. The method according to claim 2, wherein The states include: a state where routing is to be configured, a state where wavelength is to be configured, a state where spectrum is to be configured, a state where modulation format is to be configured, a state where IV analysis is to be performed, and a terminated state. The timely reward R0 in the state where routing is to be configured is 0. The timely rewards in other states meet one or any combination of the following requirements: The timely reward R1 in the state of the wavelength to be configured is a function of the working routing cost, and R1 is in a monotonically decreasing relationship with the function of the working routing cost; The timely reward R2 in the state of the spectrum to be configured is a function of the wavelength resource utilization rate, and R2 and the function of the wavelength resource utilization rate are in a monotonically increasing relationship; The timely reward R3 in the modulation format to be configured state is a function of the spectrum width occupied by the service, and R3 and the function of the spectrum width occupied by the service are in a monotonically decreasing relationship; The timely reward R4 in the pending IV analysis state is a function of the service spectrum efficiency, and R4 and the function of the service spectrum efficiency are in a monotonically increasing relationship; The timely reward R5 in the termination state is related to the IV analysis result. When the IV analysis result is qualified, R5 is a positive number, and when the IV analysis result is unqualified, R5 is a negative number.

4. The method according to claim 2, wherein The action strategy includes a random action strategy and a deterministic action strategy, and configuring resource parameters for the service to be configured according to the action strategy includes: configuring routing, wavelength, spectrum and modulation format for the service to be configured according to the random action strategy; The performing damage confirmation IV analysis according to the action strategy includes: performing IV analysis according to the deterministic action strategy.

5. The method according to claim 2, wherein When the route of the service to be configured includes multiple relay segments, configuring resource parameters for the service to be configured according to the action policy includes: configuring resource parameters for the service to be configured according to the action policy in each relay segment; performing damage confirmation IV analysis according to the action policy includes: performing damage confirmation IV analysis in each relay segment according to the action policy; The damage confirmation IV analysis includes: Calculating the budget value of the optical signal-to-noise ratio of each relay segment in the route of the service to be configured respectively; In response to the budgeted values ​​of the optical signal-to-noise ratios of the respective relay segments satisfying a preset condition, determining that the IV analysis result is qualified; In response to a budget value of an optical signal-to-noise ratio of at least one relay segment not satisfying a preset condition, the IV analysis result is determined to be unqualified.

6. The method according to claim 1, wherein The damage confirmation IV analysis includes: Calculating a budget value of an optical signal-to-noise ratio of the service to be configured; In response to the budget value of the optical signal-to-noise ratio meeting a preset condition, determining that the IV analysis result is qualified; In response to the budget value of the optical signal-to-noise ratio not meeting a preset condition, the IV analysis result is determined to be unqualified.

7. The method according to claim 5 or 6, wherein: The budget value of the optical signal-to-noise ratio meets preset conditions, including: OSNR 预算值 -OSNR 平坦度 ≥OSNR 传输门限 OSNR 传输门限 =OSNR B2B +OSNR 非线性 +OSNR CD +OSNR PMD +OSNR 滤波 +OSNR PDL +OSNR 波动 +OSNR 净余量 Among them, OSNR 预算值 is the budget value of optical signal-to-noise ratio, OSNR 平坦度 is the flatness of the optical signal-to-noise ratio, OSNR B2B is the back-to-back optical signal-to-noise ratio, OSNR 非线性 is the nonlinear cost of optical signal-to-noise ratio, OSNR CD is the dispersion penalty of optical signal-to-noise ratio, OSNR PMD is the optical signal-to-noise ratio (OSNR) of the polarization film dispersion penalty. 滤波 The filter cost of optical signal-to-noise ratio, OSNR PDL Polarization-dependent loss penalty for optical signal-to-noise ratio, OSNR 波动 is the fluctuation of optical signal-to-noise ratio, OSNR 净余量 The net margin required for the optical signal-to-noise ratio; OSNR 平坦度 、OSNR B2B 、OSNR 非线性 、OSNR CD 、OSNR PMD 、OSNR 滤波 、OSNR PDL 、OSNR 波动 、OSNR 净余量 is the default value.

8. The method according to any one of claims 1 to 6, wherein: The calculation and updating of the optimization target strategy parameters in each state based on the timely rewards in each state include: Calculate the expected return in the current state based on the immediate rewards in each state after the next state; Calculate and update the optimization target strategy parameters in the current state based on the expected return in the current state.

9. The method according to claim 8, wherein The expected return under the current state is calculated according to the following formula: Among them, G t State S t Next, perform action a t The expected return, γ is the discount coefficient, 0<γ<1; R is the immediate reward, t is the state S t The number of configured resource parameters, t = (0, ..., n-1), n-1 is the total number of resource parameters.

10. The method according to claim 8, wherein The optimization target strategy parameters include the state behavior value Q π (s,a), or, The optimization target strategy parameters include the state value V π (s), Where π(a|s) is the probability of taking action a according to the action strategy π(s,a) in state S, and A is the set of actions executed in each state; From state S t Starting from the action strategy π(s,a), the expected cumulative reward after executing action a is R t+k+1 For state s t+k+1 The immediate reward obtained under k R t+k+1 The discount factor.

11. The method according to claim 10, wherein When the optimization target strategy parameter is the state behavior value Q π (s, a), the Monte Carlo algorithm, the temporal difference algorithm with different strategies, or the temporal difference algorithm with the same strategy are used to calculate and update the optimization target strategy parameters in each state; The updating of the action strategy according to the optimal optimization target strategy parameters in each state includes: π (s,a) Update the action strategy.

12. The method according to claim 10, wherein When the optimization target strategy parameter is the state value V π (s), a dynamic programming algorithm is used to calculate the optimization target strategy parameters; The updating of the action strategy according to the optimal optimization target strategy parameters in each state includes: π (s) Update the action strategy.

13. A single service resource configuration device for an OTN optical transport network, comprising: a first processing module, a second processing module and an update module, The first processing module is used to configure resource parameters for the service to be configured according to the action strategy and calculate the timely reward in the current state. After all resource parameters are configured, damage confirmation IV analysis is performed according to the action strategy, and one round ends. After an action is completed, the next state is entered. The action includes configuring a resource parameter action or performing an IV analysis action; calculating and updating the optimization target strategy parameters in each state according to the timely reward in each state; iterating a preset number of rounds to calculate and update the optimization target strategy parameters in each state; The second processing module is used to determine the optimal optimization target strategy parameters in each state according to the optimization target strategy parameters in each state in the preset number of rounds; The updating module is used to update the action strategy according to the optimal optimization target strategy parameters in each state.

14. A computer device comprising: one or more processors; a storage device having one or more programs stored thereon; When the one or more programs are executed by the one or more processors, the one or more processors implement the single-service resource configuration method for an OTN optical transport network according to any one of claims 1 to 12.

15. A computer-readable medium having a computer program stored thereon, wherein: When the program is executed, the single-service resource configuration method for an OTN optical transport network according to any one of claims 1 to 12 is implemented.

Citation Information

Patent Citations

  • System and method for optimizing communications using reinforcement learning

    US20180082210A1

  • Method and apparatus for automatically generated curriculum sequence based reinforcement learning for autonomous vehicles

    US20190278282A1