OTN network resource optimization method, device, computer equipment and medium

By optimizing the OTN network resource creation sequence through a reinforcement learning algorithm, the problem of comprehensive optimization of multi-dimensional quantitative indicators in OTN network resource allocation is solved, achieving efficient resource utilization and reducing operation and maintenance costs, thereby improving transmission performance and revenue.

CN114125593BActive Publication Date: 2025-10-03ZTE CORP
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202010899110.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2020-08-31
Publication Date
2025-10-03
Estimated Expiration
2040-08-31

AI Technical Summary

Technical Problem

It is difficult to achieve comprehensive optimization of multi-dimensional quantitative indicators in existing OTN network resource allocation, resulting in insufficient network resource utilization and high operation and maintenance costs, affecting transmission performance and revenue.

Method used

The reinforcement learning algorithm is used to optimize the action strategy by introducing the parameter vector θ. The comprehensive global optimization objective function of multiple quantitative indicators is combined to design the action strategy πθ(s,a). The creation order of OTN network resources is optimized by iteratively updating the quantitative indicator weight vector.

Benefits of technology

It achieves comprehensive global optimization of multiple indicators of OTN network resources, improves resource utilization, reduces operation and maintenance costs, and enhances transmission performance and revenue.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114125593B_ABST
    Figure CN114125593B_ABST
Patent Text Reader

Abstract

The present disclosure provides an OTN network resource optimization method. The method determines a pending service in a current service establishment state based on an action strategy, creates the pending service, calculates a timely reward for the current service establishment state, and enters the next service establishment state. This process continues until a round ends. Comprehensive optimization parameters for each service establishment state are calculated based on the timely rewards for each service establishment state, and a quantitative indicator weight vector is updated based on the comprehensive optimization parameters for each service establishment state. The action strategy is a probability function associated with the quantitative indicator weight vector, and the quantitative indicator weight vector corresponds to multiple quantitative indicators. The method iterates a preset number of rounds to obtain an optimal quantitative indicator weight vector. The method updates the action strategy based on the optimal quantitative indicator weight vector, and obtains an optimized action strategy through improvements to achieve global optimization of OTN network resources. The present disclosure also provides an OTN network resource optimization device, computer equipment, and computer-readable medium.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of automatic control technology, and in particular to an OTN network resource optimization method, apparatus, computer equipment, and computer-readable medium. Background Art

[0002] With the development of artificial intelligence (AI), reinforcement learning (RL) is gaining increasing attention across various fields and industries. Reinforcement learning (RL), also known as re-enforcement learning or evaluation learning, is an important machine learning method with numerous applications in areas such as intelligent robotic control and network analysis and prediction. Within the connectionist school of machine learning, learning algorithms are categorized into three types: unsupervised learning, supervised learning, and reinforcement learning.

[0003] Reinforcement learning involves an agent learning through trial and error. Rewards earned through interaction with the environment guide behavior, with the goal of maximizing the reward. Reinforcement learning differs from supervised learning in connectionist learning primarily in its reinforcement signals. In reinforcement learning, the reinforcement signals provided by the environment evaluate the quality of an action (typically a scalar signal), rather than telling the reinforcement learning system (RLS) how to perform the correct action. Because the external environment provides little information, RLS must rely on its own experience to learn. In this way, RLS acquires knowledge in an action-evaluation environment and refines its action plans to adapt to the environment.

[0004] In recent years, with the application and promotion of reinforcement learning technology, how to apply the advantages of this technology to the field of intelligent management, control and operation and maintenance of OTN (Optical Transport Network), especially the application of reinforcement learning in OTN network resource optimization, has received widespread attention from OTN experts.

[0005] OTN network resource optimization (Global Co-current Optimization, GCO) solutions based on SDON (Software Defined Optical Network) architecture, such as Figure 1As shown, the primary purpose of GCO is to ensure that, during the OTN network resource allocation process, when planning or batch creating OTN network service provisioning, the calculated route and resource usage for each service on the OTN network is optimized to maximize the satisfaction of the user (network service operator)'s established resource allocation optimization goals for overall network services, without hindering individual service route calculation and resource allocation. OTN network resource optimization technology can minimize user operation and maintenance costs (CAPEX) and OPEX), increase O&M revenue, and optimize transmission performance and quality. This technology is directly related to the economic benefits of user network operations, and therefore has received significant attention from users. Implementing OTN network resource optimization technology is of great significance. Summary of the Invention

[0006] In view of the above-mentioned deficiencies in the prior art, the present disclosure provides an OTN network resource optimization method, apparatus, computer equipment, and computer-readable medium.

[0007] In a first aspect, embodiments of the present disclosure provide an OTN network resource optimization method, which includes determining a pending service in a current service establishment state based on an action strategy, creating the pending service, calculating a timely reward in the current service establishment state, entering the next service establishment state, and continuing until a round ends. Comprehensive optimization parameters are calculated for each service establishment state based on the timely reward in each service establishment state, and a quantitative indicator weight vector is calculated and updated based on the comprehensive optimization parameters in each service establishment state. The action strategy is a probability function associated with the quantitative indicator weight vector, and the quantitative indicator weight vector corresponds to multiple quantitative indicators.

[0008] Iterate a preset number of rounds to obtain the optimal quantitative indicator weight vector;

[0009] The action strategy is updated according to the optimal quantitative indicator weight vector.

[0010] In another aspect, the present disclosure further provides an OTN network resource optimization device, comprising: a first processing module, a second processing module, and an update module.

[0011] The first processing module is configured to determine a pending business in a current business establishment state according to an action strategy, create the pending business, calculate a timely reward in the current business establishment state, enter a next business establishment state, and continue until a round ends; calculate a comprehensive optimization parameter in each business establishment state according to the timely reward in each business establishment state; and calculate and update a quantitative indicator weight vector according to the comprehensive optimization parameter in each business establishment state, wherein the action strategy is a probability function associated with the quantitative indicator weight vector, and the quantitative indicator weight vector corresponds to a plurality of quantitative indicators;

[0012] The second processing module is used to iterate a preset number of rounds to obtain the optimal quantitative indicator weight vector;

[0013] The updating module is used to update the action strategy according to the optimal quantitative indicator weight vector.

[0014] In another aspect, an embodiment of the present disclosure further provides a computer device, comprising:

[0015] one or more processors;

[0016] a storage device having one or more programs stored thereon;

[0017] When the one or more programs are executed by the one or more processors, the one or more processors implement the OTN network resource optimization method as described above.

[0018] In yet another aspect, an embodiment of the present disclosure further provides a computer-readable medium having a computer program stored thereon, wherein when the program is executed, the OTN network resource optimization method as described above is implemented.

[0019] The disclosed embodiments provide an OTN network resource optimization method and apparatus. The method includes: determining a pending service in a current service establishment state based on an action strategy, creating the pending service, calculating a timely reward for the current service establishment state, entering the next service establishment state, and continuing until a round ends; calculating comprehensive optimization parameters for each service establishment state based on the timely rewards for each service establishment state; and updating a quantitative indicator weight vector based on the comprehensive optimization parameters for each service establishment state. The action strategy is a probability function associated with the quantitative indicator weight vector, and the quantitative indicator weight vector corresponds to multiple quantitative indicators; iterating a preset number of rounds to obtain an optimal quantitative indicator weight vector; and updating the action strategy based on the optimal quantitative indicator weight vector. The disclosed embodiments utilize a reinforcement learning algorithm's reward and punishment mechanism to optimize the ranking of OTN network service creation. The resulting action strategy has good convergence, rigor, and reliability, reducing the OTN network resource optimization problem to the problem of ranking OTN network service creation. Furthermore, a parameter vector is introduced into the reinforcement learning action strategy design. An optimized action strategy is obtained by improving , thereby achieving global optimization of OTN network resources. BRIEF DESCRIPTION OF THE DRAWINGS

[0020] Figure 1 A schematic diagram of OTN network resource optimization under the SDON architecture provided by an embodiment of the present disclosure;

[0021] Figure 2 A schematic diagram of the OTN network resource optimization process provided by an embodiment of the present disclosure;

[0022] Figure 3 A schematic diagram of a process for determining a service to be established in a current service establishment state provided by an embodiment of the present disclosure;

[0023] Figure 4 A schematic diagram of a process for calculating comprehensive optimization parameters provided in an embodiment of the present disclosure;

[0024] Figure 5 A schematic diagram of the structure of an OTN network resource optimization device provided in an embodiment of the present disclosure. DETAILED DESCRIPTION

[0025] Example embodiments will be described more fully hereinafter with reference to the accompanying drawings, but the example embodiments may be embodied in different forms and should not be construed as limited to the embodiments set forth herein. Rather, these embodiments are provided so that this disclosure will be thorough and complete and will fully convey the scope of this disclosure to those skilled in the art.

[0026] As used herein, the term "and / or" includes any and all combinations of one or more of the associated listed items.

[0027] The terms used herein are used only to describe specific embodiments and are not intended to limit the present disclosure. As used herein, the singular forms "a," "an," and "the" are also intended to include the plural forms, unless the context clearly indicates otherwise. It will also be understood that when the terms "comprising" and / or "made of" are used in this specification, the presence of the features, wholes, steps, operations, elements, and / or components is specified, but the presence or addition of one or more other features, wholes, steps, operations, elements, components, and / or groups thereof is not excluded.

[0028] The embodiments described herein may be described with reference to plan views and / or cross-sectional views, with the aid of idealized schematic diagrams of the present disclosure. Thus, the example illustrations may be modified based on manufacturing techniques and / or tolerances. Therefore, the embodiments are not limited to the embodiments shown in the accompanying drawings, but include modifications of the configurations formed based on the manufacturing process. Therefore, the regions illustrated in the accompanying drawings are schematic in nature, and the shapes of the regions shown in the drawings illustrate specific shapes of the regions of the elements, but are not intended to be limiting.

[0029] Unless otherwise defined, all terms (including technical and scientific terms) used herein have the same meaning as commonly understood by one of ordinary skill in the art. It will also be understood that terms such as those defined in commonly used dictionaries should be interpreted as having a meaning consistent with their meaning in the context of the relevant art and the present disclosure, and will not be interpreted as having an idealized or overly formal meaning unless expressly defined as such herein.

[0030] In existing OTN network service deployments, each service is typically allocated corresponding OTN network resources (such as bandwidth, spectrum, wavelength, modulation format, and routing) based on operational needs. This requires optimizing resources across the entire service under a specified optimization strategy, which includes minimizing latency and routing costs. Furthermore, to maximize operational revenue, optimize service performance, and minimize CAPEX / OPEX, OTN service operations must adhere to established optimization strategies to achieve overall optimization of OTN network resource allocation and utilization. These strategies include minimizing latency, routing costs, and maximizing bandwidth utilization. This requires that during the creation of OTN services, both the optimization of their own service resources and the global optimization of their use of OTN network resources by orchestrating the creation sequence of all services.

[0031] The OTN network service creation process typically uses a concurrent creation method, meaning that multiple services are created in batches at a specific time. The service creation process essentially determines the order in which all services are created. The order in which OTN network services are created determines the usage of OTN network resources and the optimal state of OTN network resource allocation. The order in which OTN network services are created is called the service creation orchestration strategy (or action strategy). A good service creation orchestration strategy can meet the requirements for optimizing network resource usage for OTN network services.

[0032] However, in actual OTN network resource allocation and utilization, network resource optimization is often considered across multiple dimensions. Optimizing only one dimension of a network resource's quantitative indicators will inevitably impact the utilization and tuning of other quantitative indicators. Therefore, users need to comprehensively optimize multiple quantitative indicators of network resources to arrive at the optimal combination. This process requires both ensuring the best possible global optimization of individual quantitative indicators and achieving comprehensive global optimization of all quantitative indicators of OTN network resources. This ensures maximum utilization of OTN network resources, maximized revenue, and optimized transmission performance.

[0033] To address these issues, this paper introduces a parameter vector θ in the design of action strategies using reinforcement learning. By continuously improving θ, the optimal action strategy is obtained, thereby achieving the goal of comprehensive, global optimization of multiple metrics for OTN network resources. Common quantitative metrics in OTN networks include cost, latency, BER (Bit Error Rate), Q margin, spectral efficiency, hop count, spectrum width, and transmission rate. These can all be considered as quantitative indicators for comprehensive, global optimization of multiple metrics for OTN network resources, depending on user needs.

[0034] During the initialization phase, n OTN services are created based on the environmental conditions of the OTN network graph topology (including mesh, star, and other topological structures). The network environment state, action space, action optimization target strategy, and action strategy are initialized. The relevant parameters of the reinforcement learning algorithm are defined as follows.

[0035] 1. Define the objective function for optimizing OTN network comprehensive indicators

[0036] The objective function for optimizing the comprehensive index of OTN network can be a reward for the comprehensive quantitative index of OTN network resource occupation w i maximum,

[0037] 2. Define the characteristic vector of the service establishment state S

[0038] The service establishment state is described by the feature vector φ(s). The feature vector φ(s) is used to indicate which services have been created and which services have not yet been created. When a service to be established is created, the next service establishment state is entered.

[0039] The characteristic vector φ(s) of the service establishment state S is described as follows:

[0040] {StateID; SvcNum; ... SvcID i ;SvcCost i ;SvcDelay i ;SvcQR i ;SvcFB i ;...SvcIndexh i ;

[0041] SvcSeqID i ;SvcRtID i ;SrcNdID i ;DstNdID i ;...};

[0042] in,

[0043] StateID is the business establishment state ID;

[0044] SvcNum is the total number of all services in the OTN network, which is the sum of the number of established services and the number of services to be established;

[0045] The following characteristic vector elements are used to represent a set of service establishment status attribute sequences for the i-th service in the network. The leading and trailing ellipsis represent the service establishment status attribute sequences for the first i-1 and next ni services with the same definition. The ellipsis in the middle represents the optimized quantitative index of the i-th service that is omitted.

[0046] SvcID i is the business ID of the i-th business;

[0047] SvcCost i is the routing cost of the i-th service. If the service has not been created, the routing cost is 0.

[0048] SvcDelay i is the delay of the i-th service. If the service has not been created, the delay is 0.

[0049] SvcQ i is the Q value margin of the i-th service. If the service has not been created, the Q value margin is 0;

[0050] SvcFBi The spectrum width occupied by the i-th service. If the service has not been created, the spectrum width is 0.

[0051] SvcIndexh i is the hth optimized quantitative index of the i-th service. If the service has not been created, it is 0;

[0052] SvcSeqID i The serial ID of the i-th service in the OTN network service. If the service has not been created, the serial ID of the service is 0.

[0053] SvcRtID i The routing ID occupied by the i-th service. If the service has not been created, the routing ID of the service is 0.

[0054] SrcNdID i is the source node ID of the i-th service;

[0055] DstNdID i The node ID of the i-th business purpose.

[0056] 3. Define Episode

[0057] An episode is defined as the process of establishing OTN network services in sequence using a certain action strategy.

[0058] 4. Define action a t and action strategies

[0059] An action involves selecting a pending service as the next service to be created within the current network topology and selecting a resource route from multiple candidate routes (routes with allocated network resources) for the pending service to complete the creation of the service. Multiple candidate resource routes for the pending service can be calculated using KSP (K-optimal path algorithm) + RWA (routing wavelength allocation algorithm) + RSA (asymmetric encryption algorithm), allocated with corresponding network resources, and a single candidate route must meet the threshold requirements for each quantitative indicator.

[0060] Action strategy π θ (s,a) represents the order in which services to be built (including their routes) are created. It is a probability function related to the quantitative indicator weight vector θ and is used to reflect the degree of comprehensive global optimization of multiple indicators of OTN network resources. The evaluation of comprehensive indicators of OTN network services is expressed in the form of comprehensive quantitative indicator scores. The higher the comprehensive quantitative indicator score, the higher the degree of comprehensive global optimization of multiple indicators of OTN network resources.

[0061] 5. Define quantitative indicators

[0062] Quantitative indicators include the first type of quantitative indicators, the second type of quantitative indicators and the third type of quantitative indicators. The value of the first type of quantitative indicators is inversely proportional to the score of the first type of quantitative indicators. The value index of the quantitative indicators is expressed in the form of reciprocal sum. ijk and quantitative index score w ijh1 The value of the second type of quantitative indicator is inversely proportional to the score of the second type of quantitative indicator, and the score of the third type of quantitative indicator is obtained after the last business of a round is established.

[0063] Comprehensive quantitative index score w ij is the sum of the first type of quantitative indicator scores w ijh1 , the sum of the second type of quantitative index scores w ijh2 And the sum of the third category quantitative index scores w ijh3 The sum of w ij =w ijh1 +w ijh2 +w ijh3 Among them, h1 is the number of the first type of quantitative indicators, h2 is the number of the second type of quantitative indicators, and h3 is the number of the third type of quantitative indicators.

[0064] 6. Define the quantitative indicator evaluation system

[0065] The disclosed embodiment defines different action strategies π for different quantitative indicator evaluation systems. θ (s,a), which are explained below.

[0066] (1) All businesses share the same set of indicator evaluation system

[0067] Assume that the number of services to be built in the OTN network is m, and the entire indicator evaluation system is represented by the indicator weight vector θ, θ=(θ1,θ2,...,θ h ), h is the total number of quantitative indicators, h=h1+h2+h3.

[0068] Definition of comprehensive quantitative index threshold of OTN network: index threshold =(index 1threshold ,index 2threshold ,...index hthreshold ).

[0069] The quantitative indicator scores of each alternative service route can be divided into three categories based on the classification of the quantitative indicators:

[0070] a. For the quantitative index value index ijk and quantitative index score w ijh1 In the case of inverse proportion, the relationship between the two is expressed in the form of reciprocal sum:

[0071] Among them, index ijk is the kth quantitative index of the jth alternative resource route of the i-th business, index kthreshold is the threshold of the kth quantitative index, the first type of quantitative index index ijk The smaller the value, the higher the first-class quantitative index score w corresponding to the alternative route ijh1 The higher.

[0072] b. For the quantitative index value index ijk and quantitative index score w ijh2 In the case of direct proportionality, the relationship between the two can be expressed as:

[0073]

[0074] c. The evaluation score w can only be obtained when the i-th business is the last business to be built in the round ijh3 Quantitative index ijk , the relationship between the two can be expressed as:

[0075]

[0076] Among them, index ijk is the kth quantitative index of the jth alternative resource route of the i-th business, index kthreshold is the threshold of the kth quantitative index,! (mi) is the logical NOT operation of (mi), f(index ijk ,index kthreshold )·θ k The indicator scoring function for the kth quantitative indicator of the jth alternative resource route of the i-th business is required to be proportional to the optimization degree of the quantitative indicator.

[0077] The greed coefficient of the jth alternative resource route for the i-th business is: Among them, i n is the number of alternative routes for the i-th service.

[0078] Let t be the business establishment state S t The number of services that have been created under this condition, the selection probability of the service to be built is m is the total number of services to be built in the OTN network, and the probability of each service to be built being selected is The action strategy π for each alternative route θ (s,a) is

[0079] (2) Each business has its own indicator evaluation system

[0080] Assume that the number of services to be built in the OTN network is m, and the entire indicator evaluation system is represented by the indicator weight vector θ, θ=(θ1,θ2,...θ m ), the index parameter vector of the i-th business can be defined as θ i =(θ i1 ,θ i2 ,...,θ ih ), h is the total number of quantitative indicators, h=h1+h2+h3.

[0081] Definition of comprehensive quantitative index threshold of OTN network: index threshold =(index 1threshold ,index 2threshold ,...index mthreshold ), index threshold Each element in represents the threshold vector of each service, so the index threshold vector of the i-th service can be defined as: index threshold =(index 1threshold ,index 2threshold ,...index mthreshold ).

[0082] The indicator evaluation of each service alternative resource route can be divided into three cases:

[0083] a. For the quantitative index value index ijk and quantitative index score w ijh1 In the case of inverse proportion, the relationship between the two is expressed in the form of reciprocal sum: Among them, index ijk is the kth quantitative index of the jth alternative resource route of the i-th business, index ikthreshold is the threshold of the kth quantitative index of the i-th business, the first type of quantitative index index ijk The smaller the value, the higher the first-class quantitative index score w corresponding to the alternative route ijh1 The higher.

[0084] b. For the quantitative index value index ijk and quantitative index score w ijh2 In the case of direct proportionality, the relationship between the two can be expressed as:

[0085] c. The evaluation score w can only be obtained when the i-th business is the last business to be built in the round ijh3 Quantitative index ijk , the relationship between the two can be expressed as:

[0086]

[0087] Among them, index ijk is the kth quantitative index of the jth alternative resource route of the i-th business, index ikthreshold is the threshold of the kth quantitative index of the i-th service,! (mi) is the logical NOT operation of (mi), f(index ijk ,index ikthreshold )·θ ik The indicator scoring function for the kth quantitative indicator of the jth alternative resource route of the i-th business is required to be proportional to the optimization degree of the quantitative indicator.

[0088] The greed coefficient of the jth alternative resource route for the i-th business is: Among them, i n is the number of alternative routes for the i-th service.

[0089] Let t be the business establishment state S t The number of services that have been created under this condition, the selection probability of the service to be built is m is the total number of services to be built in the OTN network, and the probability of each service to be built being selected is The action strategy π for each alternative route θ (s,a) is

[0090] (3) All businesses share the same set of indicator evaluation system and consider the ranking weight of each business

[0091] Assume that the number of services to be built in the OTN network is m, and the entire indicator evaluation system is represented by the indicator weight vector θ, θ=(θ1,θ2,...,θ m ,θ m+1 ,...θ m+h ), where θ1,...θ m is the ranking weight of each business, θ m+1 ,...θ m+h is the indicator weight for comprehensive evaluation of OTN network, m is the total number of services to be built in the OTN network, and h is the total number of quantitative indicators, h = h1 + h2 + h3.

[0092] Definition of comprehensive quantitative index threshold of OTN network: index threshold =(index 1threshold ,index 2threshold ,...index hthreshold ).

[0093] The quantitative indicator scores of each alternative service route can be divided into three categories based on the classification of the quantitative indicators:

[0094] a. For the quantitative index value index ijk and quantitative index score w ijh1 In the case of inverse proportion, the relationship between the two is expressed in the form of reciprocal sum:

[0095]

[0096] b. For the quantitative index value index ijk and quantitative index score w ijh2 In the case of direct proportionality, the relationship between the two can be expressed as:

[0097]

[0098] c. The evaluation score w can only be obtained when the i-th business is the last business to be built in the round ijh3 Quantitative index ijk , the relationship between the two can be expressed as:

[0099]

[0100] Among them, index ijk is the kth quantitative index of the jth alternative resource route of the i-th business, index kthreshold is the threshold of the kth quantitative index,! (mi) is the logical NOT operation of (mi), f(index ijk ,index kthreshold )·θ m+k is the indicator scoring function of the kth quantitative indicator of the jth alternative resource route of the i-th business.

[0101] The greed coefficient of the jth alternative resource route for the i-th business is: Among them, i n is the number of alternative routes for the i-th service.

[0102] The action strategy π is calculated according to the following formula θ (s,a): Among them, ξ ij is the greed coefficient of the jth alternative resource routing for the i-th business, The sorting weight is θ i The selection probability of the business i to be built, state s t There are (mt) businesses to be built, t is the number of businesses that have been built, is the ranking weight set of the services to be built, N t A set of ranking weights for established services.

[0103] (4) Each business has its own indicator evaluation system and the ranking weight of each business is taken into account

[0104] Assume that the number of services to be built in the OTN network is m, and the entire indicator evaluation system is represented by the indicator weight vector θ, θ=(θ1,θ2,...θ m ),θ i =(θ i0 ,θ i1 ,θ i2 ,...,θ ih ),θ i0 is the ranking weight of the i-th business, θ i1 ,...θ ih is the indicator weight of the ith service, m is the total number of services to be built in the OTN network, and h is the total number of quantitative indicators, h = h1 + h2 + h3.

[0105] Definition of comprehensive quantitative index threshold of OTN network: index threshold =(index 1threshold ,index 2threshold ,...index mthreshold ).

[0106] The quantitative indicator scores of each alternative service route can be divided into three categories based on the classification of the quantitative indicators:

[0107] a. For the quantitative index value index ijk and quantitative index score w ijh1 In the case of inverse proportion, the relationship between the two is expressed in the form of reciprocal sum:

[0108]

[0109] b. For the quantitative index value index ijk and quantitative index score w ijh2 In the case of direct proportionality, the relationship between the two can be expressed as:

[0110]

[0111] c. The evaluation score w can only be obtained when the i-th business is the last business to be built in the round ijh3 Quantitative index ijk , the relationship between the two can be expressed as:

[0112]

[0113] Among them, index ijk is the kth quantitative index of the jth alternative resource route of the i-th business, index ikthresholdis the threshold of the kth quantitative index of the i-th service,! (mi) is the logical NOT operation of (mi), f(index ijk ,index ikthreshold )·θ ik is the indicator scoring function of the kth quantitative indicator of the jth alternative resource route of the i-th business.

[0114] The greed coefficient of the jth alternative resource route for the i-th business is: Among them, i n is the number of alternative routes for the i-th service.

[0115] The action strategy π is calculated according to the following formula θ (s,a): Among them, ξ ij is the greed coefficient of the jth alternative resource routing for the i-th business, The sorting weight is θ i0 The selection probability of the business i to be built, state s t There are (mt) businesses to be built, t is the number of businesses that have been built, is the ranking weight set of the services to be built, N t A set of ranking weights for established services.

[0116] The present disclosure provides an OTN network resource optimization method, such as Figure 2 As shown, the method includes the following steps:

[0117] Step 11: Determine the business to be established in the current business establishment state according to the action strategy, create the business to be established, calculate the timely reward in the current business establishment state, enter the next business establishment state, and continue until the end of one round. Calculate the comprehensive optimization parameters in each business establishment state according to the timely reward in each business establishment state, and calculate and update the quantitative indicator weight vector according to the comprehensive optimization parameters in each business establishment state.

[0118] As mentioned above, the action strategy is a probability function related to a quantization indicator weight vector, and the quantization indicator weight vector corresponds to a plurality of quantization indicators.

[0119] In this step, within a round, a service to be established is determined according to the action strategy (including determining the route for the service to be established). After the service to be established is created, the immediate reward for the service establishment state is calculated, the current service establishment state ends, and the next service establishment state is entered. Following the above steps, for each service establishment state in a round, a service to be established is created and the immediate reward for the corresponding service establishment state is calculated. This continues until the end of the round. The comprehensive optimization parameters for each service establishment state are calculated and updated based on the immediate reward for each service establishment state.

[0120] In this step, different algorithms may be used to calculate and update the comprehensive optimization parameters. It should be noted that different algorithms are used, and the comprehensive optimization parameters are also different. Various algorithms will be described in detail later.

[0121] Step 12: Iterate a preset number of times to obtain the optimal quantitative indicator weight vector.

[0122] In this step, step 11 is repeated for a preset number of rounds, and the comprehensive optimization parameters for each service establishment state in each round are calculated and updated. Through this step, the optimal comprehensive optimization parameters for all service establishment states corresponding to all pending services in the OTN network can be obtained.

[0123] Step 13: Update the action strategy according to the optimal comprehensive optimization parameters under the establishment status of each service.

[0124] Comprehensive optimization parameters are used to characterize the service establishment status S t and action a t Once the optimal comprehensive optimization parameters for a business establishment state are determined, the optimal action a for the business establishment state can be determined. t , the optimal action a t That is, the action of creating the optimal service to be established in the service establishment state can determine the optimal service to be established in the service establishment state (including the route of the service to be established), thereby obtaining the services to be established sorted according to the service establishment status, and the sorting of the services to be established is the optimized action strategy.

[0125] The disclosed embodiments provide an OTN network resource optimization method and apparatus. The method includes: determining a pending service in a current service establishment state based on an action strategy, creating the pending service, calculating a timely reward for the current service establishment state, entering the next service establishment state, and continuing until a round ends; calculating comprehensive optimization parameters for each service establishment state based on the timely rewards for each service establishment state; and updating a quantitative indicator weight vector based on the comprehensive optimization parameters for each service establishment state. The action strategy is a probability function associated with the quantitative indicator weight vector, and the quantitative indicator weight vector corresponds to multiple quantitative indicators; iterating a preset number of rounds to obtain an optimal quantitative indicator weight vector; and updating the action strategy based on the optimal quantitative indicator weight vector. The disclosed embodiments utilize a reinforcement learning algorithm's reward and punishment mechanism to optimize the ranking of OTN network service creation. The resulting action strategy has good convergence, rigor, and reliability, reducing the OTN network resource optimization problem to the problem of ranking OTN network service creation. Furthermore, a parameter vector is introduced into the reinforcement learning action strategy design. An optimized action strategy is obtained by improving , thereby achieving global optimization of OTN network resources.

[0126] In some embodiments, the comprehensive optimization parameter can be a state behavior value Indicates starting from state S, according to strategy π θ The expected cumulative reward after taking action a. Where γ is the discount coefficient, 0<γ<1; R is the immediate reward, and t is the business establishment status S t The number of services created under t=(0, ..., m), where m is the total number of services to be built in the OTN network.

[0127] In some embodiments, the comprehensive optimization parameter can also be a state value Represents all state behavior values ​​under state S The weighted sum of . Among them, π θ (a|s) is the action strategy π under the service establishment state S θ (a|s) is the probability of taking action a, where A is the set of actions executed under each service establishment state.

[0128] When the comprehensive optimization parameter is a state behavior value When the comprehensive optimization parameters under the establishment state of each business are calculated and updated, the quantitative index weight vector is updated, including the following steps: using the actor-critic algorithm, according to the neural network model, the gradient of the action strategy and the state behavior value Calculate and update the quantitative index weight vector.

[0129] In some embodiments, according to the feature vector function of state S and action a, the parameterized state behavior value function Q ω (s, a) and the neural network model trains the neural network layer parameter vector ω, which is the feature vector function of the state S and the behavior action a As the input of the neural network model, the parameterized state behavior value function Q ω (s,a) is used as the output of the neural network model to train the neural network layer parameter vector ω. ω (s,a) according to Obtain Q ω (s,a)=φ(s,a) T ω. Update the state behavior value according to the neural network layer parameter vector ω And according to the state behavior value And the gradient of the action policy updates the indicator weight vector θ.

[0130] When the comprehensive optimization parameter is a state value When the comprehensive optimization parameters under the establishment state of each business are calculated and updated, the quantitative index weight vector is updated, including the following steps: using the policy gradient (PG) algorithm, according to the gradient of the action strategy and the state value Calculate and update the quantitative index weight vector.

[0131] In some embodiments, as Figure 3 As shown, the process of determining the service to be established in the current service establishment state according to the action strategy includes the following steps:

[0132] Step 21: Calculate the probability of selecting each service to be established in the current service establishment state.

[0133] In this step, according to the selected quantitative index evaluation system, a corresponding algorithm is determined to calculate the probability of each service to be established being selected. The probability of selecting each service to be established under different quantitative index evaluation systems is as described above and will not be repeated here.

[0134] Step 22: determining a service to be established according to the probability of selecting each service to be established in the current service establishment state.

[0135] It should be noted that based on the exploration idea of ​​reinforcement learning, the selection of services to be established follows the randomness of the strategy.

[0136] Step 23: rank the determined candidate routes for the service to be established according to the preset OTN network comprehensive index optimization objective function.

[0137] The OTN network comprehensive index optimization objective function is the OTN network resource occupancy comprehensive quantitative index reward w i Maximum, w i That is, the reward value for being selected as the working route among multiple alternative routes for the i-th service to be established.

[0138] Step 24: Calculate the selection probability of each candidate route in the ranking.

[0139] Step 25: Determine an alternative route according to the selection probability of each alternative route in the ranking as the route for the service to be established in the current service establishment state.

[0140] In some embodiments, the OTN network resource occupancy comprehensive quantitative indicator reward w is calculated according to the following formula: i :w i =w ih1 +w ih2 +w ih3 ;

[0141] Among them, w ih1 is the sum of the reward values ​​of the first type of quantitative indicators, w ih2 is the sum of the reward values ​​of the second type of quantitative indicators, w ih3 is the sum of the reward values ​​of the third category of quantitative indicators, λ is the reward coefficient vector of the quantitative indicator, λ=(λ1,λ2,...,λ h ), h is the total number of quantitative indicators, h=h1+h2+h3.

[0142] R t+1 Indicates state S t Next, perform action a t The immediate reward obtained, R t+1 =w t+1 , which is equal to the comprehensive quantitative index reward of the t+1th business. The higher the reward value, the higher the R t+1 The larger the value, the greater the value. In the S0 state, R0=0.

[0143] In some embodiments, as Figure 4 As shown, the method of calculating and updating the comprehensive optimization parameters in each business establishment state according to the timely reward in each business establishment state includes the following steps:

[0144] Step 31 : Calculate the expected return in the current business establishment state based on the timely rewards in each business establishment state after the next business establishment state.

[0145] In some embodiments, the expected return under the current service establishment state can be calculated according to the following formula:

[0146]

[0147] Among them, G t Establish state S for the business t Next, perform action a t The expected return, γ is the discount coefficient, 0<γ<1; R t+1 Establish state S for the business t Next, perform action a t The immediate reward obtained, R t+1 =w t+1 , t is the business establishment status S t The number of services created under t=(0, ..., m), where m is the total number of services to be built in the OTN network.

[0148] It should be noted that the expected return under the last business establishment state is the immediate reward under the business establishment state.

[0149] Step 32: Calculate and update comprehensive optimization parameters in the current business establishment state according to the expected return in the current business establishment state.

[0150] Through steps 31-32, the reward and punishment mechanism of the enhanced algorithm is used to optimize the comprehensive optimization parameters.

[0151] The following describes how the Q-Based Actor-Critic algorithm and the PG algorithm optimize OTN network resources.

[0152] (1) The process of using the Q-Based Actor-Critic algorithm to optimize OTN network resources is as follows:

[0153] Initialize the entire network topology environment, including initializing s∈S and the strategy parameter vector θ;

[0154] Initialize sampling actions a~π according to the strategy θ ;

[0155] Let Q ω (s,a)=φ(s,a) T ω

[0156] For each step of the sampling action do:

[0157] Sampling timely reward Adopt next state transition

[0158] Sample the next action a'~π according to the strategy θ (s',a');

[0159] δ=r+γQω (s',a')-Q ω (s,a);

[0160] θ=θ+α▽ θ logπ θ (s,a)Q ω (s,a);

[0161] ω←ω+βδφ(s,a);

[0162] a←a',s←s';

[0163] End for;

[0164] End processing;

[0165] (2) The process of using the PG algorithm to optimize OTN network resources is as follows:

[0166] Initialize the entire network topology environment, for all s∈S,a∈A(s), Q(s,a)←0;

[0167] Initialize θ;

[0168] For each Episode {s1,a1,r2,...,s T-1 ,a T-1 ,r T}~π θ Each step of the loop repeats the following process:

[0169] For t=1 to T-1 do

[0170] θ←θ+α▽ θ logπ θ (s t ,a t )v t ;

[0171] End for

[0172] End for

[0173] Return θ and update the policy π θ (s,a);

[0174] Based on the same technical concept, the embodiment of the present disclosure also provides an OTN network resource optimization device, such as Figure 5 As shown, the OTN network resource optimization device includes: a first processing module 101, a second processing module 102 and an update module 103,

[0175] The first processing module is used to determine the business to be established in the current business establishment state according to the action strategy, create the business to be established, and calculate the timely reward in the current business establishment state, enter the next business establishment state, until the end of a round, calculate the comprehensive optimization parameters in each business establishment state according to the timely reward in each business establishment state, and calculate and update the quantitative indicator weight vector according to the comprehensive optimization parameters in each business establishment state, wherein the action strategy is a probability function related to the quantitative indicator weight vector, and the quantitative indicator weight vector corresponds to multiple quantitative indicators.

[0176] The second processing module is used to iterate a preset number of rounds to obtain the optimal quantitative indicator weight vector.

[0177] The updating module is used to update the action strategy according to the optimal quantitative indicator weight vector.

[0178] In some embodiments, the first processing module is further used to calculate a comprehensive quantitative indicator score based on multiple quantitative indicators and indicator weight vectors; calculate a greediness coefficient based on the comprehensive quantitative indicator score; and determine the action strategy based on the selection probability of the service to be established and the greediness coefficient.

[0179] In some embodiments, the quantitative indicators include first-category quantitative indicators, second-category quantitative indicators, and third-category quantitative indicators, wherein the value of the first-category quantitative indicator is inversely proportional to the score of the first-category quantitative indicator, the value of the second-category quantitative indicator is inversely proportional to the score of the second-category quantitative indicator, and the score of the third-category quantitative indicator is obtained after the last business of a round is established.

[0180] The comprehensive quantitative index score w ij is the sum of the first type of quantitative indicator scores w ijh1 , the sum of the second type of quantitative index scores w ijh2 And the sum of the third category quantitative index scores w ijh3 The sum of , where h1 is the number of the first type of quantitative indicators, h2 is the number of the second type of quantitative indicators, and h3 is the number of the third type of quantitative indicators.

[0181] In some embodiments, the first processing module 101 is further configured to calculate the greediness coefficient according to the following formula: Among them, ξ ij is the greed coefficient of the jth alternative resource routing of the i-th business, w ij Score the comprehensive quantitative indicators, i n is the number of alternative routes for the i-th service.

[0182] In some embodiments, the indicator weight vector θ=(θ1, θ2, ..., θ h), h is the total number of quantitative indicators, h = h1 + h2 + h3; the first module 101 is used to calculate the sum of the first type of quantitative indicator scores w according to the following formula ijh1 : The sum of the second category quantitative indicator scores w is calculated according to the following formula ijh2 : The sum of the third category quantitative indicator scores w is calculated according to the following formula ijh3 : Among them, index ijk is the kth quantitative index of the jth alternative resource route of the i-th business, index kthreshold is the threshold of the kth quantitative index,! (mi) is the logical NOT operation of (mi), f(index ijk ,index kthreshold )·θ k is the index scoring function of the kth quantitative index of the jth alternative resource route of the i-th business; the action strategy π is calculated according to the following formula θ (s,a): Among them, ξ ij is the greed coefficient of the jth alternative resource routing for the i-th business, is the selection probability of the service to be established, t is the service establishment status S t The number of services created under m is the total number of services to be built in the OTN network.

[0183] In some embodiments, the indicator weight vector θ=(θ1, θ2, ...θ m ),θ i =(θ i1 ,θ i2 ,...,θ ih ), h is the total number of quantitative indicators, h = h1 + h2 + h3; the sum of the scores of the first category of quantitative indicators w is calculated according to the following formula ijh1 : The sum of the second category quantitative indicator scores w is calculated according to the following formula ijh2 : The sum of the third category quantitative indicator scores w is calculated according to the following formula ijh3 : Among them, index ijk is the kth quantitative index of the jth alternative resource route of the i-th business, index ikthreshold is the threshold of the kth quantitative index of the i-th service,! (mi) is the logical NOT operation of (mi), f(index ijk ,index ikthreshold )·θ ikis the index scoring function of the kth quantitative index of the jth alternative resource route of the i-th business; the action strategy π is calculated according to the following formula θ (s,a): Among them, ξ ij is the greed coefficient of the jth alternative resource routing for the i-th business, is the selection probability of the service to be established, t is the service establishment status S t The number of services created under m is the total number of services to be built in the OTN network.

[0184] In some embodiments, the indicator weight vector θ=(θ1, θ2, ..., θ m ,θ m+1 ,...θ m+h ), m is the total number of services to be built in the OTN network, h is the total number of quantitative indicators, h = h1 + h2 + h3; the first processing module 101 is used to calculate the sum of the first type of quantitative indicator scores w according to the following formula ijh1 : The sum of the second category quantitative indicator scores w is calculated according to the following formula ijh2 : The sum of the third category quantitative indicator scores w is calculated according to the following formula ijh3 : Among them, index ijk is the kth quantitative index of the jth alternative resource route of the i-th business, index kthreshold is the threshold of the kth quantitative index,! (mi) is the logical NOT operation of (mi), f(index ijk ,index kthreshold )·θ m+k is the indicator scoring function of the kth quantitative indicator of the jth alternative resource route of the i-th service;

[0185] The action strategy π is calculated according to the following formula θ (s,a): Among them, ξ ij is the greed coefficient of the jth alternative resource routing for the i-th business, The sorting weight is θ i The selection probability of the business i to be built, state s t There are (mt) businesses to be built, t is the number of businesses that have been built, is the ranking weight set of the services to be built, N t A set of ranking weights for established services.

[0186] In some embodiments, the indicator weight vector θ=(θ1, θ2, ...θ m ),θ i=(θ i0 ,θ i1 ,θ i2 ,...,θ ih ), m is the total number of services to be built in the OTN network, h is the total number of quantitative indicators, h = h1 + h2 + h3; the first processing module 101 is used to calculate the sum of the first type of quantitative indicator scores w according to the following formula ijh1 : The sum of the second category quantitative indicator scores w is calculated according to the following formula ijh2 : The sum of the third category quantitative indicator scores w is calculated according to the following formula ijh3 : Among them, index ijk is the kth quantitative index of the jth alternative resource route of the i-th business, index ikthreshold is the threshold of the kth quantitative index of the i-th service,! (mi) is the logical NOT operation of (mi), f(index ijk ,index ikthreshold )·θ ik is the index scoring function of the kth quantitative index of the jth alternative resource route of the i-th business; the action strategy π is calculated according to the following formula θ (s,a): Among them, ξ ij is the greedy coefficient of the jth alternative resource routing for the i-th business The sorting weight is θ i0 The selection probability of the business i to be built, state s t There are (mt) businesses to be built, t is the number of businesses that have been built, is the ranking weight set of the services to be built, N t A set of ranking weights for established services.

[0187] In some embodiments, the comprehensive optimization parameters include state behavior values Where γ is the discount coefficient, 0<γ<1; R is the timely reward, and t is the business establishment status S t The number of services created under t = (0, ..., m), where m is the total number of services to be built in the OTN network; or

[0188] The comprehensive optimization parameters include state values Among them, π θ (a|s) is the action strategy π under the service establishment state S θ (a|s) is the probability of taking action a, where A is the set of actions executed under each service establishment state.

[0189] In some embodiments, when the comprehensive optimization parameter is a state behavior value When the update module 103 is used to adopt the actor-critic algorithm, according to the neural network model, the gradient of the action strategy and the state behavior value Calculate and update the quantitative index weight vector.

[0190] In some embodiments, when the comprehensive optimization parameter is a state value When the update module 103 is used to adopt the policy gradient algorithm, according to the gradient of the action strategy and the state value Calculate and update the quantitative index weight vector.

[0191] In some embodiments, the first processing module 101 is used to calculate the probability of selecting each to-be-established service in the current service establishment state; determine a to-be-established service based on the probability of selecting each to-be-established service in the current service establishment state; sort the determined alternative routes for the to-be-established service based on a preset OTN network comprehensive indicator optimization objective function; calculate the selection probability of each alternative route in the sorting; and determine an alternative route based on the selection probability of each alternative route in the sorting as the route for the to-be-established service in the current service establishment state.

[0192] In some embodiments, the OTN network comprehensive index optimization objective function is the OTN network resource occupancy comprehensive quantitative index reward w i maximum.

[0193] In some embodiments, the first processing module 101 is used to calculate the OTN network resource occupation comprehensive quantitative indicator reward w according to the following formula: i :w i =w ih1 +w ih2 +w ih3 ; Among them, w ih1 is the sum of the reward values ​​of the first type of quantitative indicators, w ih2 is the sum of the reward values ​​of the second type of quantitative indicators, w ih3 is the sum of the reward values ​​of the third category of quantitative indicators, λ is the reward coefficient vector of the quantitative indicator, λ=(λ1,λ2,...,λ h ), h is the total number of quantitative indicators, h=h1+h2+h3.

[0194] In some embodiments, the first processing module 101 is used to calculate the expected return in the current business establishment state based on the timely rewards in each business establishment state after the next business establishment state; and calculate and update the comprehensive optimization parameters in the current business establishment state based on the expected return in the current business establishment state.

[0195] In some embodiments, the first processing module 101 is configured to calculate the expected return under the current service establishment state according to the following formula:

[0196] Among them, G t Establish state S for the business t Next, perform action a t The expected return, γ is the discount coefficient, 0<γ<1; R t+1 Establish state S for the business t Next, perform action a t The immediate reward obtained, R t+1 =w t+1 , t is the business establishment status S t The number of services created under t=(0, ..., m), where m is the total number of services to be built in the OTN network.

[0197] An embodiment of the present disclosure further provides a computer device, comprising: one or more processors and a storage device; wherein the storage device stores one or more programs, and when the one or more programs are executed by the one or more processors, the one or more processors implement the OTN network resource optimization method provided in the aforementioned embodiments.

[0198] The embodiments of the present disclosure further provide a computer-readable medium having a computer program stored thereon, wherein when the computer program is executed, the OTN network resource optimization method provided in the aforementioned embodiments is implemented.

[0199] It will be appreciated by those skilled in the art that all or some of the steps in the method disclosed above, and the functional modules / units in the device can be implemented as software, firmware, hardware, and appropriate combinations thereof. In a hardware implementation, the division between the functional modules / units mentioned in the above description does not necessarily correspond to the division of physical components; for example, a physical component may have multiple functions, or a function or step may be performed by several physical components in cooperation. Some or all physical components may be implemented as software executed by a processor, such as a central processing unit, a digital signal processor, or a microprocessor, or implemented as hardware, or implemented as an integrated circuit, such as an application-specific integrated circuit. Such software may be distributed on a computer-readable medium, which may include a computer storage medium (or non-transitory medium) and a communication medium (or temporary medium). As is well known to those skilled in the art, the term computer storage medium includes volatile and non-volatile, removable and non-removable media implemented in any method or technology for storing information (such as computer-readable instructions, data structures, program modules, or other data). Computer storage media include, but are not limited to, RAM, ROM, EEPROM, flash memory or other memory technology, CD-ROM, digital versatile disks (DVD) or other optical disk storage, magnetic cassettes, magnetic tape, magnetic disk storage or other magnetic storage devices, or any other medium that can be used to store the desired information and can be accessed by a computer. In addition, it is well known to those skilled in the art that communication media generally embodies computer-readable instructions, data structures, program modules, or other data in a modulated data signal such as a carrier wave or other transport mechanism, and may include any information delivery media.

[0200] Example embodiments have been disclosed herein, and although specific terms are employed, they are used and should be interpreted only in a general illustrative sense and not for purposes of limitation. In some instances, it will be apparent to those skilled in the art that, unless otherwise expressly indicated, features, characteristics, and / or elements described in conjunction with a particular embodiment may be used alone or in combination with features, characteristics, and / or elements described in conjunction with other embodiments. Therefore, it will be understood by those skilled in the art that various changes in form and detail may be made without departing from the scope of the invention as set forth in the appended claims.

Claims

1. A method for optimizing OTN network resources, characterized in that: include: Determining a pending service in a current service establishment state according to an action strategy, creating the pending service, calculating a timely reward in the current service establishment state, entering a next service establishment state, and continuing until a round ends; calculating comprehensive optimization parameters in each service establishment state according to the timely reward in each service establishment state, and updating a quantitative indicator weight vector according to the comprehensive optimization parameters in each service establishment state, wherein the action strategy is a probability function associated with the quantitative indicator weight vector, and the quantitative indicator weight vector corresponds to the quantitative indicators of a plurality of network resources; Iterate a preset number of rounds to obtain the optimal quantitative indicator weight vector; The action strategy is updated according to the optimal quantitative indicator weight vector.

2. The method according to claim 1, wherein Before determining the service to be established in the current service establishment state according to the action strategy, the method further includes: Calculate the comprehensive quantitative index score based on multiple quantitative indicators and indicator weight vectors; Calculate the greed coefficient based on the comprehensive quantitative indicator score; The action strategy is determined according to the selection probability of the service to be established and the greed coefficient.

3. The method according to claim 2, wherein The quantitative indicators include first-category quantitative indicators, second-category quantitative indicators, and third-category quantitative indicators. The value of the first-category quantitative indicator is inversely proportional to the score of the first-category quantitative indicator, the value of the second-category quantitative indicator is inversely proportional to the score of the second-category quantitative indicator, and the score of the third-category quantitative indicator is obtained after the last business of a round is established. The comprehensive quantitative index score w ij is the sum of the first type of quantitative indicator scores w ijh1 , the sum of the second type of quantitative index scores w ijh2 And the sum of the third category quantitative index scores w ijh3 The sum of , where h1 is the number of the first type of quantitative indicators, h2 is the number of the second type of quantitative indicators, and h3 is the number of the third type of quantitative indicators.

4. The method according to claim 3, wherein The greediness coefficient is calculated according to the following formula: Among them, ξ ij is the greed coefficient of the jth alternative resource routing of the i-th business, w ij Score the comprehensive quantitative indicators, i n is the number of alternative routes for the i-th service.

5. The method according to claim 4, wherein The indicator weight vector θ=(θ1, θ2, ..., θ h ), h is the total number of quantitative indicators, h = h1 + h2 + h3; The sum of the first type of quantitative indicator scores w is calculated according to the following formula ijh1 : The sum of the second category quantitative indicator scores w is calculated according to the following formula ijh2 : The sum of the third category quantitative indicator scores w is calculated according to the following formula ijh3 : Among them, index ijk is the kth quantitative index of the jth alternative resource route of the i-th business, index kthreshold is the threshold of the kth quantitative index,! (mi) is the logical NOT operation of (mi), f(index ijk ,index kthreshold )·θ k is the indicator scoring function of the kth quantitative indicator of the jth alternative resource route of the i-th service; The action strategy π is calculated according to the following formula θ (s,a): Among them, ξ ij is the greed coefficient of the jth alternative resource routing for the i-th business, is the selection probability of the service to be established, t is the service establishment status S t The number of services created under m is the total number of services to be built in the OTN network.

6. The method according to claim 4, wherein The indicator weight vector θ=(θ1, θ2, ...θ m ),θ i =(θ i1 ,θ i2 ,...,θ ih ), h is the total number of quantitative indicators, h = h1 + h2 + h3; The sum of the first type of quantitative indicator scores w is calculated according to the following formula ijh1 : The sum of the second category quantitative indicator scores w is calculated according to the following formula ijh2 : The sum of the third category quantitative indicator scores w is calculated according to the following formula ijh3 : Among them, index ijk is the kth quantitative index of the jth alternative resource route of the i-th business, index ikthreshold is the threshold of the kth quantitative index of the i-th service,! (mi) is the logical NOT operation of (mi), f(index ijk ,index ikthreshold )·θ ik is the indicator scoring function of the kth quantitative indicator of the jth alternative resource route of the i-th service; The action strategy π is calculated according to the following formula θ (s,a): Among them, ξ ij is the greed coefficient of the jth alternative resource routing for the i-th business, is the selection probability of the service to be established, t is the service establishment status S t The number of services created under m is the total number of services to be built in the OTN network.

7. The method according to claim 4, wherein The indicator weight vector θ=(θ1, θ2, ..., θ m ,θ m+1 ,...θ m+h ), m is the total number of services to be built in the OTN network, h is the total number of quantitative indicators, h = h1 + h2 + h3; The sum of the first type of quantitative indicator scores w is calculated according to the following formula ijh1 : The sum of the second category quantitative indicator scores w is calculated according to the following formula ijh2 : The sum of the third category quantitative indicator scores w is calculated according to the following formula ijh3 : Among them, index ijk is the kth quantitative index of the jth alternative resource route of the i-th business, index kthreshold is the threshold of the kth quantitative index,! (mi) is the logical NOT operation of (mi), f(index ijk ,index kthreshold )·θ m+k is the indicator scoring function of the kth quantitative indicator of the jth alternative resource route of the i-th service; The action strategy π is calculated according to the following formula θ (s,a): Among them, ξ ij is the greed coefficient of the jth alternative resource routing for the i-th business, The sorting weight is θ i The selection probability of the business i to be built, state s t There are (mt) businesses to be built, t is the number of businesses that have been built, is the ranking weight set of the services to be built, N t A set of ranking weights for established services.

8. The method according to claim 4, wherein The indicator weight vector θ=(θ1, θ2, ...θ m ),θ i =(θ i0 ,θ i1 ,θ i2 ,...,θ ih ), m is the total number of services to be built in the OTN network, h is the total number of quantitative indicators, h = h1 + h2 + h3; The sum of the first type of quantitative indicator scores w is calculated according to the following formula ijh1 : The sum of the second category quantitative indicator scores w is calculated according to the following formula ijh2 : The sum of the third category quantitative indicator scores w is calculated according to the following formula ijh3 : Among them, index ijk is the kth quantitative index of the jth alternative resource route of the i-th business, index ikthreshold is the threshold of the kth quantitative index of the i-th service,! (mi) is the logical NOT operation of (mi), f(index ijk ,index ikthreshold )·θ ik is the indicator scoring function of the kth quantitative indicator of the jth alternative resource route of the i-th service; The action strategy π is calculated according to the following formula θ (s,a): Among them, ξ ij is the greed coefficient of the jth alternative resource routing for the i-th business, The sorting weight is θ i0 The selection probability of the business i to be built, state s t There are (mt) businesses to be built, t is the number of businesses that have been built, is the ranking weight set of the services to be built, N t A set of ranking weights for established services.

9. The method according to claim 1, wherein The comprehensive optimization parameters include state behavior values in, Indicates starting from the business establishment state S, according to the action strategy π θ The expected cumulative reward after taking action a; γ is the discount coefficient, 0<γ<1; R is the immediate reward, t is the business establishment state S t The number of services created under t = (0, ..., m), where m is the total number of services to be built in the OTN network; or The comprehensive optimization parameters include state values Among them, π θ (a|s) is the action strategy under the service establishment state S, and its value is the probability of taking action a under the service establishment state S. A is the set of actions executed under each service establishment state.

10. The method according to claim 9, wherein When the comprehensive optimization parameter is a state behavior value , the updating of the quantitative index weight vector according to the comprehensive optimization parameters of each service establishment state, including: The actor-critic algorithm is used to calculate the state behavior value based on the neural network model, the gradient of the action strategy and the state behavior value. Calculate and update the quantitative index weight vector.

11. The method according to claim 9, wherein When the comprehensive optimization parameter is a state value , the updating of the quantitative index weight vector according to the comprehensive optimization parameters of each service establishment state, including: Using the policy gradient algorithm, according to the gradient and state value of the action strategy Calculate and update the quantitative index weight vector.

12. The method according to any one of claims 3 to 11, wherein: The determining of the service to be established in the current service establishment state according to the action strategy includes: Calculate the probability of selecting each service to be established under the current service establishment status; Determine a service to be established according to the probability of selecting each service to be established under the current service establishment state; Optimize the objective function based on the preset OTN network comprehensive indicators and sort the candidate routes for the services to be established; Calculating the selection probability of each alternative route in the sorting; An alternative route is determined according to the selection probability of each alternative route in the ranking as the route of the service to be established in the current service establishment state.

13. The method according to claim 12, wherein: The OTN network comprehensive index optimization objective function is the OTN network resource occupancy comprehensive quantitative index reward w i maximum.

14. The method according to claim 13, wherein The OTN network resource utilization comprehensive quantitative indicator reward w is calculated according to the following formula i :w i =w ih1 +w ih2 +w ih3 ; Among them, w ih1 is the sum of the reward values ​​of the first type of quantitative indicators, w ih2 is the sum of the reward values ​​of the second type of quantitative indicators, w ih3 is the sum of the reward values ​​of the third category of quantitative indicators, λ is the reward coefficient vector of the quantitative indicator, λ=(λ1,λ2,...,λ h ), h is the total number of quantitative indicators, h=h1+h2+h3; index ik is the kth quantitative indicator of the i-th business, index kthreshold is the threshold of the kth quantitative index,! (mi) is the logical NOT operation of (mi), f(index ik ,index kthreshold )·λ k is the indicator scoring function of the kth quantitative indicator of the i-th business.

15. The method according to any one of claims 3 to 11, wherein: The calculation of comprehensive optimization parameters for each business establishment state based on the timely rewards for each business establishment state includes: Calculate the expected return in the current business establishment state based on the timely rewards in each business establishment state after the next business establishment state; The comprehensive optimization parameters under the current business establishment state are calculated and updated according to the expected return under the current business establishment state.

16. The method according to claim 15, wherein Calculate the expected return based on the current business establishment status using the following formula: Among them, G t Establish state S for the business t Next, perform action a t The expected return, γ k R t+k+1 Discount coefficient, 0<γ k <1; R t+k+1 To establish status in business t+k+1 The timely reward obtained under the business establishment state S t The number of services created under t=(0, ..., m), where m is the total number of services to be built in the OTN network.

17. An OTN network resource optimization device, comprising: a first processing module, a second processing module and an update module, The first processing module is configured to determine, based on an action strategy, a pending service in a current service establishment state, create the pending service, calculate a timely reward in the current service establishment state, enter a next service establishment state, and continue until a round ends; calculate a comprehensive optimization parameter in each service establishment state based on the timely reward in each service establishment state; and calculate and update a quantitative indicator weight vector based on the comprehensive optimization parameter in each service establishment state, wherein the action strategy is a probability function associated with the quantitative indicator weight vector, and the quantitative indicator weight vector corresponds to the quantitative indicators of multiple network resources; The second processing module is used to iterate a preset number of rounds to obtain the optimal quantitative indicator weight vector; The updating module is used to update the action strategy according to the optimal quantitative indicator weight vector.

18. A computer device comprising: one or more processors; a storage device having one or more programs stored thereon; When the one or more programs are executed by the one or more processors, the one or more processors implement the OTN network resource optimization method according to any one of claims 1 to 16.

19. A computer-readable medium having a computer program stored thereon, wherein: When the program is executed, the OTN network resource optimization method according to any one of claims 1 to 16 is implemented.

Citation Information

Patent Citations

  • Implementing method for optimizing multiple services in optical synchronization data transportation network

    CN1540934A

  • Reinforcement learning for autonomous telecommunications networks

    US20190138948A1