Optical network spectrum allocation network acquisition method, storage medium and computer device

By training and optimizing the parameters of the spectrum allocation network for multi-domain elastic optical networks, the problem of inefficiency in multi-domain elastic optical networks when processing cross-domain service requests is solved, achieving more efficient processing and lower probability of obstruction.

CN115988364BActive Publication Date: 2025-06-06SOUTH CHINA NORMAL UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211694487.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-12-28
Publication Date
2025-06-06
Estimated Expiration
2042-12-28

AI Technical Summary

Technical Problem

Multi-domain elastic optical networks are inefficient when processing cross-domain service requests, with high probability of blocking the processing results and low probability of smooth execution.

Method used

By obtaining the training service request and network environment status, input it to the inter-domain spectrum allocation network for parameter training, the target inter-domain spectrum allocation network is obtained, and combined with the in-domain spectrum allocation network, the spectrum allocation action is optimized to improve processing efficiency.

Benefits of technology

The multi-domain elastic optical network has improved the processing efficiency of cross-domain service requests, reduced the probability of blockage of processing results, and increased the probability of smooth execution of processing results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115988364B_ABST
    Figure CN115988364B_ABST
Patent Text Reader

Abstract

The present application provides a spectrum allocation network acquisition method, storage medium and computer device for an optical network, which are applied to a multi-domain elastic optical network including multiple optical network autonomous domains, and the method includes: when obtaining a training service request, the inter-domain spectrum allocation network is parameter-trained according to the network environment status, the inter-domain spectrum allocation action and the first action reward to obtain a target inter-domain spectrum allocation network; the intra-domain spectrum allocation network is parameter-trained according to the intra-domain environment status, the intra-domain spectrum allocation action and the second action reward to obtain a target intra-domain spectrum allocation network for the relevant domain; the target inter-domain spectrum allocation network and the target intra-domain spectrum allocation network are determined as the target spectrum allocation network. The present application can improve the processing efficiency of the multi-domain elastic optical network for cross-domain service requests, reduce the probability of obstruction of the processing results, and increase the probability of smooth execution of the processing results.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical field of spectrum allocation, and in particular to a spectrum allocation network acquisition method, storage medium and computer device for an optical network. Background Art

[0002] In order to improve the scalability of the network, large elastic optical networks are usually divided into multiple autonomous domains, which are called multi-domain elastic optical networks. However, since each autonomous domain in a multi-domain elastic optical network has a high degree of autonomy and privacy, the coordination between each autonomous domain is poor, resulting in low efficiency in processing cross-domain service requests in the multi-domain elastic optical network, a high probability of obstruction in processing results, and a low probability of smooth execution. Summary of the invention

[0003] The purpose of this application is to overcome the shortcomings and deficiencies in the prior art and to provide a spectrum allocation network acquisition method, storage medium and computer equipment for an optical network, which can improve the processing efficiency of a multi-domain elastic optical network for cross-domain service requests, reduce the probability of obstruction of processing results, and increase the probability of smooth execution of processing results.

[0004] A first aspect of an embodiment of the present application provides a spectrum allocation network acquisition method for an optical network, which is applied to a multi-domain elastic optical network, wherein the multi-domain elastic optical network includes multiple optical network autonomous domains, including:

[0005] Acquire a plurality of training service requests, and when acquiring each of the training service requests, acquire a network environment status of the multi-domain elastic optical network;

[0006] Inputting each of the training service requests and each of the network environment states into an inter-domain spectrum allocation network to obtain a corresponding plurality of inter-domain spectrum allocation actions;

[0007] According to a preset first action reward rule, obtaining a first action reward corresponding to each of the inter-domain spectrum allocation actions;

[0008] According to the network environment status, the inter-domain spectrum allocation action and the first action reward corresponding to each of the training service requests, parameter training is performed on the inter-domain spectrum allocation network to obtain a target inter-domain spectrum allocation network;

[0009] According to the inter-domain spectrum allocation action, a plurality of related domains, an intra-domain environment state of each related domain, and a starting node and a terminal node in each related domain that are related to the inter-domain spectrum allocation action are acquired from the optical network autonomous domain;

[0010] Inputting the start node and the end node of each of the related domains and the intra-domain environment state of each of the related domains into the corresponding intra-domain spectrum allocation network to obtain the corresponding intra-domain spectrum allocation action;

[0011] According to a preset second action reward rule, obtaining a second action reward corresponding to each of the intra-domain spectrum allocation actions;

[0012] According to the intra-domain environment state of each relevant domain corresponding to each training service request, the intra-domain spectrum allocation action and the second action reward, parameter training is performed on the intra-domain spectrum allocation network of each relevant domain to obtain a target intra-domain spectrum allocation network of each relevant domain;

[0013] The target inter-domain spectrum allocation network and the target intra-domain spectrum allocation network are determined as target spectrum allocation networks.

[0014] A second aspect of an embodiment of the present application provides a computer-readable storage medium, wherein the computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps of the spectrum allocation network acquisition method for an optical network are implemented as described above.

[0015] A third aspect of an embodiment of the present application provides a computer device, including a memory, a processor, and a computer program stored in the memory and executable by the processor, wherein the processor implements the steps of the spectrum allocation network acquisition method for the optical network as described above when executing the computer program.

[0016] Compared with the related art, the present application inputs several network environment states corresponding to several training service requests into the inter-domain spectrum allocation network, obtains multiple inter-domain spectrum allocation actions and corresponding first action rewards, and then performs parameter training on the inter-domain spectrum allocation network according to the network environment state, inter-domain spectrum allocation action and first action reward corresponding to each training service request to obtain a target inter-domain spectrum allocation network; obtains the intra-domain spectrum allocation action and the corresponding second action reward of each related domain according to the intra-domain environment state of each related domain corresponding to the inter-domain spectrum allocation action, and performs parameter training on the intra-domain spectrum allocation network of each related domain according to the intra-domain environment state, intra-domain spectrum allocation action and second action reward of each related domain to obtain the target intra-domain spectrum allocation network of each related domain, and then determines the target inter-domain spectrum allocation network and the target intra-domain spectrum allocation network as the target spectrum allocation network. The obtained target spectrum allocation network can improve the processing efficiency of the multi-domain elastic optical network receiving cross-domain service requests, reduce the probability of obstruction of the processing results, and increase the probability of smooth execution of the processing results.

[0017] In order to provide a clearer understanding of the present application, the specific implementation of the present application will be described below in conjunction with the accompanying drawings. BRIEF DESCRIPTION OF THE DRAWINGS

[0018] Figure 1 This is a flow chart of a spectrum allocation network acquisition method for an optical network according to an embodiment of the present application.

[0019] Figure 2 A schematic diagram of a multi-domain elastic optical network of a spectrum allocation network acquisition method for an optical network according to an embodiment of the present application. DETAILED DESCRIPTION

[0020] In order to make the objectives, technical solutions and advantages of the present application more clear, the embodiments of the present application will be further described in detail below with reference to the accompanying drawings.

[0021] It should be clear that the described embodiments are only part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the embodiments of the present application.

[0022] When the following description relates to the accompanying drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. In the description of the present application, it should be understood that the terms "first", "second", "third", etc. are only used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence, nor can they be understood as indicating or implying relative importance. For those of ordinary skill in the art, the specific meanings of the above terms in the present application can be understood according to the specific circumstances. The singular forms of "a", "said" and "the" used in the present application and the appended claims are also intended to include the majority form, unless the context clearly indicates other meanings. The words "if" / "if" used herein can be interpreted as "at the time of" or "when" or "in response to determination".

[0023] In addition, in the description of this application, unless otherwise specified, "plurality" means two or more. "And / or" describes the association relationship of associated objects, indicating that three relationships may exist. For example, A and / or B can mean: A exists alone, A and B exist at the same time, and B exists alone. The character " / " generally indicates that the associated objects are in an "or" relationship.

[0024] See also Figure 1 , which is a flowchart of a spectrum allocation network acquisition method for an optical network of an embodiment of the present application. The spectrum allocation network acquisition method for an optical network is applied to a multi-domain elastic optical network, and the multi-domain elastic optical network includes multiple optical network autonomous domains.

[0025] Among them, each optical network autonomous domain is a large-scale local area network that uses optical fiber as the main transmission medium. The optical fiber is connected through nodes, and multiple adjacent nodes are connected to form an optical network link. The frequency slot is the unit for storing and transmitting data in the link, and the commonly used bandwidth of a single frequency slot is 12.5GHz. And each optical network autonomous domain will be connected to another optical network autonomous domain through at least one inter-domain link, that is, the connection mode between optical network autonomous domains can be expressed as an optical network autonomous domain connected to multiple different optical network autonomous domains through multiple inter-domain links, or an optical network autonomous domain connected to another optical network autonomous domain through multiple inter-domain links. Among them, the inter-domain links between optical network autonomous domains can adopt preset modulation formats and spectrum allocation rules, but the modulation formats and spectrum allocation rules of all inter-domain links in the same multi-domain elastic optical network are the same.

[0026] The spectrum allocation network acquisition method of the optical network includes:

[0027] S1: Acquire a plurality of training service requests, and when acquiring each of the training service requests, acquire the network environment status of the multi-domain elastic optical network.

[0028] A training service request refers to a traffic request received by a multi-domain elastic optical network. The training service request may be for only one optical network autonomous domain in the multi-domain elastic optical network, or for multiple optical network autonomous domains in the multi-domain elastic optical network. A training service request may be expressed as TR{src, dst, B}, where src represents the source node, dst represents the destination node, and B represents the bandwidth requirement in Gbps. When a training service request is for only one optical network autonomous domain in the multi-domain elastic optical network, the source node and the destination node of the training service request are both located in the same optical network autonomous domain; when a training service request is for multiple optical network autonomous domains in the multi-domain elastic optical network, the source node and the destination node of the training service request are both located in different optical network autonomous domains.

[0029] The network environment state of the multi-domain elastic optical network includes the source optical network autonomous domain where the source node of the training service request is located and the target optical network autonomous domain where the target node is located, as well as multiple candidate network virtual paths that meet the training service request when the corresponding training service request is received and the feasible probability of each candidate network virtual path, where each candidate network virtual path is represented by a virtual node representing an optical network autonomous domain and an inter-domain link. The network environment state of the multi-domain elastic optical network can be expressed as s H =[E src , E dst , B, P 1 , P 2 ,…,P K ]; where s Hrepresents the network environment status of the multi-domain elastic optical network, E src represents the source optical network autonomous domain where the source node of the training service request is located, E dst represents the target optical network autonomous domain where the target node is located; P 1 , P 2 ,…,P K Indicates the feasible probability of the corresponding network virtual path.

[0030] See also Figure 2 , assuming that the multi-domain elastic optical network includes three optical network autonomous domains, Domain 1, Domain 2, and Domain 3, and the source node of the training service request is the node numbered 3 in the optical network autonomous domain Domain 1, which is represented as The target node of the training service request is the node numbered 3 in the optical network autonomous domain Domain 2, which is represented by The candidate network virtual paths of the network environment state of the multi-domain elastic optical network include:

[0031] Network virtual path 1:

[0032] Network Virtual Path 2:

[0033] Network Virtual Path 3:

[0034] Network Virtual Path 4:

[0035] Network Virtual Path 5:

[0036] Network Virtual Path 6:

[0037] Among them, taking network virtual path 3 as an example, Representation Node and nodes The connection relationship is the node numbered 6 in the optical network autonomous domain Domain 1. The egress node of the optical network autonomous domain Domain 1 and the node numbered 1 in the optical network autonomous domain Domain 3 It is the entry node of the optical network autonomous domain Domain 3, so Indicates an inter-domain link between optical network autonomous domains Domain 1 and Domain 3.

[0038] The feasible probability of each network virtual path is the product of the feasible probabilities of the corresponding virtual nodes and inter-domain links. Taking network virtual path 3 as an example, its feasible probability is:

[0039]

[0040] Among them, P represents the probability of the corresponding parameter. If the inter-domain link meets the training service request, its corresponding feasible probability is 1, otherwise, its corresponding feasible probability is 0. The feasible probability of the virtual node is calculated by the following formula:

[0041]

[0042] Among them, D i Indicates the optical network autonomous domain Domaini, in is used to indicate the entry node or source node of the optical network autonomous domain, The m node represents the optical network autonomous domain Domaini, and out is used to indicate the exit node or target node of the optical network autonomous domain. represents n nodes of the optical network autonomous domain Domaini, Represents the connected nodes in the optical network autonomous domain Domaini and The number of candidate paths that satisfy the training service request, Represents the connected nodes in the optical network autonomous domain Domaini and The total number of paths in the candidate domain. and When they are the same node, the feasible probability of the optical network autonomous domain Domaini is 1.

[0043] S2: Input each of the training service requests and each of the network environment states into the inter-domain spectrum allocation network to obtain a corresponding plurality of inter-domain spectrum allocation actions.

[0044] The inter-domain spectrum allocation action refers to selecting a target network virtual path from the candidate network virtual paths of the corresponding network environment state as the inter-domain path for executing the training service request. For each network virtual path, it is also necessary to calculate the number of continuous time slots required for each inter-domain link, and then check whether the inter-domain link of the network virtual path has enough available continuous time slots to support the training service request according to the spectrum continuity constraint, and judge whether the intra-domain environment state of each optical network autonomous domain corresponding to the network virtual path supports the training service request according to the two rules of spectrum continuity and adjacency constraints. If the inter-domain link or optical network autonomous domain corresponding to the network virtual path selected by the inter-domain spectrum allocation action does not support the training service request, it means that the inter-domain spectrum allocation action does not support the training service request. If the inter-domain link and optical network autonomous domain corresponding to the network virtual path selected by the inter-domain spectrum allocation action both support the training service request, it means that the inter-domain spectrum allocation action supports the training service request.

[0045] S3: According to a preset first action reward rule, obtain a first action reward corresponding to each of the inter-domain spectrum allocation actions.

[0046] Among them, the first action reward includes positive and negative values, which are used to indicate the instantaneous benefit of the inter-domain spectrum allocation action. When the inter-domain spectrum allocation action supports the training service request, the corresponding first action reward is a positive value. When the inter-domain spectrum allocation action does not support the training service request, the corresponding first action reward is a negative value.

[0047] S4: performing parameter training on the inter-domain spectrum allocation network according to the network environment status, the inter-domain spectrum allocation action and the first action reward corresponding to each of the training service requests to obtain a target inter-domain spectrum allocation network.

[0048] Through parameter training, the target inter-domain spectrum allocation network can output the inter-domain spectrum allocation action faster, improve the processing efficiency of service requests, reduce the probability of obstruction of the inter-domain spectrum allocation action, and increase the probability of smooth execution of the inter-domain spectrum allocation action.

[0049] S5: According to the inter-domain spectrum allocation action, a number of related domains, the intra-domain environmental status of each related domain, and the starting node and the terminal node related to the inter-domain spectrum allocation action in each related domain are obtained from the optical network autonomous domain.

[0050] The relevant domain refers to the autonomous domain of the optical network related to the inter-domain spectrum allocation action. The intra-domain environmental state of the relevant domain includes the starting node and the terminal node in the relevant domain pointed to by the training service request or the inter-spectrum allocation action, and when receiving the corresponding training service request or the inter-spectrum allocation action, there are several candidate intra-domain virtual paths between the starting node and the terminal node that satisfy the training service request or the inter-spectrum allocation action. Among them, the starting node in the relevant domain can be the entry node in the relevant domain or the source node of the training service request, and the terminal node in the relevant domain can be the exit node in the relevant domain or the target node of the training service request.

[0051] S6: Inputting the starting node and the terminal node of each of the related domains and the intra-domain environmental state of each of the related domains into the corresponding intra-domain spectrum allocation network to obtain the corresponding intra-domain spectrum allocation action.

[0052] The intra-domain spectrum allocation action refers to selecting an intra-domain virtual path from a number of candidate intra-domain virtual paths between the start node and the end node that meet the training service request. It is also necessary to determine whether the intra-domain virtual path supports the corresponding training service request or inter-domain spectrum allocation action based on the two rules of spectrum continuity and adjacency constraints.

[0053] S7: According to a preset second action reward rule, obtain a second action reward corresponding to each of the intra-domain spectrum allocation actions.

[0054] Among them, the second action reward includes positive and negative values, which are used to indicate the instantaneous benefit of the intra-domain spectrum allocation action. When the intra-domain spectrum allocation action supports the corresponding training service request or inter-spectrum allocation action, the second action reward is a positive value, for example, it can be 1. When the intra-domain spectrum allocation action does not support the corresponding training service request or inter-spectrum allocation action, the second action reward is a negative value, for example, it can be -1.

[0055] S8: According to the intra-domain environmental status of each relevant domain corresponding to each of the training service requests, the intra-domain spectrum allocation action and the second action reward, parameter training is performed on the intra-domain spectrum allocation network of each relevant domain to obtain the target intra-domain spectrum allocation network of each relevant domain.

[0056] Through parameter training, the target domain spectrum allocation network can output the domain spectrum allocation action more quickly, improve the processing efficiency of service requests, reduce the probability of obstruction of the domain spectrum allocation action, and increase the probability of smooth execution of the domain spectrum allocation action.

[0057] S9: Determine the target inter-domain spectrum allocation network and the target intra-domain spectrum allocation network as target spectrum allocation networks.

[0058] Compared with the related art, the present application inputs several network environment states corresponding to several training service requests into the inter-domain spectrum allocation network, obtains multiple inter-domain spectrum allocation actions and corresponding first action rewards, and then performs parameter training on the inter-domain spectrum allocation network according to the network environment states, inter-domain spectrum allocation actions and first action rewards corresponding to each training service request to obtain the target inter-domain spectrum allocation network; obtains the intra-domain spectrum allocation action and the corresponding second action reward of each relevant domain according to the intra-domain environment states of each relevant domain corresponding to the inter-domain spectrum allocation action and the corresponding starting node and terminal node; and obtains the intra-domain spectrum allocation action and the corresponding second action reward of each relevant domain according to the intra-domain environment states, intra-domain spectrum allocation action and the second action reward of each relevant domain. As a reward, parameter training is performed on the intra-domain spectrum allocation network of each relevant domain to obtain the target intra-domain spectrum allocation network of each relevant domain, and then the target inter-domain spectrum allocation network and the target intra-domain spectrum allocation network are determined as the target spectrum allocation network. The obtained target spectrum allocation network can output the inter-domain spectrum allocation action and the intra-domain spectrum allocation action more quickly, and reduce the obstruction probability of the output inter-domain spectrum allocation action and the intra-domain spectrum allocation action, and increase the probability of the smooth execution of the inter-domain spectrum allocation action and the intra-domain spectrum allocation action, thereby achieving the technical effect of improving the processing efficiency of the multi-domain elastic optical network receiving cross-domain service requests, reducing the obstruction probability of the processing results, and increasing the probability of the smooth execution of the processing results.

[0059] In a feasible embodiment, the step of S3: obtaining a first action reward corresponding to each of the inter-domain spectrum allocation actions according to a preset first action reward rule, includes:

[0060] If the inter-domain spectrum allocation action is executed smoothly, the first action reward corresponding to the inter-domain spectrum allocation action is a positive reward value; if the inter-domain spectrum allocation action is blocked, the first action reward corresponding to the inter-domain spectrum allocation action is a negative reward value.

[0061] In this embodiment, the positive reward value is a positive value, and the negative reward value is a negative value. Depending on whether the inter-domain spectrum allocation action is successfully executed, the corresponding positive reward value or negative reward value is used as the first action reward, which is beneficial to indicate the instantaneous benefit of the inter-domain spectrum allocation action according to the first action reward, and indicate the training direction of the parameter training of the inter-domain spectrum allocation network according to the first action reward.

[0062] In a feasible embodiment, the network environment state includes a plurality of inter-domain paths satisfying the corresponding training service request, and a path probability of each of the inter-domain paths;

[0063] The step of providing a first action reward corresponding to the inter-domain spectrum allocation action with a positive reward value if the inter-domain spectrum allocation action is successfully executed includes:

[0064] S301: If the path probability of the inter-domain path pointed to by the inter-domain spectrum allocation action is equal to the maximum value of the path probability in the network environment state, determine a first positive reward value as a corresponding first action reward.

[0065] S302: If the path probability of the inter-domain path pointed to by the inter-domain spectrum allocation action is less than the maximum value of the path probability in the network environment state and greater than the minimum value of the path probability in the network environment state, determine the second positive reward value as the corresponding first action reward.

[0066] S303: If the path probability of the inter-domain path pointed to by the inter-domain spectrum allocation action is equal to the minimum value of the path probability in the network environment state, determine the third positive reward value as the corresponding first action reward.

[0067] The first positive reward value is greater than the second positive reward value, and the second positive reward value is greater than the third positive reward value.

[0068] In this embodiment, steps S301-S303 do not limit the order of execution. According to the size of the path probability of the inter-domain path pointed to by the inter-domain spectrum allocation action, the value of the first action reward is determined, so that the target inter-domain spectrum allocation network trained according to the first action reward can increase the probability of using the inter-domain path that is executed smoothly and has a large path probability as the inter-domain spectrum allocation action.

[0069] In a feasible embodiment, the network environment state includes a plurality of inter-domain paths satisfying the corresponding training service request, and a path probability of each of the inter-domain paths;

[0070] The step of setting a first action reward corresponding to the inter-domain spectrum allocation action as a negative reward value if the inter-domain spectrum allocation action is blocked includes:

[0071] S311: If the path probability of the inter-domain path pointed to by the inter-domain spectrum allocation action is equal to the minimum value of the path probability in the network environment state, determine a first negative reward value as a corresponding first action reward.

[0072] S312: If the path probability of the inter-domain path pointed to by the inter-domain spectrum allocation action is greater than the minimum value of the path probability in the network environment state, determine the second negative reward value as the corresponding first action reward.

[0073] Wherein, the first negative reward value is less than the second negative reward value.

[0074] In this embodiment, steps S311-S312 do not limit the order of execution. According to the size of the path probability of the inter-domain path pointed to by the inter-domain spectrum allocation action, the value of the first action reward is determined, so that the target inter-domain spectrum allocation network obtained by training according to the first action reward can reduce the probability of using the inter-domain path with blocked execution and small path probability as the inter-domain spectrum allocation action.

[0075] In a feasible embodiment, the inter-domain spectrum allocation network includes a first spectrum allocation action network for outputting the inter-domain spectrum allocation action, and a first spectrum allocation criticism network for outputting a first value function;

[0076] The step of performing parameter training on the inter-domain spectrum allocation network according to the network environment status, the inter-domain spectrum allocation action and the first action reward corresponding to each of the training service requests to obtain a target inter-domain spectrum allocation network comprises:

[0077] S401: Acquire a plurality of first value functions output by the first spectrum allocation criticism network according to the network environment status and the inter-domain spectrum allocation action corresponding to each of the training service requests.

[0078] The first value function accumulates the first action reward starting from the network environment state corresponding to the first training service request, and the accumulation process will result in a discount calculation of the first action reward.

[0079] S402: Obtain a first advantage function according to multiple first action rewards and multiple first value functions.

[0080] Specifically, the first advantage function can be obtained by the following formula:

[0081]

[0082] in, represents the first advantage function; n-1 represents the number of training service requests; represents the first action reward of the corresponding training service request; γH represents the first discount coefficient; and They respectively represent the first value functions of the corresponding training service requests.

[0083] S403: Perform parameter training on the first spectrum allocation network and the first spectrum allocation critic network according to the first advantage function to update network parameters of the first spectrum allocation network and the first spectrum allocation critic network.

[0084] Specifically, step S403 includes:

[0085] The update target of the network parameters of the first spectrum allocation network is obtained by the following formula:

[0086]

[0087] Among them, grad(θ H ) represents the update target of the network parameters of the first spectrum allocation network, β represents the strength of the entropy regularization term, is the wedge operator, Represents the entropy of the corresponding policy distribution.

[0088] The network parameters of the first spectrum allocation network are updated by the following formula:

[0089]

[0090] in, representing updated network parameters of the first spectrum allocation network; represents the network parameters of the first spectrum allocation network before updating, α H Represents the corresponding learning rate.

[0091] The update target of the network parameters of the first spectrum allocation criticism network is obtained by the following formula:

[0092]

[0093] in, represents the update target of the network parameters of the network parameters of the first spectrum allocation critical network, t 0 Indicates the first training service request, t 0 +T-1 represents the total number of training service requests.

[0094] The network parameters of the first spectrum allocation criticism network are updated by the following formula:

[0095]

[0096] in, represents the updated network parameters of the first spectrum allocation criticism network, represents the network parameters of the first spectrum allocation criticism network before update, η H Represents the corresponding learning rate.

[0097] S404: Determine the first spectrum allocation network after updating network parameters and the first spectrum allocation criticism network after updating network parameters as the target inter-domain spectrum allocation network.

[0098] In this embodiment, the target inter-domain spectrum allocation network obtained through training in steps S401-S404 can output an inter-domain spectrum allocation action with a higher execution success rate.

[0099] In a feasible embodiment, the step of obtaining the second action reward corresponding to each of the intra-domain spectrum allocation actions according to a preset second action reward rule in S7 includes:

[0100] If the intra-domain spectrum allocation action is executed smoothly, the second action reward corresponding to the intra-domain spectrum allocation action is a positive reward value; if the intra-domain spectrum allocation action is blocked, the second action reward corresponding to the intra-domain spectrum allocation action is a negative reward value.

[0101] In this embodiment, the positive reward value is a positive value, and the negative reward value is a negative value. Depending on whether the intra-domain spectrum allocation action is successfully executed, the corresponding positive reward value or negative reward value is used as the second action reward, which is beneficial to indicate the instantaneous benefit of the intra-domain spectrum allocation action according to the second action reward, and indicate the training direction of the parameter training of the intra-domain spectrum allocation network according to the second action reward.

[0102] In a feasible embodiment, the intra-domain environmental state includes a plurality of intra-domain paths satisfying the training service request, and a path probability of each of the intra-domain paths;

[0103] The step of providing a negative reward value for the second action corresponding to the intra-domain spectrum allocation action if the execution of the intra-domain spectrum allocation action is blocked includes:

[0104] S701: If the path probability of the inter-domain path pointed to by the inter-domain spectrum allocation action corresponding to the intra-domain spectrum allocation action is equal to the maximum path probability in the network environment state, determine the third negative reward value as the corresponding second action reward.

[0105] S702: If the path probability of the inter-domain path pointed to by the inter-domain spectrum allocation action corresponding to the intra-domain spectrum allocation action is less than the maximum value of the path probability in the network environment state and greater than the minimum value of the path probability in the network environment state, the fourth negative reward value is determined as the corresponding second action reward.

[0106] S703: If the path probability of the inter-domain path pointed to by the inter-domain spectrum allocation action corresponding to the intra-domain spectrum allocation action is equal to the minimum value of the path probability in the network environment state, determine the fifth negative reward value as the corresponding second action reward.

[0107] Among them, the third negative reward value is less than the fourth negative reward value, and the fourth negative reward value is less than the fifth negative reward value. For example, the third negative reward value is -3, the fourth negative reward value is -1, and the fifth negative reward value is -0.8. This is because when the execution of the intra-domain spectrum allocation action is blocked, the greater the path probability of the inter-domain path pointed to by the inter-domain spectrum allocation action corresponding to the intra-domain spectrum allocation action, the better the inter-domain path corresponding to the corresponding inter-domain spectrum allocation action, but at this time the execution of the intra-domain spectrum allocation action is blocked, so the corresponding intra-domain spectrum allocation action should be given a smaller reward value.

[0108] In this embodiment, steps S701-S703 do not limit the priority order of execution. According to the path probability of the inter-domain path pointed to by the inter-domain spectrum allocation action corresponding to the intra-domain spectrum allocation action, the value of the second action reward corresponding to the intra-domain spectrum allocation action is determined, so that the target intra-domain spectrum allocation network trained according to the second action reward can reduce the probability of executing the blocked intra-domain path as the intra-domain spectrum allocation action.

[0109] In a feasible embodiment, the intra-domain spectrum allocation network includes a second spectrum allocation action network for outputting the intra-domain spectrum allocation action, and a second spectrum allocation criticism network for outputting a second value function;

[0110] The step S8: performing parameter training on the intra-domain spectrum allocation network of each relevant domain according to the intra-domain environment state of each relevant domain corresponding to each training service request, the intra-domain spectrum allocation action and the second action reward to obtain the target intra-domain spectrum allocation network of each relevant domain, comprises:

[0111] S801: Acquire a plurality of second value functions output by the second spectrum allocation criticism network of each relevant domain according to the intra-domain environmental state and the intra-domain spectrum allocation action of each relevant domain corresponding to each training service request.

[0112] Among them, the second value function of the brother-related domain starts to accumulate the second action reward from the domain environment state corresponding to the first training service request, and the accumulation process will result in a discount calculation of the second action reward.

[0113] S802: Obtain a second advantage function for each relevant domain according to the plurality of second action rewards and the plurality of second value functions for each relevant domain.

[0114] Specifically, the second advantage function can be obtained by the following formula:

[0115]

[0116] in, represents the second advantage function; represents the second action reward of the corresponding training service request; γ l represents the second discount factor; and They respectively represent the second value functions of the corresponding training service requests.

[0117] S803: Perform parameter training on the second spectrum allocation network and the second spectrum allocation criticism network of each relevant domain according to the second advantage function of each relevant domain to update network parameters of the second spectrum allocation network and the second spectrum allocation criticism network of each relevant domain.

[0118] Specifically, step S403 includes:

[0119] The update target of the network parameters of the second spectrum allocation network is obtained by the following formula:

[0120]

[0121] Among them, grad(θ l ) represents the update target of the network parameters of the second spectrum allocation network, and β represents the strength of the entropy regularization term; Represents the entropy of the corresponding policy distribution.

[0122] The network parameters of the second spectrum allocation network are updated by the following formula:

[0123]

[0124] in, representing updated network parameters of the second spectrum allocation network; represents the network parameters of the second spectrum allocation network before updating; α l Represents the corresponding learning rate.

[0125] The update target of the network parameters of the second spectrum allocation criticism network is obtained by the following formula:

[0126]

[0127] in, represents the update target of the network parameters of the second spectrum allocation critic network, t 0 Indicates the first training service request, t 0 +T-1 represents the total number of training service requests.

[0128] The network parameters of the second spectrum allocation criticism network are updated by the following formula:

[0129]

[0130] in, represents the updated network parameters of the second spectrum allocation criticism network, represents the network parameters of the second spectrum allocation criticism network before update, η l Represents the corresponding learning rate.

[0131] S804: Determine the second spectrum allocation network after updating the network parameters of each related domain and the second spectrum allocation criticism network after updating the network parameters as the target domain intra-domain spectrum allocation network of each related domain.

[0132] In this embodiment, the target domain spectrum allocation network obtained by training in steps S801-S804 can output an intra-domain spectrum allocation action with a higher execution success rate.

[0133] A second aspect of an embodiment of the present application provides a computer-readable storage medium, wherein the computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps of the spectrum allocation network acquisition method for an optical network are implemented as described above.

[0134] A third aspect of an embodiment of the present application provides a computer device, including a memory, a processor, and a computer program stored in the memory and executable by the processor, wherein the processor implements the steps of the spectrum allocation network acquisition method for the optical network as described above when executing the computer program.

[0135] The device embodiments described above are merely illustrative, wherein the components described as separate parts may or may not be physically separated, and the parts displayed as units may or may not be physical units, that is, they may be located in one place, or they may be distributed on multiple network units. Some or all of the modules may be selected according to actual needs to achieve the purpose of the present application. Ordinary technicians in this field can understand and implement it without creative work.

[0136] Those skilled in the art will appreciate that the embodiments of the present application may be provided as methods, systems, or computer program products. Therefore, the present application may adopt the form of a complete hardware embodiment, a complete software embodiment, or an embodiment in combination with software and hardware. Moreover, the present application may adopt the form of a computer program product implemented in one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) that include computer-usable program code.

[0137] The present application is described with reference to the flowcharts and / or block diagrams of the methods, devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each process and / or box in the flowchart and / or block diagram, as well as the combination of the processes and / or boxes in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 A process or multiple processes and / or boxes Figure 1 These computer program instructions can also be stored in a computer-readable memory that can guide a computer or other programmable data processing device to work in a specific way, so that the instructions stored in the computer-readable memory produce a product including an instruction device, which implements the function selected in the process. Figure 1 A process or multiple processes and / or boxes Figure 1 function selected in a box or multiple boxes.

[0138] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operating steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing instructions for implementing the process. Figure 1 A process or multiple processes and / or boxes Figure 1 steps for the function selected in a box or multiple boxes.

[0139] In a typical configuration, a computing device includes one or more processors (CPU), input / output interfaces, network interfaces, and memory.

[0140] The memory may include non-permanent memory in a computer-readable medium, random access memory (RAM) and / or non-volatile memory in the form of read-only memory (ROM) or flash RAM. The memory is an example of a computer-readable medium.

[0141] Computer readable media include permanent and non-permanent, removable and non-removable media that can be implemented by any method or technology to store information. Information can be computer readable instructions, data structures, program modules or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technology, compact disk read-only memory (CD-ROM), digital versatile disk (DVD) or other optical storage, magnetic cassettes, magnetic tape magnetic disk storage or other magnetic storage devices or any other non-transmission media that can be used to store information that can be accessed by a computing device. As defined herein, computer readable media does not include temporary computer readable media (transitory media), such as modulated data signals and carrier waves.

[0142] It should also be noted that the terms "include", "comprises" or any other variations thereof are intended to cover non-exclusive inclusion, so that a process, method, commodity or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, commodity or device. In the absence of more restrictions, the elements defined by the sentence "comprises a ..." do not exclude the existence of other identical elements in the process, method, commodity or device including the elements.

[0143] The above are only embodiments of the present application and are not intended to limit the present application. For those skilled in the art, the present application may have various changes and variations. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application should be included within the scope of the claims of the present application.

Claims

1. A spectrum allocation network acquisition method for an optical network, It is characterized in that Applied to a multi-domain elastic optical network, the multi-domain elastic optical network includes multiple optical network autonomous domains, including: Acquire a plurality of training service requests, and when acquiring each of the training service requests, acquire a network environment status of the multi-domain elastic optical network; Inputting each of the training service requests and each of the network environment states into an inter-domain spectrum allocation network to obtain a corresponding plurality of inter-domain spectrum allocation actions; According to a preset first action reward rule, obtaining a first action reward corresponding to each of the inter-domain spectrum allocation actions; According to the network environment status, the inter-domain spectrum allocation action and the first action reward corresponding to each of the training service requests, parameter training is performed on the inter-domain spectrum allocation network to obtain a target inter-domain spectrum allocation network; According to the inter-domain spectrum allocation action, a plurality of related domains, an intra-domain environment state of each related domain, and a starting node and an end node in each related domain that are related to the inter-domain spectrum allocation action are acquired from the optical network autonomous domain; Inputting the start node and the end node of each of the related domains and the intra-domain environment state of each of the related domains into the corresponding intra-domain spectrum allocation network to obtain the corresponding intra-domain spectrum allocation action; According to a preset second action reward rule, obtaining a second action reward corresponding to each of the intra-domain spectrum allocation actions; According to the intra-domain environment state of each relevant domain corresponding to each training service request, the intra-domain spectrum allocation action and the second action reward, parameter training is performed on the intra-domain spectrum allocation network of each relevant domain to obtain a target intra-domain spectrum allocation network of each relevant domain; The target inter-domain spectrum allocation network and the target intra-domain spectrum allocation network are determined as target spectrum allocation networks.

2. The spectrum allocation network acquisition method of an optical network according to claim 1, It is characterized in that The step of obtaining a first action reward corresponding to each of the inter-domain spectrum allocation actions according to a preset first action reward rule includes: If the inter-domain spectrum allocation action is executed smoothly, the first action reward corresponding to the inter-domain spectrum allocation action is a positive reward value; if the inter-domain spectrum allocation action is blocked, the first action reward corresponding to the inter-domain spectrum allocation action is a negative reward value.

3. The spectrum allocation network acquisition method for an optical network according to claim 2, It is characterized in that The network environment state includes a plurality of inter-domain paths satisfying the corresponding training service request, and a path probability of each of the inter-domain paths; The step of providing a first action reward corresponding to the inter-domain spectrum allocation action with a positive reward value if the inter-domain spectrum allocation action is successfully executed includes: If the path probability of the inter-domain path pointed to by the inter-domain spectrum allocation action is equal to the maximum value of the path probability in the network environment state, determining a first positive reward value as a corresponding first action reward; If the path probability of the inter-domain path pointed to by the inter-domain spectrum allocation action is less than the maximum value of the path probability in the network environment state and greater than the minimum value of the path probability in the network environment state, determining the second positive reward value as the corresponding first action reward; If the path probability of the inter-domain path pointed to by the inter-domain spectrum allocation action is equal to the minimum value of the path probability in the network environment state, determining a third positive reward value as the corresponding first action reward; The first positive reward value is greater than the second positive reward value, and the second positive reward value is greater than the third positive reward value.

4. The spectrum allocation network acquisition method for an optical network according to claim 2, It is characterized in that The network environment state includes a plurality of inter-domain paths satisfying the corresponding training service request, and a path probability of each of the inter-domain paths; The step of setting a first action reward corresponding to the inter-domain spectrum allocation action as a negative reward value if the inter-domain spectrum allocation action is blocked includes: If the path probability of the inter-domain path pointed to by the inter-domain spectrum allocation action is equal to the minimum value of the path probability in the network environment state, determining a first negative reward value as a corresponding first action reward; If the path probability of the inter-domain path pointed to by the inter-domain spectrum allocation action is greater than the minimum value of the path probability in the network environment state, determining the second negative reward value as the corresponding first action reward; Wherein, the first negative reward value is less than the second negative reward value.

5. The spectrum allocation network acquisition method for an optical network according to claim 1, It is characterized in that The step of obtaining the second action reward corresponding to each of the intra-domain spectrum allocation actions according to a preset second action reward rule comprises: If the intra-domain spectrum allocation action is executed smoothly, the second action reward corresponding to the intra-domain spectrum allocation action is a positive reward value; if the intra-domain spectrum allocation action is blocked, the second action reward corresponding to the intra-domain spectrum allocation action is a negative reward value.

6. The method for acquiring spectrum allocation network of an optical network according to claim 5, It is characterized in that The intra-domain environment state includes a plurality of intra-domain paths satisfying the training service request, and a path probability of each of the intra-domain paths; The step of providing a negative reward value for the second action corresponding to the intra-domain spectrum allocation action if the execution of the intra-domain spectrum allocation action is blocked includes: If the path probability of the inter-domain path pointed to by the inter-domain spectrum allocation action corresponding to the intra-domain spectrum allocation action is equal to the maximum value of the path probability in the network environment state, determining the third negative reward value as the corresponding second action reward; If the path probability of the inter-domain path pointed to by the inter-domain spectrum allocation action corresponding to the intra-domain spectrum allocation action is less than the maximum value of the path probability in the network environment state and greater than the minimum value of the path probability in the network environment state, determining the fourth negative reward value as the corresponding second action reward; If the path probability of the inter-domain path pointed to by the inter-domain spectrum allocation action corresponding to the intra-domain spectrum allocation action is equal to the minimum value of the path probability in the network environment state, determining the fifth negative reward value as the corresponding second action reward; The third negative reward value is smaller than the fourth negative reward value, and the fourth negative reward value is smaller than the fifth negative reward value.

7. The method for acquiring spectrum allocation network of an optical network according to claim 1, It is characterized in that The inter-domain spectrum allocation network includes a first spectrum allocation action network for outputting the inter-domain spectrum allocation action, and a first spectrum allocation criticism network for outputting a first value function; The step of performing parameter training on the inter-domain spectrum allocation network according to the network environment status corresponding to each of the training service requests, the inter-domain spectrum allocation action and the first action reward to obtain a target inter-domain spectrum allocation network includes: Acquire a plurality of first value functions output by the first spectrum allocation criticism network according to the network environment status and the inter-domain spectrum allocation action corresponding to each of the training service requests; Obtaining a first advantage function according to a plurality of the first action rewards and a plurality of the first value functions; Performing parameter training on the first spectrum allocation network and the first spectrum allocation criticism network according to the first advantage function to update network parameters of the first spectrum allocation network and network parameters of the first spectrum allocation criticism network; The first spectrum allocation network after updating the network parameters and the first spectrum allocation criticism network after updating the network parameters are determined as the target inter-domain spectrum allocation network.

8. The method for acquiring spectrum allocation network of an optical network according to claim 1, It is characterized in that The intra-domain spectrum allocation network includes a second spectrum allocation action network for outputting the intra-domain spectrum allocation action, and a second spectrum allocation criticism network for outputting a second value function; The step of performing parameter training on the intra-domain spectrum allocation network of each relevant domain according to the intra-domain environmental state of each relevant domain corresponding to each training service request, the intra-domain spectrum allocation action and the second action reward to obtain the target intra-domain spectrum allocation network of each relevant domain comprises: According to the intra-domain environmental state and the intra-domain spectrum allocation action of each relevant domain corresponding to each training service request, acquiring a plurality of second value functions output by the second spectrum allocation criticism network of each relevant domain; Obtaining a second advantage function for each relevant domain according to the plurality of second action rewards and the plurality of second value functions for each relevant domain; Performing parameter training on the second spectrum allocation network and the second spectrum allocation criticism network of each relevant domain according to the second advantage function of each relevant domain, so as to update the network parameters of the second spectrum allocation network and the network parameters of the second spectrum allocation criticism network of each relevant domain; The second spectrum allocation network after updating the network parameters of each related domain and the second spectrum allocation criticism network after updating the network parameters are determined as the target domain intra-domain spectrum allocation network of each related domain.

9. A computer-readable storage medium storing a computer program. Features: When the computer program is executed by a processor, the steps of the method for acquiring a spectrum allocation network for an optical network as claimed in any one of claims 1 to 8 are implemented.

10. A computer device, Features: It comprises a storage, a processor and a computer program stored in the storage and executable by the processor, and when the processor executes the computer program, the steps of the spectrum allocation network acquisition method for an optical network as described in any one of claims 1 to 8 are implemented.

Citation Information

Patent Citations

  • Optical network dynamic spectrum partitioning method and device, storage medium and computer equipment

    CN113965837A

  • Spectrum allocation method and device for elastic optical network, storage medium and equipment

    CN114584871A