Traffic routing method, apparatus and device in segment routing network, and storage medium
By dynamically adjusting routing configuration strategies when network traffic changes, the congestion problem caused by uneven network traffic is solved, achieving balanced traffic distribution and improved resource utilization.
Patent Information
- Application Number
- CN202311419435.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-10-30
- Publication Date
- 2025-10-24
- Estimated Expiration
- 2043-10-30
AI Technical Summary
When network traffic changes, the existing routing configuration strategy is not compatible with the traffic, resulting in network congestion and uneven traffic distribution.
By obtaining the initial traffic matrix and routing configuration strategy, the link utilization of the target traffic matrix is calculated using the routing decision model, the routing configuration strategy with the minimum maximum link utilization is determined, and network traffic routing is dynamically adjusted.
It achieves a balanced distribution of network traffic, avoids network congestion, and improves the utilization rate of network resources.
Smart Images

Figure CN119922111B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application belongs to the technical field of network traffic engineering, and particularly relates to a traffic routing method in a segment routing network, a device, equipment and a storage medium. BACKGROUND
[0002] With the rapid development of the Internet, the Internet has problems such as explosive growth of network traffic. In addition, the continuous development of audio and video services also puts forward the requirement of quality of service for the Internet. Limited by routing algorithms and scheduling strategies, network traffic is prone to uneven distribution on links, resulting in network congestion and degradation of network service quality. Traffic engineering is a technology for optimizing network traffic distribution, which can optimize the scheduling of network traffic, thereby achieving network traffic load balancing, reducing congestion and improving the utilization rate of network resources. Traffic in the network is constantly changing over time, and it is of great significance to consider traffic changes in network traffic engineering.
[0003] However, in the related art, when the network traffic changes, the original routing configuration strategy in the network traffic engineering is still used, or the new routing configuration strategy used is not adapted to the changed network traffic, so that the network traffic is unevenly distributed on the links, resulting in network congestion. SUMMARY
[0004] The embodiments of the present application provide a traffic routing method in a segment routing network, a device, equipment and a storage medium, which can avoid network congestion in the segment routing network and achieve balanced distribution of network traffic.
[0005] In a first aspect, the embodiments of the present application provide a traffic routing method in a segment routing network, the method comprising: obtaining an initial traffic matrix corresponding to a first time period and an initial routing configuration strategy corresponding to the initial traffic matrix, the initial traffic matrix being a traffic matrix formed by traffic demand in the segment routing network in the first time period; in a case where a difference exists between a target traffic matrix corresponding to a second time period and the initial traffic matrix, calculating an initial link utilization rate of the target traffic matrix under the initial routing configuration strategy, wherein the second time period is a next time period corresponding to the first time period; inputting the initial link utilization rate into a routing decision model, and determining a target routing configuration strategy from a plurality of routing configuration strategies based on the initial link utilization rate by the routing decision model, wherein the target routing configuration strategy is a routing configuration strategy with the minimum maximum link utilization rate; and routing the target traffic matrix based on the target routing configuration strategy.
[0006] In a second aspect, an embodiment of the present application provides a traffic routing device in a segment routing network, the device comprising: a policy obtaining module configured to obtain an initial traffic matrix corresponding to a first time period and an initial routing configuration policy corresponding to the initial traffic matrix, the initial traffic matrix being a traffic matrix formed by traffic demand in the segment routing network in the first time period; a link calculating module configured to, in a case where a difference exists between a target traffic matrix corresponding to a second time period and the initial traffic matrix, calculate an initial link utilization of the target traffic matrix under the initial routing configuration policy, wherein the second time period is a next time period corresponding to the first time period; a routing decision module configured to input the initial link utilization into a routing decision model, and determine a target routing configuration policy from a plurality of routing configuration policies based on the initial link utilization by using the routing decision model, wherein the target routing configuration policy is a routing configuration policy with a minimum maximum link utilization; and a traffic routing module configured to route the target traffic matrix based on the target routing configuration policy.
[0007] In a third aspect, an embodiment of the present application provides an electronic device, the electronic device comprising: a processor and a memory storing computer program instructions; and the processor implements the traffic routing method in a segment routing network as described in the first aspect when executing the computer program instructions.
[0008] In a fourth aspect, an embodiment of the present application provides a computer readable storage medium, the computer readable storage medium storing computer program instructions, and the computer program instructions are executed by a processor to implement the traffic routing method in a segment routing network as described in the first aspect.
[0009] In a fifth aspect, an embodiment of the present application provides a computer program product, instructions in the computer program product are executed by a processor of an electronic device to cause the electronic device to perform the traffic routing method in a segment routing network as described in the first aspect.
[0010] From the above, when detecting that the traffic matrix changes, first, the link utilization of the changed traffic matrix under the initial routing configuration policy corresponding to the changed traffic matrix is calculated, and then the routing decision model is used to process the link utilization to determine the routing configuration policy with the minimum maximum link utilization, so as to obtain the target routing configuration policy corresponding to the changed traffic matrix. That is, in the embodiment of the present application, when detecting that the traffic matrix changes, the target routing configuration policy corresponding to the changed traffic matrix can be determined in time, so as to avoid the problem of network congestion caused by the delay of updating the routing configuration policy.
[0011] In addition, in the embodiments of the present application, the maximum link utilization rate can represent the balance of the traffic distribution in the network, and the smaller the maximum link utilization rate is, the more balanced the traffic distribution in the network is after the corresponding routing configuration strategy is executed. In the embodiments of the present application, the routing configuration strategy with the minimum maximum link utilization rate is taken as the routing strategy of the changed traffic matrix, which can avoid the problem of network congestion caused by the inadaptation of the routing configuration strategy to the traffic matrix, and achieve the balanced distribution of network traffic. BRIEF DESCRIPTION OF DRAWINGS
[0012] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the drawings needed to be used in the embodiments of the present application will be briefly introduced as follows, and other drawings can also be obtained by those of ordinary skill in the art without creative labor on the basis of these drawings.
[0013] Figure 1 is a traffic routing schematic diagram corresponding to SR provided by an embodiment of the present application;
[0014] Figure 2 is a flow schematic diagram of a traffic routing method in a segment routing network provided by an embodiment of the present application;
[0015] Figure 3 is a traffic routing schematic diagram provided by an embodiment of the present application;
[0016] Figure 4 is a routing decision schematic diagram provided by an embodiment of the present application;
[0017] Figure 5 is a structural schematic diagram of an action network provided by an embodiment of the present application;
[0018] Figure 6 (a) in is a cumulative probability distribution schematic diagram of a performance ratio obtained by an Abilene method on a small test set of traffic matrices provided by an embodiment of the present application;
[0019] Figure 6 (b) in is a cumulative probability distribution schematic diagram of a performance ratio obtained by a CERNET method on a small test set of traffic matrices provided by an embodiment of the present application;
[0020] Figure 7 (a) in is a cumulative probability distribution schematic diagram of a performance ratio obtained by an Abilene method on a large test set of traffic matrices provided by an embodiment of the present application;
[0021] Figure 7 (b) in is a cumulative probability distribution schematic diagram of a performance ratio obtained by a CERNET method on a large test set of traffic matrices provided by an embodiment of the present application;
[0022] Figure 8 is a structural diagram of a traffic routing device in a segment routing network according to another embodiment of the present application;
[0023] Figure 9 is a structural diagram of an electronic device according to yet another embodiment of the present application. DETAILED DESCRIPTION
[0024] The features and exemplary embodiments of the various aspects of the present application will be described in detail below with reference to the drawings. The following detailed description is merely intended to explain the present application, and is not intended to limit the present application. The present application can be implemented without some of the specific details, which are well known to those skilled in the art. The following description of the embodiments is merely provided to give a better understanding of the present application by showing examples of the present application.
[0025] It should be noted that the terms such as first and second, etc., are merely used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply that there is any such actual relationship or order between these entities or operations. Moreover, the terms "include", "contain" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements, but also includes other elements not explicitly listed or inherent to such a process, method, article or device. Without more limitations, the elements defined by the statement "include" do not exclude the presence of other identical elements in the process, method, article or device including the elements.
[0026] For the sake of understanding, before the scheme provided by the embodiments of the present application is described, the background of the scheme provided by the embodiments of the present application is first described.
[0027] Segment Routing (SR) is a source routing mechanism that can be applied to IP (Internet Protocol) / MPLS (Multi-Protocol Label Switching) or IPv6 (Internet Protocol Version 6) networks. In SR, a segment is associated with a topology or service based instruction and is represented in the form of a Segment Identifier (SID). Segments come in various types, and the segment commonly used in traffic engineering is the node segment, which can be used to indicate that a routing node forwards a data packet to a certain node via the shortest path of ECMP (Equal-Cost Multi-Path Routing), i.e., specifies that the route of the data packet must pass through a certain routing node. When SR routing is used, a Segment List is encapsulated in the packet header of a data packet to indicate the route of the data packet. For example, in the traffic routing diagram corresponding to the SR shown in Figure 1 , the routing network includes 8 routing nodes, A, B, C, D, E, F, G, and H, and the routing node A is the originating node of the data packet and the routing node H is the destination node of the data packet. The routing node A and the routing node H have multiple links therebetween, and the multiple links have the same IGP (Interior Gateway Protocol) link weight. In Figure 1 , the head of the data packet (Packet) sent from the routing node A encapsulates a Segment List [E, H], which indicates that the data packet needs to pass through the routing node E and the routing node H in sequence during the routing process. The data packet is first transmitted along the shortest routing path A-C-E of the routing node A to the routing node E. When the data packet reaches the routing node E, the routing node E in the Segment List is popped. Then, the data packet is routed along the shortest routing paths E-F-H and E-G-H of the routing node E to the routing node H until the destination node H is reached, and the routing node H in the Segment List is popped, so that the routing node H can obtain the complete data packet.
[0028] From the above description, it can be seen that Figure 1 , the routing process uses 2 segments to indicate an intermediate routing node (i.e., the routing node E in Figure 1 ) between the source node and the destination node, and this method can be referred to as 2-SR routing. From Figure 1It can be seen that the SR protocol controls the traffic path through the segment list encapsulated in the data packet header, and the intermediate forwarding node does not need to store the path state information, thereby reducing the routing configuration strategy and update overhead of the SR protocol. Therefore, the SR can be applied to online traffic engineering of the routing configuration strategy dynamically changed with traffic.
[0029] It should be noted that in the embodiments of the present application, the network topology is G=(V, E), where V is a node set, E is a set of unidirectional links, and each routing node v∈V supports the SR protocol. For each link e∈E, the weight w e and the bandwidth (i.e., the link capacity c e ) are known numbers. In the embodiments of the present application, the traffic demand in the network changes over time, and the time is divided into continuous time slots τ=1, 2, …, each of which can represent a small period of time, for example, 5 minutes. In each time slot τ, the traffic demand in the network forms a traffic matrix TM (τ) , which contains |V|(|V|-1) traffic demands between each pair of routing nodes in the network, denotes the traffic demand size between routing node i and routing node j at time slot τ.
[0030] The embodiments of the present application aim to find the optimal routing configuration strategy RC (τ) for the traffic matrix TM (τ) in each time slot τ under the considered network topology and the time-varying network traffic demand, so as to minimize the current maximum link utilization MLU (τ) .
[0031] To achieve the above goal, the embodiments of the present application route the traffic using the 2-SR mode and limit each traffic to use only one SR path for routing, rather than splitting on multiple flows. The SR routing configuration strategy of all traffic in the network can be represented as RC={m i,j |i,j∈V,i≠j}, where the variable m i,j denotes the intermediate routing node through which the traffic from routing node i to routing node j is routed.
[0032] It should be noted that in the embodiments of the present application, any routing node in the network can be selected as an intermediate routing node, that is, m i,j may be any routing node in V. Since m i,j =i and m i,j =j correspond to the same routing path, that is, the traffic from routing node i to routing node j is routed on the IGP shortest routing path between routing node i and routing node j without passing through any intermediate routing node, therefore, in the embodiments of the present application, it is set that m i,j≠i. In addition, variable denotes the intermediate routing node between routing node i and routing node j at time slot τ; denotes the proportion of traffic passing through link e when the traffic is routed on the IGP shortest routing path between routing node i and routing node k, routing node k and routing node j. Wherein, the network topology G and the weight w of all links e Given the time The shortest path algorithm such as Floyd algorithm can be used to calculate, and the SR-TE problem to be solved at each time slot τ can be modeled as an integer linear programming problem as shown below, that is, 2SR-TE-ILP problem:
[0033] min MLU (τ) (1)
[0034]
[0035]
[0036]
[0037] In the above formula (1) to formula (4), is a binary variable, which indicates whether the traffic from routing node i to routing node j is routed through routing node k as an intermediate routing node at time slot τ; if then
[0038] In addition, in the above 2SR-TE-ILP problem, formula (1) represents the optimization goal of 2SR-TE-ILP problem, that is, minimizing the maximum link utilization rate; formula (2) is used to constrain the size of the routed traffic on each link to be less than the bandwidth of the link; formula (3) is used to constrain each traffic demand to select only one 2-SR path when routing; formula (4) is used to constrain the value of variable to be 0 or 1.
[0039] In the case of real-time traffic matrix MLU (τ) Given, the optimal routing configuration strategy of each time slot τ can be obtained by solving problem 2SR-TE-ILP using linear programming solving software such as Gurobi. However, in practical applications, the measurement of real-time traffic matrix is difficult, and the prediction of traffic matrix is also not accurate, but the network link utilization rate is relatively easy to measure.f e (TM, RC) denotes the size of the routed traffic on link e when the traffic matrix TM in the network is routed according to the routing configuration strategy RC, that is:
[0040]
[0041] In formula (5), the link utilization rate u e may be defined as the proportion of the link bandwidth occupied, that is:
[0042]
[0043] In formula (6), U={u e |e∈E} is a set of utilization rates of all links in the network, and the maximum link utilization rate
[0044] As can be seen from formula (5) and formula (6), the real-time link utilization rate information U (τ) is relatively simple to measure, and U (τ) may be calculated according to the traffic transmission rate of the network link bandwidth and the router port. Therefore, the target corresponding to the scheme provided in the embodiments of the present application is to give the best routing configuration strategy RC (τ) under the condition that the traffic matrix TM (τ) is unknown and only the link utilization rate U (τ) is known. In the case where the traffic matrix information is unknown, the relationship between the routing configuration strategy of the traffic and the network link utilization rate is implicit. When the SR routing configuration strategy is taken as the decision variable of the problem, and the link utilization rate is taken as the input and the optimization target of the problem at the same time, a direct mathematical relationship between the two cannot be established by linear programming or other methods. To this end, the embodiments of the present application use deep reinforcement learning to learn the implicit relationship between the input (link utilization rate) and the output (routing configuration strategy), and solve the above problem.
[0045] To solve the above problem, the embodiments of the present application provide a traffic routing method in a segment routing network, a device, equipment and a storage medium. First, the traffic routing method in a segment routing network provided by the embodiments of the present application is introduced.
[0046] Figure 2 The flowchart of the traffic routing method in a segment routing network provided by an embodiment of the present application is shown. As Figure 2 shown, the method comprises the following steps:
[0047] Step S201, obtaining an initial traffic matrix corresponding to a first time period and an initial routing configuration strategy corresponding to the initial traffic matrix, the initial traffic matrix being a traffic matrix formed by traffic demand in the segment routing network in the first time period.
[0048] In step S201, the first time period is the time slot described above, and in the embodiments of the present application, the first time period is represented by τ-1.
[0049] As an example, in Figure 3As shown in the flow routing schematic diagram, 6 routing nodes are deployed in the segmented routing network, the straight line between the routing nodes represents the network link, the curve with arrow represents the network traffic demand and its routing path in the network, and the curve width represents the size of the traffic demand. In the embodiments of the present application, U(TM, RC) represents the set of link utilization of the network when the traffic matrix TM in the network is routed according to the routing configuration policy RC, and MLU(TM, RC) represents the maximum link utilization corresponding to the network when the traffic matrix TM in the network is routed according to the routing configuration policy RC.
[0050] In Figure 3 , in time slot τ-1, the initial traffic matrix in the network is TM (τ-1) , the routing decision model can output the initial routing configuration policy RC (τ-1) corresponding to the traffic matrix TM (τ-1) , so as to obtain the initial traffic matrix and the initial routing configuration policy corresponding to the initial traffic matrix.
[0051] It should be noted that for each time slot τ, the routing decision model can interact with the network, wherein the routing decision model can determine the routing configuration policy corresponding to the current traffic matrix according to the link utilization corresponding to the traffic matrix, and then use the routing configuration policy to route the current traffic matrix to minimize the maximum link utilization. The model structure and training process of the routing decision model are described below, which are not described here.
[0052] In step S202, in the case that the difference exists between the target traffic matrix corresponding to the second time period and the initial traffic matrix, the initial link utilization of the target traffic matrix under the initial routing configuration policy is calculated.
[0053] In step S202, the second time period is the next time period corresponding to the first time period, and in the embodiments of the present application, the second time period is represented by τ. In addition, in the case that the difference exists between the traffic matrices corresponding to the two time periods, it indicates that the traffic matrices corresponding to the two time periods have changed, at this time, the method provided in the embodiments of the present application can be used to route the traffic in the network. In addition, the network manager can also make routing decisions for the traffic in the network at other times according to actual needs, for example, the network manager can use the measurement platform to monitor the network link utilization in real time, and when the link utilization meets certain trigger conditions (for example, the maximum link utilization is greater than a certain threshold), the method provided in the embodiments of the present application is used to update the routing configuration policy in the network.
[0054] As an example, in the routing decision schematic diagram shown in Figure 4 , when the time comes to time slot τ, the traffic matrix in the network has changed, for example, from TM(τ-1) Change to TM (τ) , that is, the traffic matrix after the change is TM (τ) At the initial moment of time slot τ, the routing configuration strategy has not changed, and the routing configuration strategy in the network is still RC (τ-1) , but the routing configuration strategy corresponding to τ-1 may cause network congestion. In this regard, in the embodiment of the present application, in the routing configuration strategy RC (τ-1) Lower Convection MatrixTM (τ) Perform traffic routing and measure the network link utilization (i.e. initial link utilization) U(TM) through the measurement platform (τ) ,RC (τ-1) ).
[0055] It should be noted that the aforementioned measurement platform can quickly calculate the current utilization of each link in the network based on the traffic transmission and reception rate of each router port and the bandwidth of the network link. The measurement platform can be any device capable of measuring link utilization, and this embodiment of the application does not specifically limit this.
[0056] Step S203: input the initial link utilization into the routing decision model, and determine the target routing configuration strategy from multiple routing configuration strategies based on the initial link utilization through the routing decision model, wherein the target routing configuration strategy is the routing configuration strategy with the minimum maximum link utilization.
[0057] In one example, if Figure 4 As shown, after obtaining the routing configuration policy RC (τ-1) Lower Convection MatrixTM (τ) After the initial link utilization is obtained by traffic routing, the initial link utilization U(TM (τ) ,RC (τ-1) ) as the input of the routing decision model, the routing decision model can output the target routing configuration strategy RC (τ) , that is, determine TM (τ) The 2-SR path for each traffic demand in .
[0058] Step S204: routing the target traffic matrix based on the target routing configuration policy.
[0059] It should be noted that the update of the routing configuration strategy in the network can be achieved through the above steps S201 to S203. After the routing configuration strategy in the network is updated, the traffic routing of the target traffic matrix can be performed based on the updated routing configuration strategy. For example, Figure 4 In the network, after the routing configuration strategy is updated, the traffic matrix TM in the time slot τ is (τ) According to the new routing configuration policy RC (τ)When routing in the network, the link utilization changes to U(TM (τ) , RC (τ) ), and the corresponding maximum link utilization is MLU(TM (τ) , RC (τ) ). Meanwhile, the maximum link utilization MLU(TM (τ) , RC (τ) ) is also input to the routing decision model as feedback of the traffic engineering performance.
[0060] In addition, the steps S201 to S204 are repeatedly performed over time, thereby constituting a workflow of the traffic engineering provided in the embodiments. In the above workflow, the traffic engineering solution performs link utilization measurement at the initial moment of each time slot (i.e., at the end of the previous time slot), and uses the measurement as an input of the routing decision model to determine the routing configuration strategy in the current time slot. The network manager can also perform link utilization measurement and run the traffic engineering algorithm at other time according to the requirement.
[0061] Based on the scheme defined in the steps S201 to S204, it can be known that, in the embodiments, when the traffic matrix is detected to change, the link utilization of the changed traffic matrix under the initial routing configuration strategy corresponding to the traffic matrix before the change is calculated first, and then the routing decision model is used to process the link utilization to determine the routing configuration strategy with the minimum maximum link utilization, thereby obtaining the target routing configuration strategy corresponding to the changed traffic matrix. That is, in the embodiments, when the traffic matrix is detected to change, the target routing configuration strategy corresponding to the changed traffic matrix can be determined in time, thereby avoiding the problem of network congestion caused by untimely updating of the routing configuration strategy.
[0062] In addition, in the embodiments, the maximum link utilization can represent the balance of the traffic distribution in the network, and the smaller the maximum link utilization, the more balanced the traffic distribution in the network after the corresponding routing configuration strategy is executed. In the embodiments, the routing configuration strategy with the minimum maximum link utilization is used as the routing strategy of the changed traffic matrix, which can avoid the problem of network congestion caused by inadaptation of the routing configuration strategy to the traffic matrix, and achieve balanced distribution of network traffic.
[0063] The routing decision model used to implement routing decision in the embodiments is introduced below.
[0064] The routing decision model provided in the embodiment of the present application combines action branching technology with the reinforcement learning algorithm PPO (Proximal Policy Optimization) to interact with the network environment and learn how to determine the routing configuration strategy corresponding to each time slot.
[0065] In reinforcement learning, the agent can continuously interact with the environment through state, action and reward. For each step (step) t, the agent observes the state s from the environment t , and according to the state s t Select an action t Action a t is executed in the environment, and the state of the environment is transferred to s t+1 , while the reward value r t is sent to the agent, which is in state s t 、Action a t , reward value r t And the state after transfer s t+1 The formed quaternion (s t ,a t ,r t ,s t+1 ) is called state transition, or simply transition. Reinforcement learning can be modeled as a Markov decision process (MDP) Among them, S is the state set, is a set of actions, is the transition probability matrix, It represents the probability of state transfer to s′ after taking action a in state s; r is the reward function, r(s t ,a t )(or abbreviated as r t ) means in state s t Take action a t The immediate reward obtained; γ∈[0,1] represents the discount coefficient.
[0066] The agent's behavior is based on the policy π θ (·|s), the strategy π θ (·|s) is used to represent the probability distribution of choosing each action in a given state, where θ is the parameter corresponding to the strategy, The agent interacts with the environment over multiple rounds to collect experience, where experience is the state transition data. Each round consists of multiple steps, starting from the initial state and ending at the terminal state, or ending after a fixed number of steps. The goal of reinforcement learning is to learn the optimal strategy through the collected experience, thereby maximizing the expected cumulative discounted reward.
[0067] Proximal Policy Optimization (PPO) is a policy gradient reinforcement learning algorithm that can be applied to environments with continuous or discrete action spaces. PPO uses an action-criteria method to train the policy π. θ In the action-evaluation method, there are two neural networks, namely the action network and the evaluation network. Among them, the action network represents the strategy π θ , which takes state s as input and outputs the probability distribution π on the action θ (·|s); the evaluation network takes state s as input and outputs the state value function V with φ as parameter φ (s), the state value function is used to approximate the state s according to the policy π θ The expected cumulative reward that can be obtained by taking an action; Q value Q(s,a) represents the expected cumulative reward that can be obtained after taking action a in state s; advantage function A θ (s t ,a t ) is used to indicate that in state s t Next take action a t How much better or worse is it than randomly choosing an action? It can be defined as:
[0068]
[0069] Set the PPO algorithm to interact with the environment for multiple rounds and collect a set of trajectories Where each trajectory ρ={s1,a1,r1,…,s T ,a T ,r T The PPO algorithm uses a generalized advantage estimation method to calculate the advantage value corresponding to each state and action in the trajectory using the following formula:
[0070]
[0071] δ t =r t +γV φ (s t+1 )-V φ (s t ) (9)
[0072] In formula (8), λ∈[0,1] is the discount coefficient. The gradient of the objective function J(θ) can be approximated as:
[0073]
[0074] In order to make full use of the collected data, in the PPO algorithm, formula (10) is not directly used as the gradient to update the parameters θ of the action network, but formula (11) is used as the gradient and the same data is used. Train the action network for multiple rounds to update the parameters θ:
[0075]
[0076] In formula (11), π θ′ Indicates data collection The old strategy used when θ Indicates the new strategy being updated. ∈∈(0,1) is used to control the π θ and π θ′ Parameters of the degree of difference between Used to limit the extent of each update, that is, to ensure that each update is not too large, where the function It can be determined by formula (12):
[0077]
[0078] At the beginning of each training, π θ =π θ′ For collecting data π θ′ As the number of training times increases, π θ It is continuously updated during the training process and becomes θ′ After training, the action network of the PPO algorithm is based on the updated strategy π θ Interact with the environment again and collect trajectories; evaluate the parameters of the network by minimizing the state value function V output by the network φ (s) and Usage Data The mean square error of the expected cumulative reward is calculated to update the expected cumulative reward, which can be calculated by formula (13):
[0079]
[0080] In formula (13), L(φ) is the objective function corresponding to the evaluation network, and φ is the target network parameter of the evaluation network; is the state transition data set, ρ is any state transition data set in the state transition data set, T is the number of training times of the action network; r t′is the reward value corresponding to the t′th step in each training; γ t′-t is the discount coefficient of the t′th step relative to the tth step in each training; V φ (s t ) is state s t The corresponding state value function.
[0081] Based on the above-mentioned reinforcement learning training method, the embodiment of the present application also trains the routing decision model according to the state, action and reward in reinforcement learning.
[0082] First, for the state, in the embodiment of the present application, the traffic matrix that needs to be routed at step t is defined as TM (t) , the old routing configuration in the network environment is RC (t-1) The routing decision model is based on the new traffic matrix TM (t) Configure RC based on the old route (t-1) The link utilization U(TM) obtained by routing (t) ,RC (t-1) )={u e (TM (t) ,RC (t-1) )|e∈E} as input. In this embodiment of the application, the elements in the link utilization set U are normalized, and the vector of 1×|E| shown in formula (14) is used as the state:
[0083]
[0084] For actions, the routing decision model takes the routing configuration strategy RC as the action, which can be expressed as a=[a i,j |i,j∈V,i≠j], the action is defined in a discrete space. For problems with discrete action spaces, the dimension of the action network output layer, i.e., the number of neurons, should be equal to the number of all possible actions. Each neuron in the output layer outputs the probability of selecting the corresponding action. In the SR-TE problem, there are L = |V| (|V|-1) traffic demands. Each traffic demand can select one of K = |V|-1 routing nodes as an intermediate routing node for 2-SR routing. The number of all possible routing configuration strategies is L K =[|V|(|V|-1)] |V|-1 , even if the number of nodes in the network |V| is small, L K The corresponding value will also be very large. Therefore, a large number of action spaces is difficult for most reinforcement learning algorithms to handle. Moreover, the action space in the embodiment of the present application is not only discrete but also high-dimensional. The action space has L dimensions, and each dimension a i,j =m i,jis a sub-action, corresponding to the selected intermediate routing node of the traffic from routing node i to routing node j, and there are L possible selections, so there is an explosion of the number of actions using the existing scheme.
[0085] To solve the problem of explosion of the number of actions, the action branching technology is adopted in the embodiments of the present application to solve the reinforcement learning problem with a high-dimensional discrete action space by changing the neural network structure of the action network.
[0086] Specifically, in the embodiments of the present application, the routing decision model at least includes an action network and an evaluation network, the action network is used to determine the probability distribution of the target state data on the routing strategy, and the target state data is used to represent the link utilization rate of the target traffic matrix under the initial routing configuration strategy; the evaluation network is used to determine the expected cumulative reward obtained by routing the target traffic matrix according to the first routing configuration strategy, and the first routing configuration strategy is any one of the plurality of routing configuration strategies.
[0087] The evaluation network at least includes an input layer, a plurality of hidden layers and an output layer; the action network at least includes an input layer, a plurality of hidden layers and a plurality of branch output layers, the plurality of branch output layers are separated from each other, each branch output layer corresponds to a routing configuration strategy, and each branch output layer is used to output the intermediate routing node through which the traffic from the first routing node flows to the second routing node under the corresponding routing configuration strategy.
[0088] As an example, Figure 5 The structure diagram of the action network is shown as follows, Figure 5 As shown in the figure, the input layer and the hidden layer of the changed neural network are the same as those of the ordinary neural network, but the output layer uses not a complete output layer but a plurality of separated branch output layers. In Figure 5 , the input layer with a size of 1x|E| in the action network takes the state s as input. After this input layer, there are a plurality of fully connected hidden layers. Among them, the input layer and the hidden layer jointly constitute the latent representation of the input state s, which is shared by the L branch output layers. Each branch output layer has a size of 1xK, is fully connected with the neurons in the last hidden layer, and outputs the selection probability i,j (·|s) of the K possible intermediate routing nodes of the traffic from routing node i to routing node j, which is used to select the sub-action a i,j , i.e., the selected intermediate routing node of the traffic demand from routing node i to routing node j.
[0089] In the model training process, in order to make the algorithm fully explore the entire action space and learn how to select the best action, the sub-action a i,j is selected according to the probability distribution i,j(·|s) sampling selection. When running online, the routing decision model directly outputs the sub-action with the largest corresponding probability, i.e. After selecting the sub-action a i,j according to the branch output layer, all sub-actions a i,j are combined into a complete action a.
[0090] In this way, the representation of an action is distributed on L branch output layers, and the total number of neurons of the output layers is reduced from LK to LK. K Since all branch output layers share the same latent representation, the selection of sub-actions corresponding to each traffic demand is not completely independent, but is coordinated by the shared neural network input layer and hidden layer, so as to not only make the routing decision model output a more suitable target routing configuration strategy, but also help improve the stability of model training.
[0091] For the evaluation network, since there is no neuron number explosion problem, in the embodiments of the present application, the evaluation network uses a general feedforward neural network structure. The evaluation network also takes the state as input, the input layer size is 1x|E|, there are multiple hidden layers, and the output layer size is 1x1 and outputs the state value function V φ (s t ). Among them, the neurons between adjacent layers are connected in a fully connected manner.
[0092] For the reward, the routing decision model takes MLU(TM (t) ,a t ) as feedback from the network environment, and in the embodiments of the present application, MLU(TM (t) ,a t ) is simply denoted as MLU. In the offline training phase, the routing decision model is trained using the collected historical traffic matrix information, TM (t) is a known traffic matrix, which can be obtained by solving the 2SR-TE-ILP problem (i.e. formulas (1) to (4)), and the optimal maximum link utilization ILP of each traffic matrix under the 2-SR routing mode can be obtained. In this scenario, the reward value can be determined by formula (15):
[0093]
[0094] In formula (15), is used to describe the gap between traffic engineering performance and the optimal solution that can be obtained by the integer linear programming problem.
[0095] As can be seen from formula (15), when calculating the reward value, first subtract 1 from , and then multiply by 2, so that scaled to a suitable range; then an exponential function is used to give a large penalty when traffic engineering performs poorly; finally, the resulting value is added by 1 to limit the possible range of the reward value to the range of (-∞, 0], where the closer the reward value is to 0, the better the performance of the routing decision model.
[0096] It is easy to note that since the routing decision model uses a branching structured action network, the objective function used for neural network parameter update needs to be modified accordingly. According to equation (11), when calculating the objective function J PPO (θ) of action network parameter update, the values of advantage A t (s t ,a θ′ ), probability π t (a t |s θ ) and probability π t (a t |s θ′ ) corresponding to each state s t and action a t in the trajectory need to be known. The value of A θ′ (s t ,a t ) is calculated using the r t value in the trajectory and the output V φ (s t ) of the evaluation network, and is independent of the action network. The values of π θ (a t |s t ) and π θ′ (a t |s t ) are direct outputs of the action network.
[0097] It should be noted that the branching action network used in the embodiments of the present application outputs not the probability π(·|s) of complete action, but the probability π i,j (·|s) of sub-action. Although all branching output layers share the same input layer and hidden layer, the selection of sub-action a i,j only depends on the probability value output by the corresponding branching output layer. Therefore, in the case where the selection of sub-action is an independent event, the probability of complete action can be calculated by equation (16):
[0098]
[0099] In addition, in order to enhance the exploration efficiency of the routing decision model, the embodiment of the present application also adds an entropy reward on the basis of the original objective function, thereby enhancing the diversity of actions output by the neural network. For the branch action network used in the routing decision model, the embodiment of the present application uses the average entropy of each branch output layer as the entropy of the entire action network. The objective function of the action network parameter update of the embodiment of the present application is shown in formula (17):
[0100]
[0101] In formula (17), is the entropy of the branch output layer, and its calculation formula is shown in formula (18):
[0102]
[0103] It should be noted that in formula (17), is the objective function of the action network, θ is the target network parameter of the action network, and θ′ is the network parameter before update; is the state transition data set, ρ is any state transition data set in the state transition data set, T is the number of training times of the action network, a i,j is the intermediate routing node that the traffic demand from routing node i to routing node j passes through; π θ (a i,j |s) is the intermediate routing node a under the target network parameter θ of the action network. i,j Routing configuration strategy in state s; π θ′ (a i,j |s) is the intermediate routing node a under the network parameters θ′ before the update i,j In state a i,j Routing configuration policy under s t is the state corresponding to the tth step in each training; a t is the action used in step t of each training; A θ′ (s t ,a t ) is the network parameter θ′ before updating, under s t In state a t The corresponding advantage function; ∈ is a preset constant; β is a preset coefficient used to balance the proportion of entropy rewards in the objective function; L is the number of branch output layers; The entropy function corresponding to the branch output layer, is the target network parameter θ of the action network, in state s t The probability that the traffic from routing node i to routing node j flows through the intermediate routing nodes is .
[0104] It should be noted that, since the routing decision model provided by the embodiments of the present application does not use a special evaluation network structure, the loss function of the evaluation network is as shown in formula (13) in the embodiments of the present application.
[0105] In addition, in the embodiments of the present application, the training of the routing decision model includes two stages of offline training and network updating. In the offline training stage, the embodiments of the present application use historical traffic matrix data to simulate network traffic changes and train the routing decision model. In the offline training stage, the number of rounds (i.e. the number of training times) M of offline running is determined by the network manager, and the number of steps T in each round is the number of historical traffic matrices. At each step of the running of the reinforcement learning algorithm, one historical traffic matrix is used, and all historical traffic matrix data is used in sequence according to its time sequence in a round. At the tth step, the traffic matrix TM (t) is used in the network environment, and the reinforcement learning agent takes MLU(TM (t) ,a t-1 ) as the state s t and gives the routing configuration a t . In the above mode, the number of steps of each traffic matrix participating in running is the number of rounds of the routing decision model running.
[0106] It should be noted that the training of the above routing decision model usually only needs to run several tens or hundreds of rounds to train a reinforcement learning model with good traffic engineering performance, so that thousands of historical traffic matrix data can be effectively utilized within an acceptable time. In addition, since the complete traffic matrix data set is used for training in each round of the routing decision model running, each historical traffic matrix is fully utilized.
[0107] In addition, in order to shorten the training time and improve the traffic engineering performance of the algorithm, the embodiments of the present application also use parallel environments in the offline training stage. The routing decision model does not interact with a simulated network environment in the offline training stage, but interacts with N independent parallel network environments simultaneously through multi-processes and collects trajectories. At each step, the routing decision model observes N different states from N parallel network environments and selects N corresponding actions according to the strategy, so as to execute in the corresponding network environment, and then receives N corresponding reward values. In each round, there are T steps (i.e. T historical traffic matrix data), and the routing decision model learns (trains and updates the neural network parameters) once every T learn steps, at which time N trajectories of length T learn are collected (i.e. N*T learnIn this way, the routing decision model collects and learns experiences irrelevant to each other from independent network environments (i.e., multiple parallel network environments) at the same time, which helps the convergence of the neural network. In addition, the speed of data collection through interaction with the environment is also accelerated due to the parallelism brought by multi-process.
[0108] Specifically, in the embodiments of the present application, the network parameters of the action network and the network parameters of the evaluation network can be determined by the following steps:
[0109] Step S301, obtaining initial network parameters of the action network and initial network parameters of the evaluation network;
[0110] Step S302, for each training, determining state transition data of each historical traffic matrix in the set of historical traffic matrices in multiple parallel network environments, obtaining multiple sets of state transition data corresponding to each historical traffic matrix, wherein the multiple parallel network environments have the same network topology result, and the multiple parallel network environments correspond to different routing configuration strategies respectively;
[0111] Step S303, storing the multiple sets of state transition data corresponding to each historical traffic matrix into a target buffer to obtain a set of state transition data; wherein the set of state transition data is composed of state transition data, which can include state transition data before policy update, routing configuration strategy before update, reward value, and state transition data after policy update;
[0112] Step S304, updating the initial network parameters of the action network and the initial network parameters of the evaluation network based on the sets of state transition data in the set of state transition data respectively, to obtain target network parameters of the action network and target network parameters of the evaluation network.
[0113] In one example, in the process of determining the set of state transition data, first, the maximum link utilization of the first historical traffic matrix under the last routing configuration strategy in the first parallel network environment is determined to obtain the first state transition data, and the current routing configuration strategy is determined according to the first state transition data; then, the maximum link utilization of the first historical traffic matrix under the current routing configuration strategy in the first parallel network environment is determined to obtain the second state transition data, and the reward value of the first historical traffic matrix under the current routing configuration strategy is calculated; finally, the first state transition data, the current routing configuration strategy, the reward value, and the second state transition data are combined to form the set of state transition data corresponding to the first historical traffic matrix.
[0114] It should be noted that the first parallel network environment described above is any one of the multiple parallel network environments, and the first historical traffic matrix is any one of the set of historical traffic matrices.
[0115] As an example, the offline training of the routing decision model can be implemented by the following pseudo code:
[0116]
[0117]
[0118] It should be noted that the input of the offline training stage is the network topology G=(V, E), the historical traffic matrix data {TM (t) |t=1,…,T} and the maximum link utilization {ILP (t) |t=1,…,T} obtained by solving the 2-SR-ILP problem, the number of parallel network environments N, the number of rounds M and the training interval step T during the algorithm running learn .
[0119] From the pseudo code of the above offline training stage, in the offline training stage, first, the parameters θ of the action network and the parameters φ of the evaluation network are randomly initialized, and the trajectory buffer is emptied Then, M rounds are run, each round has T steps, where T is the number of historical traffic matrices. At each step, the routing decision model interacts with N simulated parallel network environments and collects state transitions. At the t-th step, the traffic matrix of the routing in the network is TM (t) , the routing decision model observes the traffic matrix TM (t) from each parallel network environment, and the maximum link utilization MLU(TM t-1 , a (t) ) under the routing configuration a t-1 at the previous step is calculated, and the state s t is obtained, the action a t is selected according to the state, and then the reward value r t is received. The state transition data set s t , a t , r t , s t+1 ) is stored in the trajectory buffer for training of the model. In the algorithm implementation, the present embodiment uses multi-process parallel running of each simulated parallel network environment and makes the agent interact with it, and the loop form is used in the pseudo code for the sake of simplicity of expression.
[0120] It should be noted that although the traffic matrix TM (t) is used to calculate the state s t and the reward r t at the t-th step in all network environments, the action a t given by the agent is according to the probability distribution πθ (·|s t ) sampling, the action a executed in each parallel network environment t are not the same, therefore, the state s of each parallel network environment t , reward r t and the trajectories experienced are different. When the trajectory buffer The number of transitions in is equal to N*T learn When , update the parameters of the action network and evaluation network, and clear the trajectory buffer after the update
[0121] This completes the offline training of the network model.
[0122] In the network update phase, first, the advantage value corresponding to each state transition data group in the target buffer is calculated, and based on the advantage value, the network parameters when the function value of the objective function corresponding to the action network is maximized are calculated to obtain the target network parameters of the action network; based on the advantage value, the network parameters when the function value of the objective function corresponding to the evaluation network is minimized are calculated to obtain the target network parameters of the evaluation network.
[0123] As an example, the network update of the routing decision model can be implemented by the following pseudo code:
[0124]
[0125] It should be noted that the input of the network update phase is the network parameters θ and φ before the update, and the trajectory buffer Parameters γ, λ, ∈, β used to calculate the objective function and loss function, and the number of training rounds and the size of the mini-batch
[0126] From the pseudo code of the above network update phase, we can see that in the network update phase, we first use The advantage value A corresponding to the action selected in each state is calculated according to the transfer data in formula (8) θ (s t ,a t ), and then update the parameters of the neural network In each round of update, The transfers in are randomly divided into multiple small batches, each of which has Using the state transition data in each small batch of data, the parameters θ of the action network are optimized by maximizing the objective function shown in formula (17) The parameters φ of the evaluation network are updated by minimizing the loss function L(φ) shown in formula (13).
[0127] To verify the effectiveness of the training method provided by the embodiments of the present application, the performance of the online traffic engineering method Adpt-SRTE (i.e., the method provided by the embodiments of the present application) is evaluated through simulation experiments. The process mainly includes three stages: an experimental setup stage, a traffic engineering performance comparison stage, and a time overhead comparison stage.
[0128] In the experimental setup stage, two network topologies are used for experiments: topology Abilene has 12 nodes and 30 links, and topology CERNET has 14 nodes and 32 links. The weights and bandwidths of the network links are provided by the topology data. The experiment uses real traffic matrix data of the two topologies, and the traffic matrix is measured every 5 minutes. To fully evaluate the performance of each method, two traffic matrix data sets, small and large, are used. In the small traffic matrix data set, 10 traffic matrices are used for training, and another 100 traffic matrices are used for testing. In the large traffic matrix data set, 15 days of traffic matrix data are used for training, and the next 15 days of traffic matrix data are used for testing, with 4320 traffic matrices in the training set and the test set.
[0129] The Adpt-SRTE method is implemented using Python language and the PyTorch deep learning framework in the experiment. The action network uses three layers of hidden layers with 512 neurons, and the evaluation network uses two layers of hidden layers with 256 neurons. The parameter update of the neural network uses the Adam optimizer, and the learning rate is set to 0.0003, where γ = 0.99, λ = 0.95, ∈ = 0.2, β = 0.01, the number of parallel network environments N = 16, and the neural network parameter update is performed every T learn = 32 steps, and the number of training rounds is 100. The size of the small batch data set For the small traffic matrix data set, Adpt-SRTE is trained offline for M = 500 rounds, and the number of steps T = 100 in each round. For the large traffic matrix data set, Adpt-SRTE is trained offline for M = 50 rounds, and the number of steps T = 4320 in each round. For the small traffic matrix data set and the large traffic matrix data set, the training set is first used to train Adpt-SRTE offline, and then the traffic matrices in the test set are used to test and evaluate the online traffic engineering performance of Adpt-SRTE.
[0130] In the experiment, Adpt-SRTE is compared with the following traffic engineering methods using reinforcement learning:
[0131] (1) MARL (Multiagent Reinforcement Learning) + GNN (Graph Neural Networks). This method combines multi-agent reinforcement learning and graph neural networks to decide the weights of network links, and performs online traffic engineering optimization in traditional IP networks using IGP protocol routing.
[0132] (2) DATE. This method uses multi-agent reinforcement learning to decide the split ratio of network traffic on pre-determined paths, and performs online traffic engineering optimization in networks based on explicit path routing.
[0133] Since the offline training of the above two methods is very time-consuming, for small traffic matrix data sets, the experiment uses all 100 traffic matrices in the training set to train these two methods; for large traffic matrix data sets, 100 or 200 traffic matrices are selected from the 4320 traffic matrices in the training set for training. For MARL + GNN, the experiment selects the last 100 or 200 traffic matrices in the 4320 traffic matrices in the training set in chronological order; for DATE, the experiment first uses a 15th degree polynomial function to fit the size of each traffic demand in all traffic matrices in the training set, and obtains the fitted traffic matrix data to eliminate outliers, and then extracts 100 or 200 traffic matrices from all 4320 fitted traffic matrices in uniform intervals. After training is completed, the experiment uses all 100 or 4320 traffic matrices in the test set to evaluate the traffic engineering performance of the method.
[0134] In addition, the experiment also compares the following non-machine learning methods:
[0135] (1) 2-SR-ILP, for each traffic matrix in the test set, solve the 2-SR-ILP problem and obtain the theoretically lowest maximum link utilization of the traffic matrix in the 2-SR routing mode.
[0136] (2) SPF (Shortest Path First), for each traffic matrix in the test set, the traffic demand follows the IGP shortest path and obtains the corresponding maximum link utilization.
[0137] The performance ratio is used to evaluate the traffic engineering performance of each method. For a method, the performance ratio under a certain traffic matrix refers to the ratio of the maximum link utilization obtained by the method to the maximum link utilization obtained by solving the MCF (network service based technology) problem (i.e. the theoretically optimal maximum link utilization), and the closer the performance ratio is to 1, the better the traffic engineering performance of the method.
[0138] For comparison of traffic engineering performance, the experiment evaluates the traffic engineering performance on both large and small traffic matrix datasets.
[0139] First, the experiment evaluates the traffic engineering performance of each method on small traffic matrix datasets. The experiment trains the methods using reinforcement learning on traffic matrices in the small training set, then tests and records the maximum link utilization achieved by all methods on traffic matrices in the small test set, and calculates the performance ratio. Figure 6 Figure 6 shows the cumulative probability distribution of the performance ratio achieved by all methods on traffic matrices in the small test set, where, Figure 6 Figure 6(a) shows the cumulative probability distribution of the performance ratio achieved by the Abilene method on traffic matrices in the small test set. Figure 6 Figure 6(b) shows the cumulative probability distribution of the performance ratio achieved by the CERNET method on traffic matrices in the small test set.
[0140] As can be seen from Figure 6 In topologies Abilene and CERNET, the 2-SR-ILP method achieves the same maximum link utilization as the theoretical optimal solution of the MCF problem for more than 80% of the traffic matrices (i.e., the performance ratio is 1), which indicates that the 2-SR routing mode can provide sufficient traffic engineering capability in topologies Abilene and CERNET.
[0141] It should be noted that 2-SR-ILP calculates the corresponding routing configuration under the premise of knowing the traffic matrix, while Adpt-SRTE first trains on traffic matrices in the training set, and then outputs the routing configuration for traffic matrices in the test set only knowing the link utilization. Therefore, Figure 6 The results in Figure 6 reflect that the maximum link utilization achieved by Adpt-SRTE is not as good as that achieved by 2-SR-ILP. As can be seen from the experimental results, the traffic engineering performance of Adpt-SRTE on small traffic matrix datasets is satisfactory.
[0142] For the three online traffic engineering methods using reinforcement learning, on topology Abilene, the traffic engineering performance of Adpt-SRTE, MARL+GNN and DATE is similar, and their performance ratios are better than that of the shortest route SPF, where DATE performs poorly in traffic engineering when the performance ratio is large. On topology CERNET, Adpt-SRTE performs well, but MARL+GNN and DATE perform poorly in traffic engineering. Similarly, the performance ratios of the three methods are better than that of the shortest route SPF, except that MARL+GNN performs slightly worse in traffic engineering when the performance ratio is large.
[0143] For the performance evaluation of each method on the large traffic matrix dataset, the experiment was also conducted according to the aforementioned experimental setup, the methods using reinforcement learning were trained using the traffic matrices in the large training set, and then the maximum link utilization obtained by all methods on the traffic matrices in the large test set was tested and recorded, and the performance ratio was calculated. Among them, Figure 7 Fig. 8 shows the cumulative probability distribution of the performance ratio obtained by all methods on the traffic matrices in the large test set, Figure 7 Fig. 8(a) is the cumulative probability distribution of the performance ratio obtained by the Abilene method on the traffic matrices in the large test set, Figure 7 Fig. 8(b) is the cumulative probability distribution of the performance ratio obtained by the CERNET method on the traffic matrices in the large test set.
[0144] It can be seen from Figure 7 that the performance ratio obtained by the 2-SR-ILP method on more than 80% of the traffic matrices in the topologies Abilene and CERNET is 1, further verifying that the 2-SR routing mode can provide sufficient traffic engineering capability in these two topologies. On the large traffic matrix dataset, Adpt-SRTE also has good traffic engineering performance, and its performance ratio is much better than that of the shortest route SPF, and the gap between the traffic engineering performance of the 2-SR-ILP method is within a reasonable range.
[0145] As for MARL+GNN and DATE, due to the time-consuming offline training time of these two methods, this experiment only selects 100 or 200 traffic matrices from the total 4320 traffic matrices in the large traffic matrix training set for offline training, Figure 7 the number of traffic matrices used for training corresponding to the curve is marked in the legend of Fig. 9. On the large traffic matrix dataset, the traffic engineering performance of these two methods is obviously worse than that of Adpt-SRTE, and the performance of these two methods on different network topologies is inconsistent. In the topology Abilene, the performance of MARL+GNN is not good, and the performance ratio of DATE is worse than that of the shortest route SPF. But in the topology CERNET, the performance of DATE is not good, but the performance of MARL+GNN is much worse than that of the shortest route SPF. In addition, the number of traffic matrices used in the offline training stage has no obvious effect on the traffic engineering performance of these two methods, as can be seen from Figure 7 Fig. 9, there is no obvious difference in the online traffic engineering performance of these two methods after offline training using 100 and 200 traffic matrices. Compared with MARL+GNN and DATE, Adpt-SRTE also has good traffic engineering performance on the large traffic matrix dataset, and in the case of a large amount of historical traffic matrix data, Adpt-SRTE can be applied to long-term online traffic engineering optimization.
[0146] For comparison of time overhead, the experiment evaluates the time overhead on large and small traffic matrix datasets, as shown in Table 1 and Table 2, where Table 1 is the training time of each method on the small traffic matrix training set, and Table 2 is the training time of each method on the large traffic matrix training set.
[0147] Table 1
[0148]
[0149] Table 2
[0150]
[0151] As shown in Table 1 and Table 2, on the small dataset, the offline training time of the Adpt-SRTE method is much less than the other two methods, where the training time of Adpt-SRTE is less than one hour, while the offline training of MARL+GNN takes several hours, and the offline training time of DATE is more than one day. However, the traffic engineering performance of Adpt-SRTE on the small traffic matrix test set is better than the other two methods. On the large dataset, Adpt-SRTE uses 4320 traffic matrices for offline training, while MARL+GNN and DATE only use 100 or 200 traffic matrices selected from them. However, the offline training time of Adpt-SRTE is still relatively short, where the offline training time of Adpt-SRTE is about 4 hours, and the training time of MARL+GNN is 3 to 4 hours if 100 traffic matrices are used. In addition, the training time of Adpt-SRTE is longer than that of MARL+GNN and DATE. The traffic engineering performance of Adpt-SRTE on the large traffic matrix test set is better than the other two methods. Increasing the number of traffic matrices used for offline training of MARL+GNN and DATE can improve the traffic engineering performance of the two methods on the test set, but the huge training time overhead that comes with it is unbearable.
[0152] Table 3 gives the average value of the calculation time required by each method to make online routing decisions for each traffic matrix in the test set. As shown in Table 3, the time of all methods for online routing decision is about tens of milliseconds.
[0153] Table 3
[0154]
[0155] Based on the above, the experimental results of the experiment show that Adpt-SRTE can fully utilize a large amount of historical traffic matrix data for long-term online traffic engineering optimization after a short offline training time, and the online decision time is short, which meets the requirements of online traffic engineering.
[0156] The embodiment of the present application also provides a traffic routing device in a segmented routing network, such as Figure 8 As shown, the device 800 includes: a policy acquisition module 801, a link calculation module 802, a routing decision module 803 and a traffic routing module 804.
[0157] A policy acquisition module 801 is configured to acquire an initial traffic matrix corresponding to a first time period and an initial routing configuration policy corresponding to the initial traffic matrix, wherein the initial traffic matrix is a traffic matrix formed by traffic demands in the segment routing network during the first time period;
[0158] a link calculation module 802 configured to calculate an initial link utilization of the target traffic matrix under the initial routing configuration policy when a difference is detected between the target traffic matrix corresponding to a second time period and the initial traffic matrix, wherein the second time period is a time period next to the first time period;
[0159] A routing decision module 803 is configured to input the initial link utilization into a routing decision model, and determine a target routing configuration strategy from among multiple routing configuration strategies based on the initial link utilization using the routing decision model, wherein the target routing configuration strategy is the routing configuration strategy with the lowest maximum link utilization;
[0160] The traffic routing module 804 is configured to route the target traffic matrix based on the target routing configuration policy.
[0161] In one example, a routing decision model includes at least an action network and an evaluation network, wherein the action network is used to determine the probability distribution of target state data on the routing strategy, and the target state data is used to characterize the link utilization of the target traffic matrix under the initial routing configuration strategy; the evaluation network is used to determine the expected cumulative reward obtained by routing the target traffic matrix according to a first routing configuration strategy, and the first routing configuration strategy is any one of a plurality of routing configuration strategies; the evaluation network includes at least an input layer, a plurality of hidden layers, and an output layer; the action network includes at least an input layer, a plurality of hidden layers, and a plurality of branch output layers, wherein the plurality of branch output layers are separated from each other, each branch output layer corresponds to a routing configuration strategy, and each branch output layer is used to output the intermediate routing nodes through which the traffic flowing from the first routing node to the second routing node flows under the corresponding routing configuration strategy.
[0162] In one example, the traffic routing device in the segment routing network comprises a model training module configured to determine the network parameters of the action network and the network parameters of the evaluation network by: obtaining initial network parameters of the action network and initial network parameters of the evaluation network; determining, for each training, state transition data of each historical traffic matrix in the set of historical traffic matrices in a plurality of parallel network environments, to obtain a plurality of state transition data groups corresponding to each historical traffic matrix, wherein the plurality of parallel network environments have the same network topology result, and the plurality of parallel network environments correspond to different routing configuration strategies respectively; storing the plurality of state transition data groups corresponding to each historical traffic matrix in a target buffer to obtain a set of state transition data; and updating the initial network parameters of the action network and the initial network parameters of the evaluation network based on the state transition data groups in the set of state transition data respectively, to obtain target network parameters of the action network and target network parameters of the evaluation network.
[0163] In one example, the model training module is specifically configured to determine a maximum link utilization rate of a first historical traffic matrix under a previous routing configuration strategy in a first parallel network environment, to obtain first state transition data, wherein the first parallel network environment is any one of the plurality of parallel network environments, and the first historical traffic matrix is any one of the set of historical traffic matrices; determine a current routing configuration strategy according to the first state transition data; determine a maximum link utilization rate of the first historical traffic matrix under the current routing configuration strategy in the first parallel network environment, to obtain second state transition data; calculate a reward value of the first historical traffic matrix under the current routing configuration strategy; and compose a state transition data group corresponding to the first historical traffic matrix based on the first state transition data, the current routing configuration strategy, the reward value, and the second state transition data.
[0164] In one example, the model training module is specifically configured to calculate an advantage value corresponding to each state transition data group in the target buffer; calculate the network parameters of the action network corresponding to the maximum function value of the target function of the action network based on the advantage value, to obtain the target network parameters of the action network; and calculate the network parameters of the evaluation network corresponding to the minimum function value of the target function of the evaluation network based on the advantage value, to obtain the target network parameters of the evaluation network.
[0165] In one example, the model training module determines the target network parameters of the action network by the following function:
[0166]
[0167] wherein, is the target function of the action network, θ is the target network parameters of the action network, and θ' is the network parameters before updating; is a state transition data set, p is any one of the state transition data sets in the state transition data set, T is the number of training times of the action network, a i,j is an intermediate routing node through which the traffic demand from routing node i flows to routing node j; p θ is a routing configuration policy in state s; p i,j is an intermediate routing node a i,j under the target network parameter θ of the action network; p θ′ is a routing configuration policy in state s; p i,j is an intermediate routing node a i,j under the network parameter θ' before updating; s i,j is a routing configuration policy in state s; s t is the state corresponding to the t-th step in each training; a t is the action adopted in the t-th step in each training; A θ′ (s t ,a t ) is the advantage function corresponding to action a t in state s t under the network parameter θ' before updating; is a preset constant; β is a preset coefficient; L is the number of branch output layers; is an entropy function corresponding to the branch output layer, is the probability of the traffic flowing from routing node i to routing node j through the intermediate routing node in state s t under the target network parameter θ of the action network.
[0168] In an example, the model training module determines the target network parameter of the evaluation network through the following function:
[0169]
[0170] wherein L (φ) is a target function corresponding to the evaluation network, φ is the target network parameter of the evaluation network; D is a state transition data set, p is any one of the state transition data sets in the state transition data set, T is the number of training times of the action network; r t′ is a reward value corresponding to the t'-th step in each training; γ t′-t is a discount coefficient of the t'-th step relative to the t-th step in each training; V φ (s t ) is a state value function corresponding to state s t .
[0171] The traffic routing device in the segment routing network provided by the embodiments of the present application can implement each process implemented by the foregoing method embodiments, and thus details are not repeated here.
[0172] Those skilled in the art can clearly understand that, for the convenience and brevity of description, only the division of the above functional units and modules is taken as an example, and in actual application, the above functions can be completed by different functional units and modules according to needs, that is, the internal structure of the device is divided into different functional units or modules to complete all or part of the functions described above. Each functional unit and module in the embodiment can be integrated in one processing unit, or each unit can exist physically, or two or more units can be integrated in one unit. The integrated unit can be realized in the form of hardware or in the form of a software functional unit. In addition, the specific name of each functional unit and module is only for convenient distinction, and does not limit the protection scope of the present application. The specific working process of the unit and module in the system can refer to the corresponding process in the foregoing method embodiments, and will not be repeated here.
[0173] Figure 9 A hardware structure schematic diagram of an electronic device provided in the embodiment of the present application is shown.
[0174] The electronic device can include a processor 901 and a memory 902 having computer program instructions stored therein.
[0175] Specifically, the processor 901 can include a central processing unit (CPU), or an application specific integrated circuit (ASIC), or can be configured as one or more integrated circuits that implement one or more embodiments of the present application.
[0176] The memory 902 can include a mass storage for data or instructions. By way of example and not limitation, the memory 902 can include a hard disk drive (HDD), a floppy disk drive, a flash memory, an optical disk, a magneto-optical disk, a magnetic tape, or a universal serial bus (USB) drive or a combination of two or more of these. Where appropriate, the memory 902 can include removable or non-removable (or fixed) media. Where appropriate, the memory 902 can be internal or external to the integrated gateway disaster recovery device. In a particular embodiment, the memory 902 is a non-volatile solid-state memory.
[0177] The memory can include read-only memory (ROM), random-access memory (RAM), magnetic disk storage mediums, optical storage mediums, flash memory devices, electrical, optical, or other physical / tangible memory storage devices. Accordingly, generally, the memory includes one or more tangible (non-transitory) computer-readable storage media (e.g., a memory device) encoded with software that, when executed (by one or more processors), is operable to perform the operations described with reference to the methods according to an aspect of the present disclosure.
[0178] The processor 901 implements the traffic routing method in any of the above-described embodiments of the segment routing network by reading and executing computer program instructions stored in the memory 902.
[0179] In one example, the electronic device can further include a communication interface 903 and a bus 910. As shown, the processor 901, the memory 902, and the communication interface 903 are connected through the bus 910 and complete communication with each other. Figure 9
[0180] The communication interface 903 is mainly used to realize the communication between the modules, devices, units, and / or equipment in the embodiments of the present application.
[0181] The bus 910 includes hardware, software, or both, that couples components of the electronic device to each other. By way of example, and not limitation, the bus can include an Accelerated Graphics Port (AGP) or other graphics bus, an Enhanced Industry Standard Architecture (EISA) bus, a Front Side Bus (FSB), a HyperTransport (HT) interconnect, an Industry Standard Architecture (ISA) bus, an InfiniBand (IB) interconnect, a Low Pin Count (LPC) bus, a memory bus, a Micro Channel Architecture (MCA) bus, a Peripheral Component Interconnect (PCI) bus, a PCI-Express (PCI-X) bus, a Serial Advanced Technology Attachment (SATA) bus, a Video Electronics Standards Association Local (VLB) bus, or another suitable bus or combination of two or more of these. Where appropriate, the bus 910 can include one or more buses. Although the present application describes and illustrates a particular bus, the present application contemplates any suitable bus or interconnect.
[0182] In addition, in combination with the traffic routing method in the segment routing network in the above-described embodiments, the embodiments of the present application can provide a computer-readable storage medium to implement. The computer-readable storage medium has computer program instructions stored thereon; the computer program instructions are executed by the processor to implement any of the traffic routing methods in the segment routing network in the above-described embodiments.
[0183] In addition, in combination with the traffic routing method in the segment routing network in the above embodiments, an embodiment of the present application can provide a computer program product for implementation. Instructions in the computer program product are executed by a processor of an electronic device, so that the electronic device executes the traffic routing method in the segment routing network as described in any of the above embodiments.
[0184] It should be noted that the present application is not limited to the specific configurations and processes described above and shown in the drawings. For the sake of brevity, detailed descriptions of well-known methods are omitted. In the above embodiments, several specific steps are described and shown as examples. However, the method processes of the present application are not limited to the specific steps described and shown, and those skilled in the art can make various changes, modifications and additions, or change the order of the steps, after understanding the spirit of the present application.
[0185] The functional modules shown in the structural block diagram described above can be implemented as hardware, software, firmware or a combination thereof. When implemented in hardware, it can be, for example, an electronic circuit, an application specific integrated circuit (ASIC), appropriate firmware, a plug-in, a functional card, etc. When implemented in software, the elements of the present application are program or code segments used to perform the required tasks. The program or code segments can be stored in a machine-readable medium or transmitted through a data signal carried in a carrier wave over a transmission medium or communication link. The "machine-readable medium" can include any medium capable of storing or transmitting information. Examples of machine-readable media include electronic circuits, semiconductor memory devices, ROM, flash memory, erasable ROM (EROM), floppy disks, CD-ROMs, optical disks, hard disks, optical fiber media, radio frequency (RF) links, etc. The code segments can be downloaded via a computer network such as the Internet, an intranet, etc.
[0186] It should also be noted that the exemplary embodiments mentioned in the present application describe some methods or systems based on a series of steps or devices. However, the present application is not limited to the order of the above steps, that is, the steps can be executed in the order mentioned in the embodiments, or in an order different from the embodiments, or several steps can be executed simultaneously.
[0187] The above described aspects of the present disclosure are described with reference to flowcharts and / or block diagrams of methods, apparatuses, devices, and storage media for traffic routing in a segment routing network according to embodiments of the present disclosure. It will be understood that each block of the flowchart and / or block diagrams, and combinations of blocks in the flowchart and / or block diagrams, can be implemented by computer program instructions. Such computer program instructions can be provided to a processor of a general purpose computer, special purpose computer, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, implement the functions / acts specified in the flowchart and / or block diagram block or blocks. Such a processor can be, but is not limited to, a general purpose processor, a special purpose processor, a special purpose application specific processor, or a field programmable logic array. It will be further understood that each block of the flowchart and / or block diagrams, and combinations thereof, can be implemented by special purpose hardware-based computer systems which perform the specified functions or acts, or combinations of special purpose hardware and computer instructions. The embodiments described herein can be implemented in hardware, software, firmware, or any combination thereof.
[0188] The above description merely describes specific implementation of the present application. The skilled in the art can clearly understand that, for the convenience and brevity, the specific working process of the system, module and unit described above can refer to the corresponding process in the foregoing method embodiments, which will not be described herein. It should be understood that the protection scope of the present application is not limited to this. Any skilled person in the art can easily think of various equivalent modifications or replacements within the technical range disclosed in the present application, and these modifications or replacements should be covered within the protection scope of the present application.
Claims
1. A method of traffic routing in a segment routing network, the method comprising: The method comprises: obtaining an initial traffic matrix corresponding to a first time period and an initial routing configuration strategy corresponding to the initial traffic matrix, the initial traffic matrix being a traffic matrix formed by traffic demand in a segment routing network within the first time period; in a case where a target traffic matrix corresponding to a second time period is detected to be different from the initial traffic matrix, calculating an initial link utilization rate of the target traffic matrix under the initial routing configuration strategy, wherein the second time period is a next time period corresponding to the first time period; inputting the initial link utilization rate into a routing decision model, and determining a target routing configuration strategy from a plurality of routing configuration strategies based on the initial link utilization rate by the routing decision model, wherein the target routing configuration strategy is a routing configuration strategy with the minimum maximum link utilization rate; routing the target traffic matrix based on the target routing configuration strategy; the routing decision model comprises at least an action network and an evaluation network, the action network is configured to determine a probability distribution of target state data on a routing strategy, the target state data is configured to represent a link utilization rate of the target traffic matrix under the initial routing configuration strategy; and the evaluation network is configured to determine an expected cumulative reward obtained by routing the target traffic matrix according to a first routing configuration strategy, the first routing configuration strategy being any one of the plurality of routing configuration strategies; the evaluation network comprises at least an input layer, a plurality of hidden layers, and an output layer; the action network comprises at least an input layer, a plurality of hidden layers, and a plurality of branch output layers, wherein the plurality of branch output layers are separated from each other, each branch output layer corresponds to a routing configuration strategy, and each branch output layer is configured to output an intermediate routing node through which traffic flows from a first routing node to a second routing node under the corresponding routing configuration strategy.
2. The method of claim 1, wherein, The network parameters of the action network and the network parameters of the evaluation network are determined in the following manner: obtaining initial network parameters of the action network and initial network parameters of the evaluation network; for each training, determining state transition data of each historical traffic matrix in a historical traffic matrix set in a plurality of parallel network environments, to obtain a plurality of state transition data groups corresponding to the each historical traffic matrix, wherein the plurality of parallel network environments have the same network topology, and the plurality of parallel network environments correspond to different routing configuration strategies respectively; storing the plurality of state transition data groups corresponding to the each historical traffic matrix in a target buffer to obtain a state transition data set; updating the initial network parameters of the action network and the initial network parameters of the evaluation network based on the state transition data groups in the state transition data set respectively, to obtain target network parameters of the action network and target network parameters of the evaluation network.
3. The method of claim 2, wherein, determining state transition data of each historical traffic matrix in a historical traffic matrix set in a plurality of parallel network environments, to obtain a plurality of state transition data groups corresponding to the each historical traffic matrix, comprises: Determine the maximum link utilization of the first historical traffic matrix under the last routing configuration strategy in the first parallel network environment, to obtain first state transition data, wherein the first parallel network environment is any one of the plurality of parallel network environments, and the first historical traffic matrix is any one of the historical traffic matrix set; Determine the current routing configuration strategy according to the first state transition data; Determine the maximum link utilization of the first historical traffic matrix under the current routing configuration strategy in the first parallel network environment, to obtain second state transition data; Calculate the reward value of the first historical traffic matrix under the current routing configuration strategy; Based on the first state transition data, the current routing configuration strategy, the reward value, and the second state transition data, form a state transition data group corresponding to the first historical traffic matrix.
4. The method of claim 2, wherein, Based on the state transition data group in the target buffer, update the initial network parameters of the action network and the initial network parameters of the evaluation network respectively, to obtain the target network parameters of the action network and the target network parameters of the evaluation network, including: Calculate the advantage value corresponding to each state transition data group in the target buffer; Based on the advantage value, calculate the network parameters when the function value of the target function corresponding to the action network is maximum, to obtain the target network parameters of the action network; Based on the advantage value, calculate the network parameters when the function value of the target function corresponding to the evaluation network is minimum, to obtain the target network parameters of the evaluation network.
5. The method of claim 4, wherein, Based on the advantage value, calculate the network parameters when the function value of the target function corresponding to the action network is maximum, to obtain the target network parameters of the action network, including: Determine the target network parameters of the action network through the following function: wherein, is a target function of the action network, θ is a target network parameter of the action network, θ ′ is a network parameter before update; is a set of state transition data, ρ is any one of the set of state transition data, T is a training number of the action network, a i,j is an intermediate routing node through which traffic flows from routing node i to routing node j; π θ (a i,j |s) is a routing configuration policy of the intermediate routing node a i,j in state s under the target network parameter θ of the action network; s θ′ (a i,j |s) is a routing configuration policy of the intermediate routing node a i,j in state s under the network parameter θ' before update; s t is a state corresponding to the t-th step in each training; a t is an action adopted in the t-th step in each training; A θ′ (s t ,a t ) is an advantage function corresponding to the action a t adopted in state s t under the network parameter θ' before update; is a preset constant; β is a preset coefficient; L is a number of branch output layers; is an entropy function corresponding to the branch output layer, is a probability that traffic flows through the intermediate routing node from routing node i to routing node j in state s t under the target network parameter θ of the action network.
6. The method of claim 4, wherein, Based on the advantage value, calculate the network parameters when the function value of the target function corresponding to the evaluation network is minimum, to obtain the target network parameters of the evaluation network, including: Determine the target network parameters of the evaluation network through the following function: Wherein, L(φ) is a target function corresponding to the evaluation network, φ is a target network parameter of the evaluation network; is the state transition data set, ρ is any one state transition data group in the state transition data set, T is the number of training of the action network; r t is a reward value corresponding to the t ′ step in each training; γ t′-t is a discount coefficient of the t ′ step relative to the t step; V φ (s t ) is a state value function corresponding to the state s t .
7. A traffic routing apparatus in a segment routing network, characterized in that, Including: The policy acquisition module is used to acquire an initial traffic matrix corresponding to a first time period and an initial routing configuration strategy corresponding to the initial traffic matrix, wherein the initial traffic matrix is a traffic matrix formed by traffic demand in the segment routing network in the first time period; The link calculation module is used to calculate an initial link utilization of a target traffic matrix under the initial routing configuration strategy when it is detected that there is a difference between the target traffic matrix corresponding to a second time period and the initial traffic matrix, wherein the second time period is a next time period corresponding to the first time period; The routing decision module is used to input the initial link utilization into a routing decision model, and determine a target routing configuration strategy from a plurality of routing configuration strategies based on the initial link utilization through the routing decision model, wherein the target routing configuration strategy is a routing configuration strategy with the minimum maximum link utilization; The traffic routing module is used to route the target traffic matrix based on the target routing configuration strategy. The routing decision model comprises at least an action network and an evaluation network, the action network is used to determine a probability distribution of target state data on a routing strategy, the target state data is used to represent link utilization of the target traffic matrix under the initial routing configuration strategy; the evaluation network is used to determine an expected cumulative reward obtained by routing the target traffic matrix according to a first routing configuration strategy, the first routing configuration strategy is any one of the plurality of routing configuration strategies; The evaluation network comprises at least an input layer, a plurality of hidden layers and an output layer; The action network comprises at least an input layer, a plurality of hidden layers and a plurality of branch output layers, wherein the plurality of branch output layers are separated from each other, each branch output layer corresponds to a routing configuration strategy, and each branch output layer is used to output an intermediate routing node through which traffic flows from a first routing node to a second routing node under the corresponding routing configuration strategy.
8. An electronic device, comprising: The electronic device comprises a processor and a memory storing computer program instructions; The processor executes the computer program instructions to implement the traffic routing method in the segment routing network according to any one of claims 1-6.
9. A computer-readable storage medium, characterized in that, The computer program instructions stored on the computer readable storage medium are executed by the processor to implement the traffic routing method in the segment routing network according to any one of claims 1-6.
10. A computer program product, characterised in that, The instructions in the computer program product are executed by the processor of the electronic device to enable the electronic device to perform the traffic routing method in the segment routing network according to any one of claims 1-6.