An Online Routing Method and System Based on Deep Reinforcement Learning

By combining the routing generation model of the messaging neural network and the soft actor-critician algorithm, dynamically adjusting the network index weights, the problem of insufficient flexibility in the dynamic network environment is solved, and the effect of quickly adapting to network changes and improving transmission performance is achieved.

CN119743420BActive Publication Date: 2025-07-18NANJING UNIV OF INFORMATION SCI & TECH
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202510228649.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-02-28
Publication Date
2025-07-18
Estimated Expiration
2045-02-28

AI Technical Summary

Technical Problem

The existing routing methods based on deep reinforcement learning cannot effectively reflect the importance differences of different metrics in dynamic network environments, resulting in lack of flexibility in routing strategies and inability to respond to network performance fluctuations caused by traffic fluctuations and network equipment adjustments in a timely manner.

Method used

Using a routing generation model based on messaging neural network (MPNN) prediction performance indicators and combining with the soft actor-criticist (SAC) algorithm, the weights of end-to-end delay, bandwidth and packet loss rate are dynamically adjusted, and the optimal path is selected through the intelligent optimization module, a flow table is generated and sent to the switch device.

Benefits of technology

In complex networks, it can quickly adapt to sudden service surges or link failures, improve transmission performance, significantly improve network throughput and reduce end-to-end delay, and achieve a more stable routing solution.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119743420B_ABST
    Figure CN119743420B_ABST
Patent Text Reader

Abstract

The present invention discloses an online routing method and system based on deep reinforcement learning, specifically as follows: 1. Calculate K paths between the source node and the destination node; 2. Predict the prediction performance index of the optimal path p obtained in the (n-1)-th period; 3. Calculate the actual performance index of the optimal path y obtained in the (n-2)-th period; 4. Calculate the relative difference between the prediction performance index in step 2 and the actual performance index in step 3, and update the weights of the end-to-end delay, the end-to-end remaining bandwidth, and the end-to-end packet loss rate according to the relative difference; 5. Update the reward function of the n-th period based on the weights obtained in step 4. Based on the reward function of the n-th period, use the SAC algorithm to calculate the optimal path of the n-th period, and then go to step 2. The present invention can adapt to the changes of the network environment faster and obtain a stable and optimal routing scheme.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of traffic control, and particularly relates to an online routing method and system based on deep reinforcement learning. Background Art

[0002] In recent years, there have been heterogeneous mixed traffic such as voice, telegrams, images, and videos in communication networks. Different services have different requirements for service quality, making it difficult for heterogeneous services to match the optimal path. On the other hand, the strong burstiness of network services may cause a sudden increase in the load of the selected link, unable to meet the traffic demand in time, thereby affecting the network stability and data transmission rate. Moreover, in the communication environment, adverse factors such as electromagnetic interference and complex terrain may cause link failures and even communication interruptions. However, the optimization objectives of most current routing methods based on Deep Reinforcement Learning (DRL) are static, while the network environment is dynamic. Traffic fluctuations or readjustments of network devices often lead to periodic fluctuations in network performance, such as throughput degradation or latency peaks. Existing technologies cannot effectively reflect the importance differences of different metrics in a dynamic network environment, resulting in a lack of sufficient flexibility in routing strategies. Summary of the Invention

[0003] To solve the problems existing in the above-mentioned prior art, the present invention provides an online routing method and system based on deep reinforcement learning.

[0004] Technical Solution: The present invention provides an online routing method based on deep reinforcement learning, specifically including the following steps:

[0005] Step 1: Calculate K paths between the source node and the destination node;

[0006] Step 2: Predict the performance metrics of the optimal path p obtained in the (n - 1)-th period, where the performance metrics include end-to-end latency, end-to-end throughput, and end-to-end packet loss rate; n is a positive integer greater than 2;

[0007] Step 3: Calculate the actual performance metrics of the optimal path y obtained in the (n - 2)-th period;

[0008] Step 4: Calculate the relative difference between the predicted performance metrics in Step 2 and the actual performance metrics in Step 3, and update the weights of end-to-end latency , end-to-end remaining bandwidth and end-to-end packet loss rate ;

[0009] Step 5: Based on the weights updated in Step 4, update the reward function for the nth cycle. Based on the reward function for the nth cycle, use the SAC algorithm to calculate the optimal path for the nth cycle, generate a flow table, and send it to the corresponding switch device for path installation and traffic forwarding, then go to Step 2.

[0010] Further, in Step 1, the K - shortest path algorithm is used to calculate K paths between the source node and the destination node.

[0011] Further, in Step 2, a message - passing neural network is used to predict the prediction performance metrics of the optimal path obtained in the (n - 1)th cycle. The training of the message - passing neural network is specifically as follows:

[0012] Step 2.1: Construct the topological graph of the communication network;

[0013] Step 2.2: Simulate the constructed communication network topological graph under different traffic conditions to generate a data set;

[0014] Step 2.3: Pre - process the data of link - level features and path - level features in the data set, and initialize the hidden state of the link and the hidden state of the path;

[0015] Step 2.4: Perform T loop iterations on the message - passing neural network to update the hidden state of the link and the hidden state of the path, specifically as follows:

[0016] Message - passing process of the path: When performing the (t + 1)th iteration, use a recurrent neural network to encode the hidden state of the link at the tth iteration to obtain a link state sequence related to this path , and input and into the attention network to obtain the dynamically weighted aggregated information of the link state . Input the hidden state of the path at the tth iteration and into the neural network μ to obtain the hidden state of the path at the (t + 1)th iteration output by the neural network μ ;

[0017] Message - passing process of the link: When performing the (t + 1)th iteration, use the summation method to aggregate the state information of all paths in the link at the (t + 1)th iteration to obtain the transition state information of the path , and then input the hidden state of the link and into the recurrent neural network to obtain the hidden state of the link at the (t + 1)th iteration ;

[0018] Step 2.5: Finally, input the hidden state of the path obtained in the T-th loop , into the readout function to obtain the predicted performance metrics of the optimal path.

[0019] Furthermore, the preprocessing in Step 2.3 is as follows: perform normalization processing on data represented by numerical values, and perform one-hot encoding processing on data represented by categories.

[0020] Furthermore, in Step 3, the actual end-to-end delay , actual end-to-end throughput and actual end-to-end packet loss rate of the optimal path y are calculated according to the following formula:

[0021] ;

[0022] ;

[0023] ;

[0024] where i represents the source node, j represents the destination node, e ij represents the link in the optimal path y, d(e ij ) represents the delay of the link e ij , min represents the minimum function, thr(e ij ) is the throughput of the link e ij , is the packet loss rate of the link e ij .

[0025] Furthermore, Step 4 is specifically as follows:

[0026] Step 4.1: Calculate the relative difference between the predicted performance metrics and the actual performance metrics:

[0027] ;

[0028] ;

[0029] ;

[0030] where rel_delta_delay is the relative difference between the end-to-end delays, rel_delta_thr is the relative difference between the end-to-end throughputs, rel_delta_loss is the relative difference between the end-to-end packet loss rates, p_delay is the predicted end-to-end delay of the optimal path p, p_thr is the predicted end-to-end throughput of the optimal path p, p_loss is the predicted end-to-end packet loss rate of the optimal path p; is the actual end-to-end delay of the optimal path y, is the actual end-to-end throughput of the optimal path y, is the actual end-to-end packet loss rate of the optimal path y, is a constant;

[0031] Step 4.2: Calculate the weights of end-to-end delay, end-to-end remaining bandwidth, and end-to-end packet loss rate:

[0032] ;

[0033] ;

[0034] ;

[0035] ;

[0036] ;

[0037] ;

[0038] where k1 represents a scaling factor, , and are all intermediate quantities.

[0039] Furthermore, the specific expression of the reward function r in the SAC algorithm is:

[0040] ;

[0041] where, represents the th path, ; represents the value of the end-to-end delay of the th path after normalization, represents the value of the end-to-end remaining bandwidth of the th path after normalization, represents the value of the end-to-end packet loss rate of the th path after normalization.

[0042] An online routing system based on deep reinforcement learning includes a network awareness module, a network monitoring module, a data processing module, a prediction module, an intelligent optimization module, and a path installation module;

[0043] The network monitoring module periodically sends status request information to the forwarding device, asynchronously receives the port status information of the forwarding device; and transmits it to the data processing module and the prediction module. The port status information includes the enabled status, traffic load, packet loss rate, and delay of the port;

[0044] The network perception module collects network topology, global routing scheme, path-level information, and link-level information, and transmits them to the data processing module and the prediction module.

[0045] The data processing module calculates the actual values of K paths and performance metrics.

[0046] The prediction module predicts performance metrics based on network topology, global routing scheme, flow-level information, and path-level information.

[0047] The intelligent optimization module calculates the weights of end-to-end delay, end-to-end remaining bandwidth, and end-to-end packet loss rate in the reward function, and selects the optimal path among the K paths as the global routing scheme.

[0048] The path installation module obtains the optimal path calculated by the intelligent optimization module, generates a flow table, and downloads it to the switch device for path installation and traffic forwarding.

[0049] An electronic device / system for an online routing method includes a processor and a memory. The memory stores execution instructions for the processor, and the processor is configured to execute the execution instructions to implement the above online routing method.

[0050] A computer-readable storage medium is used to store a program, and executing the program implements the above online routing method.

[0051] Beneficial effects: The present invention proposes a brand-new online routing method (abbreviated as GSAC-P), which combines a performance prediction model based on a Message Passing Neural Network (MPNN) and a routing generation model based on a Soft Actor-Critic (SAC) algorithm, and at the same time introduces a reward function that can dynamically adjust metric weights. This method effectively solves the problem that some existing online routing algorithms cannot respond to network changes in a timely manner when dealing with strong bursty traffic and strong adversarial network environments. Through this innovative design, in a complex network, when encountering a sudden increase in traffic or a link failure, the present invention can not only improve the transmission performance of different quality-of-service services, but also adapt to network state changes more quickly and accurately, significantly improve network throughput, and effectively reduce end-to-end delay. BRIEF DESCRIPTION OF THE DRAWINGS

[0052] Figure 1 is a flowchart of the present invention.

[0053] Figure 2 is a schematic diagram of the architecture of the present invention.

[0054] Figure 3It is a graph of the graph neural network model based on MPNN.

[0055] Figure 4 It is a communication network topology graph.

[0056] Figure 5 It is a communication network topology graph after simulating partial link failures.

[0057] Figure 6 It is a curve graph of the average end-to-end delay under traffic changes.

[0058] Figure 7 It is a curve graph of the average network throughput ratio under traffic changes.

[0059] Figure 8 It is a box plot of the average end-to-end delay under traffic changes.

[0060] Figure 9 It is a bar chart of the average network throughput under traffic changes.

[0061] Figure 10 It is a curve graph of the average end-to-end delay before and after partial link failures.

[0062] Figure 11 It is a curve graph of the average network throughput ratio before and after partial link failures.

[0063] Figure 12 It is a box plot of the average end-to-end delay before and after partial link failures.

[0064] Figure 13 It is a bar chart of the average network throughput before and after partial link failures. Detailed implementation manners

[0065] The accompanying drawings that form a part of the present invention are used to provide a further understanding of the present invention. The schematic embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute an improper limitation of the present invention.

[0066] The specific process of the method of the present invention is as Figure 1 shown:

[0067] Step 1, define the topology graph of the communication network as , where G represents the topology graph of the entire communication network. The topology graph includes a set V of communication entities, an edge set E1 of communication links between entities, a transmission flow set F, a link set L, and a path set P1.

[0068] Regard the communication entities as nodes, and set the set , where k2 represents the total number of nodes, represents the k2-th node, and the link set , represents the total number of links, Indicates the th link; transmission flow set , represents the total number of transmission flows, Indicates the th transmission flow, path set , represents the total number of paths, Indicates the th path. The path is the transmission path that the flow follows from source to destination. Therefore, the path is defined as a sequence of flows and traversed links , where , respectively represent the q-th indices of the flow and link along the path , and J represents the total number of indices.

[0069] Step 2: Use the simulation platform to build the communication network topology diagram to be predicted, simulate 50 traffic situations, and generate a dataset.

[0070] The present invention simulates 50 random traffic flows, and these random traffic flows have different traffic characteristics, including the bandwidth size of the traffic, the arrival time distribution of data packets, the size of data packets, and the source-destination nodes.

[0071] Step 3: Preprocess the data on link-level features and path-level features in the dataset, and initialize the hidden states of links and paths through the feature embedding H function.

[0072] First, the data in the dataset is classified according to link-level features and path-level features. Among them, link-level features include bandwidth size, link utilization, and load, etc. Path-level features include the bandwidth size of the traffic and the size of data packets, etc.

[0073] For data represented by numerical values, subtract its actual value from its average value and then divide by the standard deviation for normalization processing.

[0074] For data represented by categories, it can be encoded in one-hot encoding form for subsequent matrix operations.

[0075] Finally, initialize the hidden state by passing the link and path features through an input layer.

[0076] Step 4: Propagate and update information between states through the message passing mechanism of the message passing neural network. The structure diagram of this network is as Figure 2 shown.

[0077] First, input the link set L, path set P, initial link hidden state, and initial path hidden state.

[0078] Message passing process of the path: At the current (t + 1)-th iteration, a recurrent neural network (RNN) is used to encode the hidden state of the link at the t-th iteration to obtain a link state sequence related to this path , and and are input into the attention network to obtain the dynamically weighted aggregated information of the link state . The hidden state of the path at the t-th iteration and are input into the neural network μ to obtain the hidden state of the path at the (t + 1)-th iteration output by the neural network μ ;

[0079] Message passing process of the link: When performing the (t + 1)-th iteration, for any link, the state information of all paths passing through this link at the (t + 1)-th iteration is aggregated in a summation manner to obtain the transition state information of the path , and then the link hidden state and are input into the RNN network to obtain the hidden state of the link at the (t + 1)-th iteration .

[0080] Step 5: Loop the three-stage message passing process in Step 4 for T times to achieve message passing and update in farther nodes.

[0081] In a graph neural network, a node updates its own features through interactions with neighboring nodes. In each iteration, a node obtains information from its neighbor nodes and adjusts its own features based on this information. The depth of information transmission is determined by the number of iterations. Increasing the number of iteration rounds can make the information spread farther, so that each node can integrate feature data from more nodes.

[0082] Step 6: The readout function Readout outputs the path-level hidden state after the T-th loop to predict some path-level information, that is, performance metrics, including end-to-end delay, end-to-end throughput, and end-to-end packet loss rate; and triples are used to store the performance metric prediction information on the path from source node i to destination node j.

[0083] Step 7: Using the smooth L1 loss as the loss function, train the overall model, iterate to obtain the convergence value, and store the trained model. The above Steps 1 to 7 are all the training process of the message passing neural network.

[0084] Step 8: The system initializes the network perception module, network monitoring module, data processing module, prediction module, intelligent optimization module, and path installation module;

[0085] Figure 3 It is a schematic diagram of the architecture of the present invention, including three-layer structures of a data layer, a control layer, and an application layer. The control layer further includes a network awareness module, a network monitoring module, a data processing module, a prediction module, an intelligent optimization module, and a path installation module.

[0086] The data layer consists of programmable switches and communication links, accepts the policies of the control layer through the southbound interface and processes data packets, and is mainly used for modeling complex networks.

[0087] The control layer is used to connect the data layer and the application layer, maintain a global view, pass control policies to the data layer downward, provide network resource information to the application layer upward, and periodically send query messages to the data layer.

[0088] The application layer communicates with the control layer through the northbound interface, uses the network status information provided by the data layer, and supports various application services.

[0089] Step 9, the data processing module executes the K Shortest Paths Algorithm (KSP) according to the topological structure and source-destination node information to calculate K feasible paths.

[0090] Step 10, the network awareness module and the network monitoring module periodically collect the original data of the network to detect network state changes, including network topological structure, global routing scheme, port status information, path-level information, link-level information, etc., and store these data.

[0091] The network awareness module periodically sends feature request information to the forwarding devices (such as switches) of the data layer to obtain the latest status of the network topological structure, which includes the connection relationships of each node, link status, etc. In addition, the network awareness module also collects global routing scheme information, and tracks the path selection and traffic distribution of each traffic.

[0092] The network monitoring module periodically sends status request information to the forwarding devices of the data layer and asynchronously receives the status statistical information of the devices. It mainly focuses on the port status information of the devices, such as the enabled status of the port, traffic load, packet loss rate, delay, etc. Through these real-time monitoring data, the network monitoring module can comprehensively reflect the running health status of the network.

[0093] Step 11, the data processing module calculates the actual performance indicators of the optimal path in the (n - 2)th cycle according to the network topological structure and port status information. The performance indicators include end-to-end delay , end-to-end throughput and end-to-end packet loss rate (due to the influence of delay, in this embodiment, the performance indicators of the optimal path in the first cycle are calculated starting from the third cycle):

[0094] ;

[0095] ;

[0096] ;

[0097] where y represents the optimal path obtained in the (n - 2)-th period, i represents the source node, j represents the destination node, e ij represents the link in the optimal path y, and d(e ij ) represents the delay of link e ij ; min represents the minimum function, and thr(e ij ) is the throughput of link e ij , is the packet loss rate of link e ij .

[0098] Step 12: Input the network topology structure, global routing scheme, link-level information and path-level information into the prediction module (i.e., the message passing neural network), and the prediction module outputs the end-to-end delay, end-to-end throughput and end-to-end packet loss rate of the predicted optimal path obtained in the (n - 1)-th period.

[0099] The intelligent optimization module updates the weights of the end-to-end delay, end-to-end remaining bandwidth and end-to-end packet loss rate in the reward function according to the predicted performance metrics and the actual performance metrics calculated in Step 11.

[0100] By calculating the relative differences between the predicted values and the actual values, and using the Softmax function to weight these differences, the weights of the end-to-end delay, end-to-end remaining bandwidth and end-to-end packet loss rate are obtained, thereby updating the reward function. First, calculate their relative differences according to the following formula:

[0101] ;

[0102] ;

[0103] ;

[0104] where rel_delta_delay is the relative difference between the end-to-end delays, rel_delta_thr is the relative difference between the end-to-end throughputs, rel_delta_loss is the relative difference between the end-to-end packet loss rates, p_delay is the predicted end-to-end delay of the optimal path p in the (n - 1)-th period, p_thr is the predicted end-to-end throughput of the optimal path p, p_loss is the predicted end-to-end packet loss rate of the optimal path p, is a constant.

[0105] Calculate the weights of end-to-end delay, remaining bandwidth, and packet loss rate:

[0106] ;

[0107] ;

[0108] ;

[0109] ;

[0110] ;

[0111] ;

[0112] Among them, is the weight of end-to-end delay, is the weight of end-to-end remaining bandwidth, is the weight of end-to-end packet loss rate, k1 represents the scaling factor, , and are both intermediate quantities.

[0113] Step 13: Construct a deep reinforcement learning agent based on the SAC algorithm. The intelligent optimization module inputs the information obtained by the data processing module into the agent, and the agent obtains the network state at the current moment , selects the optimal path action for it , then obtains the network state at the next moment , and simultaneously obtains the current reward according to the latest reward function , and stores it in the experience replay pool in the form of a quadruple for model training.

[0114] The SAC algorithm is a deep reinforcement learning algorithm based on the maximum entropy idea. Different from traditional deep reinforcement learning algorithms, it uses a stochastic policy instead of a deterministic policy. At the same time, the SAC algorithm shows significant advantages in off-policy learning and exploration ability, and can better balance exploration and exploitation. Its core idea is to maximize the long-term cumulative reward while also maximizing the entropy of the policy, encouraging the policy to maintain sufficient randomness. Specifically, it achieves this by optimizing the following objective function:

[0115] ;

[0116] Among them, represents the optimized policy, π represents the policy network, E represents the expectation function, is the discount factor, is the entropy regularization coefficient, Indicates that the input state is , and the action is The reward value at this time; is the entropy value. Introducing the concept of entropy helps to keep the policy sufficiently random, promotes the agent to explore the state space more comprehensively, and avoids the policy falling into local optimal solutions. At the same time, the agent can explore multiple possible solutions, enhancing its adaptability to environmental changes and improving the robustness and anti-interference ability of the policy. max represents the maximum function.

[0117] The SAC algorithm is an off-policy algorithm that uses an experience replay mechanism to eliminate data correlation, reuses historical data, and improves the utilization rate of samples; it uses a double Q-value network structure to avoid overfitting of Q-values and ensures the stability of training. Let Q(.) represent the source Q-value network, represents the parameters of the policy network. The source Q-value network learns by minimizing the flexible Bellman residual:

[0118]

[0119] where, is the loss function of the source Q-value network, is the intermediate function, , represents the policy network with parameters , is the target Q-value network, and its parameters are softly updated from the parameters of the source Q-value network by the method of Polyak averaging:

[0120]

[0121] where, is the soft update parameter, , and the parameters of the target network are updated smoothly, improving the stability of the learning process.

[0122] The policy function is obtained by minimizing the KL divergence of the following formula:

[0123] ;

[0124] where, is the objective function of the policy network, s~D is the state in the sampling distribution, is the action sampled according to the probability distribution of the policy network .

[0125] The entropy regularization coefficient is updated by the automatic entropy adjustment method:

[0126] ;

[0127] Among them, is the hyperparameter of the target entropy.

[0128] State space. The state space consists of all states that the agent can observe. Each state reflects the situation of all traffic requests in the network at the current moment (including information about source nodes and destination nodes), as well as the feasible path states related to these requests. The state information is organized into a state matrix , and is represented by the state . N represents the source nodes and destination nodes of all traffic requests in the network at the current moment. PM consists of K pieces of feasible path state information (delay, remaining bandwidth, and packet loss rate), and the specific form is as follows:

[0129] ;

[0130] Among them, represents the end-to-end delay of the th path, represents the end-to-end remaining bandwidth of the th path, represents the end-to-end packet loss rate of the th path.

[0131] Since there are large differences in the values of the elements in the state matrix, it causes fluctuations in the training of the agent, and it is difficult for the model to converge. To solve this problem, the Min—Max normalization method is used to standardize each element in PM, and the value of each element can be adjusted to the range of .

[0132] Action space. The action space refers to a set of actions that the agent can choose to execute in each state. Each action corresponds to a set of feasible paths from the source node to the destination node in a given state. The agent selects an optimal path from this set to achieve the best routing of network traffic.

[0133] Reward function. The reward function is used to measure the feedback obtained by the agent after taking a certain action, so as to help it optimize the decision-making strategy. In reinforcement learning, the goal of the agent is to continuously maximize the obtained reward value. Therefore, the calculation method of the reward value r is as follows:

[0134] ;

[0135] Among them, represents the th path, ; represents the value of the end-to-end delay of the kth path after normalization, It represents the normalized value of the end-to-end remaining bandwidth of the k-th path. It represents the normalized value of the end-to-end packet loss rate of the path. The weights of the end-to-end delay in the initial period and the second period are both 0.3, the weights of the end-to-end remaining bandwidth in the initial period and the second period are both 0.4, and the weights of the end-to-end packet loss rate in the initial period and the second period are both 0.4. Starting from the third period, the weights of each period are dynamically updated according to step 12.

[0136] Step 14: When the experience replay pool is stored to the set threshold, the agent performs model training and parameter update, trains iteratively until the model converges, and then the intelligent optimization module calculates and stores the optimal path of the global network node pair; first, sample N tuples from the experience replay pool , calculate the target value using the target Q-value network, update the source Q-value network by minimizing the Bellman residual, optimize the policy network by minimizing the KL divergence, and use the reparameterization trick to ensure that the gradient can be backpropagated, so as to directly optimize the network parameters using the gradient descent method. For the stability of training, the parameters of the two target Q-value networks are softly updated by Polyak averaging.

[0137] After updating the neural network parameters at each step, update the entropy regularization coefficient through the automatic entropy adjustment method , to balance the exploration (maximizing entropy) and exploitation (maximizing Q-value) of the policy.

[0138] Repeat the above training process, continuously optimize the parameters of the neural network until the algorithm converges and the learned policy reaches the expected performance level.

[0139] The values of the parameters in the SCA algorithm of this embodiment are shown in Table 1 below:

[0140] Table 1

[0141] Parameter Value Optimizer Adam Entropy regularization coefficient 0.008 Discount factor 0.95 Soft update parameter 0.01 Experience pool size 15000 Sampling number 32 Target entropy -2 Gradient clipping parameter 1 Temperature parameter in softmax 2 Total number of training rounds 120 Number of feasible paths 10

[0142] Step 15: The path installation module obtains the optimal path result calculated by the intelligent optimization module, generates a flow table, and issues it to the corresponding switch device in the data layer for path installation and traffic forwarding. In the next network period, the prediction model predicts the performance metrics of the optimal path calculated by the intelligent optimization module in the current period.

[0143] According to the optimal path generated by the intelligent optimization module, the system first locates the corresponding host node, and then, combines the optimal path and the host node information to construct the best traffic transmission path and update the corresponding flow table entries. Finally, these updated flow table entries will be issued to the data layer to ensure that network traffic can be efficiently transmitted according to the optimal path.

[0144] Figure 4 is the communication network topology of this embodiment, which consists of a backbone network and an access network. Its topology contains 44 switches and 61 links. Among them, the numbers 1-47 all represent nodes, and each node represents a switch supporting the OpenFlow protocol. One host is connected under each switch, which is responsible for the output and reception of traffic. Nodes 18, 19, 20 and nodes 32, 33, 34 simulate sensor nodes, while nodes 42, 43, 44 simulate command and control nodes. Nodes 11, 12, 13, nodes 25, 26, 27 and nodes 39, 40, 41 represent composite nodes. The traffic transmission of the network follows the path planning from "sensor" to "command and control" and then to "composite". Therefore, this embodiment selects 6 sensor nodes as source nodes and 9 composite nodes as destination nodes. The link connections of the command and control network adopt different types of heterogeneous links such as microwave, optical fiber, regional broadband, VHF, and UHF. In order to simulate these heterogeneous links, different transmission bandwidths are set in the Mininet simulation platform. The operating system for the experimental operation of this embodiment is Ubuntu 18.04 with 8GB of memory and a 2-core processor. The network simulation is implemented through the Mininet 2.3.0 platform, and the Ryu 4.34 controller is used to control the entire network. The communication network topology diagram after simulating partial link failures in this embodiment is as Figure 5 shown.

[0145] Figure 6 , Figure 7 , Figure 8 , Figure 9 are the comparison diagrams of the performance of each algorithm during the process of the traffic surging from 100 kbps to 500 kbps and then to 1000 kbps. In order to verify the advantages of the algorithm of the present invention, the comparison algorithms adopted in this embodiment are as follows:

[0146] SPR algorithm: The shortest path routing algorithm, which obtains the weight information of each link in the network through the SDN measurement mechanism and calculates the path with the shortest link weight.

[0147] DRL-ST algorithm: A Dueling DQN is used to construct an end-to-end transmission path decision model, and the sampling mechanism is optimized using a tree storage structure. For fair comparison, DRL-ST maintains the same training settings (such as state, action, and initial reward) as the algorithm of the present invention.

[0148] This embodiment uses network throughput and end-to-end delay to evaluate the impact of different routing algorithms on network performance.

[0149] Figure 6Shows the end-to-end delay of each algorithm under traffic changes. When using the DRL-ST algorithm, since this algorithm calculates the optimal path based on the periodically detected network state, when the traffic surges, the agent cannot respond to the changes in the network environment in a timely manner, resulting in a sharp increase in delay. However, as the next cycle arrives, the delay will gradually decrease. However, since the algorithm uses a fixed reward function, after entering the stable state, the delay cannot be reduced to the optimal level. When the environment changes, the SPR algorithm always selects the shortest path. At the first change, due to sufficient bandwidth, the performance of SPR is hardly affected, but when the second change occurs, the bandwidth of the link cannot meet the high-load demand, resulting in a sharp increase in delay. In contrast, the GSAC-P algorithm proposed by the present invention can quickly adapt to the changes in the network environment by introducing a prediction model, predicting the performance of the next cycle based on the current cycle's network topology, traffic, and routing scheme, and timely adjusting the reward function, and maintaining a low delay after the system stabilizes. Therefore, GSAC-P can obtain a stable and optimal routing scheme faster.

[0150] Figure 7 Shows the network throughput ratio of each algorithm under traffic changes. The throughput ratios of the two deep reinforcement learning (DRL)-based routing algorithms (DRL-ST and GSAC-P) both decrease after the traffic changes, but the GSAC-P algorithm of the present invention still has smaller changes and higher throughput after routing stabilizes, indicating that the method of the present invention is more stable.

[0151] Figure 8 The end-to-end delay under traffic changes is shown by a box plot. It can be seen that there are a large number of outliers in the DRL-ST algorithm after the traffic changes, indicating that the delay fluctuates greatly, indicating that the algorithm cannot respond to environmental changes in a timely manner.

[0152] Figure 9 The average network throughput under traffic changes is shown by a bar chart. It can be seen that under the three network states, the throughput corresponding to the present invention is the largest.

[0153] Figure 10 Shows the end-to-end delay of each algorithm before and after partial link failures at a traffic transmission rate of 1000 kbps. From Figure 10 it can be seen that the present invention can better and more quickly adapt to the changes in the network environment and maintain a low delay after the system stabilizes.

[0154] Figure 11It shows the network throughput ratios of each algorithm before and after partial link failures at a traffic transmission rate of 1000 kbps. Except for the SPR algorithm, the throughput ratios of the two deep reinforcement learning-based routing algorithms (DRL-ST and GSAC-P) decreased after the traffic change. However, the present invention still has a small change and a high throughput after the routing is stabilized, indicating that the present invention is more stable.

[0155] Figure 12 The box plot is used to show the end-to-end delay before and after partial link failures at a traffic transmission rate of 1000 kbps. It can be seen that a large number of outliers appear in the DRL-ST algorithm after the traffic change, indicating that the delay fluctuates greatly and the algorithm cannot respond to environmental changes in a timely manner.

[0156] Figure 13 The bar chart is used to show the average network throughput before and after partial link failures at a traffic transmission rate of 1000 kbps. It can be seen that the throughput corresponding to the present invention is the largest in the two network states.

[0157] In addition, it should be noted that, in the above specific embodiments, the various specific technical features described can be combined in any suitable manner without contradiction. To avoid unnecessary repetition, the present invention will not separately describe various possible combination methods.

Claims

1. An online routing method based on deep reinforcement learning, characterized in that, Specifically, it includes the following steps: Step 1: Calculate K paths between the source node and the destination node; Step 2: Predict the performance metrics of the optimal path p obtained in the (n - 1)-th period, where the performance metrics include end-to-end delay, end-to-end throughput, and end-to-end packet loss rate; n is a positive integer greater than 2; Step 3: Calculate the actual performance metrics of the optimal path y obtained in the (n - 2)-th period; Step 4: Calculate the relative difference between the predicted performance metrics in Step 2 and the actual performance metrics in Step 3, and update the weight w of the end-to-end delay according to the relative difference delay , the weight w of the end-to-end remaining bandwidth rbw and the weight w of the end-to-end packet loss rate loss ; Step 5: Based on the weights updated in Step 4, update the reward function in the n-th period. Based on the reward function in the n-th period, use the SAC algorithm to calculate the optimal path in the n-th period, generate a flow table, and send it to the corresponding switch device for path installation and traffic forwarding, then go to Step 2; The specific content of Step 4 is as follows: Step 4.1: Calculate the relative difference between the predicted performance metrics and the actual performance metrics; Among them, rel_delta_delay is the relative difference between end-to-end delays, rel_delta_thr is the relative difference between end-to-end throughputs, rel_delta_loss is the relative difference between end-to-end packet loss rates, p_delay is the predicted end-to-end delay of the optimal path p, p_thr is the predicted end-to-end throughput of the optimal path p, p_loss is the predicted end-to-end packet loss rate of the optimal path p; y pD is the actual end-to-end delay of the optimal path y, y pThr is the actual end-to-end throughput of the optimal path y, y pL is the actual end-to-end packet loss rate of the optimal path y, and ε is a constant; Step 4.2: Calculate the weights of end-to-end delay, end-to-end remaining bandwidth, and end-to-end packet loss rate; z delay = k1 * rel_delta_delay; z thr = k1 * rel_delta_thr; z loss = k1 * rel_delta_loss; wherein, k1 represents a scaling factor, z delay , z thr and z loss are all intermediate quantities; The specific expression of the reward function r in the SAC algorithm is: Among them, k represents the k-th path, where k = 1, 2,..., K; represents the value of the end-to-end delay of the k-th path after normalization, represents the value of the end-to-end remaining bandwidth of the k-th path after normalization, represents the value of the end-to-end packet loss rate of the k-th path after normalization.

2. The online routing method based on deep reinforcement learning according to claim 1, characterized in that In Step 1, the K shortest path algorithm is used to calculate K paths between the source node and the destination node.

3. An online routing method based on deep reinforcement learning according to claim 1, characterized in that In Step 2, a message passing neural network is used to predict the predicted performance metrics of the optimal path obtained in the (n - 1)-th period. The training of the message passing neural network is specifically as follows: Step 2.1: Construct the topology graph of the communication network; Step 2.2: Simulate the constructed communication network topology graph under different traffic conditions to generate a dataset; Step 2.3: Preprocess the data of link-level features and path-level features in the dataset, and initialize the hidden states of links and paths; Step 2.4: Perform T loop iterations on the message passing neural network to update the hidden states of links and paths, specifically: Message passing process of the path: When performing the (t + 1)-th iteration, the hidden state of the link at the t-th iteration is encoded using a recurrent neural network to obtain the link state sequence m related to this path p,l , and then and m p,l are input into the attention network to obtain the dynamically weighted link state aggregation information n p,l . The hidden state of the path at the t-th iteration and n p,l are input into the neural network μ to obtain the hidden state of the path at the (t + 1)-th iteration output by the neural network μ Message passing process of the link: When performing the (t + 1)-th iteration, the state information of all paths in the link at the (t + 1)-th iteration is aggregated using summation to obtain the path transition state information m l . Then, the hidden state of the link and m l are input into the recurrent neural network to obtain the hidden state of the link at the (t + 1)-th iteration Step 2.5: Finally, the hidden state of the path obtained from the T-th loop is input into the readout function to obtain the prediction performance metric of the optimal path 4. An online routing method based on deep reinforcement learning according to claim 3, characterized in that The preprocessing in Step 2.3 is: perform normalization processing on the data represented by numerical values, and perform one-hot encoding processing on the data represented by categories.

5. The online routing method based on deep reinforcement learning according to claim 1, characterized in that Step 3 calculates the actual end-to-end delay y pD of the optimal path y, the actual end-to-end throughput y pThr , and the actual end-to-end packet loss rate y pL according to the following formula: pD and the actual end-to-end throughput y pThr pThr and the actual end-to-end packet loss rate y pL pL : Among them, i represents the source node, j represents the destination node, and e ij represents the link in the optimal path y, and d(e ij ) represents the delay of the link e ij , min represents the minimum function, and thr(e ij ) is the throughput of the link e ij , and l(e ij ) is the packet loss rate of the link e ij .

6. A system for implementing the online routing method based on deep reinforcement learning according to claim 1, characterized in that, It includes a network awareness module, a network monitoring module, a data processing module, a prediction module, an intelligent optimization module, and a path installation module; the network monitoring module periodically sends status request messages to the forwarding device and asynchronously receives the port status information of the forwarding device; and transmits it to the data processing module and the prediction module. The port status information includes the enabled status of the port, traffic load, packet loss rate, and delay; The network awareness module collects the network topology structure, global routing scheme, path-level information, and link-level information; and transmits it to the data processing module and the prediction module; The data processing module calculates the actual values of K paths and performance metrics; The prediction module predicts performance metrics based on the network topology structure, global routing scheme, flow-level information, and path-level information; The intelligent optimization module calculates the weights of end-to-end delay, end-to-end remaining bandwidth, and end-to-end packet loss rate in the reward function, and selects the optimal path among the K paths as the global routing scheme; The path installation module obtains the optimal path calculated by the intelligent optimization module, generates a flow table, and sends it to the switch device for path installation and traffic forwarding.

7. A computer device, comprising a memory, a processor, and a computer program stored in the memory and capable of running on the processor, characterized in that When the processor executes the computer program, the steps of the online routing method according to any one of claims 1 to 5 are implemented.

8. A computer-readable storage medium for storing a program, characterized in that, Execute the program to implement the online routing method according to any one of claims 1 to 5.

Citation Information

Patent Citations

  • Flexible in-line routing method and device based on graph neural network

    CN117041128A

  • Space-ground integrated load balancing routing method based on deep reinforcement learning

    CN117395188A

  • Key flow routing optimization method based on deep reinforcement learning

    CN119363647A