A Path Intelligent Optimization Method and System Based on Enhanced Link State Awareness
By adopting the PPO algorithm, GRU network and GAT network with AC architecture in the intelligent routing algorithm, combined with the attention mechanism, the spatiotemporal characteristics of link state information are captured, and the problem that existing intelligent routing algorithms are difficult to make precise decisions when the network topology changes dynamically, achieving more efficient network service performance.
Patent Information
- Application Number
- CN202411480526.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-10-23
- Publication Date
- 2025-06-24
- Estimated Expiration
- 2044-10-23
AI Technical Summary
The existing intelligent routing algorithm model has a single structure and lacks consideration of the spatial and temporal characteristics of link state information, which makes it difficult to accurately implement routing decisions when the network topology structure changes dynamically.
An intelligent path decision model is constructed using the PPO algorithm based on AC architecture, combining the GRU network and the GAT network to capture the spatiotemporal and spatial-temporal characteristics of link state information, and perceive the important characteristics of the network state sequence through the attention mechanism to generate the optimal path selection strategy.
It realizes more accurate routing decisions when the network topology changes dynamically, and improves network service performance. Compared with traditional routing algorithms and existing intelligent routing algorithms, the average end-to-end delay and packet loss rate are reduced by at least 14.42% and 14.66% respectively.
Smart Images

Figure CN119011463B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of network traffic control, and particularly relates to a method for intelligent path optimization enhanced by link state awareness, and also relates to a system for intelligent path optimization enhanced by link state awareness. Background Art
[0002] In the global network path planning algorithm, traditional routing algorithms usually calculate the shortest path based on limited link state information to achieve service transmission, which makes it difficult for the network to adapt to the rapid changes in service traffic and thus unable to meet the service QoS requirements; in addition, the continuous growth of service traffic and the diversity of application programs in complex communication networks also greatly reduce the efficiency and accuracy of traditional routing algorithms based on limited information decision-making.
[0003] With the development of Software Defined Network (SDN), the emergence of this new network architecture has realized the decoupling of network control and packet forwarding, improving the programmability, global view, and logical centralized control of the network, providing a new opportunity for the implementation of traffic engineering and end-to-end QoS research. In the existing research field of intelligent routing algorithms based on the SDN architecture, heuristic algorithms are still the main basis: reliable nodes are identified and the optimal path is formed by considering the performance of the routing node itself (direct information) and the performance of adjacent nodes (indirect information); and by comprehensively considering factors such as path distance, transmission direction, and link load, a path with lower delay and overhead is found to achieve network load balancing. However, heuristic algorithms have relatively strict requirements for network scenarios, and changes in network topology and link state may cause large fluctuations and errors in the algorithm, resulting in problems such as insufficient scalability of the algorithm, thus affecting network performance.
[0004] With the rapid development of deep reinforcement learning technology, intelligent algorithms based on Deep Reinforcement Learning (DRL) have shown significant advantages in solving complex high-dimensional space problems, making their applications in fields such as network optimization and resource allocation more and more extensive; in the prior art, online routing decisions under the SDN architecture are achieved by comprehensively considering multiple QoS metrics such as delay and bandwidth. Or the DDPG algorithm is combined with the SDN architecture to achieve global and real-time network intelligent management and control, effectively improving network throughput and reducing average delay.
[0005] The above intelligent routing algorithm based on deep reinforcement learning can perceive complex and dynamic network states, achieve self-adaptive adjustment of routing strategies, and effectively improve the quality of service. However, the topology of communication networks is dynamically variable, and link states mutate instantaneously. Most existing intelligent algorithms use fixed neural network structures to learn network states, resulting in limited model perception ability and difficulty in adapting to dynamically changing network environments. Therefore, existing intelligent routing algorithms are difficult to apply to optimal path decision-making in high-mobility complex environments with strong adversaries. Summary of the Invention
[0006] Objective of the present invention: Aiming at the problem that the existing intelligent routing algorithm has a single model structure and lacks consideration of the spatio-temporal correlation characteristics of link state information, resulting in difficulty in accurately implementing routing decisions when the network topology structure changes dynamically, a path intelligent optimization method and system based on enhanced link state perception (DRL-SGA) is proposed. This method constructs an intelligent path decision model based on the PPO algorithm of the AC architecture to generate an optimal path selection strategy. For a network with a dynamically changing topology, it accurately implements routing decisions and improves network service performance.
[0007] To achieve the above functions, the present invention designs a path intelligent optimization method based on enhanced link state perception. For the target network, the following steps S1 - S4 are executed to complete the path selection of business data forwarding in the target network:
[0008] Step S1: Collect the topology information and port status information of the target network, calculate the communication link state information and end-to-end path state information of the target network, including service requests, remaining bandwidth, delay, and packet loss rate of the global network, form the network state of the target network, and calculate the average throughput, average end-to-end delay, and average packet loss rate of the global network;
[0009] Step S2: The intelligent agent in the target network constructs an intelligent path decision model based on the PPO algorithm of the AC architecture, inputs the network state sequence obtained in Step S1 into the intelligent agent, and the intelligent agent s t executes a path selection action a t and then obtains the network state sequence at the next moment s t+1 Meanwhile, obtain the reward at the current moment r t and form a sample in the form of a quadruple ( s t , a t , r t , s t+1 ) and store it in the experience replay pool;
[0010] Step S3: Extract samples from the experience replay pool, iteratively train the intelligent path decision-making model, update the path selection strategy until the intelligent path decision-making model converges. For the target network at the current moment, use the converged intelligent path decision-making model to generate and store the path selection strategy, and take the path between the target network node pairs corresponding to the path selection strategy as the optimal path;
[0011] Step S4: Generate a flow table based on the optimal path and send it to the switch device of the target network for path installation and service data forwarding.
[0012] The present invention also designs a path intelligent optimization system based on enhanced link state awareness, including three-layer structures of a data layer, a control layer, and an application layer to implement the described path intelligent optimization method based on enhanced link state awareness:
[0013] The data layer includes various routing nodes and communication links of the target network, transmits the topology information and port status information of the target network to the control layer through the southbound interface, and at the same time receives the path selection strategy issued by the control layer and completes the processing and forwarding operations of service data;
[0014] The control layer includes five modules: a network awareness module, a network monitoring module, a data processing module, an intelligent optimization module, and a path installation module;
[0015] The control layer periodically sends preset request instructions to the data layer through the southbound interface, obtains the topology information and port status information of the target network in real time, and transmits the path selection strategy to the data layer;
[0016] Among them, the network awareness module periodically sends characteristic request instructions to the data layer to obtain the topology information of the target network; the network monitoring module periodically sends status request instructions to the data layer and asynchronously receives status reply messages to obtain the port status information of the routing nodes in the target network; the data processing module uses the topology information and port status information collected by the network awareness module and the network monitoring module to calculate the link state information and end-to-end path state information, and then statistically calculates and stores the global network average throughput, average end-to-end delay, and average packet loss rate; the intelligent optimization module constructs an intelligent path decision-making model, and according to the network state sequence at the current moment s t , perform a path selection action a t , obtain the network state sequence at the next moment s t+1 , and at the same time obtain the reward at the current moment r t ; the path installation module performs path installation according to the path selection action at Generate the corresponding flow table and send it to the data layer for forwarding business data;
[0017] The application layer includes various services and applications of the target network, and conducts information interaction with the control layer through the northbound interface.
[0018] Beneficial effects: Compared with the prior art, the advantages of the present invention include:
[0019] The present invention proposes a method and system for intelligent path optimization based on enhanced link state awareness (DRL-SGA). The method and system designed by the present invention construct an intelligent path decision model based on the PPO algorithm of the AC architecture, and introduce the GRU network and the GAT network to capture the spatio-temporal correlation characteristics of the network state sequence composed of multiple link states respectively. At the same time, the attention mechanism is used to perceive the important features of the network state sequence, and the optimal path selection strategy is generated through the training mechanism based on deep reinforcement learning; the experimental results show that the DRL-SGA method reduces the average end-to-end delay and packet loss rate by at least 14.42% and 14.66% compared with the traditional routing algorithm OSPF; compared with the existing intelligent routing algorithm DRL-ST, the average end-to-end delay and packet loss rate are reduced by at least 2.07% and 1.65% respectively, and the average throughput is maximally increased by 2.59% under high-intensity traffic, and the method and system designed by the present invention have stronger adaptability to dynamic changes in the network structure. Description of the Drawings
[0020] Figure 1 is a flowchart of a method for intelligent path optimization based on enhanced link state awareness according to an embodiment of the present invention;
[0021] Figure 2 is a structural diagram of a method for intelligent path optimization based on enhanced link state awareness according to an embodiment of the present invention;
[0022] Figure 3 is a schematic diagram of a policy network according to an embodiment of the present invention;
[0023] Figure 4 is a schematic diagram of a system for intelligent path optimization based on enhanced link state awareness according to an embodiment of the present invention;
[0024] Figure 5 is a command and control network topology and a set of damaged link diagrams according to an embodiment of the present invention;
[0025] Figure 6(a) is a performance comparison diagram of average network throughput according to an embodiment of the present invention;
[0026] Figure 6(b) is a performance comparison diagram of average end-to-end delay according to an embodiment of the present invention;
[0027] Figure 6 (c) is a performance comparison diagram of the average packet loss rate provided by an embodiment of the present invention;
[0028] Figure 7 (a) is a relationship diagram between the number of link damages and the average network throughput when the traffic intensity is 50 kbps provided by an embodiment of the present invention;
[0029] Figure 7 (b) is a relationship diagram between the number of link damages and the average network throughput when the traffic intensity is 100 kbps provided by an embodiment of the present invention;
[0030] Figure 7 (c) is a relationship diagram between the number of link damages and the average end-to-end delay when the traffic intensity is 50 kbps provided by an embodiment of the present invention;
[0031] Figure 7 (d) is a relationship diagram between the number of link damages and the average end-to-end delay when the traffic intensity is 100 kbps provided by an embodiment of the present invention;
[0032] Figure 7 (e) is a relationship diagram between the number of link damages and the average packet loss rate when the traffic intensity is 50 kbps provided by an embodiment of the present invention;
[0033] Figure 7 (f) is a relationship diagram between the number of link damages and the average packet loss rate when the traffic intensity is 100 kbps provided by an embodiment of the present invention. Detailed implementation manners
[0034] The following further describes the present invention with reference to the accompanying drawings. The following embodiments are only used to more clearly illustrate the technical solutions of the present invention and cannot be used to limit the protection scope of the present invention.
[0035] A path intelligent optimization method based on enhanced link state awareness provided by an embodiment of the present invention, referring to Figure 1 、 Figure 2 , perform the following steps S1 - step S4 to complete the path selection for the forwarding of service data in the target network:
[0036] Step S1: Collect the topology information and port status information of the target network, calculate the communication link state information and end-to-end path state information of the target network, including service requests, remaining bandwidth, delay, and packet loss rate of the global network, form the network state of the target network, and calculate the average throughput, average end-to-end delay, and average packet loss rate of the global network;
[0037] Step S2: The intelligent agent in the target network constructs an intelligent path decision model based on the PPO algorithm of the AC architecture, inputs the network state sequence obtained in step S1 into the intelligent agent, and the intelligent agent based on the network state sequence at the current moment s t , perform the path selection action at , and then obtain the network state sequence at the next moment s t+1 , and at the same time obtain the reward at the current moment r t , and form a sample in the form of a quadruple ( s t , a t , r t , s t+1 ) and store it in the experience replay pool;
[0038] As a framework deep reinforcement learning algorithm, the PPO algorithm needs to design different state spaces, action spaces and reward functions for different problems and application scenarios. The present invention designs the state space, action space and reward function based on the PPO algorithm for the target communication network scenario, specifically as follows:
[0039] Step S2.1: The agent collects the network state at t time to form the state space, and the network state at x t is as follows: t time x t is as follows:
[0040] ;
[0041] In the formula, , , , respectively represent the service request, link remaining bandwidth, delay, and packet loss rate of the global network at t time; f = 4 represents the network state feature dimension, and n represents the total number of target network links;
[0042] Among them, the expression of the global network service request is as follows:
[0043] ;
[0044] In the formula, represents t the service request on the link i between node j and node e ij in the target network at m time, i is the total number of nodes; when there is a traffic flow on the link between node j and node For the links without traffic flow and the links composed of non - adjacent nodes, assign their elements to 0;
[0045] The remaining bandwidth of the links in the global network The expression is as follows:
[0046]
[0047] In the formula, represents t At time, the remaining bandwidth of the link between node i and node j in the target network; e ij ;
[0048] The delay of the global network The expression is as follows:
[0049]
[0050] In the formula, represents t At time, the delay of the link between node i and node j in the target network; e ij ;
[0051] The packet loss rate of the global network The expression is as follows:
[0052] ;
[0053] In the formula, represents t At time, the packet loss rate of the link between node i and node j in the target network; e ij ;
[0054] Adopt the Min - Max method to normalize each element in , , , ; Based on the network state at time t and the network states in the previous t -1 time steps before time l , combine to form the network state sequence at time t s t as follows:
[0055] ;
[0056] Step S2.2: The agent selects a path according to the network state sequence s t and forms an action space. Assume that there are a t feasible paths between the global network source-destination node pairs, forming a set of feasible paths k . Each path corresponds to a path weight, forming a set of path weights ; Define each path selection action as ; where ; , , w ij represents the source node i and the destination node j the path weight of selecting path p between them, and ;
[0057] Step S2.3: The agent collects the real-time network performance metrics after executing the path selection action a t , sets a reward function to calculate the reward value, and feeds the reward value back to the agent. The reward function is as follows:
[0058] ;
[0059] where r is the reward value, α , β , γ are reward weights, and the value range is . Each reward weight is used to define the importance of different network performance metrics; To avoid the impact of network performance metric differences on the algorithm convergence performance, normalization processing is performed on them, , , represents the network performance metric after normalization processing.
[0060] The network performance metrics mentioned above are the global network average throughput, average end-to-end delay, and average packet loss rate.
[0061] The AC (Actor-Critic) architecture on which the intelligent path decision-making model is based includes a policy (Actor) network and an evaluation (Critic) network. The policy network outputs a path selection action x t according to the network state a t , and the evaluation network outputs the evaluation value of the network state;
[0062] Refer to Figure 3, the policy network includes a time feature extraction module, a self-attention mechanism module, a spatial feature extraction module, and a multi-layer perceptron module;
[0063] The time feature extraction module is constructed based on the GRU network and includes an update gate, a reset gate, and a hidden layer. When considering the time correlation of link state information in the present invention, a network state sequence with a time step of l needs to be input for the agent; when obtaining the initial network state sequence s t , the network states of the previous l -1 historical moments cannot be obtained, so their sequence elements are assigned 0. As the reinforcement learning iteratively updates, the network state sequence is also updated accordingly; among them, the update gate takes the network state x t at the current moment and the hidden state h t-1 at the previous moment as inputs, and calculates a value between [0,1] through the sigmoid function to determine the retention degree of information in the update gate; specifically as follows:
[0064] ;
[0065] In the formula, Z t is the update gate, W z is the update gate weight matrix, h t-1 is the hidden state at the previous moment, U z is h t-1 's update gate weight matrix;
[0066] The reset gate is similar in function to the update gate. It takes the network state x t at the current moment and the hidden state h t-1 at the previous moment as inputs, and calculates a value between [0,1] by the sigmoid function to determine the retention degree of information in the reset gate; specifically as follows:
[0067] ;
[0068] In the formula, R t is the reset gate, W r is the reset gate weight matrix, U r is h t-1 's reset gate weight matrix;
[0069] The formula for resetting the memory information is as follows. Take the output of the reset gate at the current moment r t and multiply it element-wise with the hidden state at the previous moment h t-1 Then, use the result of this calculation and the network state at the current moment x t as the concatenated input, and then pass it through tanh a function to calculate the candidate hidden state at the current moment :
[0070] ;
[0071] In the formula, is the weight matrix, is h t-1 's weight matrix, represents the Hadamard product;
[0072] The hidden state at the current moment h t has the following output formula. Use the update gate at the current moment z t to combine the hidden state at the previous moment h t-1 and the candidate hidden state at the current moment to obtain the next hidden state h t , and pass it to the neurons in the next time step, continuously updating the weight parameters of the GRU network model to extract the time features:
[0073] ;
[0074] Since the network state sequence s t input by the agent l contains l time steps, the GRU network contains
[0075] units, and the output of the GRU network is as follows:
[0076] In the formula, H R represents the output of the GRU network, represents the output of the hidden layer of the i th unit in the GRU network, ; f represents the dimension of the network state features, and n represents the total number of target network links;
[0077] The self-attention mechanism is used to capture the dependencies and importance among different samples in the network state sequence. It determines the importance of each sample in the model calculation by calculating the correlation degree (i.e., attention weights) among different samples in the network state sequence, thereby enhancing the model's perception ability of the important parts in the input network state sequence. The output of the self-attention mechanism module is input into the GRU network. H R , and calculate the query matrix, key matrix, and value matrix according to the following formula:
[0078] ;
[0079] ;
[0080] ;
[0081] where Q, K, and V are the query matrix, key matrix, and value matrix respectively. W Q , W K , W V are the weights of the query matrix, key matrix, and value matrix respectively.
[0082] Calculate the attention weight of each unit and generate the attention matrix H α as follows:
[0083] ;
[0084] ;
[0085] where M represents the dimension of the weight matrix, is the scaling factor, represents the attention weight of the i th unit; represents the transpose of the K th key vector in the key matrix i , q i represents the i th query vector in the query matrix Q, v i represents the
[0086] Finally, calculate the Hadamard product of H α and H R , and perform a weighted average operation on the resulting matrix in the l dimension to obtain the output of the self-attention mechanism module As follows:
[0087] ;
[0088] In the formula, W R is the weight matrix, represents the weighted average operation; represents the eigenvector in the output of the self-attention mechanism module ;
[0089] The spatial feature extraction module is constructed based on the GAT network. The GAT network can focus on the adjacent node information with different weights and has excellent spatial information capture ability. Therefore, the present invention uses the GAT network to extract the spatial correlation features of the link state information in the complex communication network. Since the present invention studies the link state characteristics in the communication network, the links are corresponding to the nodes in the graph structure to achieve feature extraction. The GAT network takes the element in the output of the self-attention mechanism module as the node feature of the i th node in the graph structure, and defines c ij as the attention coefficient of node j to node i as follows:
[0090] ;
[0091] Among them, represents the adjacency matrix, A ij represents the connectivity of node i and node j in the graph structure. If there is an edge connecting node i and node j , then A ij = 1, is the transformation function, and the symbol represents vector concatenation, W e is the weight matrix, ;
[0092] Introduce the softmax function to normalize the attention coefficient c ij as follows:
[0093] ;
[0094] Among them, N i represents the first-order neighbor nodes of node i , aij is the attention coefficient after standardization;
[0095] Use the attention coefficient after standardization to linearly accumulate the neighborhood representation of node i to obtain the final output feature of this node:
[0096] ;
[0097] where σ represents the non-linear activation function, represents node i 's final output feature;
[0098] Since single-head attention has the defect of instability in the learning process, multi-head attention is introduced to improve the representation ability of the model and enhance the algorithm stability. Specifically, use K to represent the number of heads, perform the above operations on each head, splice the final results, then average the spliced results, and delay the use of the non-linear function to obtain the final representation. The multi-head attention is as follows:
[0099] ;
[0100] In the formula, K represents the number of heads; represents the k th attention head, the standardized attention coefficient of node j to node i , is the final output feature of node i after multi-head attention; σ is the non-linear activation function;
[0101] GRU network output characteristic matrix , the specific form is as follows:
[0102] ;
[0103] The multi-layer perceptron module realizes the model output function. Input the characteristic matrix output by the GRU network into the multi-layer perceptron (MLP). Each layer in the MLP consists of multiple neurons. The layers are fully connected through the weight matrix and the bias vector, and use softmax function for processing, and finally output the path selection action t at time a t :
[0104] ;
[0105] Path selection action , where , is the number of neurons in the output layer.
[0106] Step S3: Extract samples from the experience replay pool, iteratively train the intelligent path decision-making model, and update the path selection strategy until the intelligent path decision-making model converges. For the target network at the current moment, use the converged intelligent path decision-making model to generate and store the path selection strategy, and take the path between the target network node pairs corresponding to the path selection strategy as the optimal path;
[0107] The advantage function is used in the update stage of the path selection strategy to measure the quality of each path selection action. The advantage function is defined as follows:
[0108]
[0109] where r t is the reward obtained by executing the path selection action s t under the network state sequence a t , represents the reward discount factor, and respectively represent the evaluation values of the network state sequence s t at the current moment and the network state sequence s t+1 at the next moment;
[0110] Due to the inaccurate estimation of the advantage function by the traditional policy gradient update algorithm, the execution policy of the intelligent agent will deviate seriously. Therefore, the present invention adopts the importance sampling method to adjust the policy update amplitude, improve the efficiency of policy update and sample utilization in the training process, and adopts the importance sampling method to adjust the update amplitude of the path selection strategy. The importance sampling is as follows:
[0111]
[0112] where represents the ratio of the probability of the current path selection policy π θ taking the path selection action s t under the network state sequence a t to the probability of the old path selection policy taking the path selection action s t under the network state sequence a t ;
[0113] To better adapt to the dynamic changes of the network topology and the characteristics of uneven sample distribution, the present invention adopts the gradient clipping method shown below to define the objective function , so as to limit the update amplitude of the path selection strategy:
[0114] ;
[0115] In the formula, E t represents the average value of evaluating the expression in the brackets over multiple time steps;
[0116] The parameter update of the path selection strategy is as follows. While limiting the update amplitude of the path selection strategy, it maximizes the expected cumulative revenue to improve the convergence and stability of the algorithm:
[0117] ;
[0118] Among them, represents the updated path selection strategy parameter; clip represents the clipping operation, ε represents the clipping factor, usually represented by a small positive number, which limits the update amplitude of the path selection strategy within the range of [1 - ε, 1 + ε].
[0119] Step S4: Generate a flow table based on the optimal path and send it to the switch device of the target network for path installation and service data forwarding.
[0120] The embodiment of the present invention also provides a path intelligent optimization system based on enhanced link state awareness. Referring to Figure 4 , it includes three-layer structures: a data layer, a control layer, and an application layer, to implement the described path intelligent optimization method based on enhanced link state awareness:
[0121] The data layer contains various routing nodes and communication links of the target network, and transmits the topology information and port status information of the target network to the control layer through the southbound interface. At the same time, it receives the path selection strategy issued by the control layer and completes the processing and forwarding operations of service data;
[0122] The control layer includes five modules: a network awareness module, a network monitoring module, a data processing module, an intelligent optimization module, and a path installation module;
[0123] The control layer periodically sends preset request instructions to the data layer through the southbound interface to obtain the topology information and port status information of the target network in real time, and transmits the path selection strategy to the data layer;
[0124] Among them, the network perception module periodically sends feature request instructions to the data layer to obtain the topology information of the target network; the network monitoring module periodically sends status request instructions to the data layer and asynchronously receives status reply messages to obtain the port status information of the routing nodes in the target network; the data processing module uses the topology information and port status information collected by the network perception module and the network monitoring module to calculate the link status information and end-to-end path status information, and then statistics the global network average throughput, average end-to-end delay and average packet loss rate and stores them; the intelligent optimization module constructs an intelligent path decision model, and according to the network state sequence at the current moment s t , performs a path selection action a t , obtains the network state sequence at the next moment s t+1 , and at the same time obtains the reward at the current moment r t ; the path installation module generates corresponding flow tables according to the path selection action a t and distributes them to the data layer for forwarding service data;
[0125] The application layer includes various services and applications of the target network, and conducts information interaction with the control layer through the northbound interface.
[0126] The following is an application embodiment of the present invention:
[0127] In this embodiment, the command and control network is used as the target network, Figure 5It is a diagram of the command and control network topology and the set damaged links. The numbers 1 to 47 in the figure are node numbers. The command and control network is an integrated tactical communication network for joint operations, including 47 nodes and 61 links. Among them, nodes 18, 19, 20 and nodes 32, 33, 34 simulate sensor nodes, nodes 42, 43, 44 simulate command and control nodes, nodes 11, 12, 13, nodes 25, 26, 27 and nodes 39, 40, 41 simulate fire strike nodes, and the traffic transmission path follows the principle of "sensor - command and control - fire strike". There are more than 10 heterogeneous links in the tactical communication network. To simulate the impact of a strong confrontation environment on the link state, experiments are carried out by setting different link bandwidths in the Mininet simulation software. The experiment sets 4 different intensities of traffic for testing, namely low intensity (flow transmission rate is 25 kbps, 50 kbps) and high intensity (flow transmission rate is 75 kbps, 100 kbps). For each traffic intensity, an Iperf script is written to achieve one-to-one or one-to-many traffic transmission from the "sensor node - fire strike node". To verify the adaptability of the DRL-SGA algorithm designed in the present invention to the dynamically changing network topology, as marked in the figure, some links in the backbone network are disconnected in sequence to simulate the link damage scenario, and experiments are carried out respectively under the low-intensity traffic of 50 kbps and the high-intensity traffic of 100 kbps, and the network performance indicators of each routing algorithm are statistically analyzed and compared.
[0128] To verify the advantages of the algorithm of the present invention, the following comparison algorithms are adopted:
[0129] (1) OSPF: Open Shortest Path First algorithm, which obtains the weight information of each link in the network through the SDN measurement mechanism and calculates the path with the shortest link weight.
[0130] (2) DQN: A deep reinforcement learning routing algorithm based on traditional DQN. The agent performs perception training according to the link state information, and the reward function is set the same as that of the algorithm of the present invention.
[0131] (3) DDPG: A deep reinforcement learning routing algorithm based on traditional DDPG. The agent adopts a fully connected feedforward neural network structure, interacts with the network environment, and learns the routing strategy using the link state information. The reward function is set the same as that of the algorithm of the present invention.
[0132] (4) DRL-ST: An intelligent path optimization algorithm based on reinforcement learning. It uses Dueling DQN to construct an end-to-end transmission path decision model, optimizes the sampling mechanism using SumTree, and makes routing decisions according to the link state information. The reward function is set the same as that of the algorithm of the present invention.
[0133] Figure 6(a) - Figure 6(c)It is a comparison chart of the average network throughput, average end-to-end delay, and average packet loss rate of five algorithms under different traffic intensities. Figure 6(a) shows the comparison of the average network throughput under different traffic intensities. In the low-intensity traffic scenario, the advantage of the DRL-SGA algorithm in improving the network throughput compared to other algorithms is not obvious. The reason is that the service traffic intensity is small and the link bandwidth resources are sufficient, resulting in no significant network congestion. Therefore, the throughput indicators of each algorithm do not differ much. However, as the traffic intensity gradually increases, the throughput improvement effect of the DRL-SGA algorithm is more obvious compared to other algorithms. In the high-intensity traffic scenario, the network throughput of the DRL-SGA algorithm compared to other algorithms always remains at a relatively high level. When the traffic intensity reaches 100 kbps, the throughput is maximally increased by 23.48% compared to the OSPF algorithm and maximally increased by 2.59% compared to the better-performing DRL-ST algorithm. This indicates that the DRL-SGA algorithm can formulate a more optimal routing strategy according to the network load situation and link state changes. Figure 6(b) shows the comparison of the average end-to-end delay under different traffic intensities. The average end-to-end delay of the DRL-SGA algorithm under different traffic intensities is lower than that of the compared routing algorithms, and as the traffic intensity continues to increase, the DRL-SGA algorithm performs more excellently in ensuring the end-to-end delay performance. The average end-to-end delay of the DRL-SGA algorithm is minimally reduced by 14.42% and maximally reduced by 33.57% compared to the traditional routing algorithm OSPF; compared to the DQN algorithm, the average end-to-end delay is minimally reduced by 5.69% and maximally reduced by 25.04%; compared to the DDPG algorithm, the average end-to-end delay is minimally reduced by 7.08% and maximally reduced by 22.44%; compared to the better-performing DRL-ST algorithm, the average end-to-end delay is minimally reduced by 2.07% and maximally reduced by 16.88%. Figure 6(c) shows the comparison of the average packet loss rate under different traffic intensities. The average packet loss rate of the DRL-SGA algorithm under different traffic intensities is lower than that of the compared routing algorithms, and as the traffic intensity continues to increase, the change in the packet loss performance of the DRL-SGA algorithm is more stable. Compared to the traditional routing algorithm OSPF, the average packet loss rate of the DRL-SGA algorithm is reduced by at least 14.66%; compared to the intelligent algorithms DQN, DDPG, and DRL-ST, the average packet loss rate of the DRL-SGA algorithm is reduced by at least 9.03%, 8.73%, and 1.65% respectively.
[0134] Figure 7(a) - Figure 7(f)It is a comparison chart of the average network throughput, average end-to-end delay, and average packet loss rate of five algorithms under different topological structures. Figures 7(a) and 7(b) show the comparison results of the average network throughput of each algorithm when the traffic intensity is 50 kbps and 100 kbps respectively under the condition of link damage. As the number of damaged links increases, the network throughput of each algorithm gradually decreases, but the throughput index of the DRL-SGA algorithm is always the best. When 3 backbone network links are damaged, under the low-intensity traffic of 50 kbps, the average throughput of the DRL-SGA algorithm is 8.76% higher than that of the traditional routing algorithm OSPF, and is 5.34%, 4.25%, and 2.77% higher than those of the intelligent routing algorithms DQN, DDPG, and DRL-ST respectively; under the high-intensity traffic of 100 kbps, the average throughput of the DRL-SGA algorithm is 35.29%, 25.47%, 20.18%, and 4.39% higher than that of these four algorithms respectively. Figures 7(c) and 7(d) show the comparison results of the average end-to-end delay of each algorithm when the traffic intensity is 50 kbps and 100 kbps respectively under the condition of link damage. As the number of damaged links increases, the average end-to-end delay of each algorithm gradually increases. When the link is damaged, under the low-intensity traffic of 50 kbps, the average end-to-end delay of the DRL-SGA algorithm is at least 12.99%, 8.06%, 6.41%, and 3.15% lower than that of the other four routing algorithms respectively; under the high-intensity traffic of 100 kbps, compared with the other four routing algorithms, the average end-to-end delay of the DRL-SGA algorithm is at least 24.25%, 16.95%, 14.43, and 6.44% lower respectively. Figures 7(e) and 7(f) show the comparison results of the average packet loss rate of each algorithm when the traffic intensity is 50 kbps and 100 kbps respectively under the condition of link damage. As the number of damaged links increases, the average packet loss rate of each algorithm gradually increases. When the link is damaged, under the low-intensity traffic of 50 kbps, compared with the other four routing algorithms, the average packet loss rate of the DRL-SGA algorithm is at least 21.92%, 15.08%, 12.52%, and 1.65% lower respectively; under the high-intensity traffic of 100 kbps, compared with the other four routing algorithms, the average packet loss rate of the DRL-SGA algorithm is at least 35.19%, 26.44%, 21.16%, and 9.63% lower respectively.
[0135] The embodiments of the present invention have been described in detail above with reference to the accompanying drawings. However, the present invention is not limited to the above embodiments. Within the scope of knowledge possessed by those of ordinary skill in the art, various changes can be made without departing from the gist of the present invention.
Claims
1. A path intelligent optimization method based on link state perception enhancement, characterized in that: For the target network, the following steps S1 to S4 are executed to complete the path selection for forwarding service data in the target network: Step S1: Collect the topology information and port status information of the target network, calculate the communication link status information and end-to-end path status information of the target network, including the service request, remaining bandwidth, delay, and packet loss rate of the global network, constitute the network status of the target network, and calculate the average throughput, average end-to-end delay, and average packet loss rate of the global network; Step S2: The intelligent agent in the target network builds an intelligent path decision model based on the PPO algorithm of the AC architecture, and inputs the network state sequence obtained in step S1 into the intelligent agent. The intelligent agent s t , execute path selection action a t , and then obtain the network status sequence at the next moment s t+1 , and get the current moment's reward r t , and in the form of a quaternion ( s t , a t , r t , s t+1 ) constitute samples and store them in the experience replay pool; The specific steps of step S2 are as follows: Step S2.1: Agent Collection t Network status at all times x t Constructing the state space, t Network status at all times x t As follows: ; In the formula, , , , Respectively t The service requests, link remaining bandwidth, latency, and packet loss rate of the global network at the moment; f=4, indicating the characteristic dimension of the network status, and n indicates the total number of target network links; Among them, the business request of the global network The expression is as follows: ; In the formula, express t Nodes in the target network at the moment i With Node j Links e ij Business requests on m is the total number of nodes; The remaining bandwidth of the global network link The expression is as follows: ; In the formula, express t Nodes in the target network at the moment i With Node j Links e ij The remaining bandwidth; Global network latency The expression is as follows: ; In the formula, express t Nodes in the target network at the moment i With Node j Links e ij Delay Global network packet loss rate The expression is as follows: ; In the formula, express t Nodes in the target network at the moment i With Node j Links e ij Packet loss rate; use Min-Max Method , , , Normalize the elements in t Time and t Before l - 1 time step of network state, l is the preset network status sequence length, combined to form t Network status sequence at each moment s t as follows: ; Step S2.2: The agent follows the network state sequence s t Path selection actions taken a t Constructing the action space, assuming that the global network source-destination node pair contains k feasible paths, forming a feasible path set , where each path corresponds to a path weight, forming a path weight set ; Define each path selection action as ;in , , w ij Represents the source node i With the destination node j The path weight of path p is selected between ; Step S2.3: Agent collects and executes path selection actions a t The real-time network performance indicators after that are set, and the reward function is set to calculate the reward value, and the reward value is fed back to the agent. The reward function is as follows: ; in, r is the reward value, α , β , γ is the reward weight, and its value range is , , , Represents the normalized network performance index, where is the global network average throughput after normalization, is the normalized average end-to-end delay, is the average packet loss rate after normalization; Step S3: extract samples from the experience replay pool, iteratively train the intelligent path decision model, and update the path selection strategy until the intelligent path decision model reaches convergence. For the target network at the current moment, the converged intelligent path decision model is used to generate and store the path selection strategy, and the path between the target network node pairs corresponding to the path selection strategy is taken as the optimal path; Step S4: Generate a flow table based on the optimal path and send it to the switch device of the target network to perform path installation and service data forwarding.
2. According to claim 1, a path intelligent optimization method based on link state perception enhancement is characterized in that: The AC architecture based on the intelligent path decision model includes a policy network and an evaluation network. The policy network x t Output path selection action a t , the evaluation network outputs the evaluation value of the network state; The policy network includes a temporal feature extraction module, a self-attention mechanism module, a spatial feature extraction module, and a multi-layer perceptron module; The temporal feature extraction module is built based on the GRU network, including an update gate, a reset gate, and a hidden layer. The update gate is based on the network state at the current moment. x t and the hidden state at the previous moment h t-1 As input, the specific formula is as follows: ; In the formula, Z t To update the gate, W z To update the gate weight matrix, h t-1 is the hidden state at the previous moment, U z for h t-1 The update gate weight matrix of Reset the gate to the current network status x t and the hidden state at the previous moment h t-1 As input, the specific formula is as follows: ; In the formula, R t To reset the gate, W r To reset the gate weight matrix, U r for h t-1 Reset gate weight matrix of; According to the reset memory information formula, calculate the candidate hidden state at the current moment : ; In the formula, is the weight matrix, for h t-1 The weight matrix of represents the Hadamard product; The hidden state at the current moment h t As follows: ; Since the network state sequence input by the agent s t Include l time steps, so the GRU network contains l units, the output of the GRU network is as follows: ; In the formula, H R represents the output of the GRU network, Indicates the first i The hidden layer output of units, ; f represents the characteristic dimension of the network status, and n represents the total number of target network links; The self-attention mechanism module inputs the output of the GRU network H R , the query matrix, key matrix and value matrix are calculated according to the following formula: ; ; ; Among them, Q, K, and V are query matrix, key matrix, and value matrix respectively. W Q , W K , W V are the weights of the query matrix, key matrix, and value matrix respectively; Calculate the attention weight of each unit and generate the attention matrix H α As follows: ; ; Among them, M represents the dimension of the weight matrix, is the scaling factor, Indicates i The attention weight of each unit; Represents the key matrix K Middle i The transpose of the key vector, q i Indicates the query matrix Q i query vector, v i Represents the i-th value vector in the value matrix V; The output of the self-attention mechanism module As follows: ; In the formula, W R is the weight matrix, represents weighted average operation; Represents the output of the self-attention mechanism module The eigenvectors in ; The spatial feature extraction module is built on the GAT network, which takes the elements in the output of the self-attention mechanism module As a graph structure i The node characteristics of the nodes are defined as c ij For Node j For Node i The attention coefficient is as follows: ; in, represents the adjacency matrix, A ij Represents a node in a graph structure i and nodes j connectivity, is the conversion function, symbol represents vector concatenation, W e is the weight matrix, ; Introduction softmax Function for attention coefficient c ij Standardization is performed as follows: ; in, N i Representation Node i The first-order neighbor nodes of a ij is the standardized attention coefficient; Use the standardized attention coefficient to i The final output feature of the node is obtained by linear accumulation of the neighborhood representation: ; Among them, σ represents the nonlinear activation function, Representation Node i The final output features of The introduction of multi-head attention is as follows: ; In the formula, K Indicates the number of heads; Indicates k In an attention head, the node j For Node i The standardized attention coefficient, is the node after multi-head attention i The final output feature of; σ is a nonlinear activation function; GRU network output feature matrix , the specific form is as follows: ; The feature matrix output by the multilayer perceptron module and the GRU network For input, output t Path selection action at the moment a t : ; Path selection action ,in , is the number of neurons in the output layer.
3. The method for intelligent path optimization based on enhanced link state perception according to claim 1, characterized in that: In step S3, the advantage function is used in the update phase of the path selection strategy To measure the quality of each path selection action, the advantage function is defined as follows: ; in, r t In the network status sequence s t Execute path selection action a t Rewards, represents the reward discount factor, and Respectively represent the network status sequence at the current moment s t and the network state sequence at the next moment s t+1 the assessed value; The importance sampling method is used to adjust the update amplitude of the path selection strategy. The importance sampling is as follows: ; in, Indicates the current path selection strategy π θ In the network status sequence s t Take path selection action a t The probability of the old path selection strategy In the network status sequence s t Take path selection action a t The ratio of the probabilities of Define the objective function using the gradient clipping method As follows: ; In the formula, E t It means evaluating the average of the expression in brackets over multiple time steps; The parameter update of the path selection strategy is as follows: ; in, represents the updated path selection strategy parameters; clip Represents a clipping operation. ε Represents the crop factor.
4. A path intelligent optimization system based on enhanced link state perception, characterized in that: The invention comprises a three-layer structure of a data layer, a control layer and an application layer, so as to realize a path intelligent optimization method based on enhanced link state perception as described in any one of claims 1 to 3: The data layer includes various routing nodes and communication links of the target network, transmits the topology information and port status information of the target network to the control layer through the southbound interface, receives the path selection strategy issued by the control layer, and completes the processing and forwarding operations of the business data; The control layer includes five modules: network perception module, network monitoring module, data processing module, intelligent optimization module and path installation module; The control layer periodically sends preset request instructions to the data layer through the southbound interface, obtains the topology information and port status information of the target network in real time, and passes the path selection strategy to the data layer; Among them, the network perception module periodically sends feature request instructions to the data layer to obtain the topology information of the target network; the network monitoring module periodically sends status request instructions to the data layer and asynchronously receives status reply messages to obtain the port status information of the routing nodes in the target network; the data processing module uses the topology information and port status information collected by the network perception module and the network monitoring module to calculate the link status information and end-to-end path status information, and then calculates and stores the global network average throughput, average end-to-end delay and average packet loss rate; the intelligent optimization module constructs an intelligent path decision model, based on the network status sequence at the current moment s t , execute path selection action a t , get the network status sequence at the next moment s t+1 , and get the current moment's reward r t ; The path installation module selects actions based on the path a t Generate the corresponding flow table and send it to the data layer for forwarding business data; The application layer includes various services and applications of the target network, and exchanges information with the control layer through the northbound interface.
Citation Information
Patent Citations
Dynamic topology network intelligent routing method
CN114051272A
KR20230090200A