SDN routing method and architecture based on graph attention mechanism and dynamic priority playback

The GAT-DPR SDN routing method addresses the challenges of dynamic tactical networks by leveraging graph attention and dynamic prioritization, achieving improved throughput and reduced delay and packet loss.

CN120321167AActive Publication Date: 2025-07-15NANJING UNIV OF INFORMATION SCI & TECH

Patent Information

Application Number
CN202510796335.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-16
Publication Date
2025-07-15
Estimated Expiration
2045-06-16

AI Technical Summary

Technical Problem

Existing routing algorithms are difficult to cope with frequent changes in network topology in tactical communication networks. Deep reinforcement learning routing algorithms have poor unstructured data representation capabilities for network topology, and insufficient priority adjustment in experience playback leads to inefficient experience playback.

Method used

Using the SDN routing method based on graph attention mechanism and dynamic priority playback, the graph neural network deep reinforcement learning model (GAT-DPR) is constructed, and the graph attention network is used to enhance node and link feature learning, and combining dynamic priority experience playback and priority weighted sampling mechanism to optimize routing decisions.

Benefits of technology

It improves the throughput of tactical communication networks, reduces end-to-end delay and packet loss rate, and enhances the algorithm's convergence speed and routing decision efficiency in complex and changing networks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120321167A_ABST
    Figure CN120321167A_ABST
Patent Text Reader

Abstract

The invention discloses an SDN (Software Defined Network) routing method and architecture based on a graph attention mechanism and dynamic priority playback. The method comprises the following steps: acquiring network equipment and link states in real time; generating an alternative path set based on a shortest path algorithm; constructing a GAT-DPR model structure, taking the state information of the alternative paths as input, training the model structure, fusing a graph attention mechanism and a dynamic priority playback technology, optimizing the training efficiency of a strategy network, and obtaining an optimal routing strategy; the optimal routing strategy is converted into a flow table to be issued to a data layer, and flow forwarding is completed; and a network topology sensing module, a network state detection module, an initial path calculation module, a path optimization module and a path installation module in the control layer complete corresponding operation steps. Compared with an existing algorithm, the GAT-DPR algorithm provided by the invention has obvious advantages in the aspects of improving throughput and reducing time delay and packet loss rate, and an innovative solution is provided for efficient routing decision-making of a tactical communication network.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of information engineering, and particularly relates to an SDN routing method and architecture based on a graph attention mechanism and dynamic priority replay. Background Art

[0002] With the accelerated evolution of the information warfare form from network-centric warfare to intelligent warfare, higher requirements are put forward for the efficient transmission and reliable sharing of battlefield information, and routing decision-making is an important technology to optimize this goal. By improving the routing strategy to optimize the transmission path of battlefield information and achieve stable and reliable transmission of battlefield information, it can not only ensure the efficient cooperation between various combat units, but also enhance the commander's comprehensive perception and dynamic control ability of the battlefield situation. Currently, routing algorithms are mainly divided into traditional routing algorithms, heuristic routing algorithms, and deep reinforcement learning routing algorithms.

[0003] Traditional routing algorithms such as the Shortest Path First (SPF) algorithm, the Bellman-Ford (BF) algorithm, and the Equal Cost Multipath (ECMP) algorithm are applicable to static and stable civilian communication networks and are difficult to adapt to the high dynamics and complexity of tactical communication networks, resulting in rigid resource allocation and large fluctuations in end-to-end delay.

[0004] To cope with dynamic network requirements, heuristic algorithms such as the greedy algorithm (GA), the simulated annealing algorithm (SA), and the ant colony optimization algorithm (ACO) generate approximate optimal solutions through local search. However, heuristic algorithms rely on the determined network state and are prone to falling into local optima and lacking the ability to perceive global topological features when facing tactical communication networks with strong uncertainty factors and high complexity and dimensionality.

[0005] In recent years, deep reinforcement learning routing algorithms such as Deep Q Network (DQN), Proximal Policy Optimization (PPO), and Deep Deterministic Policy Gradient (DDPG) can better adapt to dynamically changing tactical communication networks by training agents to continuously interact with the environment and storing combat experience in the experience pool. Although deep learning routing algorithms show potential in dynamic networks, their main drawbacks are as follows: it is difficult to effectively model non-Euclidean topological relationships using traditional neural networks, resulting in poor generalization ability for node and link features; the experience pool uses random experience replay technology, ignoring the priority of key combat experience.

[0006] Therefore, the challenges faced by existing routing algorithms in tactical communication networks are as follows: it is difficult to cope with the frequent changes in network topology; the deep reinforcement learning routing algorithms have poor ability to represent unstructured data of network topology, resulting in uneven communication delays between nodes; insufficient priority adjustment in experience replay leads to inefficient experience replay. Summary of the Invention

[0007] Aiming at the problems of high dynamicity of node traffic load and strong time-variability of link state in tactical communication networks, such as large routing calculation overhead, long forwarding strategy period, and poor timeliness of replay pool samples, an SDN (Software-Defined Networking) routing method and architecture based on Graph Attention Networks with Dynamic Prioritization Replay (GAT-DPR) are proposed.

[0008] To solve the above problems, the present invention adopts the following technical solutions: An SDN routing method based on graph attention mechanism and dynamic priority replay, used for the transmission of information in tactical communication networks, includes the following steps: Step 1: Use a network simulation tool to construct a virtual tactical communication network topology, including defining terminal nodes and forwarding nodes, establishing logical connection links between node devices, and configuring parameters of links between node devices, including bandwidth and delay; Step 2: Periodically collect network topology data of the data plane based on an event-driven topology state monitoring mechanism, including the number of nodes, the number of links, and connection relationships, and construct a network topology graph; Step 3: Periodically collect and calculate data of each forwarding node device port in the data plane based on the southbound interface protocol, including the transmission rate and reception rate of the forwarding node device port, to obtain network status information; Step 4: Based on the network topology map and network status information, use the shortest path algorithm to generate an alternative forwarding path set for each pair of source-destination nodes in the tactical communication network, and calculate the performance metrics of each forwarding path, including the remaining link bandwidth, delay, packet loss rate, and the variance of the feasible path node load, and output a structured path description file; Step 5: Construct a deep reinforcement learning model structure based on graph neural network (GAT-DPR model structure), including a graph attention network, a policy network, a target network, and an experience replay pool; Use the structured path description file as the input, aggregate the node and link features in the network topology graph based on the graph attention network to obtain the global feature vector of the network topology graph; Input the global feature vector of the network topology graph into the policy network, output the action value of the current network state, calculate the action with the highest value corresponding to the current network state, interact with the network environment to obtain the next network state and the reward value at the current moment, form a quadruple from the current network state, the action with the highest value corresponding to the current network state, the next network state, and the reward value at the current moment, and store it as a sample in the experience replay pool; According to the dynamic priority replay and sampling mechanism, iteratively optimize the policy network parameters, and periodically synchronize the policy network parameters to the target network until the reward value reaches the maximum, and obtain the optimal routing policy in the current network state and output it; Use the structured path description file as the input and the optimized routing policy as the output to train the GAT-DPR model structure to obtain the trained GAT-DPR model; Step 6: Use the structured path description file of the tactical communication network to be optimized as the input, apply the trained GAT-DPR model to obtain the optimized routing policy; Generate a flow table from the optimized routing policy and send it to the forwarding node device in the data plane to finally complete the forwarding of traffic.

[0009] Furthermore, the graph attention network encodes the status information of the input alternative forwarding path into a node feature vector through a linear transformation, calculates the attention weights between the node and its neighbor nodes through the self-attention mechanism, then generates the features of each node after weighted aggregation of the neighbor node features, and finally aggregates all the node features to obtain the global feature vector of the network topology graph as the input of the policy network; Based on the global feature vector of the network topology graph, the policy network outputs the action value of the current network state through the fully connected layer of the policy network, calculates the action with the highest value corresponding to the current network state, interacts with the network environment to obtain the next network state and the reward value at the current moment; Form a quadruple from the current network state, the action with the highest value, the next network state, and the reward value at the current moment, and store it as a sample in the experience replay pool; Until the number of stored samples reaches the capacity threshold of the experience replay pool, samples of a preset batch size are randomly drawn. Using the drawn samples as the input for the policy network and the target network, the policy network and the target network are combined to calculate the temporal difference (TD) error of each sample. An attenuation factor is introduced to dynamically adjust the priorities of the samples in the experience replay pool. Based on the priority weights, the sample sampling probability is adjusted. The policy network parameters are updated based on the mean squared error loss; periodically synchronize the policy network parameters to the target network until the model structure converges, that is, the reward value reaches the maximum, and output the optimal routing policy.

[0010] The present invention also protects an architecture for implementing the above SDN routing method, including an application layer, a control layer, and a data layer; the application layer is connected to the control layer through a northbound interface; the control layer mainly includes a network topology awareness module, a network status detection module, an initial path calculation module, a path optimization module, and a path installation module, and the network topology awareness module, the network status detection module, the initial path calculation module, the path optimization module, and the path installation module respectively execute steps 2, 3, 4, 5, and 6 in the SDN routing method; the data layer constructs a virtual network topology through a network simulation tool and is connected to the control layer through a southbound interface to receive various control policies issued by the control layer.

[0011] Compared with the prior art, the beneficial effects of the present invention are as follows: The SDN routing method based on graph attention mechanism and dynamic priority replay provided by the present invention, firstly, uses a graph attention neural network to replace the traditional neural network in the reinforcement learning model to enhance the model's learning and representation capabilities of node and edge features in the network; secondly, uses a dynamic priority experience replay and priority weighted sampling mechanism to weaken the correlation between data in the experience pool and eliminate the bias caused by uneven sampling. This improvement can optimize the problem of insufficient adjustment of sample priorities in the experience replay pool and also ensure that in a complex and changeable tactical communication network, the algorithm can converge faster and make routing decisions efficiently; compared with the OSPF, DQN, and Dueling DQN algorithms, the GAT-DPR algorithm of the present invention increases the throughput by 24.28%, 13.12%, and 7.17% respectively, reduces the end-to-end delay by 37.21%, 26.49%, and 23.02% respectively, and reduces the packet loss rate by 13.16%, 8.08%, and 7.62% respectively. Description of the Drawings

[0012] Figure 1 is a flowchart of the method provided by the present invention.

[0013] Figure 2 is an overall SDN architecture diagram.

[0014] Figure 3 is a schematic diagram of the GAT-DPR algorithm.

[0015] Figure 4 It is the schematic diagram of the GAT module extracting the features of the network topology graph.

[0016] Figure 5 It is the ablation experiment diagram of the GAT-DPR algorithm.

[0017] Figure 6 It is the hyperparameter training experiment diagram of the GAT-DPR algorithm.

[0018] Figure 7 It is the real tactical communication network topology graph of the simulation of the present invention.

[0019] Figure 8 It is the performance analysis diagram of the GAT-DPR algorithm. Detailed implementation manners

[0020] The technical solution of the present invention will be further described in detail below in conjunction with the accompanying drawings and specific implementation manners.

[0021] As Figure 1 shown, the SDN routing method based on graph attention mechanism and dynamic priority replay provided by the present invention is used for the transmission of information in a tactical communication network, and includes the following steps: Step 1: Use the mininet network simulation tool to construct a virtual network topology for tactical communication, including defining host nodes, switch nodes, and router nodes, establishing logical connection links between node devices, and configuring parameters such as bandwidth and delay for the links between node devices; represent the constructed network topology as a graph , where is the set of nodes, representing switches and routers in the network, is the set of edges, representing the links in the network topology; Step 2: Use event listeners such as EventSwitchEnter and EventLinkAdd to periodically call the get_switch() and get_link() functions to collect network topology data on the data plane, including the number of nodes, the number of links, and the connection relationship, and construct a network topology graph; Step 3: Based on the port statistics request of the OpenFlow protocol, through the OFPPortStatsRequest message and the response event EventOFPPortStatsReply, periodically collect and calculate the data of each switch port in the data plane to obtain network status information; Step 4: Based on the network topology graph and network status information, use the dijskra shortest path algorithm to calculate a set of K alternative forwarding paths for each pair of source-destination nodes in the tactical communication network , and calculate the performance metrics of each alternative path, including the remaining link bandwidth, delay, packet loss rate, and the variance of the feasible path node load, and save them as a JSON file in the form of a dictionary; Step 5. Construct a deep reinforcement learning model structure based on a graph neural network, namely the GAT-DPR model structure, including a graph attention network, a policy network, a target network, and an experience replay pool; use the JSON file as the input and the optimal routing policy as the output to train the constructed GAT-DPR model structure to obtain the trained GAT-DPR model; as Figure 3 and Figure 4 shown, the specific training process is as follows: Step 5.1 According to the state information of the input alternative path , use the graph attention network to perform a linear transformation on the input state to obtain the re-encoded node features ; calculate the attention weights between the node features and the neighbor node features through the self-attention mechanism; then use the attention weights to perform a weighted sum of the features of the neighbor nodes to obtain the aggregated node feature information ; fuse the neighbor aggregation information with the current node feature information and execute it T times to obtain the node features after message passing; aggregate and output the features of all nodes as the feature vector of the entire network topology graph; the specific calculation formula is as follows: , , , , , , , , , , , , , , , , where Indicates the source to destination node A set of candidate paths; Indicates the structure of each candidate path, consisting of the source node , passing through A number of intermediate links to reach the destination node ; 、 and respectively represent the state matrices storing the remaining bandwidth, delay, and packet loss rate of these paths; Indicates the traffic load weight matrix of the th node in the network; Indicates the adjacency relationship matrix of all nodes between the source-destination node pair; Indicates the remaining bandwidth of the link, determined by the minimum value of the bandwidth of all links in the path ; Indicates the link delay, calculated by summing the delays of all links in the path ; Indicates the link packet loss rate, obtained by calculating 1 minus the product of the successful transmission probabilities of all links; Indicates the traffic load weight of the th node, calculated by the ratio of the current traffic of this node to the historical maximum traffic ; Indicates the variance of the node load of the feasible path. The smaller the variance, the more evenly distributed the node load and the lower the congestion probability of the feasible path; Is the initial input feature, respectively are the feature representations of the th node and the th node, is the set of neighbor nodes of the th node, Indicates the weight matrix, represents the relu activation function, represents the feature concatenation operation, Indicates the attention weight of the neighbor nodes after the T-th message aggregation; Step 5.2 Input the feature vector of the extracted network topology graph into the policy network, and output the action value under the current network state , by calculating the action corresponding to the highest value in the current network state , interacting with the network environment to obtain the next moment network state and the reward value at the current moment , by Form quadruples and store them as samples in the experience replay pool; when the number of stored samples reaches the capacity threshold of the experience replay pool, randomly draw N samples from the experience replay pool ; Use the drawn samples as the input of the policy network and the target network, and use the temporal difference TD error function to obtain the error values of each sample , introduce a decay factor to dynamically adjust the priority of the samples ; Then use priority weighting to calculate the sample sampling probability , combine the mean squared error loss function to calculate the descent gradient , and through the backpropagation process of the neural network, update the policy network parameters according to the descent gradient; periodically synchronize the policy network parameters to the target network until the model structure converges, that is, the reward value reaches the maximum, and the path calculated at this time is the optimal forwarding path; the specific calculation formula is as follows: , , , , , , where represents the state value function, represents the advantage function, represents the average value of the advantage functions of all possible actions, represents the number of all possible actions, represents the parameters of the policy network; represents the TD error, represents the current moment 's immediate reward; represents the discount factor, which measures the importance of the current moment's immediate reward and future rewards; represents the initial priority of the sample, , represents the modified sample 、 sample 's priority weight; represents the decay factor, represents the number of time steps the sample is stored in the experience replay pool, represents the sample priority adjustment coefficient corresponding to the storage time step; represents the sample sampling probability after priority weighting; represents the Q value calculated by the policy function, represents the Q value calculated by the target function; denote the remaining bandwidth, delay, packet loss rate of the time link, and the variance of the load of the feasible path nodes denote the weight matrix

[0022] Step 6: Taking the structured path description file of the tactical communication network to be optimized as the input, applying the trained GAT-DPR model to obtain the optimized routing strategy; generating the openflow flow table from the optimized routing strategy and distributing it to each switch in the data plane to complete the forwarding of traffic

[0023] As Figure 2 shown, the architecture for implementing the above SDN routing method includes an application layer, a control layer, and a data layer; the application layer is connected to the control layer through a northbound interface; the control layer mainly includes a network topology awareness module, a network state detection module, an initial path calculation module, a path optimization module, and a path installation module, and the network topology awareness module, the network state detection module, the initial path calculation module, the path optimization module, and the path installation module respectively execute steps 2, 3, 4, 5, and 6 in the above SDN routing method; the data layer constructs a virtual tactical communication network topology through the mininet network simulation tool and is connected to the control layer through a southbound interface, receiving and executing each control policy issued by the control layer

[0024] Figure 5 is the ablation experiment on the attention mechanism designed by the present invention. The experiment shows that adding the attention mechanism to the model structure can ensure that the model has a faster convergence speed and a greater decline under the condition that other experimental conditions are the same

[0025] As Figure 6 shown in (6-1), (6-2), (6-3), (6-4), (6-5), (6-6), (6-7) of, the hyperparameter training results of the GAT-DPR algorithm are as follows: The optimal hyperparameters of the GAT-DPR algorithm are shown in Table 1 Table 1 Optimal Hyperparameters of the GAT-DPR Algorithm

[0026] Figure 7 is the real tactical communication network topology diagram adopted by the present invention, which is a complex tactical communication network composed of multiple link types including optical fiber, wired Ethernet, microwave link, and regional broadband network. The main difference between each link lies in the communication bandwidth. By setting different bandwidth parameters for different links in the mininet network simulation software, the simulation of the real network environment is realized

[0027] Figure 8 is based on the present invention Figure 7 ​The network topology structure, the performance comparison experimental results of four different routing algorithms under different traffic demand scenarios; as Figure 8 As shown in (8-1), (8-2), and (8-3) in Figure 8 , the GAT-DPR algorithm of the present invention is superior to the other three algorithms in the core indicators including delay, average throughput, and packet loss rate, verifying its robustness and efficiency under dynamic traffic loads.

[0028] The above is only the preferred specific implementation manner of the present invention and is not used to limit the present invention. Any modification, equivalent replacement, improvement, etc. made by any person skilled in the art within the technical scope disclosed by the present invention according to the technical solution and inventive concept of the present invention shall be covered within the protection scope of the present invention. Therefore, the protection scope of the present invention shall be subject to the protection scope of the claims.

Claims

1. An SDN routing method based on graph attention mechanism and dynamic priority replay is used for information transmission in a tactical communication network, and is characterized in that It includes the following steps: Step 1: Use a network simulation tool to construct a virtual tactical communication network topology, including defining terminal nodes and forwarding nodes, establishing logical connection links between node devices, and configuring parameters of the links between node devices, including bandwidth and delay; Step 2: Periodically collect network topology data of the data plane based on an event-driven topology state monitoring mechanism, including the number of nodes, the number of links, and connection relationships, and construct a network topology graph; Step 3: Periodically collect and calculate data of each forwarding node device port in the data plane based on the southbound interface protocol, including the transmission rate and reception rate of the forwarding node device port, to obtain network status information; Step 4: Based on the network topology map and network status information, use the shortest path algorithm to generate an alternative forwarding path set for each source-destination node pair in the tactical communication network, and calculate performance metrics of each forwarding path, including remaining link bandwidth, delay, packet loss rate, and variance of the load of feasible path nodes, and output a structured path description file; Step 5: Construct a GAT-DPR model structure, including a graph attention network, a policy network, a target network, and an experience replay pool; using the structured path description file as input, aggregate the node and link features in the network topology graph based on the graph attention network to obtain the global feature vector of the network topology graph; Input the global feature vector of the network topology graph into the policy network, output the action value of the network state at the current moment, calculate the action with the highest value corresponding to the network state at the current moment based on this, interact with the network environment to obtain the network state at the next moment and the reward value at the current moment, form a quadruple from the network state at the current moment, the action with the highest value corresponding to the network state at the current moment, the network state at the next moment, and the reward value at the current moment, and store it as a sample in the experience replay pool; According to the dynamic priority replay and sampling mechanism, iteratively optimize the parameters of the policy network, and periodically synchronize the parameters of the policy network to the target network until the reward value reaches the maximum, obtain the optimal routing policy in the current network state and output it; Use the structured path description file as input and the optimized routing policy as output to train the GAT-DPR model structure to obtain the trained GAT-DPR model; Step 6: Use the structured path description file of the tactical communication network to be optimized as input, apply the trained GAT-DPR model to obtain the optimized routing policy; generate a flow table from the optimized routing policy and send it to the forwarding node devices in the data plane to finally complete the forwarding of traffic.

2. The SDN routing method based on graph attention mechanism and dynamic priority replay according to claim 1, characterized in that The graph attention network encodes the state information of the input alternative forwarding path into a node feature vector through a linear transformation, calculates the attention weights between the node and its neighbor nodes through a self-attention mechanism, then generates the features of each node after weighted aggregation of the neighbor node features, and finally aggregates all node features to obtain the global feature vector of the network topology graph as the input of the policy network; Based on the global feature vector of the network topology graph, the policy network outputs the action value in the current network state through the fully connected layer of the policy network, calculates the action with the highest value corresponding to the current network state, and interacts with the network environment to obtain the next network state and the reward value at the current moment; A quadruple is formed by the network state at the current moment, the action with the highest value, the network state at the next moment, and the reward value at the current moment, and is stored as a sample in the experience replay pool; Until the number of stored samples reaches the capacity threshold of the experience replay pool, a preset batch size of samples is randomly drawn. The drawn samples are used as the input of the policy network and the target network. The policy network and the target network are combined to calculate the temporal difference TD error of each sample. The priority of the samples in the experience replay pool is dynamically adjusted by introducing a decay factor. The sample sampling probability is adjusted based on the priority weight, and the policy network parameters are updated based on the mean square error loss; the policy network parameters are periodically synchronized to the target network until the model structure converges, that is, the reward value reaches the maximum, and the optimal routing policy is output.

3. The architecture for implementing the SDN routing method based on graph attention mechanism and dynamic priority replay as claimed in claim 1, characterized in that It includes an application layer, a control layer, and a data layer; the application layer is connected to the control layer through a northbound interface; the control layer mainly includes a network topology awareness module, a network state detection module, an initial path calculation module, a path optimization module, and a path installation module, and the network topology awareness module, the network state detection module, the initial path calculation module, the path optimization module, and the path installation module respectively execute steps 2, 3, 4, 5, and 6 in the SDN routing method; the data layer constructs a virtual network topology through a network simulation tool and is connected to the control layer through a southbound interface to receive each control policy issued by the control layer.

Citation Information

Patent Citations

  • Construction method and application of distributed route planning model

    CN114697229A

  • SDN multipath routing method based on attention mechanism and deep reinforcement learning

    CN116170370A

  • Multi-agent collaborative decision-making method based on deep reinforcement learning under limited communication resources

    CN116456480A

  • Deterministic end-to-end slice traffic arrangement strategy based on edge graph attention

    CN116980298A

  • Intelligent routing system and method based on knowledge definition network and graph reinforcement learning

    CN118282918A

Cited By

  • Power robot hot-line work task planning method and system

    CN121635017A