QoS routing optimization method and system, computer and readable storage medium
Through the integrated graph neural network and deep reinforcement learning method, the routing protocol is optimized to adjust the routing strategy in real time in the dynamic network environment, solving the problem that routing protocols in the existing technology are difficult to adapt to dynamic network changes, and achieving more efficient network resource allocation and scheduling.
Patent Information
- Application Number
- CN202510485293.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-17
- Publication Date
- 2025-05-30
AI Technical Summary
Existing routing protocols are difficult to perceive network state changes in real time when dealing with dynamic and unpredictable network environments, cannot effectively adjust routing policies to meet quality of service (QoS) requirements, and lack robustness in the face of network link failures or topological changes, affecting network performance and reliability of data transmission.
The QoS routing optimization method of integrated graph neural network and deep reinforcement learning is adopted. By obtaining real-time network performance data sets, modeling the network topology as a graph structure, building state space and action space, selecting candidate paths and evaluating their performance value, and updating the routing table to optimize the path selection strategy.
It realizes adaptive adjustment of routing policies in a dynamic network environment, improve network performance and stability, optimize routing path selection, improve network throughput, reduce end-to-end delay, and enhance routing decision-making capabilities and resource utilization efficiency.
Smart Images

Figure CN120075126A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer technology, and particularly to a QoS routing optimization method, system, computer, and readable storage medium, specifically a QoS routing optimization method, system, computer, and readable storage medium integrating graph neural network and deep reinforcement learning.
Background Art
[0002] With the rapid development of emerging technologies such as 5G, Internet of Things (IoT), and edge computing, as well as the popularization of cloud computing services, network traffic has shown an explosive growth trend, and the requirements for quality of service (QoS) have become increasingly stringent. Especially in key application fields such as real-time communication, intelligent transportation systems, telemedicine, and industrial Internet, different service types have very different requirements for QoS parameters such as latency, jitter, and packet loss rate. In this context, network routing technology is facing unprecedented challenges:
[0003] 1) Dynamics of network traffic:
[0004] The widespread use of cloud computing services has led to an increase in the volatility of network traffic in terms of time and space.
[0005] The proliferation of IoT devices has made network connections more intensive and traffic patterns more diverse.
[0006] The popularity of mobile applications has made user behavior more unpredictable and network traffic changes more frequent.
[0007] 2) Complexity of network topology:
[0008] Modern network structures are becoming increasingly complex, including various network forms such as data centers, wide area networks, and local area networks.
[0009] There is a wide variety of network devices, including routers, switches, servers, etc., which increases the difficulty of network management and maintenance.
[0010] 3) Requirements for quality of service:
[0011] Users' requirements for network service quality (QoS) are constantly increasing, especially in key business scenarios such as financial transactions and healthcare.
[0012] Low latency, high throughput, and low packet loss rate have become important indicators for evaluating network performance.
[0013] Common routing protocols such as OSPF (Open Shortest Path First) and BGP (Border Gateway Protocol) have achieved success in many stable network environments, but they have limitations in dealing with highly dynamic and unpredictable network environments. These algorithms often struggle to perceive changes in network state in real time and are also difficult to make effective routing adjustments according to changes in services to meet QoS requirements:
[0014] 1) Static environment assumption:
[0015] Traditional routing protocols such as OSPF and BGP are based on the assumption of a static network environment and cannot quickly adapt to changes in network traffic and topology. When network conditions change, these protocols require manual intervention or reconfiguration, which is complex and has a long response time.
[0016] 2) Weak feature capture ability:
[0017] Although deep reinforcement learning (DRL) methods can continuously optimize routing strategies through interaction with the environment, in actual dynamic environments, business requirements change dynamically, and these algorithms are difficult to quickly capture global network features and often require a large amount of training to converge to the optimal routing strategy that meets the current network conditions.
[0018] 3) Lack of robustness:
[0019] Existing methods lack sufficient robustness when facing network link failures or topology changes, which may lead to a decline in network performance and affect the reliability and stability of data transmission. For example, when a critical link fails, traditional routing protocols may take a long time to find an alternative path, resulting in network interruption or a decline in service quality.
[0020] 4) Weak generalization ability:
[0021] Existing DRL solutions often struggle to generalize to unseen network topologies or traffic patterns and may not be able to effectively capture complex network structure information. This limits their performance in highly dynamic and unpredictable network environments.
[0022] To overcome the above problems, researchers have begun to explore the application of new methods combining machine learning and artificial intelligence in dynamic network routing. It is hoped that these methods can adaptively adjust routing strategies by learning network states and traffic patterns, improving network performance and stability.
[0023] A large number of existing studies have deeply explored dynamic network routing, mainly focusing on the following aspects:
[0024] Deep Reinforcement Learning (DRL) Method: The DRL method continuously optimizes the routing policy through interaction with the environment. However, in an actual dynamic environment, when network services change, these algorithms usually require a large amount of retraining to perceive the optimal routing policy under the current service.
[0025] Traditional Routing Protocols (such as OSPF and BGP): These protocols are based on the assumption of a static network environment and cannot quickly adapt to changes in network service requirements. When network conditions change, these protocols require manual intervention or reconfiguration, which is complex to operate and has a long response time.
[0026] Graph Neural Networks: Graph neural networks have become a research hotspot because they can effectively process graph-structured data. However, there is relatively little research on applying them to dynamic network routing.
[0027] Traditional routing algorithms and heuristic algorithms usually cannot adapt to rapid changes in network conditions, resulting in suboptimal decisions. Existing deep reinforcement learning-based methods have poor perception ability when network services change dynamically and usually require a large amount of training to converge. This process is both time-consuming and computationally intensive, making it impractical in real environments.
[0028] Therefore, it is necessary to study a QoS routing optimization method, system, computer, and readable storage medium to address the deficiencies of the existing technology and solve or mitigate one or more of the above problems.
Summary of the Invention
[0029] In view of this, the present invention provides a QoS routing optimization method, system, computer, and readable storage medium, which can find the best path that meets the QoS requirements for data streams in different service environments of a dynamic network environment.
[0030] On the one hand, the present invention provides a QoS routing optimization method integrating graph neural networks and deep reinforcement learning. The QoS routing optimization method includes the following steps:
[0031] S1: Obtain a real-time network performance data set, where the real-time network performance data set includes link bandwidth, latency, and packet loss rate;
[0032] S2: Model the network topology as a graph structure based on the real-time network performance data set to generate a network state feature matrix;
[0033] S3: Construct a state space and an action space according to the network state feature matrix. The action space includes multiple candidate paths from the source node to the target node;
[0034] S4: Select a candidate path from the action space through the main network, evaluate the performance value of the candidate path using the target network, update the main network according to the performance value, and generate an experience data set;
[0035] S5: Optimize the path selection strategy according to the empirical data set and update the routing table.
[0036] For the aspects and any possible implementation manners as described above, a further implementation manner is provided. The obtaining of the real-time network performance data set includes:
[0037] Collect link bandwidth, latency, and packet loss rate data through a network monitoring tool to generate an initial network performance data set;
[0038] Define low-latency and high-throughput metrics according to the video stream and file transfer service requirements;
[0039] Adopt a timed task mechanism to periodically collect network environment data and update the real-time network performance data set;
[0040] Verify the integrity of the real-time network performance data set through a data verification mechanism to obtain a verified real-time network performance data set.
[0041] For the aspects and any possible implementation manners as described above, a further implementation manner is provided. The modeling of the network topology as a graph structure based on the real-time network performance data set to generate a network state feature matrix includes:
[0042] Model the network topology as a graph structure, where nodes represent network devices, edges represent links, and the edge weights are the bandwidth, latency, and packet loss rate in the real-time network performance data set;
[0043] Learn the node and edge relationships through a graph neural network, and use an attention mechanism to calculate the edge weights related to the low-latency and high-throughput metrics to generate a set of network state feature vectors;
[0044] Aggregate the neighbor node feature vectors according to the network state feature vector set, update the node feature vectors, and generate the network state feature matrix.
[0045] For the aspects and any possible implementation manners as described above, a further implementation manner is provided. The selection of candidate paths from the action space by the main network and the evaluation of the performance value of the candidate paths by the target network includes:
[0046] The main network outputs an action value estimate according to the network state feature matrix and selects candidate paths from the action space;
[0047] Use the target network to calculate the target value estimate of the candidate path;
[0048] Adjust the parameters of the main network according to the difference between the action value estimate and the target value estimate to generate an empirical data set including states, actions, and values.
[0049] For the aspects and any possible implementation described above, a further implementation is provided. Optimizing the path selection strategy according to the empirical data set includes:
[0050] Sort the data in the empirical data set through a priority sorting algorithm, and the sorting basis is the absolute value of the temporal difference error;
[0051] Extract a batch of data from the sorted empirical data set according to the priority ratio based on the sampling batch size;
[0052] Weight according to the weight of each piece of experience in the batch data, and use the weighted batch data to update the parameters of the main network to optimize the path selection strategy.
[0053] For the aspects and any possible implementation described above, a further implementation is provided. The updating of the routing table includes:
[0054] Send the optimized path selection strategy to the network device through the network environment interaction module;
[0055] Generate device response data according to the network device's response to the path switching operation;
[0056] Update the routing table according to the device response data to generate a routing table update data set;
[0057] Extract the bandwidth, delay, and packet loss rate data of the actual path from the routing table update data set, compare it with the real-time network performance data set, and adjust the path selection strategy of the main network.
[0058] For the aspects and any possible implementation described above, a further implementation is provided. The QoS routing optimization system includes:
[0059] A real-time data acquisition module that acquires a real-time network performance data set, and the real-time network performance data set includes link bandwidth, delay, and packet loss rate;
[0060] A data modeling module that models the network topology as a graph structure based on the real-time network performance data set to generate a network state feature matrix;
[0061] A space construction module that constructs a state space and an action space according to the network state feature matrix, and the action space includes multiple candidate paths from the source node to the target node;
[0062] A performance value update module that selects candidate paths from the action space through the main network, evaluates the performance value of the candidate paths using the target network, updates the main network according to the performance value, and generates an empirical data set;
[0063] The routing update module optimizes the path selection strategy according to the experience data set and updates the routing table.
[0064] For the aspects and any possible implementation manners as described above, a further implementation manner is provided. The QoS routing optimization system is used for optimizing the network link path.
[0065] For the aspects and any possible implementation manners as described above, a further implementation manner is provided. A computer includes a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, the QoS routing optimization method as described above is implemented.
[0066] For the aspects and any possible implementation manners as described above, a further implementation manner is provided. A readable storage medium stores a computer program thereon. When the program is executed by a processor, the QoS routing optimization method as described above is implemented.
[0067] Compared with the prior art, the present invention can achieve the following technical effects:
[0068] The present invention discloses an intelligent routing optimization method based on a graph neural network and reinforcement learning. An heterogeneous graph network model is constructed through distributed topology discovery and edge traffic collection. Global state features are extracted by using graph convolution and self-attention mechanisms. A dynamic feature sequence is generated based on a temporal graph neural network and graph contrast learning. A reward function is designed by constructing a state space and an action space. Routing decisions are implemented by using a policy network and a value network. The training process is optimized through experience replay and a target network. The routing policy is dynamically adjusted according to topology changes and traffic prediction, realizing intelligent allocation and scheduling of network resources. The present invention can adapt to network topology changes, optimize routing path selection, improve network throughput, reduce end-to-end delay, and enhance routing decision-making ability and resource utilization efficiency in a complex network environment.
[0069] Of course, when implementing any product of the present invention, it is not necessarily required to achieve all the technical effects described above at the same time.
BRIEF DESCRIPTION OF THE DRAWINGS
[0070] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings required to be used in the embodiments will be briefly introduced below. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0071] Figure 1 It is a flowchart of the QoS routing optimization method provided by an embodiment of the present invention;
[0072] Figure 2It is the structural diagram of the QoS routing optimization system provided by an embodiment of the present invention.
Specific Embodiment
[0073] For a better understanding of the technical solution of the present invention, the embodiments of the present invention will be described in detail below with reference to the accompanying drawings.
[0074] It should be clear that the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative work belong to the scope of protection of the present invention.
[0075] The terms used in the embodiments of the present invention are only for the purpose of describing specific embodiments, and are not intended to limit the present invention. The singular forms of "a", "the", and "said" used in the embodiments of the present invention and the appended claims are also intended to include the plural forms, unless the context clearly indicates otherwise.
[0076] The present invention provides a QoS routing optimization method integrating graph neural network and deep reinforcement learning. The QoS routing optimization method includes the following steps:
[0077] S1: Obtain a real-time network performance data set, where the real-time network performance data set includes link bandwidth, delay, and packet loss rate;
[0078] S2: Model the network topology as a graph structure based on the real-time network performance data set to generate a network state feature matrix;
[0079] S3: Construct a state space and an action space according to the network state feature matrix. The action space includes multiple candidate paths from the source node to the target node;
[0080] S4: Select a candidate path from the action space through the main network, evaluate the performance value of the candidate path using the target network, update the main network according to the performance value, and generate an experience data set;
[0081] S5: Optimize the path selection strategy according to the experience data set and update the routing table.
[0082] The obtaining of the real-time network performance data set includes:
[0083] Collect link bandwidth, delay, and packet loss rate data through a network monitoring tool to generate an initial network performance data set;
[0084] Define low-latency and high-throughput metrics according to the video stream and file transfer service requirements;
[0085] Adopt a timed task mechanism to regularly collect network environment data and update the real-time network performance data set;
[0086] Verify the integrity of the real-time network performance dataset through a data verification mechanism to obtain a verified real-time network performance dataset.
[0087] Model the network topology as a graph structure based on the real-time network performance dataset to generate a network state feature matrix, including:
[0088] Model the network topology as a graph structure, where nodes represent network devices, edges represent links, and the edge weights are the bandwidth, latency, and packet loss rate in the real-time network performance dataset;
[0089] Learn the relationships between nodes and edges through a graph neural network, and use an attention mechanism to calculate the edge weights related to low-latency and high-throughput metrics to generate a set of network state feature vectors;
[0090] Aggregate the neighbor node feature vectors according to the network state feature vector set, update the node feature vectors, and generate the network state feature matrix.
[0091] The main network selects candidate paths from the action space, and the target network evaluates the performance value of the candidate paths, including:
[0092] The main network outputs an action value estimate according to the network state feature matrix and selects candidate paths from the action space;
[0093] Use the target network to calculate the target value estimate of the candidate path;
[0094] Adjust the parameters of the main network according to the difference between the action value estimate and the target value estimate to generate an experience dataset containing states, actions, and values.
[0095] Optimizing the path selection strategy according to the experience dataset includes:
[0096] Sort the data in the experience dataset through a priority sorting algorithm, and the sorting basis is the absolute value of the temporal difference error;
[0097] Extract a batch of data from the sorted experience dataset according to the priority ratio according to the sampling batch size;
[0098] Weight each experience in the batch data, and use the weighted batch data to update the parameters of the main network to optimize the path selection strategy.
[0099] Updating the routing table includes:
[0100] Send the optimized path selection strategy to the network device through the network environment interaction module;
[0101] Generate device response data according to the network device response path switching operation;
[0102] Update the routing table according to the device response data to generate a routing table update data set;
[0103] Extract the bandwidth, latency, and packet loss rate data of the actual path from the routing table update data set, compare it with the real-time network performance data set, and adjust the path selection strategy of the main network.
[0104] The QoS routing optimization system includes:
[0105] A real-time data acquisition module that acquires a real-time network performance data set, where the real-time network performance data set includes link bandwidth, latency, and packet loss rate;
[0106] A data modeling module that models the network topology as a graph structure based on the real-time network performance data set to generate a network state feature matrix;
[0107] A space construction module that constructs a state space and an action space according to the network state feature matrix, where the action space includes multiple candidate paths from the source node to the target node;
[0108] A performance value update module that selects a candidate path from the action space through the main network, evaluates the performance value of the candidate path using the target network, updates the main network according to the performance value, and generates an experience data set;
[0109] A routing update module that optimizes the path selection strategy according to the experience data set and updates the routing table.
[0110] The QoS routing optimization system is used for network link path optimization.
[0111] The present invention also provides a computer, including a memory, a processor, and a computer program stored on the memory and executable on the processor, where when the processor executes the computer program, it implements the QoS routing optimization method as described.
[0112] The present invention also provides a readable storage medium, on which a computer program is stored, and when the program is executed by a processor, it implements the QoS routing optimization method as described.
[0113] Embodiment 1:
[0114] As Figure 1 shown, the QoS routing optimization method and system of this embodiment specifically include:
[0115] Step S101, collect link bandwidth, latency, and packet loss rate data from the network environment to generate a network performance dataset. Define low-latency and high-throughput performance requirement metrics for video streaming and file transfer services, and update the real-time network performance dataset by regular collection.
[0116] Use a network monitoring tool to collect link bandwidth, latency, and packet loss rate data to generate an initial network performance dataset. If the data collection is completed, define low-latency and high-throughput performance requirement metrics according to the requirements of video streaming and file transfer services. Use a timed task mechanism to regularly collect network environment data and update the real-time network performance dataset. If the data update is completed, store the updated dataset. Use a data verification mechanism to check the data integrity. If the data is complete, proceed to the next step.
[0117] Step S102, model the network topology as a graph structure, where nodes represent network devices, edges represent links, and the edge weights are the bandwidth, latency, and packet loss rate in the real-time network performance dataset. Use a graph neural network to learn the relationships between nodes and edges, and calculate the edge weights related to low-latency and high-throughput metrics through an attention mechanism to generate a network state feature vector set.
[0118] Based on the network state feature vector set, aggregate the neighbor node feature vectors for each node, calculate the attention weights of neighbor nodes for this node, and update the node feature vectors according to the weights to generate a network state feature matrix.
[0119] Step S103, based on the network state feature vector set, aggregate the neighbor node feature vectors for each node, calculate the attention weights of neighbor nodes for this node, and update the node feature vectors according to the weights to generate a network state feature matrix.
[0120] Use a graph neural network to learn the relationships between the feature vectors of nodes and neighbor nodes. Calculate the attention weights of neighbor nodes for the current node through an attention mechanism, perform weighted aggregation on the neighbor node feature vectors according to the attention weights to obtain an aggregated feature vector, fuse the aggregated feature vector with the current node feature vector to update the node feature vector. Arrange the updated feature vectors of all nodes in order to generate a network state feature matrix. Associate the network state feature matrix with the low-latency and high-throughput metrics in the real-time network performance dataset to obtain an optimized network state feature matrix.
[0121] Step S104, construct a state space based on the network state feature matrix, and define an action space that includes multiple candidate paths from the source node to the target node. The candidate paths are generated from the real-time network performance dataset.
[0122] Establish the mapping relationship between the network state feature matrix and the topological structure, implement the development of a dynamic path evaluation function based on business requirements, develop a topology change-triggered path recalculation module, deploy a link quality evaluation component with a sliding window filter, construct a path encoder supporting multiple service types, design a joint representation method for the state space and the action space, integrate a path screening mechanism for link health detection, implement a double-buffered action space update service, and develop a state-action space dimension matching validator
[0123] Specifically, when constructing the state space based on the network state feature matrix, it is first necessary to extract key features from the real-time network performance dataset, such as link delay, bandwidth utilization, packet loss rate, etc., to form a multi-dimensional feature matrix. Suppose there are 10 nodes in the network and each link contains 3 features, then the dimension of the feature matrix is 10×10×3. The feature matrix is reduced in dimension through the principal component analysis (PCA) algorithm, and the first two principal components are retained to map the state space to a two-dimensional plane for subsequent analysis. The definition of the action space requires generating multiple candidate paths from the source node to the target node. For example, from node A to node F, the shortest path is calculated through the Dijkstra algorithm, and at the same time, the first 5 candidate paths are generated in combination with the K shortest path algorithm. The path selection criteria include minimum delay and maximum bandwidth. The network performance data of each path is obtained through real-time monitoring. For example, the delay of path 1 is 50ms and the bandwidth utilization is 70%, and the delay of path 2 is 60ms and the bandwidth utilization is 65%. Based on this data, the Q-learning algorithm is used to optimize the action space, with the learning rate set to 0.1 and the discount factor set to 0.9. By iteratively updating the Q-value table, the optimal path is selected. In each iteration, according to the current network state and the action selection result, the reward function is calculated. For example, the reward function can be defined as the weighted sum of the negative value of the delay and the bandwidth utilization, with the weights being 0.6 and 0.4 respectively. Through multiple iterations, the algorithm gradually converges, and finally the optimal path is determined. For example, the delay of path 3 is 55ms and the bandwidth utilization is 75%, and the comprehensive performance is the best. The whole process realizes a complete logical chain from state space construction to action space optimization, ensuring the efficient use of network resources.
[0124] Step S105, the main network selects a candidate path according to the network state feature matrix, evaluates the performance value of this path through the target network, updates the main network parameters based on the value difference, and generates an empirical dataset containing states, actions, and values.
[0125] Model the network topology as a graph structure, where nodes represent network devices, edges represent links, and edge weights are bandwidth, latency, and packet loss rate in the real-time network performance dataset. Use a graph neural network to learn the relationships between nodes and edges, calculate the edge weights related to low latency and high throughput metrics through an attention mechanism, generate a set of network state feature vectors. Based on the set of network state feature vectors, aggregate the neighbor node feature vectors for each node, calculate the attention weights of neighbor nodes for this node, update the node feature vectors according to the weights, generate a network state feature matrix. Based on the network state feature matrix, construct a state space, define an action space that includes multiple candidate paths from the source node to the target node, and the candidate paths are generated from the real-time network performance dataset. Use the main network to select a candidate path according to the network state feature matrix, evaluate the performance value of this path through the target network, update the main network parameters based on the value difference, generate an experience dataset containing states, actions, and values. Sample data from the experience dataset through a prioritized experience replay mechanism, update the main network parameters, and optimize the path selection strategy. The experience dataset is continuously filled by the experience dataset. Obtain the optimized path selection strategy from the main network, send the path selection information to the network device through the network environment interaction module, update the routing table according to the device response, generate a routing table update dataset, monitor the routing table update dataset, collect the feedback data of bandwidth, latency, and packet loss rate of the actual path, adjust the main network path selection strategy according to the feedback data, and update the experience dataset.
[0126] Step S106: Sample data from the experience dataset through a prioritized experience replay mechanism, update the main network parameters, and optimize the path selection strategy. The experience dataset is continuously filled by the experience dataset.
[0127] Use a prioritized sorting algorithm to sort the data in the experience dataset. The sorting basis is the absolute value of the temporal difference error. Extract a batch of data from the sorted experience dataset according to the priority ratio based on the sampling batch size. Calculate the weights of each experience in the batch data. The weights have a non-linear relationship with the priorities. Input the weighted batch data into the main network for forward propagation to obtain the action value estimation output by the main network. Calculate the target value estimation through the target network. Compare the difference between the action value estimation and the target value estimation. Adjust the parameters of the main network according to the difference value. If the target network update period arrives, copy the main network parameters to the target network.
[0128] Step S107: Obtain the optimized path selection strategy from the main network, send the path selection information to the network device through the network environment interaction module, update the routing table according to the device response, and generate a routing table update dataset.
[0129] The network environment interaction module is used to send the path selection information to the network device. The network device receives the path selection information and performs a path switching operation. The device responds to the path switching operation and generates device response data. The routing table is updated according to the device response data. The routing table update operation generates a routing table update data set. The bandwidth, latency, and packet loss rate data of the actual path are extracted from the routing table update data set. The bandwidth, latency, and packet loss rate data of the actual path are compared with the data in the real-time network performance data set. If there are differences between the bandwidth, latency, and packet loss rate data of the actual path and the data in the real-time network performance data set, the main network path selection strategy is adjusted and the experience data set is updated.
[0130] Specifically, when obtaining the optimized path selection strategy from the main network, path calculation based on the Dijkstra algorithm can be adopted. By calculating the shortest paths between various nodes in the network, an optimal path set is generated.
[0131] Step S108: Monitor the routing table update data set, collect the feedback data of the bandwidth, latency, and packet loss rate of the actual path, adjust the main network path selection strategy according to the feedback data, and update the experience data set.
[0132] The ipRouteTable object in the MIB library of the network device is polled using the SNMP protocol to obtain the OID values of each field of the routing table. If the device supports the NETCONF protocol, then call <get-config>Operate to retrieve routing configuration data, parse the response message in XML format, and extract routing entries for the IPv4 / v6 address family.
[0133] Specifically, during the process of monitoring the routing table update dataset, the system first periodically collects the routing table information of the router through the SNMP protocol, for example, once every 5 minutes, to ensure the real-time nature of the data. The collected data includes the destination network address, next-hop address, routing metric value, etc. For example, the destination network address is 192.168.1.0 / 24, the next-hop address is 192.168.2.1, and the routing metric value is 10. The system stores this data in the database and preprocesses it through the data analysis module, such as removing duplicate data and outliers, to ensure the data quality. Next, the system collects feedback data such as bandwidth, latency, and packet loss rate of the actual path through the ICMP protocol and iperf tool. For example, the bandwidth is 100Mbps, the latency is 20ms, and the packet loss rate is 0.1%. These data are processed through time series analysis algorithms, such as using the ARIMA model to predict the change trend of future path performance. According to the feedback data, the system adopts a reinforcement learning algorithm based on Q-learning to adjust the main network path selection strategy. For example, when the latency of a certain path exceeds 50ms, the system automatically switches to a path with lower latency. The updated path selection strategy is applied to the network in real time, and at the same time, the new experience data is stored in the experience dataset. For example, the data such as bandwidth, latency, and packet loss rate of the path 192.168.1.0 / 24 are updated to the database. The system optimizes the path selection strategy by periodically analyzing the experience dataset, for example, once a week for data analysis, to ensure the continuous improvement of network performance.
[0134] By capturing the characteristics of network performance metrics in different business scenarios, so as to select routes with different QoS requirements for different data streams, the present invention designs a graph neural network (GNN) embedded with an attention mechanism and integrates it into the routing optimization system architecture of a deep reinforcement learning (DRL) agent. These modules work together to adapt to the dynamic changes of the network environment, provide differentiated service quality guarantees for different business types, and ensure efficient QoS routing in a complex and changing network environment. The system architecture is as Figure 2 shown.
[0135] 1) Network environment interaction module
[0136] When running different network services in a network topology, the network environment interaction module is mainly responsible for collecting and processing data from network devices in real time, including service types, link status, traffic information, etc. By directly communicating with network devices such as routers, this module comprehensively understands the current network topology structure and the dynamic changes of service traffic in it. Specifically, the network environment interaction module obtains key performance indicators such as link bandwidth, latency, and packet loss rate from network devices, and regularly sends the latest network status information to other modules to ensure that the system can respond to network changes in a timely manner. It also updates the routing according to the routing policy generated by the intelligent routing decision module to ensure the immediacy and accuracy of routing decisions. In addition, this module is responsible for monitoring the effect of routing decisions, evaluating the current policy, and returning the results to the intelligent routing decision module in the form of rewards or punishments, so that the intelligent agent can learn more superior policies. In this way, the network environment interaction module acts as a bridge between the network environment and the QoS routing algorithm, ensuring the smooth flow of data streams and the efficient operation of the routing system.
[0137] 2) Network feature capture module
[0138] Use GNN to model the network and capture the dynamic changes of the network topology structure and different service performance indicators. To achieve this goal, the present invention introduces an attention mechanism in the graph neural network to identify the performance indicators that are most critical to a specific service type, providing accurate input information for intelligent routing decisions. The present invention uses GNN to learn the relationships between nodes and edges, as well as the network performance indicators loaded on links and node routers, to capture the complex characteristics of the network. Through the attention mechanism, it focuses on the performance indicators that are most critical to a specific service type, such as latency, jitter, packet loss rate, etc. This not only improves the generalization ability of the model but also better reflects the QoS constraints, providing differentiated services for different types of services.
[0139] 3) Intelligent routing decision module
[0140] Based on the DDQN algorithm of reinforcement learning, it formulates the optimal routing policy according to the information provided by the network state representation module. Specifically, the routing decision intelligent agent uses the network feature capture module designed based on the graph attention neural network as the Q-value prediction network, learns how to select the optimal path in the action space according to the QoS constraints under the current service, and adjusts the routing policy according to the real-time feedback generated by the interaction with the network environment interaction module.
[0141] The objective of the present invention is to find the optimal path that meets the QoS requirements for data streams in different service environments in a dynamic network environment. In the multi-path QoS optimization task, the present invention selects a routing strategy based on the DDQN (Double Deep Q-Network) algorithm to address the complex state space and dynamic characteristics of this task. The present invention defines the state space, action space, and reward function of the method, which are the key to the routing optimization strategy:
[0142] State space: The present invention defines the state of the network according to the link characteristics in the topology, and the state space will be part of the data in the input layer of the neural network. Specifically, the state space includes network performance metrics in the current network environment, such as delay, jitter, packet loss rate, bandwidth, etc. Different upper-layer services have different requirements for QoS services, which is the key for the graph attention network to capture service differences.
[0143] Action space: Considering that the task requires finding a balance between the shortest path and the best QoS constraints, and finding the routing selections for all source nodes and destination nodes will lead to a sharp reduction in computational efficiency, especially in the case of large-scale graph topologies. To avoid the high-dimensional calculations caused by estimating the Q values of all possible actions, and to take into account both the flexibility of routing selection and the comprehensiveness of evaluating all actions, the present invention limits the action space to the K shortest paths for each source-destination node pair, which not only reduces the computational cost but also ensures the consistency of the action space during the training and evaluation processes. In other words, this design ensures the universality of the model under different graph topologies.
[0144] Reward function: The objective of the present invention is to find the shortest route that meets the QoS constraints of the current service. Therefore, the reward function is set as the network performance metric including the QoS preference weights of the current service. The more the average performance metric on the path meets the preference of the current service, the greater the corresponding reward value.
[0145] Since the designed action space is discrete, DDQN provides an effective framework to handle this situation. In the traditional Q-learning method (DQN strategy), using the same network to select actions and evaluate actions may lead to overestimating the value of some actions, ultimately affecting the accuracy of learning. In the application of the present invention, the network state changes frequently, and the accurate evaluation of QoS parameters (such as delay, jitter, and packet loss rate) of different paths is crucial to avoid selecting paths with too high costs and to achieve service quality guarantee. DDQN significantly reduces the overestimation bias of the model by introducing a strategy of separating the target network from the main network. Specifically, the algorithm first uses the main network to select actions, and then the target network calculates the Q value of this action. This structure enables the present invention to perform more accurate value evaluation among multiple paths.
[0146] In addition, the feature space of the present invention is complex and high-dimensional, involving the state features of nodes and edges in the network graph. To accurately estimate the action value, the present invention designs a network feature capture module based on the Graph Attention Neural Network (GAT) and the Message Passing Neural Network (MPNN). This module considers the interaction between nodes (i.e., routers) and edges (i.e., links connecting routers) in the network. Each node feature vector contains network state information such as latency, jitter, packet loss rate, bandwidth, etc. For each node in the network, GAT can capture the importance of its neighbor nodes. In the current task, the importance can be represented by the correlation between the network state information on adjacent routers and the QoS constraints. GAT is used to calculate the attention weights of neighbor nodes for the current QoS service constraints, and MPNN updates the final output features of each node based on this. The module realizes receiving the features of the node as input and finally outputs the Q value representing the value of taking a certain action.
[0147] The present invention combines the prioritized experience replay mechanism into DDQN to improve the efficiency and stability of the learning process. In a network environment, since the event of a path that meets the QoS constraints appears with higher sparsity in the reduced action space, but once it occurs, it will have a significant impact on the final reward. In ordinary uniform experience replay, the sampling probability of each experience sample is the same. This setting may lead to some important information (such as rare but critical path selections) not being fully utilized, thus reducing the learning efficiency. In contrast, prioritized experience replay assigns a priority to each experience sample (the present invention uses the TD error), which ensures that important experiences can be updated more frequently. This strategy enables the agent to quickly adapt to complex and changing network conditions and efficiently learn more valuable strategies, thereby improving the overall performance. Once the model is loaded into a network environment where the upper-layer service requirements change frequently, DDQN and the prioritized experience replay strategy endow learning weights that match the current environmental state and task requirements by dynamically adjusting the priorities of experience samples. This flexibility enables the agent to quickly update the learning strategy and effectively adapt to the new network environment, ultimately achieving more efficient QoS routing selection.
[0148] The above has introduced in detail a QoS routing optimization method, system, computer, and readable storage medium provided by the embodiments of the present application. The description of the above embodiments is only used to help understand the method and its core idea of the present application; at the same time, for those of ordinary skill in the art, according to the idea of the present application, there will be changes in the specific implementation manner and application scope. In summary, the content of this specification should not be construed as a limitation to the present application.
[0149] As used in the specification and claims, certain terms are used to refer to specific components. Those skilled in the art should understand that hardware manufacturers may use different terms to refer to the same component. The specification and claims do not distinguish components by the difference in their names, but by the difference in their functions. As used throughout the specification and claims, the terms "comprising" and "including" are open-ended terms and should be interpreted as "comprising / including but not limited to". "Substantially" means within an acceptable error range. Those skilled in the art can solve the technical problems within a certain error range and basically achieve the technical effects. The following description in the specification is the preferred embodiment for implementing the present application, but the description is for the purpose of explaining the general principles of the present application and not for limiting the scope of the present application. The protection scope of the present application shall be subject to what is defined by the appended claims.
[0150] It should also be noted that the term "including", "comprising" or any other variant thereof is intended to cover non-exclusive inclusion, such that a good or system including a series of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such good or system. Without further limitation, an element defined by the statement "including a..." does not exclude the presence of another identical element in the good or system including the said element.
[0151] It should be understood that the term "and / or" used herein is only an associative relationship describing associated objects, indicating that three relationships may exist. For example, A and / or B may represent: A exists alone, A and B exist simultaneously, and B exists alone. In addition, the character " / " herein generally represents an "or" relationship between the associated objects before and after.
[0152] The above description shows and describes several preferred embodiments of the present application. However, as mentioned above, it should be understood that the present application is not limited to the form disclosed herein, should not be regarded as excluding other embodiments, but can be used in various other combinations, modifications and environments, and can be changed within the scope of the application concept described herein through the above teachings or the technology or knowledge in the relevant field. Any changes and variations made by those skilled in the art without departing from the spirit and scope of the present application shall fall within the protection scope of the appended claims of the present application.
Claims
1. A QoS routing optimization method, characterized in that: The QoS routing optimization method comprises the following steps: S1: Acquire a real-time network performance data set, where the real-time network performance data set includes link bandwidth, delay, and packet loss rate; S2: Modeling the network topology as a graph structure based on the real-time network performance data set to generate a network status feature matrix; S3: constructing a state space and an action space according to the network state feature matrix, wherein the action space includes multiple candidate paths from a source node to a target node; S4: selecting a candidate path from the action space through the main network, evaluating the performance value of the candidate path using the target network, updating the main network according to the performance value, and generating an experience data set; S5: Optimize the path selection strategy according to the empirical data set and update the routing table.
2. The QoS routing optimization method according to claim 1, characterized in that: The obtaining of the real-time network performance data set includes: Collect link bandwidth, latency, and packet loss rate data through network monitoring tools to generate an initial network performance data set; Define low latency and high throughput metrics based on video streaming and file transfer business requirements; Using a scheduled task mechanism to regularly collect network environment data and update the real-time network performance data set; The integrity of the real-time network performance data set is verified through a data verification mechanism to obtain a verified real-time network performance data set.
3. The QoS routing optimization method according to claim 1, characterized in that: The network topology modeling based on the real-time network performance data set is a graph structure, and a network status feature matrix is generated, including: The network topology is modeled as a graph structure, where nodes represent network devices, edges represent links, and edge weights are bandwidth, delay, and packet loss rate in the real-time network performance dataset; The graph neural network is used to learn the node and edge relationships, and the attention mechanism is used to calculate the edge weights related to low latency and high throughput indicators to generate a network status feature vector set. The neighbor node feature vectors are aggregated according to the network state feature vector set, the node feature vectors are updated, and the network state feature matrix is generated.
4. The QoS routing optimization method according to claim 3, characterized in that: The selecting a candidate path from the action space by using the main network and evaluating the performance value of the candidate path by using the target network includes: Outputting action value estimates according to the network state feature matrix through the main network, and selecting candidate paths from the action space; Calculating a target value estimate of the candidate path using the target network; According to the difference between the action value estimate and the target value estimate, the parameters of the main network are adjusted to generate an experience data set including states, actions and values.
5. The QoS routing optimization method according to claim 4, characterized in that: The optimizing the path selection strategy according to the empirical data set includes: Sorting the data in the empirical data set by a priority sorting algorithm, wherein the sorting basis is the absolute value of the temporal difference error; Extract batch data from the sorted empirical data set in priority proportion according to the sampling batch size; The weighting is performed according to the weight of each experience in the batch data, and the weighted batch data is used to update the parameters of the main network to optimize the path selection strategy.
6. The QoS routing optimization method according to claim 5, characterized in that: The updating routing table comprises: Sending the optimized path selection strategy to the network device through the network environment interaction module; Generate device response data according to the network device response path switching operation; Update the routing table according to the device response data to generate a routing table update data set; The bandwidth, delay and packet loss rate data of the actual path are extracted from the routing table update data set, compared with the real-time network performance data set, and the path selection strategy of the main network is adjusted.
7. A QoS routing optimization system, characterized in that: The QoS routing optimization system comprises: A real-time data acquisition module, which acquires a real-time network performance data set, wherein the real-time network performance data set includes link bandwidth, delay and packet loss rate; A data modeling module, which models the network topology into a graph structure based on the real-time network performance data set and generates a network status feature matrix; A space construction module, constructing a state space and an action space according to the network state feature matrix, wherein the action space includes multiple candidate paths from a source node to a target node; a performance value updating module, which selects a candidate path from the action space through a main network, uses a target network to evaluate the performance value of the candidate path, updates the main network according to the performance value, and generates an experience data set; The routing update module optimizes the path selection strategy according to the experience data set and updates the routing table.
8. The QoS routing optimization system according to claim 7, characterized in that: The QoS routing optimization system is used for network link path optimization.
9. A computer comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the computer program, the QoS routing optimization method according to any one of claims 1 to 6 is implemented.
10. A readable storage medium having a computer program stored thereon, characterized in that: When the program is executed by a processor, the QoS routing optimization method according to any one of claims 1 to 6 is implemented.
Citation Information
Patent Citations
Routing optimization method and system based on graph neural network and deep reinforcement learning
CN113194034A
Intelligent network path optimization method and system based on deep reinforcement learning
CN116527567A
Internet of Things defense method of reinforcement learning based on graph attention enhancement
CN119449457A
Cited By
Cooperative data migration scheduling method based on topology awareness and deep reinforcement learning
CN120469982A
Optimization method of high-performance 5G communication module
CN120676386A
Self-adaptive Mesh network architecture construction method for hybrid networking of industrial Internet of Things
CN120769326A
Adaptive mesh network architecture construction method for industrial internet of things hybrid networking
CN120769326B
Communication path matching method and system
CN120785756A