Multi-path Routing Selection Method for Communication Networks Based on Multi-agent Deep Reinforcement Learning

Through the multi-path routing method of communication networks with deep reinforcement learning of multi-agents, the real-time reliable transmission problem of communication networks in complex environments is solved, efficient path decision-making of heterogeneous services is achieved, and network performance and anti-interference ability are improved.

CN120075121BActive Publication Date: 2025-07-22NANJING UNIV OF INFORMATION SCI & TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510518347.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-04-24
Publication Date
2025-07-22
Estimated Expiration
2045-04-24

AI Technical Summary

Technical Problem

Existing communication network routing methods are difficult to achieve real-time and reliable service transmission in complex environments, especially in complex environments such as meteorology, terrain, and electromagnetic interference. Traditional, heuristic and deep reinforcement learning methods have limitations in adaptability and resource utilization, and cannot meet the QoS needs of heterogeneous services.

Method used

The multipath routing method of communication network based on deep reinforcement learning of multi-agents is adopted, and software-defined networks are built using Mininet and Ryu software, and a layered Actor-Critic network is deployed, including high- and low-level Actor networks and Critic networks. Through interactive training between agents and software-defined networks, a multipath routing algorithm model is formed to realize path decisions for heterogeneous services.

Benefits of technology

Under delay-sensitive, bandwidth-sensitive and reliability-sensitive services, the average end-to-end delay is significantly reduced, network throughput is improved, packet loss rate is reduced, anti-interference ability is achieved, and real-time and reliable transmission in complex environments is ensured.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120075121B_ABST
    Figure CN120075121B_ABST
Patent Text Reader

Abstract

The present invention discloses a multi - path routing selection method for communication networks based on multi - agent deep reinforcement learning, including constructing a software - defined network using Mininet and Ryu software, simulating a communication network using this network, where the topological structure of the communication network includes nodes and links between nodes; establishing a multi - path routing algorithm model for the communication network, and the model includes agents deployed at nodes, and the agents are multi - agent proximal policy optimization agents; through data interaction between the agents and the software - defined network, training the agents to obtain trained agents, and further obtaining a trained multi - path routing algorithm model for the communication network; using the trained multi - path routing algorithm model for the communication network to M path decision - making for heterogeneous services, and completing the selection of multi - path routing for the communication network. The present invention has a certain anti - interference ability for different service types in the communication network and can ensure the real - time and reliable transmission of services in complex environments.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of information engineering technology, and particularly to a multipath routing selection method for a communication network based on multi-agent deep reinforcement learning. Background Art

[0002] There are as many as seven types of heterogeneous services in a communication network. Due to the lack of an intelligent QoS (Quality of Service) control strategy, traditional routing methods are difficult to achieve the optimal matching of service types and forwarding paths. Especially in complex environments with meteorology, terrain, and electromagnetic interference, the transmission bandwidth of large-capacity services such as picture, video, and file transmission is limited, resulting in the inability to guarantee their QoS requirements. Existing routing methods do not consider the characteristic of coexistence of multiple links between nodes in a communication network, and cannot comprehensively utilize various communication means and give full play to communication network resources to ensure the efficient transmission of large-capacity services. In order to ensure real-time and reliable communication in complex environments, it is necessary to adopt multipath routing technology to rationally utilize these redundant links. Multipath routing refers to finding multiple paths from a source to a destination in a network through certain rule constraints, which requires multiple nodes to transmit data packets and reasonably distribute the load to the found multiple paths. Multipath routing can be divided into two categories: flow-by-flow and packet-by-packet. Packet-by-packet can route data packets from the same flow to different paths, with finer granularity. However, the out-of-order rearrangement problem caused by packet-by-packet will greatly affect network performance, and it is obviously not applicable to strongly adversarial and highly mobile environments, which is a problem that does not exist in flow-by-flow. Currently, the flow-by-flow multipath routing methods are mainly divided into traditional methods, heuristic methods, and deep reinforcement learning methods.

[0003] (1) Traditional Methods

[0004] Traditional multipath routing methods, such as equal-cost multipath, play an important role in achieving load balancing and improving network reliability. For example, equal-cost multipath is widely used in data centers and Internet backbones to effectively utilize redundant links by distributing data packets on paths with equal costs. However, these methods often face challenges in adapting to dynamic network changes, resulting in poor link utilization and congestion, so they have limitations in scenarios that require high adaptability.

[0005] (2) Heuristic Methods

[0006] To overcome the limitations of traditional methods, researchers have explored heuristic methods to provide more flexible and adaptable routing solutions. Heuristic methods, such as ant colony optimization, genetic algorithms, etc., aim to find approximate optimal solutions using exhaustive search methods while reducing computational overhead. These methods are inspired by natural processes or mathematical models and can better adapt to changing network conditions. Related literature 1 (Gu Bin, Lu Jie, Zhang Ce. Hybrid Multipath QoS Routing Protocol Based on Ant Colony Optimization [J]. Mobile Communications, 2023, 47(10): 24-31) proposed a hybrid multipath QoS routing protocol based on ant colony optimization. This protocol introduced multi-constraint QoS metrics and fully considered various service requirements. The simulation results showed that the proposed protocol was superior to other traditional routing protocols in terms of end-to-end delay and packet delivery ratio. Related literature 2 (Kumar R. Hybrid Marine Predators and Border Collie Optimization algorithm for multipath routing in IoT [J]. International Journal of Communication Systems, 2023, 36(15): e5567) proposed a hybrid method combining the marine predator optimization algorithm and the border collie optimization algorithm for optimizing multipath routing in the Internet of Things. This method combines global search and local search, can improve the efficiency of routing, reduce latency and energy consumption, and enhance the robustness of the network, showing better performance than traditional methods. Despite many advantages, heuristic methods also face challenges in terms of computational complexity and scalability. For example, heuristic methods may have a huge computational amount when applied to large-scale networks with a large number of nodes and potential paths. In addition, heuristic methods do not guarantee an optimal solution and may face convergence problems especially in highly dynamic network environments.

[0007] (3) Deep Reinforcement Learning Methods

[0008] The deep reinforcement learning method combines the adaptability of reinforcement learning and the function approximation ability of deep neural networks, and can characterize the characteristics of the continuously changing network state in a complex environment. Research shows that in terms of minimizing latency and improving bandwidth utilization, the traffic engineering framework based on DRL (Deep Reinforcement Learning), the SDN (Software Defined Network) routing method based on deep reinforcement learning, and the critical flow rerouting method based on reinforcement learning are superior to heuristic methods. However, for the changes in network structure and link state in a complex environment, the dimensionality of the network state space that the DRL agent needs to perceive increases sharply, and the existing DRL routing methods face the problem of dimensionality explosion. The training process of a single agent leads to an increase in decision-making complexity due to the exponential growth of the action space, thereby reducing the accuracy of routing inference. Summary of the Invention

[0009] The technical problem to be solved by the present invention is: aiming at the problem that the existing communication network routing methods are difficult to meet the requirements of real-time and reliable communication in a complex environment, a multi-path routing selection method for communication networks based on multi-agent deep reinforcement learning is provided, aiming to implement an intelligent QoS (Quality of Service) control strategy, form a certain anti-interference ability, so as to achieve the optimal matching of service types and forwarding paths, and ensure the real-time and reliable transmission of services in a complex environment.

[0010] To solve the above technical problems, the present invention adopts the following technical solutions:

[0011] A multi-path routing selection method for communication networks based on multi-agent deep reinforcement learning, comprising the following steps:

[0012] S1. Use Mininet and Ryu software to construct a software-defined network, and use this network to simulate a communication network. The topological structure of the communication network includes nodes and links between nodes.

[0013] S2. Based on the routing problem of the communication network, establish a multi-path routing algorithm model for the communication network. The model includes agents deployed at nodes, and the agents are multi-agent proximal policy optimization agents.

[0014] S3. Obtain node states, link states, adjacency matrices, and link matrices. After being processed by the multi-path routing algorithm model of the communication network, obtain corresponding paths. Through the data interaction between the agent and the software-defined network, train the agent to obtain a trained agent, and then obtain a trained multi-path routing algorithm model for the communication network.

[0015] S4. Use the trained communication network multipath routing algorithm model to make path decisions for M heterogeneous services in the communication network, and complete the selection of communication network multipath routing.

[0016] Further, in step S2, the agent includes a hierarchical Actor-Critic network, which includes a high-level Actor network, a low-level Actor network, a high-level Critic network, and a low-level Critic network.

[0017] The high-level Actor network includes an input layer, a general feature extraction layer, and a heterogeneous service-specific policy layer; the low-level Actor network includes a first deep neural network; the high-level Critic network includes a graph neural network and a second deep neural network; the low-level Critic network includes a third deep neural network.

[0018] Construct a high-level reward function R based on path state information h , and the specific expression is:

[0019] ;

[0020] where b represents the throughput value of the heterogeneous service, d represents the delay value of the heterogeneous service, l represents the packet loss rate of the heterogeneous service, , and respectively represent the importance weights of b, d, and l.

[0021] Construct a low-level reward function based on link bandwidth utilization and the high-level reward function , and the specific expression is:

[0022] ;

[0023] where represents the discount factor; u represents the link bandwidth utilization, , represents the link bandwidth capacity, represents the used bandwidth of the link.

[0024] Further, the heterogeneous services include delay-sensitive services, bandwidth-sensitive services, and reliability-sensitive services.

[0025] Further, in step S3, training the agent includes the following:

[0026] S301. Initialize the communication network multipath routing algorithm model, set the initial time step to 0; obtain the heterogeneous service to be selected for multipath routing, the source node of this service is , and the destination node is ; the path starts from the source node .

[0027] S302. The agent at the current node obtains the node state, link state, adjacency matrix, and link matrix at the t-th time step.

[0028] S303. Input the node state at the t-th time step into the hierarchical Actor-Critic network. The node state is transmitted to the general feature extraction layer through the input layer of the high-level Actor network, and feature extraction is performed using a deep neural network to obtain a general node feature vector. This vector passes through the heterogeneous service-specific policy layer, and the attention distribution corresponding to the heterogeneous service is obtained using the graph multi-head attention mechanism.

[0029] The attention distribution includes the weight of the current node selecting the j-th node of the neighbor as the next hop. The calculation formula of the weight is:

[0030] ;

[0031] where represents the attention weight of the current node to the j-th node, represents the total number of attention heads, represents the set of neighbor nodes of the current node, n represents the n-th attention head, , both represent the learning parameters of the n-th attention head of the current node for the -th heterogeneous service type, represents transpose of, represents the activation function, represents the feature vector of the current node, represents the feature vector of the j-th node, represents the feature vector of the -th node, exp represents the natural exponential function operation.

[0032] According to the attention distribution and the corresponding adjacency matrix, the high-level action corresponding to the heterogeneous service and the log probability corresponding to this action are obtained, and this action is used as the next-hop node .

[0033] The link state at the t-th time step and the next-hop node jointly pass through the low-level Actor network, feature extraction is performed using the first deep neural network, and the probability distribution of the link type is calculated. According to this probability distribution and the corresponding link matrix, the low-level action and the log probability corresponding to this action are obtained, and this action is used as the first link .

[0034] S304. The next-hop node and the first link Combine them to obtain the agent action at the current node , and the agent moves to the next-hop node .

[0035] S305. The agent at the next-hop node repeats steps S302 - S304 until it reaches the destination node and then stops, thus forming path Q, .

[0036] Send path Q to the software-defined network to transmit the traffic flow, measure the throughput, delay, and packet loss rate of the traffic flow, and calculate the high-level reward value and low-level reward value using the high-level reward function and low-level reward function; use the software-defined network to update the node state, link state, adjacency matrix, and link matrix to obtain the node state, link state, adjacency matrix, and link matrix at the (t + 1)-th time step, and then obtain the high-level decision sequence and low-level decision sequence corresponding to the agent in path Q. The specific expressions are as follows:

[0037] ;

[0038] ;

[0039] where represents the high-level decision sequence of the i-th agent, , , N represents the total number of agents, represents the end point of the path, represents the node state at the t-th time step, represents the logarithm of the probability of the high-level action at the t-th time step, represents the high-level action of the (i + 1)-th agent, represents the high-level reward value at the t-th time step, represents the node state at the (t + 1)-th time step, represents the low-level decision sequence of the i-th agent, represents the link state at the t-th time step, represents the logarithm of the probability of the low-level action at the t-th time step, represents the low-level action of the (i + 1)-th agent, represents the low-level reward value at the t-th time step, represents the link state at the (t + 1)-th time step.

[0040] S306. The and in the low-level decision sequence jointly pass through the low-level Critic network, and use the third deep neural network to obtain the corresponding low-level action value estimate and , and calculate the temporal difference error of the low-level action at the t-th time step , and the specific calculation formula is:

[0041] .

[0042] Recursively calculate the advantage value of the low-level action at the t-th time step , and the specific calculation formula is:

[0043] ;

[0044] Among them, , T represents the number of training batches; represents the generalized advantage estimation decay factor; represents the advantage value of the low-level action at the (t + 1)-th time step; represents the advantage value of the low-level action at the T-th time step.

[0045] S307. Based on the advantage value of the low-level action at the t-th time step , calculate the loss gradient and update the model parameters of the low-level Actor network and the low-level Critic network.

[0046] S308. Update the time step, and repeat steps S306 to S307 until the set maximum number of training times is reached and stop, to obtain the trained low-level Actor network and low-level Critic network.

[0047] S309. Freeze the model parameters of the trained low-level Actor network and low-level Critic network, set the time step to 0; and in the high-level decision sequence pass through the high-level Critic network together, and use the graph neural network and the second deep neural network to obtain the corresponding high-level action value estimation and , and calculate the temporal difference error of the high-level action at the t-th time step , and the specific calculation formula is:

[0048] .

[0049] Recursively calculate the advantage value of the high-level action at the t-th time step , and the specific calculation formula is:

[0050] ;

[0051] Among them, represents the advantage value of the high-level action at the (t + 1)-th time step, represents the advantage value of the high-level action at the T-th time step.

[0052] S310. Advantage value based on the high-level action at the t-th time step , calculate the loss gradient and update the model parameters of the high-level Actor network and the high-level Critic network.

[0053] S311. Update the time step, repeat steps S309 to S310 until the set maximum number of training times is reached and stop, obtaining the trained high-level Actor network and high-level Critic network; complete the training of the hierarchical Actor-Critic network.

[0054] Furthermore, in step S3, the node state is a matrix with the shape of , where V represents the total number of nodes; the state vector of the j-th node of the i-th agent is expressed as , and the specific expression is:[[]]

[0055] ;

[0056] ;

[0057] ;

[0058] ;

[0059] ;

[0060] ;

[0061] ;

[0062] ;

[0063] ;

[0064] where S Nij1 represents the normalized value of the service flow bandwidth requirement, S Nij2 represents the normalized value of the path hop count obtained by the shortest path algorithm, S Nij3 represents the normalized value of the maximum available bandwidth of the path obtained by the shortest path algorithm, S Nij4 represents the normalized value of the minimum cumulative packet loss rate of the path obtained by the shortest path algorithm, S Nij5 represents the normalized value of the path hop count obtained by the widest path algorithm, S Nij6 represents the normalized value of the maximum available bandwidth of the path obtained by the widest path algorithm, S Nij7 represents the normalized value of the minimum cumulative packet loss rate of the path obtained by the widest path algorithm, S Nij8 represents the shortest path indicator, S Nij9Indicates the widest path indicator, Indicates the sending rate of the service request, Indicates the path distance from the i-th agent to the j-th node obtained by using the minimum hop count algorithm, Indicates the maximum available bandwidth of the path from the i-th agent to the j-th node obtained by using the minimum hop count algorithm, Indicates the minimum cumulative packet loss rate of the path from the i-th agent to the j-th node obtained by using the minimum hop count algorithm, Indicates the lowest packet loss rate of the link between the i-th agent and the j-th node, Indicates the path distance from the i-th agent to the j-th node obtained by using the widest path algorithm, Indicates the maximum available bandwidth of the path from the i-th agent to the j-th node obtained by using the widest path algorithm, Indicates the minimum cumulative packet loss rate of the path from the i-th agent to the j-th node obtained by using the widest path algorithm, Indicates the path distance from the i-th agent to the m-th node obtained by using the minimum hop count algorithm, Indicates the maximum available bandwidth of the path from the i-th agent to the m-th node obtained by using the widest path algorithm, and min() represents the minimum value function.

[0065] The link state is a matrix of shape , where L represents the total number of link types; the state vector of the k-th link type of the i-th agent Is expressed as , and the specific calculation formula is:

[0066] ;

[0067] ;

[0068] ;

[0069] Among them, Indicates the normalized value of the remaining bandwidth of the link, Indicates the normalized value of the link delay, Indicates the normalized value of the link packet loss rate, Indicates the bandwidth capacity of the k-th link type connected to the i-th agent, Indicates the bandwidth usage of the k-th link type connected to the i-th agent, Indicates the delay of the k-th link type connected to the i-th agent, Indicates the packet loss rate of the k-th link type connected to the i-th agent.

[0070] The adjacency matrix is of shape matrix

[0071] The link matrix is a matrix with a shape of matrix

[0072] Furthermore, in step S307, updating the model parameters of the lower-level network includes the following:

[0073] Calculate the lower-level action loss gradient , and the specific formula is:

[0074] ;

[0075] where clip represents the clipping function; represents the probability ratio of the new and old lower-level actions at the t-th time step, , represents the lower-level action log probability at the (t - 1)-th time step; represents the clipping parameter.

[0076] Calculate the lower-level value loss gradient , and the specific formula is:

[0077] .

[0078] Calculate the total loss gradient , and backpropagate and store it in the agent. The specific formula is:

[0079] ;

[0080] where represents the value loss coefficient, represents the entropy regularization weight, represents the first policy distribution entropy.

[0081] According to , use the Adam optimizer to inversely update the model parameters of the lower-level Actor network and the lower-level Critic network.

[0082] Furthermore, in step S310, updating the model parameters of the higher-level network includes the following:

[0083] Calculate the higher-level action loss gradient , and the specific formula is:

[0084] ;

[0085] where represents the probability ratio of the new and old higher-level actions at the t-th time step, , represents the higher-level action log probability at the (t - 1)-th time step.

[0086] Calculate the high-level value loss gradient , and the specific formula is:

[0087] .

[0088] Calculate the high-level total loss gradient , and backpropagate and store it in the agent. The specific formula:

[0089] ;

[0090] where represents the second policy distribution entropy.

[0091] According to , use the Adam optimizer to inversely update the model parameters of the high-level Actor network and the high-level Critic network.

[0092] Furthermore, in step S4, the multi-path routing selection of the communication network includes the following:

[0093] Based on the high-level decision sequence and the low-level decision sequence obtained in step S305, and are input into the trained hierarchical Actor-Critic network, and steps S303 - S305 are repeated until routing paths are obtained for M heterogeneous services and then stop.

[0094] Furthermore, the present invention also proposes an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, it implements the steps of the communication network multi-path routing selection method based on multi-agent deep reinforcement learning.

[0095] Furthermore, the present invention also proposes a computer-readable storage medium. The computer-readable storage medium stores a computer program, and when the computer program is run by a processor, it executes the communication network multi-path routing selection method based on multi-agent deep reinforcement learning.

[0096] When the present invention adopts the above technical solutions compared with the prior art, it has the following technical effects:

[0097] Under time-delay sensitive services, the average end-to-end delay of the present invention is significantly reduced; under large-capacity services, i.e., bandwidth-sensitive services, the average network throughput is significantly improved; under reliability-sensitive services, the packet loss rate is significantly reduced. In a bandwidth-limited scenario, the throughput of large-capacity services remains at a relatively high level. Therefore, the present invention has a good strategy for the QoS requirements of heterogeneous services in a communication network, has a certain anti-interference ability, and can ensure the real-time and reliable transmission of services in a complex environment. Description of the Drawings

[0098] Figure 1 is the overall implementation flowchart of the present invention.

[0099] Figure 2 is the Actor network model diagram of the present invention.

[0100] Figure 3 is the Critic network model diagram of the present invention.

[0101] Figure 4 is the communication network topology diagram generated based on Mininet in the embodiment of the present invention.

[0102] Figure 5 is the ablation experiment comparative analysis diagram in the embodiment of the present invention.

[0103] Figure 6 is the comparison diagram of the average end-to-end delay of delay-sensitive services in the embodiment of the present invention.

[0104] Figure 7 is the comparison diagram of the average network throughput of bandwidth-sensitive services in the embodiment of the present invention.

[0105] Figure 8 is the comparison diagram of the average packet loss rate of reliability-sensitive services in the embodiment of the present invention.

[0106] Figure 9 is the comparison diagram of the network throughput of bandwidth-sensitive services in the bandwidth-limited scenario in the embodiment of the present invention. Detailed Implementation Manner

[0107] The present invention will be further described below with reference to the accompanying drawings. The following embodiments are only used to more clearly illustrate the technical solution of the present invention and cannot be used to limit the protection scope of the present invention.

[0108] To achieve the above object, the present invention proposes a multi-path routing selection method for a communication network based on multi-agent deep reinforcement learning, as Figure 1 shown, and the specific steps are as follows:

[0109] S1. Use Mininet and Ryu software to construct an SDN (Software Defined Network), which includes a series of programmable switches, routers, and links. Use this network to simulate a complex communication network, and the topological structure of the communication network includes nodes and links between nodes.

[0110] S2. Based on the routing problem of the communication network, a multi-path routing algorithm model of the communication network is established. The model includes an intelligent agent for node deployment, and the intelligent agent is a MAPPO (Multi-Agent Proximal Policy Optimization) intelligent agent. The specific content is as follows:

[0111] The intelligent agent includes a hierarchical Actor-Critic network, which includes a high-level Actor network, a low-level Actor network, a high-level Critic network, and a low-level Critic network. The high-level Actor network is used to select the next-hop node of the current node, the low-level Actor network is used to select the link type between the current node and the next-hop node, the high-level Critic network is used to output the value of the decision of the high-level Actor network, and the low-level Critic network is used to output the value of the decision of the low-level Actor network.

[0112] The high-level Actor network includes an input layer, a general feature extraction layer, and a heterogeneous service-specific policy layer. The low-level Actor network includes a first deep neural network. The high-level Critic network includes a graph neural network and a second deep neural network. The low-level Critic network includes a third deep neural network.

[0113] Among them, according to the differences in the Quality of Service (QoS) requirements of heterogeneous services in the communication network, the heterogeneous services are divided into three categories: the first category is delay-sensitive services, the second category is large-capacity services, i.e., bandwidth-sensitive services, and the third category is reliability-sensitive services.

[0114] Construct a high-level reward function R based on the path state information h , and the specific expression is:

[0115] ;

[0116] Among them, b represents the throughput value of the heterogeneous service, d represents the delay value of the heterogeneous service, l represents the packet loss rate of the heterogeneous service, 、 and respectively represent the importance weights of b, d, and l.

[0117] The values of the importance weights depend on the types of heterogeneous services in the communication network. For delay-sensitive services, 、 and are 0, 1, and 0 respectively; for bandwidth-sensitive services, 、 and are 1, 0, and 0 respectively; for reliability-sensitive services, 、 and are 0, 0, and 1 respectively.

[0118] Construct the low - level reward function based on the link bandwidth utilization and the high - level reward function , and the specific expression is:

[0119] ;

[0120] Among them, represents the discount factor; u represents the link bandwidth utilization, , represents the link bandwidth capacity, represents the used bandwidth of the link.

[0121] S3. Obtain the node status, link status, adjacency matrix, and link matrix. After being processed by the communication network multipath routing algorithm model, the corresponding path is obtained. Through the data interaction between the agent and the SDN, the agent is trained to obtain the trained agent, and then the trained communication network multipath routing algorithm model is obtained. The specific content is as follows:

[0122] S301. Initialize the communication network multipath routing algorithm model, and set the initial time step to 0; obtain the heterogeneous services to be selected for multipath routing. The source node of this service is , and the destination node is ; The path starts from the source node .

[0123] S302. The agent at the current node obtains the node status, link status, adjacency matrix, and link matrix at the t - th time step.

[0124] Among them, the node status is a matrix with the shape of , V represents the total number of nodes; the state vector of the j - th node of the i - th agent is , and the specific expression is:

[0125] ;

[0126] ;

[0127] ;

[0128] ;

[0129] ;

[0130] ;

[0131] ;

[0132] ;

[0133] ;

[0134] Among them, S Nij1 represents the normalized value of the bandwidth requirement of the service flow, S Nij2 represents the normalized value of the path hop count obtained by the shortest path algorithm, S Nij3 represents the normalized value of the maximum available bandwidth of the path obtained by the shortest path algorithm, S Nij4 represents the normalized value of the minimum cumulative packet loss rate of the path obtained by the shortest path algorithm, S Nij5 represents the normalized value of the path hop count obtained by the widest path algorithm, S Nij6 represents the normalized value of the maximum available bandwidth of the path obtained by the widest path algorithm, S Nij7 represents the normalized value of the minimum cumulative packet loss rate of the path obtained by the widest path algorithm, S Nij8 represents the shortest path indicator, S Nij9 represents the widest path indicator, represents the sending rate of the service request, represents the path distance from the i-th agent to the j-th node obtained by using the minimum hop count algorithm, represents the maximum available bandwidth of the path from the i-th agent to the j-th node obtained by using the minimum hop count algorithm, represents the minimum cumulative packet loss rate of the path from the i-th agent to the j-th node obtained by using the minimum hop count algorithm, represents the lowest packet loss rate of the link between the i-th agent and the j-th node, represents the path distance from the i-th agent to the j-th node obtained by using the widest path algorithm, represents the maximum available bandwidth of the path from the i-th agent to the j-th node obtained by using the widest path algorithm, represents the minimum cumulative packet loss rate of the path from the i-th agent to the j-th node obtained by using the widest path algorithm, represents the path distance from the i-th agent to the m-th node obtained by using the minimum hop count algorithm, represents the maximum available bandwidth of the path from the i-th agent to the m-th node obtained by using the widest path algorithm, min() represents the minimum value function.

[0135] S Nij8 reflects whether the current link is part of the shortest path routing, S Nij9 reflects whether the current link is part of the widest path routing. If there is no connection relationship between the i-th agent and the j-th node, then .

[0136] The link state is a matrix in the form of , where L represents the total number of link types; the state vector of the k-th link type of the i-th agent is denoted as , and the specific calculation formula is:

[0137] ;

[0138] ;

[0139] ;

[0140] where represents the normalized value of the remaining bandwidth of the link, represents the normalized value of the link delay, represents the normalized value of the link packet loss rate, represents the bandwidth capacity of the k-th link type connected to the i-th agent, represents the bandwidth usage of the k-th link type connected to the i-th agent, represents the delay of the k-th link type connected to the i-th agent, represents the packet loss rate of the k-th link type connected to the i-th agent.

[0141] The adjacency matrix is a matrix in the form of , where each element represents whether there is a connection relationship between two nodes. If there is a connection relationship, the value of this element is 1, otherwise it is 0. The connection relationship between a certain node and other nodes can be obtained by indexing this node.

[0142] The link matrix is a matrix in the form of , where each element is a vector of length L, representing which links exist between two nodes. For the existing links, the corresponding bit value is 1, otherwise it is 0. The link information between a certain node and other nodes can be obtained by indexing this node.

[0143] S303. Input the node state at the t-th time step into the hierarchical Actor-Critic network, as Figure 2 shown. Through the input layer of the high-level Actor network, this node state is transmitted to the general feature extraction layer, and feature extraction is performed using a deep neural network to obtain the general node feature vector. This vector passes through the heterogeneous service-specific policy layer, and the attention distribution corresponding to the heterogeneous service is obtained using the graph multi-head attention mechanism.

[0144] The attention distribution includes the weight of the j-th node selected by the current node as the next hop, and the calculation formula of the weight is:

[0145] ;

[0146] Among them, represents the attention weight of the current node to the j-th node, represents the total number of attention heads, represents the set of neighbor nodes of the current node, and n represents the n-th attention head, , both represent the learning parameters of the n-th attention head under the current node for the -th heterogeneous service type, represents the transpose of, represents the activation function, represents the feature vector of the current node, represents the feature vector of the j-th node, represents the -th node's feature vector, and exp represents the natural exponential function operation.

[0147] Obtain the high-level action corresponding to the heterogeneous service and the logarithmic probability corresponding to this action according to the attention distribution and the corresponding adjacency matrix, and use this action as the next-hop node .

[0148] The link state at the t-th time step and the next-hop node jointly pass through the low-level Actor network, use the first deep neural network for feature extraction, and calculate the probability distribution of the link type. According to this probability distribution and the corresponding link matrix, obtain the low-level action and the logarithmic probability corresponding to this action, and use this action as the first link .

[0149] S304. Combine the next-hop node and the first link to obtain the agent action at the current node, and the agent moves to the next-hop node .

[0150] S305. The agent at the next-hop node repeats steps S302 - S304 until it reaches the destination node and then stops, thus forming a path Q, .

[0151] Send path Q to the SDN (Software Defined Network) to transmit the service flow, measure the throughput, latency, and packet loss rate of the service flow, and calculate the high-level reward value and low-level reward value using the high-level reward function and low-level reward function; use the SDN to update the node state, link state, adjacency matrix, and link matrix to obtain the node state, link state, adjacency matrix, and link matrix at the (t + 1)-th time step, and then obtain the high-level decision sequence and low-level decision sequence corresponding to the agent in path Q. The specific expressions are as follows:

[0152] ;

[0153] ;

[0154] where, represents the high-level decision sequence of the i-th agent, , , N represents the total number of agents, represents the end point of the path, represents the node state at the t-th time step, represents the log probability of the high-level action at the t-th time step, represents the high-level action of the (i + 1)-th agent, represents the high-level reward value at the t-th time step, represents the node state at the (t + 1)-th time step, represents the low-level decision sequence of the i-th agent, represents the link state at the t-th time step, represents the log probability of the low-level action at the t-th time step, represents the low-level action of the (i + 1)-th agent, represents the low-level reward value at the t-th time step, represents the link state at the (t + 1)-th time step.

[0155] S306. As shown in Figure 3 , in the low-level decision sequence, and both pass through the low-level Critic network, and the corresponding low-level action value estimates and are obtained using the third deep neural network, and the temporal difference error of the low-level action at the t-th time step is calculated. The specific calculation formula is:

[0156] .

[0157] Recursively calculate the advantage value of the low-level action at the t-th time step. The specific calculation formula is:

[0158] ;

[0159] Among them, , T represents the number of training batches, taking 128; represents the generalized advantage estimation decay factor, taking 0.5; represents the advantage value of the lower-layer action at the t+1 time step; represents the advantage value of the lower-layer action at the T time step.

[0160] S307. Based on the advantage value of the lower-layer action at the t time step , calculate the loss gradient and update the model parameters of the lower-layer Actor network and the lower-layer Critic network; the specific content is as follows:

[0161] Calculate the lower-layer action loss gradient , and the specific formula is:

[0162] ;

[0163] Among them, clip represents the clipping function, used to limit the range of a value; represents the probability ratio of the lower-layer old and new actions at the t time step, , represents the logarithmic probability of the lower-layer action at the t-1 time step; represents the clipping parameter, taking 0.1.

[0164] Calculate the lower-layer value loss gradient , and the specific formula is:

[0165] .

[0166] Calculate the total loss gradient , and backpropagate and store it in the agent. The specific formula is:

[0167] ;

[0168] Among them, represents the value loss coefficient, taking 0.5; represents the entropy regularization weight, taking 0.01; represents the first policy distribution entropy.

[0169] According to , use the Adam optimizer to inversely update the model parameters of the lower-layer Actor network and the lower-layer Critic network.

[0170] S308. Update the time step, repeat steps S306 to S307, and save the model parameters of the lower-level Actor network and the lower-level Critic network every time a specific number of steps are updated until the set maximum number of training times is reached and then stop, obtaining the trained lower-level Actor network and lower-level Critic network.

[0171] S309. Freeze the model parameters of the trained lower-level Actor network and lower-level Critic network, and set the time step to 0; in the high-level decision sequence, and jointly pass through the high-level Critic network, and use the graph neural network and the second deep neural network to obtain the corresponding high-level action value estimate and , and calculate the temporal difference error of the high-level action at the t-th time step. The specific calculation formula is:

[0172] .

[0173] Recursively calculate the advantage value of the high-level action at the t-th time step. The specific calculation formula is:

[0174] ;

[0175] where, represents the advantage value of the high-level action at the (t + 1)-th time step, represents the advantage value of the high-level action at the T-th time step.

[0176] S310. Based on the advantage value of the high-level action at the t-th time step, calculate the loss gradient and update the model parameters of the high-level Actor network and the high-level Critic network; the specific content is:

[0177] Calculate the high-level action loss gradient . The specific formula is:

[0178] ;

[0179] where, represents the probability ratio of the new and old high-level actions at the t-th time step, , represents the logarithmic probability of the high-level action at the (t - 1)-th time step.

[0180] Calculate the high-level value loss gradient . The specific formula is:

[0181] .

[0182] Calculate the high-level total loss gradient , and backpropagate and store it in the agent. The specific formula:

[0183] ;

[0184] where represents the second policy distribution entropy.

[0185] According to , use the Adam optimizer to reversely update the model parameters of the high-level Actor network and the high-level Critic network.

[0186] S311. Update the time step, repeat steps S309 to S310, and save the model parameters of the high-level Actor network and the high-level Critic network every time a specific number of steps are updated until the set maximum number of training times is reached and stop, obtaining the trained high-level Actor network and high-level Critic network; thus, the training of the hierarchical Actor-Critic network is completed.

[0187] S4. Based on the high-level decision sequence and the low-level decision sequence obtained in step S305, input and into the trained hierarchical Actor-Critic network, repeat steps S303 - S305, and use the trained communication network multipath routing algorithm model to make path decisions for M heterogeneous services in the communication network until routing paths are obtained for the M heterogeneous services and stop, completing the selection of the communication network multipath routing.

[0188] Example:

[0189] The communication network topology diagram is as shown in Figure 4 , with a total of 20 nodes, where S1 and D1 are source nodes, Z1 - Z9 are destination nodes, D2 is the core node, D3 - D5 are first-level control nodes, D6 - D8 are second-level control nodes, and F1 and F2 are auxiliary nodes.

[0190] There are three links between every two nodes. Specifically:

[0191] (1) Three types of links, namely scattering, VHF (Very High Frequency), and UHF (UltraHigh Frequency), are set between S1 and D2 - D8, with bandwidths of 8Mbps, 9.6Kbps, and 1Mbps respectively.

[0192] (2) There are three types of links, namely wired, satellite, and UHF, between D2 and D3 - D8, F1, F2, with bandwidths of 1 Mbps, 8 Mbps, and 1 Mbps respectively. There are three types of links, namely wired, satellite, and microwave, between D2 and D1, with bandwidths of 1 Mbps, 8 Mbps, and 8 Mbps respectively.

[0193] (3) There are three types of links, namely wired, satellite, and microwave, between D3 - D5, with bandwidths of 1 Mbps, 8 Mbps, and 8 Mbps respectively. There are three types of links, namely area-wide, VHF, and UHF, between D3 and Z1 - Z2, D4 and Z3 - Z4, and D5 and Z5 - Z6, with bandwidths of 10 Mbps, 9.6 Kbps, and 1 Mbps respectively.

[0194] (4) There are three types of links, namely wired, satellite, and microwave, between D6 - D8, with bandwidths of 1 Mbps, 8 Mbps, and 8 Mbps respectively. There are three types of links, namely area-wide, VHF, and UHF, between D6 and Z7, D7 and Z8, and D8 and Z9, with bandwidths of 10 Mbps, 9.6 Kbps, and 1 Mbps respectively.

[0195] (5) There are three types of links, namely wired, satellite, and area-wide, between D3 - D8, F1, F2 and D1, with bandwidths of 1 Mbps, 8 Mbps, and 10 Mbps respectively.

[0196] As Figure 5 shown, in deep reinforcement learning, the reward value is an important indicator parameter reflecting the method performance. In the experiment, three types of services are set as delay-sensitive service, bandwidth-sensitive service, and reliability-sensitive service. Among them, the sending rate of the delay-sensitive service is 100 Kbps, the sending rate of the bandwidth-sensitive service is 1500 Kbps, and the sending rate of the reliability-sensitive service is 500 Kbps, and their distribution is [0.3, 0.4, 0.3]. The hierarchical Actor-Critic network proposed in the present invention and the network after removing the heterogeneous service dedicated policy layer are trained respectively, and the average reward values of the high layer are compared. By Figure 5It can be seen that as the number of iterations increases, the reward values obtained by the two networks in each round gradually increase until convergence. Initially, the hierarchical Actor-Critic network proposed in the present invention has a more obvious improvement in the convergence speed compared to the network after removing the heterogeneous service dedicated policy layer. The network after removing the heterogeneous service dedicated policy layer gradually converges to a stable state after 100 rounds, while the hierarchical Actor-Critic network proposed in the present invention reaches convergence only after 50 to 60 rounds, and the convergence speed is increased by nearly 50%. This is because the graph multi-head attention mechanism in the model increases the more fine-grained attention to neighbor nodes, enabling the network to better understand different information in the state space during the initial training stage, thus accelerating the improvement of the network; in terms of the reward level after final convergence, it can be seen that the reward value level at which the hierarchical Actor-Critic network proposed in the present invention finally converges is significantly higher than that of the network after removing the heterogeneous service dedicated policy layer. The convergence reward value of the hierarchical Actor-Critic network proposed in the present invention stabilizes at around -0.5, while that of the network after removing the heterogeneous service dedicated policy layer stabilizes at around -1.0, indicating that the effect of the network with the added heterogeneous service dedicated policy layer has been improved, which can better aggregate graph structure information, avoid falling into local optimal solutions, and improve the quality of policy selection.

[0197] To verify the advantages of the method proposed in the present invention, experiments on the QoS requirements of delay-sensitive services, bandwidth-sensitive services, and reliability-sensitive services were conducted using the traditional communication network traditional routing method, i.e., the OSPF (Open Shortest Path First) method, the traditional multi-path routing method, i.e., the ECMP (Equal-cost multi-path routing) method, and the method proposed in the present invention, as Figure 6 , Figure 7 and Figure 8 shown. The experimental scenario is that S1 and D1 send these three services to Z1-Z9 respectively. The sending rate of the delay-sensitive service is 50 Kbps, the sending rate of the bandwidth-sensitive service is 1500 Kbps, and the sending rate of the reliability-sensitive service is 500 Kbps. In addition, the continuous time slots of the services are set to 10, 30, and 50 respectively to simulate three scenarios of light load, medium load, and heavy load.

[0198] Figure 6 is the comparison graph of the average end-to-end delay of the three methods under the delay-sensitive service. It can be seen that the end-to-end delay of the method of the present invention remains at a small value in the three scenarios, with the maximum reduction of 89.03% and 75.87% compared to the other two methods. Figure 7Comparison chart of the average network throughput of three methods under bandwidth-sensitive services. It can be seen that the method of the present invention has been improved compared with the other two methods in the three scenarios, with the maximum increases being 18.44% and 11.21% respectively. Figure 8 It is a comparison chart of the average packet loss rate of three methods under reliability-sensitive services. It can be seen that the packet loss rate of the method of the present invention remains at a relatively low level in the three scenarios. Compared with the other two methods, the maximum decreases are 94.12% and 86.59% respectively. Thus, it can be seen that the method of the present invention has a good strategy for the QoS requirements of different service types and can meet the QoS requirements of various heterogeneous services in the communication network.

[0199] To further verify the anti-interference ability of the method proposed by the present invention, a scenario where the link bandwidth becomes narrower under interference is simulated, and the QoS requirements of high-capacity services are experimented under this scenario. The experimental scenario is as follows: Set the transmission rate of the bandwidth-sensitive service to 1500 Kbps. S1 continuously sends bandwidth-sensitive services to Z1. At the 4th time slot, the scattering links between S1 and D2-D5 and the satellite links between D2 and D3-D5 are reduced to less than the transmission rate of the bandwidth-sensitive service, which is 1000 Kbps, so as to simulate the situation where the transmission bandwidth of the bandwidth-sensitive service is limited. Use the OSPF method, the ECMP method, and the method proposed by the present invention to obtain the network throughput of the bandwidth-sensitive service in the above experimental scenario and conduct a result comparison, as Figure 9 shown.

[0200] From Figure 9 it can be seen that after the link bandwidth becomes narrower, the network throughput of all three methods has decreased. Among them, the decrease amplitude of the method proposed by the present invention is much smaller than that of the other two methods, and the throughput of the method proposed by the present invention remains at a relatively high level. Thus, it can be seen that the method proposed by the present invention has a good effect on the transmission of bandwidth-sensitive services in the scenario of limited bandwidth and forms a certain anti-interference ability.

[0201] The embodiment of the present invention also proposes an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. It should be noted that when the processor executes the computer program, it corresponds to the specific steps of the method provided by the embodiment of the present invention and has the corresponding functional modules and beneficial effects of the execution method. For technical details not described in detail in this embodiment, reference can be made to the method provided by the embodiment of the present invention.

[0202] An embodiment of the present invention also provides a computer-readable storage medium storing a computer program. It should be noted that when the computer program is run by a processor, it corresponds to the specific steps of the method provided by the embodiment of the present invention, and has the corresponding functional modules and beneficial effects of the execution method. For technical details not described in detail in this embodiment, reference may be made to the method provided by the embodiment of the present invention.

[0203] The above are only the preferred embodiments of the present invention. It should be pointed out that for those of ordinary skill in the art, without departing from the technical principle of the present invention, several improvements and modifications can be made, and these improvements and modifications should also be regarded as the protection scope of the present invention.

Claims

1. A multi-path routing selection method for a communication network based on multi-agent deep reinforcement learning, characterized in that Including: S1. Use Mininet and Ryu software to construct a software-defined network, and use this network to simulate a communication network. The topological structure of the communication network includes nodes and links between nodes. S2. Based on the routing problem of the communication network, establish a multi-path routing algorithm model for the communication network. This model includes an agent for node deployment, and the agent is a multi-agent proximal policy optimization agent. Specifically: The agent includes a hierarchical Actor-Critic network, which includes a high-level Actor network, a low-level Actor network, a high-level Critic network, and a low-level Critic network. The high-level Actor network includes an input layer, a general feature extraction layer, and a heterogeneous service-specific policy layer; the low-level Actor network includes a first deep neural network. The high-level Critic network includes a graph neural network and a second deep neural network; the low-level Critic network includes a third deep neural network. Construct the high-level reward function R h , and the specific expression is as follows: R h = w1 log b - w2 d - w3 l Among them, b represents the throughput value of heterogeneous services, d represents the delay value of heterogeneous services, l represents the packet loss rate of heterogeneous services, and w1, w2, and w3 respectively represent the importance weights of b, d, and l. Construct the low-level reward function R l , and the specific expression is as follows: R l = γ·R h -(1 - γ)·log(1 + u) Among them, γ represents the discount factor; u represents the link bandwidth utilization rate, b capa represents the link bandwidth capacity, and b use represents the used bandwidth of the link; S3. Obtain heterogeneous services to be selected for multi-path routing. The path starts from the source node of this service. The agent at the current node obtains the node state, link state, adjacency matrix, and link matrix, and after being processed by the multi-path routing algorithm model of the communication network, obtains the corresponding path. Through the data interaction between the agent and the software-defined network, obtain the advantage value of the high-level action and the advantage value of the low-level action. Use the advantage value of the low-level action to train the low-level network, freeze the model parameters of the trained low-level network, use the advantage value of the high-level action to train the high-level network, obtain the trained agent, and then obtain the trained multi-path routing algorithm model of the communication network. S4. Use the trained multi-path routing algorithm model of the communication network to make path decisions for M heterogeneous services of the communication network, and complete the selection of multi-path routing of the communication network.

2. The multi-path routing selection method for a communication network based on multi-agent deep reinforcement learning according to claim 1, wherein, Heterogeneous services include delay-sensitive services, bandwidth-sensitive services, and reliability-sensitive services.

3. The multi-path routing selection method for a communication network based on multi-agent deep reinforcement learning according to claim 1, characterized in that In step S3, training the agent includes the following: S301. Initialize the communication network multipath routing algorithm model, set the initial time step to 0; obtain the heterogeneous services to be selected for multipath routing, where the source node of this service is g0 and the destination node is g κ ; The path starts from the source node g0; S302. The agent at the current node obtains the node state, link state, adjacency matrix, and link matrix at the t-th time step. S303. Input the node state at the t-th time step into the hierarchical Actor-Critic network. The node state is transmitted to the general feature extraction layer through the input layer of the high-level Actor network, and feature extraction is performed using a deep neural network to obtain a general node feature vector. This vector passes through the heterogeneous service-specific policy layer, and the attention distribution corresponding to the heterogeneous service is obtained using the graph multi-head attention mechanism. The attention distribution includes the weight of the current node selecting the j-th node of the neighbor as the next hop. The calculation formula of the weight is: Among them, α 当,j represents the attention weight of the current node to the j-th node, N h represents the total number of attention heads, represents the set of neighbor nodes of the current node, n represents the n-th attention head, both represent the learning parameters of the n-th attention head under the current node for the η-th heterogeneous service type, represents the transpose of, LeakyReLU represents the activation function, X 当 represents the feature vector of the current node, X j represents the feature vector of the j-th node, X θ represents the feature vector of the θ-th node, exp represents the natural exponential function operation; Obtain the high-level action corresponding to the heterogeneous service and the logarithmic probability corresponding to this action according to the attention distribution and the corresponding adjacency matrix, and use this action as the next-hop node g1. The link state at the t-th time step and the next-hop node g1 jointly pass through the low-level Actor network, extract features using the first deep neural network, and calculate the probability distribution of the link type. Based on this probability distribution and the corresponding link matrix, the low-level action and the log probability corresponding to this action are obtained, and this action is used as the first link a1; S304. Combine the next-hop node g1 and the first link a1 to obtain the agent action (g1, a1) at the current node, and the agent moves to the next-hop node g1; S305. The agent at the next-hop node g1 repeats steps S302 - S304 until it reaches the destination node g and then forms a path Q, where Q = {g0 → (g1, a1) → … → g κ} κ ; Send the path Q to the software-defined network to transmit the traffic flow, measure the throughput, delay, and packet loss rate of the traffic flow, and calculate the high-level reward value and the low-level reward value using the high-level reward function and the low-level reward function; Update the node state, link state, adjacency matrix, and link matrix using the software-defined network to obtain the node state, link state, adjacency matrix, and link matrix at the (t + 1)-th time step, and then obtain the high-level decision sequence and the low-level decision sequence corresponding to the agent in the path Q. The specific expressions are as follows: Among them, D i h represents the high-level decision-making sequence of the \(i\)-th agent, \(i\in\{0,1,\ldots,N\}\), \(N = \kappa - 1\), \(N\) represents the total number of agents, and \(\kappa\) represents the path end point. represents the node state at the \(t\)-th time step. represents the logarithmic probability of the high-level action at the \(t\)-th time step, \(g\) i+1 represents the high-level action of the \((i + 1)\)-th agent. represents the high-level reward value at the \(t\)-th time step. represents the node state at the \((t + 1)\)-th time step, \(D\) i l represents the low-level decision-making sequence of the \(i\)-th agent. represents the link state at the \(t\)-th time step. represents the logarithmic probability of the low-level action at the \(t\)-th time step, \(a\) i+1 represents the low-level action of the \((i + 1)\)-th agent. represents the low-level reward value at the \(t\)-th time step. represents the link state at the \((t + 1)\)-th time step; S306. In the low-level decision sequence, and jointly pass through the low-level Critic network, and use the third deep neural network to obtain the corresponding low-level action value estimate and and calculate the temporal difference error of the low-level action at the t-th time step The specific calculation formula is: Recursively calculate the advantage value of the low-level action at the t-th time step The specific calculation formula is as follows: where \(t\in [T - 1,T - 2,\cdots,0]\), \(T\) represents the number of training batches; \(\lambda\) represents the generalized advantage estimation decay factor; represents the advantage value of the low-level action at the \((t + 1)\)-th time step; represents the advantage value of the low-level action at the \(T\)-th time step; S307. Advantage value based on the low-level action at the t-th time step Calculate the loss gradient and update the model parameters of the low-level Actor network and the low-level Critic network; S308. Update the time step, and repeat steps S306 to S307 until the set maximum number of training times is reached and stop to obtain the trained low-level Actor network and low-level Critic network; S309. Freeze the model parameters of the trained lower-level Actor network and lower-level Critic network, and set the time step to 0; in the high-level decision sequence, and jointly pass through the high-level Critic network, and use the graph neural network and the second deep neural network to obtain the corresponding high-level action value estimate and and calculate the temporal difference error of the high-level action at the t-th time step The specific calculation formula is: Recursively calculate the advantage value of the high-level action at the t-th time step The specific calculation formula is as follows: wherein, represents the advantage value of the high-level action at the (t + 1)-th time step, represents the advantage value of the high-level action at the T-th time step; S310. Advantage value based on the high-level action at the t-th time step Calculate the loss gradient and update the model parameters of the high-level Actor network and the high-level Critic network; S311. Update the time step, and repeat steps S309 to S310 until the set maximum number of training times is reached and stop to obtain the trained high-level Actor network and high-level Critic network; Complete the training of the hierarchical Actor-Critic network.

4. The multi-path routing selection method for a communication network based on multi-agent deep reinforcement learning according to claim 3, wherein In step S3, the node state is a matrix of shape N×V×9, where V represents the total number of nodes; the state vector S of the j-th node of the i-th agent Nij is denoted as S Nij =[S Nij1 ,S Nij2 ,S Nij3 ,S N i j4 ,S Nij5 ,S Nij6 ,S Nij7 ,S Nij8 ,S Nij9 , and the specific expression is: Among them, S Nij1 represents the normalized value of the service flow bandwidth requirement, S Nij2 represents the normalized value of the path hop count obtained by the shortest path algorithm, S Nij3 represents the normalized value of the maximum available bandwidth of the path obtained by the shortest path algorithm, S Nij4 represents the normalized value of the minimum cumulative packet loss rate of the path obtained by the shortest path algorithm, S Nij5 represents the normalized value of the path hop count obtained by the widest path algorithm, S Nij6 represents the normalized value of the maximum available bandwidth of the path obtained by the widest path algorithm, S Nij7 represents the normalized value of the minimum cumulative packet loss rate of the path obtained by the widest path algorithm, S Nij8 represents the shortest path indicator, S Nij9 represents the widest path indicator, τ represents the transmission rate of the service request, D ij SPR represents the path distance from the i-th agent to the j-th node obtained by the minimum hop count algorithm, C ij SPR represents the maximum available bandwidth of the path from the i-th agent to the j-th node obtained by the minimum hop count algorithm, L ij SPR represents the minimum cumulative packet loss rate of the path from the i-th agent to the j-th node obtained by the minimum hop count algorithm, L ij m in represents the minimum packet loss rate of the link between the i-th agent and the j-th node, D ij WP represents the path distance from the i-th agent to the j-th node obtained by the widest path algorithm, C ij wp represents the maximum available bandwidth of the path from the i-th agent to the j-th node obtained by the widest path algorithm, L ij w p represents the minimum cumulative packet loss rate of the path from the i-th agent to the j-th node obtained by the widest path algorithm, D im SPR represents the path distance from the i-th agent to the m-th node obtained by the minimum hop count algorithm, C im WP represents the maximum available bandwidth of the path from the i-th agent to the m-th node obtained by the widest path algorithm, min() represents the minimum value function; The link state is a matrix of shape N×L×3, where L represents the total number of link types; the state vector S of the k-th link type of the i-th agent Lik is denoted as S Lik = [S Lik1 , S Lik2 , S Lik3 , and the specific calculation formula is as follows: Among them, S Lik1 represents the normalized value of the remaining bandwidth of the link, S Lik2 represents the normalized value of the link delay, S Lik3 represents the normalized value of the link packet loss rate, C ik represents the bandwidth capacity of the k-th link type connected to the i-th agent, U ik represents the bandwidth usage of the k-th link type connected to the i-th agent, D ik represents the delay of the k-th link type connected to the i-th agent, L ik represents the packet loss rate of the k-th link type connected to the i-th agent; The adjacency matrix is a matrix with a shape of N×N; The link matrix is a matrix with a shape of N×N×L.

5. The multi-path routing selection method for a communication network based on multi-agent deep reinforcement learning according to claim 3, characterized in that In step S307, updating the model parameters of the low-level network includes the following content: Calculate the gradient of the low-level action loss L action_low , and the specific formula is as follows: where clip represents the clipping function; represents the probability ratio of the old and new low-level actions at the t-th time step, represents the log probability of the low-level action at the (t - 1)-th time step; ε represents the clipping parameter; Calculate the low-level value loss gradient L value_low , and the specific formula is as follows: Calculate the total loss gradient L total_low , and backpropagate it and store it in the agent. The specific formula is as follows: L total_low = c1·L value_low + L action_low - c2·H l Among them, c1 represents the value loss coefficient, c2 represents the entropy regularization weight, and H l represents the entropy of the first policy distribution; According to L total_low , the model parameters of the lower-level Actor network and the lower-level Critic network are updated backward using the Adam optimizer.

6. The multi-path routing selection method for a communication network based on multi-agent deep reinforcement learning according to claim 3, characterized in that In step S310, updating the model parameters of the high-level network includes the following content: Calculate the gradient L of the high-level action loss action_high , and the specific formula is as follows: where clip represents the clipping function; represents the probability ratio of the high-level old and new actions at the t-th time step, represents the logarithmic probability of the high-level action at the (t - 1)-th time step; ε represents the clipping parameter; Calculate the high-level value loss gradient L value_high , and the specific formula is as follows: Calculate the total loss gradient L of the high layer total_high , and backpropagate it to be stored in the agent. The specific formula is: L total_high = c1·L value_high + L action_high - c2·H h Among them, c1 represents the value loss coefficient, c2 represents the entropy regularization weight, and H h represents the second policy distribution entropy; According to L total_high , the model parameters of the high-level Actor network and the high-level Critic network are updated backward using the Adam optimizer.

7. The multi-path routing selection method for a communication network based on multi-agent deep reinforcement learning according to claim 3, characterized in that In step S4, the multi-path routing selection of the communication network includes the following content: Based on the high-level decision sequence and low-level decision sequence obtained in step S305, and are input into the trained hierarchical Actor-Critic network, and steps S303 - S305 are repeated until routing paths are obtained for M heterogeneous services and then stopped.

8. An electronic device, comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the steps of the multi-path routing selection method for a communication network based on multi-agent deep reinforcement learning according to any one of claims 1 to 7.

9. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is run by the processor, it executes the multi-path routing selection method for a communication network based on multi-agent deep reinforcement learning according to any one of claims 1 to 7.

Citation Information

Patent Citations

  • Network autonomous intelligent management and control method based on deep reinforcement learning

    CN113328938A

  • Intelligent path optimization method and system based on link state perception enhancement

    CN119011463A