Communication network multipath routing selection method based on multi-agent deep reinforcement learning
By deploying a multi-path routing method with deep reinforcement learning in multiple agents in the communication network, the problem that communication network routing methods in complex environments are difficult to achieve real-time and reliable communication, and the optimal matching of service types and forwarding paths is achieved, ensuring real-time and reliable transmission of services in complex environments is ensured.
Patent Information
- Application Number
- CN202510518347.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-24
- Publication Date
- 2025-05-30
- Estimated Expiration
- 2045-04-24
AI Technical Summary
Existing communication network routing methods are difficult to achieve real-time and reliable communication in complex environments, especially in complex environments such as meteorology, terrain, and electromagnetic interference, and it is difficult to ensure efficient transmission of large-capacity services.
The multi-path routing method of communication network based on multi-agent deep reinforcement learning is adopted. By deploying multi-agent proximity strategy in the communication network, and using a hierarchical Actor-Critic network to make path decisions, we realize intelligent QoS control strategies.
Under delay-sensitive services, the average end-to-end delay is significantly reduced; under large-capacity services, the average network throughput is significantly improved; under reliability-sensitive services, the packet loss rate is significantly reduced. This method can effectively ensure the throughput of large-capacity services in scenarios with bandwidth limitations and has good anti-interference capabilities.
Smart Images

Figure CN120075121A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of information engineering, and particularly to a multi-path routing selection method for a communication network based on multi-agent deep reinforcement learning. Background Art
[0002] There are as many as seven types of heterogeneous services in a communication network. Due to the lack of an intelligent QoS (Quality of Service) control strategy, traditional routing methods are difficult to achieve the optimal matching of service types and forwarding paths. Especially in complex environments such as meteorology, terrain, and electromagnetic interference, the transmission bandwidth of large-capacity services such as picture, video, and file transmission is limited, resulting in the inability to guarantee their QoS requirements. Existing routing methods do not consider the characteristic of coexistence of multiple links between nodes in a communication network and cannot comprehensively utilize various communication means and give full play to the communication network resources to ensure the efficient transmission of large-capacity services. In order to ensure real-time and reliable communication in complex environments, it is necessary to adopt multi-path routing technology to rationally utilize these redundant links. Multi-path routing refers to finding multiple paths from the source to the destination in the network through certain rule constraints, which requires multiple nodes to transmit data packets and reasonably distribute the load to the found multiple paths. Multi-path routing can be divided into two categories: flow-by-flow and packet-by-packet. Packet-by-packet can route data packets from the same flow to different paths and has a finer granularity. However, the out-of-order rearrangement problem caused by packet-by-packet will greatly affect the network performance and is obviously not applicable to environments with strong confrontation and high mobility, which is a problem that does not exist in flow-by-flow. Currently, the flow-by-flow multi-path routing methods are mainly divided into traditional methods, heuristic methods, and deep reinforcement learning methods.
[0003] (1) Traditional Methods
[0004] Traditional multi-path routing methods, such as equal-cost multi-path, play an important role in achieving load balancing and improving network reliability. For example, equal-cost multi-path is widely used in data centers and Internet backbones to effectively utilize redundant links by distributing data packets on paths with equal costs. However, these methods often face challenges in adapting to dynamic network changes, resulting in poor link utilization and congestion, so they have limitations in scenarios that require high adaptability.
[0005] (2) Heuristic Methods
[0006] To overcome the deficiencies of traditional methods, researchers have explored heuristic methods to provide more flexible and adaptable routing solutions. Heuristic methods, such as ant colony optimization, genetic algorithms, etc., aim to find approximate optimal solutions using exhaustive search methods while reducing computational overhead. These methods are inspired by natural processes or mathematical models and can better adapt to changing network conditions. Related literature 1 (Gu Bin, Lu Jie, Zhang Ce. Hybrid Multipath QoS Routing Protocol Based on Ant Colony Optimization [J]. Mobile Communications, 2023, 47(10): 24-31) proposed a hybrid multipath QoS routing protocol based on ant colony optimization. This protocol introduced multi-constrained QoS metrics and fully considered various service requirements. The simulation results showed that the proposed protocol was superior to other traditional routing protocols in terms of end-to-end delay and packet delivery ratio. Related literature 2 (Kumar R. Hybrid Marine Predators and Border Collie Optimization algorithm for multipath routing in IoT [J]. International Journal of Communication Systems, 2023, 36(15): e5567) proposed a hybrid method combining the marine predator optimization algorithm and the border collie optimization algorithm for optimizing multipath routing in the Internet of Things. This method combines global search and local search, can improve the efficiency of routing, reduce latency and energy consumption, and enhance the robustness of the network, showing better performance than traditional methods. Despite many advantages, heuristic methods also face challenges in terms of computational complexity and scalability. For example, heuristic methods may have a huge computational amount when applied to large-scale networks with a large number of nodes and potential paths. In addition, heuristic methods do not guarantee an optimal solution and may face convergence problems especially in highly dynamic network environments.
[0007] (3) Deep reinforcement learning method
[0008] The deep reinforcement learning method combines the adaptability of reinforcement learning and the function approximation ability of deep neural networks, and can characterize the characteristics of the continuously changing network state in complex environments. Research shows that in terms of minimizing latency and improving bandwidth utilization, the traffic engineering framework based on DRL (Deep Reinforcement Learning), the SDN (Software Defined Network) routing method based on deep reinforcement learning, and the critical flow rerouting method based on reinforcement learning are superior to heuristic methods. However, for the changes in network structure and link state in complex environments, the dimensionality of the network state space that the DRL agent needs to perceive increases sharply, and existing DRL routing methods face the problem of the curse of dimensionality. The training process of a single agent leads to an increase in decision complexity due to the exponential growth of the action space, thereby reducing the accuracy of routing inference. Summary of the Invention
[0009] The technical problem to be solved by the present invention is: aiming at the problem that the existing communication network routing methods are difficult to meet the requirements of real-time and reliable communication in complex environments, a multi-path routing selection method for communication networks based on multi-agent deep reinforcement learning is provided, aiming to implement an intelligent QoS (Quality of Service) control strategy, form a certain anti-interference ability, so as to achieve the optimal matching of service types and forwarding paths, and ensure the real-time and reliable transmission of services in complex environments.
[0010] To solve the above technical problems, the present invention adopts the following technical solutions:
[0011] A multi-path routing selection method for communication networks based on multi-agent deep reinforcement learning includes the following steps:
[0012] S1. Use Mininet and Ryu software to build a software-defined network, and use this network to simulate a communication network. The topological structure of the communication network includes nodes and links between nodes.
[0013] S2. Based on the routing problem of the communication network, establish a multi-path routing algorithm model for the communication network. The model includes agents deployed at nodes, and the agents are multi-agent proximal policy optimization agents.
[0014] S3. Obtain node status, link status, adjacency matrix, and link matrix. After being processed by the multi-path routing algorithm model of the communication network, obtain the corresponding paths. Through the data interaction between the agent and the software-defined network, train the agent to obtain the trained agent, and then obtain the trained multi-path routing algorithm model of the communication network.
[0015] S4. Use the trained communication network multipath routing algorithm model to make path decisions for M heterogeneous services in the communication network, and complete the selection of communication network multipath routing.
[0016] Further, in step S2, the agent includes a hierarchical Actor-Critic network, which includes a high-level Actor network, a low-level Actor network, a high-level Critic network, and a low-level Critic network.
[0017] The high-level Actor network includes an input layer, a general feature extraction layer, and a heterogeneous service-specific policy layer; the low-level Actor network includes a first deep neural network; the high-level Critic network includes a graph neural network and a second deep neural network; the low-level Critic network includes a third deep neural network.
[0018] Construct a high-level reward function R based on the path state information h , and the specific expression is: ;
[0019] where b represents the throughput value of the heterogeneous service, d represents the delay value of the heterogeneous service, l represents the packet loss rate of the heterogeneous service, , and respectively represent the importance weights of b, d, and l.
[0020] Construct a low-level reward function based on the link bandwidth utilization and the high-level reward function , and the specific expression is: ;
[0021] where represents the discount factor; u represents the link bandwidth utilization, , represents the link bandwidth capacity, represents the used bandwidth of the link.
[0022] Further, the heterogeneous services include delay-sensitive services, bandwidth-sensitive services, and reliability-sensitive services.
[0023] Further, in step S3, training the agent includes the following:
[0024] S301. Initialize the communication network multipath routing algorithm model, set the initial time step to 0; obtain the heterogeneous service to be selected for multipath routing, the source node of this service is , and the destination node is ; the path starts from the source node .
[0025] S302. The agent at the current node obtains the node state, link state, adjacency matrix, and link matrix at the t-th time step.
[0026] S303. Input the node state at the t-th time step into the hierarchical Actor-Critic network. The node state is transmitted to the general feature extraction layer through the input layer of the high-level Actor network, and feature extraction is performed using a deep neural network to obtain a general node feature vector. This vector passes through the heterogeneous service-specific policy layer, and the attention distribution corresponding to the heterogeneous service is obtained using the graph multi-head attention mechanism.
[0027] The attention distribution includes the weight of the current node selecting the j-th node of the neighbor as the next hop. The calculation formula of the weight is: ;
[0028] where, represents the attention weight of the current node to the j-th node, represents the total number of attention heads, represents the set of neighbor nodes of the current node, n represents the n-th attention head, , both represent the learning parameters of the n-th attention head of the current node for the -th heterogeneous service type, represents transpose of, represents the activation function, represents the feature vector of the current node, represents the feature vector of the j-th node, represents the -th node's feature vector, exp represents the natural exponential function operation.
[0029] According to the attention distribution and the corresponding adjacency matrix, the high-level action corresponding to the heterogeneous service and the logarithmic probability corresponding to this action are obtained, and this action is used as the next-hop node .
[0030] The link state at the t-th time step and the next-hop node jointly pass through the low-level Actor network, feature extraction is performed using the first deep neural network, and the probability distribution of the link type is calculated. According to this probability distribution and the corresponding link matrix, the low-level action and the logarithmic probability corresponding to this action are obtained, and this action is used as the first link .
[0031] S304. Combine the next-hop node and the first link to obtain the agent action at the current node , the agent moves to the next-hop node .
[0032] S305. The agent at the next-hop node repeats steps S302 - S304 until it reaches the destination node and then stops, thereby forming path Q. .
[0033] Send path Q to the software-defined network to transmit the traffic flow, measure the throughput, latency, and packet loss rate of the traffic flow, and calculate the high-level reward value and low-level reward value using the high-level reward function and low-level reward function; update the node state, link state, adjacency matrix, and link matrix using the software-defined network to obtain the node state, link state, adjacency matrix, and link matrix at the (t + 1)-th time step, and then obtain the high-level decision sequence and low-level decision sequence corresponding to the agent in path Q. The specific expressions are: ; ;
[0034] where represents the high-level decision sequence of the i-th agent, , , N represents the total number of agents, represents the path end point, represents the node state at the t-th time step, represents the high-level action log probability at the t-th time step, represents the high-level action of the (i + 1)-th agent, represents the high-level reward value at the t-th time step, represents the node state at the (t + 1)-th time step, represents the low-level decision sequence of the i-th agent, represents the link state at the t-th time step, represents the low-level action log probability at the t-th time step, represents the low-level action of the (i + 1)-th agent, represents the low-level reward value at the t-th time step, represents the link state at the (t + 1)-th time step.
[0035] S306. The and in the low-level decision sequence jointly pass through the low-level Critic network, and use the third deep neural network to obtain the corresponding low-level action value estimates and , and calculate the temporal difference error of the low-level action at the t-th time step. The specific calculation formula is:
[0036] 。
[0037] Recursively calculate the advantage value of the low-level action at the t-th time step , and the specific calculation formula is: ;
[0038] Among them, , T represents the number of training batches; represents the generalized advantage estimation decay factor; represents the advantage value of the low-level action at the (t + 1)-th time step; represents the advantage value of the low-level action at the T-th time step.
[0039] S307. Based on the advantage value of the low-level action at the t-th time step , calculate the loss gradient and update the model parameters of the low-level Actor network and the low-level Critic network.
[0040] S308. Update the time step, and repeat steps S306 to S307 until the set maximum number of training times is reached and stop, to obtain the trained low-level Actor network and low-level Critic network.
[0041] S309. Freeze the model parameters of the trained low-level Actor network and low-level Critic network, set the time step to 0; in the high-level decision sequence and pass through the high-level Critic network together, and use the graph neural network and the second deep neural network to obtain the corresponding high-level action value estimation and , and calculate the temporal difference error of the high-level action at the t-th time step , and the specific calculation formula is:
[0042] .
[0043] Recursively calculate the advantage value of the high-level action at the t-th time step , and the specific calculation formula is: ;
[0044] Among them, represents the advantage value of the high-level action at the (t + 1)-th time step, represents the advantage value of the high-level action at the T-th time step.
[0045] S310. Based on the advantage value of the high-level action at the t-th time step , calculate the loss gradient and update the model parameters of the high-level Actor network and the high-level Critic network.
[0046] S311. Update the time step, and repeat steps S309 to S310 until the set maximum number of training times is reached and stop, to obtain the trained high-level Actor network and high-level Critic network; complete the training of the hierarchical Actor-Critic network.
[0047] Further, in step S3, the node state is a matrix with a shape of , where V represents the total number of nodes; the state vector of the j-th node of the i-th agent is expressed as , and the specific expression is: ; ; ; ; ; ; ; ; ; Among them, S Nij1 represents the normalized value of the service flow bandwidth requirement, S Nij2 represents the normalized value of the path hop count obtained by the shortest path algorithm, S Nij3 represents the normalized value of the maximum available bandwidth of the path obtained by the shortest path algorithm, S Nij4 represents the normalized value of the minimum cumulative packet loss rate of the path obtained by the shortest path algorithm, S Nij5 represents the normalized value of the path hop count of the path obtained by the widest path algorithm, S Nij6 represents the normalized value of the maximum available bandwidth of the path obtained by the widest path algorithm, S Nij7 represents the normalized value of the minimum cumulative packet loss rate of the path obtained by the widest path algorithm, S Nij8 represents the shortest path indicator, S Nij9 represents the widest path indicator, represents the transmission rate of the service request, represents the path distance from the i-th agent to the j-th node obtained by using the minimum hop count algorithm, represents the maximum available bandwidth of the path from the i-th agent to the j-th node obtained by using the minimum hop count algorithm, represents the minimum cumulative packet loss rate of the path from the i-th agent to the j-th node obtained by using the minimum hop count algorithm, represents the lowest packet loss rate of the link between the i-th agent and the j-th node represents the path distance from the i-th agent to the j-th node obtained by using the widest path algorithm, represents the maximum available bandwidth of the path from the i-th agent to the j-th node obtained by using the widest path algorithm, represents the minimum cumulative packet loss rate of the path from the i-th agent to the j-th node obtained by using the widest path algorithm, represents the path distance from the i-th agent to the m-th node obtained by using the minimum hop count algorithm, represents the maximum available bandwidth of the path from the i-th agent to the m-th node obtained by using the widest path algorithm, and min() represents the minimum value function.
[0048] The link state is a matrix with the shape of , where L represents the total number of link types; the state vector of the k-th link type of the i-th agent is represented as , and the specific calculation formula is: ; ; ; Among them, represents the normalized value of the remaining bandwidth of the link, represents the normalized value of the link delay, represents the normalized value of the link packet loss rate, represents the bandwidth capacity of the k-th link type connected to the i-th agent, represents the bandwidth usage of the k-th link type connected to the i-th agent, represents the delay of the k-th link type connected to the i-th agent, represents the packet loss rate of the k-th link type connected to the i-th agent.
[0049] The adjacency matrix is a matrix with the shape of .
[0050] The link matrix is a matrix with the shape of .
[0051] Furthermore, in step S307, updating the model parameters of the lower-layer network includes the following:
[0052] Calculating the gradient of the lower-layer action loss , and the specific formula is: ;
[0053] Among them, clip represents the clipping function; Denote the probability ratio of the old and new low-level actions at the $t$-th time step. , Denote the log probability of the low-level action at the $(t - 1)$-th time step; Denote the clipping parameter.
[0054] Calculate the low-level value loss gradient , and the specific formula is:
[0055] .
[0056] Calculate the total loss gradient , and backpropagate and store it in the agent. The specific formula is: ;
[0057] where, Denote the value loss coefficient, Denote the entropy regularization weight, Denote the entropy of the first policy distribution.
[0058] According to , use the Adam optimizer to inversely update the model parameters of the low-level Actor network and the low-level Critic network.
[0059] Furthermore, in step S310, updating the model parameters of the high-level network includes the following:
[0060] Calculate the high-level action loss gradient , and the specific formula is: ;
[0061] where, Denote the probability ratio of the old and new high-level actions at the $t$-th time step, , Denote the log probability of the high-level action at the $(t - 1)$-th time step.
[0062] Calculate the high-level value loss gradient , and the specific formula is:
[0063] .
[0064] Calculate the high-level total loss gradient , and backpropagate and store it in the agent. The specific formula: ;
[0065] where, Denote the entropy of the second policy distribution.
[0066] According to , the model parameters of the high-level Actor network and the high-level Critic network are updated backward using the Adam optimizer.
[0067] Further, in step S4, the multi-path routing selection of the communication network includes the following:
[0068] Based on the high-level decision sequence and the low-level decision sequence obtained in step S305, and are input into the trained hierarchical Actor-Critic network, and steps S303 - S305 are repeated until routing paths are obtained for M heterogeneous services and then stop.
[0069] Further, the present invention also proposes an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, the steps of the communication network multi-path routing selection method based on multi-agent deep reinforcement learning are implemented.
[0070] Further, the present invention also proposes a computer-readable storage medium. The computer-readable storage medium stores a computer program, and when the computer program is run by a processor, it executes the communication network multi-path routing selection method based on multi-agent deep reinforcement learning.
[0071] Compared with the prior art, the present invention adopts the above technical solutions and has the following technical effects:
[0072] Under delay-sensitive services, the average end-to-end delay of the present invention is significantly reduced; under large-capacity services, i.e., bandwidth-sensitive services, the average network throughput is significantly improved; under reliability-sensitive services, the packet loss rate is significantly reduced. In a bandwidth-constrained scenario, the throughput of large-capacity services remains at a relatively high level. Therefore, the present invention has a good strategy for the QoS requirements of heterogeneous services in a communication network, has a certain anti-interference ability, and can ensure the real-time and reliable transmission of services in a complex environment. BRIEF DESCRIPTION OF THE DRAWINGS
[0073] Figure 1 is the overall implementation flowchart of the present invention.
[0074] Figure 2 is the Actor network model diagram of the present invention.
[0075] Figure 3 is the Critic network model diagram of the present invention.
[0076] Figure 4 is the communication network topology diagram generated based on Mininet in the embodiment of the present invention.
[0077] Figure 5It is a comparative analysis diagram of ablation experiments in an embodiment of the present invention.
[0078] Figure 6 It is a comparison diagram of the average end-to-end delay of delay-sensitive services in an embodiment of the present invention.
[0079] Figure 7 It is a comparison diagram of the average network throughput of bandwidth-sensitive services in an embodiment of the present invention.
[0080] Figure 8 It is a comparison diagram of the average packet loss rate of reliability-sensitive services in an embodiment of the present invention.
[0081] Figure 9 It is a comparison diagram of the network throughput of bandwidth-sensitive services in a bandwidth-constrained scenario in an embodiment of the present invention. Detailed implementation manners
[0082] The present invention will be further described below with reference to the accompanying drawings. The following embodiments are only used to more clearly illustrate the technical solutions of the present invention and cannot be used to limit the protection scope of the present invention.
[0083] To achieve the above object, the present invention proposes a multi-path routing selection method for a communication network based on multi-agent deep reinforcement learning, as Figure 1 shown, and the specific steps are as follows:
[0084] S1. Use Mininet and Ryu software to construct an SDN (Software Defined Network), which includes a series of programmable switches, routers and links, and use this network to simulate a complex communication network. The topological structure of the communication network includes nodes and links between nodes.
[0085] S2. Based on the routing problem of the communication network, establish a multi-path routing algorithm model for the communication network. The model includes agents deployed at nodes, and the agents are MAPPO (Multi-Agent Proximal Policy Optimization) agents. The specific content is as follows:
[0086] The agent includes a hierarchical Actor-Critic network, which includes a high-level Actor network, a low-level Actor network, a high-level Critic network and a low-level Critic network. The high-level Actor network is used to select the next-hop node of the current node, the low-level Actor network is used to select the link type between the next-hop node, the high-level Critic network is used to output the value of the decision of the high-level Actor network, and the low-level Critic network is used to output the value of the decision of the low-level Actor network.
[0087] The high-level Actor network includes an input layer, a general feature extraction layer, and a heterogeneous service-specific policy layer. The low-level Actor network includes a first deep neural network. The high-level Critic network includes a graph neural network and a second deep neural network. The low-level Critic network includes a third deep neural network.
[0088] Among them, according to the differences in the Quality of Service (QoS) requirements of heterogeneous services in the communication network, the heterogeneous services are divided into three categories: the first category is delay-sensitive services, the second category is large-capacity services, i.e., bandwidth-sensitive services, and the third category is reliability-sensitive services.
[0089] Construct a high-level reward function R based on the path state information h , and the specific expression is: ;
[0090] where b represents the throughput value of the heterogeneous service, d represents the delay value of the heterogeneous service, l represents the packet loss rate of the heterogeneous service, , and respectively represent the importance weights of b, d, and l.
[0091] The values of the importance weights depend on the types of heterogeneous services in the communication network. For delay-sensitive services, , and are 0, 1, and 0 respectively; for bandwidth-sensitive services, , and are 1, 0, and 0 respectively; for reliability-sensitive services, , and are 0, 0, and 1 respectively.
[0092] Construct a low-level reward function based on the link bandwidth utilization rate and the high-level reward function , and the specific expression is: ;
[0093] where represents the discount factor; u represents the link bandwidth utilization rate, , represents the link bandwidth capacity, represents the used bandwidth of the link.
[0094] S3. Obtain the node status, link status, adjacency matrix, and link matrix. After being processed by the communication network multipath routing algorithm model, the corresponding paths are obtained. Through the data interaction between the agent and SDN, the agent is trained to obtain the trained agent, and then the trained communication network multipath routing algorithm model is obtained. The specific content is as follows:
[0095] S301. Initialize the communication network multipath routing algorithm model, and set the initial time step to 0; obtain the heterogeneous services to be selected for multipath routing. The source node of this service is , and the destination node is ; The path starts from the source node .
[0096] S302. The agent at the current node obtains the node status, link status, adjacency matrix, and link matrix at the t-th time step.
[0097] Among them, the node status is a matrix with a shape of , where V represents the total number of nodes; the status vector of the j-th node of the i-th agent is , and the specific expression is: ; ; ; ; ; ; ; ; ; Among them, S Nij1 represents the normalized value of the service flow bandwidth requirement, S Nij2 represents the normalized value of the path hop count obtained by the shortest path algorithm, S Nij3 represents the normalized value of the maximum available bandwidth of the path obtained by the shortest path algorithm, S Nij4 represents the normalized value of the minimum cumulative packet loss rate of the path obtained by the shortest path algorithm, S Nij5 represents the normalized value of the path hop count obtained by the widest path algorithm, S Nij6 represents the normalized value of the maximum available bandwidth of the path obtained by the widest path algorithm, S Nij7 represents the normalized value of the minimum cumulative packet loss rate of the path obtained by the widest path algorithm, S Nij8 represents the shortest path indicator, and S Nij9 represents the widest path indicator. represents the transmission rate of the service request, represents the path distance from the i-th agent to the j-th node obtained by the minimum-hop algorithm, represents the maximum available bandwidth of the path from the i-th agent to the j-th node obtained by the minimum-hop algorithm, represents the minimum cumulative packet loss rate of the path from the i-th agent to the j-th node obtained by the minimum-hop algorithm, represents the minimum packet loss rate of the link between the i-th agent and the j-th node, represents the path distance from the i-th agent to the j-th node obtained by the widest-path algorithm, represents the maximum available bandwidth of the path from the i-th agent to the j-th node obtained by the widest-path algorithm, represents the minimum cumulative packet loss rate of the path from the i-th agent to the j-th node obtained by the widest-path algorithm, represents the path distance from the i-th agent to the m-th node obtained by the minimum-hop algorithm, represents the maximum available bandwidth of the path from the i-th agent to the m-th node obtained by the widest-path algorithm, and min() represents the minimum value function.
[0098] S Nij8 reflects whether the current link is part of the shortest path routing, S Nij9 reflects whether the current link is part of the widest path routing. If there is no connection relationship between the i-th agent and the j-th node, then .
[0099] The link state is a matrix of shape , where L represents the total number of link types; the state vector of the k-th link type of the i-th agent is represented as , and the specific calculation formula is: ; ; ; where, represents the normalized value of the remaining bandwidth of the link, represents the normalized value of the link delay, represents the normalized value of the link packet loss rate, represents the bandwidth capacity of the k-th link type connected to the i-th agent, represents the bandwidth usage of the k-th link type connected to the i-th agent, represents the delay of the k-th link type connected to the i-th agent, Represents the packet loss rate of the k-th link type connected to the i-th agent.
[0100] The adjacency matrix is a matrix of shape , where each element represents whether there is a connection relationship between two nodes. If there is a connection relationship, the value of the element is 1; otherwise, it is 0. The connection relationship between a node and other nodes can be obtained by indexing a certain node.
[0101] The link matrix is a matrix of shape , where each element is a vector of length L, representing which links exist between two nodes. For the existing links, the corresponding bit value is 1; otherwise, it is 0. The link information between a node and other nodes can be obtained by indexing a certain node.
[0102] S303. Input the node state at the t-th time step into the hierarchical Actor-Critic network, as shown in Figure 2 . Through the input layer of the high-level Actor network, the node state is transmitted to the general feature extraction layer, and feature extraction is performed using a deep neural network to obtain a general node feature vector. This vector passes through the heterogeneous service-specific policy layer, and the attention distribution corresponding to the heterogeneous service is obtained using the graph multi-head attention mechanism.
[0103] The attention distribution includes the weight of the current node selecting the j-th node of the neighbor as the next hop. The calculation formula of the weight is: ;
[0104] where, represents the attention weight of the current node to the j-th node, represents the total number of attention heads, represents the set of neighbor nodes of the current node, n represents the n-th attention head, , both represent the learning parameters of the n-th attention head of the current node for the -th heterogeneous service type, represents transpose, represents the activation function, represents the feature vector of the current node, represents the feature vector of the j-th node, represents the -th node's feature vector, and exp represents the natural exponential function operation.
[0105] According to the attention distribution and the corresponding adjacency matrix, obtain the high-level action corresponding to the heterogeneous service and the logarithmic probability corresponding to this action, and use this action as the next-hop node .
[0106] Link state and next-hop node at the t-th time step They jointly pass through the lower-level Actor network, use the first deep neural network for feature extraction, calculate the probability distribution of link types, obtain the lower-level action and the logarithmic probability corresponding to this action according to this probability distribution and the corresponding link matrix, and take this action as the first link .
[0107] S304. Combine the next-hop node and the first link to obtain the agent action at the current node , and the agent moves to the next-hop node .
[0108] S305. The agent at the next-hop node repeats steps S302 - S304 until it stops when reaching the destination node , thus forming a path Q .
[0109] Send the path Q to the SDN (Software Defined Network), transmit the service flow, measure the throughput, delay, and packet loss rate of the service flow, and calculate the high-level reward value and the low-level reward value using the high-level reward function and the low-level reward function; use the SDN to update the node state, link state, adjacency matrix, and link matrix to obtain the node state, link state, adjacency matrix, and link matrix at the (t + 1)-th time step, and further obtain the high-level decision sequence and low-level decision sequence corresponding to the agent in the path Q. The specific expressions are as follows: ; ; where represents the high-level decision sequence of the i-th agent , , N represents the total number of agents represents the end point of the path represents the node state at the t-th time step represents the logarithmic probability of the high-level action at the t-th time step represents the high-level action of the (i + 1)-th agent represents the high-level reward value at the t-th time step represents the node state at the (t + 1)-th time step represents the low-level decision sequence of the i-th agent represents the link state at the t-th time step represents the logarithmic probability of the low-level action at the t-th time step Represents the low-level action of the (i + 1)-th agent, Represents the low-level reward value at the t-th time step, Represents the link state at the (t + 1)-th time step.
[0110] S306, as Figure 3 shown, in the low-level decision sequence, and jointly pass through the low-level Critic network, and use the third deep neural network to obtain the corresponding low-level action value estimate and , and calculate the temporal difference error of the low-level action at the t-th time step , and the specific calculation formula is:
[0111] .
[0112] Recursively calculate the advantage value of the low-level action at the t-th time step , and the specific calculation formula is: ; where, , T represents the number of training batches, taking 128; represents the generalized advantage estimation decay factor, taking 0.5; represents the advantage value of the low-level action at the (t + 1)-th time step; represents the advantage value of the low-level action at the T-th time step.
[0113] S307, based on the advantage value of the low-level action at the t-th time step, calculate the loss gradient and update the model parameters of the low-level Actor network and the low-level Critic network; the specific content is:
[0114] Calculate the low-level action loss gradient , and the specific formula is: ; where, clip represents the clipping function, used to limit the range of a value; represents the probability ratio of the low-level old and new actions at the t-th time step, , represents the logarithm probability of the low-level action at the (t - 1)-th time step; represents the clipping parameter, taking 0.1.
[0115] Calculate the low-level value loss gradient , and the specific formula is:
[0116] .
[0117] Calculate the total loss gradient , and backpropagate and store in the agent. The specific formula is: ; Among them, represents the value loss coefficient, taking 0.5; represents the entropy regularization weight, taking 0.01; represents the first policy distribution entropy.
[0118] According to , use the Adam optimizer to inversely update the model parameters of the lower-level Actor network and the lower-level Critic network.
[0119] S308. Update the time step, repeat steps S306 to S307, and save the model parameters of the lower-level Actor network and the lower-level Critic network every time after a specific number of steps of update until the set maximum number of training times is reached and stop, obtaining the trained lower-level Actor network and lower-level Critic network.
[0120] S309. Freeze the model parameters of the trained lower-level Actor network and the lower-level Critic network, set the time step to 0; in the high-level decision sequence, and jointly pass through the high-level Critic network, and use the graph neural network and the second deep neural network to obtain the corresponding high-level action value estimate and , and calculate the temporal difference error of the high-level action at the t-th time step. The specific calculation formula is:
[0121] .
[0122] Recursively calculate the advantage value of the high-level action at the t-th time step. The specific calculation formula is:
[0123] ; Among them, represents the advantage value of the high-level action at the (t + 1)-th time step, represents the advantage value of the high-level action at the T-th time step.
[0124] S310. Based on the advantage value of the high-level action at the t-th time step, calculate the loss gradient and update the model parameters of the high-level Actor network and the high-level Critic network; the specific content is:
[0125] Calculate the high-level action loss gradient The specific formula is:
[0126] ; Among them, represents the probability ratio of the high-level old and new actions at the t-th time step, , represents the logarithmic probability of the high-level action at the (t - 1)-th time step.
[0127] Calculate the high-level value loss gradient , and the specific formula is:
[0128] .
[0129] Calculate the high-level total loss gradient , and backpropagate and store it in the agent. The specific formula:
[0130] ; Among them, represents the second policy distribution entropy.
[0131] According to , use the Adam optimizer to inversely update the model parameters of the high-level Actor network and the high-level Critic network.
[0132] S311. Update the time step, repeat steps S309 to S310, and save the model parameters of the high-level Actor network and the high-level Critic network every time a specific number of steps are updated, until the set maximum number of training times is reached and stop, to obtain the trained high-level Actor network and high-level Critic network; thus, the training of the hierarchical Actor-Critic network is completed.
[0133] S4. Based on the high-level decision sequence and the low-level decision sequence obtained in step S305, input and into the trained hierarchical Actor-Critic network, repeat steps S303 - S305, and use the trained communication network multipath routing algorithm model to make path decisions for M heterogeneous services in the communication network until routing paths are obtained for the M heterogeneous services and stop, to complete the selection of the communication network multipath routing.
[0134] Example:
[0135] The communication network topology diagram is as Figure 4 shown, with a total of 20 nodes, where S1 and D1 are source nodes, Z1 - Z9 are destination nodes, D2 is the core node, D3 - D5 are first-level control nodes, D6 - D8 are second-level control nodes, and F1 and F2 are auxiliary nodes.
[0136] There are three links between every two nodes. Specifically:
[0137] (1)Three types of links, namely scattering, VHF (Very High Frequency), and UHF (Ultra High Frequency), are set between S1 and D2 - D8, with bandwidths of 8 Mbps, 9.6 Kbps, and 1 Mbps respectively.
[0138] (2)Three types of links, namely wired, satellite, and UHF, are set between D2 and D3 - D8, F1, F2, with bandwidths of 1 Mbps, 8 Mbps, and 1 Mbps respectively. Three types of links, namely wired, satellite, and microwave, are set between D2 and D1, with bandwidths of 1 Mbps, 8 Mbps, and 8 Mbps respectively.
[0139] (3)Three types of links, namely wired, satellite, and microwave, are set between D3 - D5, with bandwidths of 1 Mbps, 8 Mbps, and 8 Mbps respectively. Three types of links, namely area - wide, VHF, and UHF, are set between D3 and Z1 - Z2, D4 and Z3 - Z4, and D5 and Z5 - Z6, with bandwidths of 10 Mbps, 9.6 Kbps, and 1 Mbps respectively.
[0140] (4)Three types of links, namely wired, satellite, and microwave, are set between D6 - D8, with bandwidths of 1 Mbps, 8 Mbps, and 8 Mbps respectively. Three types of links, namely area - wide, VHF, and UHF, are set between D6 and Z7, D7 and Z8, and D8 and Z9, with bandwidths of 10 Mbps, 9.6 Kbps, and 1 Mbps respectively.
[0141] (5)Three types of links, namely wired, satellite, and area - wide, are set between D3 - D8, F1, F2 and D1, with bandwidths of 1 Mbps, 8 Mbps, and 10 Mbps respectively.
[0142] As Figure 5 shown, in deep reinforcement learning, the reward value is an important indicator parameter reflecting the method performance. In the experiment, three types of services are set as delay - sensitive service, bandwidth - sensitive service, and reliability - sensitive service. The transmission rate of the delay - sensitive service is 100 Kbps, the transmission rate of the bandwidth - sensitive service is 1500 Kbps, and the transmission rate of the reliability - sensitive service is 500 Kbps. Their distribution is [0.3, 0.4, 0.3]. The hierarchical Actor - Critic network proposed in the present invention and the network after removing the heterogeneous service - specific policy layer are trained respectively, and the average reward values of the high - level are compared. By Figure 5It can be seen that as the number of iterations increases, the reward values obtained by the two networks in each round gradually increase until convergence. Initially, the hierarchical Actor-Critic network proposed in the present invention has a more obvious improvement in the convergence speed compared to the network after removing the heterogeneous service dedicated policy layer. The network after removing the heterogeneous service dedicated policy layer gradually converges to a stable state after 100 rounds, while the hierarchical Actor-Critic network proposed in the present invention reaches convergence in only 50 to 60 rounds, and the convergence speed is increased by nearly 50%. This is because the graph multi-head attention mechanism in the model increases the finer-grained attention to neighbor nodes, enabling the network to better understand different information in the state space during the initial training stage, thus accelerating the improvement of the network; in terms of the reward level after final convergence, it can be seen that the reward value level at which the hierarchical Actor-Critic network proposed in the present invention finally converges is significantly higher than that of the network after removing the heterogeneous service dedicated policy layer. The convergence reward value of the hierarchical Actor-Critic network proposed in the present invention is stable at about -0.5, while the network after removing the heterogeneous service dedicated policy layer is stable at around -1.0, indicating that the effect of the network with the added heterogeneous service dedicated policy layer has been improved, which can better aggregate graph structure information, avoid falling into local optimal solutions, and improve the quality of policy selection.
[0143] To verify the advantages of the method proposed in the present invention, experiments on the QoS requirements of delay-sensitive services, bandwidth-sensitive services, and reliability-sensitive services were conducted using the traditional communication network traditional routing method, i.e., the OSPF (Open Shortest Path First) method, the traditional multi-path routing method, i.e., the ECMP (Equal-cost multi-path routing) method, and the method proposed in the present invention, as Figure 6 , Figure 7 and Figure 8 shown. The experimental scenario is that S1 and D1 send these three services to Z1-Z9 respectively. The sending rate of the delay-sensitive service is 50 Kbps, the sending rate of the bandwidth-sensitive service is 1500 Kbps, and the sending rate of the reliability-sensitive service is 500 Kbps. In addition, the continuous time slots of the services are set to 10, 30, and 50 respectively to simulate three scenarios of light load, medium load, and heavy load.
[0144] Figure 6 is the comparison graph of the average end-to-end delay of the three methods under the delay-sensitive service. It can be seen that the end-to-end delay of the method of the present invention remains at a small value in the three scenarios, with the maximum reduction of 89.03% and 75.87% compared to the other two methods. Figure 7Comparison chart of average network throughput of three methods under bandwidth-sensitive services. It can be seen that the method of the present invention has been improved compared with the other two methods in the three scenarios, with the maximum increases of 18.44% and 11.21% respectively. Figure 8 It is a comparison chart of average packet loss rates of three methods under reliability-sensitive services. It can be seen that the packet loss rates of the method of the present invention remain at a low level in the three scenarios. Compared with the other two methods, the maximum decreases are 94.12% and 86.59% respectively. Thus, it can be seen that the method of the present invention has good strategies for the QoS requirements of different service types and can meet the QoS requirements of various heterogeneous services in the communication network.
[0145] To further verify the anti-interference ability of the method proposed by the present invention, a scenario where the link bandwidth becomes narrower under interference is simulated, and the QoS requirements of high-capacity services are experimentally tested in this scenario. The experimental scenario is as follows: Set the transmission rate of the bandwidth-sensitive service to 1500 Kbps. S1 continuously sends bandwidth-sensitive services to Z1. At the 4th time slot, the scattering links between S1 and D2-D5 and the satellite links between D2 and D3-D5 are reduced to less than the transmission rate of the bandwidth-sensitive service, which is 1000 Kbps, so as to simulate the situation where the transmission bandwidth of the bandwidth-sensitive service is limited. Use the OSPF method, the ECMP method, and the method proposed by the present invention to obtain the network throughput of the bandwidth-sensitive service in the above experimental scenario and compare the results, as Figure 9 shown.
[0146] From Figure 9 it can be seen that after the link bandwidth becomes narrower, the network throughputs of the three methods all decrease. Among them, the decrease amplitude of the method proposed by the present invention is much smaller than that of the other two methods, and the throughput of the method proposed by the present invention remains at a relatively high level. Thus, it can be seen that the method proposed by the present invention has good effects on the transmission of bandwidth-sensitive services in the scenario of limited bandwidth and forms a certain anti-interference ability.
[0147] The embodiment of the present invention also proposes an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. It should be noted that when the processor executes the computer program, it corresponds to the specific steps of the method provided by the embodiment of the present invention and has the corresponding functional modules and beneficial effects of the execution method. For the technical details not described in detail in this embodiment, reference can be made to the method provided by the embodiment of the present invention.
[0148] An embodiment of the present invention also provides a computer-readable storage medium, which stores a computer program. It should be noted that when the computer program is run by a processor, it corresponds to the specific steps of the method provided by the embodiment of the present invention, and has the corresponding functional modules and beneficial effects of the execution method. For technical details not described in detail in this embodiment, reference may be made to the method provided by the embodiment of the present invention.
[0149] The above are only the preferred embodiments of the present invention. It should be pointed out that for those of ordinary skill in the art, without departing from the technical principle of the present invention, several improvements and modifications can be made, and these improvements and modifications should also be regarded as the protection scope of the present invention.
Claims
1. A communication network multipath routing selection method based on multi-agent deep reinforcement learning, characterized in that: include: S1. Use Mininet and Ryu software to build a software-defined network, and use the network to simulate a communication network. The topology of the communication network includes nodes and links between nodes. S2. Based on the routing problem of the communication network, a communication network multipath routing algorithm model is established, wherein the model includes an intelligent agent deployed at a node, and the intelligent agent is a multi-agent proximal strategy optimization intelligent agent; S3, obtaining the node state, link state, adjacency matrix and link matrix, and obtaining the corresponding path through the processing of the communication network multipath routing algorithm model, training the agent through data interaction between the agent and the software defined network, obtaining the trained agent, and then obtaining the trained communication network multipath routing algorithm model; S4. Use the trained communication network multipath routing algorithm model to make path decisions for M heterogeneous services in the communication network and complete the selection of multipath routing in the communication network.
2. The communication network multipath routing selection method based on multi-agent deep reinforcement learning according to claim 1 is characterized in that: In step S2, the agent includes a hierarchical Actor-Critic network, which includes a high-level Actor network, a low-level Actor network, a high-level Critic network, and a low-level Critic network; The high-level Actor network includes an input layer, a general feature extraction layer, and a heterogeneous business-specific strategy layer; the low-level Actor network includes the first deep neural network; The high-level critic network includes a graph neural network and a second deep neural network; the low-level critic network includes a third deep neural network; Constructing a high-level reward function R h , the specific expression is: ; Wherein, b represents the throughput value of heterogeneous services, d represents the delay value of heterogeneous services, and l represents the packet loss rate of heterogeneous services. , and Represent the importance weights of b, d, and l respectively; Building a low-level reward function , the specific expression is: ; in, represents the discount factor; u represents the link bandwidth utilization, , represents the link bandwidth capacity, Indicates the used bandwidth of the link.
3. The communication network multipath routing selection method based on multi-agent deep reinforcement learning according to claim 1 is characterized in that: Heterogeneous services include delay-sensitive services, bandwidth-sensitive services, and reliability-sensitive services.
4. The communication network multipath routing selection method based on multi-agent deep reinforcement learning according to claim 2 is characterized in that: In step S3, training the agent includes the following: S301, initialize the communication network multipath routing algorithm model, set the initial time step to 0; obtain the heterogeneous service to be selected by multipath routing, the source node of the service is , the destination node is ; Path from source node start; S302, the agent at the current node obtains the node state, link state, adjacency matrix and link matrix at the tth time step; S303, input the node state at the tth time step into the hierarchical Actor-Critic network, transmit the node state to the general feature extraction layer through the input layer of the high-level Actor network, use the deep neural network to extract features, and obtain a general node feature vector. This vector passes through the heterogeneous business dedicated strategy layer, and uses the graph multi-head attention mechanism to obtain the attention distribution corresponding to the heterogeneous business; The attention distribution includes the weight of the current node selecting the jth node of its neighbor as the next hop. The weight is calculated as: ; in, represents the attention weight of the current node to the jth node, represents the total number of attention heads, represents the set of neighbor nodes of the current node, n represents the nth attention head, , Both represent the attention head of the nth node at the current node to the learning parameters for heterogeneous service types, express The transpose of represents the activation function, represents the feature vector of the current node, represents the feature vector of the jth node, Indicates The feature vector of a node, exp represents the natural exponential function operation; According to the attention distribution and the corresponding adjacency matrix, the high-level action corresponding to the heterogeneous business and the logarithmic probability corresponding to the action are obtained, and the action is used as the next hop node ; Link status and next hop node at time step t Through the low-level Actor network, the first deep neural network is used to extract features and calculate the probability distribution of link types. According to the probability distribution and the corresponding link matrix, the logarithmic probability of the low-level action and the action is obtained, and the action is used as the first link. ; S304: Next hop node and the first link Combine to get the agent action at the current node , the agent moves to the next hop node ; S305, at the next hop node The agent repeats steps S302-S304 until it reaches the destination node When it stops, it forms a path Q. ; Send the path Q to the software-defined network, transmit the business flow, measure the throughput, delay and packet loss rate of the business flow, and use the high-level reward function and the low-level reward function to calculate the high-level reward value and the low-level reward value; use the software-defined network to update the node state, link state, adjacency matrix and link matrix, and obtain the node state, link state, adjacency matrix and link matrix at the t+1 time step, and then obtain the high-level decision sequence and low-level decision sequence corresponding to the agent in the path Q. The specific expression is: ; ; in, represents the high-level decision sequence of the ith agent, , , N represents the total number of agents, Indicates the end point of the path. represents the node state at the tth time step, represents the log-probability of the high-level action at the t-th time step, represents the high-level action of the i+1th agent, represents the high-level reward value at the t-th time step, represents the node state at the t+1th time step, represents the low-level decision sequence of the ith agent, represents the link status at the tth time step, represents the logarithmic probability of the low-level action at the t-th time step, represents the low-level action of the i+1th agent, represents the low-level reward value at the t-th time step, represents the link status at the t+1th time step; S306, Low-level decision sequence and Through the low-level Critic network, the corresponding low-level action value estimation is obtained using the third deep neural network and , and calculate the time difference error of the low-level action at the tth time step , the specific calculation formula is: ; Recursively calculate the advantage value of the low-level action at the tth time step , the specific calculation formula is: ; in, , T represents the number of training batches; represents the attenuation factor of the generalized advantage estimate; represents the advantage value of the low-level action at the t+1 time step; represents the advantage value of the low-level action at the Tth time step; S307, advantage value based on the low-level action at time step t , calculate the loss gradient and update the model parameters of the low-level Actor network and the low-level Critic network; S308, update the time step, repeat steps S306 to S307, and stop until the set maximum number of training times is reached, to obtain the trained low-level Actor network and low-level Critic network; S309, freeze the model parameters of the low-level Actor network and the low-level Critic network after training, set the time step to 0; and Through the high-level Critic network, the graph neural network and the second deep neural network are used to obtain the corresponding high-level action value estimation. and , and calculate the time difference error of the high-level action at the tth time step , the specific calculation formula is: ; Recursively calculate the advantage value of the high-level action at the tth time step , the specific calculation formula is: ; in, represents the advantage value of the high-level action at the t+1 time step, represents the advantage value of the high-level action at the Tth time step; S310, advantage value based on high-level action at time step t , calculate the loss gradient and update the model parameters of the high-level Actor network and the high-level Critic network; S311, update the time step, repeat steps S309 to S310, and stop when the set maximum number of training times is reached, to obtain the trained high-level Actor network and high-level Critic network; complete the training of the hierarchical Actor-Critic network.
5. The communication network multipath routing selection method based on multi-agent deep reinforcement learning according to claim 4 is characterized in that: In step S3, the node state is of shape , V represents the total number of nodes; the state vector of the jth node of the i-th agent Expressed as , the specific expression is: ; ; ; ; ; ; ; ; ; Among them, S Nij1 Indicates the normalized value of the service flow bandwidth requirement, S Nij2 represents the normalized value of the path hop count obtained by the shortest path algorithm, S Nij3 represents the normalized value of the maximum available bandwidth of the path obtained by the shortest path algorithm, S Nij4 represents the normalized value of the minimum cumulative packet loss rate of the path obtained by the shortest path algorithm, S Nij5 represents the normalized value of the path hop count obtained by the widest path algorithm, S Nij6 represents the normalized value of the maximum available bandwidth of the path obtained by the widest path algorithm, S Nij7 represents the normalized value of the minimum cumulative packet loss rate of the path obtained by the widest path algorithm, S Nij8 Represents the shortest path indicator, S Nij9 represents the widest path indicator, Indicates the sending rate of service requests. represents the path distance from the ith agent to the jth node obtained using the minimum hop count algorithm, represents the maximum available bandwidth of the path from the ith agent to the jth node obtained using the minimum hop count algorithm, represents the minimum cumulative packet loss rate of the path from the ith agent to the jth node obtained using the minimum hop count algorithm, represents the minimum packet loss rate of the link between the ith agent and the jth node, represents the path distance from the ith agent to the jth node obtained using the widest path algorithm, represents the maximum available bandwidth of the path from the ith agent to the jth node obtained using the widest path algorithm, represents the minimum cumulative packet loss rate of the path from the i-th agent to the j-th node obtained by the widest path algorithm, represents the path distance from the ith agent to the mth node obtained using the minimum hop count algorithm, It represents the maximum available bandwidth of the path from the ith agent to the mth node obtained by the widest path algorithm, and min() represents the minimum value function; The link state is of the form , L represents the total number of link types; the state vector of the kth link type of the ith agent is Expressed as , the specific calculation formula is: ; ; ; in, Represents the normalized value of the link remaining bandwidth, represents the normalized value of link delay, Indicates the normalized value of the link packet loss rate, represents the bandwidth capacity of the kth link type connected to the ith agent, represents the bandwidth usage of the kth link type connected to the ith agent, represents the delay of the kth link type connected to the ith agent, represents the packet loss rate of the kth link type connected to the i-th agent; The adjacency matrix is of shape Matrix of The link matrix is of shape The matrix of .
6. The communication network multipath routing selection method based on multi-agent deep reinforcement learning according to claim 4 is characterized in that: In step S307, updating the model parameters of the lower layer network includes the following: Compute low-level action loss gradients , the specific formula is: ; Among them, clip represents the clipping function; represents the probability ratio of the low-level new and old actions at the t-th time step, , represents the logarithmic probability of the low-level action at the t-1th time step; Indicates the cropping parameters; Calculate low-level value loss gradients , the specific formula is: ; Calculate the total loss gradient , and back propagated and stored in the agent. The specific formula is: ; in, represents the value loss coefficient, represents the entropy regularization weight, represents the first strategy distribution entropy; according to , use the Adam optimizer to reversely update the model parameters of the low-level Actor network and the low-level Critic network.
7. The communication network multipath routing selection method based on multi-agent deep reinforcement learning according to claim 4 is characterized in that: In step S310, updating the model parameters of the high-level network includes the following: Compute high-level action loss gradients , the specific formula is: ; Among them, clip represents the clipping function; represents the probability ratio of high-level new and old actions at the t-th time step, , represents the logarithmic probability of the high-level action at the t-1th time step; Indicates the cropping parameters; Calculate high-level value loss gradient , the specific formula is: ; Calculate the total loss gradient at high level , and back propagated and stored in the agent, the specific formula is: ; in, represents the value loss coefficient, represents the entropy regularization weight, represents the second strategy distribution entropy; according to , use the Adam optimizer to reversely update the model parameters of the high-level Actor network and the high-level Critic network.
8. The communication network multipath routing selection method based on multi-agent deep reinforcement learning according to claim 4 is characterized in that: In step S4, the communication network multipath routing selection includes the following contents: Based on the high-level decision sequence and low-level decision sequence obtained in step S305, and The information is input into the trained hierarchical Actor-Critic network, and steps S303 to S305 are repeated until routing paths are obtained for M heterogeneous services.
9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the computer program, the steps of the communication network multipath routing selection method based on multi-agent deep reinforcement learning as described in any one of claims 1 to 8 are implemented.
10. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the communication network multipath routing selection method based on multi-agent deep reinforcement learning described in any one of claims 1 to 8 is executed.
Citation Information
Patent Citations
Network autonomous intelligent management and control method based on deep reinforcement learning
CN113328938A
Resource optimization method and system based on deep reinforcement learning under SDN architecture
CN113518039A
Intelligent network path optimization method and system based on deep reinforcement learning
CN116527567A
Software defined network routing method based on deep reinforcement learning
CN116599885A
Adaptive QoS intelligent routing method based on deep reinforcement learning
CN116614437A
Cited By
Cooperative data migration scheduling method based on topology awareness and deep reinforcement learning
CN120469982A
Intelligent beam field multi-torpedo-tank dynamic collaborative material distribution control method and intelligent beam field multi-torpedo-tank dynamic collaborative material distribution control system
CN122222003A