Intelligent cross-domain multicast routing method based on multi-agent deep reinforcement learning
By deploying multi-agent deep reinforcement learning agents in multi-domain software-defined wireless networks, the problem of dynamic changes in multicast group members in large-scale network environments is solved, efficient cross-domain multicast routing construction and optimization is achieved, and network performance and reliability are improved.
Patent Information
- Application Number
- CN202510136753.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-07
- Publication Date
- 2025-05-09
AI Technical Summary
The prior art is difficult to apply to the optimal multicast tree algorithm in multi-domain software-defined wireless networks where multicast group members change dynamically in large-scale network environments.
Using an intelligent cross-domain multicast routing method based on multi-agent deep reinforcement learning, the cross-domain multicast tree is constructed and optimized by deploying inter-domain and intra-domain software-defined wireless networks using multi-agent collaborative learning and policy coordination.
It realizes efficient construction and optimization of cross-domain multicast routing in a large-scale dynamically changing network environment, improves network performance and reliability, and adapts to complex network topology and dynamic traffic requirements.
Smart Images

Figure CN119967538A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of SDWN (Software-Defined Wireless Network), and in particular to an intelligent cross-domain multicast routing method based on multi-agent deep reinforcement learning. Background Art
[0002] With the continuous development of wireless communication technology and the popularization of application scenarios, wireless multicast communication is widely used in various scenarios such as multimedia conferencing, real-time video transmission, team collaboration and distributed computing. Compared with point-to-point communication, multicast communication can transmit the same data to multiple receivers, realize efficient information distribution and sharing in wireless networks, and can effectively save network bandwidth and reduce network load. Designing multicast routing in multicast communication essentially requires constructing an optimal multicast tree from the source node to all destination nodes to achieve preset performance indicators and improve the utilization of network resources.
[0003] SDWN technology separates the network control plane from the data plane, and achieves the acquisition of network status information and global optimization configuration of network resources through centralized management and flexible programming mechanisms, which can well solve the defects of traditional wireless network management methods. However, with the increase of network scale and complexity, the single-controller management mode (single-domain software-defined wireless network) is prone to single point failure, poor scalability and difficulty in adapting to heterogeneity and other performance bottlenecks. In order to overcome these limitations and improve network performance and reliability, SDWN networks are usually expanded from a single-controller management mode to a multi-controller management mode (multi-domain software-defined wireless network). The cross-domain multicast routing problem in multi-domain software-defined wireless networks is not only an NP-hard combinatorial optimization problem, but also with the increase of network scale and the dynamic changes of multicast group members, the construction of efficient cross-domain multicast routing paths requires the design of flexible and timely acquisition and maintenance of global network status information and efficient optimal cross-domain multicast tree solution algorithms.
[0004] At present, some classic optimization methods for solving the optimal multicast tree include approximate optimal solution methods based on pruning strategy and greedy search strategy (such as KMB and SCTF), heuristic swarm intelligence optimization methods (such as genetic algorithm GA), but these methods are not flexible enough to adapt to the needs of high-speed dynamic network traffic, and usually only focus on a few optimization performance indicators, lacking global optimization capabilities. As a data-driven solution, deep reinforcement learning algorithm has good learning ability, can extract features and patterns from a large amount of network data, generate optimal multicast routing strategies, thereby improving the performance and effect of multicast routing, and has stronger flexibility than traditional methods to adapt to complex network topologies and dynamically changing traffic requirements. At present, deep reinforcement learning has been used to solve multicast problems (such as routing methods based on reinforcement learning mechanism Q-RL, DRL-M4MR, MADRL-MR methods), but these algorithms discuss single-domain multicast problems, and cross-domain multicast problems need further discussion.
[0005] In addition, most of the deep reinforcement learning algorithms currently used to solve routing optimization problems are online reinforcement learning methods, and their training strategies are divided into two methods: on-policy and off-policy. Both strategies require interaction with the environment. However, in a large-scale network environment, the real-time interaction cost between the agent and the environment is relatively high, which requires a lot of trial and error and exploration during the training process. In addition, online reinforcement learning can usually only use current real-time data, and the utilization rate of historical data is low, which may lead to high resource consumption and training costs. Although offline reinforcement learning usually does not require interaction with the environment, it can better adapt to the scenario of large-scale network environment, reduce the cost of real-time interaction with the environment, and can improve learning effect and training speed by effectively using historical data. However, for a real-time dynamically changing network environment, if offline reinforcement learning is used to solve the routing problem, it is also necessary to solve the sample bias and lack of exploration that may be caused by simply using offline data. Summary of the invention
[0006] The present invention aims to solve the problem that the existing optimal multicast tree algorithm is difficult to apply to multi-domain software-defined wireless networks with dynamically changing multicast group members in a large-scale network environment, and provides an intelligent cross-domain multicast routing method based on multi-agent deep reinforcement learning.
[0007] To solve the above problems, the present invention is achieved through the following technical solutions:
[0008] The intelligent cross-domain multicast routing method based on multi-agent deep reinforcement learning includes the following steps:
[0009] Step 1: deploy an inter-domain agent on the root controller of the multi-domain software-defined wireless network and deploy an intra-domain agent on the local controller of each domain;
[0010] Step 2: Initialize the Actor network parameters θ between domains and within domains int ,θ intra 、Critic network parametersω int ,ω intra and data buffer area B int ,B intra ;
[0011] Step 3: The root controller and each local controller communicate and collect information of the multi-domain software-defined wireless network through the controller communication mechanism, and analyze the information collected by the corresponding controller through the multicast group management module to determine the domains where the source node and each destination node are located, as well as the forwarding boundary nodes of each domain;
[0012] Step 4: First, the inter-domain agent interacts with the environment to generate an inter-domain network link information matrix and the inter-domain multicast tree state matrix And stack the inter-domain network link information matrix and the inter-domain multicast tree state matrix Get the current state s of the inter-domain agent t ; Then, the inter-domain agent changes from the current state s t Sample the current action a from the output action set t , execute the current action a t Get the current reward value r t and the next state s t+1 , where the action set of the inter-domain agent is the set of edges between all domains, and each action is selected as one of the edges; finally, the reinforcement learning data (s t ,a t ,r t ,s t+1 ) is stored in the data cache area B of the inter-domain agent int ;
[0013] Step 5: When the inter-domain agent repeats the action of step 4, the inter-domain agent is trained by offline and online hybrid training, that is:
[0014] Every set period of time, i.e., the inter-domain period, the data cache area b is used first. int The inter-domain agent is trained offline without interacting with the environment using the offline reinforcement learning data in , so as to update the Actor network parameters θ of the inter-domain agent. int and Critic network parameter ω int, and then enter the next inter-domain agent action process;
[0015] When the set inter-domain update frequency is reached, the online reinforcement learning data obtained from this action is first used to conduct online training for the inter-domain agent to interact with the environment, so as to update the Actor network parameters θ of the inter-domain agent. int and Critic network parameter ω int , and then enter the next inter-domain agent action process;
[0016] Step 6: In the inter-domain multicast tree T int When the construction is completed, the inter-domain agent will construct the inter-domain multicast tree T int Synchronize to each agent in the domain;
[0017] Step 7: First, each agent in the domain interacts with the environment to generate the network link information matrix in its own domain. and the intra-domain multicast tree state matrix And stack its own intra-domain network link information matrix and the intra-domain multicast tree state matrix Get the current state s of the agent in the domain t ; Then, each agent ring in the domain starts from the current state s t Sample the current action a from the output action set t , execute the current action a t Get the current reward value r t and the next state s t+1 , where the action set of the agent in the domain is the set of nodes in the domain, and each action is selected as the next hop node; finally, the reinforcement learning data (s t ,a t ,r t ,s t+1 ) is stored in the data buffer area B of the agent in the domain intra ;
[0018] Step 8: When each agent in the domain repeats the action of step 7, each agent in the domain is trained by offline and online hybrid training, that is:
[0019] Every set period of time, i.e., the domain cycle, the data buffer area B is used first during this period of time. intra The offline reinforcement learning data in the domain is used to train the in-domain agent offline without interacting with the environment to update the Actor network parameters θ of the in-domain agent. intra and Critic network parameter ω intra , and then enter the next action process of the agent in the domain;
[0020] When the set domain update frequency is reached, the online reinforcement learning data obtained from this action is first used to conduct online training on the domain agent to interact with the environment, so as to update the Actor network parameters θ of each domain agent. intra and Critic network parameter ω intra , and then enter the next action process of each agent in the domain;
[0021] Step 9: Multicast tree T in each domain intra When all are constructed, the combined inter-domain multicast tree T int and all intra-domain multicast trees T intra Combined into a cross-domain multicast tree;
[0022] Step 10: Determine whether the cross-domain multicast tree converges or reaches a preset number of iterations: If so, the multi-domain software-defined wireless network performs intelligent cross-domain multicast routing based on the current cross-domain multicast tree; otherwise, return to step 4 and perform the next round of iteration.
[0023] In the above steps 4 and 7, the inter-domain network link information matrix and intra-domain network link information matrix The network link information in the ,includes the remaining bandwidth, delay, packet loss rate, packet error rate and the,distance between APs.
[0024] In the above steps 4 and 7, the current action a is performed for the inter-domain agent and the intra-domain agent t Get the current reward value r t :
[0025] If the current action a is executed t After that, the added next hop node or link just adds a normal node to the multicast tree, so the current reward value r t is the single-step reward R part :
[0026] R part =β1bw ij +β2(1-delay ij )+β3(1-loss ij )+β4(1-err ij )+β5(1-dist ij )
[0027] If the next hop node or link added is to add a destination node to the multicast tree k , then the current reward value r t Reward R for the subtask end :
[0028] R end =β1bw k+β2(1-delay k )+β3(1-loss k )+β4(1-err k )+β5(1-dist k )
[0029] If the added next-hop node or link causes the multicast tree to form a loop, the current reward value r t If the current reward value is r t is the penalty value R loop :
[0030] R loop =C1
[0031] If a t is an invalid action, which is neither the next hop node of the current node nor the edge connected to the current domain neighborhood. Then the current reward value r t is the penalty value R hell :
[0032] R hell =C2
[0033] Where β1, β2, β3, β4, and β5 represent the weights of remaining bandwidth, delay, packet loss rate, packet error rate, and distance between APs, respectively; bw ij 、delay ij 、loss ij 、err ij 、dist ij They represent the links e added to the multicast tree. ij The remaining bandwidth, latency, packet loss rate, packet error rate and distance between APs; bw k 、delay k 、loss k 、err k 、dist k Respectively represent the source node src to the destination node d k The remaining bandwidth, delay, packet loss rate, packet error rate and distance between APs of the entire link; C1 and C2 represent two constants respectively.
[0034] In the above steps 5 and 8, the process of using reinforcement learning data to train the agent is as follows: first, the reinforcement learning data (s t ,a t ,r t ,s t+1 ) in the current state s t and the next state s t+1 Input the agent's Critic network to calculate the current expected value V ω (s t) and the next expected value V ω (s t+1 ); Then, based on the reinforcement learning data (s t ,a t ,r t ,s t+1 ) in the current reward value r t , and the current expected value V calculated above ω (s t ) and the next expected value V ω (s t+1 ) Calculate the time series difference residual ψ t , and adopt the learning method of time series difference residual based on the time series difference residual ψ t Update the agent's Critic network parameters ω and Actor network parameters θ.
[0035] In the above steps 5 and 8, the inter-domain period and the intra-domain period are the same or different, and the inter-domain update frequency and the intra-domain update frequency are the same or different.
[0036] Compared with the prior art, the multi-agent deep reinforcement learning intelligent cross-domain multicast routing method (MA-CDMR) proposed in the present invention has the following characteristics:
[0037] 1. In view of the characteristics of multi-domain scenarios of software-defined wireless networks (SDWNs), the present invention conducts a theoretical analysis of the multicast routing problem in software-defined networks (SDWNs) multi-domain scenarios, proposes a multicast problem-solving framework based on network state information perception and multi-agent deep reinforcement learning, decomposes the cross-domain multicast tree solution problem into two sub-problems of inter-domain multicast tree construction and intra-domain multicast tree construction, and designs collaborative multi-agent reinforcement learning solution algorithms for these sub-problems to solve the multicast tree problem in SDWN multi-domain scenarios.
[0038] 2. Based on the characteristics of SDWN control logic centralization and programmability, the present invention utilizes a multi-controller communication mechanism and a multicast group management module to respectively realize the transmission and synchronization of network information between different control domains of SDWN, as well as the discovery and effective management of cross-domain multicast group members; this controller communication and management mechanism flexibly and conveniently obtains global network status information, realizes optimized distribution and resource utilization of multicast traffic, achieves better coordination and collaboration between controller domains, and improves the overall performance and efficiency of the network.
[0039] 3. The present invention comprehensively considers the state space composed of network link information and multicast tree state, so that the intelligent agent can better perceive the changes in network link state information and the changes in the multicast tree construction process; and according to the characteristics of the decomposed inter-domain multicast tree and intra-domain multicast tree problems, the corresponding action strategies are designed to improve the exploration efficiency of the intelligent agent, and different reward functions are designed for different action strategies taken by the intelligent agent to guide the intelligent agent to build efficient inter-domain and intra-domain multicast trees;
[0040] 4. The present invention realizes the construction and optimization of cross-domain multicast trees through collaborative learning and strategy coordination of multiple agents, adopts a completely decentralized solution paradigm to improve the stability of multi-agent collaboration, and designs a training method that combines offline and online training to reduce the frequency of interaction with the environment and dependence on the real-time environment, effectively improving the convergence speed of multi-agents. BRIEF DESCRIPTION OF THE DRAWINGS
[0041] Figure 1 It is a multicast tree in a multi-domain scenario.
[0042] Figure 2 It is an inter-domain multicast tree.
[0043] Figure 3 There is no destination node in the domain.
[0044] Figure 4 There are multiple destination nodes in the domain.
[0045] Figure 5 It is the multicast tree within the domain.
[0046] Figure 6 It is the SDWN multi-agent cross-domain multicast routing structure.
[0047] Figure 7 Build for inter-domain multicast tree.
[0048] Figure 8 Build a multicast tree for the domain.
[0049] Fig. 9 It is a cross-domain multicast tree.
[0050] Fig.10 This is the flow chart of the MA-CDMR algorithm.
[0051] Fig.11 is the state matrix of the agent.
[0052] Fig.12 Schematic diagram of four training strategies: (a) online on-policy, (b) online off-policy, (c) offline learning, and (d) offline-online learning.
[0053] Fig.13 It is a multi-domain network topology.
[0054] Fig.14 This is the simulated traffic distribution diagram.
[0055] Fig.15 Comparison plot of results for agent training policy settings.
[0056] Fig.16 Comparison results of different multicast algorithms: (a) average bottleneck bandwidth, (b) average delay, (c) average packet loss rate, (d) average multicast tree length, and (e) average link distance. DETAILED DESCRIPTION
[0057] In order to make the objectives, technical solutions and advantages of the present invention more clearly understood, the present invention is further described in detail below in conjunction with specific examples and with reference to the accompanying drawings.
[0058] Multicast is a communication method that transmits data to multiple destination nodes simultaneously in a computer network. The goal of the multicast problem is to build a multicast tree to minimize the transmission cost. A multicast tree is a tree structure with a multicast source as the root node and covering all destination nodes. By building a multicast tree with the minimum cost, data can be effectively transmitted to all destination nodes while reducing network bandwidth and transmission delay. The minimum cost multicast tree problem actually corresponds to the minimum Steiner tree problem in graph theory.
[0059] Given a wireless weighted graph G(V,E,w), where V is the set of nodes in G, E is the set of edges in G, and w is the weight of the edge. ij represents the edge from node i to node j, e ij ∈E, the edge weight is w(e ij ). Given a set of multicast nodes And M = {src}∪D, where src is the source node, D is the set of multicast destination nodes, and D = {d1, d2, …, d n}. Graph G′ is a subgraph of graph G that contains the multicast node set M, where graph G′ also contains some nodes that are not in M. These nodes are called Steiner nodes. The goal of the minimum Tanner tree problem is to find a spanning tree T = (V T ,E T ), and minimize the sum of the weights of its edges, that is:
[0060]
[0061] Among them, V T Represents all nodes in the tree T, E T Represents all the edges of the tree T.
[0062] The present invention considers the following cross-domain multicast problem of software-defined wireless network in multi-controller scenario and its solution method: Assume that the subsequent work is carried out in a multi-domain scenario, that is, the entire wireless network has been reasonably divided into m control domains N i ,i=1,…,m(the division of software defined network control domain belongs to the existing technology in the field of software defined network. The cross-domain multicast problem discussed in the following of this invention is assumed to be carried out after the entire network is determined to be divided, and the occurrence of dynamic migration of SDN switches in each control domain after division is not considered). It is also assumed that the inter-domain path cost between any two domains is greater than the path cost between any two nodes in the domain. To summarize the above description of the problem background, we have the following two assumptions:
[0063] Assumption 1: After the software-defined wireless network is reasonably divided, the SDN switch nodes and their connection relationships in each domain are fixed, and there is no need to consider the migration of switch nodes in and out;
[0064] Assumption 2: The inter-domain path cost of the link between any two domains is greater than the path cost between any two nodes within the domain.
[0065] The cross-domain multicast problem involved in the present invention is that the multicast source is represented by src and the destination node of the multicast is d i , i = 1, 2, ..., n, some of these n multicast destination nodes may be in the same control domain as the src, or none of them may be in the same control domain. The multicast routing path includes the intra-domain routing path part and the inter-domain routing path part. Figure 1 As shown in, src is the source node, d i is the destination node, BN i,a is the boundary node between domains, representing domain N i The a-th boundary node in the i →{d j} indicates that any domain N i The destination node d contained in j The set of mapping functions Domain(d i ):d i →N j For any destination node d i The corresponding domain N j , using the mapping function BND:N i →{BN ia} indicates that any domain N i The corresponding boundary node BN ia The three mapping functions dm(N i )、Domain(d i ) and BND(N i) is determined after the division of multiple control domains and can be considered known.
[0066] Since the multicast routing path consists of two parts: the intra-domain routing path and the inter-domain routing path, the problem of the optimal cross-domain multicast tree can be expressed by the following definitions:
[0067] Definition 1: From the source node src to the destination node d in the multicast tree i Inter-domain Path PN i : Indicates reaching the destination node d from the source node src i The domain sequence of each control domain passed in turn in Indicates the inter-domain path PN i The jth domain passed through, Indicates the inter-domain path PN i The set of all domains passed, ε i express The connectivity relationship between the domains in the , that is, the inter-domain path PN i The adjacent domain N i,j-1 and N i,j The edge between <N i,j-1 ,N i,j >. Figure 1 The three inter-domain paths in are: PN2 =<N1,N2> ,PN3=PN4=<N1,N3> ,PN5=PN6=<N1,N2,N4> .
[0068] Definition 2: Inter-domain multicast tree: source node src to all destination nodes d i Inter-domain path PN i The corresponding tree, i = 1, ..., n. This tree only contains links between domains and can be regarded as a multicast tree in a graph with domains as "nodes" and "edges" representing adjacent connectivity between domains.
[0069] There are n inter-domain paths from source node src to destination node d. i Inter-domain Path PN i ,i=1,…,n, represents any inter-domain multicast tree T int , that is, T int = {PN1,…,PN i ,…PN n}; The inter-domain multicast tree T can also be represented by the boundary nodes BN connecting the domains. int , whose edge set is P int ={(BN ia ,BN jb )|BN ia ∈BND(N i),BN ja ∈BND(N j )}, where (BN ia ,BN jb ) indicates the connection domain N i and domain N j The edge of BN ia Indicates domain N i The ath boundary node in BN ja Indicates domain N j The bth boundary node in T int =P int .like Figure 2 The inter-domain multicast tree shown in Figure 1 is a multicast tree with BN 11 , BN 13 , BN 21 , BN 24 , BN 34 , BN 43 is the selected boundary node, N1 is the source domain, N2, N3, N4 are the destination domains, and the inter-domain multicast tree can be expressed as: int ={(BN 11 ,BN 21 ),(BN 13 ,BN 21 ),(BN 24 ,BN 41 ),(BN 34 ,BN 43 )}. In general,
[0070] T int = {PN1,…,PN i ,…PN n}or
[0071] T int =P int ={(BN ia ,BN jb )|
[0072] BN ia ∈BND(N i ),BN jb ∈BND(N j )}
[0073] The inter-domain multicast tree defined in Definition 2 is actually the inter-domain routing path part of the cross-domain multicast tree. Corresponding to this is the intra-domain routing path part of the cross-domain multicast tree. Regarding the intra-domain routing path part, we have the following definitions of the intra-domain multicast tree and the intra-domain multicast forest.
[0074] Definition 3: Domain N i Intra-domain multicast tree and intra-domain multicast forest Ti : From the source node src to each destination node d i , i = 1,…, n, the intra-domain routing path part passing through each domain is represented by T i Indicates domain N i Obviously, as long as there is a path passing through one of the domains N i , then there must be T i It could be a single tree or a forest of multiple trees, which can be defined in the following ways:
[0075] (1) Domain N i There is no destination node in the multicast tree. In this case, the multicast path only passes through the domain, from the border node BN ia Enter the domain N i , from other border nodes BN ib Leave the domain N i , if only one enters domain N i Boundary nodes BN ia , then domain N i The path part inside has only one connected component, and T i It can be regarded as a BN ia As the root, leaving the boundary node BN ib The tree with leaf nodes is called the intra-domain multicast tree T i ,like Figure 3 (a) As shown; if there are multiple entry domains N i Boundary nodes BN ia , then domain N i The path part within may have a connected component, such as Figure 3 As shown in (b), there may also be multiple connected components, such as Figure 3 As shown in (c), at this time T i It can be seen as a multiple BN ia The forest composed of trees with root T is called intra-domain multicast forest T i .
[0076] (2) Domain N i There is a destination node d to be reached by the multicast tree i ∈dn(N i ), if there is only one such destination node, that is, |dm(N i )|=1, the multicast path will only pass through a single border node BN ia Enter this domain i , at this time T i It is a BN ia The root tree (if you do not leave the domain Ni The boundary nodes of the tree T i is a chain, such as Figure 4 As shown in (a), if there are still N people leaving the domain i Reach the border nodes BN of other domains ib , then the tree T i It is a star with a d i and BN ib The ordinary tree with leaf nodes is called the intra-domain multicast tree T i ,like Figure 4 (b) If there is more than one destination node, that is, |dm(N i )|>1, the multicast path may pass through more than one border node BN ia Enter this domain i Reach these destination nodes respectively, then domain N i The path part within may have a connected component, such as Figure 4 As shown in (c), it is also possible that there are connected components, such as Figure 4 As shown in (d), at this time T i It can be seen as a multiple BN ia The forest composed of trees with root T is called intra-domain multicast forest T i .
[0077] For the domain N defined in Definition 3 i Intra-domain multicast tree and intra-domain multicast forest T i , we simply use edge sets to represent it:
[0078] T i ={(v1,v2),(v2,v3),…(v i,l ,v i,l-1 )}
[0079] Among them, (v i,l ,v i,l-1 ) represents domain N i Node v in the connection domain i,l-1 and the node v in the domain i,l edge.
[0080] For domain N that does not have an inter-domain multicast tree passing through i , which can be considered as an intra-domain multicast tree or intra-domain multicast forest Then all domains N i The set T of intra-domain multicast trees (i=1,…,m) intra It is expressed as:
[0081]
[0082] Here, m represents the number of domains.
[0083] Figure 5 An example of an intra-domain multicast tree is given, where four intra-domain multicast trees are represented by T1, T2, T3, and T4. Figure 5 As shown, src is the source node, d1, d2, d3, d4, d5 and d6 are the destination nodes, and v nor1 and v nor2 For ordinary nodes, T1={(src,BN 11 ),(src,v nor1 ),(v nor ,d1),(d1,BN 13 )}, T2={(BN 21 ,BN 24 ),(BN 21 ,d2)},T3={(BN 31 ,BN 32 ),(BN 32 ,d3),(BN 31 ,d4)},T4={(BN 41 ,d5),(BN 41 ,BN 42 ),(BN 42 ,d6)}.
[0084] Definition 4: Cross-domain multicast tree T cd :The inter-domain multicast tree T int and all intra-domain multicast trees T intra All path parts of the combined form the cross-domain multicast tree T cd :
[0085] T cd =T int ∪T intra
[0086] Definition 5: Optimal cross-domain multicast tree: Each multicast tree can be defined as a set of nodes from the source node src to n destination nodes d i The path p of (i=1,…,n) i The set of (i=1,…,n) (similar to the definition of multicast tree in our previous work MADRL-MR) can be represented by T={p1,…,p k ,…,p n}, where p k From the source node src to the destination node d in the multicast tree T k , n is the number of destination nodes. If each path p k ∈T has a minimum cost, then the multicast tree T is an end-to-end minimum cost tree. Define each path p kThe cost is c k , whose optimization goal is to maximize p k The bottleneck bandwidth bw k , minimize p k Delay k , Packet loss rate loss k 、Error packet rate err k The distance between the wireless access point and AP is dist k to calculate, and all indicator parameters are normalized to [0,1] through Max-Min. Each path p k The construction cost can be expressed as c k The definition is as follows:
[0087] cost(p k )=c k =β1(1-bw k )+β2delay k +β3loss k +β4err k +β5dist k
[0088] in, p k The source nodes src to d in the multicast tree T k The path, bw ij Indicates e ij The remaining bandwidth, delay ij Indicates e ij Delay, loss ij Indicates e ij Packet loss rate, err ij Indicates e ij The packet error rate, dist ij Indicates e ij distance.
[0089] The optimal cross-domain multicast tree is to find such a Steiner tree The path cost c from the source node of this Steiner tree to all the destination nodes k The sum is the smallest, that is:
[0090]
[0091] The multicast tree T with the lowest cost is the optimal cross-domain multicast tree, and the following formula holds true:
[0092] cost(T cd )=cost(T int ∪T intra )
[0093] Through Definition 1 to Definition 5, the present invention provides the definition of the optimal cross-domain multicast tree. Based on these definitions, we know that the cross-domain multicast tree can be decomposed into an inter-domain multicast tree and an intra-domain multicast tree for solution. Based on the assumption that "the path cost between any two domains is greater than the path cost between any two nodes in the domain", we can obtain the following properties of the optimal cross-domain multicast tree: There is no inter-domain path loop in the optimal multicast tree routing path.
[0094] Since the communication of inter-domain paths involves the coordinated operation of different controllers, increasing the number of inter-domain communications will increase the cost of collaborative interaction. The inter-domain path overhead is much greater than the intra-domain path overhead, and the optional paths within each domain have sufficient redundancy to ensure the connectivity between network nodes within the domain. There is only one inter-domain link between any two domains crossed by the multicast path, which reduces multicast overhead such as network traffic.
[0095] Assume that the inter-domain multicast tree and each intra-domain multicast tree are minimum cost trees. The definition of the minimum cost tree is the same as that of the optimal inter-domain multicast tree in Definition 5. They are both represented as a set of nodes from the source node src to n destination nodes d i The path p of (i=1,…,n) i The set of (i=1,…,n), T={p1,…,p k ,…,p n}. Find each path p in T k The minimum cost c k , thus obtaining the minimum cost of T. Assume that the intra-domain multicast tree T in each domain i The construction cost is denoted as C i , domain N i The destination node in the domain (if the domain has a border node that leaves the domain, then the border node also becomes the multicast tree T in the domain i The number of destination nodes is n i If it is expressed, then:
[0096]
[0097] The cost of the inter-domain multicast tree is C int The calculation method is the same as that of the intra-domain multicast tree. Then the total cost of the inter-domain multicast tree (T cd ) consists of the cost of the intra-domain and inter-domain multicast trees, and minimizing its total cost can be expressed as follows:
[0098]
[0099] in, represents the sum of the construction costs of all multicast trees in the domain, m represents the number of domains, C int Indicates the construction cost of the inter-domain multicast tree.
[0100] The SDWN multi-agent cross-domain multicast routing strategy is to decompose the construction of the cross-domain multicast tree into the construction of the inter-domain multicast tree and multiple intra-domain multicast trees. The cross-domain multicast routing is realized through multi-agent collaboration. Its architecture is as follows: Figure 6 shown.
[0101] Each domain in the control plane has a local controller. The local controller periodically obtains the network status information of the corresponding domain, and synchronizes the information collected by the local controller to the root controller through the controller communication mechanism. The root controller manages the global network resources. ② The global network link information (NLI, Network Link Information) collected by the controller plane is processed and stored in the knowledge plane. ③ The multiple agents in the knowledge plane are trained by perceiving and learning the NLI at different times. ④ The inter-domain agent decides the inter-domain multicast tree through global network information, that is, the selection of paths and boundary nodes between multiple domains. Then notify the agent in each controller domain to complete the construction of the intra-domain multicast tree. Finally, multiple agents collaborate to complete the construction of the optimal cross-domain multicast tree. ⑤ The cross-domain multicast tree decided is synchronized to the root controller and the local controller. ⑥ Before the next network traffic arrives, the control plane sends and installs the flow table to the wireless access node of the data plane. Finally, the data plane completes the forwarding of traffic.
[0102] (1) Data plane
[0103] The data plane is responsible for processing and forwarding data packets in the network, which contains multiple network subdomains. Each domain consists of wireless access nodes (AP, Access Point) and stations (STA, Station). The APs in each domain form a multi-hop wireless network through a wireless mesh method. APs are divided into intra-domain nodes (IN, Intra-domain node) and boundary nodes (BN, Boundary node). Each AP in the domain is connected to a STA. The AP receives wireless data packets from terminal devices, selects the best routing path according to the instructions and policies issued by the control plane, and forwards the data packets from the source device to the target device. Each domain periodically interacts with the local controller to pass the wireless network link information of the current domain to the control plane.
[0104] (2) Control Plane
[0105] The control plane is responsible for collecting and analyzing network status information, making decisions, and issuing commands and policies to the APs in the data plane, which includes a root controller and multiple local controllers. The local controller communicates with the APs in the data plane through the southbound interface, and synchronizes the network status information of the corresponding domain to the root controller through the controller communication mechanism (CCM). The root controller builds a global view of the network and manages and schedules global network resources. The control plane interacts with the knowledge plane through the northbound interface to facilitate the issuance and deployment of knowledge plane policies. These include network topology discovery, link information detection, controller communication mechanism, multicast group management, and flow table installation.
[0106] Network topology discovery and link information detection are performed by sending data packets from the controller to the data plane to obtain relevant information. Network topology discovery is to periodically send LLDP (Link Layer Discovery Protocol, LLDP) request packets. When the AP replies, it will encapsulate the device's port connection, ID and other information into the Reply packet, and the controller will eventually parse and build the network topology. Link information detection is to obtain port information by sending PortStatsRequest request packets. The controller parses the reply message to obtain the number of packets sent tx p , Number of received packets rx p 、Number of bytes sent tx b 、Number of received bytes rx b 、Number of sent error packets tx err And the number of received error packets rx err , and the duration t of the port sending bytes dur . The remaining bandwidth bw is calculated using the above parameters ij , Packet loss rate loss ij and error rate err ij The network link information is calculated as follows:
[0107]
[0108] Among them, bw max Indicates the maximum bandwidth, tx b* 、rx b* and t dur* Represents the number of bytes sent, the number of bytes received and the duration of the node*. p* ,tx err* and rx p* 、rx err* They represent the number of data packets and error packets sent by node* and the number of data packets and error packets received by node*.
[0109] In SDN, since the communication between two switches needs to be forwarded by the controller, the network link delay between the two switches needs to be approximately calculated. The controller parses the timestamp information of the data packets passing through the link to obtain the round-trip delays RTT1 and RTT2 from the controller to the two switches, and calculates the forward transmission delay T of the two switches based on the LLDP message. fwd and reply transmission delay T re , and thus the link delay is approximately calculated ij , the calculation formula is as follows:
[0110]
[0111] In addition, the distance between the two APs is calculated using the deployment coordinates of the wireless AP. ij .
[0112] The network link information (NLI) calculated above is the remaining bandwidth bw ij , Packet loss rate loss ij 、Error packet rate err ij , link delay ij and the distance between two APs dist ij After Max-Min normalization, it is stored in the NLIStorage of the knowledge plane.
[0113] The controller communication mechanism (CCM) ensures stable and fast communication between the local controller and the root controller. Traditional SDN multi-domain communication is through the MBGP protocol, but the protocol is cumbersome to configure in the SDN environment, and the message update may cause inconsistencies between multiple controllers. The present invention designs a communication mechanism between a local controller and a root controller based on the RESTful API. CCM has good flexibility, so that the root controller and the local controller can interact in a unified manner and adapt to different network environments and requirements. And the RESTful API is standardized, with clearly defined specifications and constraints, reducing compatibility issues. In addition, it is scalable and can easily expand and enhance communication capabilities. CCM also achieves loose coupling, and each controller can be independently developed and deployed, which improves maintainability and scalability. At the same time, it is portable and can run in different environments and platforms.
[0114] The multicast group management (MGM) function is mainly to analyze the domains where the multicast source and multicast destination nodes are located, and to manage the joining and leaving of multicast nodes. The multicast group management function determines the domains where the multicast source and multicast destination nodes are located by monitoring network traffic and analyzing data packets, and coordinates the communication between domains to ensure that the multicast stream can be correctly transmitted across domains. At the same time, it can communicate with the controllers in each domain, provide them with relevant information about the multicast stream, and coordinate the transmission path of the multicast stream between domains. In addition, it can determine joining and leaving based on the requests of multicast users, and update the multicast tree.
[0115] The flow table installation function receives instructions from the knowledge plane through the SDN northbound interface and sends the multicast flow table. When receiving node join and leave requests, it modifies or deletes the existing group table entries to implement the joining and deletion of group members and ensure accurate data forwarding.
[0116] (3) Knowledge Plane
[0117] The knowledge plane is an important component newly added to the SDN architecture. The multi-agent deep reinforcement learning intelligent cross-domain multicast routing algorithm of the present invention runs on this plane. It includes NLI Storage, which stores the network link information of the data plane collected by the control plane; and it is necessary to process the NLI into a traffic matrix TM and provide it to the agent training and learning. After training, multiple agents collaborate to complete the construction of the cross-domain multicast tree, and then send instructions to the control plane, which sends flow table items to the data plane.
[0118] Based on the description and modeling of the cross-domain multicast problem above, we solve the multicast problem by constructing a minimum cost Steiner tree, and decompose the cross-domain multicast tree into multiple multicast tree construction problems, such as inter-domain multicast trees and multiple intra-domain multicast trees. In order to distinguish it from the agent that constructs the intra-domain multicast tree, we call the agent that constructs the inter-domain multicast tree an inter-domain agent. The state space and reward function of the inter-domain agent are the same as those of the intra-domain agent, but the action space is different. The action space of the inter-domain agent is the set of connecting edges between different domains, while the action space of the intra-domain agent is the set of nodes within the domain. This is because when constructing the inter-domain multicast tree, we abstract the domains as nodes, but there may be multiple connecting edges between adjacent domains, so the node set cannot be used as the action space.
[0119] The following two simple examples introduce the process of multiple intelligent agents collaborating to build a cross-domain multicast tree.
[0120] Example 1: Figure 7 It is a multi-domain network topology consisting of 4 domains. Adjacent domains are connected through border nodes BN. The overall representation is the construction process of the inter-domain multicast tree. Figure 7In (a), the source node and the destination node are marked, and they are in different domains. Figure 7 In (b), we abstract domains into nodes, N1 is the source domain, N2, N3, N4 are the target domains, and only retain the edges connecting adjacent domains and boundary nodes, and set the weight of each edge. We regard the edges between adjacent domains as the actions of the inter-domain agent, with the minimum cost as the goal. The inter-domain agent will choose (BN 11 ,BN 21 )、(BN 13 ,BN 21 )、(BN 24 ,BN 41 ) and (BN 34 ,BN 43 ) and other four edges to connect the source domain and the target domain. The cost of the inter-domain multicast tree constructed in this way is the minimum C int =6, and the boundary nodes of each domain to other domains are determined. The minimum cost inter-domain multicast tree T is finally constructed int ,like Figure 7 (c) as shown.
[0121] Example 2: Figure 8 The figure shows the process of building a multicast tree within a domain. First, we build an inter-domain multicast tree and select the cross-domain boundary nodes. Figure 8 (a) shows that according to Figure 7 In the inter-domain multicast tree constructed in , we mark the selected boundary nodes with depth. These boundary nodes can be regarded as source nodes and destination nodes in each domain. Figure 8 As shown in (b), we use BN 11 ,BN 13 Equivalent to the destination node of domain N1; BN 21 Equivalent to the source node in N2, BN 24 Equivalent to the destination node; BN 31 Equivalent to the source node of domain N3; BN 11 Equivalent to the source node of domain N4. Then the corresponding domain agent constructs the domain multicast trees T1, T2, T3 and T4 from the source node to the destination node. Finally, the domain multicast tree constructed by each domain is as follows: Figure 8 (c) as shown.
[0122] The above two simple examples show the construction process of inter-domain multicast trees and multiple intra-domain multicast trees. First, the inter-domain multicast tree is constructed by the inter-domain agent. After the boundary nodes of each domain are determined, the corresponding domains construct the intra-domain multicast tree. Finally, the inter-domain multicast tree is combined with multiple intra-domain multicast trees to obtain a cross-domain multicast tree, such as Fig. 9 shown.
[0123] According to the design of the multicast tree solution idea in the knowledge plane, the flowchart of the inter-domain multicast tree solution algorithm MA-CDMR provided by the present invention is as follows: Fig.10 As shown. First, the global network topology and link information are obtained through the multi-controller SDWN architecture, and then the network information is stored in the NLI Storage of the knowledge plane and provided to the intelligent agent for training and learning. Secondly, the global network information is first perceived by the inter-domain intelligent agent, and the optimal inter-domain multicast tree is determined to determine the boundary nodes leading to the neighboring domains from each domain. Then the information is synchronized to the intra-domain intelligent agent of each domain, and the intra-domain intelligent agent of the corresponding domain completes the construction of the optimal intra-domain multicast tree. Finally, the optimal inter-domain multicast tree and multiple intra-domain multicast trees are combined to generate the optimal cross-domain multicast tree, and the forwarding path of the optimal cross-domain routing of the global network is obtained.
[0124] The following is a detailed introduction to the MA-CDMR algorithm, which includes the design of the agent’s state space, action space, reward function, and multi-agent training strategy.
[0125] (1) Agent Design
[0126] In the MA-CDMR algorithm, we designed two types of agents, one is the inter-domain agent that builds the inter-domain multicast tree, and the other is the intra-domain agent that builds the intra-domain multicast tree. There is only one inter-domain agent, while there is one intra-domain agent in each domain. The state space and reward function of the two agents are the same, but the action space is different. The following is a detailed introduction.
[0127] 1) State Space
[0128] The state space of the agent determines its ability to perceive and understand the environment, and should contain enough environmental information for the agent to make meaningful decisions. Therefore, we use the traffic matrix of the network link's remaining bandwidth bw, delay delay, packet loss rate loss, and error rate err, as well as the multicast tree construction state matrix M T A multi-channel matrix X = [bw, delay, loss, err, dist] is formed as the state s of the agent. Its multi-channel state matrix is as follows Fig.11 shown.
[0129] The set of all possible changes of the multi-channel matrix X is the state space of the agent For example, if an edge is selected to join the multicast tree, the state of the agent changes from s t becomes s t+1 The state space designed in this way can provide important information about network links and nodes, as well as the specific situation of multicast tree construction, so that the intelligent agent can make corresponding decisions based on the current network conditions and select the best path to build the multicast tree.
[0130] 2) Action Space
[0131] The action space determines the range and granularity of operations that the agent can perform. The design of the action space should include as many possible actions as possible that the agent needs to take and adapt to the characteristics of the environment and the requirements of the task. Therefore, we design the action space for inter-domain agents and intra-domain agents separately.
[0132] Action Space of Inter-Domain Agents When we construct the inter-domain multicast tree, we abstract each domain as a node, and this node may have multiple edges connecting to adjacent nodes. Therefore, we use the set of edges connecting all domains as the action space of the inter-domain agent. Among them (BN ia ,BN jb ) indicates the connection domain N i and domain N j The edge of BN ia Indicates domain N i The ath boundary node in , k is the total number of edges between different domains. The edges connecting the current domain and the domain are valid actions, and the rest are invalid actions. This design enables the agent to select the most suitable boundary nodes and paths across domains, thereby constructing the best inter-domain multicast tree.
[0133] Action space of agents in the domain Since the connections between nodes in the domain are complex and there are many connected edges, in order to reduce the complexity of the action space, we use the set of nodes in the domain as the action space of the agent in the domain. Node v and action a correspond one to one. The next hop node of the current node is the valid action, and the rest are invalid actions. This design can reduce the invalid exploration of the intelligent agent and make the optimal path decision more quickly.
[0134] 3) Reward Function
[0135] The reward function determines how the agent evaluates the quality of actions during the learning process. It can provide clear feedback to the agent's behavior and guide the agent towards the desired goal. Our optimization goal is to maximize the remaining bandwidth of the multicast tree and minimize the latency, packet loss rate, packet error rate and the distance between wireless access points AP. To this end, we design the reward function based on network link information parameters. Since we design the action space based on a collection of links or nodes, valid actions and invalid actions will be generated for the state s at different times. When the agent performs a valid action a, the state s t Transformed into t+1, there may be two different positive feedbacks and one negative feedback. Performing an invalid action will produce a negative feedback. The following are the reward functions designed for different situations.
[0136] If the current action a is executed t After that, the added next-hop node or link just adds a normal node to the multicast tree. For this positive feedback, we designed a single-step reward R part , to join the link e of the multicast tree ij The remaining bandwidth bw ij , delay ij , Packet loss rate loss ij 、Error packet rate err ij The distance between the AP and ij To calculate the reward, as shown below:
[0137] R part =β1bw ij +β2(1-delay ij )+β3(1-loss ij )+β4(1-err ij )+β5(1-dist ij )
[0138] If the next hop node or link added is to add a destination node to the multicast tree k For this positive feedback, we designed a subtask reward R end . From the source node src to the destination node d k The remaining bandwidth bw of the entire link k , delay k , Packet loss rate loss k 、Error packet rate err k and the average distance of the link dist k To calculate the reward, as shown below:
[0139] R end =β1bw k +β2(1-delay k )+β3(1-loss k )+β4(1-err k )+β5(1-dist k )
[0140] If the added next hop node or link causes the multicast tree to form a loop, a penalty value R is set for this negative feedback. loop =C1.
[0141] If a tis an invalid action, which is neither the next hop node of the current node nor the edge connected to the current domain neighborhood. For this kind of negative feedback, we also give a penalty value R hell =C2.
[0142] 4) Agent strategy update
[0143] All agents in this algorithm use the Actor-Critic framework. The framework is divided into two parts: Actor (strategy network π θ (a|s), parameter is θ) and Critic (value network V ω , with parameter ω). The policy network Actor interacts with the environment and uses policy gradient to learn a better policy under the guidance of the Critic value function. It uses policy gradient update, as shown in the following formula:
[0144]
[0145] Where T is the maximum number of steps to interact with the environment, π θ (a|s) is the policy function, a t is the current action, s t is the current state. t is the temporal difference residual, which is used to guide policy gradient learning and is calculated as follows:
[0146] ψ t =r t +γV ω (s t+1 )-V ω (s t )
[0147] Among them, r t V is the reward value obtained in this round, γ is the decay factor, γ∈[0,1]. ω (s t ) is in state s t Next, perform action a t The expected value generated.
[0148] Critic Value Network V ω Using the learning method of time series difference residual, the loss function of the following value function is defined for a single data, as shown in the following formula:
[0149]
[0150] In the above formula, r t +γV ω (s t+1) as a temporal difference target, no gradient is generated to update the value function. Therefore, the gradient of the value function is as follows:
[0151]
[0152] (2) Design of multi-agent training strategies
[0153] The training of multiple agents is more complicated than that of a single agent, because each agent interacts with the environment while also interacting directly or indirectly with other agents, and each agent continuously learns and updates its own strategy. Therefore, for each agent, the environment is non-steady-state. On the other hand, for a multi-domain environment, the training objectives of the agents in each domain are different, and different agents need to maximize their own rewards, which may require large-scale distributed training to improve efficiency. In response to the above problems, the present invention adopts a fully decentralized training method, which is also called independent learning (IL).
[0154] In the MA-CDMR algorithm, for each intra-domain agent, the global network environment is non-steady-state, but the sub-domain network environment it is responsible for is steady-state. For inter-domain agents, their policy execution takes precedence over intra-domain agents, and they face the global network, so their environment is also steady-state. Therefore, by adopting the IL training method, each agent can learn independently in its own environment without considering the changes of other agents. And it has good scalability as the network expands and the number of agents increases.
[0155] In order to improve the convergence speed of multi-agents and reduce the training cost in a large-scale network environment, the present invention designs a training method that combines offline and online learning based on online reinforcement learning and the advantages of offline reinforcement learning. To show the difference between the other three training strategies and the training strategy designed by the present invention, we show them in the form of pictures, as shown in the figure below: Fig.12 shown. Fig.12 (a) and (b) are the on-policy and off-policy of online reinforcement learning, respectively. Both require real-time interaction with the environment. The on-policy interaction with the environment and the strategy of using data update are the same strategy π k The off-policy interacts with the environment to generate data strategy π β and strategies for using data k For different strategies. Fig.12(c) is the training strategy for offline reinforcement learning, which collects data sets before training the model, and then uses these data sets to train the agent offline. It is the same as off-policy, but does not interact with the environment to update the strategy. Fig.12 (d) The training strategy designed by the present invention combines the advantages of online and offline. The same strategy interacts with the environment and uses data. First, collect a certain amount of data, then use the offline data to train an initial strategy, and then use the on-policy in online reinforcement learning to interact with the environment and improve the strategy in the actual environment. The offline and online hybrid training strategy can reduce the demand for online learning, improve data utilization efficiency, and better cope with dynamic changes in the actual environment.
[0156] (3) MA-CDMR algorithm design
[0157] The MA-CDMR algorithm divides cross-domain multicast routing into two stages: inter-domain multicast tree construction and intra-domain multicast tree construction. First, the inter-domain intelligent agent perceives the global network environment and decides on the inter-domain multicast tree, selects the best inter-domain forwarding path and the boundary node of each domain forwarding. Then the information is synchronized to each intra-domain intelligent agent, and the intra-domain intelligent agent starts to build the intra-domain multicast tree. Finally, the inter-domain multicast tree and multiple intra-domain multicast trees are combined to generate the best cross-domain multicast tree. For the joining and leaving of multicast users, the multicast path is constructed and the multicast tree is pruned respectively.
[0158] Accordingly, the intelligent cross-domain multicast routing method based on multi-agent deep reinforcement learning proposed in the present invention includes the following steps:
[0159] Step 1: deploy an inter-domain agent on the root controller of the multi-domain software-defined wireless network and deploy an intra-domain agent on the local controller of each domain;
[0160] Step 2: Initialize the Actor network parameters θ between domains and within domains int ,θ intra 、Critic network parametersω int ,ω intra and data buffer area B int ,B intra ;
[0161] Step 3: The root controller and each local controller communicate and collect information of the multi-domain software-defined wireless network through the controller communication mechanism, and analyze the information collected by the corresponding controller through the multicast group management module to determine the domains where the source node and each destination node are located, as well as the forwarding boundary nodes of each domain;
[0162] Step 4: First, the inter-domain agent interacts with the environment to generate an inter-domain network link information matrix and the inter-domain multicast tree state matrix And stack the inter-domain network link information matrix and the inter-domain multicast tree state matrix Get the current state s of the inter-domain agent t ; Then, the inter-domain agent changes from the current state s t Sample the current action a from the output action set t , execute the current action a t Get the current reward value r t and the next state s t+1 , where the action set of the inter-domain agent is the set of edges between all domains, and each action is selected as one of the edges; finally, the reinforcement learning data (s t ,a t ,r t ,s t+1 ) is stored in the data cache area B of the inter-domain agent int ;
[0163] Step 5: When the inter-domain agent repeats the action of step 4, the inter-domain agent is trained by offline and online hybrid training, that is:
[0164] Every set period of time, i.e., the inter-domain period, the data cache area B is used first during this period of time. int The inter-domain agent is trained offline without interacting with the environment using the offline reinforcement learning data in , so as to update the Actor network parameters θ of the inter-domain agent. int and Critic network parameter ω int , and then enter the next inter-domain agent action process;
[0165] When the set inter-domain update frequency is reached, the online reinforcement learning data obtained from this action is first used to conduct online training for the inter-domain agent to interact with the environment, so as to update the Actor network parameters θ of the inter-domain agent. int and Critic network parameter ω int , and then enter the next inter-domain agent action process;
[0166] Step 6: In the inter-domain multicast tree T int When the construction is completed, the inter-domain agent will construct the inter-domain multicast tree T int Synchronize to each agent in the domain;
[0167] Step 7: First, each agent in the domain interacts with the environment to generate the network link information matrix in its own domain. and the intra-domain multicast tree state matrix And stack its own intra-domain network link information matrix and the intra-domain multicast tree state matrix Get the current state s of the agent in the domain t ; Then, each agent ring in the domain starts from the current state s t Sample the current action a from the output action set t , execute the current action a t Get the current reward value r t and the next state s t+1 , where the action set of the agent in the domain is the set of nodes in the domain, and each action is selected as the next hop node; finally, the reinforcement learning data (s t ,a t ,r t ,s t+1 ) is stored in the data buffer area B of the agent in the domain intra ;
[0168] Step 8: When each agent in the domain repeats the action of step 7, each agent in the domain is trained by offline and online hybrid training, that is:
[0169] Every set period of time, i.e., the domain cycle, the data buffer area B is used first during this period of time. intra The offline reinforcement learning data in the domain is used to train the in-domain agent offline without interacting with the environment to update the Actor network parameters θ of the in-domain agent. intra and Critic network parameter ω intra , and then enter the next action process of the agent in the domain;
[0170] When the set domain update frequency is reached, the online reinforcement learning data obtained from this action is first used to conduct online training on the domain agent to interact with the environment, so as to update the Actor network parameters θ of each domain agent. intra and Critic network parameter ω intra , and then enter the next action process of each agent in the domain;
[0171] Step 9: Multicast tree T in each domain intra When all are constructed, the combined inter-domain multicast tree T int and all intra-domain multicast trees T intra Combined into a cross-domain multicast tree;
[0172] Step 10: Determine whether the cross-domain multicast tree converges or reaches a preset number of iterations: If so, the multi-domain software-defined wireless network performs intelligent cross-domain multicast routing based on the current cross-domain multicast tree; otherwise, return to step 4 and perform the next round of iteration.
[0173] The server software system used in the experiment is Ubuntu 18.04.6, and the hardware configuration is a 64-core processor and a GeForce RTX 3090 graphics card. Mininet-WIFI 2.3.1b is installed on the server as the simulation platform of SDWN. The controller uses Ryu 4.3.4. The Iperf tool is used to send streams to simulate real network traffic.
[0174] The multi-domain network topology used in the experiment is as follows Fig.13 The link parameters of the global network are uniformly distributed and randomly generated within a certain range. The random ranges of link bandwidth and latency are 5-40 Mbps and 1-10 ms respectively, and the distance between wireless APs is set within 30-120 m. The distribution of simulated traffic is shown in Fig.14 shown.
[0175] In the process of deep reinforcement learning model training, the setting of hyperparameters has a significant impact on the experimental results. In the experiment, we selected a group of cross-domain multicast nodes with complex optional paths between and within domains in a multi-domain environment as representatives, M = {src} ∪ D = {3} ∪ {6, 15, 18, 19, 26, 28}.
[0176] ①Set the positive feedback reward value R part , R end and negative feedback penalty value R hell , R loop According to our previous work, MADRL-MR contains a large number of R part , R end The optimal weights of the remaining bandwidth, delay, packet loss rate, packet error rate and distance between APs are [0.7, 0.3, 0.1, 0.1, 0.1]. hell , R loop The experiment showed that R hell = -0.7 and R loop = -0.5 works best. On this basis, modify R part and R end The agent can reach the destination node more quickly by setting the reward ratio to 1:0.1, 1:0.3, 1:0.5 and 1:0.7 for experiments. We also tried setting the reward ratio to 1:1, but the agent failed to converge. The reward for a single step cannot be set too large, as it will induce the agent to take more steps and thus fail to reach the destination node. The experimental results show that 1:0.1 is the best.
[0177] ② The learning rate α1 of the actor and the learning rate α2 of the critic both have an important impact on the training process and performance. α1 determines the speed and amplitude of the policy update. α2 determines the update speed and amplitude of the value function. A higher learning rate can speed up the update and convergence of the policy, but may lead to instability and overfitting. A lower learning rate improves stability, but the learning speed is slow, and more training iterations may be required to achieve convergence. To this end, experiments are conducted on the settings of α1 and α2. First, experiments are conducted on α1. The experimental results show that setting α1 to 1e-4 has the best convergence effect. Then, on the basis of α1 being 1e-4, the value of α2 is adjusted. The experimental results show that the best convergence effect is achieved when α2 is set to 3e-4.
[0178] ③ Set the decay factor γ of the reward value. The decay factor determines the relative importance of the current moment reward and the future reward. A higher decay factor will pay more attention to future rewards, allowing the agent to consider decisions in the long term, but may lead to slower convergence. A lower decay factor will pay more attention to immediate rewards, and the training process may be faster, but may lead to insufficient long-term planning capabilities. For this reason, this paper experiments on the setting of γ. The experimental results show that when γ is set to 0.9, its convergence speed is slower, but the reward value converges most stably.
[0179] ④Set the update frequency of online learning n update .n update A higher value can speed up the convergence speed and improve the stability of training, reduce the number of iterations required for training, and reduce jitter. However, too high a value may increase the computational overhead and lead to the risk of overfitting. update When setting n update = 1000, the agent cannot converge, so we set it to 1, 5, 10, and 100. It can be observed that when n update =100 is the best, but n update = 10, the convergence speed is only slower than when it is set to 100. In order to prevent overfitting, we finally set n update Set to 10.
[0180] ⑤ Set the batch size k in offline learning. The dataset in offline reinforcement learning is usually divided into batches for training, and the batch size k of each batch is an important hyperparameter. A larger k can speed up the training and reduce the variance of parameter updates, but it will increase memory consumption. However, a larger k may also cause the model to overfit the training data and the generalization ability of the model will deteriorate. For this reason, this paper experiments on the setting of k. The experimental results show that the reward values converged when k is set to 16, 32, 64, and 128 are close, but when k = 32, the convergence speed is faster and the convergence effect is the best. Therefore, we set k to 32. .
[0181] We use the Actor-Critic network as the core framework of the agent. Each agent was originally trained using the on-policy of online learning. In order to verify the effect of the offline-online learning combined training method we designed, we will use the agents of online learning on-policy and offline-online learning combined training methods for comparative experiments. The results are as follows: Fig.15 As shown. Fig.15 The results show that the on-policy training method of online reinforcement learning tends to converge, but the convergence speed is slow. This is because in this training method, the agent needs to interact with the environment in real time and needs to continuously update the agent's strategy, which results in a long training time before convergence. However, the learning method we designed, which combines offline and online learning, converges faster and reaches the convergence value in a shorter training time. This is because the interaction between the agent and the environment is reduced. The pre-collected data set is used for offline training first, and then the current agent's strategy is adjusted by interacting with the environment, which reduces the training cost of the agent's interaction with the environment, thereby accelerating the convergence speed of the agent.
[0182] In order to evaluate the performance of the algorithm of the present invention, MA-CDMR is compared with four SDN multicast methods, including the classic multicast tree optimization methods KMB and SCTF, and the recently proposed reinforcement learning-based multicast routing methods DRL-M4MR and MADRL-MR. We have implemented the various SDN multicast algorithms mentioned above in a multi-domain environment. Specifically, given the source node and the destination node, a cross-domain multicast tree is generated at each moment according to various algorithms. The average value of the bottleneck bandwidth of each corresponding path, the average delay and average packet loss rate of each link, as well as the number of links and the average distance of the links are calculated respectively. Then they are averaged at each moment to calculate performance indicators such as average bottleneck bandwidth, average delay, average packet loss rate, average cross-domain multicast tree length and average link distance. The results are shown in the figure. Fig.16 shown.
[0183] Fig.16 (a) is a comparison of the average bottleneck bandwidth of the cross-domain multicast tree constructed by MA-CDMR and DRL-M4MR, MADRL-MR, KMB and SCTF. The results show that MA-CDMR is significantly better than SCTF in the average bottleneck bandwidth of the cross-domain multicast tree, with an average performance improvement of 46.01%. It is slightly better than DRL-M4MR, MADRL-MR and KMB, with improvements of 9.61%, 10.11% and 7.09% respectively.
[0184] Fig.16(b) is a comparison of the average delay of the cross-domain multicast tree constructed by MA-CDMR and DRL-M4MR, MADRL-MR, KMB and SCTF. The results show that the average delay of MA-CDMR is smaller than that of KMB and SCTF, and the average performance is improved by 26.39% and 78.74% respectively. Compared with DRL-M4MR, the performance is improved by 12.47%, which is very close to the value of MADRL-MR, which is only improved by 7.17%. Therefore, it shows that MA-CDMR has a better effect on average delay.
[0185] Fig.16 (c) is a comparison of the average packet loss rate of the cross-domain multicast tree constructed by MA-CDMR and DRL-M4MR, MADRL-MR, KMB and SCTF. The results show that the average packet loss rate of all algorithms is very small. The average packet loss rate of MA-CDMR is better than that of DRL-M4MR and KMB in overall performance, with average performance improvements of 1.76% and 26.94% respectively, which is very close to the values of MADRL-MR and SCTF.
[0186] Fig.16 (d) is a comparison of the average length of the cross-domain multicast tree constructed by MA-CDMR and algorithms such as DRL-M4MR, MADRL-MR, KMB and SCTF. The results in the figure show that the length of the cross-domain multicast tree constructed by MA-CDMR is shorter than that of DRL-M4MR and MADRL-MR, which reflects the advantage of MA-CDMR strategy in multi-domain environments. However, it has a longer length than KMB and SCTF, which indicates that the algorithm considers more parameters in the construction of the multicast tree and considers more nodes when selecting nodes to join the multicast path. In the case of balancing length and performance, the MA-CDMR algorithm will make a compromise choice.
[0187] Fig.16 (e) is a comparison of the average distance between wireless AP nodes in the cross-domain multicast tree constructed by MA-CDMR and DRL-M4MR, MADRL-MR, KMB and SCTF. The results show that the average distance between AP nodes in the multicast tree constructed by MA-CDMR has achieved better results, and is better than the other four algorithms overall. Although the average length of the multicast tree of MA-CDMR is longer than that of KMB and SCTF, Fig.16 (d) shows that the distance between AP nodes does not show the same trend. This shows that the algorithm takes the distance between wireless AP nodes into consideration and achieves better results. In addition, the cross-domain multi-agent strategy is better than the single-domain multi-agent subtask strategy of the MADRL-MR algorithm.
[0188] The MA-CDMR algorithm proposed in this invention is an intelligent cross-domain multicast routing method based on multi-agent deep reinforcement learning in the SDWN multi-controller domain. First, we analyze and model the cross-domain multicast routing in the multi-controller domain, decompose the cross-domain multicast tree into an inter-domain multicast tree and multiple intra-domain multicast trees, and construct them in two stages to find the approximate optimal solution. Secondly, two kinds of agents are designed, namely, the inter-domain agent and the intra-domain agent. The inter-domain agent is responsible for constructing the inter-domain multicast tree, and each intra-domain agent is responsible for constructing the intra-domain multicast tree of the corresponding domain. The state space of the agent is designed through the traffic data collected by the SDWN multi-controller domain network architecture; and according to the characteristics of the two agent tasks, the corresponding action space is designed respectively. The action space of the inter-domain agent is designed as the set of all the edges between the domains, and each action is to select an inter-domain path of the adjacent domain; the action space of the intra-domain agent is designed as the combination of the intra-domain nodes of the corresponding domain, and each action is to select a node to join the intra-domain multicast tree. Finally, the corresponding reward function is designed according to different situations.
[0189] Based on a large number of comparative experiments, it is verified that the MA-CDMR algorithm has better performance in multi-controller domains than the classic multicast tree optimization methods KMB and SCTF, as well as algorithms such as DRL-M4MR and MADRL-MR that use reinforcement learning to build multicast trees.
[0190] The present invention utilizes a multi-controller communication mechanism and a multicast groups management module to respectively realize the transmission and synchronization of network information between different control domains of SDWN, decomposes the optimal cross-domain multicast tree into two parts: an inter-domain multicast tree and an intra-domain multicast tree, and designs an assisted multi-agent reinforcement learning solution algorithm for each controller. The multi-agent reinforcement learning training method combining online and offline is used to reduce the dependence on the real-time environment and accelerate the convergence speed of the multi-agent.
[0191] It should be noted that although the embodiments of the present invention described above are illustrative, they are not intended to limit the present invention, and therefore the present invention is not limited to the above specific embodiments. Without departing from the principles of the present invention, any other embodiments obtained by those skilled in the art under the guidance of the present invention are deemed to be within the protection of the present invention.
Claims
1. An intelligent cross-domain multicast routing method based on multi-agent deep reinforcement learning, characterized by: The steps include: Step 1: deploy an inter-domain agent on the root controller of the multi-domain software-defined wireless network and deploy an intra-domain agent on the local controller of each domain; Step 2: Initialize the Actor network parameters θ between domains and within domains int ,θ intra 、Critic network parametersω int ,ω intra and data buffer area B int ,B intra ; Step 3: The root controller and each local controller communicate and collect information of the multi-domain software-defined wireless network through the controller communication mechanism, and analyze the information collected by the corresponding controller through the multicast group management module to determine the domains where the source node and each destination node are located, as well as the forwarding boundary nodes of each domain; Step 4: First, the inter-domain agent interacts with the environment to generate an inter-domain network link information matrix and the inter-domain multicast tree state matrix And stack the inter-domain network link information matrix and the inter-domain multicast tree state matrix Get the current state s of the inter-domain agent t ; Then, the inter-domain agent changes from the current state s t Sample the current action a from the output action set t , execute the current action a t Get the current reward value r t and the next state s t+1 , where the action set of the inter-domain agent is the set of edges between all domains, and each action is selected as one of the edges; finally, the reinforcement learning data (s t ,a t ,r t ,s t+1 ) is stored in the data cache area B of the inter-domain agent int ; Step 5: When the inter-domain agent repeats the action of step 4, the inter-domain agent is trained by offline and online hybrid training, that is: Every set period of time, i.e., the inter-domain period, the data cache area B is used first during this period of time. int The inter-domain agent is trained offline without interacting with the environment using the offline reinforcement learning data in , so as to update the Actor network parameters θ of the inter-domain agent. int and Critic network parameter ω iny , and then enter the next inter-domain agent action process; When the set inter-domain update frequency is reached, the online reinforcement learning data obtained from this action is first used to conduct online training for the inter-domain agent to interact with the environment, so as to update the Actor network parameters θ of the inter-domain agent. int and Critic network parameter ω int , and then enter the next inter-domain agent action process; Step 6: In the inter-domain multicast tree T int When the construction is completed, the inter-domain agent will construct the inter-domain multicast tree T int Synchronize to each agent in the domain; Step 7: First, each agent in the domain interacts with the environment to generate the network link information matrix in its own domain. and the intra-domain multicast tree state matrix And stack its own intra-domain network link information matrix and the intra-domain multicast tree state matrix Get the current state s of the agent in the domain t ; Then, each agent ring in the domain starts from the current state s t Sample the current action a from the output action set t , execute the current action a t Get the current reward value r t and the next state s t+1 , where the action set of the agent in the domain is the set of nodes in the domain, and each action is selected as the next hop node; finally, the reinforcement learning data (s t ,a t ,r t ,s t+1 ) is stored in the data buffer area B of the agent in the domain intra ; Step 8: When each agent in the domain repeats the action of step 7, each agent in the domain is trained by offline and online hybrid training, that is: Every set period of time, i.e., the domain cycle, the data buffer area B is used first during this period of time. intra The offline reinforcement learning data in the domain is used to train the in-domain agent offline without interacting with the environment to update the Actor network parameters θ of the in-domain agent. intra and Critic network parameter ω intra , and then enter the next action process of the agent in the domain; When the set domain update frequency is reached, the online reinforcement learning data obtained from this action is first used to conduct online training on the domain agent to interact with the environment, so as to update the Actor network parameters θ of each domain agent. intra and Critic network parameter ω intra , and then enter the next action process of each agent in the domain; Step 9: Multicast tree T in each domain imtra When all are constructed, the combined inter-domain multicast tree T int and all intra-domain multicast trees T intra Combined into a cross-domain multicast tree; Step 10: Determine whether the cross-domain multicast tree converges or reaches a preset number of iterations: If so, the multi-domain software-defined wireless network performs intelligent cross-domain multicast routing based on the current cross-domain multicast tree; otherwise, return to step 4 and perform the next round of iteration.
2. According to claim 1, the intelligent cross-domain multicast routing method based on multi-agent deep reinforcement learning is characterized in that: In steps 4 and 7, the inter-domain network link information matrix and intra-domain network link information matrix The network link information in the ,includes the remaining bandwidth, delay, packet loss rate, packet error rate and the,distance between APs.
3. According to claim 1, the intelligent cross-domain multicast routing method based on multi-agent deep reinforcement learning is characterized in that: In step 4 and step 7, the current action a is executed for the inter-domain agent and the intra-domain agent t Get the current reward value r t : If the current action a is executed t After that, the added next hop node or link just adds a normal node to the multicast tree, so the current reward value r t is the single-step reward R part : R part =β1bw ij +β2(1-delay ij )+β3(1-loss ij )+β4(1-err ij )+β5(1-dist ij ) If the next hop node or link added is to add a destination node to the multicast tree k , then the current reward value r t Reward R for the subtask end : R end =β1bw k +β2(1-delay k )+β3(1-loss k )+β4(1-err k )+β5(1-dist k ) If the added next-hop node or link causes the multicast tree to form a loop, the current reward value r t If the current reward value is r t is the penalty value R loop : R loop =C1 If a t is an invalid action, which is neither the next hop node of the current node nor the edge connected to the current domain neighborhood. Then the current reward value r t is the penalty value R hell : R hell =C2 Where β1, β2, β3, β4, and β5 represent the weights of remaining bandwidth, delay, packet loss rate, packet error rate, and distance between APs, respectively; bw ij 、delay ij 、loss ij 、err ij 、dist ij They represent the links e added to the multicast tree. ij The remaining bandwidth, latency, packet loss rate, packet error rate and distance between APs; bw k 、delay k 、loss k 、err k 、dist k Respectively represent the source node src to the destination node d k The remaining bandwidth, delay, packet loss rate, packet error rate and distance between APs of the entire link; C1 and C2 represent two constants respectively.
4. The intelligent cross-domain multicast routing method based on multi-agent deep reinforcement learning according to claim 1 is characterized in that: In step 5 and step 8, the process of training the agent using reinforcement learning data is: First, the reinforcement learning data (s t ,a t ,r t ,s t+1 ) in the current state s t and the next state s t+1 Input the agent's Critic network to calculate the current expected value V ω (s t ) and the next expected value V ω (s t+1 ); Then, based on the reinforcement learning data (s t ,a t ,r t ,s t+1 ) in the current reward value r t , and the current expected value V calculated above ω (s t ) and the next expected value V ω (s t+1 ) Calculate the time series difference residual ψ t , and adopt the learning method of time series difference residual based on the time series difference residual ψ t Update the agent's Critic network parameters ω and Actor network parameters θ.
5. The intelligent cross-domain multicast routing method based on multi-agent deep reinforcement learning according to claim 1 is characterized in that: In step 5 and step 8, the inter-domain period and the intra-domain period are the same or different, and the inter-domain update frequency and the intra-domain update frequency are the same or different.
Citation Information
Cited By
Many-to-many communication routing method based on multi-agent graph reinforcement learning in SDWN
CN119966873A