Multi-controller deployment method and system based on SDN network and deep reinforcement learning

By adopting deep reinforcement learning algorithms in the SDN network and optimizing the deployment of multiple controllers, the problems of single point failure and limited resources in the SDN network are solved, and the effects of reducing latency, improving network performance and enhancing destruction resistance are achieved.

CN114355775BActive Publication Date: 2025-05-09AEROSPACE SCI & ENG NETWORK INFORMATION DEV CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202111641069.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-12-29
Publication Date
2025-05-09
Estimated Expiration
2041-12-29

AI Technical Summary

Technical Problem

In SDN networks, single point of failure and limited controller resources lead to increased communication consumption, and unreasonable deployment of multiple controllers may lead to network congestion or paralysis, affecting the scalability of the network.

Method used

Using a multi-controller deployment method based on SDN network and deep reinforcement learning, by obtaining optimization goals such as communication delay between the switch and the controller, communication delay between the controller, synchronization overhead, minimum security, load constraints and bandwidth constraints, the deep reinforcement learning algorithm is used to deploy controllers with a certain number and near-optimal locations for a given network topology.

Benefits of technology

It realizes that in complex SDN networks and special application scenarios, reduce latency, improve network performance, avoid control node crashes, enhance network damage resistance and reliability, and realize precise network management.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114355775B_ABST
    Figure CN114355775B_ABST
Patent Text Reader

Abstract

The present invention provides a multi-controller deployment method based on SDN network and deep reinforcement learning, comprising: obtaining a first performance optimization index for multi-controller deployment according to the SDN network structure; the first performance optimization index includes: average propagation delay between a switch and a controller, average propagation delay between controllers, synchronization overhead between controllers and minimum security of synchronization between controllers; establishing a first objective function according to the first performance optimization index; obtaining a first constraint condition for multi-controller deployment; the first constraint condition includes controller load constraint, mapping relationship constraint between switch and controller and control layer synchronization link bandwidth constraint; constructing a multi-controller deployment model according to the first objective function and the first constraint condition; solving the multi-controller deployment model based on a Markov decision model and using a deep reinforcement learning algorithm to obtain a multi-controller deployment solution.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of controller deployment technology, and in particular to a method and system for deploying multiple controllers based on SDN network and deep reinforcement learning. Background Art

[0002] With the increase of network traffic and the continuous expansion of network scale, the inherent defects of single controllers such as single point failure and limited controller resources are becoming increasingly prominent, which will increase the communication consumption of the control link. In addition, if multiple controllers are used to manage the network, the unreasonable deployment of multiple controllers may cause network congestion or paralysis when meeting network service needs, which will have a great impact on the scalability of the SDN network, so the deployment of multiple controllers is particularly important. The deployment of controllers will also have an important impact on the performance, reliability and network cost of the SDN network. Therefore, the present invention proposes a multi-controller deployment method and system based on SDN network and deep reinforcement learning. Summary of the invention

[0003] The purpose of the present invention is to provide a method and system for deploying multiple controllers based on SDN networks and deep reinforcement learning. On the control plane, according to optimization objectives such as communication delay between switches and controllers, communication delay between controllers, synchronization overhead between controllers, minimum security of controllers, controller load constraints and bandwidth constraints of control links, a deep reinforcement learning algorithm is used to deploy controllers with a certain number and near-optimal positions for a given network topology.

[0004] To achieve the above object, the present invention provides the following solutions:

[0005] A multi-controller deployment method based on SDN network and deep reinforcement learning, comprising:

[0006] Obtaining a first performance optimization index for multi-controller deployment according to the SDN network structure; the first performance optimization index includes: an average propagation delay between a switch and a controller, an average propagation delay between controllers, a synchronization overhead between controllers, and a minimum security of synchronization between controllers;

[0007] Establishing a first objective function according to the first performance optimization indicator;

[0008] Obtaining a first constraint condition for multi-controller deployment; the first constraint condition includes a controller load constraint, a mapping relationship constraint between a switch and a controller, and a control layer synchronization link bandwidth constraint;

[0009] Constructing a multi-controller deployment model according to the first objective function and the first constraint condition;

[0010] Based on the Markov decision model, a deep reinforcement learning algorithm is used to solve the multi-controller deployment model to obtain a multi-controller deployment solution.

[0011] A system based on SDN network and deep reinforcement learning multi-controller deployment method, comprising:

[0012] A first performance optimization index acquisition module is used to acquire a first performance optimization index for multi-controller deployment according to the SDN network structure; the first performance optimization index includes: an average propagation delay between a switch and a controller, an average propagation delay between controllers, a synchronization overhead between controllers, and a minimum security of synchronization between controllers;

[0013] A first objective function establishing module, used to establish a first objective function according to the first performance optimization index;

[0014] A first constraint condition acquisition module, used to acquire a first constraint condition for multi-controller deployment; the first constraint condition includes a controller load constraint, a mapping relationship constraint between a switch and a controller, and a control layer synchronization link bandwidth constraint;

[0015] A multi-controller deployment model building module, used to build a multi-controller deployment model according to the first objective function and the first constraint condition;

[0016] The solution module is used to solve the multi-controller deployment model based on the Markov decision model using a deep reinforcement learning algorithm to obtain a multi-controller deployment solution.

[0017] According to the specific embodiments provided by the present invention, the present invention discloses the following technical effects:

[0018] The present invention provides a method and system for deploying multiple controllers based on SDN network and deep reinforcement learning, including: obtaining a first performance optimization index for multiple controller deployment according to the SDN network structure; the first performance optimization index includes: average propagation delay between switch and controller, average propagation delay between controllers, synchronization overhead between controllers and minimum security of synchronization between controllers; establishing a first objective function according to the first performance optimization index; obtaining a first constraint condition for multiple controller deployment; the first constraint condition includes controller load constraint, mapping relationship constraint between switch and controller and control layer synchronization link bandwidth constraint; constructing a multiple controller deployment model according to the first objective function and the first constraint condition; solving the multiple controller deployment model using a deep reinforcement learning algorithm based on a Markov decision model to obtain a multiple controller deployment solution. For special application scenarios such as complex SDN networks and battlefields, a multi-controller deployment mechanism is proposed to reduce latency, improve network performance, and avoid the control node crash caused by frequent service flow from the control node. And according to the interactive information of the control layer in special application scenarios, the synchronization data packet field format of the control layer is flexibly designed. The present invention establishes an optimized deployment model for the cluster, so that when a controller in the network is damaged and stops working, other controllers can take over all nodes under the control of the faulty controller to ensure that the communication of all nodes is not interrupted and enhance the anti-destruction capability of the control nodes. At the same time, the deep reinforcement learning algorithm is used to establish a reliable and stable data transmission channel to achieve precise management of the network. BRIEF DESCRIPTION OF THE DRAWINGS

[0019] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.

[0020] Figure 1 A flow chart of a method for deploying multiple controllers based on SDN network and deep reinforcement learning provided in Example 1 of the present invention;

[0021] Figure 2 A flow chart of synchronization between controllers provided in Embodiment 1 of the present invention;

[0022] Figure 3 A neural network structure diagram provided in Example 1 of the present invention;

[0023] Figure 4 A block diagram of a system for deploying multiple controllers based on SDN networks and deep reinforcement learning provided in Example 2 of the present invention. DETAILED DESCRIPTION

[0024] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0025] The purpose of the present invention is to provide a method and system for deploying multiple controllers based on SDN networks and deep reinforcement learning. On the control plane, according to optimization objectives such as communication delay between switches and controllers, communication delay between controllers, synchronization overhead between controllers, minimum security of controllers, controller load constraints and bandwidth constraints of control links, a deep reinforcement learning algorithm is used to deploy controllers with a certain number and near-optimal positions for a given network topology.

[0026] In order to make the above-mentioned objects, features and advantages of the present invention more obvious and easy to understand, the present invention is further described in detail below with reference to the accompanying drawings and specific embodiments.

[0027] Most of the models established by the prior art are based on single-objective optimization. The multi-objective optimization model uses the delay between the switch and the controller and the delay between the controllers as optimization indicators, and the controller load as a constraint. There are few existing models that involve the overhead of synchronization information between controllers as an optimization indicator. Although this method is simple to implement, the optimization effect is not good. At the same time, most deployment schemes do not consider the optimized deployment of clusters. There are few existing technologies for Atomix to assist in the deployment of ONOS controller clusters, and no optimization is performed for the deployment method in which Atomix nodes are physically separated from ONOS controllers. The deployment problem of multiple controllers is an NP-hard problem and the calculation is very time-consuming. The current solution algorithms mainly include integer linear programming algorithms, heuristic algorithms, etc. These algorithms have high complexity, poor scalability, and are prone to local optimality. Therefore, the present invention studies the use of a distributed coordination framework Atomix to assist ONOS controllers in establishing clusters. The present invention proposes a method for establishing a full-network control layer by using the reasonable deployment of distributed controllers in complex network environments such as SDN scenarios and changing battlefields. The present invention focuses on the design of the synchronization data packet format at the control layer, and the reasonable and efficient implementation of the deployment of distributed controllers under the constraints of the synchronization overhead of the control layer. The design plan mainly includes three aspects: control layer synchronization message design, establishment of a flat multi-controller deployment model, and establishment of a cluster deployment optimization model of the distributed coordination framework Atomix.

[0028] Example 1

[0029] like Figure 1As shown, this embodiment provides a multi-controller deployment method based on SDN network and deep reinforcement learning, including:

[0030] S1: Obtaining a first performance optimization index for multi-controller deployment according to the SDN network structure; the first performance optimization index includes: average propagation delay between the switch and the controller, average propagation delay between controllers, synchronization overhead between controllers, and minimum security of synchronization between controllers;

[0031] Specifically, the expression of the average propagation delay between the switch and the controller is: Where N is the number of switches, d ij is the shortest link delay between the switch and the controller, x ij It is a binary number. When the value is 1, it indicates that the switch i is successfully connected to the controller j. i is switch i, S is the set of switches, c j is controller j, and C is the controller set.

[0032] The expression of the average propagation delay between the controllers is: c k is the controller k, K is the number of controllers, is the shortest link delay between controllers;

[0033] The expression of the synchronization overhead between controllers is: Among them, l jk is the length of the data packet synchronized between controller j and controller k, p sjk is the synchronization frequency between controller j and controller k.

[0034] S2: Establishing a first objective function according to the first performance optimization indicator;

[0035] Specifically, the first objective function is:

[0036] minimize(αT scavg +βT ccavg +ρC cc )+μ·K

[0037] Among them, T scavg represents the average propagation delay between the switch and the controller; T ccavg represents the average propagation delay between controllers; C cc represents the synchronization overhead between controllers; μ is the safety factor for achieving the minimum security of synchronization between controllers; α, β, ρ are the T scavg , T ccavg , C cc The weight of , α+β+ρ=1.

[0038] S3: Obtain a first constraint condition for multi-controller deployment; the first constraint condition includes a controller load constraint, a mapping relationship constraint between switches and controllers, and a control layer synchronization link bandwidth constraint;

[0039] Specifically, the controller load constraint is:

[0040]

[0041] Among them, L opt is the expected optimal controller load quantity; L c is the actual controller load quantity; L j is the load quantity of controller j, LBI is used to measure the load quantity difference among all controllers in the network, and is defined as the load deviation index; ΔLB is the difference in the number of switches managed by each controller, and is defined as the average value of the load difference; LB C is the threshold value of ΔLB;

[0042] The mapping relationship between the switch and the controller is constrained as follows:

[0043]

[0044] Where V represents the set of switches.

[0045] The control layer synchronization link bandwidth constraint is:

[0046]

[0047] Among them, B W The bandwidth provided to the control layer physical link.

[0048] S4: constructing a multi-controller deployment model according to the first objective function and the first constraint condition;

[0049] S5: Based on the Markov decision model, a deep reinforcement learning algorithm is used to solve the multi-controller deployment model to obtain a multi-controller deployment solution.

[0050] Wherein, step S5 specifically includes:

[0051] Obtaining the state space of the Markov decision model according to the controller placement of each node in the network at the current time t and the network topology information;

[0052] Obtaining the action space of the Markov decision model according to the number of controllers deployed in the network, the node locations where the controllers are deployed, and the mapping relationship between the switches and the controllers;

[0053] Obtaining the state transition probability of the Markov decision model according to the probability of transferring to the next state after executing a certain action in the current state;

[0054] Obtaining a reward function of the Markov decision model according to an average communication delay between the switch and the controller, an average communication delay between the controllers, a synchronization overhead between the controllers, and a minimum safety of the controller;

[0055] A multi-controller deployment scheme is obtained based on the state space, the action space, the state transition probability and the reward function of the Markov decision.

[0056] As another optional implementation, an Atomix node is embedded on the basis of the multi-controller deployment model, and the Atomix node is optimally deployed, specifically including:

[0057] (1) obtaining a second performance optimization indicator for Atomix node deployment; the second performance optimization indicator includes an average synchronization delay between Atomix nodes and an average synchronization delay between an Atomix node and a controller node;

[0058] Specifically, the average synchronization delay between the Atomix nodes is: Where A is the number of Atomix nodes, is the shortest link delay between Atomix nodes, a j ,a k ,a i is Atomix node j, k, i;

[0059] The average synchronization delay between the Atomix node and the controller node is:

[0060]

[0061] is the shortest link delay between the Atomix node and the controller node, z ij It is a binary number. When the value is 1, it indicates that the Atomix node i is successfully connected to the controller j.

[0062] (2) constructing a second objective function according to the second performance optimization index;

[0063] Specifically, the second objective function is:

[0064] minimizeρ1T aaavg +ρ2T acavg , ρ1, ρ2 are respectively the aaavg , T acavg The weight of , ρ1+ρ2=1.

[0065] (3) obtaining a second constraint condition for Atomix node deployment; the second constraint condition is that the number of Atomix nodes deployed in the network is B, and there is at least one mapping relationship between the Atomix node and the controller node;

[0066] Specifically, there is at least one mapping relationship between the Atomix node and the controller node:

[0067]

[0068] (4) constructing an Atomix node deployment model according to the first objective function and the first constraint condition;

[0069] (5) Use deep reinforcement learning algorithm to solve the Atomix node deployment model.

[0070] For the solution method, please refer to the part of solving the multi-controller deployment model.

[0071] In this embodiment, based on the changeable and complex SDN scenarios and actual battlefield and other network environments, a multi-controller deployment mechanism based on SDN network and deep reinforcement learning is proposed, and multiple controllers are efficiently and reasonably placed in the network to achieve collaborative management of the entire network by multiple controllers, reduce network latency and bandwidth overhead in special application scenarios, improve the utilization of network resources, balance the load under each controller, enhance the robustness of the network, and thus reduce the probability of network failures caused by controller crashes. Taking into account network latency, overhead, load and other performance, a control layer management method for special application scenarios is established to achieve distributed management of each node in the network. In summary, the innovative points of the present invention are:

[0072] (1) Establish a multi-controller deployment model. According to the network requirements in special application scenarios, multiple network performance optimization indicators are selected comprehensively. In order to enable the control layer to collaboratively manage the network and not affect the normal operation of the network when a controller fails, in addition to considering latency and load, the synchronization overhead between controllers, the minimum security of the controllers, and the bandwidth constraints of the control layer links are used as network performance optimization indicators. Considering multiple factors affecting deployment, the special requirements in special application scenarios are met and the controllers are deployed reasonably.

[0073] Design a multi-controller deployment model. Select and define the network performance optimization indicators that need to be considered when deploying controllers. By analyzing the relationship between the optimization indicators and the deployment problems that the model needs to solve, derive the objective function, complete the model establishment, and solve it through a deep reinforcement learning algorithm.

[0074] (2) In order to optimize the controller synchronization overhead in the multi-controller deployment model, the present invention designs the data packet field format for synchronization between controllers based on the controller synchronization requirements of special application scenarios, thereby optimizing the network performance indicators of the controller deployment model in a targeted manner and establishing a more accurate and flexible deployment model.

[0075] Design control plane synchronization data packet information. By designing the information that needs to be synchronized between controllers according to the requirements of different application scenarios, a flexible optimization target is provided for the controller deployment solution.

[0076] (3) For collaborative management of the control layer, an optimized deployment model of the controller cluster is established. The network control layer established based on the ONOS controller needs to use the distributed coordination framework Atomix. The present invention proposes an Atomix node deployment model to optimize the number and location of Atomix deployment.

[0077] In order to make the existing ONOS controller version and Atomix nodes physically separate and make cluster deployment more flexible, an Atomix node deployment optimization solution is designed to improve network synchronization performance.

[0078] (4) Aiming at the NP-Hard problem of multi-controller deployment, the present invention uses a deep reinforcement learning algorithm to solve the model. When deployed in actual network application scenarios, the distance between nodes is relatively far, and the optimal result needs to be calculated within a limited time. The deep reinforcement learning algorithm has low complexity and integrates historical network data learning into controller deployment and switch controller mapping decisions to adapt to the network environment.

[0079] For the established model, the deep reinforcement learning algorithm is applied to solve the deployment model, making the results of the present invention more efficient and reasonable.

[0080] In order to enable those skilled in the art to more clearly understand the solution of this embodiment, a detailed description will be given below:

[0081] (I) Design of control plane synchronization message based on ONOS controller

[0082] Due to the complexity and variability of application scenarios, the present invention needs to establish a large-scale complex SDN network platform. At the same time, in a complex network environment, the controller needs to frequently send control information such as flow tables and service flows. However, a single controller is difficult to carry functions such as sending control information for all nodes in the network, which may lead to conflicts in the sending of control information and cause the crash of the control node. Therefore, in order to avoid single point failures and improve the response speed and overall performance of the network, it is necessary to deploy multiple controllers throughout the network to manage the network.

[0083] In the mechanism of the present invention, the entire network is divided into multiple subdomains, and each controller controls an area in the network. Since each controller needs to assume the task of calculating the strategy for the switch it controls, the controller of each domain needs to master not only the switch topology relationship within its own control range, but also the switch topology relationship within the control range of other controllers. Therefore, it is necessary to establish a synchronous communication mechanism between the domain controllers to synchronize the topology information of each domain with each other. At the same time, in order to avoid the loss of control of the switch caused by controller failure, when the corresponding controller of the switch fails, other non-faulty controllers can still send control information to the out-of-control switch quickly and promptly after taking over the out-of-control switch, and the domain controllers also need to synchronize the calculated strategy and other information.

[0084] To this end, take two controllers as an example to illustrate the synchronization process between controllers. Figure 2 As shown, it is divided into three steps:

[0085] Each domain controller collects topology information at the data level;

[0086] The controllers establish a connection through the TCP three-way handshake protocol;

[0087] Controller A and controller B calculate strategies separately;

[0088] The controllers send synchronization data packets to each other to synchronize information;

[0089] After the synchronization is completed, controller A and controller B send ACK confirmation messages to each other to terminate the synchronization.

[0090] In view of step ②, the present invention designs the format of the synchronization data packet. In the present invention, the information synchronized between controllers includes topology information, flow table information, and specific service flow information. The three synchronization data packet field formats are shown in Table 1, Table 2, and Table 3.

[0091] Table 1 Topology information synchronization data packet field format

[0092]

[0093] Table 2 Flow table information synchronization data packet field format

[0094]

[0095] Table 3 Specific service flow information synchronization data packet field format

[0096]

[0097] 2. Establishment of a distributed flat multi-controller deployment model

[0098] The deployment of multiple controllers needs to focus on three key issues:

[0099] Given a network topology, calculate the number of controllers that need to be deployed;

[0100] Determine the best location for the deployment controller;

[0101] Which controller should manage each switch?

[0102] Therefore, starting from the above three key issues, the present invention establishes a deployment model of the controller according to large-scale SDN and complex network application scenarios.

[0103] Physical network: The network topology is composed of an undirected graph G(V,E), where V represents the set of switches and E represents the set of physical links. K represents the number of controllers in the network, and C = {c1,...,c k} represents a controller set. In the present invention, each controller is placed at a certain switch position in the network. θ Indicates the deployment location of the controller. Indicates that the controller θ i The mapping relationship between the switch and the controller can be expressed as a set

[0104]

[0105] The average propagation delay between the switch and the controller and the average propagation delay between controllers are related to the deployment location of the controller. Based on this, the network performance optimization indicators are selected as follows:

[0106] ① Average propagation delay between the switch and the controller: represents the average propagation delay between the switch and the controller. As shown in (Equation 1). Where N = |V| is the number of switches, d ij is the shortest link delay between the switch and the controller, x ij It is a binary number. When the value is 1, it indicates that the switch i is successfully connected to the controller j.

[0107]

[0108] ② Average propagation delay between controllers: represents the average value of propagation delay between controllers. As shown in (Equation 2). Where K is the number of controllers, is the shortest link delay between controllers.

[0109]

[0110] ③ Inter-controller synchronization overhead: refers to the communication overhead generated when the controllers are synchronized. As shown in (Equation 3). It is related to the data packet format and synchronization frequency of the inter-controller synchronization. jk is the length of the data packet synchronized between controller j and controller k, p sjk is the synchronization frequency between controller j and controller k.

[0111]

[0112] ④Minimum security of synchronization between controllers:

[0113] Since the controller can obtain the whole network topology information through communication with other controllers, in order to reduce the probability of information leakage when the network is attacked and enhance the security of the network, it is necessary to minimize the number of deployed controllers. In this model, the safety factor μ is set to constrain the number of deployed controllers.

[0114] In view of the above description of the optimization index selected by the model, a multi-controller deployment model based on delay and synchronization overhead is established in the present invention. The controller deployment is realized by comprehensively considering the delay and synchronization overhead, combining the constraints such as the controller load and the minimum security of synchronization. The optimization target is shown in (Formula 4).

[0115] minimize(αT scavg +βT ccavg +ρC cc )+μ·K(4)

[0116] In this model, there is a contradictory relationship between the average propagation delay between the switch and the controller and the average propagation delay between the controllers. That is to say, these two network performance optimization indicators restrict each other. In order to minimize the average propagation delay between the switch and the controller, the controller will be deployed at a position close to each switch, which will increase the distance between the controllers and increase the average propagation delay between the controllers. And vice versa. Under this relationship, there is no solution that makes all performance optimization indicators reach the optimal solution. Usually, if one optimization indicator is improved, the performance of other optimization indicators will be sacrificed. Therefore, in the present invention, weights are set for each network performance optimization indicator, where α+β+ρ=1. According to the focus of each optimization indicator in the specific application scenario, an efficient and flexible deployment model is established.

[0117] This model needs to meet the following constraints:

[0118] ① Controller load constraints:

[0119] During the controller deployment process, this model needs to meet the controller load limit, that is, the number of switches managed by each controller cannot exceed a specific threshold. The difference in the number of switches managed by each controller is defined as the load difference average value ΔLB, as shown in (Equation 7). This value cannot exceed the specified threshold LB C .L opt is the expected optimal number of controller loads. The load deviation index LBI is defined as the load quantity difference between all controllers in the network, as shown in (Equation 6). The smaller the index is, the better the network load balancing performance is after the controller is reasonably deployed.

[0120]

[0121] ②Mapping relationship between switch and controller:

[0122] In this model, the number of controllers deployed in the network is K, and there is only one mapping relationship between the switch and the controller. At the same time, it is necessary to ensure that each switch has a controller to which it belongs. Therefore, the mapping relationship needs to satisfy the following equation (8).

[0123]

[0124] x sc ≤y c (8)

[0125]

[0126] ③Control layer synchronization link bandwidth constraint: Controllers are deployed in the nodes

[0127] In specific application environments such as large-scale SDN or complex networks, the bandwidth requirement for synchronization information between controllers cannot exceed the bandwidth resources provided by the physical link of the control layer. The bandwidth provided by the physical link of the control layer is set to B. W , then the bandwidth constraint is as shown in (Equation 9).

[0128]

[0129] (III) Cluster deployment optimization model based on the distributed coordination framework Atomix

[0130] Since the present invention adopts the ONOS controller, the management of the cluster in the environment of the ONOS controller needs to adopt the distributed coordination framework Atomix, which physically separates the functions such as cluster management, service discovery and persistent data storage from the ONOS node itself. According to the understanding of Atomix, before deploying the ONOS controller cluster, an Atomix cluster must be formed first for data storage and coordination, and then the ONOS node is configured using the list of Atomix nodes to be connected. At the same time, in the past version of the ONOS controller, the Atomix node needs to be embedded in it to form a cluster and synchronize the state. In the existing version of the ONOS controller, functions such as synchronization status are moved to a separate Atomix cluster. According to the study of the Atomix framework, it can be deployed on non-control nodes or embedded in control nodes. However, according to the reference to relevant information, the deployment information of Atomix nodes is very scarce. Therefore, it is also particularly important to effectively select the number and location of Atomix nodes in large-scale SDN. Therefore, based on the model established by embedding the Atomix nodes in (1) and (2) into the ONOS controller, the Atomix nodes are optimized for deployment. The optimized deployment is used to establish a model to determine the number and location of the Atomix nodes, thereby achieving synchronization of the status information of the entire network.

[0131] In the model for optimizing Atomix node deployment, since there is a Raft protocol to maintain strong consistency between Atomix nodes and ONOS nodes and between Atomix nodes, the present invention selects and defines performance optimization indicators.

[0132] ① Average synchronization delay between Atomix nodes: represents the average propagation delay of synchronization information between Atomix nodes. As shown in (Equation 10). Where A is the number of Atomix nodes, is the shortest link delay between Atomix nodes.

[0133]

[0134] ② Average synchronization delay between Atomix nodes and ONOS nodes: represents the average propagation delay of synchronization information between Atomix nodes and ONOS nodes. As shown in (Equation 11). Where K is the number of ONOS controllers, is the shortest link delay between the Atomix node and the ONOS node, z ij A binary number with a value of 1 indicating a successful connection between Atomix node i and ONOS controller j.

[0135]

[0136] In view of the above description of the optimization index of the model selection, the present invention establishes the Atomix optimization deployment model based on the time delay distributed coordination framework. The optimization target is shown in (Formula 12).

[0137] minimizeρ1T aaavg +ρ2T acavg (12)

[0138] In this model, the average synchronization delay between Atomix nodes and the average synchronization delay between Atomix nodes and ONOS nodes constrain each other. In order to reduce the average synchronization delay between Atomix nodes, the deployment of Atomix nodes will be more clustered, which will increase the average synchronization delay between Atomix nodes and ONOS nodes. And vice versa. Therefore, in the present invention, in order to balance the relationship between the two, by setting the weight, where ρ1+ρ2=1, according to the specific application scenario, an efficient and flexible Atomix node deployment model is established.

[0139] The constraints that this model needs to meet are:

[0140] Since in this model, the number of Atomix nodes deployed in the network is B, there is at least one mapping relationship between the Atomix nodes and the ONOS nodes, and the mapping relationship needs to satisfy the following equation (13).

[0141]

[0142] Design of adaptive multi-controller deployment algorithm framework

[0143] In view of the established multi-controller deployment model, the present invention proposes an adaptive multi-controller deployment algorithm based on deep reinforcement learning, so that the present invention can realize the deployment of controllers more efficiently and accurately. The controller deployment problem is converted into an MDP model for solution. The state space, action space, state transition probability and reward function are set to a four-tuple (S, A, P, R). The definitions are as follows:

[0144] (1) State space S

[0145] In the present invention, the state space can be expressed as the controller placement of each node in the network and the network topology information at the current time t. It is expressed as follows:

[0146]

[0147] The meanings of each element are as follows:

[0148] Represents the physical network topology information at time t, including t t , c t .

[0149] t t : represents the delay of each link at time t.

[0150] c t : represents the synchronization overhead of each control link at time t.

[0151] b t : Represents the load situation of the deployed control nodes at time t.

[0152] f t : represents the probability of failure of each node at time t.

[0153] ω t : represents the controller placement of each node at time t, including

[0154] Represents the number of controller deployments at time t.

[0155] Represents the deployment location of the controller at time t.

[0156] δ t : represents the Atomix placement of each node at time t, including

[0157] Represents the number of Atomix deployments at time t.

[0158] Represents the deployment position of Atomix at time t.

[0159] (2) Action Space A

[0160] In the present invention, the action space A is represented by the number of controllers that need to be deployed in the network, the location of the appropriate node where the controller is deployed, and the mapping relationship between the switch and the controller. It is represented as follows:

[0161] a t =(p θ ,S C ),θ∈(1,2,...,K)

[0162] a t ′=p c ,c∈(1,2,...,B)

[0163] The meanings of each element are as follows:

[0164] p θ : Indicates the appropriate location where the controller is deployed.

[0165] p c : Indicates the location where Atomix is ​​deployed.

[0166] S C : Represents the mapping relationship between the switch and the controller.

[0167] K: represents the number of controllers that need to be deployed in the network.

[0168] B: represents the number of Atomix that needs to be deployed in the network.

[0169] (3) State transition probability P

[0170] In the present invention, the state transition probability represents the transition from the current state s t Perform an action t Then transfer to the next state s t+1 The probability of . It is expressed as follows:

[0171]

[0172] The meanings of each element are as follows:

[0173] s t : Represents the current status.

[0174] s t+1 : Represents the next state.

[0175] a t : Represents an action in the current state.

[0176] (4) Reward function R

[0177] In the present invention, each action will generate a reward value according to the set reward function. The larger the reward, the higher the value of the action, and the better the performance that can be achieved by the controller deployment. Therefore, the average communication delay between the switch and the controller, the average communication delay between the controllers, the synchronization overhead between the controllers, and the minimum security of the controller are set as the reward function. It is expressed as follows:

[0178] r1=-((αT 1t +βT 2t +ρC cc )+μ·K)

[0179] r2=-(ρ1T 3t +ρ2T 4t )

[0180] The meanings of each element are as follows:

[0181] α: represents the proportion of the average communication delay between the switch and the controller in the reward and penalty measures during the deployment process.

[0182] β: represents the proportion of the average communication delay between controllers in the reward and penalty measures during the deployment process.

[0183] ρ: represents the proportion of synchronization overhead between controllers in the reward and penalty measures during the deployment process.

[0184] μ: represents the safety factor of the controller during the deployment process.

[0185] ρ1: represents the proportion of the communication delay between Atomix in the reward and penalty measures during the deployment process.

[0186] ρ2: represents the proportion of the communication delay between Atomix and the controller in the reward and penalty measures during the deployment process.

[0187] T 1t : Represents the average communication delay between the switch and the controller during the deployment process.

[0188] T 2t : Represents the average communication delay between controllers during the deployment process.

[0189] T 3t : Represents the communication delay between Atomix during the deployment process.

[0190] T 4t : Represents the communication delay between Atomix and the controller during the deployment process.

[0191] C cc : Represents the synchronization overhead between controllers during the deployment process.

[0192] K: represents the number of controllers in the deployment process.

[0193] In the present invention, the reward function takes into account the importance of link delay for deployment, and assigns different weights to the delay between the switch and the controller and the delay between controllers according to the actual application scenario; at the same time, considering the synchronization overhead and minimum security, the more the deployment result meets the above four goals, the greater the reward value.

[0194] In order to enable the model to select the optimal action under a certain network state and maximize the cumulative reward value, so that the deployment result is more accurate, the present invention establishes a neural network structure to enable the intelligent agent to better perceive environmental information such as the network topology state, thereby better interacting and learning with the environment to generate a better strategy.

[0195] The state at each moment is used as the input of the neural network, and the dimension of the input state determines the number of neurons in the input layer. The two middle layers of the neural network are fully connected, and the output is the Q value of all possible actions executed under the input state. The number of output neurons is determined by the size of the action set.

[0196] The DQN algorithm used in the present invention is an offline learning method, which sets the parameters of the neural network, and deploys each control node in the network through the results of the training output model. Figure 3 As shown, the structure of the neural network is given. The following is the multi-controller deployment decision algorithm training process design and Atomix deployment decision algorithm training process design.

[0197]

[0198]

[0199]

[0200] Example 2

[0201] like Figure 4 As shown, this embodiment provides a system based on SDN network and deep reinforcement learning multi-controller deployment method, including:

[0202] A first performance optimization indicator acquisition module M1 is used to acquire a first performance optimization indicator for multi-controller deployment according to the SDN network structure; the first performance optimization indicator includes: average propagation delay between the switch and the controller, average propagation delay between controllers, synchronization overhead between controllers, and minimum security of synchronization between controllers;

[0203] A first objective function establishing module M2, used to establish a first objective function according to the first performance optimization index;

[0204] A first constraint condition acquisition module M3 is used to acquire a first constraint condition for multi-controller deployment; the first constraint condition includes a controller load constraint, a mapping relationship constraint between a switch and a controller, and a control layer synchronization link bandwidth constraint;

[0205] A multi-controller deployment model building module M4, used to build a multi-controller deployment model according to the first objective function and the first constraint condition;

[0206] The solving module M5 is used to solve the multi-controller deployment model based on the Markov decision model using a deep reinforcement learning algorithm to obtain a multi-controller deployment solution.

[0207] As for the system disclosed in the embodiment, since it corresponds to the method disclosed in the embodiment, the description is relatively simple, and the relevant parts can be referred to the description of the method part.

[0208] The principles and implementation methods of the present invention are described in this article using specific examples. The description of the above embodiments is only used to help understand the method and core idea of ​​the present invention. At the same time, for those skilled in the art, according to the idea of ​​the present invention, there will be changes in the specific implementation methods and application scope. In summary, the content of this specification should not be understood as limiting the present invention.

Claims

1. A multi-controller deployment method based on SDN network and deep reinforcement learning, characterized in that: include: Obtaining a first performance optimization indicator for multi-controller deployment according to the SDN network structure; The first performance optimization index includes: average propagation delay between the switch and the controller, average propagation delay between controllers, synchronization overhead between controllers, and minimum security of synchronization between controllers; Establishing a first objective function according to the first performance optimization indicator; Obtaining a first constraint condition for multi-controller deployment; the first constraint condition includes a controller load constraint, a mapping relationship constraint between a switch and a controller, and a control layer synchronization link bandwidth constraint; Constructing a multi-controller deployment model according to the first objective function and the first constraint condition; Based on the Markov decision model, a deep reinforcement learning algorithm is used to solve the multi-controller deployment model to obtain a multi-controller deployment solution; in, The expression of the average propagation delay between the switch and the controller is: Where N is the number of switches, d ij is the shortest link delay between the switch and the controller, x ij It is a binary number. When the value is 1, it means that the switch i is successfully connected to the controller j. i is switch i, S is the set of switches, c j is controller j, C is the controller set, The expression of the average propagation delay between the controllers is: c k is the controller k, K is the number of controllers, is the shortest link delay between controllers, The expression of the inter-controller synchronization overhead is: Among them, l jk is the length of the data packet synchronized between controller j and controller k, p sjk is the frequency of synchronization between controller j and controller k; The first objective function is: minimize(αT scavg +βT ccavg +ρC cc )+μ·K Among them, T scavg represents the average propagation delay between the switch and the controller, T ccavg represents the average propagation delay between controllers, C cc represents the synchronization overhead between controllers, μ is the safety factor for achieving the minimum security of synchronization between controllers, α, β, ρ are the T scavg , T ccavg , C cc The weight of , α+β+ρ=1; The first constraint condition is: The controller load constraint is: Among them, L opt is the desired optimal controller load quantity, L c is the actual controller load quantity, L j is the load quantity of controller j, LBI is used to measure the load quantity difference among all controllers in the network, defined as the load deviation index, ΔLB is the difference in the number of switches managed by each controller, defined as the average value of the load difference, LB C is the threshold value of ΔLB, The mapping relationship between the switch and the controller is constrained as follows: Where V represents the set of switches, The control layer synchronization link bandwidth constraint is: Among them, B W The bandwidth provided for the control layer physical link; The method further includes: embedding Atomix nodes based on the multi-controller deployment model, and optimizing the deployment of the Atomix nodes, specifically including: Obtain a second performance optimization indicator for Atomix node deployment, where the second performance optimization indicator includes an average synchronization delay between Atomix nodes and an average synchronization delay between an Atomix node and a controller node. Constructing a second objective function according to the second performance optimization index, A second constraint condition for Atomix node deployment is obtained, where the second constraint condition is that the number of Atomix nodes deployed in the network is B, and there is at least one mapping relationship between the Atomix node and the controller node. Constructing an Atomix node deployment model according to the second objective function and the second constraint condition, Using deep reinforcement learning algorithm to solve the Atomix node deployment model; The average synchronization delay between the Atomix nodes is: Where A is the number of Atomix nodes, is the shortest link delay between Atomix nodes, a j ,a k ,a i For Atomix nodes j, k, i, The average synchronization delay between the Atomix node and the controller node is: is the shortest link delay between the Atomix node and the controller node, z ij It is a binary number. When the value is 1, it indicates the successful connection between Atomix node i and controller j. The second objective function is: minimizeρ1T aaavg +ρ2T acavg , ρ1, ρ2 are respectively the aaavg , T acavg The weight of , ρ1+ρ2=1; There is at least one mapping relationship between the Atomix node and the controller node:

2. The method according to claim 1, characterized in that Solving the multi-controller deployment model using a deep reinforcement learning algorithm based on a Markov decision model specifically includes: Acquire the state space of the Markov decision model according to the controller placement of each node in the network at the current moment and the network topology information; Obtaining the action space of the Markov decision model according to the number of controllers deployed in the network, the node locations where the controllers are deployed, and the mapping relationship between the switches and the controllers; Obtaining the state transition probability of the Markov decision model according to the probability of transferring to the next state after executing a certain action in the current state; Obtaining a reward function of the Markov decision model according to an average communication delay between the switch and the controller, an average communication delay between the controllers, a synchronization overhead between the controllers, and a minimum safety of the controller; A multi-controller deployment scheme is obtained based on the state space, the action space, the state transition probability and the reward function of the Markov decision.

3. A system based on the SDN network and deep reinforcement learning multi-controller deployment method according to claim 1 or 2, characterized in that: include: A first performance optimization indicator acquisition module, used to acquire a first performance optimization indicator for multi-controller deployment according to the SDN network structure; The first performance optimization index includes: average propagation delay between the switch and the controller, average propagation delay between controllers, synchronization overhead between controllers, and minimum security of synchronization between controllers; A first objective function establishing module, used to establish a first objective function according to the first performance optimization index; A first constraint condition acquisition module, used to acquire a first constraint condition for multi-controller deployment; the first constraint condition includes a controller load constraint, a mapping relationship constraint between a switch and a controller, and a control layer synchronization link bandwidth constraint; A multi-controller deployment model building module, used to build a multi-controller deployment model according to the first objective function and the first constraint condition; The solution module is used to solve the multi-controller deployment model based on the Markov decision model using a deep reinforcement learning algorithm to obtain a multi-controller deployment solution.

Citation Information

Patent Citations

  • OpenFlow multi-controller system and management method thereof

    CN104980296A

  • Multi-controller dynamic deployment method of software defined spatial information network

    CN107276662A