SDN data center network intelligent multi-controller flow load balancing method

By adopting reinforcement learning algorithms and multi-agent models in the data center, the problem of traditional SDN controllers dealing with bottlenecks in large-scale data centers is solved, adaptive traffic load balancing and real-time optimization are achieved, and the stability and efficiency of the system are improved.

CN120128541APending Publication Date: 2025-06-10SOUTH CHINA UNIV OF TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510325855.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-19
Publication Date
2025-06-10

AI Technical Summary

Technical Problem

In hyper-large data centers, traditional centralized SDN controllers are prone to surge in control signaling processing delays, lag in network status updates due to processing bottlenecks, and risk of a single point of failure. The existing SDN multi-controller load balancing method has the problems of polarization of resource utilization and high migration decision-making.

Method used

The reinforcement learning algorithm is adopted to pre-learn routing and load balancing strategies through the intelligent multi-controller model, and real-time optimization is performed based on the pre-trained model in the online environment. Combined with the reinforcement learning method of multi-agents, the global optimal path is selected to improve load balancing effect and stability.

Benefits of technology

Adaptive traffic load balancing is realized, and the network environment is optimized in real time according to the service traffic state, which reduces training time, improves the effect and stability of load balancing, and a security mechanism is designed to prevent the emergence of network loops.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120128541A_ABST
    Figure CN120128541A_ABST
Patent Text Reader

Abstract

The invention discloses an SDN (Software Defined Network) data center network intelligent multi-controller flow load balancing method. The method comprises the following steps: S1, establishing an SDN service architecture model according to a data center server architecture; s2, establishing an intelligent multi-controller model according to the SDN service architecture model environment; each SDN controller is instantiated into an intelligent agent model, and the step of establishing the intelligent multi-controller model comprises initializing parameters of the intelligent agent model, a strategy network and an evaluation network; s3, the intelligent agent learns a routing and load balancing strategy in advance, and the intelligent multi-controller model is pre-trained; s4, acquiring network state information from an online environment, and training the intelligent multi-controller model based on the pre-training model; and S5, applying the trained intelligent multi-controller model to the SDN network, and carrying out load balancing. The distributed SDN data center network flow characteristics are learned through the reinforcement learning algorithm, flow load balancing can be carried out in a self-adaptive mode, and the network environment is optimized in real time according to the service flow state.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of computer networks, and particularly relates to an intelligent multi - controller traffic load - balancing method for SDN data center networks. Background Art

[0002] Software - Defined Networking (SDN) logically decouples the data forwarding function (data plane) and the routing decision - making function (control plane) of network devices, thereby constructing a centralized network intelligent center. Its core value lies in the visualization of the global topology of the control layer and the programmable traffic management ability.

[0003] In the application scenario of ultra - large - scale data centers, network traffic exhibits typical spatio - temporal heterogeneity characteristics: there are sudden traffic spikes in the time dimension and non - uniform distribution characteristics in the space dimension. Such multi - dimensional dynamic characteristics make traditional centralized SDN controllers prone to processing bottlenecks, specifically manifested as: a sharp increase in control signaling processing delay, a lag in network state updates, and the potential risk of a single point of failure. In response to the above challenges, the distributed SDN control architecture proposed jointly by the academic community and the industrial community realizes the horizontal expansion ability of the control plane.

[0004] In the distributed SDN architecture, network performance optimization essentially depends on the effectiveness of the multi - controller load - balancing mechanism, and its strategy design is directly related to the overall system efficiency. Existing SDN multi - controller load - balancing methods have significant defects: when local network nodes are hit by sudden traffic surges, traditional static allocation strategies will cause asymmetric distribution of data loads between controller sub - domains, resulting in polarization of controller resource utilization. More seriously, existing iterative optimization methods relying on heuristic algorithms need to repeatedly reconstruct the mapping relationship between the control plane and the data plane, resulting in high time consumption for a single migration decision. Frequent migration operations not only cause flow table configuration errors but may also lead to key service flow interruption time exceeding the threshold. Summary of the Invention

[0005] The main purpose of the present invention is to overcome the defects and deficiencies of the prior art and propose an intelligent multi - controller traffic load - balancing method for SDN data center networks.

[0006] To achieve the above purpose, the present invention adopts the following technical solutions:

[0007] An intelligent multi - controller traffic load - balancing method for SDN data center networks, comprising:

[0008] S1. Establish an SDN service architecture model based on the data center server architecture, including an SDN controller, an SDN switch, an SDN host, a network link discovery module, and a network status monitoring module;

[0009] S2. Establish an intelligent multi - controller model according to the SDN service architecture model environment; each SDN controller is instantiated as an agent model, and establishing the intelligent multi - controller model includes initializing the parameters of the agent model, the policy network, and the evaluation network;

[0010] S3. The agent pre - learns routing and load - balancing policies to pre - train the intelligent multi - controller model;

[0011] S4. Obtain network status information from the online environment and train the intelligent multi - controller model based on the pre - trained model;

[0012] S5. Apply the trained intelligent multi - controller model to the SDN network for load balancing. Compared with the prior art, the present invention has the following advantages and beneficial effects:

[0013] 1. The present invention learns the network traffic characteristics of the distributed SDN data center through the reinforcement learning algorithm, can perform traffic load balancing adaptively, and optimize the network environment in real - time according to the business traffic status.

[0014] 2. Aiming at the problems of complex network traffic environment and long training convergence time, the present invention proposes a method for pre - training a reinforcement learning model based on the shortest - path algorithm, which can accelerate the online training process and reduce the training time.

[0015] 3. Based on the reinforcement learning method of multi - agents, considering the local state of each agent and the overall state of the network environment comprehensively, the global optimal path is selected to improve the effect and stability of load balancing.

[0016] 4. Aiming at the complex environment of the SDN data center network, a security mechanism is designed, which can prevent the appearance of network loops during training and increase the stability of the network environment. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] Figure 1 is the flowchart of the method of the present invention;

[0018] Figure 2 is the architecture diagram of the load - balancing optimization method of the present invention;

[0019] Figure 3 is the schematic diagram of the security learning mechanism of the present invention. DETAILED DESCRIPTION OF THE INVENTION

[0020] The present invention will be further described in detail below with reference to the embodiments and the accompanying drawings, but the embodiments of the present invention are not limited thereto.

[0021] Embodiment

[0022] As Figure 1 and Figure 2 shown, the present invention, an SDN data center network intelligent multi - controller traffic load balancing method, includes:

[0023] S1. Establish an SDN service architecture model according to the data center server architecture, including an SDN controller, an SDN switch, an SDN host, a network link discovery module, and a network status monitoring module;

[0024] In this embodiment, the SDN service architecture model is specifically:

[0025] Establish a graph structure G = {V, E} according to the data center server architecture, where there are M SDN switches sw, and the M switches together form a set V, that is, sw i ∈V, i ∈ [1, M]; the N edges e formed by the connection of the switches form a set E, where e i ∈E, i ∈ [1, N];

[0026] Deploy a reinforcement learning agent for the SDN switches directly controlled by the SDN controller, and the edge - layer switches are directly connected to the SDN hosts; the network link discovery module establishes a network global topology by sending traffic detection packets based on the Link Layer Discovery Protocol (LLDP); the network status monitoring module collects network raw information in real - time and calculates information such as link average delay, link available bandwidth, and packet loss rate;

[0027] In order to accelerate the subsequent network training process, normalize the link average delay, link available bandwidth, and packet loss rate; the overall goal of load balancing is to make the value of the following function as large as possible:

[0028]

[0029] where, U t is the utility function, α t 、β t 、ψ t respectively represent the normalized link average delay, link available bandwidth, and packet loss rate at time t, ω 1 、ω 2 、ω 3 respectively represent the preferences for network optimization, which are divided into four cases: low - delay preference, high - bandwidth preference, low - delay and high - bandwidth preference, and low - delay and low - packet - loss preference, ω 1 、ω 2 、ω 3They are represented as [0 1 0], [1 0 0], [0.5 0.5 0] and [0 0.5 0.5] respectively.

[0030] S2. Establish an intelligent multi - controller model according to the SDN service architecture model environment; each SDN controller is instantiated as an agent model, and establishing the intelligent multi - controller model includes initializing the parameters of the agent model, the policy network, and the evaluation network.

[0031] S21. Initialize the agent model, and set the initialization of a quadruple composed of the current environment state, action, next - moment environment state, and reward. Initialize the experience replay buffer.

[0032] S22. Initialize the parameters of the policy network and the evaluation network, and repeat steps S21 and S22 until all the multi - controller agents in all networks are initialized.

[0033] In this embodiment, the way the intelligent multi - controller model generates the optimal routing path follows the MAMDP process (Multi - Agent Markov Decision Process); the MAMDP process is an extended model of the classic MDP (Markov Decision Process) in a multi - agent scenario, and is formally represented as a tuple where n ∈ [1, M] is the number of agents in the model; is the global environment input of the i - th agent in n; A i is the action decision made by a single agent, and the joint action space A = Π i∈n A i ; is the state - transition policy of MAMDP; represents the global reward function; γ is the discount factor.

[0034] S3. The agents pre - learn the routing and load - balancing policies to pre - train the intelligent multi - controller model; specifically including:

[0035] S31. Obtain information including network - link load rate, link delay, and packet - loss rate from each module of the SDN service architecture model.

[0036] S32. Train the policy network and the evaluation network of the agents based on the shortest - routing policy, provide prior knowledge for the agents to avoid the agents making unsafe decisions (such as routing loops, over - utilization of links, etc.) at the initial stage, and store the calculated shortest paths of nodes in the path table for pre - training use.

[0037] S33. Select the action a based on the node shortest - path table t, according to this action a t Execute the interaction with the SDN network environment to obtain the reward r t , the environmental state s of this round of interaction t , action a t , reward r t and the environmental state s at the next moment t+1 to form the experience {s t , a t , r t , s t+1} and store it in the experience replay buffer

[0038] S34. Take a batch of experiences from the experience replay buffer to train the global policy of the intelligent multi - controller model and update the model parameters of each agent. Repeat steps S31 to S34 until the number of iteration rounds reaches the pre - training set upper limit to obtain the pre - trained intelligent multi - controller model.

[0039] Facing the environmental state Only the nodes where the source switch sw src and the destination switch sw dst form a path need to make a decision a ∈ A. Assume that the path is {sw src , a 1 ,..., a n-1 , sw dst}. The decision - making process of the agent is represented by the following formula:

[0040] P(a|s) = P(a 1 |s)P(a 2 |s, a 1 )P(a 3 |s, a 1 , a 2 )...P(a n |s, a 1 , a 2 ...a n-1 )

[0041] Build a single - agent model with a policy network Actor and a critic network Critic; the Actor network of the i - th agent is Use to estimate the agent's decision - making process P(a n |s, a 1 , a 2 ...a n-1 ), where θ i represents the neural network parameters of the policy network Actor, and c i is the conditional state based on the global state s.

[0042] S4. Obtain network status information from the online environment and train the intelligent multi - controller model based on the pre - trained model, including:

[0043] S41. Obtain information including network link load rate, link delay, and packet loss rate from each module of the SDN service architecture model;

[0044] S42. Random exploration: Randomly sample an action target from the action space following a probability distribution, calculate the advantage function, and evaluate the quality of the selected action;

[0045] S43. Optimize the policy function, and update the old policy function with the policy function updated using the clip method;

[0046] S44. Use the method of gradient ascent to update the policy, update the parameters of the policy network and the evaluation network, and repeat steps S41 to S44 until the policy performance no longer improves or reaches the pre - set number of training steps.

[0047] During step S4, the proximal policy optimization (PPO) method is used to update the network parameters during training:

[0048]

[0049] Among them When the i - th agent is on the path, m i = 1, otherwise 0; τ represents the environment - action pair (s, a), and the corresponding action pair is matched by sampling A(s, a) represents the advantage value estimation, which is used to judge the advantage of choosing action a based on the network policy Π θ compared with a random action in the environment state s; During the training process, if there is a loop or the time threshold is exceeded, the policy takes punitive measures, and at the same time, a path is generated for the service flow based on the shortest - path method, as Figure 3 shown.

[0050] S5. Apply the trained intelligent multi - controller model to the SDN network for load balancing, including:

[0051] S51. Obtain information including network link load rate, link delay specification, and packet loss rate from each module of the SDN service architecture model;

[0052] S52. The intelligent multi - controller model calculates and selects action a according to the input network status information, updates the local policy based on the clipping function and the minimization function, and uses J clip (θ i ) to represent the local objective function of the agent:

[0053]

[0054] where r′(θ i ) = clip(rθ i ), 1 - ε, 1 + ε), ε is a clipping hyperparameter, and 0.2 is adopted in this embodiment;

[0055] Obtain the global optimal path p, and complete the load balancing decision of the SDN environment by issuing flow tables.

[0056] It should also be noted that in this specification, terms such as "including", "comprising" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements, but also includes other elements not explicitly listed, or also includes elements inherent to such a process, method, article or device. Without further limitation, an element defined by the statement "including one..." does not exclude the existence of another identical element in the process, method, article or device including the said element.

[0057] The above description of the disclosed embodiments enables those skilled in the art to implement or use the present invention. Various modifications to these embodiments will be obvious to those skilled in the art, and the general principles defined herein can be implemented in other embodiments without departing from the spirit or scope of the present invention. Therefore, the present invention will not be limited to these embodiments shown herein, but rather to the widest scope consistent with the principles and novel features disclosed herein.

Claims

1. A method for intelligent multi-controller traffic load balancing in an SDN data center network, characterized in that: include: S1. Establish an SDN service architecture model based on the data center server architecture, including an SDN controller, an SDN switch, an SDN host, a network link discovery module, and a network status monitoring module; S2. Establish an intelligent multi-controller model according to the SDN service architecture model environment; each SDN controller is instantiated as an intelligent agent model, and establishing the intelligent multi-controller model includes initializing the parameters of the intelligent agent model, the policy network, and the evaluation network; S3, the agent pre-learns routing and load balancing strategies and pre-trains the intelligent multi-controller model; S4, obtaining network status information from the online environment and training the intelligent multi-controller model based on the pre-trained model; S5. Apply the trained intelligent multi-controller model to the SDN network for load balancing.

2. According to claim 1, a SDN data center network intelligent multi-controller traffic load balancing method is characterized in that: In step S1, the SDN service architecture model is specifically: According to the data center server architecture, a graph structure G = {V, E} is established, in which there are M SDN switches sw, and the M switches together form a set V, that is, sw i ∈V,i∈[1,M]; the N edges e formed by the switch connections form a set E, where e i ∈E,i∈[1,N]; The SDN switches directly controlled by the SDN controller deploy reinforcement learning agents, and the edge layer switches are directly connected to the SDN hosts. The network link discovery module establishes the global network topology by sending traffic detection packets based on the link layer discovery protocol LLDP. The network status monitoring module collects network raw information in real time and calculates the average link delay, link available bandwidth and packet loss rate; In order to speed up the subsequent network training process, the average link delay, link available bandwidth and packet loss rate are normalized; the overall goal of load balancing is to make the value of the following function as large as possible: Among them, U t is the utility function, α t , β t , t represent the normalized average link delay, available link bandwidth and packet loss rate at time t respectively. ω1, ω2 and ω3 represent the preferences for network optimization, which are divided into four cases: low delay preference, high bandwidth preference, low delay and high bandwidth preference and low delay and low packet loss rate preference. ω1, ω2 and ω3 are expressed as [0 1 0], [1 0 0], [0.5 0.5 0] and [0 0.5 0.5] respectively.

3. According to claim 1, a SDN data center network intelligent multi-controller traffic load balancing method is characterized in that: Step S2 includes: S21. Initialize the agent model and set the initialization based on the current environment state, action, next moment environment state and reward. Initialize the experience replay buffer S22, initialize the parameters of the policy network and the evaluation network, and repeat steps S21 and S22 until the multi-controller agents in all networks are initialized.

4. According to claim 1, a SDN data center network intelligent multi-controller traffic load balancing method is characterized in that: The way the intelligent multi-controller model generates the optimal routing path follows the MAMDP process; the MAMDP process is an extended model of the classic MDP in the multi-agent scenario, formally represented as a tuple Where n∈[1,M] is the number of agents in the model; A is the global environment input of agent i∈n; i The action decision made by a single agent, the joint action space A = Π i∈n A i ; is the state transition strategy of MAMDP; represents the global reward function; γ is the discount factor.

5. According to claim 3, a method for balancing traffic load of SDN data center network intelligent multi-controllers is characterized in that: Step S3 specifically includes: S31, obtaining information including network link load rate, link delay and packet loss rate from each module of the SDN service architecture model; S32, training the agent's strategy network and evaluation network based on the shortest routing strategy, providing the agent with prior knowledge, and storing the calculated node shortest path in the path table for subsequent training; S33, select action a based on the node shortest path table t , according to this action a t Execute interaction with the SDN network environment and get reward r t , the environment state s of this round of interaction t 、Action a t , Reward t and the next moment environment state s t+1 Composition of Experience t 、a t 、r t 、s t+1 }Store in experience replay buffer S34. Replaying from the Experience Buffer Take a batch of experience, train the global strategy of the intelligent multi-controller model, update the model parameters of each agent, repeat steps S31 to S34 until the iteration round reaches the pre-training setting upper limit, and obtain the pre-trained intelligent multi-controller model.

6. A method for balancing traffic load of SDN data center network intelligent multi-controllers according to claim 5, characterized in that: Facing the environmental status Only on the source switch sw src and the destination switch sw dst The nodes in the path need to make a decision a∈A. Assume that the path {sw src ,a1,...,a n-1 ,sw dst }, the agent decision process is expressed using the following formula: P(a|s)=P(a1|s)P(a2|s,a1)P(a3|s,a1,a2)...P(a n |s,a1,a2...a n-1 ) Establish a single agent model with a strategy network Actor and an evaluation network Critic; the Actor network of the i-th agent is use For the agent decision process P(a n |s,a1,a2...a n-1 ) is estimated, where θ i Represents the neural network parameters of the policy network Actor, c i is a conditional state based on the global state s.

7. A method for balancing traffic load of SDN data center network intelligent multi-controllers according to claim 6, characterized in that: Step S4 includes: S41, obtaining information including network link load rate, link delay and packet loss rate from each module of the SDN service architecture model; S42, random exploration, randomly sampling an action target from the action space that follows the probability distribution, calculating the advantage function, and evaluating the quality of the selected action; S43, optimizing the policy function, and using the policy function updated in the clip mode to update the old policy function; S44. Use the gradient ascent method to update the strategy, update the parameters of the strategy network and the evaluation network, and repeat steps S41 to S44 until the strategy performance no longer improves or reaches the preset number of training steps.

8. The method for balancing traffic load of an SDN data center network intelligent multi-controller according to claim 7, characterized in that: In step S4, the training adopts the PPO proximal strategy optimization method to update the network parameters, which is expressed as: in, When the i-th agent is on the path, m i =1, otherwise 0; τ represents the environment and action pair (s, a), through sampling Match the corresponding action pairs; A(s,a) represents the advantage value estimate, which is used to judge the network strategy Π in the environment state s θ The advantage of selecting action a over random action; if there is a loop or the time threshold is exceeded during training, the strategy takes punitive measures and generates a path for the business flow based on the shortest path method.

9. The method for balancing traffic load of an SDN data center network intelligent multi-controller according to claim 1, characterized in that: Step S5 includes: S51, obtaining information including network link load rate, link delay specification and packet loss rate from each module of the SDN service architecture model; S52, the intelligent multi-controller model calculates and selects action a according to the input network status information, updates the local strategy based on the clipping function and the minimization function, and uses H clip (θ i ) represents the local objective function of the agent: Among them, r′(θ i )=clip(r(θ i ), 1-ε, 1+ε), ε is the pruning hyperparameter; Obtain the global optimal path p and complete the load balancing decision of the SDN environment by issuing the flow table.