Optimization system and method for SDN (Software Defined Network) traffic engineering
By combining graph neural networks and an improved A2C architecture, the scalability and real-time responsiveness issues of SDN traffic engineering in large-scale network environments are solved, efficient traffic distribution and fault tolerance are achieved, and network performance and throughput are improved.
Patent Information
- Application Number
- CN202510565731.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-30
- Publication Date
- 2025-09-12
AI Technical Summary
Existing SDN traffic engineering faces problems such as insufficient scalability, real-time responsiveness, and global optimal solution calculation accuracy in large-scale network environments. It is difficult to effectively respond to link failures and dynamic traffic demands. Traditional methods have problems such as high computational complexity, delayed routing updates, and traffic black holes.
Combining graph neural networks (GNN) and an improved A2C architecture, the control plane and data execution plane are constructed through the monitoring module, traffic engineering controller, and flow rule translation module. The feature extraction model of GNN and DNN and deep reinforcement learning are used to optimize traffic distribution. Combined with ADMM processing constraint optimization, efficient traffic distribution is achieved.
It achieves efficient traffic distribution in large-scale WAN topologies, significantly improves the speed and accuracy of SDN traffic engineering decisions, and can complete optimization within a five-minute scheduling window, improving network throughput and link fault tolerance.
Smart Images

Figure CN120639696A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of network data processing, and in particular to an optimization system and method for SDN traffic engineering. Background Art
[0002] With the rapid development of global digital infrastructure, the Wide Area Network (WAN) has become a key backbone network connecting distributed data centers and serving global users. Although these large-scale cross-regional networks have achieved seamless global communication, their expensive construction and operation and maintenance costs make efficient traffic management the primary challenge facing network service providers. Traffic Engineering (TE), as a core technology for network resource optimization, can not only significantly improve network resource utilization, but also enhance the network's resilience to failures through intelligent traffic path planning and dynamic scheduling. In the increasingly complex cloud computing environment and the ever-increasing demand for data transmission, efficient traffic engineering solutions have become an indispensable strategic component in modern WAN operations, directly affecting service quality and operating costs.
[0003] Software Defined Network (SDN), with its core feature of separating the control plane from the data plane, provides network operators with a centralized network view and refined control capabilities. Compared with traditional distributed network architectures, SDN can monitor the global network status in real time and quickly respond to changes in network conditions through a centralized controller, making it an ideal platform for efficient traffic engineering. By fully leveraging the programmability of SDN, network operators can achieve dynamic scheduling and optimized allocation of traffic, effectively respond to fluctuating business demands and quickly alleviate network congestion caused by link failures. However, with the expansion of network scale and the increase in business complexity, current SDN-based traffic engineering solutions still face huge challenges in terms of scalability, real-time responsiveness and the calculation accuracy of the global optimal solution in large-scale network environments, and urgently need innovative technological breakthroughs.
[0004] Over the past decade, with the exponential expansion of networks and the dynamic evolution of traffic patterns, traffic engineering has faced unprecedentedly complex challenges. Traditional traffic engineering relies on centralized optimization solvers based on linear programming to achieve dynamic resource allocation. However, with the dramatic growth of WAN topologies and the dramatic fluctuations in traffic demand, these approaches have increasingly experienced performance and adaptability bottlenecks. The computational complexity of these systems grows superlinearly with network size, leading to configuration delays and potentially causing traffic black holes and forwarding table capacity pressure. Current mainstream traffic engineering systems pursue global optimal solutions by solving multi-commodity flow models, but face three key challenges: First, the computational complexity of solving multi-commodity flow models is difficult to meet real-time requirements, especially in emergency situations such as link failures, where delayed routing updates may cause network performance to deteriorate sharply. Second, existing acceleration schemes such as SMORE's regional partitioning and NCFlow's cluster decoupling can improve computing speed, but they sacrifice the quality of the solution by relaxing optimization constraints, which may lead to link overload or produce solutions that cannot be deployed in practice. Third, traditional traffic prediction methods rely on host agents or statistical models. Their error functions are inherently mismatched with the actual optimization goals of traffic engineering, making it difficult to effectively cope with the high uncertainty of user-driven traffic, limiting the adaptability and robustness of the overall system.
[0005] In recent years, machine learning (ML) and deep reinforcement learning (DRL) have demonstrated revolutionary potential in the field of traffic engineering. By adaptively learning complex network traffic patterns and topological features, they have significantly improved the perception and adaptability of traffic engineering systems to dynamic network environments. However, the actual application of these advanced technologies in production environments still faces severe challenges: traditional DRL architectures rely on long-term cumulative reward signals for policy optimization, which not only has slow convergence speed and is difficult to meet real-time decision-making requirements, but is also prone to falling into local optimal solutions. At the same time, existing methods perform poorly in handling network hard constraints (such as link capacity limits) and may produce unfeasible solutions. In addition, while supervised learning-based models can imitate optimal solutions, their generalization capabilities are limited, making it difficult to effectively cope with dynamic topological changes (such as unexpected link failures) or unprecedented traffic patterns, restricting their large-scale deployment in high-reliability network environments. Summary of the Invention
[0006] In response to the technical problems of low real-time responsiveness and optimization accuracy of existing SDN traffic engineering and difficulty in coping with link failures, the present invention discloses an optimization system and method for SDN traffic engineering, which combines graph neural networks (GNN) and an improved A2C (Actor-Critic) architecture. It can achieve efficient traffic distribution in large-scale WAN topologies while meeting time and space requirements, thereby significantly improving the speed and accuracy of SDN traffic engineering decisions.
[0007] In order to achieve the above object, the present invention provides the following technical solutions:
[0008] An optimization system for SDN traffic engineering includes a control plane and a data execution plane, wherein the control plane and the data execution plane are bidirectionally connected;
[0009] The control plane includes a monitoring module, a traffic engineering controller and a flow rule translation module;
[0010] Among them, the monitoring module is used to collect the traffic matrix of the network;
[0011] Traffic engineering controller, used to generate traffic distribution strategy based on the collected traffic matrix;
[0012] The flow rule translation module is used to convert the traffic distribution policy into flow table rules that can be executed by the data execution plane;
[0013] The data execution plane includes at least one device that actually carries network traffic, which is used to execute traffic distribution strategies according to received flow table rules, adjust traffic forwarding and distribution, and feed back real-time network status to the monitoring module of the control plane.
[0014] The present invention also provides an optimization method for SDN traffic engineering, which specifically includes the following steps:
[0015] S1: Obtain the traffic matrix in the network and represent the network topology as a graph structure;
[0016] S2: Input the traffic matrix and graph structure into the constructed feature extraction model and output the feature vector;
[0017] S3: Input the feature vector into the deep reinforcement learning model for iterative optimization, and output the initial traffic allocation ratio strategy;
[0018] S4: Perform constraint optimization on the initial traffic allocation ratio strategy and output the final traffic allocation strategy.
[0019] In summary, due to the adoption of the above technical solution, compared with the prior art, the present invention has at least the following beneficial effects:
[0020] This paper combines GNN with an improved A2C architecture, modeling the WAN as a bipartite graph between paths and links. Graph Isomorphism Network Convolution (GINConv) is used to achieve efficient information exchange between links and paths, and Multilayer Perceptrons (MLP) are employed to handle the collaboration and competition between multiple paths under the same traffic demand. Therefore, the model considers both the network topology and the real-time traffic demand characteristics at the embedding layer.
[0021] On this basis, the present invention adopts an improved A2C method to output the initial traffic allocation strategy, and combines ADMM to fine-tune any potential overcapacity or traffic imbalance problems in the post-processing stage, thereby achieving global optimal traffic allocation.
[0022] Experimental results on multiple real-world network topologies (e.g., B4, UsCarrier, KDL, and ASN) demonstrate that the proposed Graph-Based Reinforcement Learning for Traffic Engineering (GRL-TE) solution offers significant advantages in improving network throughput and tolerant to link faults, with all optimizations completed within a five-minute scheduling window. These results demonstrate its feasibility, efficiency, and scalability in real-world production environments, providing a new technical path for next-generation SDN traffic engineering. BRIEF DESCRIPTION OF THE DRAWINGS
[0023] Figure 1 Schematic diagram of an optimization system for SDN traffic engineering according to an exemplary embodiment of the present invention.
[0024] Figure 2 Schematic diagram of an optimization method for SDN traffic engineering according to an exemplary embodiment of the present invention.
[0025] Figure 3 Schematic diagram of a bipartite graph structure according to an exemplary embodiment of the present invention.
[0026] Figure 4 FIG. 1 is a schematic diagram of experimental results of average satisfaction of requirements according to an exemplary embodiment of the present invention.
[0027] Figure 5 FIG. 1 is a schematic diagram of a link failure on B4 according to an exemplary embodiment of the present invention.
[0028] Figure 6 Schematic diagram of a link failure on UsCarrier according to an exemplary embodiment of the present invention.
[0029] Figure 7 FIG. 1 is a schematic diagram of a link failure on Kd1 according to an exemplary embodiment of the present invention.
[0030] Figure 8 FIG. 1 is a schematic diagram of a link failure on an ASN according to an exemplary embodiment of the present invention.
[0031] Figure 9 FIG. 4 is a schematic diagram of average calculation time according to an exemplary embodiment of the present invention.
[0032] Figure 10 Schematic diagram of a feature extraction model according to an exemplary embodiment of the present invention. DETAILED DESCRIPTION
[0033] The present invention will be further described in detail below with reference to the examples and specific implementation methods. However, this should not be understood as limiting the scope of the present invention to the following examples, as all technologies implemented based on the present invention fall within the scope of the present invention.
[0034] like Figure 1 As shown, the present invention provides an optimization system for SDN traffic engineering, including a control plane and a data execution plane, and the control plane and the data execution plane are bidirectionally connected.
[0035] In this embodiment, the control plane includes a monitoring module, a traffic engineering controller and a flow rule translation module; the input end of the monitoring module is connected to the output end of the data execution plane, the output end of the monitoring module is connected to the input end of the traffic engineering controller, the output end of the traffic engineering controller is connected to the input end of the flow rule translation module, and the output end of the flow rule translation module is connected to the input end of the data execution plane.
[0036] The monitoring module is responsible for collecting the traffic matrix (TM) of the global network. This traffic matrix represents the traffic demand between source and destination nodes and is generated by a traffic measurement or prediction platform deployed in the network. It reflects the real-time traffic demand in the network and serves as the basic data for the traffic engineering controller to perform routing optimization.
[0037] The traffic engineering controller runs an optimization method for SDN traffic engineering. By combining deep reinforcement learning (DRL) and graph neural networks (GNNs), it generates traffic distribution policies based on collected information such as traffic matrices, network topology, and link status. The combination of DRL and GNNs gives the control plane powerful feature representation and generalization capabilities, enabling it to perform well in complex network environments.
[0038] The flow rule translation module is used to convert the traffic distribution policy into flow table rules that can be executed by the data execution plane and send them to each network device of the data execution plane.
[0039] The data execution plane includes at least one device that actually carries network traffic. It executes traffic distribution policies based on received flow table rules, dynamically adjusting traffic forwarding paths and distribution ratios to meet optimized network performance requirements. It also provides real-time network status feedback to the control plane's monitoring module, serving as a basis for subsequent optimization. Due to the separation of control and data in the SDN architecture, routing optimization and policy distribution in the control plane can be performed in parallel, improving system processing efficiency and fault response speed.
[0040] Based on the above, an optimization system for SDN traffic engineering, such as Figure 2 As shown, the present invention also provides an optimization method for SDN traffic engineering, referred to as GRL-TE, which specifically includes the following steps:
[0041] S1: Obtain the traffic matrix in the network (e.g., wide area network) and represent the network topology as a graph structure.
[0042] The capacity of a wide area network is inherently limited by physical factors such as link bandwidth and node processing power. These limitations require the use of traffic engineering (TE) techniques to efficiently allocate resources, prevent link congestion, and minimize resource waste.
[0043] In this embodiment, in this context, the WAN topology is typically represented as a graph structure G = (V, E, c), where V represents the set of network nodes responsible for forwarding data packets, E represents the set of physical links connecting network nodes and supporting data transmission, and c represents the maximum data transmission rate supported by each link, i.e., the capacity allocated to each link. Let n represent the total number of nodes in the network.
[0044] Based on the graph structure G, a traffic matrix D is introduced to represent the set of all traffic demands, where d∈D represents the traffic demand from a source node to a destination node. These traffic demands must be transmitted within a fixed time window, such as 5 or 10 minutes. The monitoring module continuously tracks historical traffic data and network status, periodically evaluating the traffic demand between each pair of nodes for future time periods. To simplify the optimization process, the predicted traffic demands are typically treated as fixed values and provided to the traffic engineering method in short time intervals.
[0045] The traffic matrix D is independent of the topology. This separate representation aligns with real-world network operations: the network topology G = (V, E, c) is typically relatively stable, while traffic demand D changes dynamically over time, requiring continuous monitoring and updating. This modeling approach enables the system to handle diverse traffic demands across different time windows within a fixed topology.
[0046] In this embodiment, each traffic demand d passes through a set of predefined paths P d For transmission, these paths are generated by a shortest path algorithm based on the network topology before TE optimization. This method of pre-defining paths is called path formulation. In practical WAN deployments, path formulation is widely adopted because it only requires optimization on a limited set of pre-defined paths, without having to consider all possible link combinations in the network. This avoids the computational complexity associated with arbitrary traffic distribution across all possible links. Furthermore, because path formulation involves fewer path entries, it simplifies path management in switches and reduces hardware complexity, thereby facilitating easier network deployment.
[0047] The main goal of traffic distribution is to distribute traffic demand to predefined paths. Specifically, F d (p) represents the proportion of traffic demand d allocated to path p, ranging from 0 to 1. The sum of the proportions of traffic demand d allocated to all paths cannot exceed 1. i ) represents the traffic distribution ratio in the i-th period.
[0048] In this embodiment, traffic distribution must comply with two key constraints. The first is the demand constraint: for any traffic demand d, the total allocation ratio on all paths p must satisfy This ensures that the allocated traffic does not exceed the actual demand. The second is the capacity constraint: for any link e∈E in the network, the total traffic on the link must satisfy Where c(e) represents the capacity of link e, ensuring that the load on any link does not exceed its capacity. By enforcing these two constraints, traffic engineering can optimize traffic distribution along the path while ensuring network stability and performance, thereby maximizing the utilization of network resources.
[0049] Traffic engineering involves multiple objectives, including maximizing network throughput to allow more traffic to pass through the network, minimizing data transmission time to reduce latency, and alleviating network congestion by maximizing minimum link utilization. This application chooses maximizing network throughput as the optimization objective. The core goal is to optimize the distribution of traffic demand along paths, ensuring that as much traffic as possible passes through the network while satisfying all constraints. This approach has been widely adopted in production traffic engineering systems.
[0050] To maximize network throughput, the traffic engineering model includes the following three key constraints: demand constraint, capacity constraint, and non-negativity constraint, as shown in formulas (1), (2), and (3), respectively. The non-negativity constraint ensures that the traffic allocation ratio is non-negative, thereby maintaining the rationality and physical feasibility of the allocation.
[0051]
[0052] In formulas (1), (2), and (3), P d represents the predefined path of traffic demand d; F d (p) represents the proportion of traffic demand d allocated to path p, ranging from 0 to 1; c(e) represents the capacity of link e, e∈E, E represents the set of physical links connecting network nodes and supporting data transmission; D is the set of all traffic demands.
[0053] By combining these goals and constraints, traffic engineering methods can optimize the distribution of traffic along paths, ensuring network stability and performance while maximizing network throughput:
[0054]
[0055] In formula (4), η represents the network throughput; maximize represents the maximization function.
[0056] S2: Input the traffic matrix and graph structure into the constructed feature extraction model and output the feature vector.
[0057] S2-1: Model the WAN topology as a bipartite graph, such as Figure 3 As shown, the nodes are divided into two disjoint subsets: edge nodes and path nodes. Edge nodes represent links (link V between node 1 and node 2). 12 ), and the path node represents the path that the traffic demand passes through (for example, the path P passing through nodes 1, 2, and 4 124 ). In this bipartite graph, edge nodes and path nodes are connected only when the link belongs to the path, e.g., path P 124 Including Link V 12 and V 24 , but does not include V 34 , so P 124 Connect V 12 and V 24 , but do not connect V 34 .
[0058] In the GNN layer, the embeddings of edge nodes and path nodes are initialized based on their respective attributes. The embeddings of edge nodes are initialized using the capacity of the corresponding link (e.g., V 12The embedding of represents the link capacity connecting node 1 and node 2), while the embedding of the path nodes is initialized using the traffic demand (e.g., P 124 The embedding of represents the flow demand from node 1 to node 4). This initialization captures the initial capacity and demand information, laying the foundation for subsequent information exchange.
[0059] S2-2: The core challenge of traffic engineering lies in optimizing traffic distribution to improve network performance and satisfy various network constraints. This process involves handling complex network topologies, dynamic traffic demands, and global resource scheduling. Graph neural networks are a powerful tool for addressing these challenges because they can directly model complex relationships in network topologies and capture both global and local information through efficient message passing mechanisms. By incorporating dynamic traffic demands and network constraints, GNNs can generate efficient and real-time traffic distribution strategies, making them a promising candidate for solving wide area network traffic optimization problems.
[0060] However, traditional graph neural network designs primarily focus on nodes and edges in the graph, which is insufficient in the context of TE. TE needs to deal with contention between flows, where multiple flows compete for limited capacity on the same link. This makes traditional graph neural network designs unable to meet the specific needs of TE.
[0061] The GNN layer of this application adopts the GINConv module, which promotes the flow of information between edge nodes and path nodes through a message passing mechanism. For example, when multiple paths share the same link, the embedding of the edge node corresponding to the link will be affected by the adjacent path nodes, thereby reflecting the bottleneck state of the link. Similarly, the embedding of the path node is updated based on the state of the edge node, capturing the load of the path in the current traffic distribution. However, the GNN layer only allows path nodes to exchange information with adjacent edge nodes, and cannot directly enable multiple path nodes related to the same traffic demand to interact, such as P 134 、P 124 、P 1234 and P 1324 These are all traffic demands from vertex 1 to vertex 4.
[0062] In order to solve the above limitations, this application constructs a feature extraction model by integrating the advantages of graph neural networks (GNN) and deep neural networks (DNN).
[0063] In this embodiment, Figure 10As shown in the figure, the feature extraction model includes at least one feature extraction unit (each feature extraction unit is connected in sequence). The feature extraction unit includes a GNN layer and a DNN layer, and the output of the GNN layer is connected to the input of the DNN layer. To enhance the learning ability of the feature extraction model, the feature extraction unit also introduces a residual connection mechanism: the output of the GNN layer is directly connected to the output of the DNN layer, and the input of a single feature extraction unit is also directly connected to the output of the DNN layer.
[0064] In the feature extraction model, the GINConv module of the GNN layer is used to model the interactions between links to ensure that the allocated traffic does not exceed the link capacity, thereby satisfying the capacity constraint. The GINConv module is based on the Weisfeiler-Lehman graph isomorphism test, enabling it to capture structural information and learn features for edge nodes and path nodes; within each GNN layer, edge nodes and path nodes exchange information and update their embeddings, thereby achieving dynamic perception of traffic bottlenecks.
[0065] On the other hand, the MLP module in the introduced DNN is used to process the embeddings of all path nodes related to the same demand, transform and update them to ensure reasonable traffic distribution and meet demand constraints.
[0066] For example, for a demand from node 1 to node 4, the path node P 134 、P 124 、P 1234 and P 1324 The embeddings of are processed through the MLP to generate new embeddings. These updated embeddings capture the collaboration information between path nodes and are stored back to the corresponding path nodes.
[0067] The feature extraction model of this application effectively combines the advantages of GNN and DNN by alternating between GINConv and MLP modules, leveraging their complementary strengths in feature extraction and constraint modeling. The GINConv module effectively captures the complex relationships between links and paths through a message passing mechanism, while the MLP module models the collaborative relationships between path nodes related to the same traffic demand. Ultimately, the output feature vector encodes the network's topological structure and dynamic traffic characteristics, providing key support for subsequent traffic allocation tasks and enhancing the model's ability to achieve global optimization under various network constraints.
[0068] S3: Input the feature vector into the deep reinforcement learning model for iterative optimization and output the initial traffic distribution ratio strategy.
[0069] In this example, a deep reinforcement learning model uses an A2C architecture for policy optimization: the actor generates a traffic allocation strategy based on feature vectors, while the critic evaluates the quality of the current traffic allocation strategy by calculating a value function and guiding the actor to continuously optimize. This iterative optimization mechanism gradually improves the efficiency and performance of traffic allocation under complex network topologies and multiple constraints.
[0070] The core idea of A2C is to combine the policy function Actor and the value function Critic for reinforcement learning training, and to improve the stability and efficiency of training by introducing the advantage function.
[0071] This application improves the traditional A2C method, incorporating the characteristics of traffic distribution to develop a deep reinforcement learning model. Because traffic distribution only affects the current time step and does not affect future traffic demand, a single-step reward mechanism is designed. By simplifying the reward structure, the training process can focus on immediate rewards and reduce reliance on long-term cumulative rewards, thereby simplifying the overall training objectives.
[0072] In this embodiment, the pseudo code of the deep reinforcement learning model is shown in Algorithm 1.
[0073]
[0074]
[0075] When the network state changes (e.g., a link change or a traffic matrix change), the DRL model defines the traffic distribution embedding as state s. Based on the current state s, the actor selects an action a, where action a represents the traffic distribution policy. During training, the policy gradient method is used to optimize the policy network parameters θ, aiming to enable the policy to select actions that produce higher rewards in the current state.
[0076] In order to increase the exploration ability, the output of the policy network is modeled as a Gaussian distribution with mean μ θ (s) and logarithmic standard deviation logσ θ (s), introducing randomness in the action selection process. Specifically, the policy network output is:
[0077] μ θ (s),logσ θ (s)=Actor(s,θ) (5)
[0078] In formula (5), σ θ (s)=exp(logσ θ (s)) is constrained to be positive, exp is the exponential function, which is the inverse function of the natural logarithm and is used to convert the logarithmic standard deviation logσ θ(s) is converted to the actual standard deviation σ θ (s).
[0079] During training, actions a are sampled from a Gaussian distribution:
[0080] a N(μ θ (s),σ θ (s) 2 )(6)
[0081] For simplicity, let Gaussian distribution N(μ θ (s),σ θ (s) 2 ) is expressed as π θ (a|s), represents the strategy used in the reinforcement learning process; in the deployment phase, the mean μ of the Gaussian distribution θ (s) are directly used as traffic distribution plans to ensure the stability and controllability of the strategy.
[0082] The Critic's task is to output a state value function V(s), which represents the expected return from the current state s. The Critic optimizes V(s) through regression to approximate the global total flow R(s,a) and help the Actor evaluate the quality of action a. The advantage function is defined as:
[0083] A(s,a)=R(s,a)-V(s) (7)
[0084] In formula (7), R(s,a) represents the direct reward for taking action a in state s, and V(s) is the state value function predicted by the Critic, which is used to evaluate the overall expected return of the current state.
[0085] In deep reinforcement learning, especially when using policy gradient methods, the focus is on updating the parameters θ:
[0086]
[0087] In formula (8), α is the learning rate of the policy network, A(s,a) is the advantage function, is the strategy π θ The gradient of the log probability of with respect to its parameters θ indicates how a small change in the parameters affects the probability of choosing action a under the current policy.
[0088] The value network is updated separately from the policy network. The value network is updated by minimizing the error between the predicted value and the actual return:
[0089]
[0090] In formula (9), Δw is the weight update, β is the learning rate of the value network, and the loss function L is defined as:
[0091] L=MSE(V(s),R(s,a)) (10)
[0092] In formula (10), MSE is the mean square error, V(s) is the value network’s estimate of state s, and R(s,a) is the actual return.
[0093] In the actual implementation, the feature extraction model, policy network, and value network are combined and trained end-to-end using the improved A2C method. In this architecture, θ and w represent the parameters that need to be learned in the policy network and value network, respectively, and the training goal is to optimize these two components simultaneously.
[0094] Based on A2C, this paper further considers the characteristics of TE and proposes a single-step reward mechanism that focuses on immediate feedback. The Actor and Critic networks are optimized through end-to-end joint training, where the state value function V(s) output by the Critic affects all parts of the Actor network, including the feature extraction model, through gradient backpropagation. The feature extraction model helps the policy network generate a better traffic allocation solution by capturing the topological features of the network. This approach not only maintains the efficiency of A2C, but also enhances the exploration ability of the model by introducing the advantage function A(s,a), ensuring the training efficiency in the training phase and the stability and accuracy of traffic allocation in the deployment phase. By combining the dual optimization objectives of Actor and Critic, this paper provides a more targeted and effective solution for achieving optimization goals (such as maximizing total traffic).
[0095] In this example, the variables used and their descriptions are shown in Table 1.
[0096] Table 1: Variables and their descriptions
[0097]
[0098]
[0099] S4: Perform constraint optimization on the initial traffic allocation ratio strategy and output the final traffic allocation strategy.
[0100] The traffic allocation ratio policy generated by the deep reinforcement learning model may still violate certain link capacity constraints, resulting in packet loss or poor traffic engineering performance. To address these issues and further improve the quality of the solution, we introduced a classic constrained optimization method, ADMM, which effectively alleviated the constraint violation problem and further improved the overall performance of traffic allocation.
[0101] The traffic engineering problem can be viewed as a complex combinatorial optimization problem, whose core goal is to allocate traffic demand to appropriate routing strategies to maximize the throughput of the entire network. However, due to the inherent complexity of combinatorial optimization problems, deep reinforcement learning-based methods are prone to falling into suboptimal solutions. This is because when a deep reinforcement learning agent makes a wrong decision, it is usually unable to undo the decision and explore other options. To overcome this problem, this application introduces ADMM after the DRL optimization process. ADMM is particularly suitable for dealing with optimization problems with complex constraints. By decoupling constraints, each variable (such as path flow and link flow) can be optimized independently according to its corresponding constraints, and then ADMM coordinates the updates between variables. This approach effectively alleviates the constraint violation problem, improves the quality of the optimization results, and generates better solutions.
[0102] The pseudo code of traffic engineering optimization based on ADMM is shown in Algorithm 2.
[0103]
[0104] S4-1: Transform the objective function and constraints into a form suitable for the ADMM method. The goal is to maximize the network throughput while satisfying three types of constraints: demand constraints, capacity constraints, and non-negativity constraints on traffic distribution.
[0105] In the original formula (2), the path flow F d (p)·d and link capacity c(e) are coupled, making the problem difficult to solve. To decouple path flow and link flow, a new variable z is defined pe , which represents the flow F of traffic d on link e through path p d (p)·d. Thus, the capacity constraint equation (2) is rewritten as:
[0106]
[0107] In formula (11), z pe represents the flow F that is assigned to path p and passes through link e. d (p)·d; c(e) represents the capacity of link e, e∈E, E represents the set of links in the network; F d (p) represents the proportion of flow d allocated to path p, ranging from 0 to 1.
[0108] At the same time, ADMM is more inclined to deal with equality constraints. Therefore, by introducing the slack variable ξ 1d and ξ 3e , converting the inequality constraints into equality constraints. These slack variables correspond to demand constraints and capacity constraints respectively. Formula (1) and Formula (11) are rewritten as:
[0109]
[0110] In formulas (12) and (13), ξ 1d represents the slack variable of the demand constraint; ξ 3e represents the slack variable of the capacity constraint.
[0111] S4-2: The key to ADMM is to construct a new objective function, called the augmented Lagrangian function, which combines the objective function with all constraints. The set of constraint functions is defined as follows:
[0112]
[0113] G 4dpe =F d (p)·dz pe (17)
[0114] In formulas (14), (15), (16), and (17), G(F, z, ξ) represents the set of constraint functions; G 1d represents the demand constraint function; G 3e represents the capacity constraint function; G 4dpe Represents the flow consistency constraint function.
[0115] According to the above definition, the augmented Lagrangian function is expressed as:
[0116]
[0117] In formula (18), L p (F, z, ξ, λ) is the augmented Lagrangian function; represents the objective function, corresponding to the traffic reward; λG(F,z,ξ) represents the penalty constraint violation, where λ is the Lagrange multiplier; Represents the quadratic penalty term, which is scaled by ρ and enhances the effect of the Lagrangian penalty.
[0118] S4-3: The core of the ADMM method is to iteratively optimize path flow, auxiliary variables and slack variables, adjust the Lagrange multiplier to gradually approach the feasible solution, and obtain the final flow allocation strategy.
[0119] In each iteration, the variables are optimized in turn as follows:
[0120] First, update the path flow:
[0121]
[0122] In formula (19), F k+1 Indicates the traffic distribution ratio after the k+1th iteration;
[0123] Lp (F,z k ,ξ k ,λ k ) is the augmented Lagrangian function; F represents the flow distribution ratio and is the decision variable for this step optimization; z k represents the auxiliary variable value after the kth iteration; ξ k represents the slack variable value after the kth iteration; λ k represents the Lagrange multiplier value after the kth iteration;
[0124] Next, update the auxiliary variables:
[0125]
[0126] In formula (20), z k+1 represents the auxiliary variable value after the k+1th iteration; z represents the auxiliary variable, which is the decision variable optimized in this step;
[0127] Then, update the slack variables:
[0128]
[0129] In formula (21), ξ k+1 represents the slack variable value after the k+1th iteration; ξ represents the slack variable, which is the decision variable optimized in this step;
[0130] Finally, update the Lagrange multiplier:
[0131] λ k+1 =λ k +ρ·G(F k+1 , z k+1 ,ξ k+1 ) (twenty two)
[0132] In formula (22), λ k+1 represents the Lagrange multiplier value after the k+1th iteration, and λ represents the Lagrange multiplier value, which is the decision variable optimized in this step.
[0133] The initial traffic allocation ratio policy can be generated by the existing policy network (warm start), which accelerates convergence and avoids instability caused by random initialization. This process ensures that the problem gradually approaches the optimal solution while satisfying all constraints.
[0134] In this embodiment, an experimental evaluation was conducted to evaluate the practicability of the above-mentioned system and method.
[0135] 1. The following WAN dataset was used:
[0136] B4: A private wide area network connecting Google's global data centers, designed to meet the high bandwidth requirements of communications between data centers.
[0137] UsCarrier and Kdl: These topologies are part of the Internet Topology Repository, a network data repository built from public information provided by network operators.
[0138] ASN: An autonomous system-level Internet topology for wide area networks that provides a realistic representation of inter-domain routing.
[0139] Table 2: Details of each topology
[0140]
[0141]
[0142] Table 2 provides detailed information for each topology, including the number of nodes, number of edges, average shortest path length, and network diameter. Generally, as the size of a network topology increases, the average shortest path length and network diameter also increase. However, the ASN topology exhibits an exception to this trend. This can be attributed to the presence of star-shaped clusters in the ASN, which interconnect and provide strong cluster-level connectivity.
[0143] The traffic matrices and link capacities of all topologies are obtained from Teal. In addition, for each pair of nodes, we precompute four shortest paths as candidate paths for traffic distribution.
[0144] 2. Benchmark comparison
[0145] Comparison is made with several state-of-the-art traffic engineering schemes designed to quickly solve traffic engineering problems for large-scale topologies:
[0146] NCFlow (2021): NCFlow divides the network into multiple clusters, uses linear programming (LP) to solve traffic engineering subproblems in each cluster, and merges the results to generate a global solution.
[0147] Teal (2023): Teal combines multi-agent reinforcement learning with ADMM to optimize traffic engineering. It uses FlowGNN to extract features, adopts a policy network for traffic allocation, and uses ADMM to optimize the solution to maximize throughput.
[0148] 3. Experimental Setup
[0149] The operation is as follows: In each time window, GRL-TE receives the WAN topology, link capacity information, a traffic matrix containing the demand between each pair of nodes, and a set of four precomputed paths for each pair of nodes. It outputs the traffic distribution for each pair of nodes along these four paths. To implement GRL-TE, we developed its core modules using the PyTorch framework and optimized its performance through hyperparameter tuning experiments.
[0150] Modeling the Feature Extraction Model: The feature extraction model consists of six layers alternating between GNN and DNN layers. The GNN uses GINConv with an MLP to capture structural information and capacity constraints, while the DNN uses an MLP to model traffic demand constraints. Each GINConv layer contains an MLP with two linear layers and ReLU activations, gradually increasing feature representation capability. The DNN layer dynamically adjusts the input dimension based on the number of paths and the current embedding dimension to ensure compatibility with the changing feature space. The initial embedding dimension is 1, and each layer increases the dimension by 1 through skip connections, resulting in a final embedding dimension of 7. This design enhances feature representation capability while maintaining constraint awareness.
[0151] Modeling Deep Reinforcement Learning: The policy network uses a fully connected neural network as its output layer, with an output dimension equal to the number of paths defined in the network. For a setting with four paths, the policy network has an input dimension of 28 and an output dimension of 4. The value network is a three-layer fully connected neural network that estimates the value of a given global state. The value network's input dimension is the feature dimension corresponding to the global state. It has two hidden layers, each with a hidden dimension of 256. The output layer maps the hidden features to a scalar value representing the estimated value of the current state.
[0152] Modeling ADMM: For smaller topologies, such as B4, two ADMM iterations were performed. For larger topologies, including UsCarrier, Kdl, and ASN, 25 iterations were performed. The penalty factor ρ was set to 1, which determines the penalty strength for constraint violations during the ADMM optimization process.
[0153] The training process is set to 3 rounds, and the Adam optimizer is used for training. The learning rate of the policy network is set to 1×10 -4 , the learning rate of the value network is set to 1×10 -3 .
[0154] For the benchmark model NCFlow, the number of clusters specified in the original paper is used (36 for UsCarrier and 81 for Kdl). The default partitioning method (FMPartitioning) is used for the B4 and ASN datasets.
[0155] All experiments were conducted on a system equipped with Nvidia GeForce RTX 4090 GPUs to ensure computational efficiency and resource optimization. The server is equipped with an Intel Xeon Silver 4310 processor (12 cores, 2.1GHz), 32GB of DDR4 RAM (3200MHz), and four Nvidia GeForce RTX 4090 GPUs.
[0156] 4. Comparison of this method with existing technologies
[0157] In order to eliminate the interference of random initialization in the neural network on the evaluation results, a systematic random seed control strategy was adopted to ensure that the parameter initialization in each training is consistent. Specifically, 10 independent training experiments were conducted for each network topology, and the random seed was set from 100 to 1000 in sequence. This method ensures the stability and statistical significance of the results, effectively reduces random fluctuations during training, and thus makes the performance comparison of the model more reliable. The experimental results that meet the requirements on average are as follows Figure 4 shown.
[0158] In terms of network performance evaluation, GRL-TE demonstrates significant advantages:
[0159] (1) Small-scale network scenario: In the B4 topology (number of nodes: 12), GRL-TE achieved a demand satisfaction rate of 92.15%, which was 2.76 and 7.38 percentage points higher than NCFlow (89.39%) and Teal (84.77%), respectively. In the UsCarrier topology (number of nodes: 158), GRL-TE continued to show similar advantages, with a demand satisfaction rate of 91.96%, which was 25.66 and 5.15 percentage points higher than NCFlow (66.30%) and Teal (86.81%), respectively.
[0160] (2) Large-scale network scenario: In the Kdl topology (number of nodes: 754), GRL-TE maintains a demand satisfaction rate of up to 90.84%, which is 35.16 and 6.07 percentage points higher than NCFlow (55.68%) and Teal (84.77%), respectively. In the more complex ASN topology (number of nodes: 1739), GRL-TE achieves a demand satisfaction rate of 87.69%, which is 44.9 and 4.82 percentage points higher than NCFlow (42.79%) and Teal (82.87%), respectively, verifying the strong scalability of this method in large-scale networks.
[0161] Further experimental data demonstrates that GRL-TE maintains a consistent performance advantage across various network topologies: Across four benchmark topologies, GRL-TE achieved an average demand fulfillment rate of 90.53%, 3.21 percentage points higher than Teal. In particular, in ASN topologies with over 1,700 nodes, GRL-TE maintained a high demand fulfillment rate of over 87%, fully demonstrating the method's strong adaptability in complex network environments. GRL-TE's consistent performance across multiple scenarios strongly validates its enormous potential and value in practical engineering applications.
[0162] 5. Dealing with link failures
[0163] This experiment simulates various link failure scenarios to verify the robustness of GRL-TE in dynamic topologies.
[0164] like Figure 5 As shown, GRL-TE achieves a high demand fulfillment rate of 92.15% with no link failures, significantly outperforming NCFlow's 89.39% and Teal's 84.77%. Its standard deviation is only 5.52, far lower than Teal's 23.15, demonstrating the efficiency and stability of its strategy. Even when three link failures are encountered, GRL-TE maintains a 75.94% demand fulfillment rate, surpassing NCFlow's 75.35% and Teal's 69.70%. Its standard deviation remains stable at 5.62, highlighting its fault tolerance.
[0165] Although all methods experience a decrease in demand fulfillment rate when link failures increase, GRL-TE consistently outperforms other methods. Figure 6 As shown in Figure 2, under five link failures, GRL-TE achieves a satisfaction rate of 78.49%, significantly higher than NCFlow's 59.20% and Teal's 74.31%. The standard deviation of GRL-TE remains below 3.0, significantly lower than Teal's 13.76-16.78, indicating that GRL-TE's strategy is more stable.
[0166] like Figure 7 As shown in Figure 3, GRL-TE achieves the highest demand satisfaction rate and the lowest standard deviation in all test scenarios on the Kdl dataset, demonstrating its robustness in the presence of link failures. Although NCFlow exhibits some stability when failures increase, its satisfaction rate is generally lower than that of GRL-TE and Teal.
[0167] like Figure 8As shown in the figure, GRL-TE performs well in the fault-free scenario, with a satisfaction rate of 87.69% and a standard deviation of 1.59. Even under extreme conditions with up to 200 link failures, GRL-TE still maintains a satisfaction rate of 85.01% with a maximum standard deviation of 2.70, demonstrating its robustness. In addition, GRL-TE can quickly recalculate traffic distribution on the faulty topology (capacity is zero after a link failure), mitigating the impact during temporary link failures. This feature not only improves the system's response speed but also enhances its adaptability under uncertain conditions, which is the key to GRL-TE's excellent performance in dynamic network environments.
[0168] 6. Ablation Experiment
[0169] As shown in Table 3, ablation experiments on four datasets demonstrate the significant performance advantage of the full GRL-TE model in demand satisfaction.
[0170] Table 3: Ablation experiments on four datasets
[0171]
[0172] The complete GRL-TE model consistently achieved the highest demand fulfillment rates across all datasets (B4: 92.15%, UsCarrier: 91.96%, Kdl: 90.84%, ASN: 87.69%) with low standard deviations, demonstrating its superior stability and adaptability.
[0173] When GNNs are removed, the performance degrades significantly, with demand fulfillment decreasing by 4.44-7.97% and the standard deviation increasing significantly, highlighting the critical role of GNNs in modeling network topology and managing path contention, especially in large-scale networks.
[0174] Similarly, removing ADMM leads to further performance degradation, with demand fulfillment decreasing by 6.74-12.66%, but the standard deviation remains high, highlighting the importance of ADMM in ensuring feasible solutions through distributed optimization.
[0175] In contrast, the impact of removing DNN is relatively small, resulting in only a slight decrease in demand fulfillment rate (0.28-1.90%) and minimal change in standard deviation, indicating that DNN serves more as an auxiliary module for path-level optimization.
[0176] Overall, the complete GRL-TE model performs well in various network scenarios. GNN and ADMM are the core of its performance, while DNN plays a supporting role. This verifies the effectiveness and robustness of the GRL-TE model design.
[0177] 7. Computation time comparison
[0178] We measured the total allocated computation time for each solution on each TM, carefully excluding initialization and other one-time overhead to ensure that the results accurately reflect computational performance. Table 4 shows how the total allocated computation time for each solution is calculated. For example, Teal's computation time includes the total runtime (using the GPU), while NCFlow takes into account the Gurobi runtime and the additional time for merging subproblems.
[0179] Table 4: Comparison of computation time for different solutions
[0180]
[0181]
[0182] The experimental results of average computing time are as follows Figure 9 As shown in Figure 2, the computation time of each method under different topologies is compared.
[0183] While GRL-TE's computation time is slightly longer than the comparison methods, it still meets the practical requirements of production-grade WANs. According to industry standards, centralized traffic engineering controllers must complete traffic demand allocation within a 5-minute cycle. GRL-TE significantly exceeds this threshold in all topologies, demonstrating its feasibility in real-world applications.
[0184] (1) Small-scale network scenario: In the B4 topology (number of nodes: 12), the computation time of GRL-TE is 0.0122 seconds, which is 64.90% less than NCFlow. Although slightly higher than Teal (0.0098 seconds), it is still only 0.0040% of the 5-minute threshold. In the UsCarrier topology (number of nodes: 158), the computation time of GRL-TE is 0.0801 seconds, which is 59.70% better than NCFlow and only 0.8 times the time of Teal.
[0185] (2) Large-scale network scenario: In the KDL topology (number of nodes: 754), the computation time of GRL-TE is 1.8056 seconds, which is only 1.42 times that of NCFlow, while the demand satisfaction rate is improved by 35.16 percentage points. In the larger ASN topology (number of nodes: 1739), the computation time of GRL-TE is 1.5429 seconds, which is 97.7% less than that of NCFlow (67.2494 seconds), while the demand satisfaction rate is improved by 44.90 percentage points, demonstrating the excellent scalability of this method in large-scale networks.
[0186] Across all test scenarios, GRL-TE's maximum computation time was 1.8056 seconds, representing only 0.60% of the 5-minute threshold, fully meeting periodic computational requirements. Although Teal's computational time was lower in some scenarios (such as B4), its demand fulfillment rate was, on average, 6.23 percentage points lower than that of GRL-TE. This balance of performance and efficiency demonstrates that GRL-TE can effectively support the traffic engineering requirements of modern wide area networks, particularly in complex networks with over 1,000 nodes, where its computational efficiency is exponentially improved compared to traditional methods.
[0187] This paper addresses the challenge of efficient traffic engineering in large-scale wide area networks (WANs) and proposes GRL-TE, an innovative approach that integrates graph neural networks with an enhanced A2C architecture. Unlike traditional approaches that rely on centralized optimization solvers, GRL-TE models the WAN as a bipartite graph of paths and links. This approach cleverly leverages GINConv to achieve efficient information exchange between links and paths, while employing MLP to manage the collaboration and competition among multiple paths serving the same traffic demand. This design enables the model to simultaneously consider network topology and traffic demand characteristics at the embedding layer. Furthermore, the improved A2C method generates traffic allocation policies, while ADMM is used in the post-processing stage to effectively address potential overcapacity or load imbalance issues. Experiments on real-world network topologies (e.g., B4, UsCarrier, KDL, and ASN) demonstrate that GRL-TE not only maintains high throughput and strong fault tolerance in the face of link failures, but also completes all processing within a 5-minute scheduling window. The experimental results fully demonstrate the scalability and practical value of this approach in production environments.
[0188] Looking ahead, GRL-TE can be extended to multi-objective optimization, such as considering factors like latency, jitter, and energy consumption, to address a wider range of requirements in wide area network traffic scheduling. Furthermore, exploring the application of GRL-TE in multi-domain or cross-carrier collaborative environments and validating its feasibility in more complex environments, such as data center networks, remain promising directions for further research.
Claims
1. An optimization system for SDN traffic engineering, characterized in that: It includes a control plane and a data execution plane, and the control plane and the data execution plane are bidirectionally connected; The control plane includes a monitoring module, a traffic engineering controller and a flow rule translation module; Among them, the monitoring module is used to collect the traffic matrix of the network; Traffic engineering controller, used to generate traffic distribution strategy based on the collected traffic matrix; The flow rule translation module is used to convert the traffic distribution policy into flow table rules that can be executed by the data execution plane; The data execution plane includes at least one device that actually carries network traffic, which is used to execute traffic distribution strategies according to received flow table rules, adjust traffic forwarding and distribution, and feed back real-time network status to the monitoring module of the control plane.
2. An optimization method for SDN traffic engineering based on the system of claim 1, characterized in that: The specific steps include: S1: Obtain the traffic matrix in the network and represent the network topology as a graph structure; S2: Input the traffic matrix and graph structure into the constructed feature extraction model and output the feature vector; S3: Input the feature vector into the deep reinforcement learning model for iterative optimization, and output the initial traffic allocation ratio strategy; S4: Perform constraint optimization on the initial traffic allocation ratio strategy and output the final traffic allocation strategy.
3. The optimization method for SDN traffic engineering according to claim 2, characterized in that: In S1, the network topology is usually represented as a graph structure G = (V, E, c), where V represents the set of network nodes responsible for forwarding data packets, E represents the set of physical links connecting network nodes and supporting data transmission, and c represents the maximum data transmission rate that can be supported, that is, the capacity allocated to each link.
4. The optimization method for SDN traffic engineering according to claim 2, characterized in that: The SDN traffic engineering includes the following three key constraints: demand constraint, capacity constraint and non-negative constraint:
5. The optimization method for SDN traffic engineering according to claim 2, wherein: The S2 comprises the following steps: S2-1: Model the graph structure as a bipartite graph, where nodes are divided into two disjoint subsets: edge nodes and path nodes. Edge nodes represent links, while path nodes represent the paths that traffic demands pass through. S2-2: Input the edge nodes and path nodes into the constructed feature extraction model and output the feature vector.
6. The optimization method for SDN traffic engineering according to claim 5, characterized in that: In S2-2, the feature extraction model includes at least one feature extraction unit, the feature extraction unit includes a GNN layer and a DNN layer, and the output end of the GNN layer is connected to the input end of the DNN layer; The GINConv module of the GNN layer is used to model the interaction between links to ensure that the allocated traffic does not exceed the link capacity, thereby satisfying the capacity constraint; The multi-layer perceptron module of the DNN layer is used to process the path node embeddings related to the same demand, perform conversion and update, and ensure that the traffic distribution is reasonable and meets the demand constraints.
7. The optimization method for SDN traffic engineering according to claim 2, characterized in that: In S3, the deep reinforcement learning model uses the A2C architecture for policy optimization: the Actor generates a traffic allocation ratio strategy based on the feature vector, and the Critic evaluates the quality of the current traffic allocation ratio strategy by calculating the value function.
8. The optimization method for SDN traffic engineering according to claim 2, wherein: In S3, the iterative optimization method is specifically as follows: When the network state changes, the traffic distribution embedding is defined as state s. The policy network Actor selects an action a based on the current state s. The output action a represents the initial traffic distribution ratio strategy. Model the output of the policy network as a Gaussian distribution with mean μ θ (s) and logarithmic standard deviation logσ θ (s) to introduce randomness in the action selection process: m θ (s),logσ θ (s)=Actor(s,θ) (4) In formula (4), σ θ (s)=exp(logσ θ (s)) is constrained to be positive, and exp is the exponential function used to convert the logarithmic standard deviation logσ θ (s) is converted to the actual standard deviation σ θ (s); θ represents the parameters of the policy network; During training, actions a are sampled from a Gaussian distribution: a N(μ θ (s),s θ (s) 2 (5) Formula (5), for simplicity, the Gaussian distribution N(μ θ (s),σ θ (s) 2 ) is expressed as π θ (a|s), represents the strategy used in the reinforcement learning process; in the deployment phase, the mean μ of the Gaussian distribution θ (s) is directly used as a traffic distribution plan to ensure the stability and controllability of the strategy; The task of the value network Critic is to output the state value function V(s), which is the expected return of the current state s. It optimizes V(s) through regression to approximate the global total flow R(s,a) and help the Actor evaluate the quality of action a: A(s,a)=R(s,a)-V(s) (6) In formula (6), A(s,a) represents the advantage function; R(s,a) represents the direct reward for taking action a in state s, and V(s) is the state value function predicted by the critic to evaluate the overall expected return of the current state; In deep reinforcement learning (DRL), the update of the parameter θ is: In formula (7), α is the learning rate of the policy network, A(s,a) is the advantage function, is the strategy π θ The gradient of the log probability with respect to its parameters θ, which indicates how a small change in the parameters affects the probability of choosing action a under the current policy; The update of the value network is handled separately from the policy network. The value network is updated by minimizing the error between the predicted value and the actual return: In formula (8), Δw is the weight update, β is the learning rate of the value network, and the loss function L is defined as: L=MSE(V(s),R(s,a)) (9) In formula (9), MSE is the mean square error, V(s) is the value network’s estimate of state s, and R(s,a) is the actual return.
9. The optimization method for SDN traffic engineering according to claim 2, wherein: The S4 includes: S4-1: Transform the objective function and constraints into a form suitable for the ADMM method. The goal is to maximize the network throughput while satisfying three types of constraints: demand constraint, capacity constraint, and non-negativity constraint of traffic distribution. To decouple path traffic and link traffic, a new variable z is defined pe , represents the flow F that is allocated to path p and passes through link e d (p)·d, the capacity constraint equation (2) is rewritten as: In formula (10), z pe represents the flow d assigned to path p and passing through link e; c(e) represents the capacity of link e, e∈E, E represents the set of links in the network; F d (p) represents the proportion of flow demand d allocated to path p, ranging from 0 to 1; By introducing the slack variable ξ 1d and ξ 3e , converting the inequality constraints into equality constraints. These slack variables correspond to demand constraints and capacity constraints respectively. Formula (1) and Formula (10) are rewritten as: In formulas (11) and (12), ξ 1d represents the slack variable of the demand constraint; ξ 3e represents the slack variable of the capacity constraint; S4-2: Construct the augmented Lagrangian function and define the set of constraint functions as follows: G 4dpe =F d (p)·d-z pe (16) In formulas (13), (14), (15), and (16), G(F, z, ξ) represents the set of constraint functions; G 1d represents the demand constraint function; G 3e represents the capacity constraint function; G 4dpe represents the flow consistency constraint function; According to the above definition, the augmented Lagrangian function is expressed as: In formula (17), L p (F, z, ξ, λ) is the augmented Lagrangian function; represents the objective function, corresponding to the traffic reward; λG(F,z,ξ) represents the penalty constraint violation, where λ is the Lagrange multiplier; represents the quadratic penalty term, which is scaled by ρ and enhances the effect of Lagrangian penalty; S4-3: Iteratively optimize the path flow, auxiliary variables and slack variables, adjust the Lagrange multiplier to gradually approach the feasible solution, and obtain the final flow allocation strategy.
10. The optimization method for SDN traffic engineering according to claim 9, characterized in that: In the above S4-3, First, the path flow is updated: In formula (18), F k+1 Indicates the traffic distribution ratio after the k+1th iteration; L p (F,z k ,ξ k ,λ k ) is the augmented Lagrangian function; F represents the flow distribution ratio; z k represents the auxiliary variable value after the kth iteration; ξ k represents the slack variable value after the kth iteration; λ k represents the Lagrange multiplier value after the kth iteration; Next, update the auxiliary variables: In formula (19), z k+1 represents the auxiliary variable value after the k+1th iteration; z represents the auxiliary variable; Then, update the slack variables: In formula (20), ξ k+1 represents the slack variable value after the k+1th iteration; ξ represents the slack variable; Finally, update the Lagrange multiplier: l k+1 =λ k +ρ·G(F k+1 ,z k+1 ,x k+1 ) (21) In formula (21), λ k+1 Represents the Lagrange multiplier value after the k+1th iteration.
Citation Information
Cited By
Cross-environment generalization network flow optimization method based on large model agent
CN121462429A
Cross-environment generalization network traffic optimization method based on large model agent
CN121462429B