Joint optimization method for network resource partitioning and path planning based on bi-level programming
By employing a joint optimization method combining two-level programming and reinforcement learning, network resources and routing strategies are coordinated, addressing the flexibility and efficiency issues of network slicing technology when facing diverse service demands, thereby improving network performance and resource utilization.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-07-31
- Publication Date
- 2026-03-31
AI Technical Summary
Existing network slicing technology is unable to quickly adapt to changes in business needs, traditional routing strategies cannot optimize network resource utilization, and machine learning routing methods cannot meet the QoS requirements of various types of services, resulting in low network performance and resource utilization efficiency.
A joint optimization method for network resource allocation and path planning based on bi-level programming is adopted. Through iterative computation and reinforcement learning algorithms, the network resource allocation and routing strategies are coordinated to optimize network topology utilization and communication data flow satisfaction.
It effectively reduces business communication latency and packet loss rate, improves network bandwidth utilization, increases network capacity, and enhances network performance.
Smart Images

Figure CN116915622B_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of computer networks, specifically a method for joint optimization of network resource allocation and path planning based on two-layer planning. Background Technology
[0002] The rapidly developing internet has become an indispensable part of today's technological landscape. With the advancement of network technology, various applications are emerging, and future communication networks will serve diverse vertical industries, needing to meet various types of business requirements. Simultaneously, different services have varying QoS requirements, such as latency, bandwidth, and reliability, yet all must coexist within the same network architecture. To address the personalized needs of different services, communication networks need to provide more flexible and efficient demand-based multi-type services.
[0003] Network slicing technology is a fusion of Software Defined Networking (SDN) and Network Function Virtualization (NFV) technologies, such as... Figure 1 As shown, a network slice comprises a set of network resources deployed and configured into a complete logical network to meet the specific network characteristics required for the service. Each network slice is logically independent and adaptable to various business applications with specific needs. To meet the diverse and differentiated business requirements in the network, it is necessary to dynamically adjust network slice resources and flexibly manage available resources in the network, thereby improving the utilization efficiency of network resources while satisfying business needs.
[0004] In addition, network routing policies determine forwarding paths for services solely based on source and destination addresses, which cannot prevent congestion and the selected paths are often insufficient to meet the actual needs of the services. To address this issue, traditional QoS routing and other technical solutions require acquiring a large amount of link-state information to calculate feasible paths and maintain routing resources; however, routing policies do not optimize according to actual conditions, making the generation and maintenance of routing policies very difficult.
[0005] In recent years, the development of machine learning, especially reinforcement learning, has provided new opportunities for the generation of routing policies. Various machine learning-based routing algorithms have been proposed to incorporate various information such as network status and service QoS into the scope of routing decisions, thereby formulating online routing policies for various dynamic service needs.
[0006] While these solutions have improved network performance, they still face many challenges:
[0007] First, network slicing requires reconfiguring resources based on changes in business needs, which is not possible in traditional networks to quickly adapt to network conditions and make real-time decisions.
[0008] Secondly, various machine learning-based routing methods only consider the needs of a single business when formulating routing strategies, and do not fully consider the characteristics of global network services. They cannot meet the QoS requirements of multiple types of services, let alone achieve the goal of adaptively ensuring service QoS.
[0009] However, in reality, network resource allocation and service routing strategy formulation are an organic whole, and they influence each other. Summary of the Invention
[0010] To address the challenges of integrating network management and control within existing network topologies, this invention proposes a joint optimization method for network resource allocation and path planning based on bi-layer programming. This method decomposes the joint optimization scenario into a bi-layer programming model, uses iterative methods to solve the joint optimization problem in this scenario, and maximizes network topology utilization and communication data flow satisfaction. It can explicitly and rationally coordinate the allocation of network resources, while optimizing the routing strategy of communication data flows within each type of network resource.
[0011] The specific steps of the network resource allocation and path planning joint optimization method based on two-layer programming are as follows:
[0012] Step 1: Obtain the topology information of the network under test, and construct a directed graph of the communication network topology based on SDN using each link and forwarding node;
[0013] Link attributes include the two forwarding nodes connecting the link, bandwidth capacity, and transmission rate; forwarding node attributes include the node's location, etc. A unique identifier E is used. i Each link in the network under test is represented by a unique identifier V. i Indicates a forwarding node in the network;
[0014] The directed graph of the communication network topology refers to a directed graph G = (V, E) that includes a set of links and a set of forwarding nodes, where edge E represents the set of links in the network topology and node V represents the set of forwarding nodes in the network topology. m =(v i ,v j )∈E represents a node v in a directed graph. i and v j The nodes are connected, and there is a communication link between the forwarding nodes corresponding to the two nodes.
[0015] Step 2: Based on network modeling, 5G network communication traffic characteristics, and QoS quantification indicators, all communication traffic of the network under test is classified into four slice types.
[0016] The four slice types are latency-sensitive, bandwidth-sensitive, latency-bandwidth-sensitive, and latency-and-packet-loss-sensitive slices. Each slice type contains several service slices, and each slice type selects the most stringent constraint from its respective service slices as the slice type constraint.
[0017] Step 3: Based on the directed graph and slice types of the network topology model, the bandwidth resources of each link in the network under test are divided into four slices according to the service requirements.
[0018] Each slice only occupies the bandwidth resources of its own slice; the allocation method is as follows:
[0019] The resources required for the k-th slice are represented as follows:
[0020] b m λ is the bandwidth of the communication link. k The proportion allocated to the k-th slice; λ k ∈[0,1], and K represents the number of slices in the network under test.
[0021] Step 4: Design a reinforcement learning algorithm model for slice resource partitioning, which is used to partition the data stream of request communication according to slice resources and calculate the bandwidth resources of all links occupied by each type of slice.
[0022] The slice resource partitioning reinforcement learning algorithm model takes a directed graph of network topology and request communication flow data as input, and outputs the resource proportion of each slice.
[0023] The specific modeling process is as follows:
[0024] Step 401, calculate the m-th link of the k-th slice. The bandwidth utilization rate is:
[0025] Link at time t Bandwidth usage:
[0026] Indicates whether there is a service flow on this link at time t. Select as a component of its path; a value of 1 indicates yes, and a value of 0 indicates no. Indicates that the link passes through at time t. The number of business flows; This indicates that a request was made on the link at time t. Throughput.
[0027] Step 402: Calculate the bandwidth resource utilization rate of the k-th slice on each link at time t.
[0028]
[0029] Where M represents the set of all links;
[0030] Step 403: Calculate the average bandwidth resource utilization of all slices on each link within time T:
[0031]
[0032] Step 404: Use the variance of average resource utilization to measure the difference in resource utilization across all slice links;
[0033] Represented as:
[0034] Step 405, calculate the average variance of all slices within time T using the variability in resource utilization:
[0035]
[0036] Step 406: Obtain the optimization objective function of the slice resource partitioning reinforcement learning algorithm model, and allocate resources reasonably to each slice to uniformly improve the resource utilization of all links;
[0037] The objective function is expressed as:
[0038] Step 5: Based on the resource allocation of the slice reinforcement learning algorithm model and the actual bandwidth resource constraints of the data stream requesting communication, calculate the candidate path set.
[0039] First, for the k-th slice, based on the source node v s and the destination node v d Find their respective neighbor nodes v s' ,v d' ∈V, forming the logical links of this slice, i.e. All logical links constitute the candidate path set.
[0040] Step 6: Each slice selects a path from the candidate path set to provide services for its respective data flow based on the routing algorithm, according to bandwidth resource constraints, network status and service requirements.
[0041] Step 7: Based on the selected path for each slice, and in combination with the requirements for latency, throughput, and packet loss rate of data flow requests, formulate a joint optimization objective for the user satisfaction model and routing strategy.
[0042] The optimization goal of the routing strategy is to maximize the satisfaction of all users over a long period of time, expressed as:
[0043] The selected path for the k-th slice; N represents all choices. Let γ be the set of all services along the path, and let γ be a fixed constant. For the satisfaction model; it is represented as:
[0044]
[0045] β d β represents the weight of delay in the satisfaction model. th β represents the weight of bandwidth in the satisfaction model. l The weight of packet loss rate in the satisfaction model. This is the delay constraint corresponding to the slice type.
[0046] Data Stream Request The network requirements for latency, throughput, and packet loss rate at time t are respectively expressed as: and At time t, on the selected path The actual packet loss, throughput, and packet loss rate obtained are expressed as follows: and
[0047] Step 8: Using an iterative calculation method based on a two-layer model for joint optimization of network resources and routing, the optimal strategy for the joint optimization objective is solved.
[0048] The specific process is as follows:
[0049] Step 801: Utilize slice resources to divide the reinforcement learning algorithm model and the user satisfaction model, and build an optimization model based on bi-layer programming for the upper and lower layers;
[0050] The upper layer performs slicing, with the ultimate goal of maximizing network resource utilization;
[0051] The specific formula is as follows:
[0052] The lower layer performs routing and optimization, with the goal of optimizing all indicators of this data flow and maximizing service satisfaction.
[0053] The specific formula is as follows:
[0054] Step 802: Use two reinforcement learning models to make decisions at the upper and lower layers respectively, and achieve joint optimization of the two-layer problem through continuous iteration using hierarchical reinforcement learning.
[0055] The upper-level algorithm of this model can make the optimal slicing action based on the network resource status, and the strategy is represented as π. u Then at time t, its state is According to strategy π u Make an action
[0056] SDN controller module based on action Allocate network resources and provide feedback to lower-level intelligent agents.
[0057] The lower-level agent calculates the specific routing information p of the data stream requesting communication based on the allocated network resources, and then requests data stream communication.
[0058] Based on the actual network status information, the lower-level traffic satisfaction model calculates the traffic satisfaction rate. The routing strategy is optimized, and at the same time, the upper-layer algorithm aims to maximize network utilization while also calculating the agent's reward value. Optimize slicing strategy.
[0059] The upper-layer agent then reallocates resources based on the current network status and the needs reported by the lower layer. The lower layer then makes routing decisions. This process continues until the rewards for both the upper and lower layer agents reach their maximum. At this point, the optimal strategy of the joint optimization method is obtained, namely, the optimal resource allocation and the optimal routing strategy.
[0060] The advantages of this invention are:
[0061] 1) The present invention is based on a joint optimization method for network resource partitioning and path planning in a two-layer planning approach. When faced with multiple service flows with different constraints, it can effectively reduce service communication latency, packet loss rate, and jitter rate, thereby improving network performance.
[0062] 2) The present invention is based on a two-layer planning network resource allocation and path planning joint optimization method, which can improve network bandwidth utilization and increase network capacity. Attached Figure Description
[0063] Figure 1 This is a schematic diagram of network slicing in the prior art;
[0064] Figure 2 This is an overview of the architecture of the joint optimization method for network resource partitioning and path planning of this invention;
[0065] Figure 3 This is a flowchart of the network resource partitioning and path planning joint optimization method based on two-layer planning according to the present invention;
[0066] Figure 4 This is the intelligent agent and network interaction model of the present invention;
[0067] Figure 5 These are the execution steps of the path planning algorithm of the present invention;
[0068] Figure 6 These are the execution steps of the two-layer model algorithm for joint optimization of network resources and routing in this invention. Detailed Implementation
[0069] To facilitate understanding and implementation of the present invention by those skilled in the art, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments.
[0070] This invention is based on a joint optimization method for network resource partitioning and path planning using a two-layer programming approach. The architecture of the proposed method is as follows: Figure 2 As shown, firstly, a two-layer interaction model based on bi-layer programming is constructed. In the joint optimization process of resource partitioning and routing strategies, the state space and computational complexity increase exponentially compared to traditional single optimization problems, making it impossible to solve using a single optimization objective. Therefore, this invention constructs the solutions for slice partitioning and route generation into two nested layers, describing the game process between the two parties with a master-slave relationship. The objective of the upper-layer slice partitioning is to maximize the overall network utilization; the objective of the lower-layer routing strategy optimization is to optimize all indicators of the data flow and maximize service satisfaction. The purpose of this method is to find the joint optimal strategy of the two mutually influential layers.
[0071] Secondly, based on the interaction between the upper and lower layer models, this invention proposes an iterative calculation method for a two-layer model of joint optimization of network resources and routing. This method achieves iteration between the upper and lower layers through interactive feedback, optimizing the optimal strategy at each layer. Through continuous iteration, the upper and lower layer models obtain the optimal strategy. Because the two-layer programming model is combined with the joint optimization problem, the state space and action space of the problem scenario grow exponentially. Traditional solutions to the two-layer programming model are no longer applicable, and using heuristic algorithms to solve this problem can lead to the optimization strategy getting trapped in local optima.
[0072] Based on this situation, this invention designs a two-layer model that interacts cyclically, with each iteration being an iterative optimization until the joint optimization method reaches its optimum, thus solving the problem of an excessively large scenario space. Furthermore, during continuous iteration, this method considers global state information, avoiding the generation of local optima. Therefore, for this two-layer planning model, this invention uses reinforcement learning methods to solve the slice partitioning and path selection problems in the upper and lower layers respectively. Then, the reinforcement learning methods of the upper and lower layers interact and iterate. The resource partitioning actions generated by the upper layer serve as prerequisite constraints for the routing selection in the lower layer, and the routing actions generated by the lower layer affect the overall network state, becoming state feedback for the upper layer. Both the upper and lower reinforcement learning methods retain a certain exploration space when generating actions to avoid getting trapped in local optima, and continue to interact and run until both reach their optimal strategies.
[0073] The network resource partitioning and path planning joint optimization method based on two-level programming, such as Figure 3 As shown, the specific steps are as follows:
[0074] Step 1: Obtain the topology information of the network under test, and construct a directed graph of the communication network topology based on SDN using each link and forwarding node;
[0075] For traditional operator wide area networks, an SDN-based communication network architecture is constructed, which is divided into an SDN controller module, an intelligent algorithm module, and a network forwarding module. The SDN controller module is responsible for network resource management, network information collection, and slice routing deployment. The intelligent algorithm module is responsible for performing deep learning calculations on the received network information to obtain the results of slice partitioning and routing policy deployment. The network forwarding module is the network topology deployed with SDN switches.
[0076] First, obtain the topology information of the network under test, including the location, connectivity, and bandwidth capacity of links and forwarding nodes. Second, use the unique identifier E. i This represents each link in the network under test. Link attributes include the two forwarding nodes connected to the link, bandwidth capacity, and transmission rate. A unique identifier V is used. i This represents a forwarding node in the network; the attributes of a forwarding node include its location, etc.
[0077] The directed graph of the communication network topology refers to a directed graph G = (V, E) that includes a set of links and a set of forwarding nodes, where edge E represents the set of links in the network topology and node V represents the set of forwarding nodes in the network topology. m =(v i ,v j )∈E represents a node v in a directed graph. i and v j The nodes are connected, and there is a communication link between the forwarding nodes corresponding to the two nodes. The communication bandwidth is denoted as b. m .
[0078] Step 2: Based on network modeling, 5G network communication traffic characteristics, and QoS quantification indicators, all communication traffic of the network under test is classified into four slice types.
[0079] QoS quantification metrics include rate, latency, and data loss rate;
[0080] The four slice types are latency-sensitive, bandwidth-sensitive, latency-bandwidth-sensitive, and latency-and-packet-loss-sensitive slices. The actual slice classification can be determined by different strategies based on the above conditions.
[0081] Each slice type contains several business slices. Each slice type selects the most stringent constraint from its respective business slices as the slice type constraint, and the allocation is as follows:
[0082] Slice type Business Slicing downlink bandwidth Uplink bandwidth Delay latency jitter Packet loss rate LaTh AR_VR 160Mbps 160Mbps 30ms / 3e-4 LaTh Email_Web 18Mbps 2Mbps 20ms / 0.08% LaLo live_game 2Mbps 2Mbps 50ms 100ms 0.1% LaTh live video (1080p) 10Mbps 10Mbps 50ms / 0.3% LaTh movie 45Mbps 20ms / 0.05% La music 128kbps 10kbps 100ms / 0.5% LaTh video_call 10Mbps 10Mbps 50ms / 0.5% Th Download 200Mbps 20Mbps 200ms / 1% La voice_call 64Kbps 64Kbps 150ms / 0.5%
[0083] Note: La—Latency Sensitive; Th—Throughput Sensitive; LaTh—Latency Throughput Sensitive; LaLo—Latency Loss Sensitive.
[0084] Step 3: Based on the directed graph and slice types of the network topology model, the bandwidth resources of each link in the network under test are divided into four slices according to the service requirements.
[0085] The data stream requesting communication can only use the resources of its own slice and cannot occupy the resources of other slices.
[0086] In this scenario, the bandwidth resources on each link are divided into four parts according to the division ratio, and each slice only occupies the bandwidth resources (bandwidth rate) of its own slice.
[0087] To accommodate the differences in data flows across different business types, this invention maps different data flows to their respective dedicated network slices (SLs). k ∈SL, the slice set is represented as SL. A network slice is a logical network on a physical network. The resources occupied by a slice are part of the physical link network, and the total resources occupied by all slices are equal to the physical link resources.
[0088] This invention will be implemented according to the business needs in proportion to λ k Allocate the bandwidth resources required by each slice, and the resources occupied by the k-th slice are represented as follows: b m λ is the bandwidth of the communication link. k The proportion allocated to the k-th slice; λ k ∈[0,1], and K represents the number of slices in the network under test.
[0089] Step 4: Design a reinforcement learning algorithm model for slice resource partitioning, which is used to partition the data stream of request communication according to slice resources and calculate the bandwidth resources of all links occupied by each type of slice.
[0090] The slice resource partitioning reinforcement learning algorithm model receives directed graph information of network topology and data flow information of communication requests as input, outputs the resource proportion of each slice, and adjusts the algorithm parameters of the agent according to the reward value.
[0091] Given the network resources for different slices, the path planning phase begins. For the data stream requesting communication, assuming the request stream corresponds to the k-th slice, it is represented as: v s ,v d Let V represent the source and destination nodes of the data stream requesting communication, respectively. Based on the network topology and the data stream request information, output the slice resource percentage λ. k ∈[0,1], the overall action is represented as: And receive rewards based on the optimized targets within that period.
[0092] The specific modeling process is as follows:
[0093] Step 401: First, allocate bandwidth resources reasonably to provide bandwidth resource capacity for each service, and calculate the average resource utilization rate of each link in all slices over a long period of time.
[0094] Calculate the m-th link of the k-th slice The bandwidth utilization rate is:
[0095] Link at time t Bandwidth usage:
[0096] Indicates whether there is a service flow on this link at time t. Selecting it as part of its path, if its value is 1, it means that there is a service flow passing through the link at time t, and its value is 0, which means no. Indicates that the link passes through at time t. The number of business flows; Indicates a request at time t In the link Throughput.
[0097] Step 402: Calculate the bandwidth resource utilization rate of the k-th slice on each link at time t.
[0098]
[0099] Where M represents the set of all links;
[0100] Step 403: Calculate the average bandwidth resource utilization of all slices on each link within time T:
[0101]
[0102] t represents a single time slice, T represents the overall time slice; k represents a single type of slice, and K represents all slices.
[0103] Step 404: Use the variance of average resource utilization to measure the difference in resource utilization across all slice links;
[0104] Represented as:
[0105] When allocating resources for network slices, this invention considers not only the average resource utilization rate mentioned above, but also the balance of resource usage among network slices. This is because maximizing source utilization alone will lead to excessively high utilization and load on some links, causing congestion throughout the network.
[0106] Step 405, calculate the average variance of all slices within time T using the variability in resource utilization:
[0107]
[0108] Step 406: Obtain the optimization objective function of the slice resource partitioning reinforcement learning algorithm model, and allocate resources reasonably to each slice to uniformly improve the resource utilization of all links;
[0109] The objective function is expressed as:
[0110] Step 5: Based on the resource allocation of the slice reinforcement learning algorithm model and the actual bandwidth resource constraints of the data stream requesting communication, calculate the candidate path set.
[0111] Based on the slice resource ratio and slice partitioning method output by the slice partitioning reinforcement learning algorithm model, the actual bandwidth resource constraints of the requested communication data stream are obtained. Candidate path sets are calculated using basic routing algorithms such as the shortest path algorithm and the widest path algorithm. First, for the k-th slice, based on the source node v... s and the destination node v d Find their respective neighbor nodes v s' ,v d' ∈V, forming the logical links of this slice, i.e. All logical links constitute the candidate path set.
[0112] Step 6: Each slice selects a path from the candidate path set to provide services for its respective data flow based on the routing algorithm, according to bandwidth resource constraints, network status and service requirements.
[0113] For the candidate path set, a routing strategy agent is designed to acquire current network state information, request communication data flow state information, and slice resource partitioning information, which are uniformly represented as follows: The network state includes the ratio of actual bandwidth used to allocated bandwidth, average packet loss latency, average throughput, average packet loss rate, and upper-layer resource allocation ratio for each type of network flow on each link. The data flow information for requested communication includes the source and destination addresses, bandwidth, latency, and jitter constraints. Within each slice, a discrete-scene reinforcement learning algorithm is used to generate a probability distribution of candidate paths based on the above state space. The path index with the highest probability value is selected as the decision action, represented as: The routing algorithm selects the appropriate routing path P for different request flows. k The requested data stream is then routed via path P. k Transmit the final route.
[0114] Step 7: Based on the selected path for each slice, and in combination with the requirements for latency, throughput, and packet loss rate of data flow requests, formulate a joint optimization objective for the user satisfaction model and routing strategy.
[0115] To improve the network experience quality of a single data stream, this invention focuses on user needs regarding network latency, throughput, and packet loss. Due to the differences in service types, different user communication requests have different requirements for network experience quality. Based on actual data and combining the latency, throughput, and packet loss rate requirements of data stream requests, this invention develops a user satisfaction model.
[0116] In the process of providing services to users, when network metrics fail to meet user needs, the user experience is poor; conversely, when network metrics improve, user satisfaction increases. However, once network metrics meet user needs, further increases in network metrics do not improve user experience but instead lead to excessive waste of network resources. Therefore, this invention selects appropriate paths for users based on their needs and the network resource availability within the network slice to improve user experience. The optimization objective of the routing strategy is to maximize the satisfaction of all users over a long period, expressed as:
[0117] The selected path for the k-th slice; N represents all choices. Let γ be the set of all services along the path, and let γ be a fixed constant. For the satisfaction model; it is represented as:
[0118]
[0119] β d β represents the weight of delay in the satisfaction model.th β represents the weight of bandwidth in the satisfaction model. l The weight of packet loss rate in the satisfaction model. This is the delay constraint corresponding to the slice type.
[0120] Data Stream Request The network requirements for latency, throughput, and packet loss rate at time t are respectively expressed as: and At time t, on the selected path The actual packet loss, throughput, and packet loss rate obtained are expressed as follows: and
[0121] Step 8: Using an iterative calculation method based on a two-layer model for joint optimization of network resources and routing, the optimal strategy for the joint optimization objective is solved.
[0122] Based on the dynamic characteristics of network slicing and routing, this invention constructs an upper-layer network resource allocation and scheduling model and a lower-layer routing and scheduling model. To address the mutual constraints between the two models, a joint optimization model based on bi-layer programming is constructed. By constructing a cyclical interaction model between the upper-layer network resource allocation model and the lower-layer network path planning model, each cyclical interaction forms an optimization process for the joint action strategy. Through continuous iteration, the strategies of the upper and lower layer models are gradually optimized, and finally, the optimal strategy of this joint optimization model is solved.
[0123] The specific process is as follows:
[0124] Step 801: Utilize slice resources to divide the reinforcement learning algorithm model and the user satisfaction model, and build an optimization model based on bi-layer programming for the upper and lower layers;
[0125] The upper layer performs slicing, with the ultimate goal of maximizing network resource utilization. The decision results constrain the routing decisions of the lower layer.
[0126] The specific formula is as follows:
[0127] The lower layer performs routing selection and optimization, aiming to optimize various indicators of the current data flow and maximize service satisfaction; its decision results also affect the calculation of the target network resource utilization rate in the upper layer.
[0128] The specific formula is as follows:
[0129] The specific form of the model is as follows:
[0130] The nested influence of the two optimization objectives and actions conforms to a bilevel programming problem model, where the upper-level optimization problem is:
[0131] Its decision variable is x; and the lower-level optimization problem is: Its decision variable is y.
[0132] Step 802: Use two reinforcement learning models to make decisions at the upper and lower layers respectively, and achieve joint optimization of the two-layer problem through continuous iteration using hierarchical reinforcement learning.
[0133] The decision of a single model affects the computation of another model, changing the optimization strategy of the other model. Through continuous iteration, the joint optimization of the two-layer problem is achieved, which solves the problem with the idea of hierarchical reinforcement learning.
[0134] At the upper layer, the network slicing agent continuously makes optimal slicing actions based on the network's resource status, according to policy π. u Allocating network resources. Based on the Markov property of network resource allocation, this invention establishes a Markov process for resource allocation. Within a specified time slot t, when a new data stream is generated, the agent first receives current state information, including the ratio of the actual bandwidth occupied by each type of network stream on each link to the allocated bandwidth, as well as the differences between different slices, expressed as: The agent executes actions based on its state. These actions represent the resource allocation proportions of a slice, expressed as follows: And receive rewards based on the optimized targets within that period.
[0135] To enable the upper-level SAC agent to make optimal actions based on the network resource status, the algorithm continuously optimizes its resource allocation strategy, which is represented by π. u Then at time t, its state is According to strategy π u Make In different time slots, the upper-layer intelligent agent—the SDN controller module—continuously adjusts its settings according to the network resource status and follows policy π. u Network resources are allocated and fed back to the lower-level agents. The agents in each lower-level slice adjust their routing strategies according to the allocated resources.
[0136] At the lower layer, the agent selects different transmission paths for the request flow based on the current network environment. The routing strategy selects different paths based on the state of network resources, causing the network state to transition to the next state, exhibiting Markov properties; therefore, a Markov process for routing selection can be established. At any time t in a given slice, after the agent gives a network request, it first receives the current network state and the request data flow information, uniformly represented as... Network status includes the ratio of actual bandwidth used to allocated bandwidth, average packet loss latency, average throughput, average packet loss rate, and upper-layer resource allocation ratio for each type of network flow on each link. Request flow information includes the source and destination addresses of the flow, bandwidth, latency, and jitter constraints.
[0137] This invention calculates the user satisfaction value for each point-to-point path based on the aforementioned network state information. Then, it calculates the top k paths from the source node to the destination node with the smallest sum of user satisfaction values (weights) on each edge, generating a candidate path set. Based on the aforementioned state space, the lower-level agent generates a probability distribution of candidate paths and selects the path index with the highest probability value as the decision action, represented as: After selecting a path, the lower-layer agent calculates user satisfaction, i.e., its optimization target, based on actual operational conditions. This process is implemented using a discrete-state SAC agent. To achieve joint optimization of resources and routes for multiple types of services, this invention employs an interactive training method between upper and lower-layer agents. The specific process is as follows: Figure 5 As shown.
[0138] Based on the actual network status information, the lower-level traffic satisfaction model calculates the traffic satisfaction rate. The routing strategy is optimized, and at the same time, the upper-layer algorithm aims to maximize network utilization while also calculating the agent's reward value. Optimize slicing strategy.
[0139] During the iterative training process between the upper and lower layers, based on the resource allocation provided by the upper-layer agent, the lower-layer agent updates its routing strategy according to the requests of different flows and feeds back the current network status and the service requirements of each slice to the upper layer. The upper-layer agent then reallocates resources based on the current network status and the requirements reported by the lower layer, and the lower layer makes routing decisions again. This cycle continues until the rewards of both the upper and lower layer agents continuously reach their maximum. At this point, the optimal strategy of the joint optimization method is obtained, meaning that resources are optimally allocated and the routing strategy is also optimal.
[0140] Considering the dynamic changes in network requirements and network environment, this invention continuously monitors the network status, and triggers an iterative mechanism between upper and lower layers when the status changes exceed a set threshold.
[0141] In the above two-level optimization objectives, network resource utilization is determined by setting its decision variable λ. k To maximize the value of the service forwarding path, the selection of the service forwarding path is achieved through its decision variable P. k To maximize user satisfaction. Meanwhile, the parameter P involved in network resource utilization... k The goal is to optimize the solution that yields the highest level of user satisfaction across all users. However, the optimal solution for all user satisfaction levels is influenced by the higher-level decision variable λ. kThe interaction between these two optimization objectives and their corresponding actions is as follows: the lower layer selects different paths for different service flows based on the bandwidth resources allocated by the upper layer; the upper layer's optimization objective is influenced by the path selection of the lower layer, affecting the resource allocation ratio. These two optimization objectives and actions are nested and influence each other, forming a typical two-layer programming problem. To improve resource utilization, network providers will not allocate excessive network resources; however, to improve user satisfaction, they need to maximize the bandwidth resources of candidate paths during path selection. In other words, the two optimization objectives will not make decisions that benefit each other, making it a pessimistic two-layer programming problem.
[0142] This problem involves a large state space and numerous decision-making actions. Therefore, it employs an iterative approach, using two reinforcement learning models to make decisions at the upper and lower layers respectively, thus solving the problem through a hierarchical reinforcement learning strategy. In the upper-layer network resource partitioning module, after obtaining the current network state information and request flow state information, this invention uses a reinforcement learning algorithm to output a numerical value λ for each slice. k Resource allocation is performed for each slice, where the resource proportion of each slice is λ. k During the lower-layer routing strategy optimization process, after obtaining the current network state information, request flow state information, and upper-layer resource partitioning information, this invention uses a discrete scenario reinforcement learning algorithm within each slice to select a suitable routing path P for different request flows. k The requested data stream is then routed via path P. k For the final route, transmission is performed, network state is updated, and upper and lower layer reinforcement learning agents calculate the feedback of the current network state and optimize their own decisions to achieve continuous and iterative joint optimization of network resource allocation and path planning.
[0143] This invention uses a two-layer interactive reinforcement learning model to solve the aforementioned two-layer planning model. In the reinforcement learning mode, each agent takes action by observing the state and continuously optimizes the learned strategy based on historical experience data. The upper-layer agent allocates slice resources according to network load conditions, and the lower-layer agent adjusts its routing strategy in a timely manner according to the resource allocation of the upper-layer network. The iterative updates of the upper and lower-layer agents eventually reach a state where the upper-layer agent can reasonably allocate slice resources and the lower-layer agent selects a routing strategy that satisfies the user, thereby improving user satisfaction. To achieve joint optimization of resources and routes for multiple types of services, this invention adopts an interaction model between the upper and lower-layer agents and the network, as follows: Figure 4 As shown.
[0144] The flowchart of the two-layer interactive reinforcement learning implementation steps is as follows: Figure 6As shown in the flowchart, this invention explains how network resource allocation and path planning are jointly optimized in a multi-network slice network. Repeating the above steps constitutes the specific implementation process of the invention. When facing multiple service flows with different constraints, this invention increases network bandwidth utilization and reduces communication latency, packet loss rate, and jitter rate.
Claims
1. A network resource partitioning and path planning joint optimization method based on double-layer programming, characterized in that, The specific steps are as follows: Step one, obtain the topology information of the network to be tested, use each link and forwarding node to construct a directed graph of the communication network topology based on SDN; Step two, according to network modeling, 5G network communication traffic characteristics and QoS quantitative indicators, classify all communication traffic of the network to be tested into four slice types; The four slice types are delay-sensitive, bandwidth-sensitive, delay-bandwidth-sensitive and delay-packet loss-sensitive slices; Step three, for the directed graph of network topology modeling and the classified slice types, the bandwidth resources of each link in the network to be tested are divided into the occupied resources of the four slices according to the proportion of the business demand; Each slice only occupies the bandwidth resources of this slice; The division method is as follows: The resource occupied by the kth slice is expressed as: b m λ is the bandwidth of the communication link k is the proportion allocated to the kth slice; λ k ∈[0,1], and K represents the number of slices in the network under test; Step four, design a slice resource division reinforcement learning algorithm model for dividing the data flow according to the slice resources and calculating the bandwidth resources of all links occupied by each type of slice; The specific modeling process is as follows: Step 401, calculate the bandwidth utilization rate of the mth link of the kth slice is: Bandwidth usage of the link at time t: Bandwidth usage of the link at time t: indicates whether there is a traffic flow going through the link at time t selected as a component of its path, if its value is 1, it indicates yes, and if its value is 0, it indicates no; indicates the number of traffic flows going through the link at time t; indicates the throughput requested on the link at time t; Step 402, count the bandwidth resource utilization rate of the kth slice on each link at time t Where M is the set of all links; Step 403, calculate the average bandwidth resource utilization of all slices in each link within time T: Step 404, use the variance of the average resource utilization to measure the difference in resource utilization of all slice links; denoted as: Step 405, use the difference in resource utilization to calculate the average variance of all slices within time T: Step 406, obtain the optimization objective function of the slice resource division reinforcement learning algorithm model, which is to reasonably allocate resources to each slice to uniformly improve the resource utilization of all links; The optimization objective function is expressed as: Step five, according to slice resource division reinforcement learning algorithm model and request communication data flow of the actual bandwidth resource constraints, calculate candidate path set Step six, each slice selects a path from the candidate path set based on the bandwidth resource constraint, network state and business demand to provide services for its own data flow based on the routing algorithm; Step seven, based on the selected path of each slice, combined with the delay, throughput occupancy and packet loss rate requirements of the data flow request, formulate a joint optimization objective of user satisfaction model and routing strategy; The optimization goal of the routing strategy is to maximize the long-term user satisfaction, which is expressed as: selected path for the kth slice; N is the number of all selections to γ is a constant; is the satisfaction model; represents: β d is a proportion weight of the delay in the satisfaction model, β th is a proportion weight of the bandwidth in the satisfaction model, β l is a proportion weight of the packet loss rate in the satisfaction model, is a delay constraint corresponding to the slice type; Data stream request The network's requirements on latency, throughput and packet loss rate at time t are denoted as and The actual obtained packet loss, throughput and packet loss rate on the selected path at time t are denoted as and and Step eight, use the iterative calculation method of the network resource and routing joint optimization double-layer model to solve the optimal strategy of the joint optimization objective; The specific process is as follows: Step 801, use the slice resource division reinforcement learning algorithm model and the user satisfaction model to build a double-layer planning optimization model in the upper and lower layers; The upper layer performs slice division, and the ultimate goal is to maximize the resource utilization of the network; The specific formula is: The lower layer performs routing selection and optimization, and the goal is to optimize each indicator of the data flow and maximize the service satisfaction rate; The specific formula is Step 802, use two reinforcement learning models to make decisions in the upper and lower layers respectively, and achieve joint optimization of the double-layer problem through continuous iteration in a layered reinforcement learning manner; The upper-layer algorithm of the model can make optimal slice division actions according to the network resource state, and the policy is represented as π u Then at time t, its state According to the policy π u Make action The SDN controller module is configured to determine the action based on the received information allocating network resources and feeding back to the lower agent; The lower layer agent calculates the specific routing information p of the data flow according to the allocated network resources, and then requests data flow communication; According to the network actual state information, the traffic satisfaction degree model of the lower layer calculates traffic satisfaction degree Optimizing routing strategy, at the same time, the upper layer algorithm target maximizes network utilization rate and also calculates the agent reward value Optimizing slice strategy; The upper layer agent allocates resources again according to the current network state and the demand feedback from the lower layer, and the lower layer makes routing decisions again, and the cycle continues until the rewards of the upper and lower layer agents continuously reach the maximum, and then the optimal strategy of the joint optimization method is obtained, that is, the optimal resource allocation and the optimal routing strategy.
2. The network resource partitioning and path planning joint optimization method based on double-layer programming according to claim 1, wherein, The attributes of the link include two forwarding nodes connecting the link, bandwidth capacity and transmission rate; the attributes of the forwarding nodes include the locations of the nodes; The directed graph of the communication network topology refers to a directed graph G = (V, E) that includes a set of links and a set of forwarding nodes, where each edge E represents a set of links in the network topology and is uniquely identified by E. i This represents each link in the network under test, and point V represents the set of forwarding nodes in the network topology; a unique identifier V is used. i Indicates a forwarding node in the network; e m =(v i ,v j )∈E represents a node v in a directed graph. i and v j The nodes are connected, and there is a communication link between the forwarding nodes corresponding to the two nodes.
3. The network resource partitioning and path planning joint optimization method based on double-layer programming according to claim 1, wherein, Among the four slice types, each slice type has a plurality of service slices, and each slice type selects the most stringent constraint condition from the respective service slices as the slice type constraint of the type.
4. The network resource partitioning and path planning joint optimization method based on double-layer programming of claim 1, wherein, The step five is specifically: firstly, for the kth slice, according to the source node v s and the destination node v d , find the respective neighbor nodes v s' ,v d' ∈V, which constitute the logical link of the slice, that is all the logical links constitute the candidate path set
Citation Information
Patent Citations
Network slicing-based SDN joint routing and resource allocation method
CN108206790A
Method for designing spatial information network routing strategy under SDN architecture
CN110493131A