A network load balancing method, device and electronic equipment
By decomposing the load balancing task of mobile communication networks into sub-tasks of hierarchical reinforcement learning, configuring corresponding parameters and combining optimal strategies, the load imbalance problem caused by the hierarchical complexity of mobile communication networks is solved, and global load optimization is achieved.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- CHINA MOBILE GRP HENAN CO LTD
- Filing Date
- 2021-07-26
- Publication Date
- 2026-08-04
AI Technical Summary
Mobile communication networks are complex and have a wide coverage area. Adjusting load balancing at only one level cannot guarantee that the upper or lower layers of the network will also improve, and may even increase the network burden.
The network load balancing task is decomposed into hierarchical reinforcement learning subtasks. A state set, policy set, action set, termination state set, and reward function are configured. The relatively optimal policy at each level is obtained through hierarchical reinforcement learning and combined into a global policy for load adjustment.
It achieves load balancing at different granular levels, avoiding the phenomenon that local load problems lead to excessive load in other areas, and optimizes network load distribution.
Smart Images

Figure CN115701175B_ABST
Abstract
Description
Technical Field
[0001] This document relates to the field of artificial intelligence technology, and in particular to a network load balancing method, device, and electronic device. Background Technology
[0002] Currently, mobile operators primarily address the issue of insufficient mobile communication network capacity through capacity expansion or load balancing adjustments. In the context of the call for "speeding up and reducing fees," load balancing adjustments are receiving more attention from mobile operators compared to capacity expansion.
[0003] Currently, load balancing adjustments for mobile communication networks remain quite challenging. This is primarily because mobile communication networks are complex and have a wide coverage area. Adjusting the load balancing of a mobile communication network at only one level cannot guarantee that the upper or lower layer networks will also improve, and may even increase the burden on the upper or lower layer networks.
[0004] Therefore, how to correlate and weigh different levels of granularity in order to adjust the network load balance is the technical problem that this application addresses. Summary of the Invention
[0005] The purpose of this invention is to provide a network load balancing method, apparatus, and electronic device that can associate and weigh different levels of granularity to adjust the network load balancing.
[0006] To achieve the above objectives, the embodiments of the present invention are implemented as follows:
[0007] Firstly, a network load balancing method is provided, including:
[0008] The task of performing load balancing on the target network is decomposed into sub-tasks of hierarchical reinforcement learning according to the hierarchy of the target network, wherein the target network is pre-divided into multiple hierarchical levels.
[0009] Parameters for hierarchical reinforcement learning are configured for each subtask, including: a state set, a policy set, an action set, a termination state set, a state transition function, and a reward function. The states in the state set are the load balancing metric values for the corresponding level of the subtask; the termination states in the termination state set are the load balancing metric values for the corresponding level of the subtask that meet the load balancing criteria; the policies in the policy set are the load adjustment policies for the corresponding level of the subtask; the actions in the action set are the actions generated by executing the policies in the policy set; the state transition function represents the probability of performing an action in one state to enter another state; and the reward function represents the instantaneous reward value obtained by performing an action in one state to enter another state.
[0010] Hierarchical reinforcement learning is performed on the subtasks corresponding to the layers in the target network to obtain the relatively optimal policies of at least two layers of subtasks, and the determined relatively optimal policies of the at least two layers of subtasks are combined into a global policy for load balancing of the target network.
[0011] Based on the global strategy, load adjustment is performed on the target network.
[0012] Secondly, a network load balancing device is provided, comprising:
[0013] The task creation module is used to decompose the task of performing load balancing for the target network into sub-tasks of hierarchical reinforcement learning according to the hierarchy of the target network, wherein the target network is pre-divided into multiple hierarchical levels.
[0014] The parameter configuration module is used to configure the parameters for hierarchical reinforcement learning for each subtask, including: a state set, a policy set, an action set, a termination state set, a state transition function, and a reward function. The states in the state set are the load balancing metric values for the corresponding level of the subtask; the termination states in the termination state set are the load balancing metric values for the corresponding level of the subtask that meet the load balancing criteria; the policies in the policy set are the load adjustment policies for the corresponding level of the subtask; the actions in the action set are the actions generated by executing the policies in the policy set; the state transition function represents the probability of performing an action in one state to enter another state; and the reward function represents the instantaneous reward value obtained by performing an action in one state to enter another state.
[0015] The hierarchical reinforcement learning module is used to perform hierarchical reinforcement learning on the sub-tasks corresponding to the layers in the target network, obtain the relative optimal policies of at least two layers of sub-tasks, and combine the determined relative optimal policies of the at least two layers of sub-tasks into a global policy for load balancing for the target network.
[0016] The load adjustment module is used to adjust the load of the target network based on the global policy.
[0017] Thirdly, an electronic device is provided, comprising: a memory, a processor, and a computer program stored in the memory and executable on the processor, the computer program being executed by the processor.
[0018] The task of performing load balancing on the target network is decomposed into sub-tasks of hierarchical reinforcement learning according to the hierarchy of the target network, wherein the target network is pre-divided into multiple hierarchical levels.
[0019] Parameters for hierarchical reinforcement learning are configured for each subtask, including: a state set, a policy set, an action set, a termination state set, a state transition function, and a reward function. The states in the state set are the load balancing metric values for the corresponding level of the subtask; the termination states in the termination state set are the load balancing metric values for the corresponding level of the subtask that meet the load balancing criteria; the policies in the policy set are the load adjustment policies for the corresponding level of the subtask; the actions in the action set are the actions generated by executing the policies in the policy set; the state transition function represents the probability of performing an action in one state to enter another state; and the reward function represents the instantaneous reward value obtained by performing an action in one state to enter another state.
[0020] Hierarchical reinforcement learning is performed on the subtasks corresponding to the layers in the target network to obtain the relatively optimal policies of at least two layers of subtasks, and the determined relatively optimal policies of the at least two layers of subtasks are combined into a global policy for load balancing of the target network.
[0021] Based on the global strategy, load adjustment is performed on the target network.
[0022] Fourthly, a computer-readable storage medium is provided, on which a computer program is stored, the computer program performing the following steps when executed by a processor:
[0023] The task of performing load balancing on the target network is decomposed into sub-tasks of hierarchical reinforcement learning according to the hierarchy of the target network, wherein the target network is pre-divided into multiple hierarchical levels.
[0024] Parameters for hierarchical reinforcement learning are configured for each subtask, including: a state set, a policy set, an action set, a termination state set, a state transition function, and a reward function. The states in the state set are the load balancing metric values for the corresponding level of the subtask; the termination states in the termination state set are the load balancing metric values for the corresponding level of the subtask that meet the load balancing criteria; the policies in the policy set are the load adjustment policies for the corresponding level of the subtask; the actions in the action set are the actions generated by executing the policies in the policy set; the state transition function represents the probability of performing an action in one state to enter another state; and the reward function represents the instantaneous reward value obtained by performing an action in one state to enter another state.
[0025] Hierarchical reinforcement learning is performed on the subtasks corresponding to the layers in the target network to obtain the relatively optimal policies of at least two layers of subtasks, and the determined relatively optimal policies of the at least two layers of subtasks are combined into a global policy for load balancing of the target network.
[0026] Based on the global strategy, load adjustment is performed on the target network.
[0027] The solution of this invention decomposes the task of load balancing the target network into sub-tasks of hierarchical reinforcement learning according to the hierarchy of the target network. Through hierarchical reinforcement learning, the sub-tasks of different levels can solve the load adjustment strategy relatively optimally in a smaller sub-problem space. The relatively optimal strategies obtained from at least two levels are combined into a global strategy, thereby adjusting the load of the target network based on the global strategy. This avoids the phenomenon of only solving the load problem in a local area of the target network, which leads to excessive load in other areas. Attached Figure Description
[0028] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments recorded in the embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0029] Figure 1 This is a flowchart illustrating the network load balancing method provided in an embodiment of the present invention.
[0030] Figure 2 This is a schematic diagram illustrating how the network load balancing method provided in this embodiment of the invention decomposes the task of performing load balancing on a mobile communication network.
[0031] Figure 3 This is a schematic diagram of the hierarchical reinforcement learning process provided in an embodiment of the present invention.
[0032] Figure 4 This is a schematic diagram of the network load balancing device provided in an embodiment of the present invention.
[0033] Figure 5 This is a schematic diagram of the structure of an electronic device provided in an embodiment of the present invention. Detailed Implementation
[0034] To enable those skilled in the art to better understand the technical solutions in this specification, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this specification, and not all embodiments. Based on the embodiments in this specification, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of this specification.
[0035] As mentioned earlier, under the general trend of "speeding up and reducing fees," mobile operators are increasingly using load balancing to address the insufficient capacity of mobile communication networks. In the communications field, mobile communication networks are complex and have a wide coverage area. Load balancing adjustments at only one layer cannot guarantee improvements to the upper or lower layers; it may even increase the burden on those layers. Therefore, this document aims to provide a technical solution that can correlate and balance different layers to achieve network load balancing. The networks discussed here are not limited to mobile communication networks.
[0036] Figure 1 This is a flowchart of a network load balancing method according to an embodiment of the present invention, including the following steps:
[0037] S102 will perform a load balancing task on the target network, which will be decomposed into hierarchical reinforcement learning subtasks according to the layers of the target network. The target network is pre-divided into multiple layers.
[0038] Hierarchical reinforcement learning decomposes the overall task into sub-tasks at different levels, allowing each sub-task to solve for a strategy in a smaller sub-problem space, and finally combining them into a global strategy.
[0039] In this embodiment of the invention, the layers that have been divided in the target network can be directly used as the layers of hierarchical reinforcement learning. That is, by using hierarchical reinforcement learning, the task of performing load balancing on the target network is subdivided into sub-tasks corresponding to different layers, so as to request load adjustment strategies for sub-tasks at different layers, and finally combine them into a global strategy.
[0040] This article does not specifically limit the hierarchical division of the target network. As an example, if the target network is a mobile communication network, the layers from top to bottom can be: network-wide layer (relative to the target network), pool layer, network element layer, and link / virtual machine layer. Each layer is not limited to containing only one subtask. Taking the network element layer as an example, it can include multiple network elements, and each network element can correspond to a subtask, which is the task of performing load balancing for its corresponding network element.
[0041] S104, configure the parameters for hierarchical reinforcement learning for each subtask, including: state set, policy set, action set, termination state set, state transition function, and reward function. The states in the state set are the load balancing index values of the corresponding level of the subtask. The termination states in the termination state set are the load balancing index values of the corresponding level of the subtask that meet the load balancing criteria. The policies in the policy set are the load adjustment policies of the corresponding level of the subtask. The actions in the action set are the actions generated by executing the policies in the policy set. The state transition function represents the probability of performing an action in one state to enter another state. The reward function represents the instantaneous reward value obtained by performing an action in one state to enter another state.
[0042] S106, perform hierarchical reinforcement learning on the subtasks corresponding to the layers in the target network to obtain the relatively optimal policies of at least two layers of subtasks, and combine the determined relatively optimal policies of the at least two layers of subtasks into a global policy for load balancing of the target network.
[0043] In this process, hierarchical reinforcement learning solves for relatively optimal policies for sub-tasks, thereby associating load adjustment policies at different levels. It should be noted that this embodiment of the invention does not require determining the relatively optimal policy for each level in the target network based on hierarchical reinforcement learning.
[0044] S108 adjusts the load on the target network based on a global policy.
[0045] The method of this invention decomposes the task of load balancing the target network into sub-tasks of hierarchical reinforcement learning according to the hierarchy of the target network. Through hierarchical reinforcement learning, the sub-tasks of different levels can solve the load adjustment strategy relatively optimally in a smaller sub-problem space. The relatively optimal strategies obtained from at least two levels are combined into a global strategy, thereby adjusting the load of the target network based on the global strategy. This avoids the phenomenon of only solving the load problem in a local area of the target network, which leads to excessive load in other areas.
[0046] The hierarchical reinforcement learning process in the method of this invention will be described in detail below, taking into account practical application scenarios.
[0047] This application scenario performs load balancing on a 5G mobile communication network. For example... Figure 2 As shown, the task of performing load balancing in mobile communication networks can be decomposed into sub-tasks of hierarchical reinforcement learning at the network-wide level, pool level, network element level, and link / virtual machine level.
[0048] The subtasks at the entire network level of the mobile communication network are used as the root tasks of the hierarchical reinforcement learning. Here, the termination condition of the hierarchical reinforcement learning can be that the subtasks at the entire network level reach the termination state, which means that the hierarchical reinforcement learning can ensure that the load at the entire network level reaches the balance standard.
[0049] Before implementing load balancing, load data can be collected at different levels and preprocessed. Specific steps include: collecting user capacity utilization and N3 / N6 interface traffic bandwidth utilization at the network element level, CPU and memory utilization of each virtual machine (VM), and signaling load of each link within a Virtual Network Feature (VNF) network element. Then, the collected load data undergoes preprocessing, specifically including: data cleaning to remove load data that does not conform to the actual network conditions, such as data with daily fluctuations in user capacity utilization exceeding 20% or daily fluctuations in N3 / N6 interface traffic bandwidth utilization exceeding 40%; null value interpolation (using median, mean, or mode); and calculating pool-level user capacity utilization and N3 / N6 interface traffic bandwidth utilization based on the pool networking relationship of each VNF network element in the 5G core network. This provides data for subsequent acquisition of the optimal load balancing strategy based on hierarchical reinforcement learning.
[0050] To ensure the safe and stable operation of the 5G core network, the goals of load balancing optimization are defined as follows: user capacity utilization at both the POOL and network element levels within the 5G core network must be below a warning threshold (e.g., set to 80%); N3 / N6 interface traffic bandwidth utilization must be below a warning threshold (e.g., 45% for load sharing in disaster recovery mode, and 80% for primary / backup mode); VM-level CPU and memory utilization must be below a warning threshold (e.g., 80%); port / link-level traffic bandwidth utilization of network elements such as switches must be below a warning threshold (e.g., set to 60%); and load balancing at the POOL and network element levels must be below a certain threshold (e.g., set to 10%). The load balancing β at a certain level of the 5G network (e.g., POOL and network element) is defined as... l The definition is shown in equation (1):
[0051]
[0052] In equation (1), y represents the average resource utilization rate of the hierarchy. L,m This represents the resource utilization rate of the m-th entity in the hierarchy, and N represents the total number of entities in the hierarchy.
[0053] Here, we take the pool level as an example. This represents the average resource utilization rate at the pool level (resources such as bandwidth, CPU, and memory are not limited here), y L,m This represents the resource utilization rate of the m-th network element in the pool hierarchy. This represents the average resource utilization rate of all network elements in the pool level, where N represents the total number of network elements in the pool level. For example, y L,m = N3 interface flow rate of the m-th network element within the POOL / N3 interface bandwidth rate * 100%.
[0054] To obtain the optimal load balancing strategy for each layer of the 5G network and improve the learning speed, a 5G network load balancing method based on the production and operation experience of the 5G core network, using the MAXQ (a typical hierarchical reinforcement learning) approach, is proposed. The corresponding process includes:
[0055] S202: Decompose the entire task W of the load balancing strategy into n sub-task sets (W1, W2, ..., Wn). n For example, based on prior knowledge, the load balancing of the entire network is taken as the root task, the second layer is the load balancing sub-task at the pool level, and the third layer is the load balancing sub-task between network elements within the pool. The entire hierarchical task is specifically as follows: Figure 3 As shown. Furthermore, the policy π is decomposed into a set of policies (π1, π2, ..., π). n ), where π i It is a subtask w i The strategy. For each subtask w i Using a six-tuple i A i ,P i ,R i ,F i ,π i > indicates that, where i represents the subtask number, a represents the action, a′ represents the currently selected action, s represents the current state, and s′ represents the other state entered by performing the action in the current state. i For subtask W i State sets (such as user capacity utilization and bandwidth utilization of a network element), A i For subtask W i Action set (can be basic actions or compound actions; basic actions include load migration in or out operations for a specific network element / VM / link; compound actions include root task W0 selecting to execute POOL-level load balancing actions, etc.), P i For subtask W i State transition probability (e.g., P) i (s′,τ|s,a) represents the probability that performing action a in state s for time τ will lead to a transition to state s′. π(s′) represents the strategy of state s′. For example, if the current state s is a network element with a capacity utilization rate of 60%, performing action a to migrate load to that network element will lead to a state s′ with a capacity utilization rate of 61%. τ is the duration of the subtask performed in state s. R i For subtask W i The reward function, where for each subtask W i W i The immediate reward r obtained by performing action a is R. i (s,a) = V(a,s). The goal of learning is to obtain an optimal policy π. * This maximizes the reward value that the action 'a' receives from the environment, as shown in formulas (2) and (3); F i For subtask W i The set of termination states (for example, in this invention, the threshold of load balancing at the POOL and network element levels being lower than 10%, and the load of capacity utilization at the POOL and network element levels being lower than the warning value are taken as the termination states of the corresponding layers, i.e., when state s∈F i When the subtask W ends i ),and Where S is the state set of the entire task (such as the load state of 5G network equipment at the POOL level, network element level, VM level, and link level), and A is the set of actions that can be executed by the entire task (including composite actions and basic actions).
[0056] π * =arg π max V π (s), s∈S (2)
[0057]
[0058] In equations (2) and (3), γ is the discount factor, ranging from [0,1], u represents the random jump coefficient, and r t This means that performing action a results in the t-th state s. t Transition to state s (t+1) t+1 The instantaneous reward value obtained afterwards, V π (s tThe sum of accumulated instantaneous reward values is represented by ). For example, in this invention, the immediate reward value for each load migration in or out operation (note: in this invention, each migration in or out operation is uniformly set to 1% of the original traffic volume of the network element / VM / link) is -1; when the state satisfies that the bandwidth utilization rate at the port / link level is less than 60% and the traffic balance of each link within the same network element is less than 10%, the immediate reward value is set to +10; when the CPU utilization rate and memory utilization rate of all VMs in each network element are less than 80%, the immediate reward value is set to +20; when the bandwidth utilization rate of each load-sharing network element within the POOL is less than 45% and the traffic balance of each network element within the POOL is less than 10%, the immediate reward value is set to +30; when the bandwidth utilization rate of each load-sharing POOL is less than 45% and the traffic balance between POOLs is less than 10%, the immediate reward value is set to +40. The value function V... π (s t ) represents the sum of accumulated immediate reward values. Therefore, the essence of searching for the optimal strategy is to maximize the reward value obtained by the execution environment a from the state.
[0059] For each subtask W i W i The immediate reward for performing action 'a' is R. i (s, a) = V(a, s), and subtask W i The corresponding state value function equation is:
[0060] V(i,s)=V(a,s)+∑ s′,τ P i (s′,τ|s,a)γ τ V(i,s′) (4)
[0061] Where V(i,s′) is the state from which subtask W is completed starting from state s′ (the state corresponding to the end of subtask a). i The expected return value.
[0062] Subtask W i In this context, the value function of a state-action pair is defined as:
[0063] Q(i,s,a)=V(a,s)+∑ s′,τ P i (s′,τ|s,a)γ τ maxQ[i,s′,π(s′)] (5)
[0064] The second term in equation (5) is called the completion function C(i, s, a), i.e.:
[0065] C(i, s, a)=∑ s′,τ P i(s′,τ|s,a)γ τ maxQ[i, s′, π(s′)] (6)
[0066] Through MAXQ value function decomposition, the value function of the state-action pair can be divided into two parts: the immediate reward V(a, s) and the completion function C(i, s, a), as shown in equation (7).
[0067] Q(i,s,a)=V(a,s)+C(i,s,a) (7)
[0068] The task of hierarchical reinforcement learning is to determine the W for each subtask. i The optimal strategy π * (i,s), that is, according to equation (8), the corresponding action a is selected, and all subtasks form a hierarchical task structure with W0 as the root node, that is Figure 2 As shown.
[0069]
[0070] Since solving the root task W0 also solves the entire learning task W, in order to solve W0, it is necessary to sequentially select and call the basic action or other subtasks.
[0071] Given the hierarchical strategy π of the task, assume that the root task W0's strategy selects a1 to execute subtasks, a1's strategy selects subtask a2, and so on, until a... n-1 The strategy selected basic action a n Then the value function V of state s in root task W0 π (0,s) can be decomposed into:
[0072] V π (0,s)=V π (a n ,s)+C π (a n-1 ,s,a n )+L+C π (a1,s,a2)+C π (0,s,a1) (9)
[0073] in,
[0074] V π (a n ,s)=∑ s′ P(s′|s,a n R(s′|s,a) n (10)
[0075] Equation (9) is the basis for hierarchical reinforcement learning using the MAXQ algorithm in this invention.
[0076] Specifically, in this application scenario, the optimal policy selection algorithm for subtasks based on hierarchical reinforcement learning selects subtasks and solves for their relatively optimal policies until the relatively optimal policy of the current subtask enables the top-level subtask to reach its termination state. The termination conditions for solving the relatively optimal policy include the current subtask reaching its termination state and the current subtask's relatively optimal policy involving actions related to the next level. The corresponding process is as follows: Figure 3 The following are included:
[0077] S304: Begin learning the optimal load balancing strategy, and set up subtask W. i The initial values are: i = 0, j = 0, where j is the hierarchical number of the subtask being processed. Initially, the load balancing root task W0 of the entire network is selected, and the process proceeds to step S306.
[0078] S306: Determine if the current state has reached the termination state of the current subtask? If the current state is the termination state of the current subtask, proceed to step S316; if the current state is not the termination state of the current subtask, proceed to step S308.
[0079] S308: Select the action a corresponding to the optimal strategy according to formula (8) (it can be a basic action or a compound action. A basic action is such as performing load migration in or out for a certain network element / VM / link. A compound action is such as the root task W0 selecting to execute a POOL-level load balancing subtask, etc.), and then proceed to step S310.
[0080] S310: Determine whether the selected action a is a compound action? If action a is a compound action, proceed to step S312; if action a is a basic action, proceed to step S314.
[0081] S312: Update i based on action a, and record.<i,s,a> Given information such as the value function, let j = i, and then proceed to step S306.
[0082] S314: Execute the basic action a, record the immediate reward R(s,a) and the subsequent state s′; and save the result.<i,s,a> The information such as the value function Q(i,s,a) is obtained, and then the process proceeds to step S316.
[0083] S316: Determine if the current state is the termination state of the root task W0 of the entire network? If the current state is the termination state of the root task, proceed to step S318; if the current state is not the termination state of the root task, proceed to step 306.
[0084] S318: End the load balancing adjustment task.
[0085] Therefore, this application scenario is based on the MAXQ hierarchical reinforcement learning method. By establishing multiple hierarchical subtasks that can be learned in parallel (such as POOL layer, network element layer, VM layer and link layer), and by using the hierarchical structure to constrain the search range of the optimization strategy, it is possible to quickly obtain the optimal strategy for load balancing at each level of the 5G core network, such as POOL, network element, VM and link.
[0086] The above application scenarios are exemplary descriptions of the methods in the embodiments of the present invention. It should be understood that appropriate changes can be made without departing from the principles described above, and these changes should also be considered within the scope of protection of the embodiments of the present invention.
[0087] In addition, corresponding to Figure 1 The method shown in this embodiment of the invention also provides a network load balancing device. Figure 4 This is a schematic diagram of the network load balancing device 400 according to an embodiment of the present invention, including:
[0088] The task creation module 410 is used to decompose the task of performing load balancing for the target network into sub-tasks of hierarchical reinforcement learning according to the hierarchy of the target network, wherein the target network is pre-divided into multiple hierarchical levels.
[0089] The parameter configuration module 420 is used to configure the parameters for hierarchical reinforcement learning for each subtask, including: a state set, a policy set, an action set, a termination state set, a state transition function, and a reward function. The states in the state set are the load balancing metric values of the corresponding level of the subtask; the termination states in the termination state set are the load balancing metric values of the corresponding level of the subtask that meet the load balancing criteria; the policies in the policy set are the load adjustment policies of the corresponding level of the subtask; the actions in the action set are the actions generated by executing the policies in the policy set; the state transition function represents the probability of performing an action in one state to enter another state; and the reward function represents the instantaneous reward value obtained by performing an action in one state to enter another state.
[0090] The hierarchical reinforcement learning module 430 is used to perform hierarchical reinforcement learning on the sub-tasks corresponding to the layers in the target network, obtain the relatively optimal policies of at least two layers of sub-tasks, and combine the determined relatively optimal policies of the at least two layers of sub-tasks into a global policy for load balancing for the target network.
[0091] The load adjustment module 440 is used to adjust the load of the target network based on the global policy.
[0092] The apparatus of this invention can decompose the task of load balancing a target network into sub-tasks of hierarchical reinforcement learning according to the hierarchy of the target network. Through hierarchical reinforcement learning, the sub-tasks of different levels can solve the load adjustment strategy relatively optimally in a smaller sub-problem space. The relatively optimal strategies of at least two levels are combined into a global strategy, thereby adjusting the load of the target network based on the global strategy. This avoids the phenomenon of only solving the load problem in a local area of the target network, which leads to excessive load in other areas.
[0093] Optionally, the target network includes a full network layer, and the subtasks of the full network layer serve as the root tasks of the hierarchical reinforcement learning. The termination condition of the hierarchical reinforcement learning is that the subtasks of the full network layer reach a termination state.
[0094] Optionally, the hierarchical reinforcement learning module 430 is specifically used for: a subtask optimal policy selection algorithm based on hierarchical reinforcement learning, selecting subtasks to solve for relatively optimal policies until the relatively optimal policy of the current subtask can enable the top-level subtask to reach a termination state. The termination conditions for solving the relatively optimal policy include the current subtask reaching a termination state and the relatively optimal policy of the current subtask involving actions related to the next level.
[0095] Optionally, the formula for the optimal strategy selection algorithm is: in:
[0096] Q(i,s,a)=V(a,s)+C(i,s,a);
[0097] Q(i,s,a)=V(a,s)+∑ s′,τ P i (s′,τ|s,a)γ τ maxQ[i,s′,π(s′)];
[0098] C(i, s, a)=∑ s′,τ P i (s′,τ|s,a)γ τ maxQ[i,s′,π(s′)];
[0099] V(i,s)=V(a,s)+∑ s′,τ P i (s′,τ|s,a)γ τ V(i,s′);
[0100] i represents the subtask number, a represents the action, a′ represents the currently selected action, and A iLet represent the action set of subtask i, s represent the current state, s′ represent the other state entered by performing an action in the current state, τ represent the duration of subtask execution in state s, γ represent the discount factor, Q(i,s,a) represent the value function corresponding to subtask i performing action a in state s, C(i,s,a) represent the completion function corresponding to subtask i performing action a in state s, V(i,s′) is the expected reward value for completing subtask i starting from state s′, and V(a,s) represents the instantaneous reward value for performing action a in state s. P i (s′,τ|s,a) represents the probability that performing action a in state s for time τ will lead to a transition to state s′, and π(s′) represents the policy of state s′.
[0101] Optionally, the function used by the hierarchical reinforcement learning to determine the global policy is π. * =arg π maxV π (s), where, r t This means that performing action a results in the t-th state s. t Transition to state s (t+1) t+1 The instantaneous reward value obtained afterwards, V π (s t ) represents the sum of accumulated instantaneous reward values, u represents the random jump coefficient, 0 <u≤1。
[0102] Optionally, the target network is a mobile communication network, and the layers of the mobile communication network, in the order from top to bottom, include: the whole network layer, the pool layer, the network element layer, and the link / virtual machine layer.
[0103] Among them, the load balancing index value of any layer in the mobile communication network in, y represents the average resource utilization rate of the hierarchy. L,m This represents the resource utilization rate of the m-th entity in the hierarchy, and N represents the total number of entities in the hierarchy.
[0104] Obviously, the embodiments of the present invention Figure 4 The network load balancing device shown can achieve the above. Figures 1 to 3 The steps and functions of the method shown are explained below. Since the principle is the same, they will not be repeated here.
[0105] Figure 5 This is a schematic diagram of the structure of an electronic device according to one embodiment of this specification. Please refer to it. Figure 5At the hardware level, the electronic device includes a processor, and optionally also includes an internal bus, a network interface, and memory. The memory may include main memory, such as high-speed random-access memory (RAM), or non-volatile memory, such as at least one disk drive. Of course, the electronic device may also include other hardware required for other business operations.
[0106] The processor, network interface, and memory can be interconnected via an internal bus, which can be an ISA (Industry Standard Architecture) bus, a PCI (Peripheral Component Interconnect) bus, or an EISA (Extended Industry Standard Architecture) bus, etc. This bus can be divided into address bus, data bus, control bus, etc. For ease of representation, Figure 5 The symbol is represented by a single double-headed arrow, but this does not mean that there is only one bus or one type of bus.
[0107] Memory is used to store programs. Specifically, the program may include program code, which includes computer operation instructions. Memory may include main memory and non-volatile memory, and provides instructions and data to the processor. The processor reads the corresponding computer program from the non-volatile memory into main memory and then runs it, forming a network load balancer at the logical level. This network load balancer can be independent of the server or it can be a component within the server. Correspondingly, the processor executes the program stored in the memory and specifically performs the following operations:
[0108] The task of performing load balancing on the target network is decomposed into sub-tasks of hierarchical reinforcement learning according to the hierarchy of the target network, wherein the target network is pre-divided into multiple hierarchical levels.
[0109] Parameters for hierarchical reinforcement learning are configured for each subtask, including: a state set, a policy set, an action set, a termination state set, a state transition function, and a reward function. The states in the state set are the load balancing metric values for the corresponding level of the subtask; the termination states in the termination state set are the load balancing metric values for the corresponding level of the subtask that meet the load balancing criteria; the policies in the policy set are the load adjustment policies for the corresponding level of the subtask; the actions in the action set are the actions generated by executing the policies in the policy set; the state transition function represents the probability of performing an action in one state to enter another state; and the reward function represents the instantaneous reward value obtained by performing an action in one state to enter another state.
[0110] Hierarchical reinforcement learning is performed on the subtasks corresponding to the layers in the target network to obtain the relatively optimal policies of at least two layers of subtasks, and the determined relatively optimal policies of the at least two layers of subtasks are combined into a global policy for load balancing of the target network.
[0111] Based on the global strategy, load adjustment is performed on the target network.
[0112] The electronic device of this invention can decompose the task of load balancing a target network into sub-tasks of hierarchical reinforcement learning according to the hierarchy of the target network. Through hierarchical reinforcement learning, the sub-tasks of different levels can solve the load adjustment strategy relatively optimally in a smaller sub-problem space. The relatively optimal strategies of at least two levels are combined into a global strategy, thereby adjusting the load of the target network based on the global strategy. This avoids the phenomenon of only solving the load problem in a local area of the target network, which leads to excessive load in other areas.
[0113] The above is as described in this instruction manual. Figure 1The methods disclosed in the illustrated embodiments can be applied to or implemented by a processor. The processor may be an integrated circuit chip with signal processing capabilities. During implementation, each step of the above methods can be completed by integrated logic circuits in the processor's hardware or by instructions in software form. The processor can be a general-purpose processor, including a Central Processing Unit (CPU), a Network Processor (NP), etc.; it can also be a Digital Signal Processor (DSP), an Application Specific Integrated Circuit (ASIC), a Field-Programmable Gate Array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components. The methods, steps, and logic block diagrams disclosed in the embodiments of this invention can be implemented or executed. The general-purpose processor can be a microprocessor or any conventional processor. The steps of the methods disclosed in the embodiments of this invention can be directly manifested as execution by a hardware decoding processor, or execution by a combination of hardware and software modules in the decoding processor. The software module can reside in a mature storage medium in the field, such as random access memory, flash memory, read-only memory, programmable read-only memory, electrically erasable programmable memory, or registers. This storage medium is located in memory, and the processor reads information from the memory and, in conjunction with its hardware, completes the steps of the above method.
[0114] It should be understood that the electronic device in the embodiments of the present invention can enable the network load balancing device to achieve the corresponding Figures 1 to 3 The steps and functions of the method shown are not repeated here as the principle is the same.
[0115] Of course, in addition to software implementation, the electronic device described in this specification does not exclude other implementation methods, such as logic devices or a combination of hardware and software. In other words, the execution subject of the following processing flow is not limited to each logic unit, but can also be hardware or logic devices.
[0116] Furthermore, embodiments of the present invention also provide a computer-readable storage medium that stores one or more programs, the one or more programs including instructions.
[0117] The task of performing load balancing on the target network is decomposed into sub-tasks of hierarchical reinforcement learning according to the hierarchy of the target network, wherein the target network is pre-divided into multiple hierarchical levels.
[0118] Parameters for hierarchical reinforcement learning are configured for each subtask, including: a state set, a policy set, an action set, a termination state set, a state transition function, and a reward function. The states in the state set are the load balancing metric values for the corresponding level of the subtask; the termination states in the termination state set are the load balancing metric values for the corresponding level of the subtask that meet the load balancing criteria; the policies in the policy set are the load adjustment policies for the corresponding level of the subtask; the actions in the action set are the actions generated by executing the policies in the policy set; the state transition function represents the probability of performing an action in one state to enter another state; and the reward function represents the instantaneous reward value obtained by performing an action in one state to enter another state.
[0119] Hierarchical reinforcement learning is performed on the subtasks corresponding to the layers in the target network to obtain the relatively optimal policies of at least two layers of subtasks, and the determined relatively optimal policies of the at least two layers of subtasks are combined into a global policy for load balancing of the target network.
[0120] Based on the global strategy, load adjustment is performed on the target network.
[0121] Those skilled in the art will understand that the embodiments of this specification can be provided as methods, systems, or computer program products. Therefore, this specification may take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, this specification may take the form of a computer program product embodied on one or more computer-usable storage media (including, but not limited to, disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0122] The foregoing has described specific embodiments of this specification. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims may be performed in a different order than that shown in the embodiments and may still achieve the desired result. Furthermore, the processes depicted in the drawings do not necessarily require the specific or sequential order shown to achieve the desired result. In some embodiments, multitasking and parallel processing are possible or may be advantageous.
[0123] The above are merely embodiments of this specification and are not intended to limit the scope of this specification. Various modifications and variations can be made to this specification by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this specification should be included within the scope of the claims of this specification. Furthermore, all other embodiments obtained by those skilled in the art without inventive effort should fall within the protection scope of this document.
Claims
1. A network load balancing method, characterized in that, include: The task of performing load balancing on the target network is decomposed into sub-tasks of hierarchical reinforcement learning according to the hierarchy of the target network, wherein the target network is pre-divided into multiple hierarchical levels. Parameters for hierarchical reinforcement learning are configured for each subtask, including: a state set, a policy set, an action set, a termination state set, a state transition function, and a reward function. The states in the state set are the load balancing metric values for the corresponding level of the subtask; the termination states in the termination state set are the load balancing metric values for the corresponding level of the subtask that meet the load balancing criteria; the policies in the policy set are the load adjustment policies for the corresponding level of the subtask; the actions in the action set are the actions generated by executing the policies in the policy set; the state transition function represents the probability of performing an action in one state to enter another state; and the reward function represents the instantaneous reward value obtained by performing an action in one state to enter another state. Hierarchical reinforcement learning is performed on the subtasks corresponding to the layers in the target network to obtain the relatively optimal policies of at least two layers of subtasks, and the determined relatively optimal policies of the at least two layers of subtasks are combined into a global policy for load balancing of the target network. Based on the global strategy, load adjustment is performed on the target network; The formula for the optimal strategy selection algorithm is: ;in: ; ; ; ; Indicates the subtask number. Indicates an action, Indicates the currently selected action. This represents the action set of subtask i. Indicates the current state. This indicates the state that is entered after performing an action in the current state. Indicates the state The duration of the subtask execution. Indicates the discount factor. Subtasks In state Execute action The corresponding value function, Subtasks In state Execute action The corresponding completion function, It is a state Start completing subtasks Expected return value, Indicates in Perform an action in a state Instantaneous reward value, Indicates the state Execute action This leads to a transition to a state. The probability, Representing state The strategy.
2. The method according to claim 1, characterized in that, The target network has a full network hierarchy, and the subtasks of the full network hierarchy serve as the root tasks of the hierarchical reinforcement learning. The termination condition of the hierarchical reinforcement learning is that the subtasks of the full network hierarchy reach the termination state.
3. The method according to claim 2, characterized in that, Hierarchical reinforcement learning is performed on the subtasks corresponding to the layers in the target network to obtain the relatively optimal policies for at least two layers of subtasks. The determined relatively optimal policies for the at least two layers of subtasks are then combined into a global policy for load balancing of the target network, including: The optimal policy selection algorithm for subtasks based on hierarchical reinforcement learning selects subtasks to solve for relatively optimal policies until the relatively optimal policy of the current subtask enables the top-level subtask to reach a termination state. The termination conditions for solving the relatively optimal policy include the current subtask reaching a termination state and the relatively optimal policy of the current subtask involving actions related to the next level.
4. The method according to claim 1, characterized in that, The function used by the hierarchical reinforcement learning to determine the global policy is: ,in, Indicates that performing action a causes the first... each state Transfer to the each state The instantaneous reward value obtained afterwards, This represents the sum of accumulated instantaneous reward values. This represents the random jump coefficient. .
5. The method according to claim 1, characterized in that, The target network is a mobile communication network, and the layers of the mobile communication network, in the order from top to bottom, include: the whole network layer, the pool layer, the network element layer, and the link / virtual machine layer.
6. The method according to claim 5, characterized in that, Load balancing metric values at any level in the mobile communication network ; in, This represents the average resource utilization rate of the hierarchy. Indicates the first in the hierarchy Resource utilization rate of an individual entity This indicates the total number of entities in the hierarchy.
7. A network load balancing device, characterized in that, include: The task creation module is used to decompose the task of performing load balancing for the target network into sub-tasks of hierarchical reinforcement learning according to the hierarchy of the target network, wherein the target network is pre-divided into multiple hierarchical levels. The parameter configuration module is used to configure the parameters for hierarchical reinforcement learning for each subtask, including: a state set, a policy set, an action set, a termination state set, a state transition function, and a reward function. The states in the state set are the load balancing metric values for the corresponding level of the subtask; the termination states in the termination state set are the load balancing metric values for the corresponding level of the subtask that meet the load balancing criteria; the policies in the policy set are the load adjustment policies for the corresponding level of the subtask; the actions in the action set are the actions generated by executing the policies in the policy set; the state transition function represents the probability of performing an action in one state to enter another state; and the reward function represents the instantaneous reward value obtained by performing an action in one state to enter another state. The hierarchical reinforcement learning module is used to perform hierarchical reinforcement learning on the sub-tasks corresponding to the layers in the target network, obtain the relative optimal policies of at least two layers of sub-tasks, and combine the determined relative optimal policies of the at least two layers of sub-tasks into a global policy for load balancing for the target network. A load adjustment module is used to adjust the load of the target network based on the global policy. The formula for the optimal strategy selection algorithm is: ;in: ; ; ; ; Indicates the subtask number. Indicates an action, Indicates the currently selected action. This represents the action set of subtask i. Indicates the current state. This indicates the state that is entered after performing an action in the current state. Indicates the state The duration of the subtask execution. Indicates the discount factor. Subtasks In state Execute action The corresponding value function, Subtasks In state Execute action The corresponding completion function, It is a state Start completing subtasks Expected return value, Indicates in Perform an action in a state Instantaneous reward value, Indicates the state Execute action This leads to a transition to a state. The probability, Representing state The strategy.
8. An electronic device comprising: A memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that the computer program is executed by the processor: The task of performing load balancing on the target network is decomposed into sub-tasks of hierarchical reinforcement learning according to the hierarchy of the target network, wherein the target network is pre-divided into multiple hierarchical levels. Parameters for hierarchical reinforcement learning are configured for each subtask, including: a state set, a policy set, an action set, a termination state set, a state transition function, and a reward function. The states in the state set are the load balancing metric values for the corresponding level of the subtask; the termination states in the termination state set are the load balancing metric values for the corresponding level of the subtask that meet the load balancing criteria; the policies in the policy set are the load adjustment policies for the corresponding level of the subtask; the actions in the action set are the actions generated by executing the policies in the policy set; the state transition function represents the probability of performing an action in one state to enter another state; and the reward function represents the instantaneous reward value obtained by performing an action in one state to enter another state. Hierarchical reinforcement learning is performed on the subtasks corresponding to the layers in the target network to obtain the relatively optimal policies of at least two layers of subtasks, and the determined relatively optimal policies of the at least two layers of subtasks are combined into a global policy for load balancing of the target network. Based on the global strategy, load adjustment is performed on the target network; The formula for the optimal strategy selection algorithm is: ;in: ; ; ; ; Indicates the subtask number. Indicates an action, Indicates the currently selected action. This represents the action set of subtask i. Indicates the current state. This indicates the state that is entered after performing an action in the current state. Indicates the state The duration of the subtask execution. Indicates the discount factor. Subtasks In state Execute action The corresponding value function, Subtasks In state Execute action The corresponding completion function, It is a state Start completing subtasks Expected return value, Indicates in Perform an action in a state Instantaneous reward value, Indicates the state Execute action This leads to a transition to a state. The probability, Representing state The strategy.
9. A computer-readable storage medium storing a computer program thereon, characterized in that, When the computer program is executed by the processor, it performs the following steps: The task of performing load balancing on the target network is decomposed into sub-tasks of hierarchical reinforcement learning according to the hierarchy of the target network, wherein the target network is pre-divided into multiple hierarchical levels. Parameters for hierarchical reinforcement learning are configured for each subtask, including: a state set, a policy set, an action set, a termination state set, a state transition function, and a reward function. The states in the state set are the load balancing metric values for the corresponding level of the subtask; the termination states in the termination state set are the load balancing metric values for the corresponding level of the subtask that meet the load balancing criteria; the policies in the policy set are the load adjustment policies for the corresponding level of the subtask; the actions in the action set are the actions generated by executing the policies in the policy set; the state transition function represents the probability of performing an action in one state to enter another state; and the reward function represents the instantaneous reward value obtained by performing an action in one state to enter another state. Hierarchical reinforcement learning is performed on the subtasks corresponding to the layers in the target network to obtain the relatively optimal policies of at least two layers of subtasks, and the determined relatively optimal policies of the at least two layers of subtasks are combined into a global policy for load balancing of the target network. Based on the global strategy, load adjustment is performed on the target network; The formula for the optimal strategy selection algorithm is: ;in: ; ; ; ; Indicates the subtask number. Indicates an action, Indicates the currently selected action. This represents the action set of subtask i. Indicates the current state. This indicates the state that is entered after performing an action in the current state. Indicates the state The duration of the subtask execution. Indicates the discount factor. Subtasks In state Execute action The corresponding value function, Subtasks In state Execute action The corresponding completion function, It is a state Start completing subtasks Expected return value, Indicates in Perform an action in a state Instantaneous reward value, Indicates the state Execute action This leads to a transition to a state. The probability, Representing state The strategy.