A load balancing method and device of a computing power network, a load balancer and a medium
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-04-20
- Publication Date
- 2026-08-11
AI Technical Summary
[0004]本发明提供了一种算力网络的负载均衡方法、装置、负载均衡器及介质,以解决现有技术存在的资源单位难统一、均衡指标难衡量以及多层结构难应对等问题
Smart Images

Figure CN116366557B_ABST
Abstract
Description
Technical Field
[0001] The embodiments of the present invention relate to the field of cloud computing technology, and in particular to a load balancing method, device, load balancer and medium for computing power networks. Background Technology
[0002] A computing power network is a new type of information infrastructure formed by multiple computing power nodes connected via a network. It can allocate and flexibly schedule computing, storage, and network resources on demand across the cloud, network, and edge, according to business needs. The load of a computing power network is mainly divided into two categories: server load, which refers to the state of server computing, storage, and I / O resources used to respond to current business needs; and network load, which refers to the state of network resources used to respond to current business needs. Load balancing refers to the even distribution of computing network services to different computing power nodes based on the network's load status, ensuring balanced utilization of computing network resources and avoiding performance degradation and service unresponsiveness caused by overload of a single node. However, in real-world scenarios, the load of a computing power network exhibits three main characteristics: diverse locations, numerous types, and different causes.
[0003] Currently, there are two load balancing solutions: one is the academic-led, hard-coupled load balancing strategy between upper and lower layers; the other is the industry-led, decoupled load balancing strategy between upper and lower layers. However, current load balancing solutions have the following problems: difficulty in unifying multi-dimensional resources, difficulty in measuring balancing metrics, and difficulty in handling multi-layered structures. Summary of the Invention
[0004] This invention provides a load balancing method, device, load balancer, and medium for computing networks, to solve problems such as difficulty in unifying resource units, difficulty in measuring balancing indicators, and difficulty in dealing with multi-layer structures in existing technologies.
[0005] According to one aspect of the present invention, a load balancing method for a computing network is provided, comprising:
[0006] Collect first resource information, which includes the remaining quantity of various resources and the subordinate relationships of various resources;
[0007] Based on the initial resource consumption mapping model of various resources and the remaining amount of various resources, resource normalization is performed to obtain the request capacity of various resources.
[0008] Based on the resource request capacity of each type, the resource bottlenecks of each node are identified hierarchically, and the resource bottlenecks of each node are summarized upwards layer by layer according to the resource hierarchy.
[0009] The request distribution weights for each layer are obtained by calculating the request distribution weights based on the resource bottlenecks of each node.
[0010] The request distribution weights of each layer are distributed to each node so that when a service request arrives, each node distributes the service request based on the request distribution weights of each layer using a preset method.
[0011] According to another aspect of the present invention, a load balancing device for a computing network is provided, comprising:
[0012] A collection device is used to collect first resource information, which includes the remaining amount of various resources and the subordinate relationships of various resources.
[0013] The normalization module is used to normalize resources based on the initial resource consumption mapping model of various resources and the remaining amount of various resources, so as to obtain the request capacity of various resources.
[0014] The identification module is used to identify the resource bottlenecks of each node in a hierarchical manner based on the capacity to accommodate various types of resource requests, and to summarize the resource bottlenecks of each node layer by layer upward according to the hierarchical relationship of various types of resources.
[0015] The calculation module is used to calculate the request distribution weights at each layer based on the resource bottlenecks of each node.
[0016] The distribution module is used to distribute the request distribution weights of each layer to each node, so that when a service request arrives, each node distributes the service request based on the request distribution weights of each layer and uses a preset method.
[0017] According to another aspect of the present invention, a load balancer is provided, comprising:
[0018] At least one processor;
[0019] and a memory communicatively connected to the at least one processor;
[0020] The memory stores a computer program that can be executed by the at least one processor, which is then executed by the at least one processor to enable the at least one processor to perform the load balancing method for the computing power network according to any embodiment of the present invention.
[0021] According to another aspect of the present invention, a computer-readable storage medium is provided, the computer-readable storage medium storing computer instructions for causing a processor to execute and implement the load balancing method of the computing network according to any embodiment of the present invention.
[0022] The technical solution of this invention collects first resource information, including the remaining amount of various resources and the subordinate relationships of various resources; performs resource normalization based on the initial resource consumption mapping model of various resources and the remaining amount of various resources to obtain the request capacity of various resources; identifies the resource bottlenecks of each node hierarchically based on the request capacity of various resources, and summarizes the resource bottlenecks of each node layer by layer according to the subordinate relationships of various resources; calculates the request distribution weight of each layer based on the resource bottleneck of each node to obtain the request distribution weight of each layer; and distributes the request distribution weight of each layer to each node so that when a business request arrives, each node distributes the business request according to the request distribution weight of each layer using a preset method. This solves the problems of difficulty in unifying resource units, difficulty in measuring balance indicators, and difficulty in dealing with multi-layer structures in the prior art.
[0023] It should be understood that the description in this section is not intended to identify key or essential features of the embodiments of the present invention, nor is it intended to limit the scope of the invention. Other features of the invention will become readily apparent from the following description. Attached Figure Description
[0024] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0025] Figure 1 This is a flowchart illustrating a load balancing method for a computing network provided in Embodiment 1 of the present invention.
[0026] Figure 2 This is a schematic diagram of the overall architecture of a load balancer provided in Embodiment 1 of the present invention;
[0027] Figure 3 This is a flowchart illustrating a load balancing method for a computing network provided in Embodiment 2 of the present invention.
[0028] Figure 4 The calculation flowchart of the load balancer dynamic adjustment module provided in Embodiment 3 of the present invention;
[0029] Figure 5 This is a schematic diagram of the structure of a load balancing device for a computing network provided in Embodiment 4 of the present invention;
[0030] Figure 6 This is a schematic diagram of the load balancer structure of a load balancing method for a computing network according to an embodiment of the present invention. Detailed Implementation
[0031] To enable those skilled in the art to better understand the present invention, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are merely some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of the present invention. It should be understood that the various steps described in the method embodiments of the present invention can be performed in different orders and / or in parallel. Furthermore, the method embodiments may include additional steps and / or omit the steps shown. The scope of the present invention is not limited in this respect.
[0032] The term "comprising" and its variations as used herein are open-ended inclusions, meaning "including but not limited to". The term "based on" means "at least partially based on". The term "one embodiment" means "at least one embodiment"; the term "another embodiment" means "at least one additional embodiment"; the term "some embodiments" means "at least some embodiments". Definitions of other terms will be given in the description below.
[0033] It should be noted that the terms "first," "second," etc., in the specification, claims, and accompanying drawings of this invention are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of the invention described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover a non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.
[0034] It should be noted that the terms "a" and "a plurality of" used in this invention are illustrative rather than restrictive. Those skilled in the art should understand that, unless otherwise expressly indicated in the context, they should be understood as "one or more".
[0035] The names of the messages or information exchanged between the multiple devices in the embodiments of the present invention are for illustrative purposes only and are not intended to limit the scope of these messages or information.
[0036] It is important to understand that the load of a computing network, or computing network load, has three main characteristics: it exists in different locations, has many different types, and has different causes. These are as follows:
[0037] 1. Different locations of computing network load: The load depends on different locations. The server load is located on the server, while the network load is located in the communication network. The different load dependencies will lead to the problem of lower-level structures sharing upper-level resources. The data transmission generated by computing network services allocated to the cluster is supported by the communication network. The cluster is composed of multiple computing nodes, and multiple computing nodes will share the cluster network resources. Therefore, the node computing resources and network resources are not on the same plane.
[0038] 2. Computing network load consists of two main categories, each of which contains several subcategories, and each subcategory also contains different resource types. Meanwhile, the diversification of current computing network services is becoming increasingly apparent, and the types of computing resources provided at the underlying level are also increasing. Load balancers must possess diversified resource management capabilities and backward compatibility with new computing resources.
[0039] 3. Different Causes of Computing and Network Load: Server load arises from responding to computing requests, while network load arises from transmitting the data required for those requests. Within server load, the causes of each type of load are different. Furthermore, different resources have different dimensions, and the diversity of computing and network services results in a non-fixed proportional relationship in their consumption.
[0040] The two existing load balancing solutions are as follows:
[0041] 1. Hard-coupled load balancing strategy between upper and lower layers
[0042] This strategy defines the load balancing problem as a task scheduling problem, based on clearly defined and given assumptions regarding metrics such as business computation volume, data transmission volume, node topology, network transmission capacity, and node computing capacity. It achieves load balancing by using optimal scheduling strategies to perform fine-grained scheduling control and resource allocation for underlying data transmission and computation. However, its underlying assumptions are extremely difficult to obtain, and the underlying scheduling control strategies are incompatible with existing industry technologies. Therefore, while it offers strong load balancing capabilities, its feasibility is low.
[0043] 2. Decoupled load balancing strategy between upper and lower layers
[0044] This strategy decouples underlying resources from upper-layer applications, focusing on the state of underlying resources and determining the load balancing strategy based on that state, offering strong operability. This type of load balancing strategy can be further divided into two subcategories: static load balancing strategies and dynamic load balancing strategies.
[0045] 1) Static load balancing strategy
[0046] Static load balancing strategies allocate service requests with fixed parameters, disregarding the current load status. They have evolved from simple random strategies, simple round-robin strategies, simple weighted round-robin strategies, and smoothed weighted round-robin strategies. Currently, static load balancing strategies employ a weighted round-robin scheduling method, assigning different weights to different nodes and using a smoothing algorithm to avoid pseudo-randomness. This type of strategy features extremely fast decision-making speed and strong implementability; however, because it does not consider the current load status and uses fixed parameters to allocate service requests, it cannot adjust the strategy in real time according to the current load status to optimize the balancing effect.
[0047] 2) Dynamic load balancing strategy
[0048] Dynamic load balancing strategies determine service allocation based on real-time server status information, primarily including least connections strategy, weighted least connections strategy, and least connections slow start time strategy. These strategies mainly rely on the number of node connections for balancing decisions, thus considering real-time server status information and allowing adjustments to the balancing strategy to optimize performance. However, these strategies fail to consider network resource conditions and focus on balancing node load while neglecting node capacity, therefore they do not truly possess the ability to balance the load across all nodes.
[0049] Existing load balancing solutions have the following problems:
[0050] 1. Difficulty in unifying multi-dimensional resources: Current load scheduling schemes based on node scoring use a linear summation method to score nodes, summing various resources according to certain weights. This results in a lack of uniformity in resource dimensions. Furthermore, when additional types of resources are added to the system, the original scoring and evaluation system will no longer be applicable, and the summation weights must be manually set.
[0051] 2. Difficulty in measuring load balancing metrics: Current load balancing schemes based on node scores evaluate each resource and then perform a weighted summation. This makes it difficult to accurately identify resource bottlenecks that affect load balancing performance. When adding excess resources, the node score will also increase accordingly, which can easily lead to overload of a single resource.
[0052] 3. Difficulty in handling multi-layered structures: Current node-based load balancing schemes assume that all resources are at the same level before load balancing decisions can be made. However, in real-world scenarios, some resources are shared by multiple load balancing targets, such as multiple nodes sharing cluster communication resources, which renders existing load balancing strategies ineffective.
[0053] In summary, existing load balancing schemes suffer from difficulties in unifying multi-dimensional resources, measuring balancing metrics, and handling multi-layered structures. Therefore, this invention proposes a load balancing method for computing power networks to address these issues.
[0054] Example 1
[0055] Figure 1 This is a flowchart illustrating a load balancing method for a computing network provided in Embodiment 1 of the present invention. This method is applicable to situations where the computing network is balanced in an environment with multiple clusters and multiple dimensions of resources. This method can be executed by a load balancing device of the computing network, which can be implemented by software and / or hardware and is generally integrated on a load balancer. In this embodiment, the load balancer includes, but is not limited to, a computing network load balancer.
[0056] like Figure 1 As shown in Embodiment 1 of the present invention, a load balancing method for a computing network includes the following steps:
[0057] S110. Collect first resource information, which includes the remaining amount of various resources and the subordinate relationships of various resources.
[0058] The first resource information can be related to various types of resources. It can include the remaining amount of various resources, the subordinate relationships of various resources, and the usage of various resources.
[0059] These resources can include computing resources, storage resources, and network resources.
[0060] In this embodiment, after the computing network service is launched and the microservice container is deployed, the first resource information can be collected periodically. The method of collecting the first resource information is not limited here.
[0061] Furthermore, the overall architecture of the load balancer includes the load balancer, cluster, and containers. Based on the hierarchical relationships of these resources, the overall architecture of the load balancer is converted into a tree diagram representation. The load balancer, cluster, compute nodes, and containers are converted into node representations. The load balancer is represented by a root node, the cluster and compute nodes are represented by intermediate nodes, and the containers are represented by leaf nodes. Each node is marked with a resource identifier, remaining resource quantity, and resource usage. Nodes at the same level have the same set of resource identifiers, representing the same type of resource.
[0062] Figure 2 This is a schematic diagram of the overall architecture of a load balancer provided in Embodiment 1 of the present invention, as shown below. Figure 2As shown, the load balancer is divided into cluster 1 and cluster 2. Within cluster 1 and cluster 2, the master node controls the scheduler to schedule containers under the compute nodes. Compute node 1 under cluster 1 contains multiple containers such as container A, container B, and container C; compute node 2 under cluster 1 contains multiple containers such as container D, container E, and container F; compute node 3 under cluster 2 contains multiple containers such as container G, container H, and container I; and compute node 4 under cluster 2 contains multiple containers such as container J, container K, and container L.
[0063] In this embodiment, the architecture diagram of the load balancer can be converted into a tree-like resource topology diagram, i.e., a tree diagram, based on various resource hierarchical relationships. Here, the load balancer, cluster, compute nodes, and containers can all be represented by nodes in the tree diagram. Specifically, the load balancer can be represented as a root node, the cluster and compute nodes as intermediate nodes, and the containers as leaf nodes.
[0064] Specifically, a node is represented as N. i,j Where j is the node index and i is the node degree. The universal set of all nodes is represented as: Right now From node N i,j To node N i+1,j′ The edge of E (i,j),(i+1,j′) The label represents node N. i+1,j′ For node N i,j The lower-level nodes; the complete set of all edges in the tree-like resource topology graph is represented as ε, i.e., E. (i,j),(i+1,j′) ∈ε. The remaining amount of various resources on the node is determined by... The tag indicates that resource usage is determined by... The label, k, represents the resource type, i.e., node N. i,j There are k types of resources, and nodes at the same level have the same type of resource, while nodes at different levels have different types of resources. For example, there are 3 nodes at the same level, and each node can include both type A and type B resources; Node N i,j The complete collection of resources is represented as After a service request enters from the root node load balancer, it is distributed to the leaf node containers via the child node cluster. A distribution weight x needs to be configured for each edge. (i,j),(i+1,j′) Furthermore, the sum of the weights distributed from a parent node to all its child nodes is 1. A compute node can be the parent node of a container and also the child node of a cluster. A cluster can be the parent node of a compute node and also the child node of a load balancer.
[0065] S120. Based on the initial resource consumption mapping model of various resources and the remaining amount of various resources, resource normalization is performed to obtain the resource request capacity of various resources.
[0066] The initial resource consumption mapping model can be understood as a resource consumption mapping model of the initial state, used to characterize the amount of each resource consumed by a single service request, and can be denoted as q. k The initial resource consumption mapping model can be obtained based on the type of computing network service or based on historical information of computing network services. The initial resource consumption mapping model is the same for all types of resources.
[0067] The capacity to accommodate various resource requests can be understood as the number of requests that the remaining resources of each type can serve.
[0068] In this embodiment, based on the initial resource consumption mapping model and the remaining amount of various resources, the remaining amount of various resources can be normalized into the request capacity of various resources.
[0069] Specifically, the request capacity of different types of resources on each node can be calculated using the following method. Taking a single node as an example, the calculation method is as follows: the ratio of the remaining resource quantity of a certain type of resource on a node to the initial resource consumption yields the request capacity of that type of resource on that node. The initial resource consumption can be obtained through an initial resource consumption mapping model, and can be the initial resource consumption of a single request.
[0070] S130. Based on the resource request capacity of the various types of resources, identify the resource bottlenecks of each node in a hierarchical manner, and summarize the resource bottlenecks of each node layer by layer upward according to the hierarchical relationship of the various types of resources.
[0071] Among them, the resource bottleneck can be understood as the minimum resource requirement.
[0072] Specifically, when a node is a leaf node, only the minimum value among the various resource request capacities of that leaf node needs to be calculated as its resource bottleneck. Since the leaf node is the lowest level node, its resource bottleneck can be aggregated upwards to the root node. When a node is a non-leaf node (i.e., an intermediate node or the root node), the same calculation method can be used to calculate the resource bottleneck of each node layer by layer upwards. The calculation method is as follows: calculate the minimum value among the various resource request capacities of the non-leaf node, and the sum of the resource bottlenecks of all its lower-level nodes. Then, compare the minimum value among the various resource request capacities of the non-leaf node with the sum of the resource bottlenecks of all its lower-level nodes, and take the smaller value as the resource bottleneck of the non-leaf node.
[0073] In this embodiment, the dependency relationships between nodes can be known based on various resource dependencies, and the resource bottlenecks of each node can be aggregated layer by layer upwards to the root node, i.e., the load balancer.
[0074] S140. Calculate the request distribution weights for each layer based on the resource bottlenecks of each node.
[0075] The request distribution weight of the current layer can include the request distribution weights of all nodes in the current layer and their subordinate nodes.
[0076] In this embodiment, the same calculation method can be used to calculate the request distribution weight in layers. Taking one request distribution weight in a layer as an example, the ratio of the resource bottleneck of a node's lower-level node to the sum of the resource bottlenecks of all its lower-level nodes is calculated as the request distribution weight between that node and that lower-level node. Following the above method, the request distribution weight between that node and each lower-level node can be calculated, and then the request distribution weight between each node in a layer and each of its lower-level nodes can be calculated to obtain the layer's request distribution weight. Each layer can be calculated in the above way to obtain the request distribution weight of each layer.
[0077] S150. The request distribution weights of each layer are distributed to each node so that when a service request arrives, each node distributes the service request based on the request distribution weights of each layer using a preset method.
[0078] The preset methods include, but are not limited to, weighted round-robin algorithms or derived strategies of weighted round-robin algorithms. Derived strategies of weighted round-robin algorithms involve using different distribution methods to balance request distribution based on a weighted round-robin distribution ratio, such as greatest common divisor weighted round-robin, smooth weighted round-robin, etc.
[0079] In this embodiment, the calculated request distribution weights for each layer are distributed to each node. When a service request arrives, each node can distribute the request using a weighted round-robin algorithm or a derivation strategy based on the distribution weight of its child nodes.
[0080] This invention provides a load balancing method for a computing network. First, it collects first resource information, including the remaining amount of various resources and their hierarchical relationships. Second, it performs resource normalization based on the initial resource consumption mapping model of each resource and the remaining amount of each resource to obtain the request capacity of each resource. Then, it identifies the resource bottlenecks of each node based on the requested capacity of each resource, and aggregates the resource bottlenecks of each node layer by layer according to the hierarchical relationships of each resource. Next, it calculates the request distribution weights at each layer based on the resource bottlenecks of each node to obtain the request distribution weights at each layer. Finally, it distributes the request distribution weights at each layer to each node to enable service requests to meet their respective requirements. Upon arrival, each node distributes the service request based on the request distribution weights of each layer using a preset method. This method establishes a resource consumption mapping model and normalizes remaining resources into service request capacity, unifying the evaluation criteria for different resources and adapting to changes in resource types. After normalizing remaining resources into service request capacity, this method accurately identifies resource bottlenecks affecting load balancing performance by horizontally comparing the current capacity of each resource, enabling load balancing based on these bottlenecks. This method employs a layered comparison and upward aggregation resource bottleneck identification strategy, accurately identifying resource bottlenecks at each level and thus addressing multi-layered resource structures.
[0081] Based on the above embodiments, modified embodiments of the above embodiments are proposed. It should be noted that, in order to keep the description brief, only the differences from the above embodiments are described in the modified embodiments.
[0082] In one embodiment, the method for determining the initial resource consumption mapping model for the various types of resources includes:
[0083] The resource consumption mapping model obtained by initializing based on the computing power network service type can be used as the initial resource consumption mapping model for various types of resources, or the resource consumption mapping model obtained by initializing based on the historical information of computing power network services can be used as the initial resource consumption mapping model for various types of resources.
[0084] In one embodiment, the resource normalization based on the initial resource consumption mapping model for various types of resources and the remaining amount of various types of resources to obtain the resource request capacity includes:
[0085] For k types of resources on a node, the ratio of the remaining resource amount of the k types of resources to the initial resource consumption of the k types of resources is calculated to obtain the request capacity of the k types of resources. The initial resource consumption of the k types of resources is calculated through the initial resource consumption mapping model.
[0086] The capacity to accommodate various resource requests is calculated for each node in the manner described above.
[0087] It is understandable that a node can have multiple types of resources, and the calculation method for the request capacity of each type of resource is the same, which will not be elaborated here. The request capacity of each type of resource can include the request capacity of each type of resource on each node.
[0088] The following formula can be used for resource normalization to obtain the request handling capacity of various resources:
[0089]
[0090] In the above formula, Represents node N i,j The remaining amount of resources of type k, q k This represents the resource consumption of resource type k calculated by the resource consumption model for resource type k. Represents node N i,j The request handling capacity of the k-th type of resources.
[0091] In one embodiment, the hierarchical identification of resource bottlenecks for each node based on the capacity to accommodate various resource requests includes:
[0092] If the nodes included in the current layer are leaf nodes, calculate the minimum value among the various resource request capacity of the leaf node as the resource bottleneck of the leaf node;
[0093] If the nodes included in the current layer are non-leaf nodes, calculate the minimum value among the resource request capacity of each type of non-leaf node as the resource request capacity of the non-leaf node, calculate the sum of the resource bottlenecks of all lower-level nodes of the non-leaf node, compare the resource request capacity of the non-leaf node with the sum of the resource bottlenecks of all lower-level nodes of the non-leaf node, and take the smaller value as the resource bottleneck of the non-leaf node.
[0094] Specifically, when a node is a leaf node, which is the lowest-level node, the resource bottleneck of the leaf node can be calculated using the following formula:
[0095]
[0096] in, Represents node N i,j The corresponding resource request capacity, Represents node N i,j The complete collection of resources on r i,j Represents node N i,j Resource bottlenecks.
[0097] Specifically, when a node is a root node or a child node, the resource bottleneck of each node can be calculated layer by layer upwards using the following formula:
[0098]
[0099] in, This represents the minimum resource request capacity among all types of resources corresponding to a non-leaf node. r represents the sum of resource bottlenecks of all lower-level nodes of a non-leaf node. i+1,j′ N represents the lower-level node of a non-leaf node. i+1,j′ Resource bottlenecks.
[0100] Furthermore, the step of calculating the request distribution weights for each layer based on the resource bottlenecks of each node includes:
[0101] For a node in a layer, calculate the ratio of the resource bottleneck of one of the node's lower-level nodes to the sum of the resource bottlenecks of all the node's lower-level nodes to obtain the request distribution weight from the node to one of its lower-level nodes; use all the request distribution weights calculated in the layer as the layer request distribution weight.
[0102] Specifically, the request distribution weight is calculated using the following formula:
[0103]
[0104] Where, x (i,j),(i+1,j′) Represents node N i,j To its lower node N i+1,j′ Request distribution weight, r i+1,j′ Represents node N i,j A lower-level node N i+1,j′ Resource bottlenecks Represents node N i,j The sum of resource bottlenecks of all lower-level nodes, where (i+1,j'') represents node N. i,j The index of any lower-level node, j'' can be j'.
[0105] Example 2
[0106] Figure 3 This is a flowchart illustrating a load balancing method for a computing network according to Embodiment 2 of the present invention. Embodiment 2 is an optimization based on the above embodiments. For details not covered in this embodiment, please refer to Embodiment 1.
[0107] like Figure 3 As shown in Embodiment 2 of the present invention, a load balancing method for a computing network includes the following steps:
[0108] S210. Collect first resource information, which includes the remaining amount of various resources and the subordinate relationships of various resources.
[0109] S220. Based on the initial resource consumption mapping model of various types of resources and the remaining amount of various types of resources, resource normalization is performed to obtain the resource request capacity of various types of resources.
[0110] S230. Based on the resource request capacity of the various types of resources, identify the resource bottlenecks of each node in a hierarchical manner, and summarize the resource bottlenecks of each node layer by layer upward according to the hierarchical relationship of the various types of resources.
[0111] S240. Calculate the request distribution weights for each layer based on the resource bottlenecks of each node.
[0112] S250. The request distribution weights of each layer are distributed to each node so that when a service request arrives, each node distributes the service request based on the request distribution weights of each layer using a preset method.
[0113] S260. Collect second resource information and total number of requests according to a preset period; wherein, the second resource information includes the remaining amount of various resources, the usage of various resources and the subordinate relationship of various resources, and the total number of requests includes the total number of incoming requests, the total number of completed service requests and the total number of service failure requests.
[0114] The preset period can be a pre-set time length, and there are no specific restrictions on the value of the preset period. The total number of arriving requests can be the total number of requests that have arrived in the system, the total number of completed service requests can be the total number of requests that have been processed in the system, and the total number of service failure requests can be the total number of requests that failed to be processed in the system.
[0115] S270. Calculate the current resource consumption mapping model for each type of resource based on the total number of requests and the usage of each type of resource at each node.
[0116] Specifically, for resource k, the sum of resource k usage of each node is calculated, and the ratio of the sum of resource k usage of each node to the total number of currently served requests is used as the current resource consumption mapping model of resource k; wherein, the total number of currently served requests is the difference between the total number of arriving requests, the total number of completed service requests, and the total number of service failure requests.
[0117] The current resource consumption mapping model for various resources can be calculated using the following formula:
[0118]
[0119] in, This represents the current resource consumption mapping for resource type k. RE represents the total usage of k types of resources on each node. total E represents the total number of requests received. completed Represents the total number of completed service requests, RE failed This indicates the total number of failed service requests.
[0120] S280. Update the initial resource consumption mapping model using the current resource consumption mapping model to obtain the updated resource consumption mapping model.
[0121] After calculating the current resource consumption mapping model for each resource, the initial resource consumption mapping model can be updated using the exponential smoothing method, and the updated resource consumption mapping model can be used as the updated resource consumption mapping model.
[0122] Specifically, the initial resource mapping model can be updated using the following formula:
[0123]
[0124] Where α represents the update factor, which can be preset according to the scenario. α determines the update rate of the current resource mapping model and has a value range of [0,1]. Represents the current resource consumption mapping for resource type k, q k Let q represent the initial resource consumption mapping model for resource class k. ′k This represents the updated resource consumption mapping model for resource type k.
[0125] S290. Using the updated resource consumption mapping model, return to execute S220 to S240 to update the request distribution weights of each layer, and distribute the updated request distribution weights of each layer to each node so that when a service request arrives, each node distributes the service request based on the updated request distribution weights of each layer using a preset method.
[0126] In this embodiment, after updating the resource consumption mapping model to obtain the updated resource consumption mapping model, resource normalization is performed based on the updated resource consumption mapping models of various resources and the remaining amount of each type of resource to obtain the request capacity of each type of resource. Based on the request capacity of each type of resource, the resource bottlenecks of each node are identified hierarchically, and the resource bottlenecks of each node are summarized layer by layer upward according to the hierarchical relationship of each type of resource. Based on the resource bottlenecks of each node, the request distribution weights are calculated hierarchically to obtain the updated request distribution weights of each layer. The updated request distribution weights of each layer are distributed to each node so that when a service request arrives, each node distributes the service request according to the updated request distribution weights of each layer using a preset method.
[0127] Embodiment 2 of this invention provides a load balancing method for computing networks, including the update process of a resource consumption mapping model. This method collects computing network resource information and request status information to calculate the average resource consumption of a request. Furthermore, it can smoothly update the mapping model by combining historical data, thereby addressing existing problems such as inaccurate and fluctuating resource indicator systems and supporting load balancing decisions.
[0128] Example 3
[0129] Based on the technical solutions of the above embodiments, this invention provides several specific implementation methods.
[0130] As a specific implementation method of this embodiment, the load balancer collects resource information from the bottom up periodically, and then updates the request distribution weight based on the collected information and its own statistical information. After forming the request distribution strategy from the request distribution weight, the request distribution strategy is passed down. When a request arrives, the service request is distributed to each cluster based on the distribution strategy. Each cluster further distributes the request to each compute node based on the received request distribution strategy. Each compute node further distributes the request to each container within the node based on the request distribution strategy.
[0131] A load balancer consists of three main modules: a periodic feedback module, a dynamic adjustment module, and a request distribution module.
[0132] The periodic feedback module periodically collects resource status information (i.e., first resource information and second resource information). It can also calculate the current total number of requests being served by comparing the number of requests flowing into the load balancer (i.e., the total number of arriving requests), the total number of completed service requests, and the total number of service failure requests. This information is then aggregated into periodic feedback information and input into the dynamic adjustment module to formulate a request distribution strategy.
[0133] Figure 4 The calculation flowchart of the load balancer dynamic adjustment module provided in Embodiment 3 of the present invention is as follows: Figure 4 As shown, after receiving the periodic feedback information summarized by the periodic feedback module, the dynamic adjustment module updates the resource consumption mapping model based on resource usage, resource remaining amount, and the current number of requests being served; it performs resource normalization based on the resource consumption mapping model and resource remaining amount, and then identifies resource bottlenecks layer by layer and summarizes resource bottlenecks layer by layer based on resource affiliation, and calculates the request distribution weight factor for each layer.
[0134] The request distribution module is used to distribute requests based on the weighting factors of each layer, employing a smooth weighted round-robin strategy. This module exists not only in the load balancer but also in the cluster's control nodes and scheduler for request distribution at each layer.
[0135] As one specific implementation method of this embodiment, the specific process includes the following steps:
[0136] Step S1: After the computing network service is launched and the microservice container is deployed, collect the remaining amount of various resources and resource ownership relationships.
[0137] Based on resource hierarchies, the load balancer architecture diagram is transformed into a tree-like resource topology diagram. Load balancers, clusters, compute nodes, and containers are all represented by nodes in the tree diagram. For clarity and conciseness, clusters, compute nodes, and containers will all be referred to as nodes in the following text. Nodes are denoted by N. i,j Let j be the node number and i be the node degree. The universal set of all nodes is represented as: Right now From node N i,j To node N i+1,j′ The edge of E (i,j),(i+1,j′) The label represents node N. i+1,j′ For node N i,j The lower-level nodes. The complete set of all edges within the topology is represented by ε, i.e., E. (i,j),(i+1,j′) ∈v. The remaining resources on the node are determined by The tag indicates that resource usage is determined by... The label, k, is the resource label, i.e., node N. i,j There are k types of resources, and nodes at the same level have the same type of resources, while nodes at different levels have different types of resources. Node N i,j The complete collection of resources is represented as Business requests originate from the root node and are distributed to leaf nodes via child nodes. A distribution weight x needs to be configured for each edge. (i,j),(i+1,j′) Furthermore, the sum of the weights distributed from a parent node to all its child nodes is 1.
[0138] Step S2: Initialize the resource consumption mapping model.
[0139] The resource consumption mapping model can be initialized based on the business type to be the default mapping for that type of business, or it can be initialized based on the historical information of that business to be the historical model for that business. The standard form of the resource consumption mapping model is the amount of each resource consumed by a single service request, denoted as q. k .
[0140] Step S3: Normalize the remaining resource information.
[0141] Based on the resource consumption mapping model and resource remaining information, the remaining amount of each resource is normalized to the request capacity of each resource. The request capacity is the number of requests that the remaining amount of this type of resource can serve.
[0142] Resource normalization is performed using equation (1). In the equation... For node N i,j The number of requests that the remaining resources k can accommodate.
[0143]
[0144] Step S4: Based on the capacity of each resource request, identify resource bottlenecks layer by layer and aggregate them upwards.
[0145] Calculate node N using equation (2) or equation (3). i,j Resource bottleneck. In the formula, r... i,j Represents node N i,j Resource bottleneck. When the node is a leaf node, use equation (2) to calculate the resource bottleneck.
[0146]
[0147] When a node is a non-leaf node, use equation (3) to calculate the resource bottleneck of each node layer by layer upwards.
[0148]
[0149] Step S5: Calculate request distribution weights based on resource bottlenecks at each layer.
[0150] The request distribution weight is calculated hierarchically using equation (4). Where x... (i,j),(i+1,j′) Represents node N i,j To its child node N i+1,j′ The request distribution weight.
[0151]
[0152] Step S6: When a business request arrives, the business request is distributed based on the request distribution weight.
[0153] The calculated hierarchical request distribution weights are distributed to each node. When a business request arrives, each node distributes the request based on the distribution weights of its child nodes using a weighted round-robin algorithm or its derived strategy.
[0154] Step S7: Periodically update the resource consumption mapping model and request distribution weights.
[0155] After a certain period of time, the usage, remaining amount, and resource affiliation of various resources are collected again, along with the total number of arriving requests, the total number of completed service requests, and the total number of failed service requests. The current resource consumption mapping model is calculated using equation (5). Where... Represents the current resource consumption mapping for resource k, RE total Represents the total number of requests received, RE completed Represents the total number of completed service requests, RE failed This indicates the total number of failed service requests.
[0156]
[0157] After calculating the current resource consumption map, the resource consumption map is updated using the exponential smoothing method. The map is updated using Equation (6). In the equation, α represents the update factor, which needs to be preset according to the scenario. This factor determines the update rate and has a value range of [0,1].
[0158]
[0159] After updating the resource consumption mapping model, repeat steps S3 to S5 to update the request distribution weight, and use step S6 to distribute the business request when it arrives.
[0160] The load balancing method for computing networks provided in Embodiment 3 of this invention solves the load balancing problem of multi-layer structures in multi-dimensional resource scenarios by using methods such as resource normalization transformation, hierarchical resource aggregation, and resource bottleneck identification. By collecting computing network resource information and request status information, the average resource consumption of requests can be calculated, and the mapping model can be smoothly updated by combining historical data to solve existing problems such as inaccurate measurement and fluctuation of resource indicator system, thus supporting load balancing decision-making. Based on the resource consumption mapping model, the resource status is uniformly transformed into request capacity, and then resource bottlenecks are found hierarchically. On this basis, resource bottlenecks are aggregated upwards in layers, realizing accurate identification of resource bottlenecks in hierarchical structures. Subsequently, request distribution weights are formulated based on the resource bottlenecks of each layer, and request distribution is completed through a round-robin strategy, solving existing problems such as difficulty in unifying resource units, difficulty in measuring balancing indicators, and difficulty in coping with multi-layer structures.
[0161] Example 4
[0162] Figure 5 This is a schematic diagram of a load balancing device for a computing network provided in Embodiment 4 of the present invention. The device is applicable to the situation of balancing the load of a computing network in an environment with multiple clusters and multiple resources. The device can be implemented by software and / or hardware and is generally integrated on a load balancer.
[0163] like Figure 5 As shown, the device includes: a collection module 110, a normalization module 120, an identification module 130, a calculation module 140, and a distribution module 150.
[0164] Collection module 110 is used to collect first resource information, which includes the remaining amount of various resources and the subordinate relationships of various resources;
[0165] The normalization module 120 is used to normalize resources based on the initial resource consumption mapping model of various resources and the remaining amount of various resources to obtain the request capacity of various resources.
[0166] The identification module 130 is used to identify the resource bottlenecks of each node in a hierarchical manner based on the capacity to accommodate various types of resource requests, and to summarize the resource bottlenecks of each node layer by layer upward according to the hierarchical relationship of various types of resources.
[0167] The calculation module 140 is used to calculate the request distribution weight of each layer based on the resource bottleneck of each node.
[0168] The distribution module 150 is used to distribute the request distribution weights of each layer to each node, so that when a service request arrives, each node distributes the service request based on the request distribution weights of each layer and uses a preset method.
[0169] In this embodiment, the device first collects first resource information through the collection module 110, which includes the remaining amount of various resources and the subordinate relationships of various resources. Next, the normalization module 120 performs resource normalization based on the initial resource consumption mapping model of various resources and the remaining amount of various resources to obtain the request capacity of various resources. Then, the identification module 130 identifies the resource bottlenecks of each node hierarchically based on the request capacity of various resources, and summarizes the resource bottlenecks of each node layer by layer according to the subordinate relationships of various resources. Afterwards, the calculation module 140 calculates the request distribution weights of each layer based on the resource bottlenecks of each node to obtain the request distribution weights of each layer. Finally, the distribution module 150 distributes the request distribution weights of each layer to each node, so that when a service request arrives, each node distributes the service request according to the request distribution weights of each layer using a preset method.
[0170] This embodiment provides a load balancing device for computing networks, which can solve existing problems such as difficulty in unifying resource units, difficulty in measuring balancing indicators, and difficulty in dealing with multi-layered structures.
[0171] Furthermore, the overall architecture of the load balancer includes a load balancer, a cluster, compute nodes, and containers. Based on the hierarchical relationships of the various resources, the overall architecture of the load balancer is converted into a tree diagram representation, and the load balancer, cluster, compute nodes, and containers are converted into node representations. The load balancer is represented by a root node, the cluster and the compute nodes are represented by intermediate nodes, and the containers are represented by leaf nodes.
[0172] Each node is marked with a resource number, remaining resource quantity, and resource usage; nodes at the same level have the same set of resource numbers, representing the same type of resource.
[0173] Furthermore, the methods for determining the initial resource consumption mapping model for each type of resource include:
[0174] The resource consumption mapping model obtained by initializing based on the computing power network service type can be used as the initial resource consumption mapping model for various types of resources, or the resource consumption mapping model obtained by initializing based on the historical information of computing power network services can be used as the initial resource consumption mapping model for various types of resources.
[0175] Based on the above optimizations, the normalization module 120 is specifically used to: calculate the ratio of the remaining resource amount of the k-type resources to the initial resource consumption of the k-type resources for a node to obtain the request capacity of the k-type resources, wherein the initial resource consumption of the k-type resources is calculated through the initial resource consumption mapping model; and calculate the request capacity of each type of resource for each node in the above manner.
[0176] Based on the above optimizations, the identification module 130 is specifically used for: if the nodes included in the current layer are leaf nodes, calculating the minimum value among the various resource request capacity of the leaf node as the resource bottleneck of the leaf node; if the nodes included in the current layer are non-leaf nodes, calculating the minimum value among the various resource request capacity of the non-leaf node as the resource request capacity of the non-leaf node, calculating the sum of the resource bottlenecks of all lower-level nodes of the non-leaf node, comparing the resource request capacity of the non-leaf node with the sum of the resource bottlenecks of all lower-level nodes of the non-leaf node, and taking the smaller value of the two as the resource bottleneck of the non-leaf node.
[0177] Furthermore, the calculation module 140 is specifically used to: for a node in a layer, calculate the ratio of the resource bottleneck of a lower node of the node to the sum of the resource bottlenecks of all lower nodes of the node, and obtain the request distribution weight from the node to a lower node of the node; and use all the request distribution weights calculated in the layer as the request distribution weight of the layer.
[0178] Furthermore, the preset method includes a weighted round-robin algorithm or a derived strategy of the weighted round-robin algorithm.
[0179] Furthermore, the device also includes an update module, which includes:
[0180] The collection unit is used to collect second resource information and the total number of requests according to a preset period; wherein, the second resource information includes the remaining amount of various resources, the usage of various resources, and the subordinate relationship of various resources, and the total number of requests includes the total number of incoming requests, the total number of completed service requests, and the total number of service failure requests.
[0181] The calculation unit is used to calculate the current resource consumption mapping model of various resources based on the total number of requests and the resource usage of each node.
[0182] The update unit is used to update the initial resource consumption mapping model using the current resource consumption mapping model to obtain the updated resource consumption mapping model;
[0183] The return unit is used to return the updated resource consumption mapping model based on various types of resources and the remaining amount of the various types of resources to perform resource normalization to obtain the request capacity of various types of resources, and to perform the subsequent steps to update the request distribution weight of each layer.
[0184] The distribution unit is used to distribute the updated request distribution weights of each layer to each node, so that when a service request arrives, each node distributes the service request based on the updated request distribution weights of each layer using a preset method.
[0185] Furthermore, the calculation unit is specifically used to: calculate the sum of the usage of k types of resources on each node for k types of resources, and use the ratio of the sum of the usage of k types of resources on each node to the total number of currently served requests as the current resource consumption mapping model of k types of resources.
[0186] The total number of currently served requests is the difference between the total number of arriving requests, the total number of completed service requests, and the total number of failed service requests.
[0187] The load balancing device of the computing power network described above can execute the load balancing method of the computing power network provided in any embodiment of the present invention, and has the corresponding functional modules and beneficial effects of executing the method.
[0188] Example 5
[0189] Figure 6 A schematic diagram of a load balancer 10, which can be used to implement embodiments of the present invention, is shown. The load balancer is intended to represent various forms of digital computers, such as laptops, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframes, and other suitable computers. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the invention described and / or claimed herein.
[0190] like Figure 6As shown, the load balancer 10 includes at least one processor 11 and a memory, such as a read-only memory (ROM) 12 or a random access memory (RAM) 13, communicatively connected to the at least one processor 11. The memory stores computer programs executable by the at least one processor. The processor 11 can perform various appropriate actions and processes based on the computer program stored in the ROM 12 or loaded from storage unit 18 into the RAM 13. The RAM 13 can also store various programs and data required for the operation of the load balancer 10. The processor 11, ROM 12, and RAM 13 are interconnected via a bus 14. An input / output (I / O) interface 15 is also connected to the bus 14.
[0191] Multiple components in the load balancer 10 are connected to the I / O interface 15, including: input units 16, such as a keyboard, mouse, etc.; output units 17, such as various types of displays, speakers, etc.; storage units 18, such as disks, optical disks, etc.; and communication units 19, such as network interface cards, modems, wireless transceivers, etc. The communication unit 19 allows the load balancer 10 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.
[0192] Processor 11 can be a variety of general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of processor 11 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various processors running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. Processor 11 performs the various methods and processes described above, such as load balancing methods for computing networks.
[0193] In some embodiments, the load balancing method for the computing power network can be implemented as a computer program tangibly contained in a computer-readable storage medium, such as storage unit 18. In some embodiments, part or all of the computer program can be loaded and / or installed on the electronic device 10 via ROM 12 and / or communication unit 19. When the computer program is loaded into RAM 13 and executed by processor 11, one or more steps of the load balancing method for the computing power network described above can be performed. Alternatively, in other embodiments, processor 11 can be configured to perform the load balancing method for the computing power network by any other suitable means (e.g., by means of firmware).
[0194] Various embodiments of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-a-chip (SoCs), payload-programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting data and instructions to the storage system, the at least one input device, and the at least one output device.
[0195] Computer programs used to implement the methods of the present invention may be written in any combination of one or more programming languages. These computer programs may be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device, such that when executed by the processor, the computer programs cause the functions / operations specified in the flowcharts and / or block diagrams to be performed. The computer programs may be executed entirely on a machine, partially on a machine, or as a standalone software package, partially on a machine and partially on a remote machine, or entirely on a remote machine or server.
[0196] In the context of this invention, a computer-readable storage medium can be a tangible medium that may contain or store a computer program for use by or in conjunction with an instruction execution system, apparatus, or device. A computer-readable storage medium may include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination thereof. Alternatively, a computer-readable storage medium may be a machine-readable signal medium. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof.
[0197] To provide user interaction, the systems and techniques described herein can be implemented on a load balancer having: a display device for displaying information to the user (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor); and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the load balancer. Other types of devices can also be used to provide user interaction; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including sound input, voice input, or tactile input).
[0198] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as data servers), or computing systems that include middleware components (e.g., application servers), or computing systems that include frontend components (e.g., user computers with graphical user interfaces or web browsers through which users can interact with implementations of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication of any form or medium (e.g., communication networks). Examples of communication networks include local area networks (LANs), wide area networks (WANs), blockchain networks, and the Internet.
[0199] A computing system can include clients and servers. Clients and servers are generally located far apart and typically interact through communication networks. The client-server relationship is created by computer programs running on the respective computers and having a client-server relationship with each other. The server can be a cloud server, also known as a cloud computing server or cloud host, which is a hosting product within the cloud computing service system to address the shortcomings of traditional physical hosts and VPS services, such as high management difficulty and weak business scalability.
[0200] It should be understood that the various forms of processes shown above can be used, with steps reordered, added, or deleted. For example, the steps described in this invention can be executed in parallel, sequentially, or in different orders, as long as the desired result of the technical solution of this invention can be achieved, and this is not limited herein.
[0201] The specific embodiments described above do not constitute a limitation on the scope of protection of this invention. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this invention should be included within the scope of protection of this invention.
Claims
1. A load balancing method for a computing network, characterized in that, Applied to a load balancer, the overall architecture of which includes a load balancer, a cluster, compute nodes, and containers, the method includes: Collect first resource information, which includes the remaining quantity of various resources and the subordinate relationships of various resources; Based on the initial resource consumption mapping model of various resources and the remaining amount of various resources, resource normalization is performed to obtain the request capacity of various resources. The initial resource consumption mapping model is used to characterize the amount of each resource consumed by a single service request. Based on the resource request capacity of each type of resource, the resource bottlenecks of each node are identified hierarchically, and the resource bottlenecks of each node are summarized upwards layer by layer according to the resource hierarchy of each type of resource. The resource bottleneck is the minimum value among the resource request capacity of each node. The request distribution weights for each layer are obtained by calculating the request distribution weights based on the resource bottlenecks of each node. The request distribution weights of each layer are distributed to each node so that when a service request arrives, each node distributes the service request based on the request distribution weights of each layer using a preset method. The resource normalization based on the initial resource consumption mapping model for various resources and the remaining amount of various resources yields the resource request capacity for various resources, including: For a node k Class resources, calculate the k The remaining amount of the resource of the type and the above k The ratio of the initial resource consumption of the resource class can be obtained. k The request handling capacity of the resource class, the k The initial resource consumption of each type of resource is calculated using the initial resource consumption mapping model; the capacity to accommodate various resource requests for each node is calculated in the manner described above.
2. The method according to claim 1, characterized in that, The method further includes: converting the overall architecture of the load balancer into a tree diagram representation based on the various resource affiliations, and converting the load balancer, cluster, compute node, and container into node representations, with the load balancer represented by a root node, the cluster and the compute node represented by intermediate nodes, and the container represented by leaf nodes; Each node is marked with a resource number, remaining resource quantity, and resource usage; nodes at the same level have the same set of resource numbers, representing the same type of resource.
3. The method according to claim 1, characterized in that, The methods for determining the initial resource consumption mapping model for each type of resource include: The resource consumption mapping model obtained by initializing based on the computing power network service type can be used as the initial resource consumption mapping model for various types of resources, or the resource consumption mapping model obtained by initializing based on the historical information of computing power network services can be used as the initial resource consumption mapping model for various types of resources.
4. The method according to claim 2, characterized in that, The method of hierarchically identifying resource bottlenecks in each node based on the capacity to accommodate various resource requests includes: If the nodes included in the current layer are leaf nodes, calculate the minimum value among the various resource request capacity of the leaf node as the resource bottleneck of the leaf node; If the nodes included in the current layer are non-leaf nodes, calculate the minimum value among the resource request capacity of each type of non-leaf node as the resource request capacity of the non-leaf node, calculate the sum of the resource bottlenecks of all lower-level nodes of the non-leaf node, compare the resource request capacity of the non-leaf node with the sum of the resource bottlenecks of all lower-level nodes of the non-leaf node, and take the smaller value as the resource bottleneck of the non-leaf node.
5. The method according to claim 1, characterized in that, The process of calculating request distribution weights at each layer based on the resource bottlenecks of each node includes: For a node in a layer, calculate the ratio of the resource bottleneck of a lower-level node of the node to the sum of the resource bottlenecks of all lower-level nodes of the node to obtain the request distribution weight from the node to a lower-level node of the node; use all the request distribution weights calculated in the layer as the layer request distribution weight.
6. The method according to claim 1, characterized in that, The preset method includes a weighted round-robin algorithm or a derived strategy of the weighted round-robin algorithm.
7. The method according to claim 1, characterized in that, The method further includes: The second resource information and the total number of requests are collected according to a preset period; wherein, the second resource information includes the remaining amount of various resources, the usage of various resources, and the subordinate relationships of various resources, and the total number of requests includes the total number of incoming requests, the total number of completed service requests, and the total number of service failure requests. Based on the total number of requests and the usage of various resources at each node, calculate the current resource consumption mapping model for each type of resource; The updated resource consumption mapping model is obtained by updating the initial resource consumption mapping model using the current resource consumption mapping model. Return to the updated resource consumption mapping model based on various resources and the remaining amount of each type of resource to perform resource normalization to obtain the request capacity of each type of resource, and then execute the subsequent steps to update the request distribution weight of each layer. The updated request distribution weights for each layer are distributed to each node so that when a service request arrives, each node distributes the service request based on the updated request distribution weights for each layer using a preset method.
8. The method according to claim 7, characterized in that, The resource consumption mapping model for calculating the current resource consumption of various resources based on the total number of requests and the resource usage of each node includes: against k Class resources, compute on each node k The sum of the usage of each type of resource will be used to calculate the total usage of each node. k The ratio of the sum of resource usage across all classes to the total number of currently served requests is used as... k A mapping model of current resource consumption for resource classes; The total number of currently served requests is the difference between the total number of arriving requests, the total number of completed service requests, and the total number of failed service requests.
9. A load balancing device for a computing network, characterized in that, The load balancing device is applied to a load balancer, the overall architecture of which includes a load balancer, a cluster, compute nodes, and containers. The device includes: The collection module is used to collect first resource information, which includes the remaining amount of various resources and the subordinate relationships of various resources. The normalization module is used to normalize resources based on the initial resource consumption mapping model of various resources and the remaining amount of various resources to obtain the request capacity of various resources. The initial resource consumption mapping model is used to characterize the amount of each resource consumed by a single service request. The identification module is used to identify the resource bottlenecks of each node in a hierarchical manner based on the resource request capacity of each type, and to summarize the resource bottlenecks of each node layer by layer upward according to the resource hierarchy of each type. The resource bottleneck is the minimum value among the resource request capacity of each node. The calculation module is used to calculate the request distribution weights at each layer based on the resource bottlenecks of each node. The distribution module is used to distribute the request distribution weights of each layer to each node, so that when a service request arrives, each node distributes the service request based on the request distribution weights of each layer and uses a preset method. Specifically, the normalization module is used for: targeting a node k Class resources, calculate the k The remaining amount of the resource of the type and the above k The ratio of the initial resource consumption of the resource class can be obtained. k The request handling capacity of the resource class, the k The initial resource consumption of each type of resource is calculated using the initial resource consumption mapping model; the capacity to accommodate various resource requests for each node is calculated in the manner described above.
10. A load balancer, characterized in that, The load balancer includes: At least one processor; and a memory communicatively connected to the at least one processor; The memory stores a computer program that can be executed by the at least one processor, which is then executed by the at least one processor to enable the at least one processor to perform the load balancing method of the computing power network according to any one of claims 1-8.
11. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer instructions that are used to cause a processor to execute the load balancing method of the computing network according to any one of claims 1-8.
Citation Information
Patent Citations
Micro service gateway optimization method and apparatus, and storage medium
CN109618002A