Node load balancing method for container cloud
By building and training neural networks, the weighting coefficients of the load resource allocation nodes of the container cloud platform are dynamically calculated, and the problems of low resource utilization and single point failure in the existing technology are solved, achieving more efficient load allocation and platform stability improvement.
Patent Information
- Application Number
- CN202510305231.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-14
- Publication Date
- 2025-07-22
AI Technical Summary
The existing load balancing method cannot adapt to the dynamic changes of container cloud platforms, resulting in low resource utilization, extending service response time, and traditional Ingress controllers are prone to becoming a single point of failure.
By building a neural network including input layer, hidden layer and output layer, training and optimizing parameters, dynamically compute the weighting coefficients of the load resource allocation node, and combining node utilization and bottleneck thresholds, dynamic load resource allocation is achieved.
It improves resource utilization and overall performance and stability of container cloud platform, avoids single point of failure, and improves the flexibility and efficiency of load allocation.
Smart Images

Figure CN120353573A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of cloud computing, and more particularly relates to a node load balancing method for container cloud. Background Art
[0002] With the rapid development of cloud computing technology, containerized applications have become the mainstream trend. By packaging application programs and their dependencies in lightweight and portable containers, containerized applications enable rapid deployment and operation in any environment. Container cloud is becoming a key technology for enterprise digital transformation and cloud-native application development, indicating the future trend of software development and operation and maintenance management.
[0003] With the wide application of container technology and the rapid growth of the number of containers, how to efficiently manage and schedule these containers in a container cloud environment has become an important challenge. To improve the availability and response speed of services and avoid the risk of single-point failures, an effective load balancing method is needed to intelligently distribute the workload to multiple container instances running the same service, ensure reasonable utilization of resources, optimize the overall performance, and enhance the fault tolerance of the system. This not only involves the deployment and expansion of containers but also includes real-time monitoring and dynamic adjustment of the resource usage of containers during runtime to adapt to the changing workload requirements.
[0004] Most existing load balancing methods rely on static configurations or fixed load balancing strategies and cannot adapt to the dynamically changing requirements of container cloud platforms, resulting in low resource utilization and extended service response times. For example, traditional Ingress controllers are usually deployed as a single instance and are prone to becoming a single-point failure point of the system. At the same time, Ingress controllers mostly rely on static configurations and are difficult to dynamically adjust according to the actual running state of the container cloud platform. This leads to inflexible load distribution and inability to effectively respond to the dynamically changing load resource allocation requirements. Summary of the Invention
[0005] To solve the problem of low resource utilization caused by the current fixed load balancing strategy, the present invention provides a node load balancing method for container cloud to improve resource utilization.
[0006] To achieve the above technical effects, the technical solution of the present invention is as follows:
[0007] S1: The container cloud platform receives a load resource allocation request from the client and issues a load allocation requirement to the load resource allocation node;
[0008] S2: Determine whether the number of load resource allocation requirements received by the load resource allocation node reaches the set number threshold. If so, execute step S3; otherwise, perform load resource allocation according to the initial priority principle;
[0009] S3: Construct a neural network, which includes an input layer, a hidden layer and an output layer. Train the neural network to optimize the neural network parameters. After the training is completed, input several resource status parameters of the load resource allocation node into the trained neural network model, and output the weighting coefficients of several resource status parameters of the load resource allocation node.
[0010] S4: Based on the weighting coefficients and several resource status parameters of the load resource allocation node, calculate the relative closeness value of the load resource allocation node, sort the relative closeness values from high to low, and sort the load resource allocation nodes in the corresponding order of the sorting.
[0011] S5: According to the sorting order, determine whether the utilization rate of the current load resource allocation node exceeds the bottleneck threshold. If so, execute S6; otherwise, use the current load resource allocation node as the optimal load resource allocation node for load resource allocation.
[0012] S6: Use the load resource allocation node ranked one position after the current load resource allocation node as the optimal load resource allocation node, use the optimal load resource allocation node as the current load resource allocation node, and return to S5.
[0013] Further, the container cloud platform includes: multiple master nodes and multiple node nodes. The master nodes are responsible for load balancing calculation, and the node nodes are responsible for allocating loads according to the instructions of the master nodes. The container cloud platform communicates with the client through the wide area network WAN.
[0014] Further, the several resource status parameters of the load resource allocation node include: cpu utilization rate, network transmission bandwidth, memory utilization rate, storage capacity utilization rate, load distribution service request volume and load distribution service intensity.
[0015] Further, the process of allocating load resources according to the initial priority principle is as follows: compare the magnitudes of each resource status parameter of the load resource allocation nodes in turn. First, select the node with the lowest CPU utilization rate to allocate the load preferentially. When the CPU utilization rates of all load resource allocation nodes are the same, select the load resource allocation node with the lowest memory utilization rate to allocate the load; when the memory utilization rates of all load resource allocation nodes are the same, select the node with the lowest load allocation service request volume to allocate the load; when the load allocation service request volumes of all load resource allocation nodes are the same, select the node with the lowest load allocation service intensity to allocate the load; when the load allocation service intensities of all load resource allocation nodes are the same, select the node with the lowest storage capacity utilization rate to allocate the load; when the storage capacity utilization rates of all load resource allocation nodes are the same, select the node with the largest network transmission bandwidth to allocate the load; when the network transmission bandwidths of all load resource allocation nodes are the same, select the node with the smallest IP address network number to allocate the load; when the IP address network numbers of all load resource allocation nodes are the same, select the node with the smaller IP address host number to allocate the load.
[0016] Further, the process of training the neural network and optimizing the neural network parameters is as follows:
[0017] S301: Set the number threshold of load resource allocation requirements, and initialize the parameters of the neural network. The parameters of the neural network include: the connection weights from the hidden layer nodes to the output layer nodes, the connection weights from the input layer nodes to the hidden layer nodes, the scale factor, and the time translation factor; the load resource allocation requirements represent the data for training the neural network, and the number threshold of load resource allocation requirements represents the upper limit of the quantity of data for training the neural network, which is used to control the number of training times;
[0018] S302: Input several resource status parameters of the load resource allocation nodes into the neural network to obtain the weighted coefficients of several resource status parameters of the load resource allocation nodes;
[0019] S303: Use the loss function to evaluate the deviation between the weighted coefficients of the network output and the actual situation;
[0020] S304: Calculate the gradients of the loss function with respect to each neural network parameter through the backpropagation algorithm, and update the neural network parameters using the optimization algorithm according to the calculated gradients;
[0021] S305: Repeat S302 - S304 until the set number threshold of load resource allocation requirements is reached.
[0022] Further, the specific process of inputting several resource status parameters of the load resource allocation nodes and outputting the weighted coefficients of several resource status parameters of the load resource allocation nodes using the trained neural network model is as follows:
[0023] First, the input layer of the neural network receives several resource status parameters of the load resource allocation node. The input layer contains 6 nodes, which respectively receive the cpu utilization rate, network transmission bandwidth, memory utilization rate, storage capacity utilization rate, load allocation service request volume, and load allocation service intensity resource status parameters;
[0024] Next, the output of the input layer is input into the hidden layer, where the input resource status parameters are weighted and summed, and a non-linear transformation is performed through an activation function to extract the high-level representation of the input resource status parameters. The hidden layer includes 9 nodes. Among them, the output of the j-th node is expressed as:
[0025]
[0026] In the formula, represents the output of the j-th node in the hidden layer, h represents the activation function, represents the weighted input of the j-th node in the hidden layer, a j represents the scale factor, b j represents the time translation factor, w ej represents the connection weight from the r-th node in the input layer to the j-th node in the hidden layer, s represents the weighted input after scale scaling and time translation, represents the input of the r-th node in the input layer;
[0027] The initial weighted coefficients of the resource status parameters are output through the output layer. The output layer includes 6 nodes, and each node outputs a weighted coefficient of a resource status parameter. The output of the k-th node in the output layer is expressed as:
[0028]
[0029] In the formula, represents the output of the k-th node in the output layer, represents the output of the j-th node in the hidden layer, represents the weighted input of the k-th node in the output layer, w jk represents the connection weight from the j-th node in the hidden layer to the k-th node in the output layer;
[0030] Finally, a Softmax function is added after the output layer to normalize the initial weighted coefficients of the resource status parameters output by the output layer, and the weighted coefficients of the resource status parameters are output. The expression is:
[0031]
[0032] In the formula, Denote the k-th output of the neural network, i.e., the weighted coefficient of the k-th resource status parameter.
[0033] According to the above technical means, by constructing a neural network and training it to optimize the parameters, it is possible to dynamically generate the corresponding weighted coefficients according to the input load resource allocation node's resource status parameters, which shows the importance of different resource status parameters in the actual load allocation scenario, realizes dynamic resource allocation, and improves the flexibility and efficiency of resource allocation.
[0034] Further, when the time interval between two times when the container cloud platform receives the load resource allocation request from the client is less than 5 minutes, use the weighted coefficient calculated for the first time in the two times and do not recalculate the weighted coefficient; when the interval between two times when the container cloud platform receives the load resource allocation request from the client is greater than 5 minutes, recalculate the weighted coefficient.
[0035] Further, the calculation process of the relative closeness value of the load resource allocation node is as follows:
[0036] First, construct a decision matrix according to several resource status parameters of the load resource allocation node, and the expression is:
[0037]
[0038] In the formula, M represents the construction of the decision matrix, n represents the number of node nodes, R represents the cpu utilization rate, B represents the network transmission bandwidth, D represents the memory utilization rate, C represents the storage capacity utilization rate, F represents the load distribution service request volume, and Q represents the load distribution service intensity;
[0039] Then, perform normalization processing on the decision matrix to obtain a normalized decision matrix, and the expression is:
[0040]
[0041] In the formula, m uv represents the value of the u-th row and v-th column in the matrix, and M uv represents the value after normalization processing of the value of the u-th row and v-th column in the matrix;
[0042] Next, weight the normalized decision matrix according to the weighted coefficients obtained by neural network training to obtain a weighted decision matrix, and the expression is:
[0043] F = W v ×M uv , u = 1, 2, 3,..., n v = 1, 2, 3, 4, 5, 6
[0044] In the formula, F represents the weighted decision matrix, W vRepresents the weighted coefficient of the resource status parameter, that is, the weighted coefficient y of the resource status parameter output by the neural network k z set;
[0045] Define the vector composed of the maximum value of each parameter in the weighted decision matrix as the positive ideal value of the weighted decision matrix, and define the vector composed of the minimum value of each parameter in the weighted decision matrix as the negative ideal value of the weighted decision matrix. The expression is:
[0046]
[0047] In the formula, F + represents the positive ideal value, F - represents the negative ideal value, m uv represents the value of the u-th row and v-th column in the matrix;
[0048] Next, calculate the distances from the load resource allocation node to the positive and negative ideal values. The expression is:
[0049]
[0050]
[0051] In the formula, represents the distance from the u-th load resource allocation node to the positive ideal value, represents the distance from the u-th load resource allocation node to the negative ideal value;
[0052] Finally, calculate the relative closeness value of each load resource allocation node according to the distances from each load resource allocation node to the positive and negative ideal values. The expression is:
[0053]
[0054] In the formula, S u represents the relative closeness value of the u-th load resource allocation node.
[0055] Furthermore, when any one of the cpu utilization rate, memory utilization rate, and storage capacity utilization rate of the load resource allocation node reaches 70%, it means that the utilization rate of the load resource allocation node reaches the bottleneck threshold; otherwise, it means that the utilization rate of the load resource allocation node does not reach the bottleneck threshold.
[0056] Furthermore, calculate the network transmission bandwidth according to the bandwidths of all transmission links connected to the load resource allocation node. The expression is:
[0057]
[0058] Wherein, l represents the transmission link, L represents the set of all transmission links connected to the load resource allocation node, and B i represents the network transmission bandwidth, and B l represents the transmission bandwidth of the transmission link l, and α l represents a binary parameter, which is 1 when transmitting through link l and 0 otherwise;
[0059] The resource service intensity is calculated according to the ratio of the time for the load resource allocation node to complete the service request within a unit time to the unit time. The expression is:
[0060]
[0061] Wherein, Q i represents the resource service intensity of the i-th load resource allocation node, t represents the time for the i-th load resource allocation node to complete the service request, and T represents the unit time.
[0062] Compared with the prior art, the beneficial effects of this method are:
[0063] The present invention provides a node load balancing method for container cloud. This method receives the load resource allocation request from the client through the container cloud platform and issues a load allocation requirement to the load resource allocation node. When the number of load allocation requirements does not reach the set number threshold, the load resources are allocated according to the initial priority principle; when the threshold is reached, a neural network including an input layer, a hidden layer, and an output layer is constructed and trained. After optimizing the parameters, the resource status parameters of the nodes are input, the corresponding weighting coefficients are output, and the relative proximity values of the nodes are calculated for sorting. According to the sorting results, combined with the judgment of the bottleneck threshold of the node utilization rate, the optimal load resource allocation node is dynamically selected to complete the load resource allocation. As a whole, the present invention performs load resource allocation through a dynamic node load balancing method, improves the resource utilization efficiency, and thus enhances the overall performance and stability of the container cloud platform. BRIEF DESCRIPTION OF THE DRAWINGS
[0064] Figure 1 represents the flowchart of the node load balancing method for container cloud proposed in the embodiment of the present invention;
[0065] Figure 2 represents the structural schematic diagram of the container cloud platform proposed in the embodiment of the present invention;
[0066] Figure 3 represents the structural diagram of the neural network proposed in the embodiment of the present invention;
[0067] Figure 4 represents the flow block diagram of training the neural network proposed in the embodiment of the present invention. DETAILED DESCRIPTION OF THE INVENTION
[0068] The accompanying drawings are only for illustrative purposes and should not be construed as limiting the patent.
[0069] To better illustrate this embodiment, some parts of the accompanying drawings are omitted, enlarged or reduced, which do not represent the actual size.
[0070] For those skilled in the art, it is understandable that some well-known content descriptions in the accompanying drawings may be omitted.
[0071] The technical solutions of the present invention will be further described below in conjunction with the accompanying drawings and embodiments.
[0072] The description of the positional relationship in the accompanying drawings is only for illustrative purposes and should not be construed as limiting the patent.
[0073] Embodiment 1
[0074] This embodiment proposes a method for node load balancing in a container cloud. As shown in the flowchart of the method, the method proposed in this embodiment generally includes the following steps: Figure 1 S1: The container cloud platform receives a load resource allocation request from the client and issues a load allocation requirement to the load resource allocation node.
[0075] S2: Determine whether the number of load resource allocation requirements received by the load resource allocation node reaches the set number threshold. If so, execute step S3; otherwise, perform load resource allocation according to the initial priority principle.
[0076] S3: Construct a neural network, which includes an input layer, a hidden layer and an output layer. Train the neural network, optimize the neural network parameters. After training is completed, input several resource status parameters of the load resource allocation node into the trained neural network model, and output the weighted coefficients of several resource status parameters of the load resource allocation node.
[0077] S3: Build a neural network, which includes an input layer, a hidden layer and an output layer. Train the neural network to optimize the neural network parameters. After training is completed, input several resource status parameters of the load resource allocation node into the trained neural network model, and output the weighted coefficients of several resource status parameters of the load resource allocation node.
[0078] S4: Calculate the relative closeness value of the load resource allocation node based on the weighted coefficients and several resource status parameters of the load resource allocation node. Sort the relative closeness values from high to low, and sort the load resource allocation nodes in the corresponding order of the sorting.
[0079] S5: According to the sorting order, determine whether the utilization rate of the current load resource allocation node exceeds the bottleneck threshold. If so, execute S6; otherwise, use the current load resource allocation node as the optimal load resource allocation node to perform load resource allocation.
[0080] S6: Use the load resource allocation node ranked one position after the current load resource allocation node as the optimal load resource allocation node, use the optimal load resource allocation node as the current load resource allocation node, and return to S5.
[0081] In this embodiment, as Figure 2 shown in the structural schematic diagram of the container cloud platform, the container cloud platform includes: a plurality of master nodes and a plurality of node nodes. The master nodes are responsible for performing load balancing calculations, while the node nodes perform load distribution according to the instructions of the master nodes, realizing the separation of calculation and execution. At the same time, the container cloud platform communicates with the client through a wide area network (WAN), enabling the container cloud platform to support cross-regional distributed deployment.
[0082] In this embodiment, several resource status parameters of the load resource allocation node include: CPU utilization rate, network transmission bandwidth, memory utilization rate, storage capacity utilization rate, load distribution service request volume, and load distribution service intensity.
[0083] The network transmission bandwidth is calculated according to the bandwidths of all transmission links connected to the load resource allocation node. The expression is:
[0084]
[0085] In the formula, l represents a transmission link, L represents the set of all transmission links connected to the load resource allocation node, B i represents the network transmission bandwidth, B l represents the transmission bandwidth of transmission link l, and α l represents a binary parameter, which is 1 when transmitting through link l, and 0 otherwise;
[0086] The resource service intensity is calculated according to the ratio of the time taken by the load resource allocation node to complete a service request in a unit time to the unit time. The expression is:
[0087]
[0088] In the formula, Q i represents the resource service intensity of the i-th load resource allocation node, t represents the time taken by the i-th load resource allocation node to complete a service request, and T represents the unit time.
[0089] The CPU utilization rate is calculated according to the CPU resources already used by the load resource allocation node and the total CPU resources of the node. The expression is:
[0090]
[0091] In the formula, R i represents the CPU utilization rate, U used represents the used CPU resources, and U total represents the total CPU resources.
[0092] Calculate the memory utilization rate based on the memory resources already used by the load resource allocation node and the total memory resources of the node. The expression is as follows:
[0093]
[0094] In the formula, D i represents the memory resource utilization rate, M used represents the used memory resources, M total represents the total memory resources.
[0095] Calculate the storage capacity utilization rate based on the storage capacity already used by the load resource allocation node and the total storage capacity of the node. The expression is as follows:
[0096]
[0097] In the formula, C i represents the storage capacity utilization rate, S used represents the used storage capacity, S total represents the total storage capacity.
[0098] Indicate the load distribution service request volume F based on the number of resource service requests received by the load resource allocation node within unit time T i .
[0099] In this embodiment, the process of load resource allocation according to the initial priority principle is as follows: successively compare the magnitudes of each resource status parameter of the load resource allocation nodes. First, select the node with the minimum CPU utilization rate to preferentially allocate the load. When the CPU utilization rates of all load resource allocation nodes are the same, select the load resource allocation node with the minimum memory utilization rate to allocate the load; when the memory utilization rates of all load resource allocation nodes are the same, select the node with the minimum load distribution service request volume to allocate the load; when the load distribution service request volumes of all load resource allocation nodes are the same, select the node with the minimum load distribution service intensity to allocate the load; when the load distribution service intensities of all load resource allocation nodes are the same, select the node with the minimum storage capacity utilization rate to allocate the load; when the storage capacity utilization rates of all load resource allocation nodes are the same, select the node with the maximum network transmission bandwidth to allocate the load; when the network transmission bandwidths of all load resource allocation nodes are the same, select the node with the minimum IP address network number to allocate the load; when the IP address network numbers of all load resource allocation nodes are the same, select the node with the minimum IP address host number to allocate the load.
[0100] In this embodiment, when any one of the CPU utilization rate, memory utilization rate, and storage capacity utilization rate of the current load resource allocation node reaches 70%, no more load is allocated to this node. If the resource service requirements of the client are not yet met, the remaining resources are allocated to the node with the highest priority among the remaining nodes according to the priority principle until all resources are allocated.
[0101] Embodiment 2
[0102] In this embodiment, the process of constructing and training a neural network is described in detail. As Figure 3 shown in the structure diagram of the neural network, the neural network includes an input layer, a hidden layer, and an output layer. The input layer of the neural network receives several resource status parameters of the load resource allocation node. The hidden layer performs weighted summation on the input resource status parameters and performs a non-linear transformation through an activation function to extract a high-level representation of the input resource status parameters. The output layer outputs the initial weighted coefficients of the resource status parameters, and finally performs normalization processing through the Softmax function to output the weighted coefficients of the resource status parameters.
[0103] In this embodiment, as Figure 4 shown in the flowchart of training the neural network, the process of training the neural network and optimizing the neural network parameters is as follows:
[0104] S301: Set the number threshold of the load resource allocation requirements, and initialize the parameters of the neural network. The parameters of the neural network include: the connection weights from the hidden layer nodes to the output layer nodes, the connection weights from the input layer nodes to the hidden layer nodes, the scale stretching factor, and the time translation factor; the load resource allocation requirements represent the data for training the neural network, and the number threshold of the load resource allocation requirements represents the upper limit of the number of data for training the neural network, which is used to control the number of training times.
[0105] Exemplarily, set the number threshold to 500 to train the neural network. Use the number threshold to control the number of training times to ensure that the neural network is trained after accumulating enough samples, thereby improving the training effect. The parameters of the neural network are randomly initialized, and the initial values are randomly selected from a uniform distribution or a normal distribution.
[0106] S302: Input several resource status parameters of the load resource allocation node into the neural network to obtain the weighted coefficients of several resource status parameters of the load resource allocation node.
[0107] S303: Use the loss function to evaluate the deviation between the weighted coefficients output by the network and the actual situation.
[0108] Evaluate the importance of each resource status parameter based on the actual load distribution effect, so as to obtain the weighted coefficients in the actual situation. Exemplarily, in actual load distribution, the load resource allocation node is selected for allocation because the CPU utilization rate of this node is the lowest. Then, the weighted coefficient of the CPU utilization rate is 1, and the weighted coefficients of other parameters are 0. If the CPU utilization rates are the same and the node with the lowest memory utilization rate is selected, then the weighted coefficients of both the CPU utilization rate and the memory utilization rate are set to 0.5, and the weighted coefficients of the remaining parameters are 0. The greater the impact of a resource status parameter on load distribution, the greater the corresponding weighted coefficient; the smaller the impact of a resource status parameter on load distribution, the smaller the corresponding weighted coefficient.
[0109] S304: Calculate the gradient of the loss function with respect to each neural network parameter through the backpropagation algorithm. According to the calculated gradient, use an optimization algorithm to update the neural network parameters.
[0110] Specifically, calculate the gradient of the loss function with respect to each neural network parameter. The backpropagation algorithm calculates the gradient layer by layer through the chain rule, propagating backward from the output layer to the input layer. Use an optimization algorithm to update the neural network parameters according to the calculated gradient. The optimization algorithms include gradient descent method, stochastic gradient descent method SGD, Adam optimization algorithm, etc.
[0111] S305: Repeat S302 - S304 until the number threshold required for the set load resource allocation is reached.
[0112] In this embodiment, repeat the steps of S302 - S304 until the accumulated number reaches the set threshold. During each training process, the parameters of the neural network will be continuously updated to improve the performance of the neural network. When the number threshold is reached, the training of the neural network is completed. The trained neural network can output the weighted coefficients according to the input resource status parameters for the calculation of the relative proximity value.
[0113] The specific process of inputting several resource status parameters of the load resource allocation node and using the trained neural network model to output the weighted coefficients of several resource status parameters of the load resource allocation node is as follows:
[0114] First, the input layer of the neural network receives several resource status parameters of the load resource allocation node. The input layer contains 6 nodes, which respectively receive the resource status parameters of CPU utilization rate, network transmission bandwidth, memory utilization rate, storage capacity utilization rate, load distribution service request volume, and load distribution service intensity.
[0115] Next, the output of the input layer is input into the hidden layer, where the weighted sum of the input resource state parameters is calculated, and a non-linear transformation is performed through an activation function to extract the high-level representation of the input resource state parameters. The hidden layer includes 9 nodes. Among them, the output of the j-th node is expressed as:
[0116]
[0117] In the formula, represents the output of the j-th node in the hidden layer, h represents the activation function, represents the weighted input of the j-th node in the hidden layer, a j represents the scaling factor, b j represents the time translation factor, w ej represents the connection weight from the r-th node of the input layer to the j-th node of the hidden layer, s represents the weighted input after scaling and time translation, represents the input of the r-th node of the input layer, and z represents the number threshold of the set load resource allocation requirements;
[0118] The initial weighted coefficients of the resource state parameters are output through the output layer. The output layer includes 6 nodes, and each node outputs a weighted coefficient of a resource state parameter. The output of the k-th node of the output layer is expressed as:
[0119]
[0120] In the formula, represents the output of the k-th node of the output layer, represents the output of the j-th node of the hidden layer, represents the weighted input of the k-th node of the output layer, w jk represents the connection weight from the j-th node of the hidden layer to the k-th node of the output layer;
[0121] Finally, a Softmax function is added after the output layer to normalize the initial weighted coefficients of the resource state parameters output by the output layer, and the weighted coefficients of the resource state parameters are output. The expression is:
[0122]
[0123] In the formula, represents the k-th output of the neural network, that is, the weighted coefficient of the k-th resource state parameter. Using the weighted coefficient indicates that different resource state parameters have different influences on the selection of the load resource allocation node.
[0124] Exemplarily, to reduce the computing burden on the master node, when the time interval between two receptions of the load resource allocation request from the client by the container cloud platform is less than 5 minutes, the weighted coefficient calculated for the first time in the two times is used, and the weighted coefficient is not recalculated; when the interval between two receptions of the load resource allocation request from the client by the container cloud platform is greater than 5 minutes, the weighted coefficient is recalculated.
[0125] Embodiment 3
[0126] In this embodiment, the calculation process of the relative closeness value of the load resource allocation node is described in detail. The relative closeness value provides a dynamic evaluation mechanism. According to the resource status and load conditions of the current load resource allocation node, the most suitable load resource allocation node is selected in real time to allocate new load tasks. This method avoids the limitations of static configuration and better adapts to the dynamically changing requirements of the container cloud platform.
[0127] Exemplarily, when any one of the cpu utilization rate, memory utilization rate, and storage capacity utilization rate of the load resource allocation node reaches 70%, it indicates that the utilization rate of the load resource allocation node reaches the bottleneck threshold; otherwise, it indicates that the utilization rate of the load resource allocation node does not reach the bottleneck threshold. By setting the resource utilization threshold, excessive load is avoided from being concentrated on a certain load resource allocation node, thereby preventing the occurrence of a single-point bottleneck.
[0128] In this embodiment, the calculation process of the relative closeness value of the load resource allocation node is as follows:
[0129] First, a decision matrix is constructed according to several resource status parameters of the load resource allocation node, and the expression is:
[0130]
[0131] In the formula, M represents the construction of the decision matrix, n represents the number of node nodes, R represents the cpu utilization rate, B represents the network transmission bandwidth, D represents the memory utilization rate, C represents the storage capacity utilization rate, F represents the load distribution service request volume, and Q represents the load distribution service intensity;
[0132] Then, the decision matrix is normalized to obtain a normalized decision matrix, and the expression is:
[0133]
[0134] In the formula, m uv represents the value of the u-th row and v-th column in the matrix, and M uv represents the value after the normalization process of the value of the u-th row and v-th column in the matrix;
[0135] Next, the normalized decision matrix is weighted according to the weighted coefficients obtained from neural network training to obtain a weighted decision matrix, and the expression is:
[0136] F = W v × M uv , u = 1, 2, 3, …, n v = 1, 2, 3, 4, 5, 6
[0137] In the formula, F represents the weighted decision matrix, and W v represents the weighted coefficient of the resource state parameter, that is, the set of weighted coefficients of the resource state parameter output by the neural network ;
[0138] Define the vector composed of the maximum value of each parameter in the weighted decision matrix as the positive ideal value of the weighted decision matrix, and define the vector composed of the minimum value of each parameter in the weighted decision matrix as the negative ideal value of the weighted decision matrix. The expression is:
[0139]
[0140] In the formula, F + represents the positive ideal value, F - represents the negative ideal value, and m uv represents the value of the u-th row and v-th column in the matrix;
[0141] Next, calculate the distances from the load resource allocation node to the positive and negative ideal values. The expression is:
[0142]
[0143] In the formula, represents the distance from the u-th load resource allocation node to the positive ideal value, represents the distance from the u-th load resource allocation node to the negative ideal value;
[0144] The expression for the set of distances from each load resource allocation node to the positive and negative ideal values is:
[0145]
[0146] In the formula, Y + represents the set of distances from the load resource allocation node to the positive ideal value, and Y - represents the set of distances from the load resource allocation node to the negative ideal value;
[0147] Finally, calculate the relative closeness value of each load resource allocation node according to the distances from each load resource allocation node to the positive and negative ideal values. The expression is:
[0148]
[0149] In the formula, Su Represents the relative proximity value of the u-th load resource allocation node.
[0150] The embodiments are merely examples given to clearly illustrate the present invention and are not intended to limit the implementation manners of the present invention. For those of ordinary skill in the art, other different forms of changes or modifications can be made based on the above description. It is not necessary and impossible to enumerate all the implementation manners here. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the claims of the present invention.
Claims
1. A method for node load balancing in a container cloud, characterized in that, It includes the following steps: S1: The container cloud platform receives the load resource allocation request from the client and issues a load allocation requirement to the load resource allocation node; S2: Determine whether the number of load resource allocation requirements received by the load resource allocation node reaches the set number threshold. If so, execute step S3; otherwise, perform load resource allocation according to the initial priority principle; S3: Construct a neural network, which includes an input layer, a hidden layer, and an output layer. Train the neural network, optimize the neural network parameters. After training, input several resource status parameters of the load resource allocation node into the trained neural network model, and output the weighted coefficients of several resource status parameters of the load resource allocation node; S4: Based on the weighted coefficients and several resource status parameters of the load resource allocation node, calculate the relative closeness value of the load resource allocation node, sort the relative closeness values from high to low, and sort the load resource allocation nodes in the corresponding sorting order; S5: According to the sorting order, determine whether the utilization rate of the current load resource allocation node exceeds the bottleneck threshold. If so, execute S6; otherwise, use the current load resource allocation node as the optimal load resource allocation node to perform load resource allocation; S6: Use the load resource allocation node ranked one position after the current load resource allocation node as the optimal load resource allocation node, use the optimal load resource allocation node as the current load resource allocation node, and return to S5.
2. The node load balancing method for a container cloud according to claim 1, wherein The container cloud platform includes: multiple master nodes and multiple node nodes. The master nodes are responsible for load balancing calculation, and the node nodes are responsible for allocating loads according to the instructions of the master nodes; the container cloud platform communicates with the client through the wide area network WAN.
3. A method for node load balancing in a container cloud according to claim 1, characterized in that, The several resource status parameters of the load resource allocation node include: cpu utilization rate, network transmission bandwidth, memory utilization rate, storage capacity utilization rate, load allocation service request volume, and load allocation service intensity.
4. A method for node load balancing in a container cloud according to claim 3, characterized in that The process of performing load resource allocation according to the initial priority principle is as follows: Compare the sizes of each resource status parameter of the load resource allocation node in turn. First, select the node with the smallest cpu utilization rate to allocate loads preferentially. When the cpu utilization rates of all load resource allocation nodes are the same, select the load resource allocation node with the smallest memory utilization rate to allocate loads; when the memory utilization rates of all load resource allocation nodes are the same, select the node with the smallest load allocation service request volume to allocate loads; when the load allocation service request volumes of all load resource allocation nodes are the same, select the node with the smallest load allocation service intensity to allocate loads; when the load allocation service intensities of all load resource allocation nodes are the same, select the node with the smallest storage capacity utilization rate to allocate loads; when the storage capacity utilization rates of all load resource allocation nodes are the same, select the node with the largest network transmission bandwidth to allocate loads; when the network transmission bandwidths of all load resource allocation nodes are the same, select the node with the smallest IP address network number to allocate loads; when the IP address network numbers of all load resource allocation nodes are the same, select the node with the smallest IP address host number to allocate loads.
5. A method for node load balancing in a container cloud according to claim 1, characterized in that, The process of training a neural network and optimizing its parameters is as follows: S301: Set the number threshold of load resource allocation requirements, and initialize the parameters of the neural network. The parameters of the neural network include: the connection weights from the hidden layer nodes to the output layer nodes, the connection weights from the input layer nodes to the hidden layer nodes, the scale factor, and the time translation factor. The load resource allocation requirements represent the data for training the neural network, and the number threshold of load resource allocation requirements represents the upper limit of the quantity of data for training the neural network, which is used to control the number of training times. S302: Input several resource status parameters of the load resource allocation node into the neural network to obtain the weighted coefficients of several resource status parameters of the load resource allocation node. S303: Use the loss function to evaluate the deviation between the weighted coefficients output by the network and the actual situation. S304: Calculate the gradient of the loss function with respect to each neural network parameter through the backpropagation algorithm, and update the neural network parameters using the optimization algorithm according to the calculated gradients. S305: Repeat S302 - S304 until the set number threshold of load resource allocation requirements is reached.
6. The node load balancing method for a container cloud according to claim 3, wherein The specific process of inputting several resource status parameters of the load resource allocation node and using the trained neural network model to output the weighted coefficients of several resource status parameters of the load resource allocation node is as follows: First, the input layer of the neural network receives several resource status parameters of the load resource allocation node. The input layer contains 6 nodes, which respectively receive the resource status parameters of cpu utilization rate, network transmission bandwidth, memory utilization rate, storage capacity utilization rate, load allocation service request volume, and load allocation service intensity. Next, the output of the input layer is input into the hidden layer, and the input resource status parameters are weighted and summed, and a non-linear transformation is performed through the activation function to extract the high-level representation of the input resource status parameters. The hidden layer includes 9 nodes. Among them, the output of the j-th node is expressed as: wherein, represents the output of the j-th node in the hidden layer, h represents the activation function, represents the weighted input of the j-th node in the hidden layer, a j represents the scaling factor, b j represents the time translation factor, w ej represents the connection weight from the r-th node in the input layer to the j-th node in the hidden layer, s represents the weighted input after scaling and time translation, represents the input of the r-th node in the input layer; The initial weighted coefficients of the resource status parameters are output through the output layer. The output layer includes 6 nodes, and each node outputs a weighted coefficient of a resource status parameter. The output of the k-th node in the output layer is expressed as: wherein, represents the output of the k-th node in the output layer, represents the output of the j-th node in the hidden layer, represents the weighted input of the k-th node in the output layer, w jk represents the connection weight from the j-th node in the hidden layer to the k-th node in the output layer; Finally, a Softmax function is added after the output layer to normalize the initial weighted coefficients of the resource status parameters output by the output layer, and the weighted coefficients of the resource status parameters are output. The expression is: In the formula, represents the k-th output of the neural network, that is, the weighting coefficient of the k-th resource state parameter.
7. A method for node load balancing in a container cloud according to claim 1, characterized in that When the time interval between two receptions of the load resource allocation request from the client by the container cloud platform is less than 5 minutes, the weighted coefficients calculated in the first of the two times are used, and the weighted coefficients are not recalculated. When the interval between two receptions of the load resource allocation request from the client by the container cloud platform is greater than 5 minutes, the weighted coefficients are recalculated.
8. A method for node load balancing of a container cloud according to claim 3, characterized in that, The calculation process of the relative closeness value of the load resource allocation node is as follows: First, construct a decision matrix according to several resource status parameters of the load resource allocation node. The expression is: Wherein, M represents the construction of the decision matrix, n represents the number of node nodes, R represents the CPU utilization rate, B represents the network transmission bandwidth, D represents the memory utilization rate, C represents the storage capacity utilization rate, F represents the service request volume of load distribution, and Q represents the service intensity of load distribution; Then, the decision matrix is normalized to obtain a normalized decision matrix, and the expression is: where m uv represents the value of the element in the u-th row and v-th column of the matrix, and M uv represents the value after normalizing the value of the element in the u-th row and v-th column of the matrix; Next, the normalized decision matrix is weighted according to the weighted coefficients obtained by neural network training to obtain a weighted decision matrix, and the expression is: F = W v × M uv , u = 1, 2, 3, …, n v = 1, 2, 3, 4, 5, 6 where F represents the weighted decision matrix and W v represents the weighted coefficient of the resource state parameter, that is, the weighted coefficient of the resource state parameter output by the neural network set; The vector composed of the maximum values of each parameter in the weighted decision matrix is defined as the positive ideal value of the weighted decision matrix, and the vector composed of the minimum values of each parameter in the weighted decision matrix is defined as the negative ideal value of the weighted decision matrix, and the expression is: Where, F + represents the positive ideal value, F - represents the negative ideal value, and m uv represents the value at the u-th row and v-th column in the matrix; Next, calculate the distances from the load resource allocation nodes to the positive and negative ideal values, and the expression is: wherein, represents the distance from the u-th load resource allocation node to the positive ideal value, represents the distance from the u-th load resource allocation node to the negative ideal value; Finally, calculate the relative closeness value of each load resource allocation node according to the distances from each load resource allocation node to the positive and negative ideal values, and the expression is: Where S u represents the relative proximity value of the u-th load resource allocation node.
9. The node load balancing method for a container cloud according to claim 2, wherein When any one of the CPU utilization rate, memory utilization rate, and storage capacity utilization rate of the load resource allocation node reaches 70%, it means that the utilization rate of the load resource allocation node reaches the bottleneck threshold; otherwise, it means that the utilization rate of the load resource allocation node does not reach the bottleneck threshold.
10. A method for node load balancing of a container cloud according to claim 3, characterized in that, Calculate the network transmission bandwidth according to the bandwidths of all transmission links connected to the load resource allocation node, and the expression is: Wherein, l represents a transmission link, L represents the set of all transmission links connected to the load resource allocation node, B i represents the network transmission bandwidth, B l represents the transmission bandwidth of the transmission link l, α l represents a binary parameter, which is 1 when transmitted through the link l and 0 otherwise; Calculate the resource service intensity according to the ratio of the time taken by the load resource allocation node to complete the service request within a unit time to the unit time, and the expression is: Where Q i represents the resource service intensity of the i-th load resource allocation node, t represents the time for the i-th load resource allocation node to complete a service request, and T represents the unit time.