A method, device, computer equipment and storage medium for scheduling container instances

By scheduling and eviction of container instances based on predicted load in the container cluster, the problem of inability to schedule in time when pressure ejection of container cluster nodes is solved, the failure probability and cost are reduced, and the scheduling efficiency and system availability are improved.

CN119065795BActive Publication Date: 2025-05-06INSPUR SUZHOU INTELLIGENT TECH CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202411537025.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-10-31
Publication Date
2025-05-06
Estimated Expiration
2044-10-31

AI Technical Summary

Technical Problem

When the container cluster node is driven out of pressure, the container instance cannot be scheduled in time, resulting in a high probability of cluster failure, high cost of container instance scheduling, and low efficiency.

Method used

By determining the total predicted load of each host node and the predicted load of each container instance in the host node of the container cluster, comparing the total predicted load and the preset load, determining whether the container instance needs to be scheduled, and expelling and scheduling according to the predicted load, ensuring that the load of the host node is within the preset threshold.

Benefits of technology

Timely scheduling is realized when pressure evicting the container cluster nodes is reduced, the probability of cluster failure is improved, the efficiency of container instance scheduling is reduced, and the cost is reduced, and the system availability is improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119065795B_ABST
    Figure CN119065795B_ABST
Patent Text Reader

Abstract

The present invention relates to the field of computer technology, and discloses a method, device, computer equipment and storage medium for scheduling container instances, the method comprising: determining the total predicted load of each host node and the predicted load of each container instance on each host node in at least one host node included in a container cluster; judging whether there is a host node whose total predicted load is greater than the preset load by comparing the total predicted load of each host node with a preset load; if there is a first host node whose total predicted load is greater than the preset load, determining the first container instance to be scheduled according to the total predicted load and the preset load of the first host node; expelling the first container instance from the first host node, and scheduling the first container instance to a second host node. The present invention can solve the problem that when container cluster nodes are expelled due to pressure, container instances cannot be scheduled in time, resulting in a high probability of cluster failure.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer technology, and in particular to a method, device, computer equipment and storage medium for scheduling a container instance. Background Art

[0002] With the development of distributed storage systems and container technology, the demand for container orchestration has become increasingly obvious, leading to the birth of container orchestration tools. All containers managed by container orchestration tools are called container clusters. In order to ensure the normal operation of each host node in the container cluster, the container orchestration tool has created a mechanism - pressure eviction mechanism, that is, when the memory, disk and other resources of a node are insufficient, one or more container instances running on this node will be randomly evicted to ensure that the operation of other container instances and container orchestration programs on the node is not affected. This random behavior may lead to the deletion of some important services, resulting in a decrease in throughput, slow response, and system jams. The evicted (i.e. deleted) container instances will randomly try to pull up on other host nodes. If the re-allocated host still has insufficient resources, it will be repeatedly evicted, which may cause the container to be unable to be scheduled, or it will take multiple evictions to find a suitable node, which is inefficient.

[0003] In order to deal with the problem of node pressure expulsion in container clusters, it is usually necessary to manually monitor the container instances running on each node. When a container instance is found to be expelled, it is necessary to observe whether the expelled container instance can be successfully pulled up again, and whether the secondary allocated host node has sufficient resources to ensure that it will not be expelled again; and when there are no suitable nodes, the cluster nodes are expanded. This method causes the container instance to be unable to be scheduled in time, resulting in a high probability of cluster failure; and there is no proactive warning in advance. Since the operation and maintenance personnel need to pay attention to the cluster expulsion status at all times, the cost of container instance scheduling is high and the efficiency is low. Summary of the invention

[0004] In view of this, the present invention provides a scheduling method, apparatus, computer equipment and storage medium for container instances to solve the problem that when container cluster nodes are pressured and evicted, container instances cannot be scheduled in time, resulting in a high probability of cluster failure, high cost and low efficiency of container instance scheduling.

[0005] In a first aspect, the present invention provides a method for scheduling container instances, the method comprising: determining, in at least one host node included in a container cluster, a total predicted load of each host node and a predicted load of each container instance on each host node, wherein each host node can run at least one container instance, each container instance corresponds to a predicted load when running, and the total predicted load is the sum of the predicted loads of all container instances running on each host node; by comparing the total predicted load of each host node with a preset load of each host node, determining whether there is a host node whose total predicted load is greater than the preset load; if so, determining, based on the total predicted load and the preset load of the first host node, a first container instance that needs to be scheduled by the first host node, the first host node being a host node in at least one host node; expelling the first container instance from the first host node, and scheduling the first container instance to a second host node other than the first host node in at least one host node, the second host node being a host node in the container cluster that can run the first container instance.

[0006] Based on the method of the first aspect above, the total predicted load of each host node included in the container cluster can be determined, and by comparing the total predicted load with the preset load, the first host node that needs to schedule the container instance and the first container instance that the first host node needs to schedule are determined, and then the first container instance is expelled from the first host node. Since the total predicted load of each host node is determined, and the total predicted load is the total predicted load of each host node at a future moment, the first host node and the first container instance that need to be scheduled can be determined in advance based on the predicted total predicted load, thereby avoiding the problem of not being able to schedule the container instance in time when the container cluster node is expelled due to pressure, thereby reducing the probability of cluster failure, improving the efficiency of container instance scheduling, reducing the cost of container instance scheduling, and improving the availability of the system.

[0007] In an optional embodiment, in at least one host node included in the container cluster, the total predicted load of each host node and the predicted load of each container instance on each host node are determined, including: obtaining the load of each container instance among all container instances running on each host node in the current cycle; inputting the load into a target prediction model and outputting the predicted load, where the target prediction model is used to predict the predicted load of each container instance in the next cycle; and calculating the sum of the predicted loads of each container instance to obtain the total predicted load of each host node.

[0008] Based on the above method, the target prediction model can be used to predict the load of each container instance running on each host node, obtain the predicted load of each container instance, and then obtain the total predicted load of each host node, which is convenient for subsequent judgment on whether each host node needs to schedule a container instance.

[0009] In an optional implementation, N container instances are running on a first host node; determining a first container instance that needs to be scheduled by the first host node according to a total predicted load and a preset load of the first host node includes: calculating a difference between the total predicted load and the preset load of the first host node to obtain a first difference; selecting M container instances from the N container instances, wherein the predicted load of each of the M container instances is greater than or equal to the first difference, N≥M; calculating a difference between the first difference and the predicted load of each of the M container instances to obtain M second differences; and using a container instance corresponding to a minimum second difference among the N second differences as the first container instance.

[0010] Based on the above method, when the total predicted load of the first host node exceeds the preset threshold, the container instance corresponding to the predicted load that is closest to the first difference but larger than the first difference can be determined, and then the first container instance can be determined. Since the predicted load of the first container instance is closest to the first difference, the scheduled first container instance not only satisfies the requirement that the total predicted load of the first host node after scheduling is less than the preset threshold, but also satisfies the requirement that the container instance with the minimum predicted load for scheduling can be determined, thereby improving the utilization rate of each node in the entire container cluster.

[0011] In an optional embodiment, after scheduling the first container instance to the second host node, the method further includes: calculating the sum of the total predicted load of the second host node and the predicted load of the first container instance to obtain the total predicted load of the second host node; comparing the total predicted load of the second host node with the preset load; if the total predicted load is greater than the preset load, determining the second container instance that needs to be scheduled for the second host node based on the total predicted load and the preset load, and the predicted load of the second container instance is less than the predicted load of the first container instance; expelling the second container instance from the second host node, and scheduling the second container instance to a host node other than the second host node among other host nodes until the total predicted load of each scheduled host node is less than or equal to the preset load.

[0012] Based on the above method, when the predicted load of the second host node is always greater than the preset load, the second container instance to be scheduled by the second host node can be determined again, so that the total predicted load of the scheduled second host node is less than the preset threshold, thereby avoiding pressure eviction of the second host node.

[0013] In an optional implementation, P container instances are running on the second host node; determining the second container instance that needs to be scheduled by the second host node according to the total predicted load and the preset load includes: calculating the difference between the total predicted load and the preset load to obtain a third difference; selecting Q container instances from the P container instances, the predicted load of each of the Q container instances being greater than or equal to the third difference, P≥Q; calculating the difference between the third difference and the predicted load of each of the Q container instances to obtain Q fourth differences; and using the container instance corresponding to the smallest fourth difference among the Q fourth differences as the second container instance.

[0014] Based on the above method, when the total predicted load of the second host node exceeds the preset threshold, the container instance corresponding to the predicted load that is closest to the third difference but larger than the third difference can be determined, and then the second container instance can be determined. Since the predicted load of the second container instance is closest to the third difference, the scheduled second container instance not only satisfies the requirement that the total predicted load of the second host node after scheduling is less than the preset threshold, but also satisfies the requirement that the container instance with the minimum predicted load scheduled on the second host node can be determined, thereby improving the utilization rate of each node in the entire container cluster.

[0015] In an optional implementation, the method further includes: if the total predicted load is less than or equal to the preset load, not triggering the second host node to evict the container instance.

[0016] Based on the above method, when the total predicted load is less than or equal to the preset load, the second host node will not be triggered to evict the container instance.

[0017] In an optional embodiment, the method further includes: if after traversing and scheduling at least one host node, there is still a third host node whose total predicted load is greater than the preset load, then the container instance is evicted from the third host node until the total predicted load of the third host node is less than or equal to the preset load.

[0018] Based on the above method, after at least one host node is scheduled, if the total predicted load of the host node is still not less than or equal to the preset load, the container instance in the third host node needs to be expelled to ensure that the load of the third host node is less than the preset load.

[0019] In an optional implementation, a container instance belongs to a service, and a container instance corresponding to a service can be deployed on multiple host nodes; a third host node runs a container instance of at least one service; and the container instance is expelled from the third host node, including: obtaining the number of container instances corresponding to each service in at least one service, the number of task requests, and the weight ratio of each service in at least one service; determining an importance score value for each service based on the number of task requests, the number of container instances, and the weight ratio; and based on the importance score value, expelling the container instance corresponding to the service with the smallest importance score value from the third host node.

[0020] Based on the above method, when evicting container instances in the third host node, the container instances to be evicted from the third host node can be comprehensively determined based on the number of container instances corresponding to each service, the number of task requests, and the weight ratio of each service in at least one service, so as to minimize the impact of the eviction and avoid evicting container instances of important services.

[0021] In an optional implementation, the importance score of each service is determined based on the number of task requests, the number of container instances, and the weight ratio, including: calculating the importance score of each service according to a preset algorithm based on the number of task requests, the number of container instances, and the weight ratio, and the preset algorithm is: D = A / B*C; wherein D represents the importance score, A represents the number of task requests, B represents the number of container instances, and C represents the weight ratio.

[0022] Based on the above method, the importance score value of each service can be calculated according to a preset algorithm, so as to determine the container instance to be expelled from the third host node based on the importance score value.

[0023] In an optional embodiment, before inputting the load into the target prediction model, the method also includes: obtaining a first historical load of each container instance in all container instances running on at least one host node within a first historical period; normalizing the first historical load to obtain a historical load data set; sliding the historical load data set through a time pane to obtain a feature data set; obtaining a pre-trained first model and adjusting the gradient parameters of the first model to obtain a second model; and training the second model according to the feature data set to obtain a target prediction model.

[0024] Based on the above method, the target prediction model can be obtained according to the historical load of each container instance and the pre-trained first model. Since the historical load has a strong correlation with the output predicted load, the prediction result is more accurate, and the gradient parameters of the first model are adjusted to further improve the accuracy of the prediction.

[0025] In an optional embodiment, the first model corresponds to a first gradient parameter; adjusting the gradient parameter of the first model to obtain the second model includes: obtaining a historical feature data set of the first model, the historical feature data set including multiple sample data; calculating a gradient estimate of the second model based on the multiple sample data and the first gradient parameter; determining a cumulative square gradient value of the second model based on the gradient estimate and the attenuation coefficient; calculating a parameter update value of the first gradient parameter based on the cumulative square gradient value and the gradient estimate; calculating the sum of the parameter update value and the first gradient parameter to obtain a second gradient parameter, and configuring the second gradient parameter to the first model to obtain the second model.

[0026] Based on the above method, the gradient parameters of the first model can be adjusted according to the historical feature data set and the first gradient parameters of the first model to obtain the second model, thereby improving the prediction accuracy of the target prediction model.

[0027] In an optional implementation, based on the total predicted load of the other host nodes except the first host node among the at least one host node, it is determined that the host node with the smallest total predicted load is the second host node.

[0028] Based on the above method, since the second host node is the host node with the smallest total predicted load among other host nodes, it means that among other host nodes, the second host node is most likely not to trigger the eviction of the container instance. Therefore, the first container instance needs to be scheduled to the second host node.

[0029] In a second aspect, the present invention provides a scheduling device for a container instance, the device comprising: a processing module, for determining, in at least one host node included in a container cluster, a total predicted load of each host node and a predicted load of each container instance on each host node, wherein at least one container instance can be run on each host node, each container instance corresponds to a predicted load when running, and the total predicted load is the sum of the predicted loads of all container instances running on each host node.

[0030] The processing module is also used to determine whether there is a host node with a total predicted load greater than the preset load by comparing the total predicted load of each host node with the preset load of each host node, and the total predicted load of the first host node is greater than the preset load.

[0031] The processing module is also used to determine, if there is a first host node whose total predicted load is greater than the preset load, a first container instance that needs to be scheduled by the first host node based on the total predicted load and the preset load of the first host node, where the first host node is a host node among at least one host node.

[0032] The scheduling module is also used to evict the first container instance from the first host node and schedule the first container instance to a second host node other than the first host node among at least one host node, where the second host node is a host node in the container cluster that can run the first container instance.

[0033] In a third aspect, the present invention provides a computer device, comprising: a memory and a processor, the memory and the processor being communicatively connected to each other, the memory storing computer instructions, and the processor executing the scheduling method for a container instance of the first aspect or any corresponding embodiment thereof by executing the computer instructions.

[0034] In a fourth aspect, the present invention provides a computer-readable storage medium having computer instructions stored thereon, the computer instructions being used to enable a computer to execute the method for scheduling a container instance of the first aspect or any corresponding embodiment thereof.

[0035] In a fifth aspect, the present invention provides a computer program product, comprising computer instructions, wherein the computer instructions are used to enable a computer to execute the method for scheduling a container instance of the first aspect or any corresponding embodiment thereof. BRIEF DESCRIPTION OF THE DRAWINGS

[0036] In order to more clearly illustrate the specific implementation methods of the present invention or the technical solutions in the prior art, the drawings required for use in the specific implementation methods or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are some implementation methods of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.

[0037] Figure 1 is a topology diagram of a scheduling system for container instances according to an embodiment of the present invention;

[0038] Figure 2 is a flowchart of a method for scheduling a container instance according to an embodiment of the present invention;

[0039] Figure 3 is a flow chart of a training method for a target prediction model according to an embodiment of the present invention;

[0040] Figure 4 is a flowchart of a scheduling method for another container instance according to an embodiment of the present invention;

[0041] Figure 5 is a structural block diagram of a scheduling device for a container instance according to an embodiment of the present invention;

[0042] Figure 6 It is a schematic diagram of the hardware structure of a computer device according to an embodiment of the present invention. DETAILED DESCRIPTION

[0043] In order to make the purpose, technical solution and advantages of the embodiments of the present invention clearer, the technical solution in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative work are within the scope of protection of the present invention.

[0044] The embodiments of the present invention are applied to a container cluster.

[0045] Before the emergence of container technology, virtual machines (VMs) were the main virtualization technology. Virtual machines provide hardware-level isolation, and each virtual machine has its own operating system and resources. However, virtual machines have some limitations, such as slow startup, high resource usage, and difficulty in migration. Container technology overcomes these limitations by providing a lightweight isolation environment at the operating system level. Because containers share the host's operating system kernel, they start quickly, occupy less resources, and are easy to migrate and expand. With the popularity of cloud computing, the demand for elastic, scalable, and distributed systems has increased. Container technology can provide a lightweight, portable execution environment, which is very suitable for deploying and scaling applications in cloud environments.

[0046] As described in the background technology, in order to deal with the problem of node pressure eviction in a container cluster, it is usually necessary to manually monitor the container instances running on each node. This method causes the container instances to be unable to be scheduled in a timely manner, resulting in a high probability of cluster failure; and there is no proactive warning in advance. Since operation and maintenance personnel need to pay attention to the cluster eviction status at all times, the cost of scheduling container instances is high and the efficiency is low.

[0047] In order to solve the above technical problems, an embodiment of the present invention provides a method for scheduling container instances, which schedules container instances by predicting the load of container instances running on host nodes, thereby achieving timely scheduling of container instances, reducing the probability of cluster failures, reducing scheduling costs, and improving scheduling efficiency.

[0048] Below Figure 1 Taking the scheduling system 100 of the container instance shown as an example, the method provided in the embodiment of the present application is described. Figure 1 It is only a schematic diagram and does not constitute a limitation on the applicable scenarios of the technical solution provided in this application.

[0049] like Figure 1 As shown, Figure 1 It is a topological diagram of a scheduling system for container instances according to an embodiment of the present invention. Figure 1In the example, the scheduling system 100 of the container instance may include a scheduling device 101 of the container instance, a first host node 102, a second host node 103 and a third host node 104.

[0050] The scheduling device 101 of the container instance in the embodiment of the present application may be a device with communication and computing functions, such as a server, a cloud server, or a virtual machine, etc. The scheduling device 101 of the container instance may be used to manage the host nodes included in the container cluster, such as the first host node 102, the second host node 103, and the third host node 104.

[0051] Optionally, the scheduling device 101 of the container instance includes a monitoring module, a prediction module, a scheduling and warning module, and an expulsion module. Among them, the monitoring module is used to monitor each host node in the container cluster. The prediction module is used to predict the load of each container instance. The scheduling and warning module is used to schedule the container instance, and when the load of the host node in the schedule cluster is greater than the preset load, a warning message is issued. The expulsion module is used to expel the container instance from the host node.

[0052] The first host node 102, the second host node 103 or the third host node 104 in the embodiment of the present application may be any host node in the container cluster. At least one container instance runs on each host node.

[0053] Among them, each host node is deployed with a NodeExporter component, and the NodeExporter component can obtain the load of each container instance among all container instances running on each host node.

[0054] It can be understood that the scheduling device of the container instance is deployed with a PrometheusServer component. The PrometheusServer component is used to periodically pull the interface of the NodeExporter component and store the acquired load of each container instance.

[0055] Figure 1 The container instance scheduling system 100 shown is only used as an example and is not used to limit the technical solution of the present application. Those skilled in the art should understand that in the specific implementation process, the container instance scheduling system 100 may also include other host nodes, and the number of host nodes may also be determined according to specific needs without limitation.

[0056] According to an embodiment of the present invention, an embodiment of a scheduling method for a container instance is provided. It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer executable instructions, and although a logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in an order different from that shown here.

[0057] In this embodiment, a method for scheduling a container instance is provided, which can be used in the above-mentioned scheduling device for the container instance. Figure 2 is a flowchart of a method for scheduling a container instance according to an embodiment of the present invention. Figure 2 As shown, the process includes the following steps:

[0058] S201: In at least one host node included in a container cluster, determine a total predicted load of each host node and a predicted load of each container instance on each host node.

[0059] The total predicted load is the sum of the predicted loads of all container instances running on each host node.

[0060] Among them, at least one container instance can run on each host node, and each container instance corresponds to a predicted load when running.

[0061] In some optional embodiments, the scheduling device of the container instance obtains the load of each container instance among all container instances running on each host node in the current cycle; inputs the load into the target prediction model and outputs the predicted load; calculates the sum of the predicted loads of each container instance to obtain the total predicted load of each host node.

[0062] The load of each container instance is the resource usage of each container instance in the current cycle. For example, the load of each container instance is the CPU usage, memory usage, disk throughput, or network throughput of each container instance in the current cycle.

[0063] In the embodiment of the present invention, the target prediction model is used to predict the predicted load of each container instance in the next cycle. The target prediction model is obtained by training the scheduling device of the container instance. Figure 3 As shown, Figure 3 is a flow chart of a training method for a target prediction model according to an embodiment of the present invention; specifically, Figure 3 In the example, the container instance scheduling device also performs the following steps:

[0064] S301: Before inputting the load into a target prediction model, obtain a first historical load of each container instance in all container instances running on at least one host node in a first historical period.

[0065] The first historical load may be a plurality of historical loads within a first historical period.

[0066] Example 1, taking the first host node as an example, the first historical load may be a plurality of historical disk throughputs in the first historical period. For example, the plurality of historical disk throughputs include 100M, 10M, 90M, and 30M.

[0067] S302: Normalize the first historical load to obtain a historical load data set.

[0068] In one example, a scheduling device of a container instance normalizes a first historical load based on a first formula, compresses the first historical load to a range of [0, 1] through minimum-maximum normalization, and obtains a historical load data set.

[0069] The first formula is .

[0070] in, is the i-th historical load; is the smallest historical load among multiple historical loads; The largest historical load among multiple historical loads.

[0071] Example 2: Taking the first historical load of 100M, 10M, 90M, and 30M in Example 1 as an example, the scheduling device of the container instance normalizes 100M, 10M, 90M, and 30M based on the first formula, and obtains a historical load data set including 1, 0, 8 / 9, and 2 / 9.

[0072] S303: Sliding processing is performed on the historical load data set through the time window to obtain a characteristic data set.

[0073] In one example, a scheduling device of a container instance normalizes the first historical load based on a second formula to obtain a historical load data set.

[0074] The second formula is .in, is the ith time pane, is the previous time pane, n is the number of data in the pane, i is the index of data in the pane, is the new feature data value in the feature data set.

[0075] S304: Obtain a pre-trained first model, and adjust the gradient parameters of the first model to obtain a second model.

[0076] Among them, the first model is a prediction model based on the Long Short-Term Memory (LSTM) network, and introduces the LSTM-Attention model with the Attention Mechanism.

[0077] The first model corresponds to the first gradient parameter.

[0078] In some optional embodiments, the scheduling device of the container instance obtains a historical feature data set of the first model, the historical feature data set includes multiple sample data; calculates a gradient estimate of the second model based on the multiple sample data and the first gradient parameter; determines a cumulative square gradient value of the second model based on the gradient estimate and the attenuation coefficient; calculates a parameter update value of the first gradient parameter based on the cumulative square gradient value and the gradient estimate; calculates the sum of the parameter update value and the first gradient parameter to obtain a second gradient parameter, and configures the second gradient parameter to the first model to obtain a second model.

[0079] In one example, a scheduling device of a container instance calculates a gradient estimate of a second model according to a plurality of sample data and a first gradient parameter and based on a third formula.

[0080] The third formula is .

[0081] in, It is the i-th sample data among multiple sample data, and can also be the i-th historical load.

[0082] is the objective function. The corresponding mapping is . is the loss function. represents the first gradient parameter.

[0083] In one example, a scheduling device of the container instance determines a cumulative square gradient value of the second model according to the gradient estimate and the attenuation coefficient and based on a fourth formula.

[0084] Among them, the fourth formula is .

[0085] Among them, β is the attenuation coefficient and α is the cumulative square gradient value.

[0086] In one example, a scheduling device of a container instance calculates a parameter update value of a first gradient parameter according to a fifth formula based on the accumulated square gradient value and the gradient estimate value.

[0087] Among them, the fifth formula is .

[0088] Among them, ε represents the global learning rate, δ is a small constant, e is a natural constant, and ∆θ is the parameter update value.

[0089] S305: Perform model training on the second model according to the feature data set to obtain a target prediction model.

[0090] The feature data set includes a training feature data set and a test feature data set. The ratio of the training feature data set to the test feature data set can be set according to actual needs, for example, the ratio can be 7:3.

[0091] In one example, the scheduling device of the container instance inputs the training feature data set into the second model for model training to obtain a third model; and inputs the test feature data set into the third model for model verification to obtain a target prediction model.

[0092] It is understandable that the target prediction model can assign different weights to different features during the prediction process to improve the accuracy of the prediction. At the same time, since the attenuation coefficient β is added when calculating the cumulative square gradient, the speed at which the target prediction model predicts the predicted load of the container instance is accelerated.

[0093] S202: By comparing the total predicted load of each host node with the preset load of each host node, it is determined whether there is a host node whose total predicted load is greater than the preset load.

[0094] A possible design, the preset load can be set according to actual needs without restriction.

[0095] It is understandable that, among at least one host node, a host node whose total predicted load is greater than a preset load needs to schedule a container instance running on the host node.

[0096] S203: If there is a first host node whose total predicted load is greater than the preset load, determine a first container instance that needs to be scheduled on the first host node according to the total predicted load and the preset load of the first host node.

[0097] Among them, N container instances are running on the first host node, and the total predicted load of the first host node is greater than the preset load. The first host node is the host node that needs to schedule the container instance.

[0098] In some optional implementations, a scheduling device for a container instance calculates a difference between a total predicted load of a first host node and a preset load to obtain a first difference; selects M container instances from N container instances; calculates a difference between the first difference and a predicted load of each of the M container instances to obtain M second differences; and uses a container instance corresponding to a minimum second difference among the N second differences as the first container instance.

[0099] The predicted load of each container instance in the M container instances is greater than or equal to the first difference, N≥M.

[0100] S204: Evict the first container instance from the first host node, and schedule the first container instance to a second host node other than the first host node among the at least one host node.

[0101] In one example, a scheduling device of a container instance determines that a host node with a minimum total predicted load is a second host node based on a total predicted load of other host nodes except a first host node among at least one host node.

[0102] It can be understood that the second host node is the host node with the smallest total predicted load among other host nodes, which means that among other host nodes, the second host node is most likely not to trigger the eviction of the container instance. Therefore, the first container instance needs to be scheduled to the second host node.

[0103] In some optional embodiments, after scheduling the first container instance to the second host node, the scheduling device of the container instance calculates the sum of the total predicted load of the second host node and the predicted load of the first container instance to obtain the total predicted load of the second host node; compares the total predicted load of the second host node with the preset load; if the total predicted load is greater than the preset load, determines the second container instance that needs to be scheduled for the second host node based on the total predicted load and the preset load; expels the second container instance from the second host node, and schedules the second container instance to a host node other than the second host node among other host nodes until the total predicted load of each scheduled host node is less than or equal to the preset load.

[0104] The predicted load of the second container instance is less than the predicted load of the first container instance. P container instances are running on the second host node.

[0105] In some optional embodiments, the scheduling device of the container instance calculates the difference between the total predicted load and the preset load to obtain a third difference; selects Q container instances from the P container instances; calculates the difference between the third difference and the predicted load of each of the Q container instances to obtain Q fourth differences; and uses the container instance corresponding to the smallest fourth difference among the Q fourth differences as the second container instance.

[0106] The predicted load of each container instance in the Q container instances is greater than or equal to the third difference, P≥Q.

[0107] Optionally, if the total predicted load is less than or equal to the preset load, the scheduling device of the container instance does not trigger the second host node to evict the container instance.

[0108] In some optional embodiments, if after traversing and scheduling at least one host node, there is still a third host node whose total predicted load is greater than the preset load, the scheduling device of the container instance expels the container instance from the third host node until the total predicted load of the third host node is less than or equal to the preset load.

[0109] In one example, a container instance belongs to a service, and the container instance corresponding to a service can be deployed on multiple host nodes.

[0110] In some optional embodiments, a scheduling device for a container instance obtains the number of container instances corresponding to each service in at least one service, the number of task requests, and the weight ratio of each service in at least one service; determines an importance score value for each service based on the number of task requests, the number of container instances, and the weight ratio; and based on the importance score value, expels the container instance corresponding to the service with the smallest importance score value from a third host node.

[0111] In one example, a scheduling device for a container instance calculates an importance score value for each service according to a preset algorithm based on the number of task requests, the number of container instances, and a weight ratio.

[0112] Among them, the preset algorithm is: D= A / B*C.

[0113] Among them, D represents the importance score, A represents the number of task requests, B represents the number of container instances, and C represents the weight ratio.

[0114] It can be understood that when evicting container instances in the third host node, the container instances to be evicted from the third host node can be comprehensively determined based on the number of container instances corresponding to each service, the number of task requests, and the weight ratio of each service in at least one service, so as to minimize the impact of the eviction and avoid evicting container instances of important services.

[0115] Based on the above methods S201 to S204, the total predicted load of each host node included in the container cluster can be determined, and by comparing the total predicted load with the preset load, the first host node that needs to schedule the container instance and the first container instance that the first host node needs to schedule are determined, and then the first container instance is expelled from the first host node. Since the total predicted load of each host node is determined, and the total predicted load is the total predicted load of each host node at a future moment, the first host node and the first container instance that need to be scheduled can be determined in advance based on the predicted total predicted load, thereby avoiding the problem of not being able to schedule the container instance in time when the container cluster node is expelled due to pressure, thereby reducing the probability of cluster failure, improving the efficiency of container instance scheduling, reducing the cost of container instance scheduling, and improving the availability of the system.

[0116] In this embodiment, a method for scheduling a container instance is also provided. Figure 4 As shown, Figure 4 is the embodiment of the present invention; Figure 4 In the container instance scheduler, the container instance also performs:

[0117] S401: Determine a total predicted load of each host node in at least one host node included in a container cluster;

[0118] S402: Determine whether the total predicted load of each host node is greater than the preset load of each host node;

[0119] S403: If not, the host node is not triggered to evict the container instance;

[0120] S404: If yes, determine the first host node that needs to schedule the container instance, and determine the first container instance that needs to be scheduled by the first host node according to the total predicted load and the preset load of the first host node;

[0121] S405: Determine whether there is a target host node with an idle load greater than or equal to the first container instance;

[0122] S406: If yes, schedule the first container instance to the target host node, and the sum of the total predicted load of the scheduled target host node and the predicted load of the first container instance is less than or equal to the preset load;

[0123] S407: If not, determine and evict the container instance to be scheduled based on the number of container instances corresponding to each service on at least one host node, the number of task requests, and the weight ratio.

[0124] In this embodiment, a scheduling device for a container instance is also provided, which is used to implement the above-mentioned embodiments and preferred implementation modes, and will not be repeated hereafter. As used below, the term "module" may be a combination of software and / or hardware that implements a predetermined function. Although the devices described in the following embodiments are preferably implemented in software, the implementation of hardware, or a combination of software and hardware, is also possible and conceivable.

[0125] This embodiment provides a scheduling device for a container instance, such as Figure 5 As shown, Figure 5 1 is a structural block diagram of a scheduling device for a container instance according to an embodiment of the present invention; the device includes:

[0126] The processing module 501 is used to determine the total predicted load of each host node and the predicted load of each container instance on each host node in at least one host node included in the container cluster, wherein at least one container instance can be run on each host node, each container instance corresponds to a predicted load when it is running, and the total predicted load is the sum of the predicted loads of all container instances running on each host node.

[0127] The processing module 501 is further used to determine whether there is a host node whose total predicted load is greater than the preset load by comparing the total predicted load of each host node with the preset load of each host node.

[0128] The processing module 501 is also used to determine, if there is a first host node whose total predicted load is greater than the preset load, a first container instance that needs to be scheduled by the first host node based on the total predicted load and the preset load of the first host node, where the first host node is a host node among at least one host node.

[0129] The scheduling module 502 is further used to evict the first container instance from the first host node, and schedule the first container instance to a second host node other than the first host node among at least one host node, where the second host node is a host node in the container cluster that can run the first container instance.

[0130] In some optional implementations, processing module 501 is specifically used to obtain the load of each container instance among all container instances running on each host node in the current cycle; processing module 501 is also specifically used to input the load into a target prediction model and output the predicted load, and the target prediction model is used to predict the predicted load of each container instance in the next cycle; processing module 501 is also specifically used to calculate the sum of the predicted loads of each container instance to obtain the total predicted load of each host node.

[0131] In some optional implementations, N container instances are running on the first host node; the processing module 501 is further specifically used to calculate the difference between the total predicted load of the first host node and the preset load to obtain a first difference; the processing module 501 is further specifically used to select M container instances from the N container instances, and the predicted load of each container instance in the M container instances is greater than or equal to the first difference, N≥M; the processing module 501 is further specifically used to calculate the difference between the first difference and the predicted load of each container instance in the M container instances to obtain M second differences; the processing module 501 is further specifically used to use the container instance corresponding to the smallest second difference among the N second differences as the first container instance.

[0132] In some optional embodiments, after scheduling the first container instance to the second host node, the processing module 501 is further used to calculate the sum of the total predicted load of the second host node and the predicted load of the first container instance to obtain the total predicted load of the second host node; the processing module 501 is further used to compare the total predicted load of the second host node with the preset load; the processing module 501 is further used to determine the second container instance that needs to be scheduled for the second host node based on the total predicted load and the preset load if the total predicted load is greater than the preset load, and the predicted load of the second container instance is less than the predicted load of the first container instance; the scheduling module 502 is further used to evict the second container instance from the second host node, and schedule the second container instance to a host node other than the second host node among other host nodes, until the total predicted load of each scheduled host node is less than or equal to the preset load.

[0133] In some optional implementations, P container instances are running on the second host node; the processing module 501 is further specifically used to calculate the difference between the total predicted load and the preset load to obtain a third difference; the processing module 501 is further specifically used to select Q container instances from the P container instances, and the predicted load of each of the Q container instances is greater than or equal to the third difference, P≥Q; the processing module 501 is further specifically used to calculate the difference between the third difference and the predicted load of each of the Q container instances to obtain Q fourth differences; the processing module 501 is further specifically used to use the container instance corresponding to the smallest fourth difference among the Q fourth differences as the second container instance.

[0134] In some optional implementations, the scheduling module 502 is further configured to not trigger the second host node to evict the container instance if the total predicted load is less than or equal to a preset load.

[0135] In some optional embodiments, the scheduling module 502 is further used to evict container instances from the third host node if, after traversing and scheduling at least one host node, there is still a third host node whose total predicted load is greater than the preset load, until the total predicted load of the third host node is less than or equal to the preset load.

[0136] In some optional implementations, a container instance belongs to a service, and a container instance corresponding to a service can be deployed on multiple host nodes; the third host node runs a container instance of at least one service; the scheduling module 502 is specifically used to obtain the number of container instances corresponding to each service in at least one service, the number of task requests and the weight ratio of each service in at least one service; the scheduling module 502 is also specifically used to determine the importance score value of each service based on the number of task requests, the number of container instances and the weight ratio; the scheduling module 502 is also specifically used to evict the container instance corresponding to the service with the smallest importance score from the third host node based on the importance score value.

[0137] In some optional implementations, the scheduling module 502 is also specifically used to calculate the importance score of each service according to the number of task requests, the number of container instances and the weight ratio according to a preset algorithm, and the preset algorithm is: D = A / B*C; wherein D represents the importance score, A represents the number of task requests, B represents the number of container instances, and C represents the weight ratio.

[0138] In some optional embodiments, before inputting the load into the target prediction model, the processing module 501 is also used to obtain the first historical load of each container instance in all container instances running on at least one host node within a first historical period; the processing module 501 is also used to normalize the first historical load to obtain a historical load data set; the processing module 501 is also used to slide the historical load data set through a time pane to obtain a feature data set; the processing module 501 is also used to obtain a pre-trained first model and adjust the gradient parameters of the first model to obtain a second model; the processing module 501 is also used to perform model training on the second model according to the feature data set to obtain a target prediction model.

[0139] In some optional embodiments, the first model corresponds to the first gradient parameter; the processing module 501 is also specifically used to obtain a historical feature data set of the first model, and the historical feature data set includes multiple sample data; the processing module 501 is also specifically used to calculate the gradient estimate of the second model based on the multiple sample data and the first gradient parameter; the processing module 501 is also specifically used to determine the cumulative square gradient value of the second model based on the gradient estimate and the attenuation coefficient; the processing module 501 is also specifically used to calculate the parameter update value of the first gradient parameter based on the cumulative square gradient value and the gradient estimate; the processing module 501 is also specifically used to calculate the sum of the parameter update value and the first gradient parameter to obtain the second gradient parameter, and configure the second gradient parameter to the first model to obtain the second model.

[0140] In some optional implementations, the processing module 501 is further configured to determine, based on the total predicted loads of the other host nodes except the first host node among the at least one host node, that the host node with the smallest total predicted load is the second host node.

[0141] The further functional description of each of the above modules and units is the same as that of the above corresponding embodiments and will not be repeated here.

[0142] The scheduling device of the container instance in this embodiment is presented in the form of a functional unit, where the unit refers to an ASIC (Application Specific Integrated Circuit) circuit, a processor and memory that executes one or more software or fixed programs, and / or other devices that can provide the above functions.

[0143] The embodiment of the present invention also provides a computer device having the above Figure 5 The scheduling device of the container instance is shown.

[0144] See also Figure 6 , Figure 6 is a schematic diagram of the structure of a computer device provided by an optional embodiment of the present invention, such as Figure 6 As shown, the computer device includes: one or more processors 10, a memory 20, and interfaces for connecting various components, including high-speed interfaces and low-speed interfaces. Various components are connected to each other using different buses for communication, and can be installed on a common mainboard or installed in other ways as needed. The processor can process the instructions executed in the computer device, including instructions stored in or on the memory to display the graphical information of the GUI on an external input / output device (such as, a display device coupled to the interface). In some optional embodiments, if necessary, multiple processors and / or multiple buses can be used together with multiple memories and multiple memories. Similarly, multiple computer devices can be connected, and each device provides some necessary operations (for example, as a server array, a group of blade servers, or a multi-processor system). Figure 6 A processor 10 is taken as an example.

[0145] The processor 10 may be a central processing unit, a network processor or a combination thereof. The processor 10 may further include a hardware chip. The hardware chip may be a dedicated integrated circuit, a programmable logic device or a combination thereof. The programmable logic device may be a complex programmable logic device, a field programmable gate array, a general purpose array logic or any combination thereof.

[0146] The memory 20 stores instructions executable by at least one processor 10, so that the at least one processor 10 executes the method shown in the above embodiment.

[0147] The memory 20 may include a program storage area and a data storage area, wherein the program storage area may store an operating system, an application required for at least one function; the data storage area may store data created according to the use of the computer device, etc. In addition, the memory 20 may include a high-speed random access memory, and may also include a non-transient memory, such as at least one disk storage device, a flash memory device, or other non-transient solid-state storage device. In some optional embodiments, the memory 20 may optionally include a memory remotely arranged relative to the processor 10, and these remote memories may be connected to the computer device via a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.

[0148] The memory 20 may include a volatile memory, such as a random access memory; the memory may also include a non-volatile memory, such as a flash memory, a hard disk or a solid state drive; the memory 20 may also include a combination of the above types of memory.

[0149] The computer device further comprises a communication interface 30 for the computer device to communicate with other devices or a communication network.

[0150] The embodiment of the present invention also provides a computer-readable storage medium. The method according to the embodiment of the present invention can be implemented in hardware, firmware, or can be implemented as a computer code that can be recorded in a storage medium, or can be implemented as a computer code that is originally stored in a remote storage medium or a non-temporary machine-readable storage medium and will be stored in a local storage medium through a network download, so that the method described herein can be stored in such software processing on a storage medium using a general-purpose computer, a dedicated processor, or programmable or dedicated hardware. Among them, the storage medium can be a magnetic disk, an optical disk, a read-only storage memory, a random access memory, a flash memory, a hard disk or a solid-state hard disk, etc.; further, the storage medium can also include a combination of the above types of memories. It can be understood that a computer, a processor, a microprocessor controller, or programmable hardware includes a storage component that can store or receive software or computer code. When the software or computer code is accessed and executed by a computer, a processor, or hardware, the method shown in the above embodiment is implemented.

[0151] Part of the present application may be applied as a computer program product, such as a computer program instruction, which, when executed by a computer, can call or provide the method and / or technical solution according to the present application through the operation of the computer. Those skilled in the art should understand that the existence of computer program instructions in computer-readable media includes but is not limited to source files, executable files, installation package files, etc., and accordingly, the way in which computer program instructions are executed by a computer includes but is not limited to: the computer directly executes the instruction, or the computer compiles the instruction and then executes the corresponding compiled program, or the computer reads and executes the instruction, or the computer reads and installs the instruction and then executes the corresponding installed program. Here, the computer-readable medium can be any available computer-readable storage medium or communication medium accessible to the computer.

[0152] Although the embodiments of the present invention have been described in conjunction with the accompanying drawings, those skilled in the art may make various modifications and variations without departing from the spirit and scope of the present invention, and such modifications and variations are all within the scope defined by the appended claims.

Claims

1. A method for scheduling a container instance, characterized in that: The method comprises: In at least one host node included in the container cluster, determine the total predicted load of each host node and the predicted load of each container instance on each host node, wherein each host node can run at least one container instance, each container instance corresponds to a predicted load when running, and the total predicted load is the sum of the predicted loads of all container instances running on each host node; By comparing the total predicted load and the preset load of each of the host nodes, determining whether there is a first host node whose total predicted load is greater than the preset load, the first host node being a host node among the at least one host node; If so, determining a first container instance that needs to be scheduled by the first host node according to the total predicted load of the first host node and the preset load; Evicting the first container instance from the first host node, and scheduling the first container instance to a second host node other than the first host node among the at least one host node, where the second host node is a host node in the container cluster that can run the first container instance; If after traversing and scheduling the at least one host node, there is still a third host node whose total predicted load is greater than the preset load, the container instance is expelled from the third host node until the total predicted load of the third host node is less than or equal to the preset load, a container instance belongs to a service, and a container instance corresponding to a service can be deployed on multiple host nodes, and the third host node runs a container instance of at least one service; Wherein, expelling the container instance from the third host node includes: Obtain the number of container instances, the number of task requests, and the weight ratio of each service in the at least one service corresponding to each service in the at least one service; Determine the importance score of each of the services according to the number of task requests, the number of container instances, and the weight ratio; Based on the importance score value, the container instance corresponding to the service with the smallest importance score value is expelled from the third host node.

2. The method according to claim 1, characterized in that Determining, in at least one host node included in the container cluster, a total predicted load of each host node and a predicted load of each container instance on each host node, comprises: Obtain the load of each container instance among all container instances running on each of the host nodes in the current cycle; Inputting the load into a target prediction model and outputting the predicted load, wherein the target prediction model is used to predict the predicted load of each container instance in the next cycle; The sum of the predicted loads of each of the container instances is calculated to obtain the total predicted load of each of the host nodes.

3. The method according to claim 2, characterized in that N container instances are running on the first host node; and determining, according to the total predicted load of the first host node and the preset load, the first container instance that needs to be scheduled by the first host node includes: Calculating a difference between a total predicted load of the first host node and the preset load to obtain a first difference; Selecting M container instances from the N container instances, wherein the predicted load of each container instance in the M container instances is greater than or equal to the first difference, N≥M; Calculate the difference between the first difference and the predicted load of each container instance in the M container instances to obtain M second difference values; The container instance corresponding to the minimum second difference value among the M second difference values ​​is used as the first container instance.

4. The method according to any one of claims 1 to 3, characterized in that: After scheduling the first container instance to the second host node, the method further includes: Calculate the sum of the total predicted load of the second host node and the predicted load of the first container instance to obtain the sum of the predicted loads of the second host node; comparing the predicted total load of the second host node with the preset load; If the total predicted load is greater than the preset load, determining a second container instance that needs to be scheduled by the second host node according to the total predicted load and the preset load, the predicted load of the second container instance being less than the predicted load of the first container instance; The second container instance is evicted from the second host node, and the second container instance is scheduled to a host node other than the second host node among other host nodes until the total predicted load of each of the scheduled host nodes is less than or equal to the preset load.

5. The method according to claim 1, characterized in that P container instances are running on the second host node; and determining the second container instance that needs to be scheduled by the second host node according to the predicted load sum and the preset load includes: Calculating a difference between the predicted load sum and the preset load to obtain a third difference; Select Q container instances from the P container instances, where the predicted load of each container instance in the Q container instances is greater than or equal to the third difference, P≥Q; Calculate the difference between the third difference value and the predicted load of each container instance in the Q container instances to obtain Q fourth difference values; The container instance corresponding to the smallest fourth difference value among the Q fourth differences is used as the second container instance.

6. The method according to claim 5, characterized in that The method further comprises: If the total predicted load is less than or equal to the preset load, the second host node is not triggered to evict the container instance.

7. The method according to claim 1, characterized in that Determining the importance score of each of the services according to the number of task requests, the number of container instances, and the weight ratio includes: According to the number of task requests, the number of container instances and the weight ratio, the importance score of each service is calculated according to a preset algorithm, and the preset algorithm is: D = A / B×C; Wherein, D represents the importance score, A represents the number of task requests, B represents the number of container instances, and C represents the weight ratio.

8. The method according to claim 2, characterized in that: Before inputting the load into the target prediction model, the method further includes: Obtaining a first historical load of each container instance in all container instances running on the at least one host node in a first historical period; Normalizing the first historical load to obtain a historical load data set; Sliding the historical load data set through a time pane to obtain a feature data set; Obtain a pre-trained first model, and adjust the gradient parameters of the first model to obtain a second model; The second model is trained according to the feature data set to obtain the target prediction model.

9. The method according to claim 8, characterized in that The first model corresponds to a first gradient parameter; and adjusting the gradient parameter of the first model to obtain the second model includes: Acquire a historical feature data set of the first model, wherein the historical feature data set includes a plurality of sample data; Calculating a gradient estimate of the second model according to the plurality of sample data and the first gradient parameter; Determining a cumulative square gradient value of the second model according to the gradient estimate and the attenuation coefficient; Calculating a parameter update value of the first gradient parameter according to the accumulated square gradient value and the gradient estimate value; The sum of the parameter update value and the first gradient parameter is calculated to obtain a second gradient parameter, and the second gradient parameter is configured to the first model to obtain the second model.

10. A container instance scheduling device, characterized in that: The device comprises: A processing module is used to determine, in at least one host node included in the container cluster, a total predicted load of each host node and a predicted load of each container instance on each host node, wherein each host node can run at least one container instance, each container instance corresponds to a predicted load when running, and the total predicted load is the sum of the predicted loads of all container instances running on each host node; The processing module is further used to determine whether there is a first host node whose total predicted load is greater than the preset load by comparing the total predicted load of each of the host nodes with the preset load of each of the host nodes, the first host node being a host node among the at least one host node; The processing module is further configured to determine, if any, a first container instance that needs to be scheduled by the first host node based on the total predicted load of the first host node and the preset load; The scheduling module is further used to evict the first container instance from the first host node, and schedule the first container instance to a second host node other than the first host node among the at least one host node, where the second host node is a host node in the container cluster that can run the first container instance; The processing module is further configured to, if after traversing and scheduling the at least one host node, there is still a third host node whose total predicted load is greater than the preset load, evict the container instance from the third host node until the total predicted load of the third host node is less than or equal to the preset load, a container instance belongs to a service, a container instance corresponding to a service can be deployed on multiple host nodes, and the third host node runs a container instance of at least one service; Wherein, expelling the container instance from the third host node includes: The processing module is specifically used to obtain the number of container instances corresponding to each service in the at least one service, the number of task requests, and the weight ratio of each service in the at least one service; The processing module is further specifically used to determine the importance score value of each of the services according to the number of task requests, the number of container instances and the weight ratio; The processing module is further specifically configured to evict the container instance corresponding to the service with the smallest importance score from the third host node based on the importance score.

11. A computer device, characterized in that: include: A memory and a processor, wherein the memory and the processor are communicatively connected to each other, the memory stores computer instructions, and the processor executes the method for scheduling a container instance according to any one of claims 1 to 9 by executing the computer instructions.

12. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores computer instructions, and the computer instructions are used to enable a computer to execute the method for scheduling container instances according to any one of claims 1 to 9.

13. A computer program product, characterized in that The method comprises computer instructions, wherein the computer instructions are used to cause a computer to execute the method for scheduling container instances according to any one of claims 1 to 9.

Citation Information

Patent Citations

  • Intelligent resource optimization method of container cloud platform based on load prediction

    CN108829494A