Method and apparatus for adjusting the number of containers
By obtaining the resource load data of the container in Kubernetes, and using the combination of time series and Kalman filtering model for prediction, the accuracy of container number adjustment is solved, and the timely adjustment of container number is achieved, ensuring the service quality of the application is ensured.
Patent Information
- Application Number
- CN202211175211.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-09-26
- Publication Date
- 2025-07-08
- Estimated Expiration
- 2042-09-26
AI Technical Summary
Kubernetes' horizontal scaling algorithm has a single measurement indicator and response delay problem, and it is impossible to accurately determine the resource load rate in the application, and the number of containers cannot be adjusted in time when the load changes, resulting in the inability to guarantee the quality of application service.
By obtaining the resource load data of the container at multiple moments, calculating the comprehensive load rate, and using a combination of the time series prediction model and the Kalman filtering model to make predictions, the number of containers is adjusted to achieve elastic scaling.
Improve the accuracy of container load rate, ensure that the application can perform elastic scaling operations in a timely and accurate manner, and ensure service quality.
Smart Images

Figure CN115525394B_ABST
Abstract
Description
Technical Field
[0001] The embodiments of the present application relate to the field of computers, and more specifically, to a method and device for adjusting the number of containers. Background Art
[0002] With the rapid development of virtualization technology and cloud computing technology, cloud computing technology based on traditional hypervisor virtualization has gradually been replaced by container technology based on container virtualization represented by Docker due to a series of problems such as low resource utilization.
[0003] Compared with traditional virtualization technology, Docker container technology increases startup speed and reduces overhead by reusing the local operating system. At the same time, because Docker container technology simplifies application deployment, it is welcomed by developers. However, when faced with large-scale cluster container groups, it is very difficult to manage these containers. Kubernetes is Google's open source version of Borg, a Docker container cluster orchestration and scheduling system that uses the Master-Slave model to provide containerized applications with resource scheduling, automated deployment, service discovery, elastic scaling, resource monitoring and other services. Among them, elastic scaling is to monitor the evaluation indicators specified by the user and use thresholds to horizontally expand and shrink the application to ensure the service quality of the application and maximize resource savings.
[0004] The current built-in elastic scaling strategy of Kubernetes is to achieve automatic scaling of Pods through HPA (Horizontal Pod Autoscaler). When deploying K8s applications, a resource monitoring indicator and the target average usage threshold TAU of the resource will be set. The scaling threshold calculation formula of this indicator is
[0005]
[0006] Tolarence is the default value of 0.1 in the system. This parameter is set to prevent the application from frequently expanding and shrinking. Up and down are the upper and lower limits of the scaling. Assuming that there are currently k pods of this application, HPA obtains the resource usage Ui of all the pods in the current set through polling, and obtains the current resource utilization and CAU as:
[0007]
[0008] Where request represents the amount of resources allocated to the pod. If k*down≤CAU≤k*up, then no scaling is required. Otherwise, scaling is required. The calculation formula is:
[0009] TPN = ceil(CAU / TAU) (3)
[0010] Where TPN represents the number of target pods, and ceil means rounding up the calculation result. To limit the number of pods, there are the following constraints:
[0011]
[0012] Where, R min is the minimum value of the number of pods, and R max is the maximum value of the number of pods.
[0013] The above is an overview of the process of Kubernetes implementing the scaling function. It can be seen from the above analysis that although the built-in horizontal scaling algorithm of Kubernetes is relatively simple, there are two obvious problems: a single measurement metric and response latency. When facing a complex application system, the consumption of multiple resource types involved in the application may change over time and with the business, and a single measurement metric cannot accurately measure the overall load of the application. When the application faces sudden load changes, the quality of application services before pod expansion cannot be guaranteed, and it is even possible that the application may crash due to too high load. Summary of the Invention
[0014] The embodiments of the present application provide a method and device for adjusting the number of containers, at least to solve the problems in the related art that the resource load rate in the application cannot be accurately determined, and the number of containers in the application cannot be changed in time after the resource load rate is determined.
[0015] An embodiment of the present application provides a method for adjusting the number of containers, including:
[0016] In an exemplary embodiment, obtain resource load data sets of any one container in a target application at multiple moments, obtain multiple resource load data sets, and calculate the resource load rate of the container for each resource according to the resource usage amount and resource occupancy amount in the resource load data set respectively, to obtain multiple sets of resource load rates. Wherein, the target application includes at least one container, and the resource load data sets of each container are the same; calculate the comprehensive load rate according to the multiple resource load rates in each set of resource load rates, to obtain the comprehensive load rate of the container at each moment; generate a load rate sequence of the container according to the comprehensive load rates of the container at each moment; input the load rate sequence into a preset prediction model, to obtain a predicted value of the comprehensive load rate of the container at the target moment, and adjust the number of containers in the target application according to the predicted value. Wherein, the preset prediction model includes a time series prediction model and a Kalman filter model, and the target moment is the maximum moment different from the multiple moments.
[0017] Optionally, calculating the comprehensive load rate based on multiple resource load rates in each group includes: when there is a resource load rate greater than or equal to the first resource load rate among the multiple resource load rates, determining the resource load rate greater than or equal to the first resource load rate as the comprehensive load rate; or when all the multiple resource load rates are less than or equal to the second resource load rate, determining the maximum load rate among the multiple resource load rates as the comprehensive load rate, where the first resource load rate is greater than the second resource load rate; or when each resource load rate is greater than the second resource load rate and less than the first resource load rate, performing a weighted sum on the multiple resource load rates to obtain the comprehensive load rate.
[0018] Optionally, performing a weighted sum on the multiple resource load rates to obtain the comprehensive load rate includes: adding the multiple resource load rates to obtain the sum of the multiple resource load rates; successively dividing each resource load rate by the sum of the multiple resource load rates to obtain the weight value of each resource load rate; and performing a weighted sum calculation through each resource load rate and the corresponding weight value to obtain the comprehensive load rate of each group of resource load rates.
[0019] Optionally, generating a load rate sequence of a container according to the comprehensive load rate of the container at each moment includes: sorting the comprehensive load rates in ascending order of time to obtain a candidate load rate sequence; determining whether the candidate load rate sequence is stationary data, where stationary data is data that fluctuates around the mean; when the candidate load rate sequence is stationary data, determining the candidate load rate sequence as the load rate sequence; and when the candidate load rate sequence is not stationary data, performing a stationary processing on the candidate load rate sequence to obtain the load rate sequence.
[0020] Optionally, adjusting the number of containers in a target application according to a predicted value includes: determining whether the predicted value is greater than a first predicted value and determining whether the predicted value is less than a second predicted value, where the first predicted value is greater than the second predicted value; when the predicted value is greater than the first predicted value, adding a preset number of containers to the target application; when the predicted value is less than the second predicted value, removing a preset number of containers from the target application; and when the predicted value is less than or equal to the first predicted value and greater than or equal to the second predicted value, keeping the number of containers in the target application unchanged.
[0021] Optionally, inputting the load rate sequence into a preset prediction model to obtain a predicted value of the comprehensive load rate of the container at a target moment includes: inputting the load rate sequence into a time series prediction model to obtain a candidate predicted value; inputting the load rate sequence and the candidate predicted value into a Kalman filter model, and correcting the candidate predicted value through the Kalman filter model and the load rate sequence to obtain the predicted value of the comprehensive load rate.
[0022] Optionally, calculate the resource load rate of each resource for the container based on the resource usage and resource occupancy in the resource load data set, and obtain multiple sets of resource load rates, including: determining the resource usage and resource occupancy of each resource occupied by the container; respectively dividing the resource usage of each resource by the resource occupancy to obtain multiple resource load rates; and grouping the multiple resource load rates to obtain a set of resource load rates of the container at the same moment.
[0023] According to another embodiment of the present application, there is provided an adjustment device for the number of containers, including: an acquisition module, configured to acquire resource load data sets of any one container in a target application at multiple moments, obtain multiple resource load data sets, and calculate the resource load rate of each resource for the container based on the resource usage and resource occupancy in the resource load data set, to obtain multiple sets of resource load rates, where the target application includes at least one container, and the resource load data sets of each container are the same; a calculation module, configured to calculate a comprehensive load rate based on the multiple resource load rates in each set of resource load rates to obtain the comprehensive load rate of the container at each moment; a generation module, configured to generate a load rate sequence of the container according to the comprehensive load rates of the container at each moment; and a prediction module, configured to input the load rate sequence into a preset prediction model to obtain a predicted value of the comprehensive load rate of the container at a target moment, and adjust the number of containers in the target application according to the predicted value, where the preset prediction model includes a time series prediction model and a Kalman filter model, and the target moment is the maximum moment different from the multiple moments.
[0024] According to still another embodiment of the present application, there is further provided a computer-readable storage medium, in which a computer program is stored, where the computer program is configured to execute the steps in any one of the above method embodiments when running.
[0025] According to still another embodiment of the present application, there is further provided an electronic device, including a memory and a processor, where a computer program is stored in the memory, and the processor is configured to run the computer program to execute the steps in any one of the above method embodiments.
[0026] Through this application, by obtaining the occupation amount and usage amount of each resource occupied by a container, the resource load rate of the container for each resource is determined. Then, according to the different importance levels among resources, a weighted sum calculation is performed on the resource load rates to obtain the accurate comprehensive load rate of the container. The comprehensive load rates at multiple historical moments are used as training data to train a combined model composed of a time series prediction model and a Kalman filter model, so as to achieve the effect of accurately predicting the comprehensive load rate of the container according to the training data. Furthermore, the number of containers in the application can be changed in advance according to the predicted value, ensuring that the application can perform elastic scaling operations in a timely and accurate manner. Therefore, it can effectively solve the problems in the related art that the resource load rate in the application cannot be accurately determined and the number of containers in the application cannot be changed in a timely manner after the resource load rate is determined. Description of the Drawings
[0027] Figure 1 is a hardware structure block diagram of a mobile terminal for a method of adjusting the number of containers according to an embodiment of the present application;
[0028] Figure 2 is a flowchart of a method for adjusting the number of containers according to an embodiment of the present application;
[0029] Figure 3 is a schematic diagram of an optional candidate load rate sequence according to an embodiment of the present application Figure 1 ;
[0030] Figure 4 is a schematic diagram of an optional candidate load rate sequence according to an embodiment of the present application Figure 2 ;
[0031] Figure 5 is a schematic diagram of an optional predicted value error degree according to an embodiment of the present application;
[0032] Figure 6 is a schematic diagram of a device for adjusting the number of containers according to an embodiment of the present application. Detailed Embodiments
[0033] In the following, embodiments of the present application will be described in detail with reference to the drawings and in conjunction with the embodiments.
[0034] It should be noted that the terms "first", "second", etc. in the description and claims of the present application and the above drawings are used to distinguish similar objects, and do not necessarily need to describe a specific order or sequence.
[0035] The method embodiments provided in the embodiments of the present application can be executed on a mobile terminal (electronic device), a computer terminal, or a similar computing device. Taking running on a mobile terminal as an example, Figure 1 is a hardware structure block diagram of a mobile terminal for a method of adjusting the number of containers according to an embodiment of the present application. AsFigure 1 As shown, the mobile terminal may include one or more ( Figure 1 only one is shown in the figure) processors 102 (the processor 102 may include, but is not limited to, processing devices such as a microprocessor MCU or a programmable logic device FPGA) and a memory 104 for storing data. Among them, the above mobile terminal may further include a transmission device 106 for communication functions and an input / output device 108. Those of ordinary skill in the art can understand that Figure 1 the structure shown in the figure is only schematic and does not limit the structure of the above mobile terminal. For example, the mobile terminal may further include more or fewer components than Figure 1 shown in the figure, or have a different configuration from Figure 1 shown in the figure.
[0036] The memory 104 can be used to store computer programs. For example, software programs and modules of application software, such as the computer program corresponding to the method for adjusting the number of containers in the embodiments of the present application. The processor 102 executes various functional applications and data processing by running the computer programs stored in the memory 104, that is, implements the above method. The memory 104 may include high-speed random access memory and may also include non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memories. In some instances, the memory 104 may further include a memory remotely set relative to the processor 102, and these remote memories can be connected to the mobile terminal through a network. Examples of the above network include, but are not limited to, the Internet, an enterprise intranet, a local area network, a mobile communication network, and combinations thereof.
[0037] The transmission device 106 is used to receive or send data via a network. Specific examples of the above network may include a wireless network provided by the communication provider of the mobile terminal. In one instance, the transmission device 106 includes a network adapter (abbreviated as NIC), which can be connected to other network devices through a base station and thus can communicate with the Internet. In one instance, the transmission device 106 may be a radio frequency (abbreviated as RF) module, which is used to communicate with the Internet wirelessly.
[0038] For the convenience of description, the following explains some nouns or terms related to the embodiments of the present application:
[0039] Kubernetes: An open-source system for automatically deploying, scaling, and managing containerized applications, abbreviated as k8s.
[0040] A Pod is a collection of containers (such as Docker containers) consisting of one or more containers, and has the ability to share storage / network / UTS / PID, as well as the specification for running containers. In Kubernetes, a Pod is the smallest schedulable atomic unit.
[0041] CLR: Comprehensive Load Rate, the comprehensive load rate.
[0042] ARIMA: Autoregressive Integrated Moving Average model, the differential autoregressive stationary moving average model.
[0043] HPA: Horizontal Pod Autoscaler, horizontal Pod auto-scaling.
[0044] In this embodiment, a method running on the above-mentioned mobile terminal (electronic device) is provided. Figure 2 It is a flowchart of the method for adjusting the number of containers according to the embodiments of the present application. As Figure 2 shown, the process includes the following steps:
[0045] Step S202, obtain resource load data sets of any one container in the target application at multiple moments, obtain multiple resource load data sets, and calculate the resource load rate of each resource of the container according to the resource usage amount and resource occupancy amount in the resource load data set respectively, to obtain multiple groups of resource load rates. Among them, the target application contains at least one container, and the resource load data sets of each container are the same.
[0046] Specifically, the target application can be a containerized application program managed by the k8s cluster. The container can be the smallest schedulable unit pod in k8s. Each pod can contain one or more containers. Therefore, each pod can also be a container group.
[0047] When there are multiple pods in the application, since the resource occupancy situations of each pod are the same, when determining the resource occupancy rate, the resource occupancy rate of any one pod in the application can be calculated only, and the resource occupancy situation of the entire application can be reflected according to the obtained resource occupancy rate result, so that the number of pods in the application can be elastically scaled and adjusted according to the resource occupancy situation.
[0048] Therefore, resource load data sets of any one container in the target application at multiple moments can be obtained, and multiple resource load data sets can be obtained. Among them, each resource load set includes the usage situation of multiple resources in the application by the pod at one moment and the allocation situation of each resource.
[0049] For example, at time A, the resources occupied by Pod 1 are memory. There are 2 pods in this application, and the memory allocation for this application is 2G. Therefore, the memory allocation for each pod is 1G. Here, the allocation amount is also the occupancy amount of the pods in this application. Thus, when calculating, the resources allocated to the pods are named resource occupancy amounts.
[0050] After obtaining the resource occupancy amounts from the resource load dataset, it is also necessary to obtain the resource usage amounts of the same pod at the same time, that is, the usage amounts of the resources that have been used. After determining the resource usage amounts and resource occupancy amounts in the resource load dataset of this pod, the resource load rates of this pod for each resource can be calculated according to the resource usage amounts and resource occupancy amounts, that is, the resource load rates of a certain pod for each resource at a certain moment, and the resource load rates of each resource are combined into a set of resource load rate information of this pod at this moment.
[0051] Step S204: Calculate the comprehensive load rate according to the multiple resource load rates in each set of resource load rates, and obtain the comprehensive load rate of the container at each moment.
[0052] Specifically, after obtaining a set of resource load rates of a certain pod at a certain moment, the comprehensive load rate of this container at this moment can be determined according to the multiple resource load rates in this set of resource load rates. Among them, the comprehensive load rate is the comprehensive usage situation obtained after comprehensively considering the resource usage situations of multiple resources. Therefore, it is necessary to calculate the comprehensive load rate according to the importance levels among the resources, so as to more accurately measure the overall load level in the application.
[0053] Step S206: Generate a load rate sequence of the container according to the comprehensive load rates of the container at each moment.
[0054] Specifically, after obtaining the comprehensive load rate of this pod at each moment, a time sequence of the comprehensive load rate of this pod can be generated in ascending order of time, and this time sequence is determined as the load rate sequence of the pod.
[0055] Step S208: Input the load rate sequence into a preset prediction model to obtain a predicted value of the comprehensive load rate of the container at the target moment, and adjust the number of containers in the target application according to the predicted value. The preset prediction model includes a time series prediction model and a Kalman filter model, and the target moment is the largest moment different from multiple moments.
[0056] Specifically, the preset prediction model is composed of a time series prediction model and a Kalman filter model. Among them, the time series prediction model is an ARIMA model (Autoregressive Integrated Moving Average model, differential autoregressive stationary moving average model), which is a time series prediction method that can transform a non-stationary time series into a stationary time series, and make the dependent variable regress on its lag values and the present and lag values of the random error terms, so as to complete the prediction of the data. The Kalman filter model is an algorithm for obtaining the best estimate of a variable, which combines the past measurement estimation errors into the new measurement errors to estimate future errors and perform an optimal estimate of the system state. By combining the two models, the error of the predicted value of the resource load rate can be reduced, thereby improving the accuracy of the predicted value.
[0057] After obtaining the predicted value, the number of pods in the application can be changed according to the predicted value. By elastically scaling the pods in the application in advance, the service quality of the application is ensured.
[0058] Through the above steps, the problem in the related technology that the resource load rate in the application cannot be accurately determined and the number of containers in the application cannot be changed in time after the resource load rate is determined is solved. The accuracy of calculating the comprehensive load rate of the containers is improved, and the effect of accurately predicting the comprehensive load rate of the containers according to the training data can be achieved. Furthermore, the number of containers in the application can be changed in advance according to the predicted value, ensuring that the application can perform elastic scaling operations in a timely and accurate manner.
[0059] Among them, the execution subject of the above steps can be a mobile terminal (electronic device), a computer terminal, or a similar computing device, etc., but not limited to this.
[0060] Optionally, calculate the resource load rate of each resource of the container according to the resource usage and resource occupancy in the resource load dataset. The obtained multiple groups of resource load rates include: determining the resource usage and resource occupancy of each resource occupied by the container; respectively dividing the resource usage of each resource by the resource occupancy to obtain multiple resource load rates; grouping the multiple resource load rates to obtain a group of resource load rates of the container at the same moment.
[0061] Specifically, when calculating the resource load rate of each resource occupied by each pod, it is necessary to determine the resource usage U i and the resource occupancy R i , where i is the i-th resource. After that, the resource usage U i and the resource occupancy R i of each resource can be divided to obtain the resource load rate of the pod for each resource.
[0062] For example, there are many factors that affect the quality of service of an application, which can include multiple basic metrics such as CPU, memory, network, disk I / O, etc. Assume that the types of resources involved in the current pod node are n, and C i represents the resource load rate of the i-th type of resource occupied by the pod. Then, the formula for calculating C i can be:
[0063] C i = U i / R i (5)
[0064] where U i is the resource usage, and R i is the resource occupancy, that is, the resource allocation amount. Thus, the resource load rate of each resource used by the pod is determined, and multiple resource load rates are grouped to obtain a set of resource load rates of the pod at this moment.
[0065] Optionally, calculating the comprehensive load rate based on multiple resource load rates in each group includes: when there is a resource load rate greater than or equal to the first resource load rate among multiple resource load rates, determining the resource load rate greater than or equal to the first resource load rate as the comprehensive load rate; or, when all multiple resource load rates are less than or equal to the second resource load rate, determining the maximum load rate among multiple resource load rates as the comprehensive load rate, where the first resource load rate is greater than the second resource load rate; or, when each resource load rate is greater than the second resource load rate and less than the first resource load rate, performing a weighted sum of multiple resource load rates to obtain the comprehensive load rate.
[0066] It should be noted that after obtaining a set of resource load rates of a certain pod at a certain moment, it is necessary to determine the comprehensive load rate of the pod at this moment according to the resource load rate, so as to comprehensively reflect the overall load level of the current pod node based on the comprehensive load rate, and further provide data support for whether to perform elastic scaling.
[0067] Specifically, denote C max as the maximum resource load rate value in this set of resource load rates. To ensure the quality of service of the application and prevent the situation where only one resource is consumed very high while the resource load rates of other resources are relatively low, resulting in the application not being properly scaled when it should be scaled due to the comprehensive load rate not reaching the scaling threshold. Therefore, it is necessary to determine whether C max in this set of resource load rates is greater than the first load rate, where the first load rate is the upper limit value of the resource load rate. When there is a resource load rate exceeding the upper limit value in this set of resource load rates, the comprehensive load rate of this group, that is, CLR, is determined as C max .
[0068] Further, to reduce the computational complexity, when the load rate of a certain resource is lower than the second resource load rate, that is, the lower threshold, it is considered that the load rate of this resource has little impact on the application load. Therefore, the load rate of this resource will not be used as a parameter for calculating the comprehensive load rate CLR. Thus, in C max is less than the second resource load rate, that is, when the load rates of all resources are less than the second resource load rate, let CLR = C max , thus ensuring that the CLR value is too low, thereby affecting the subsequent prediction of the comprehensive load rate.
[0069] In C max When it is between the first resource load rate and the second resource load rate, the resource load rates with values greater than the second resource load rate among the group of resource load rates can be weighted and summed, so as to reflect the differences between resources in the comprehensive load rate on the premise of simplifying the computational complexity and ensuring the quality of application services, and then accurately and comprehensively reflect the overall load level of the current pod node.
[0070] Optionally, obtaining the comprehensive load rate by weighted summation of multiple resource load rates includes: adding multiple resource load rates to obtain the sum of multiple resource load rates; dividing each resource load rate by the sum of multiple resource load rates in turn to obtain the weight value of each resource load rate; performing weighted summation calculation through each resource load rate and the corresponding weight value to obtain the comprehensive load rate of each group of resource load rates.
[0071] Specifically, when C max is between the first resource load rate and the second resource load rate, since the higher the utilization rate of a certain load, the greater the impact on CLR, it is necessary to first calculate the dynamic weight value of each resource load rate. The formula for the weight value is:
[0072]
[0073] where k is the number of resource load rates with load rate values between the first resource load rate and the second resource load rate in this group of resource load rates, m i is the weight value of the i-th resource load rate, and C i represents the resource load rate of the i-th type of resource occupied by the pod.
[0074] Further, after determining the dynamic weight value of each resource load rate, multiple resource load rates can be weighted and summed to obtain the comprehensive load rate CLR of the pod. The formula for the comprehensive load rate is as follows:
[0075]
[0076] According to the above formula, the CLR of the current pod is finally calculated. Under the premise of simplifying the calculation complexity and ensuring the quality of application services, this CLR metric comprehensively reflects the overall load level of the current pod.
[0077] Optionally, generating a load rate sequence for a container based on the comprehensive load rates at each moment includes: sorting the comprehensive load rates in ascending order of time to obtain a candidate load rate sequence; determining whether the candidate load rate sequence is stationary data, where stationary data is data that fluctuates around a mean; in the case where the candidate load rate sequence is stationary data, determining the candidate load rate sequence as the load rate sequence; in the case where the candidate load rate sequence is not stationary data, performing a stationary processing on the candidate load rate sequence to obtain the load rate sequence.
[0078] Specifically, after calculating the comprehensive load rate at each moment, multiple comprehensive load rates can be sorted in chronological order to obtain a time series of the comprehensive load rate changing over time, thereby obtaining a candidate load rate sequence.
[0079] Furthermore, after obtaining the candidate load rate sequence, in order to ensure the stationarity of the data, it is necessary to use the stationary test method ADF (Augmented Dickey-Fuller Test) to test the candidate load rate sequence to determine whether the candidate load rate sequence is a stationary data sequence. In the case of stationary data, the candidate load rate sequence can be directly determined as the load rate sequence; in the case of non-stationary data, it is necessary to perform a stationary processing through differencing until the ADF test meets the requirements.
[0080] For example, Figure 3 is a schematic diagram of an optional candidate load rate sequence according to an embodiment of the present application Figure 1 As Figure 3 shown, in the case where the image of the candidate load rate sequence is Figure 3 in the style of, it indicates that the data is stationary data. Figure 4 is a schematic diagram of an optional candidate load rate sequence according to an embodiment of the present application Figure 2 As Figure 4 shown, in the case where the image of the candidate load rate sequence is Figure 4 in the style of, it indicates that the data is non-stationary data. At this time, it is necessary to perform a stationary processing on the candidate load rate sequence by differencing, so that the image changes from Figure 4 to Figure 3 in the style of, so that the data meets the test of the stationary test method ADF.
[0081] Optionally, adjusting the number of containers in the target application according to the predicted value includes: determining whether the predicted value is greater than the first predicted value and determining whether the predicted value is less than the second predicted value, where the first predicted value is greater than the second predicted value; adding a preset number of containers to the target application when the predicted value is greater than the first predicted value; clearing a preset number of containers in the target application when the predicted value is less than the second predicted value; and keeping the number of containers in the target application unchanged when the predicted value is less than or equal to the first predicted value and greater than or equal to the second predicted value.
[0082] Specifically, after obtaining the predicted value, the number of pods in the application can be adjusted according to the predicted value. When the predicted value is greater than the first predicted value, it indicates that the load rate of each pod in the application exceeds the upper limit. Therefore, new pods need to be added to the application to relieve the load pressure. When the predicted value is less than the second predicted value, it indicates that the load rate of each pod in the application is less than the lower limit. Therefore, a certain pod in the application needs to be deleted and assigned to other applications that need to add new pods, so as to complete the reasonable allocation of pods. Furthermore, by accurately predicting the application load, the problem of response delay in elastic scaling of the application is solved.
[0083] Optionally, inputting the load rate sequence into a preset prediction model to obtain the predicted value of the comprehensive load rate of the container at the target moment includes: inputting the load rate sequence into a time series prediction model to obtain a candidate predicted value; inputting the load rate sequence and the candidate predicted value into a Kalman filter model, and correcting the candidate predicted value through the Kalman filter model and the load rate sequence to obtain the predicted value of the comprehensive load rate.
[0084] Specifically, the time series prediction model can be an autoregressive integrated moving average model (ARIMA), which is a time series prediction method. This model transforms a non-stationary time series into a stationary time series and regresses the dependent variable on its lagged values and the present and lagged values of the random error term. The advantage of the ARIMA model is that a relatively accurate prediction model can be established with only a limited sample sequence. However, this model has the disadvantages of low prediction accuracy for low-order models and difficult parameter estimation for high-order models.
[0085] Kalman filtering is an algorithm for obtaining the best estimate of a variable, which combines the past measurement estimation error into the new measurement error to estimate the future error and optimally estimate the system state. The Kalman filtering algorithm is in a recursive form and does not require all the data. Only the measurement value at time t k is used to correct the estimated value at time t k-1 and has the characteristic of dynamic weighted correction, with good prediction accuracy. However, the Kalman filtering algorithm requires a state equation and a measurement equation to ensure good prediction accuracy.
[0086] Therefore, based on the above description of the ARIMA prediction model and the Kalman filter model, the two prediction models can be combined by using the ARIMA-Kalman model to reduce the impact of the defects in the two models on the prediction results. The ARIMA-Kalman model uses the sequence of the earlier moments in the historical load rate sequence as input, and compares the obtained results with the comprehensive load rate in the sequence of the later moments to complete the training of the ARIMA-Kalman model, thereby obtaining a trained prediction model.
[0087] Firstly, the smoothed load rate sequence is input into the ARIMA-Kalman model, and a low-order prediction model is established using the ARIMA model. After processing the low-order model, the state equation and measurement equation of the Kalman filter model are calculated, and the Kalman iteration equation is used for prediction to obtain an accurate prediction value of the comprehensive load rate.
[0088] Through the description of the above implementation methods, those skilled in the art can clearly understand that the method according to the above embodiment can be implemented by means of software plus a necessary general hardware platform, and of course by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of the present application, or the part that contributes to the prior art, can be embodied in the form of a software product, which is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk), and includes a number of instructions for a terminal device (which can be a mobile phone, computer, server, or network device, etc.) to execute the methods described in each embodiment of the present application.
[0089] The following is an experimental process and experimental results based on the technical method in this application.
[0090] First, in order to conduct a comparative experiment, two identical Kubernetes cluster environments were built, both of which were Kubernetes version 1.20.5, and each cluster contained a Master node and two Slave nodes. One cluster used the built-in elastic scaling method, and the other cluster used the container quantity adjustment method proposed in the embodiment of the present application.
[0091] Secondly, it is necessary to select experimental data, that is, multiple resource load data sets. The data sets include resource usage in multiple dimensions such as CPU utilization, memory utilization, network and disk IO. Since the time intervals of the public data sampling are not equal, the points with close time intervals are removed and the mean difference is used to supplement the data with longer time intervals in the middle to make the time intervals roughly equal.
[0092] Furthermore, it is necessary to determine the content to be detected. The method for adjusting the number of containers proposed in the embodiments of this application needs to be experimentally verified from two aspects: functionality and accuracy. Therefore, the experiment should include two parts: 1) Use the JMeter tool to perform a stress test on the pods of Kubernetes, record the changes in the number of pods in the two clusters respectively, and make a comparison. Verify the predictive scaling effect of the pod number elastic scaling method in this article when dealing with load changes. 2) Use the publicly available container load information to calculate the CLR sequence. Use the exponential smoothing method, the ARIMA prediction model, and the ARIMA-Kalman prediction model for prediction respectively, and evaluate the prediction accuracy of these three prediction models.
[0093] When conducting the experiment, first deploy the same Web application in two experimental environments respectively, set the scaling threshold to 60%, and the tolerance to the default 0.1. Use the JMeter tool to simulate concurrent access requests, increase the number of concurrent requests every 1 minute, and check the current number of pods. Among them, in the experimental environment using the pod number elastic scaling method in this article, Prometheus is used to obtain the resource information of the pods and calculate the CLR for prediction. Table 1 is a schematic diagram of the optional changes in the number of pods according to the embodiments of this application.
[0094] Table 1
[0095]
[0096] It can be seen from Table 1 that compared with the built-in scaling policy method, the method used in this application can perform predictive elastic scaling in advance according to the change trend of the load, thus solving the response delay problem of the built-in scaling policy of Kubernetes and ensuring the quality of the service.
[0097] In order to compare the accuracy of the ARIMA-Kalman prediction model used in this article with other models, after processing multiple resource load data sets, calculate the CLR at multiple moments, and use the CLR as the input for model prediction. Use the three most common measurement indicators: absolute error (MAE), mean absolute error (MSE), and mean squared root error (MSRE) to evaluate the exponential smoothing method, the ARIMA prediction model, and the ARIMA-Kalman prediction model used in this article.
[0098] The processed sampling data for each container is approximately 650 - 800. Uniformly select the first 550 information nodes as training data, and the last 100 data as prediction. The error graphs of the predicted values of the exponential smoothing method, the ARIMA prediction model, and the ARIMA-Kalman are as Figure 5 shown.
[0099] ByFigure 5 It can be intuitively seen that the error of ARIMA-Kalman is relatively smaller than that of the other two methods. Since the process of determining the parameters of the ARIMA model involves a large amount of computation and is not suitable for dynamically updating the data model, but only for short-term prediction. In this paper, 100 data points are predicted at one time, so the overall error is relatively large. The exponential smoothing method consumes more memory resources during use because it uses the information of all historical nodes and does not take into account the regularity of data changes. The ARIMA-Kalman model determines the state transition equation by establishing a low-order ARIMA model and uses the Kalman model for iterative estimation, resulting in a significant improvement in its accuracy compared to the ARIMA and exponential smoothing methods. Table 2 shows the comparison results of the evaluation indicators of each model.
[0100] Table 2
[0101] Model / Evaluation Index MAE MSE RMSE Exponential Smoothing Method 0.02269 0.00185 0.04301 ARIMA 0.02392 0.00135 0.03681 ARIMA-Kalman 0.01208 0.00038 0.01959
[0102] It can be seen that compared with the exponential smoothing method and the ARIMA model, the prediction accuracy of the ARIMA-Kalman prediction model used in this paper is better. In the face of load changes, it can accurately make predictions.
[0103] In this embodiment, a device for adjusting the number of containers is also provided. This device is used to implement the above-mentioned embodiments and preferred implementation manners, and those that have been described will not be repeated. As used hereinafter, the term "module" can be a combination of software and / or hardware that can achieve a predetermined function. Although the devices described in the following embodiments are preferably implemented in software, implementation in hardware, or a combination of software and hardware is also possible and contemplated.
[0104] Figure 6 is a schematic diagram of the device for adjusting the number of containers according to an embodiment of the present application. As Figure 6 shown, the device includes:
[0105] An acquisition module 61, configured to acquire resource load data sets of any one container in a target application at multiple moments, obtain multiple resource load data sets, and calculate the resource load rates of the container for each resource based on the resource usage and resource occupancy in the resource load data sets respectively, to obtain multiple groups of resource load rates. Among them, the target application includes at least one container, and the resource load data sets of each container are the same.
[0106] A calculation module 62, configured to calculate a comprehensive load rate based on the multiple resource load rates in each group of resource load rates, to obtain the comprehensive load rate of the container at each moment.
[0107] A generation module 63, configured to generate a load rate sequence of the container based on the comprehensive load rates of the container at each moment.
[0108] A prediction module 64, configured to input a load rate sequence into a preset prediction model to obtain a predicted value of the comprehensive load rate of a container at a target time, and adjust the number of containers in a target application according to the predicted value. The preset prediction model includes a time series prediction model and a Kalman filter model, and the target time is the maximum time different from multiple times.
[0109] The device for adjusting the number of containers provided in the embodiment of the present application includes an acquisition module 61, configured to acquire resource load data sets of any one container in a target application at multiple times, obtain multiple resource load data sets, and calculate the resource load rate of each container for each resource according to the resource usage amount and resource occupancy amount in the resource load data sets, to obtain multiple groups of resource load rates. The target application includes at least one container, and the resource load data sets of each container are the same. A calculation module 62, configured to calculate a comprehensive load rate according to multiple resource load rates in each group of resource load rates, to obtain the comprehensive load rate of the container at each time. A generation module 63, configured to generate a load rate sequence of the container according to the comprehensive load rates of the container at each time. A prediction module 64, configured to input the load rate sequence into a preset prediction model to obtain a predicted value of the comprehensive load rate of the container at a target time, and adjust the number of containers in the target application according to the predicted value. The preset prediction model includes a time series prediction model and a Kalman filter model, and the target time is the maximum time different from multiple times. The effect of accurately predicting the comprehensive load rate of the container according to the training data can be achieved. Furthermore, the number of containers in the application can be changed in advance according to the predicted value, ensuring that the application can perform elastic scaling operations in a timely and accurate manner. Therefore, the problem in the related art that the resource load rate in the application cannot be accurately determined and the number of containers in the application cannot be changed in time after the resource load rate is determined can be effectively solved.
[0110] Optionally, in the device for adjusting the number of containers provided in the embodiment of the present application, the calculation module 62 includes: a first determination sub-module, configured to determine, when there is a resource load rate greater than or equal to a first resource load rate among multiple resource load rates, the resource load rate greater than or equal to the first resource load rate as the comprehensive load rate; or a second determination sub-module, configured to determine, when all multiple resource load rates are less than or equal to a second resource load rate, the maximum load rate among the multiple resource load rates as the comprehensive load rate, where the first resource load rate is greater than the second resource load rate; or a first calculation sub-module, configured to, when each resource load rate is greater than the second resource load rate and less than the first resource load rate, perform a weighted sum of the multiple resource load rates to obtain the comprehensive load rate.
[0111] Optionally, in the container quantity adjustment device provided by the embodiments of the present application, the first calculation sub-module includes: a first calculation unit, configured to add multiple resource load rates to obtain the sum of the multiple resource load rates; a second calculation unit, configured to divide each resource load rate by the sum of the multiple resource load rates in sequence to obtain the weight of each resource load rate; a third calculation unit, configured to perform weighted summation calculation through each resource load rate and the corresponding weight to obtain the comprehensive load rate of each group of resource load rates.
[0112] Optionally, in the container quantity adjustment device provided by the embodiments of the present application, the generation module 63 includes: a sorting sub-module, configured to sort the comprehensive load rates in ascending order of time to obtain a candidate load rate sequence; a first judgment sub-module, configured to judge whether the candidate load rate sequence is stationary data, where the stationary data is data that fluctuates around the mean; a third determination sub-module, configured to determine the candidate load rate sequence as the load rate sequence when the candidate load rate sequence is stationary data; a processing sub-module, configured to perform stationary processing on the candidate load rate sequence to obtain the load rate sequence when the candidate load rate sequence is not stationary data.
[0113] Optionally, in the container quantity adjustment device provided by the embodiments of the present application, the prediction module 64 includes: a second judgment sub-module, configured to judge whether the predicted value is greater than the first predicted value and whether the predicted value is less than the second predicted value, where the first predicted value is greater than the second predicted value; a first addition sub-module, configured to add a preset number of containers to the target application when the predicted value is greater than the first predicted value; a second addition sub-module, configured to remove a preset number of containers from the target application when the predicted value is less than the second predicted value; a maintaining sub-module, configured to keep the number of containers in the target application unchanged when the predicted value is less than or equal to the first predicted value and greater than or equal to the second predicted value.
[0114] Optionally, in the container quantity adjustment device provided by the embodiments of the present application, the prediction module 64 includes: a first input sub-module, configured to input the load rate sequence into a time series prediction model to obtain a candidate predicted value; a second input sub-module, configured to input the load rate sequence and the candidate predicted value into a Kalman filter model, and correct the candidate predicted value through the Kalman filter model and the load rate sequence to obtain the predicted value of the comprehensive load rate.
[0115] Optionally, in the apparatus for adjusting the number of containers provided in the embodiments of the present application, the obtaining module 61 includes: a fourth determining sub-module, configured to determine the resource usage amount and resource occupancy amount of each resource occupied by the container; a second calculating sub-module, configured to divide the resource usage amount of each resource by the resource occupancy amount respectively to obtain a plurality of resource load ratios; and a grouping sub-module, configured to group the plurality of resource load ratios to obtain a set of resource load ratios of the container at the same moment.
[0116] It should be noted that the above-mentioned various modules can be implemented by software or hardware. For the latter, it can be implemented in the following ways, but not limited thereto: the above-mentioned modules are all located in the same processor; or, the above-mentioned various modules are respectively located in different processors in any combination form.
[0117] The embodiments of the present application further provide a computer-readable storage medium, in which a computer program is stored, and the computer program is configured to execute the steps in any one of the above method embodiments when running.
[0118] In an exemplary embodiment, the above computer-readable storage medium may include, but is not limited to: various media such as a USB flash drive, a read-only memory (ROM), a random access memory (RAM), a mobile hard disk, a magnetic disk, or an optical disc that can store a computer program.
[0119] The embodiments of the present application further provide an electronic device, including a memory and a processor, where a computer program is stored in the memory, and the processor is configured to run the computer program to execute the steps in any one of the above method embodiments.
[0120] In an exemplary embodiment, the above electronic device may further include a transmission device and an input / output device, where the transmission device is connected to the above processor, and the input / output device is connected to the above processor.
[0121] The specific examples in this embodiment may refer to the examples described in the above embodiments and exemplary embodiments, and will not be repeated here.
[0122] Obviously, those skilled in the art should understand that the various modules or steps of the present application described above can be implemented by a general-purpose computing device. They can be concentrated on a single computing device or distributed on a network composed of multiple computing devices. They can be implemented by program codes executable by the computing device. Thus, they can be stored in a storage device and executed by the computing device. And in some cases, the steps shown or described can be executed in a sequence different from that here, or they can be separately fabricated into individual integrated circuit modules, or multiple modules or steps among them can be fabricated into a single integrated circuit module for implementation. In this way, the present application is not limited to any specific combination of hardware and software.
[0123] The foregoing is only a preferred embodiment of the present application and is not intended to limit the present application. For those skilled in the art, the present application can have various changes and modifications. Any modification, equivalent replacement, improvement, etc. made within the principle of the present application shall be included within the protection scope of the present application.
Claims
1. A method for adjusting the number of containers, characterized in that, Including: Obtain resource load data sets of any one container in a target application at multiple moments, obtaining multiple resource load data sets, and respectively calculate the resource load rates of the container for each resource based on the resource usage amount and resource occupancy amount in the resource load data sets, obtaining multiple groups of resource load rates, where the target application includes at least one container, and the resource load data sets of each container are the same; Calculate a comprehensive load rate based on the multiple resource load rates in each group of resource load rates, obtaining the comprehensive load rate of the container at each moment; Generate a load rate sequence of the container according to the comprehensive load rates of the container at each moment; Input the load rate sequence into a preset prediction model to obtain a predicted value of the comprehensive load rate of the container at a target moment, and adjust the number of containers in the target application according to the predicted value, where the preset prediction model includes a time series prediction model and a Kalman filter model, and the target moment is the largest moment different from the multiple moments; Wherein, calculating the comprehensive load rate based on the multiple resource load rates in each group of resource load rates includes: In the case where there is a resource load rate greater than or equal to a first resource load rate among the multiple resource load rates, determine the resource load rate greater than or equal to the first resource load rate as the comprehensive load rate, or; In the case where all the multiple resource load rates are less than or equal to a second resource load rate, determine the maximum load rate among the multiple resource load rates as the comprehensive load rate, where the first resource load rate is greater than the second resource load rate, or; In the case where each resource load rate is greater than the second resource load rate and less than the first resource load rate, perform a weighted sum of the multiple resource load rates to obtain the comprehensive load rate.
2. The method according to claim 1, characterized in that Performing a weighted sum of the multiple resource load rates to obtain the comprehensive load rate includes: Add the multiple resource load rates to obtain the sum of the multiple resource load rates; Successively divide each resource load rate by the sum of the multiple resource load rates to obtain the weight value of each resource load rate; Perform a weighted sum calculation through each resource load rate and the corresponding weight value to obtain the comprehensive load rate of each group of resource load rates.
3. The method according to claim 1, characterized in that, Generating the load rate sequence of the container according to the comprehensive load rates of the container at each moment includes: Sort the comprehensive load rates in ascending order of the moments to obtain a candidate load rate sequence; Determine whether the candidate load rate sequence is stationary data, where the stationary data is data that fluctuates around the mean; In the case where the candidate load rate sequence is stationary data, determine the candidate load rate sequence as the load rate sequence; In the case where the candidate load rate sequence is not stationary data, perform a stationary processing on the candidate load rate sequence to obtain the load rate sequence.
4. The method according to claim 1, characterized in that, Adjusting the number of containers in the target application according to the predicted value includes: Determine whether the predicted value is greater than a first predicted value, and determine whether the predicted value is less than a second predicted value, where the first predicted value is greater than the second predicted value; When the predicted value is greater than the first predicted value, add a preset number of containers to the target application; When the predicted value is less than the second predicted value, remove the preset number of containers from the target application; When the predicted value is less than or equal to the first predicted value and greater than or equal to the second predicted value, keep the number of containers in the target application unchanged.
5. The method according to claim 1, characterized in that Inputting the load rate sequence into a preset prediction model to obtain the predicted value of the comprehensive load rate of the container at the target moment includes: Inputting the load rate sequence into the time series prediction model to obtain a candidate predicted value; Inputting the load rate sequence and the candidate predicted value into the Kalman filter model, and correcting the candidate predicted value through the Kalman filter model and the load rate sequence to obtain the predicted value of the comprehensive load rate.
6. The method according to claim 1, characterized in that, Calculating the resource load rate of the container for each resource according to the resource usage amount and resource occupancy amount in the resource load data set, and obtaining multiple groups of resource load rates, including: Determining the resource usage amount and resource occupancy amount of each resource occupied by the container; Dividing the resource usage amount of each resource by the resource occupancy amount respectively to obtain a plurality of resource load rates; Grouping the plurality of resource load rates to obtain a group of resource load rates of the container at the same moment.
7. An adjustment device for the number of containers, characterized in that Including: An acquisition module, configured to acquire resource load data sets of any one container in the target application at multiple moments, obtain multiple resource load data sets, and calculate the resource load rate of the container for each resource according to the resource usage amount and resource occupancy amount in the resource load data set, respectively, to obtain multiple groups of resource load rates, where the target application includes at least one container, and the resource load data sets of each container are the same; A calculation module, configured to calculate a comprehensive load rate according to the multiple resource load rates in each group of resource load rates to obtain the comprehensive load rate of the container at each moment; A generation module, configured to generate a load rate sequence of the container according to the comprehensive load rate of the container at each moment; A prediction module, configured to input the load rate sequence into a preset prediction model to obtain the predicted value of the comprehensive load rate of the container at the target moment, and adjust the number of containers in the target application according to the predicted value, where the preset prediction model includes a time series prediction model and a Kalman filter model, and the target moment is the maximum moment different from the multiple moments; Wherein, the calculation module includes: A first determination sub-module, configured to, when there is a resource load rate greater than or equal to the first resource load rate among the multiple resource load rates, determine the resource load rate greater than or equal to the first resource load rate as the comprehensive load rate, or; A second determination sub-module, configured to, when all the multiple resource load rates are less than or equal to the second resource load rate, determine the maximum load rate among the multiple resource load rates as the comprehensive load rate, where the first resource load rate is greater than the second resource load rate, or; The first calculation sub-module is configured to perform a weighted sum of multiple resource load rates to obtain the comprehensive load rate when each of the resource load rates is greater than the second resource load rate and less than the first resource load rate.
8. A computer-readable storage medium, characterized in that, A computer program is stored in the computer-readable storage medium, wherein when the computer program is executed by a processor, the steps of the method described in any one of claims 1 to 6 are implemented.
9. An electronic device, comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that, When the processor executes the computer program, the steps of the method described in any one of claims 1 to 6 are implemented.
Citation Information
Patent Citations
Kubernetes-based container number flexible adjustment implementation method and device
CN108769100A
Container operation method and system of cloud platform and related device
CN114296867A