A container cluster hybrid deployment system, a container scheduling method and related devices
Through the coordinated work of load scheduler and rescheduler, dynamic allocation and scheduling of container services is solved, and the service operation quality and business continuity are improved.
Patent Information
- Application Number
- CN202510511615.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-23
- Publication Date
- 2025-07-04
- Estimated Expiration
- 2045-04-23
AI Technical Summary
The existing container cluster scheduling methods fail to effectively respond to the dynamic changes in node load, resulting in resource imbalance, affecting service operation quality and business continuity, and lacking a flexible load balancing mechanism.
Using a combination of load scheduler and rescheduler, container services are dynamically allocated by predicting the load conditions of nodes, and secondary scheduling is performed when necessary, low-priority services are turned off to alleviate overload and ensure the stable operation of high-priority services.
It realizes the balanced allocation of node resources in the container cluster, improves service operation quality, reduces resource competition, and ensures business stability and continuity.
Smart Images

Figure CN120034542B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of cloud computing technology, and in particular, to a container cluster hybrid deployment system, a container scheduling method, and related devices. Background Art
[0002] The container cluster hybrid deployment technology refers to the technology of mixing and deploying multiple different types of services in a container cluster to fully utilize the resources of the container cluster for business. Among them, a container cluster is a distributed system including nodes and containers. In a container cluster, a container is the smallest deployment unit, which can encapsulate the program code of a service and its dependencies, and realize the operation of the service by running the container service on a node.
[0003] A container cluster can schedule container services to run on different nodes, but the running resources of each node are limited, and there are often problems of multiple container services competing for running resources on the node, thus reducing the running quality of the services. Summary of the Invention
[0004] In view of the above problems, this application provides a container cluster hybrid deployment system, a container scheduling method, and related devices to achieve the purpose of improving the service running quality of the container cluster. The specific solutions are as follows:
[0005] In a first aspect of this application, a container scheduling method for a container cluster hybrid deployment system is provided. The container cluster hybrid deployment system includes a load scheduler, a rescheduler, and multiple nodes. The container scheduling method of the container cluster hybrid deployment system includes:
[0006] During the operation of the container cluster, the load scheduler predicts the load conditions of multiple nodes respectively, and allocates the nodes to at least one container service according to the load conditions;
[0007] During the operation of the container service, if a first node meets the secondary scheduling trigger condition and is not overloaded, the rescheduler determines a second node whose running load condition meets the low load condition from multiple other nodes. At least one container service is running on the first node;
[0008] The rescheduler determines a first container service from at least one container service running on the first node according to the type level of the container service, and schedules the first container service to the second node;
[0009] If the first node is overloaded, the rescheduler shuts down a second container service running on the first node, and the type level of the second container service is lower than the type levels of other container services running on the first node.
[0010] In an alternative approach, the load scheduler predicts the load conditions of multiple nodes respectively, including:
[0011] The load scheduler obtains the historical load data of each node for a historical period and the current load data at the current moment respectively;
[0012] The load scheduler inputs the historical load data and the current load data into a load prediction model, obtains the predicted load scores for a future period output by the load prediction model, and uses the predicted load scores as the load conditions of the nodes.
[0013] In an alternative approach, the historical load data includes the historical load scores for the historical period, and the current load data includes the current load scores at the current moment;
[0014] The load scheduler obtains the historical load data of each node for a historical period and the current load data at the current moment respectively, including:
[0015] The load scheduler collects the historical CPU utilization rate, historical memory usage, historical network bandwidth, and historical disk read / write times of each node for the historical period respectively, and calculates the historical load scores for the historical period according to the historical CPU utilization rate, the historical memory usage, the historical network bandwidth, and the historical disk read / write times;
[0016] The load scheduler collects the current CPU utilization rate, current memory usage, current network bandwidth, and current disk read / write times of each node at the current moment respectively, and calculates the current load scores at the current moment according to the current CPU utilization rate, the current memory usage, the current network bandwidth, and the current disk read / write times.
[0017] In an alternative approach, the load condition of the node is the predicted load score of the node;
[0018] The allocating the node for at least one container service according to the load condition includes:
[0019] For each of the container services, when the container service needs to run, the load scheduler allocates the node with the minimum current predicted load score for the container service.
[0020] In an alternative approach, the low load condition is: the predicted load score of the node does not exceed a first preset threshold and is the minimum;
[0021] The rescheduler determines a second node whose running load condition meets the low load condition from multiple other nodes, including:
[0022] The rescheduler obtains the predicted load scores of each of the other nodes, and filters out those other nodes whose predicted load scores do not exceed a first preset threshold;
[0023] Selects the other node with the smallest predicted load score from the part of the other nodes as the second node.
[0024] In an alternative manner, on the first node, the type level of the first container service is lower than the type levels of other container services.
[0025] In an alternative manner, it further includes:
[0026] If the load score of the first node is not less than a second preset threshold, it is confirmed that the first node is overloaded, and the second preset threshold is greater than the first preset threshold.
[0027] The second aspect of the present application provides a container cluster hybrid deployment system, which includes a load scheduler, a rescheduler, and multiple nodes;
[0028] The load scheduler is used to predict the load conditions of multiple nodes respectively during the operation of the container cluster, and allocate the nodes for at least one container service according to the load conditions;
[0029] The rescheduler is used to, during the operation of the container service, if a first node meets the secondary scheduling trigger condition and is not overloaded, determine a second node whose running load condition meets the low load condition from multiple other nodes, and at least one container service is running on the first node;
[0030] The rescheduler is further used to determine a first container service from at least one container service running on the first node according to the type level of the container service, and schedule the first container service to the second node;
[0031] If the first node is overloaded, the rescheduler is further used to shut down a second container service running on the first node, and the type level of the second container service is lower than the type levels of other container services running on the first node.
[0032] The third aspect of the present application provides an electronic device, including at least one processor and a memory connected to the processor, wherein:
[0033] The memory is used to store a computer program;
[0034] The processor is used to execute the computer program so that the electronic device can implement the container scheduling method of the container cluster hybrid deployment system in the above first aspect or any implementation manner of the first aspect.
[0035] A fourth aspect of the present application provides a computer program product, including computer-readable instructions, which, when running on an electronic device, enable the electronic device to implement the container scheduling method of the container cluster hybrid deployment system in the first aspect or any implementation manner of the first aspect.
[0036] By means of the above technical solutions, the present application provides a container cluster hybrid deployment system, a container scheduling method and related devices. The container cluster hybrid deployment can include multiple nodes, a load scheduler and a rescheduler. In this method, during the operation of the container cluster, the load conditions of each node are predicted respectively, and nodes for running are allocated to at least one container service according to the predicted load conditions. And during the subsequent operation of the container service, if the first node meets the secondary scheduling trigger condition and is not overloaded, a second node with a load condition meeting the low-load condition can be selected from other nodes, and the first container service running on the first node is scheduled to the second node. If the first node is overloaded, the second container service with a lower type level running on the first node is directly shut down. This method first allocates nodes to each container service according to the predicted load conditions of the nodes, ensuring that each container service can match a suitable node and preventing the problem of excessive load immediately after the node runs the container service. And when it is found that a node is in a high-load condition during the operation of the container service, by scheduling the container service to other nodes with lower loads, the container services running on the overloaded node are reduced, so as to alleviate the intensity of competition for running resources among container services. If the node is already overloaded, the second container service with a lower type level is directly selected to be shut down to ensure the operation of other container services with higher type levels. Therefore, this method can effectively improve the running quality of the service. BRIEF DESCRIPTION OF THE DRAWINGS
[0037] In combination with the accompanying drawings and with reference to the following specific embodiments, the above and other features, advantages and aspects of the embodiments of the present disclosure will become more obvious. Throughout the drawings, the same or similar reference numerals denote the same or similar elements. It should be understood that the drawings are schematic and the original elements and elements are not necessarily drawn to scale.
[0038] Figure 1 It is a schematic flowchart of a container scheduling method for a container cluster hybrid deployment system provided by an embodiment of the present application;
[0039] Figure 2 It is a hardware structure block diagram of an electronic device provided by the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0040] The following describes the embodiments of the present application in combination with the accompanying drawings in the embodiments of the present application. The terms used in the embodiments part of the present application are only used to explain the specific embodiments of the present application, and are not intended to limit the present application.
[0041] The embodiments of the present application will be described below in conjunction with the accompanying drawings. As is known to those of ordinary skill in the art, with the development of technology and the emergence of new scenarios, the technical solutions provided by the embodiments of the present application are equally applicable to similar technical problems.
[0042] The terms "first", "second", etc. in the specification of the present application and the above-mentioned drawings are used to distinguish similar objects and do not necessarily need to be used to describe a specific order or sequence. It should be understood that such terms can be interchanged under appropriate circumstances, which is only a way of distinguishing objects with the same attributes when describing the embodiments of the present application. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion, so that a process, method, system, product or device comprising a series of units does not have to be limited to those units, but may include other units not clearly listed or inherent to these processes, methods, products or devices.
[0043] In cloud computing and large-scale distributed computing environments, service containerization has become the mainstream application deployment method. Container cluster co-location means that different types of services are simultaneously and mixedly deployed on each node of the same container cluster, and the container cluster can schedule container services so that the container services run on different nodes. However, there are many key problems in the actual application of traditional container service scheduling and resource management methods.
[0044] First, the current container service scheduling method only considers static resource information during initial scheduling, fixes the node corresponding to each container service in advance, and fails to fully consider the dynamic load changes of container services on the nodes. Since the load of the nodes may fluctuate greatly when the container services are running, it may lead to resource imbalance between the nodes, with some nodes having too high resource utilization and some nodes having idle resources, thus affecting the overall operation efficiency of the container cluster.
[0045] Second, the current container service scheduling method lacks service quality guarantee. The current container service scheduling adopts a priority and preemption mechanism, allowing high-priority services to terminate low-priority services to ensure resource allocation for critical services. However, due to the priority and preemption mechanism, it is unable to actively perform load balancing, resulting in low-priority services being frequently interrupted, thus affecting business continuity and service stability in high-load or resource competition situations. And it can only control resource competition through simple priority values, making it difficult to flexibly handle complex business scenarios.
[0046] Third, in a co-location environment, when different container services share node resources, it is easy for a single container to occupy too much resources, which in turn affects the normal operation of other container services and leads to a decline in the overall performance of the entire container cluster.
[0047] Fourth, after the first scheduling of the container service, due to the dynamic changes in the load, local resource competition and bottleneck problems may occur on some nodes. The existing container scheduling mechanism mainly makes a one-time allocation when the container service starts, without considering the change of node load over time. Once the load of a node changes, it is difficult to automatically adjust the resource allocation, resulting in intensified local resource competition and difficulty in achieving the balanced distribution and continuous stable operation of the operating resources of nodes within the container cluster.
[0048] To solve the above problems, the embodiments of the present application provide a container scheduling method for a container cluster hybrid system. When making the initial allocation, this method predicts the load conditions of nodes to prevent the load from changing too much after the nodes run the container service. And after the initial allocation, if the load of a node is too high, this method can also perform rescheduling of the container service during the operation of the container cluster, and schedule the container service with a lower type level to other nodes with lower load to balance the load and operating resources of the nodes. The following will introduce in detail the container scheduling method for the container cluster hybrid system of the embodiments of the present application with reference to the accompanying drawings.
[0049] The container cluster hybrid system may include a load scheduler, a rescheduler, and multiple nodes. Figure 1 It is a schematic flowchart of a container scheduling method for a container cluster hybrid system provided by an embodiment of the present application. As Figure 1 shown, a container scheduling method for a container cluster hybrid system provided by an embodiment of the present application may include steps S10 to S13, and the following will describe these steps in detail respectively.
[0050] S10. During the operation of the container cluster, the load scheduler predicts the load conditions of multiple nodes respectively, and allocates nodes for at least one container service according to the load conditions.
[0051] Among them, the load scheduler may be a device or module that senses the node load. The load sensing of the load scheduler can monitor and analyze the load metrics of each node in the container cluster in real time, such as CPU utilization rate, memory occupancy, network bandwidth, and disk IOPS, etc. Among them, disk IOPS (Input / Output Operations Per Second, the number of input / output operations per second) represents the number of I / O (input / output) requests that the system can process per unit time, and is one of the main metrics for measuring disk performance. The container service may refer to a container image obtained by encapsulating the program code of the service and its dependencies, and is used to deploy, manage, and expand services in the container cluster. A node is the basic unit of the container cluster, and is used to provide computing resources to run the container service. Specifically, in the container cluster, a node may be a physical server (such as a host), a virtual machine, or a cloud service instance, etc.
[0052] Furthermore, whenever a container service needs to allocate nodes, it is necessary to predict the load conditions of all current nodes to allocate the most suitable nodes for running the container service at present. For example, before allocating the A container service, predict the load conditions of all nodes and allocate nodes for the A container service. The next container service of the A container service is the B container service. Then, after allocating nodes for the A container service, the A container service is already running on the nodes. Predict the load conditions of all nodes again and allocate nodes for the B container service.
[0053] Specifically, the specific process of the load scheduler predicting the load conditions of multiple nodes can be shown in Step 1 and Step 2 as follows:
[0054] Step 1: The load scheduler respectively obtains the historical load data of each node in the historical period and the current load data at the current moment;
[0055] Step 2: The load scheduler inputs the historical load data and the current load data into the load prediction model to obtain the predicted load score in the future period output by the load prediction model, and uses the predicted load score as the load condition of the node.
[0056] Among them, the load scheduler can periodically collect the load data of the nodes to meet the subsequent requirements of load prediction. The historical load data can refer to the load data of the node in the historical period, such as at least one of the historical CPU utilization rate, historical memory usage, historical network bandwidth, and historical disk read and write times. The current load data can refer to the load data of the node at the current moment, such as at least one of the current CPU utilization rate, current memory usage, current network bandwidth, and current disk read and write times. In this embodiment, the historical load data can include the historical load score in the historical period, and the historical load score can be calculated based on at least one of the historical CPU utilization rate, historical memory usage, historical network bandwidth, and historical disk read and write times. The current load data includes the current load score at the current moment, and the current load score can be calculated based on at least one of the current CPU utilization rate, current memory usage, current network bandwidth, and current disk read and write times. The load prediction model can refer to a model used for load prediction. Specifically, it can be an ARIMA model. The ARIMA model (Autoregressive Integrated Moving Average Model) is a statistical model widely used in time series analysis and prediction.
[0057] Specifically, the acquisition of the historical load score and the current load score can be shown in Step 3 and Step 4 as follows:
[0058] Step 3: The load scheduler collects the historical CPU utilization, historical memory usage, historical network bandwidth, and historical disk read / write counts of each node during a historical period, and calculates the historical load score for the historical period based on the historical CPU utilization, historical memory usage, historical network bandwidth, and historical disk read / write counts.
[0059] Step 4: The load scheduler collects the current CPU utilization, current memory usage, current network bandwidth, and current disk read / write counts of each node at the current moment, and calculates the current load score at the current moment based on the current CPU utilization, current memory usage, current network bandwidth, and current disk read / write counts.
[0060] Among them, the CPU utilization can refer to the ratio between the workload of the central processing unit (CPU) of a node and its available processing capacity, which is used to measure the CPU usage efficiency of the node; the memory usage refers to the amount of memory consumed by the node, which can reflect the utilization of the current memory resources of the node; the network bandwidth can refer to the amount of data that can be transmitted by the node network per unit time, which is used to measure the network performance of the node; the disk read / write counts can refer to the number of I / O (input / output) requests processed by the node per unit time. Further, in this embodiment, the average values of the CPU utilization, memory usage, network bandwidth, and disk read / write counts in the historical period can be calculated by taking the average, and the average values are used as the historical CPU utilization, historical memory usage, historical network bandwidth, and historical disk read / write counts for the historical period.
[0061] In this embodiment, a multi-dimensional weighted scoring function can be constructed in the load evaluation model, and the historical load score and the current load score are calculated through the load evaluation model. In the multi-dimensional weighted scoring function, the sum of the dimension weights of the parameters in each dimension is 1, and the dimension weights can be adjusted according to different application scenarios of the container cluster. Specifically, in this embodiment, the historical CPU utilization, historical memory usage, historical network bandwidth, and historical disk read / write counts of the historical period can be input into the load evaluation model, and weighted summation is performed in the load evaluation model to obtain the historical load score output by the load evaluation model; the current CPU utilization, current memory usage, current network bandwidth, and current disk read / write counts at the current moment are input into the load evaluation model, and weighted summation is performed in the load evaluation model to obtain the current load score output by the load evaluation model.
[0062] After obtaining the historical load score and the current load score of each node in this embodiment, they can be input into the load prediction model to obtain the predicted load score for the future period output by the load prediction model, and the predicted load score is used as the predicted load condition of the node. Among them, the load prediction model can also output the predicted load score at a certain future moment.
[0063] When allocating container services in this embodiment, not only static resource information is considered, but also some actual parameters during node runtime (such as CPU utilization, memory usage, network bandwidth, and disk read / write times, etc.) are taken into account. By adopting a data analysis and prediction mechanism, the load situation of the node is predicted, the future dynamic load changes of the node are understood, and then the container services are allocated according to the predicted load situation of the node. Moreover, each time a node is allocated, the load situation of the node is re-predicted, considering the dynamic load changes brought by the changes of container services on the node, realizing the dynamic policy adjustment of container service allocation, ensuring that each container service can match the most suitable node, reducing resource waste, and improving resource utilization rate.
[0064] Certainly, in another alternative embodiment, the historical CPU utilization, historical memory usage, historical network bandwidth, historical disk read / write times, current CPU utilization, current memory usage, current network bandwidth, and current disk read / write times of the node can be directly input into the load prediction model to obtain the predicted CPU utilization, predicted memory usage, predicted network bandwidth, and predicted disk read / write times of the node at a future moment or in a future time period predicted by the load prediction model, as the predicted load situation of the node. Of course, it is also possible to only select one parameter or at least two parameters among CPU utilization, memory usage, network bandwidth, and disk read / write times and input them into the load prediction model to only predict the situation of one parameter or at least two parameters at a future moment or in a future time period.
[0065] In another alternative embodiment, the historical CPU utilization, historical memory usage, historical network bandwidth, historical disk read / write times, previous CPU utilization, current memory usage, current network bandwidth, and current disk read / write times of the node can also be input into the load prediction model. The load prediction model first calculates the historical load score and the current load score, and then makes a prediction by synthesizing the historical load score and the current load score to obtain the predicted load score of the node at a future moment or in a future time period.
[0066] After this embodiment obtains the predicted load score of each node at the current moment, for each container service, when the container service needs to run, the load scheduler can allocate the node with the minimum current predicted load score to the container service. Specifically, whenever a container service needs to run in the container cluster, the load scheduler will calculate the predicted load score of each node once and select the node with the minimum predicted load score as the running node of the container service.
[0067] Further, in another alternative embodiment, the load scheduler may combine a load prediction model and a dynamic scheduling algorithm (such as an improved Best-Fit Decreasing algorithm) to allocate the currently most suitable running node for the container service. For example, for a container service to be scheduled, if it is predicted that the CPU utilization of the node will increase by 30% when it runs in the next five minutes, and there are currently two available nodes: Node A and Node B. The load data of Node A and Node B is collected: the current CPU utilization of Node A is 50%, and the current CPU utilization of Node B is 70%; the load conditions of Node A and Node B are predicted: the CPU usage of Node A will increase to 70% in the next five minutes, and the CPU usage of Node B will increase to 80% in the next five minutes. Therefore, according to the current CPU utilization of Node A and Node B, both nodes can run the container service to be scheduled. However, according to the predicted load conditions of Node A and Node B, the CPU usage of Node A is 70%, with 30% remaining, which is sufficient for the container service to be scheduled to run, while the CPU usage of Node B is 80%, with 20% remaining, which is not sufficient for the container service to be scheduled to run. Therefore, it is selected to allocate the container service to be scheduled to run on Node A.
[0068] Further, after the container service is allocated in this embodiment, in each node, the node itself can set resource isolation and rate limiting policies to prevent the situation where all container services on a node become unavailable due to vicious resource competition among containers on the node. Specifically, this embodiment realizes resource isolation through CPU isolation. Optionally, CFS (Completely Fair Scheduler) quota + CPUset binding is adopted to set the burst limit. Among them, CPUset is a resource management mechanism that allows a group of processes to be bound to specific CPUs and memory nodes. The burst limit refers to the amount of data that is allowed to be sent or received exceeding the average rate limit within a specific time window in the traffic control policy. This embodiment uses I / O rate limiting and network bandwidth control to implement the rate limiting policy to prevent system instability caused by resource overload. And this embodiment can also improve the resource utilization rate through a memory dynamic overcommitment strategy while ensuring the safe operation of the container service. Among them, the memory dynamic overcommitment strategy is a technical means to improve resource utilization, which allows physical memory resources to be allocated to nodes or container services that exceed their actual capacity under certain conditions.
[0069] S11. During the operation of the container service, if the first node meets the secondary scheduling trigger condition and is not overloaded, the rescheduler determines a second node whose running load condition meets the low load condition from multiple other nodes, and at least one container service is running on the first node;
[0070] S12. The rescheduler determines a first container service from at least one container service running on the first node according to the type level of the container service, and schedules the first container service to the second node;
[0071] S13. If the first node is overloaded, the rescheduler shuts down a second container service running on the first node, and the type level of the second container service is lower than the type levels of other container services running on the first node.
[0072] Among them, there may be no strict order of processing between step S11 and step S12 and step S13. In one embodiment, step S13 may be executed after step S11 and step S12. At this time, after allocation and secondary scheduling, the first node has an overload problem, and the rescheduler shuts down the second container service running on the first node. In another embodiment, step S13 may be directly executed after step S10. At this time, after allocation, the first node has an overload problem, and the rescheduler shuts down the second container service running on the first node.
[0073] The rescheduler may be a device or module that performs secondary scheduling on container services during the running of container services. The rescheduler may periodically monitor the load data of each node to determine whether to trigger secondary scheduling, or the monitoring system of the container cluster may monitor the load data of the nodes to determine whether to trigger secondary scheduling. The conditions for triggering secondary scheduling may include: the CPU utilization rate of a certain node exceeds a preset threshold (such as 75% or 80%), or a certain node has a high CPU utilization rate for a continuous period of time (such as the CPU utilization rate of 80% lasts for one minute). Of course, secondary scheduling may also be triggered periodically by the rescheduler (such as every five minutes) or manually forced by the operation and maintenance personnel. In some special scenarios, according to the characteristics of service running in the scenario, the monitored node load data may be correspondingly increased. For example, in services that occupy network bandwidth such as file download services or video download services, in addition to monitoring the CPU utilization rate of the node, the network bandwidth of the node may be correspondingly increased.
[0074] In this embodiment, secondary scheduling is implemented through the rescheduler. When the container cluster is running, the container services are scheduled in real time according to the dynamic load changes of the nodes, reducing the possibility of local resource competition and bottleneck problems in the nodes, realizing automatic resource allocation, balancing the running resources of the nodes in the container cluster, reducing local resource bottlenecks, and ensuring the continuous and stable operation of the container cluster.
[0075] Of course, in another alternative embodiment, the load entropy value in the container cluster can be calculated to determine whether secondary scheduling is required. The load entropy value is a physical quantity used to measure the degree of uniformity of load distribution in a network or system. When the load entropy value reaches the maximum, it indicates that the load in the container cluster is completely balanced, while when the load entropy value reaches the minimum, it indicates that the load in the container cluster has a high degree of non-uniformity, with some nodes being overloaded and some nodes being underloaded, and thus secondary scheduling is required. The calculation formula for the load entropy value can be shown as follows:
[0076]
[0077] Among them, can represent the load entropy value of the container cluster; can represent the load ratio of node , and this load ratio is: the ratio of the load of node to the total load of the cluster (the ratio of the load of node / the total load of the cluster).
[0078] When the first node needs secondary scheduling, the rescheduler needs to determine a second node from multiple other nodes whose running load conditions meet the low-load condition, and schedule the container services on the first node to run on the second node. Among them, the low-load condition can be: the predicted load score of the node does not exceed the first preset threshold and the predicted load score is the smallest. The specific process for the rescheduler to select the second node from multiple other nodes can be shown as follows:
[0079] The rescheduler obtains the predicted load scores of each other node, screens out some other nodes whose predicted load scores do not exceed the first preset threshold, and selects the other node with the smallest predicted load score from the some other nodes as the second node. Of course, in another alternative embodiment, the rescheduler can directly select the second node according to the load data of each node. For example, among all other nodes, select the other node with the lowest CPU utilization rate as the second node.
[0080] After the rescheduler determines the second node, the rescheduler can first determine the first container service from at least one container service running on the first node according to the type level of the container service, and then schedule the first container service to the second node. When the first container service is scheduled to the second node, resource isolation can be automatically set for it according to the type level of the first container service.
[0081] Among them, the type level of the container service can be determined according to the type of the container service. The type level can represent the importance level and service quality level of the container service. Of course, different business scenarios can have different level divisions. Specifically, in this embodiment, the container service can be divided into three service quality (QoS) levels, namely: SLA-A (core service), SLA-B (important service), and SLA-C (batch processing service). Different weights are assigned according to the different service quality levels of the container service as the type level of the container service. The higher the service quality level, the greater the weight, and the higher the type level. Optionally, SLA-A corresponds to a resource-guaranteed container service with a weight of 3, SLA-B corresponds to an elastic-guaranteed container service with a weight of 2, and SLA-C corresponds to a best-effort container service with a weight of 1.
[0082] Since in this embodiment, the level of the type level is determined according to the service quality of the container service, and the container services with high type levels are generally core services and are generally not easily scheduled. Therefore, the container services with low type levels are generally scheduled to preferentially meet the resource requirements of the container services with high type levels. Therefore, the first container service for secondary scheduling is on the first node, and the type level of the first container service is lower than the type levels of other container services. Of course, when all the container services running on the first node are container services with high type levels, then at this time, a container service with a high type level can be selected to be scheduled to other nodes.
[0083] In order not to affect the operation of the scheduled first container service, in this embodiment, when scheduling the first container service, an elegant migration strategy is adopted to ensure the service continuity of the first container service. Specifically, in this embodiment, the CRIU checkpoint restoration technology can be used to schedule the first container service. CRIU (Checkpoint / Restore In Userspace) checkpoint restoration technology is a tool for capturing and restoring the process state. The CRIU checkpoint restoration technology allows freezing a running process (or process group) and saving its state to a set of files on the disk. The files can contain all the context information of the process, such as memory, file descriptors, network connections, etc. After that, these files can be used to restore the running state of the process on any host that supports CRIU, so as to achieve the migration or backup of the process.
[0084] In this embodiment, by setting the type level for the container service, it is convenient to preferentially meet important container services. And when resource competition occurs, the container services with low type levels are selected for scheduling to actively perform load balancing to ensure the continuity of the container services with low type levels. And different type levels can be set according to different business scenarios, which is convenient for scheduling container services to flexibly cope with complex business scenarios.
[0085] Further, in the container cluster of this embodiment, each node has its own monitoring component. During the operation of the container cluster, the nodes can monitor their own load data more frequently, and its monitoring period can be shorter than the monitoring period of the rescheduler or the monitoring period of the monitoring system of the container cluster. For example, the rescheduler monitors once every five minutes, and the nodes can monitor once every two minutes. The nodes' monitoring of their own load conditions can determine the load conditions by calculating the load score or the current values of multiple parameters in the load data. Among them, if the load score of the first node is not less than the second preset threshold, it is confirmed that the first node is overloaded, and the second container service with the lowest type level on the first node is selected to be shut down. The second preset threshold of the node is much larger than the first preset threshold of the rescheduler. If the load score of the first node is not less than the second preset threshold, it means that the first node has a vicious resource competition, which has seriously affected the operation of the container service and even the container service has stopped running. The node is at risk of running collapse at any time. At this time, the node itself shuts down the container service with a low type level to give priority to ensuring the operation of the container service with a high type level.
[0086] A container scheduling method for container cluster hybrid deployment provided by an embodiment of the present application. The container cluster hybrid deployment may include multiple nodes, a load scheduler, and a rescheduler. This method first predicts the load conditions of each node during the operation of the container cluster, and allocates nodes for at least one container service to run according to the predicted load conditions. And during the subsequent operation of the container service, if the first node meets the secondary scheduling trigger condition and is not overloaded, a second node with a load condition meeting the low-load condition can be selected from other nodes, and the first container service running on the first node is scheduled to the second node. If the first node is overloaded, the second container service with a low type level running on the first node is directly shut down. This method first allocates nodes for each container service according to the predicted load conditions of the nodes, ensuring that each container service can match a suitable node and preventing the problem of excessive load immediately after the node runs the container service. And when it is found that a node is in a high-load situation during the operation of the container service, by scheduling the container service to other nodes with lower load, the number of container services running on the overloaded node is reduced to alleviate the intensity of the competition for running resources among the container services. If the node is already overloaded, the second container service with a low type level is directly selected to be shut down to ensure the operation of other container services with a high type level. Therefore, this method can effectively improve the operation quality of the service.
[0087] The above introduced a container scheduling method for a container cluster hybrid deployment system provided by an embodiment of the present application. Next, a system applying the above container scheduling method of the container cluster hybrid deployment system will be introduced. This container cluster hybrid deployment system includes: a load scheduler, a rescheduler, and multiple nodes;
[0088] The load scheduler is used to predict the load conditions of multiple nodes respectively during the operation of the container cluster, and allocate nodes for at least one container service according to the load conditions;
[0089] The rescheduler is used to, during the operation of the container service, if the first node meets the secondary scheduling trigger condition and is not overloaded, determine a second node whose running load condition meets the low load condition from multiple other nodes, and at least one container service is running on the first node;
[0090] The rescheduler is also used to determine a first container service from at least one container service running on the first node according to the type level of the container service, and schedule the first container service to the second node;
[0091] If the first node is overloaded, the rescheduler is also used to shut down the second container service running on the first node, and the type level of the second container service is lower than the type levels of other container services running on the first node.
[0092] In a possible implementation, the load scheduler predicting the load conditions of multiple nodes respectively can be specifically configured as:
[0093] The load scheduler respectively obtains the historical load data of each node in the historical period and the current load data at the current moment. The load scheduler inputs the historical load data and the current load data into the load prediction model, obtains the predicted load score in the future period output by the load prediction model, and uses the predicted load score as the load condition of the node.
[0094] In a possible implementation, the historical load data includes the historical load score in the historical period, and the current load data includes the current load score at the current moment;
[0095] The load scheduler respectively obtaining the historical load data of each node in the historical period and the current load data at the current moment can be specifically configured as:
[0096] The load scheduler respectively collects the historical CPU utilization rate, historical memory usage, historical network bandwidth, and historical disk read and write times of each node in the historical period, and calculates the historical load score in the historical period according to the historical CPU utilization rate, historical memory usage, historical network bandwidth, and historical disk read and write times;
[0097] The load scheduler respectively collects the current CPU utilization rate, current memory usage, current network bandwidth, and current disk read and write times of each node at the current moment, and calculates the current load score at the current moment according to the current CPU utilization rate, current memory usage, current network bandwidth, and current disk read and write times.
[0098] In a possible implementation, the load condition of a node is the predicted load score of the node;
[0099] Based on the load condition, the load scheduler assigns nodes to at least one container service, and can be specifically configured as follows:
[0100] For each container service, when the container service needs to run, the load scheduler assigns the node with the smallest current predicted load score to the container service.
[0101] In a possible implementation, the low-load condition is that the predicted load score of the node does not exceed the first preset threshold and is the smallest;
[0102] The rescheduler determines a second node whose running load condition meets the low-load condition from multiple other nodes, and can be specifically configured as follows:
[0103] The rescheduler obtains the predicted load scores of each other node and screens out some other nodes whose predicted load scores do not exceed the first preset threshold;
[0104] Selects the other node with the smallest predicted load score from the some other nodes as the second node.
[0105] In a possible implementation, on the first node, the type level of the first container service is lower than that of other container services.
[0106] In a possible implementation, it further includes:
[0107] The rescheduler is further configured to confirm that the first node is overloaded when the load score of the first node is not less than the second preset threshold, and the second preset threshold is greater than the first preset threshold.
[0108] An embodiment of the present application also provides an electronic device. Refer to Figure 2 As shown, it shows a schematic structural diagram of an electronic device suitable for implementing the electronic device in the embodiment of the present application. The electronic device in the embodiment of the present application may include, but is not limited to, fixed terminals such as mobile phones, laptop computers, PDAs (Personal Digital Assistants), PADs (Tablet Computers), desktop computers, and the like. Figure 2 The electronic device shown is only an example and should not impose any limitations on the functions and usage scope of the embodiment of the present application.
[0109] As Figure 2As shown, the electronic device may include a processing device (such as a central processing unit, a graphics processing unit, etc.) 201, which may perform various appropriate actions and processes according to the program stored in the read-only memory (ROM) 202 or the program loaded from the storage device 208 into the random access memory (RAM) 203. When the electronic device is powered on, various programs and data required for the operation of the electronic device are also stored in the RAM 203. The processing device 201, the ROM 202, and the RAM 203 are connected to each other through a bus 204. The input / output (I / O) interface 205 is also connected to the bus 204.
[0110] Generally, the following devices may be connected to the I / O interface 205: an input device 206 including, for example, a touch screen, a touchpad, a keyboard, a mouse, a camera, a microphone, an accelerometer, a gyroscope, etc.; an output device 207 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, etc.; a storage device 208 including, for example, a memory card, a hard disk, etc.; and a communication device 209. The communication device 209 may allow the electronic device to communicate with other devices wirelessly or wiredly to exchange data. Although Figure 2 an electronic device with various devices is shown, it should be understood that it is not required to implement or have all the shown devices. Instead, more or fewer devices may be implemented or had.
[0111] In an embodiment of the present application, there is also provided a computer program product including computer-readable instructions. When the computer-readable instructions run on an electronic device, the electronic device implements the container scheduling method of any container cluster hybrid system provided in the embodiment of the present application.
[0112] In an embodiment of the present application, there is also provided a computer-readable storage medium. The storage medium carries one or more computer programs. When the one or more computer programs are executed by an electronic device, the electronic device can implement the container scheduling method of any container cluster hybrid system provided in the embodiment of the present application.
[0113] In addition, it should be noted that the system embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules may be selected according to actual needs to achieve the purpose of the solution of this embodiment. In addition, in the system embodiment drawings provided in the present application, the connection relationship between modules indicates that they have a communication connection, which can be specifically implemented as one or more communication buses or signal lines.
[0114] Through the description of the above embodiments, those skilled in the art can clearly understand that this application can be implemented by means of software plus necessary general hardware. Of course, it can also be implemented by dedicated hardware including application-specific integrated circuits, dedicated CPUs, dedicated memories, dedicated components, etc. Generally, functions accomplished by computer programs can be easily implemented by corresponding hardware, and the specific hardware structures for implementing the same function can also be diverse, such as analog circuits, digital circuits, or dedicated circuits, etc. However, for this application, in more cases, software program implementation is a better embodiment. Based on such an understanding, the technical solution of this application, in essence, or the part that contributes to the prior art can be embodied in the form of a software product. This computer software product is stored in a readable storage medium, such as a floppy disk, USB flash drive, mobile hard disk, ROM, RAM, magnetic disk, or optical disc of a computer, etc., and includes several instructions to enable a computer device (which can be a personal computer, training device, or network device, etc.) to execute the methods described in various embodiments of this application.
[0115] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented in whole or in part in the form of a computer program product.
[0116] The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the processes or functions described in the embodiments of this application are generated in whole or in part. The computer can be a general-purpose computer, a dedicated computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer instructions can be transmitted from a website, computer, training device, or data center to another website, computer, training device, or data center by wire (such as coaxial cable, optical fiber, digital subscriber line (DSL)) or wirelessly (such as infrared, wireless, microwave, etc.). The computer-readable storage medium can be any available medium that a computer can store, or a data storage device such as a training device or data center that includes one or more integrated available media. The available medium can be a magnetic medium (such as a floppy disk, hard disk, magnetic tape), an optical medium (such as a DVD), or a semiconductor medium (such as a solid-state drive (SSD)), etc.
[0117] Each embodiment in this specification is described in a related manner. For the same or similar parts among the embodiments, reference can be made to each other, and the key point of each embodiment is to illustrate the differences from other embodiments. In particular, for the apparatus embodiments, since they are basically similar to the method embodiments, the description is relatively simple, and reference can be made to the relevant parts of the method embodiments for the related content.
[0118] It can be understood that before using the technical solutions disclosed in the embodiments of the present disclosure, the types, usage scopes, usage scenarios, etc. of the personal information involved in the present disclosure should be informed to the users and the authorization of the users should be obtained in an appropriate manner in accordance with relevant laws and regulations.
[0119] The above are only the embodiments of the present application and are not used to limit the present application. For those skilled in the art, various changes and modifications can be made to the present application. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application shall be included in the protection scope of the present application.
Claims
1. A container scheduling method for a container cluster hybrid deployment system, characterized in that, The container cluster hybrid scheduling system includes a load scheduler, a rescheduler, and multiple nodes. The container scheduling method of the container cluster hybrid scheduling system includes: During the operation of the container cluster, the load scheduler predicts the load conditions of multiple nodes respectively, and allocates the nodes to at least one container service according to the load conditions; During the operation of the container service, if the first node meets the secondary scheduling trigger condition and is not overloaded, the rescheduler determines a second node whose running load condition meets the low load condition from multiple other nodes. At least one container service is running on the first node; The rescheduler determines a first container service from at least one container service running on the first node according to the type level of the container service, and schedules the first container service to the second node; If the first node is overloaded, the rescheduler shuts down the second container service running on the first node. The type level of the second container service is lower than the type levels of other container services running on the first node.
2. The container scheduling method of the container cluster hybrid deployment system according to claim 1, characterized in that The load scheduler predicts the load conditions of multiple nodes respectively, including: The load scheduler obtains the historical load data of each node in the historical period and the current load data at the current moment respectively; The load scheduler inputs the historical load data and the current load data into a load prediction model, obtains the predicted load score in the future period output by the load prediction model, and uses the predicted load score as the load condition of the node.
3. The container scheduling method of the container cluster hybrid deployment system according to claim 2, characterized in that, The historical load data includes the historical load score in the historical period, and the current load data includes the current load score at the current moment. The load scheduler obtains the historical load data of each node in the historical period and the current load data at the current moment respectively, including: The load scheduler collects the historical CPU utilization rate, historical memory usage, historical network bandwidth, and historical disk read / write times of each node in the historical period respectively, and calculates the historical load score in the historical period according to the historical CPU utilization rate, the historical memory usage, the historical network bandwidth, and the historical disk read / write times; The load scheduler collects the current CPU utilization rate, current memory usage, current network bandwidth, and current disk read / write times of each node at the current moment respectively, and calculates the current load score at the current moment according to the current CPU utilization rate, the current memory usage, the current network bandwidth, and the current disk read / write times.
4. The container scheduling method of the container cluster hybrid deployment system according to claim 1 or 2, characterized in that The load condition of the node is the predicted load score of the node. The allocating the nodes to at least one container service according to the load conditions includes: For each of the container services, when the container service needs to run, the load scheduler allocates the node with the smallest current predicted load score to the container service.
5. The container scheduling method of the container cluster hybrid deployment system according to claim 1, characterized in that The low load condition is that the predicted load score of the node does not exceed a first preset threshold and is the smallest. The rescheduler determines a second node whose running load condition meets the low load condition from multiple other nodes, including: The rescheduler obtains the predicted load scores of each of the other nodes, and filters out those other nodes whose predicted load scores do not exceed the first preset threshold; Selects the other node with the minimum predicted load score from the part of the other nodes as the second node.
6. The container scheduling method of the container cluster hybrid deployment system according to claim 1, wherein On the first node, the type level of the first container service is lower than the type levels of other container services.
7. The container scheduling method of the container cluster hybrid deployment system according to claim 1, wherein It further includes: If the load score of the first node is not less than the second preset threshold, it is confirmed that the first node is overloaded, and the second preset threshold is greater than the first preset threshold.
8. A container cluster hybrid deployment system, characterized in that, The container cluster hybrid system includes a load scheduler, a rescheduler, and multiple nodes, The load scheduler is used to predict the load conditions of multiple nodes respectively during the operation of the container cluster, and allocate the nodes for at least one container service according to the load conditions; The rescheduler is used to, during the operation of the container service, if the first node meets the secondary scheduling trigger condition and is not overloaded, determine a second node whose running load condition meets the low load condition from multiple other nodes, and at least one container service is running on the first node; The rescheduler is further used to determine a first container service from at least one container service running on the first node according to the type level of the container service, and schedule the first container service to the second node; If the first node is overloaded, the rescheduler is further used to shut down the second container service running on the first node, and the type level of the second container service is lower than the type levels of other container services running on the first node.
9. An electronic device, characterized in that, It includes at least one processor and a memory connected to the processor, wherein: The memory is used to store computer programs; The processor is used to execute the computer program so that the electronic device can implement the container scheduling method of the container cluster hybrid system as described in any one of claims 1 to 7.
10. A computer program product, characterized in that, It includes computer-readable instructions, which, when running on an electronic device, enable the electronic device to implement the container scheduling method of the container cluster hybrid system as described in any one of claims 1 to 7.
Citation Information
Patent Citations
Cloud computing cluster mixing job scheduling method and device, server and storage device
CN110908795A
Resource scheduling method and server system for off-line mixing operation
CN111026553A