Resource scheduling method, device, system and equipment and storage medium

By introducing virtual machines as the second type of work nodes in the container management platform, dynamically expanding the resource pool to run container groups that cannot be scheduled to the physical machine, solving the problem of stable and rapid expansion of containers in the container management platform, and achieving high reliability and efficient operation of container groups.

CN120104295APending Publication Date: 2025-06-06HANGZHOU ALICLOUD FEITIAN INFORMATION TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202311650617.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-12-04
Publication Date
2025-06-06

AI Technical Summary

Technical Problem

The existing container management platform fails to effectively ensure the stable operation of containers, especially when the physical machine resources are insufficient, it is difficult to achieve rapid expansion of container groups and high reliability operation.

Method used

By introducing virtual machines as the second type of work node in the container management platform, the resource pool is dynamically expanded to run container groups that are not scheduled to the physical machine, and the rapid expansion of container groups and high reliability operation are achieved.

Benefits of technology

The operation reliability of container groups and the success rate of container scheduling are improved. The problem of insufficient physical machine resources is solved through rapid expansion of virtual machines, and the stable and efficient operation of container groups is achieved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120104295A_ABST
    Figure CN120104295A_ABST
Patent Text Reader

Abstract

The invention provides a resource scheduling method, device, system and equipment and a storage medium, a container management platform is used for managing a working node cluster, the working node cluster comprises at least one first-class working node, the first-class working node is a physical machine, and the physical machine is used for operating at least one container group; the method comprises the steps of receiving an operation request of at least one container group; determining whether the at least one container group can be successfully scheduled to a first type of working nodes in a working node cluster for operation or not; and in response to determining that at least one target container group which cannot be scheduled to run in the first type of working nodes exists, obtaining a virtual machine corresponding to each target container group, and adding the obtained virtual machines to the working node cluster as second type of working nodes.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of cloud computing technology, and in particular to a resource scheduling method, apparatus, system, device and storage medium. Background Art

[0002] Containerized applications refer to packaging applications and all their dependencies into an independent, standardized runtime environment to achieve rapid deployment, portability, and cross-platform nature of applications. At present, there are some container management platforms, such as Kubernetes (K8s for short, the name of a container orchestration system) or docker-swarm (container cluster, the name of a container management system). Taking Kubernetes as an example, the Master in Kubernetes can manage the working nodes in the Node cluster. Each Node is an independent hardware resource for the Master. The underlying Node may be a physical machine or a virtual machine. One or more Pods (container groups) can run on the Node. Pod is the smallest operating unit managed by Kubernetes. Each Pod can be understood as a collection of containers, that is, a Pod is composed of one or more containers. A container contains the application and all its dependencies and provides an independent, standardized runtime environment. Based on this, how to ensure the stable operation of containers is a technical problem that needs to be solved urgently. Summary of the invention

[0003] To overcome the problems existing in the related art, the present disclosure provides a resource scheduling method, apparatus, computer equipment and storage medium.

[0004] According to a first aspect of an embodiment of this specification, a resource scheduling method based on a container management platform is provided, wherein the container management platform is used to manage a work node cluster, wherein the work node cluster includes at least one first-class work node, wherein the first-class work node is a physical machine, and the physical machine is used to run at least one container group; the method includes:

[0005] Receive a running request of at least one container group;

[0006] Determine whether the at least one container group can be successfully scheduled to run in a first type of working node in the working node cluster;

[0007] In response to determining that there is at least one target container group that cannot be scheduled to run in the first type of worker node, a virtual machine corresponding to each of the target container groups is obtained, and the obtained virtual machine is added to the worker node cluster as a second type of worker node, where the second type of worker node is used to run the corresponding target container group.

[0008] According to a second aspect of an embodiment of this specification, a resource scheduling device based on a container management platform is provided, wherein the container management platform is used to manage a work node cluster, wherein the work node cluster includes at least one first-class work node, wherein the first-class work node is a physical machine, and the physical machine is used to run at least one container group; the device includes:

[0009] A receiving module, configured to receive a running request of at least one container group;

[0010] A determination module, used to determine whether the at least one container group can be successfully scheduled to run in a first type of working node in the working node cluster;

[0011] A scheduling module, configured to, in response to determining that there is at least one target container group that cannot be scheduled to run in the first-type working node, obtain a virtual machine corresponding to each of the target container groups, and add the obtained virtual machines to the working node cluster as a second-type working node, where the second-type working node is used to run the corresponding target container group.

[0012] According to a third aspect of an embodiment of this specification, a container management system is provided, the system comprising a scheduling node and a working node cluster, the working node cluster comprising at least one first-class working node, the first-class working node being a physical machine, and the physical machine being used to run at least one container group; the scheduling node is used to execute the steps of the method described in the first aspect.

[0013] According to a fourth aspect of the embodiments of this specification, a computer-readable storage medium is provided, on which a computer program is stored, and when the computer program is executed by a processor, the steps of the method embodiment described in the first aspect are implemented.

[0014] According to a fifth aspect of the embodiments of this specification, a computer device is provided, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the computer program, the steps of the method embodiment described in the first aspect are implemented.

[0015] The technical solutions provided by the embodiments of this specification may have the following beneficial effects:

[0016] In the embodiment of this specification, the working node cluster in the container management platform includes at least one first-class working node, and the first-class working node adopts a physical machine, which is used to run at least one container group; when receiving a running request containing at least one container group, it is first determined whether the container group can be successfully scheduled to the first-class working node in the working node cluster for running; in the case of insufficient physical machine resources, that is, when it is determined that there is at least one target container group that cannot be scheduled to run in the first-class working node, a virtual machine corresponding to each of the target container groups is obtained, and the obtained virtual machine is added to the working node cluster as a second-class working node, and the second-class working node is used to run the corresponding target container group. On the one hand, this embodiment adopts a physical machine as the first-class working node. Compared with a virtual machine, a physical machine as the first-class working node is preferentially used to schedule containers, and its stability is higher, which can improve the reliability of the operation of the container group; on the other hand, when the resources of the first-class working node are insufficient, considering the large-scale characteristics of the physical machine itself, its expansion speed to the cluster is slow; this embodiment adopts a virtual machine to achieve rapid expansion, and the virtual machine is only used to run the container group that cannot be currently scheduled to the first-class working node, so the reliability of the operation of the container group is also guaranteed.

[0017] It is to be understood that the foregoing general description and the following detailed description are exemplary and explanatory only and are not restrictive of the present disclosure. BRIEF DESCRIPTION OF THE DRAWINGS

[0018] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with the specification and, together with the description, serve to explain the principles of the disclosure.

[0019] Figure 1 This is a schematic diagram of the architecture of Kubernetes shown in this specification according to an exemplary embodiment.

[0020] Figure 2A It is a schematic diagram of a working node cluster shown in this specification according to an exemplary embodiment.

[0021] Figure 2B This is a flowchart of a resource scheduling method based on a container management platform according to an exemplary embodiment of this specification.

[0022] Figure 2C This is a schematic diagram of resource scheduling according to an exemplary embodiment of the present specification.

[0023] Figure 3 It is a hardware structure diagram of a computer device where a resource scheduling device based on a container management platform is located according to an exemplary embodiment of this specification.

[0024] Figure 4 This is a block diagram of a resource scheduling device based on a container management platform according to an exemplary embodiment of the present specification. DETAILED DESCRIPTION

[0025] Exemplary embodiments will be described in detail herein, examples of which are shown in the accompanying drawings. When the following description refers to the drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The implementations described in the following exemplary embodiments do not represent all implementations consistent with this specification. Instead, they are merely examples of devices and methods consistent with some aspects of this specification as detailed in the appended claims.

[0026] The terms used in this specification are for the purpose of describing specific embodiments only and are not intended to limit this specification. The singular forms "a", "the" and "the" used in this specification and the appended claims are also intended to include plural forms unless the context clearly indicates otherwise. It should also be understood that the term "and / or" used herein refers to and includes any or all possible combinations of one or more associated listed items.

[0027] It should be understood that although the terms first, second, third, etc. may be used in this specification to describe various information, this information should not be limited to these terms. These terms are only used to distinguish the same type of information from each other. For example, without departing from the scope of this specification, the first information may also be referred to as the second information, and similarly, the second information may also be referred to as the first information. Depending on the context, the word "if" as used herein may be interpreted as "at the time of" or "when" or "in response to determining".

[0028] The user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in this disclosure are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of relevant data must comply with the relevant laws, regulations and standards of relevant countries and regions, and provide corresponding operation entrances for users to choose to authorize or refuse.

[0029] The container management platform in the scheme of this embodiment includes but is not limited to systems that can perform cluster management, such as Kubernetes, Docker Swarm or Apache Mesos (a cluster management system). Among them, the container group in this embodiment can be called Pod in Kubernetes, and can be called container in other platforms such as Docker Swarm or Apache Mesos. The specific process of the resource scheduling method embodiment of this embodiment that can be applied to these different cluster management systems is relatively similar. For the sake of ease of description, the following embodiments are explained using Kubernetes as an example. It should be emphasized that although the Kubernetes architecture is used as the basis for explanation, those skilled in the art should understand that the scheme of this embodiment can be transplanted to other similar architectures such as Docker Swarm or Apache Mesos, and this embodiment does not limit this.

[0030] like Figure 1 FIG. 1 is a schematic diagram of the architecture of Kubernetes according to an exemplary embodiment of the present specification. In terms of cluster management, Kubernetes can divide the machines in the cluster into one or more master nodes and some worker nodes. Figure 1 For the sake of convenience in the example, only one Master and three Nodes (Node1, Node2, and Node3) are shown.

[0031] (1)Master

[0032] The Master in Kubernetes refers to the main node used for cluster control. In each Kubernetes cluster, there must be a Master to be responsible for the management and control of the entire cluster. Basically, all control commands of Kubernetes are sent to it, and it is responsible for the specific execution process. The Master usually occupies one or more independent servers.

[0033] A set of processes related to cluster management run on the Master. These processes implement management functions such as resource management, Pod scheduling, elastic scaling, security control, system monitoring and error correction for the entire cluster, and are all completed automatically.

[0034] (2)Node

[0035] Node is a working node in the cluster. This node can be a physical machine or a virtual machine in a private cloud or public cloud. Node can be dynamically added to the Kubernetes cluster during operation. The smallest operating unit managed by Kubernetes on the Node is Pod. A Node can run one or more Pods. For example, Figure 1 Node1 and Node2 shown in the figure run two Pods respectively, and Node3 runs one Pod. Kubernetes' kubelet and kube-proxy service processes run on the Nodes, which are responsible for creating, starting, monitoring, restarting, and destroying Pods, as well as implementing software-based load balancers.

[0036] (3)Pod

[0037] A Pod can be composed of one or more containers. For example, when at least two container applications are tightly coupled and combined into a whole to provide external services, the at least two containers can be packaged into a Pod.

[0038] As an example, suppose there is a computer program that requires a specific runtime environment, specific libraries, and configuration files to run properly. This computer program and all its dependencies can be packaged into a container, which can be packaged into a Pod, which can be scheduled to run in a Node. The Pod provides an independent runtime environment for the computer program, allowing the computer program to run in this environment without being affected by the external environment.

[0039] In other examples, a Pod can contain at least two containers; for example, when there are at least two containerized applications that are tightly coupled and combined into a whole to provide external services, the at least two containers can be packaged into a Pod. This approach is usually used for two closely related services, such as a web service program and its log collection program, or an application and its auxiliary tasks, etc. By packaging them in the same Pod, they can work better together. For example, Figure 1 Pod1 in Node1 consists of two containers.

[0040] Cloud services can be provided to users based on the above container management platform. For example, when a user needs to run an instance in the cloud, one or more Pods can be created in the container management platform and scheduled to run in the Node. The Pod contains the container corresponding to the user's instance.

[0041] The existing container management platform does not consider the specific settings of the working nodes, and fails to ensure the stable operation of the container group. Based on this, the present description embodiment provides a resource scheduling solution that can improve the reliability of resource scheduling.

[0042] like Figure 2A FIG. 1 is a schematic diagram of a working node cluster according to an exemplary embodiment of the present specification. The working node cluster managed by the container management platform in this embodiment can be understood as including the following two types of resource pools:

[0043] (1) A first resource pool, which may include at least one physical machine; each of the physical machines may be configured as a first-class working node Node in the container management platform, which is referred to as a first-class working node for the sake of distinction; the first-class working node is used to run at least one container group Pod. Figure 2A Four Nodes (Node11 to Node14) are shown.

[0044] As an example, the physical machine here can be a large-scale computer device, such as a bare metal physical machine, etc. The amount of hardware resources in a physical machine is relatively large, for example, it can contain 80 CPUs (Central Processing Units, processors) or 104 CPUs, etc. Based on the characteristics of the large amount of hardware resources of the physical machine, adding a new physical machine to the first resource pool usually requires a long preparation time when resources are insufficient; for example, when a physical machine needs to be added to the cluster, it may be necessary to manually allocate a new machine and configure it in the computer room, and it may also be necessary to apply for approval and other processes; it is also possible to allocate machines currently used for other services, but it is necessary to clean up the machine currently in use, such as moving the programs running in it to other machines, etc. Based on this, the first resource pool of this embodiment can also be understood as a fixed resource pool, and the fixed here can mean that the number of physical machines in the resource pool will not change at a high frequency, and the number is relatively stable within a certain period of time. Optionally, in actual implementation, the service party can pre-configure the first resource pool and use it as a working node cluster in the container management platform. For the first resource pool, since each physical machine serves as a working node, the capacity of the hardware resources in the first resource pool can be determined, and the usage of the hardware resources in the first resource pool can also be determined. The service provider can maximize the operation of the hardware resources therein, such as resource scheduling or breaking up and relocating Pods according to different users' resource preferences for Pods, etc.

[0045] (2) A second resource pool, wherein the second resource pool includes at least one virtual machine, which can be configured as a second type of working node in the container management platform, and the second type of working node can be used to run the container group corresponding to it. Figure 2AThe two Nodes (Node21 and Node22) shown in FIG.

[0046] Because of the above advantages of the fixed pool, this embodiment is designed to give priority to the use of resources in the first resource pool. When the first resource pool cannot meet the resource requirements, the second resource pool will deliver these resources. That is, the second resource pool of this embodiment is used to quickly supplement resources when the resources in the first resource pool are insufficient, because its specifications are small and expansion is convenient and fast. Virtual machines can be obtained from other computer devices and added to the working node cluster as working nodes, which are referred to as the second type of working nodes in this embodiment. Taking into account the characteristics of virtual machines, this embodiment designs each virtual machine to run a corresponding container group to improve the operating reliability of the container group. In this way, since virtual machines can be quickly obtained at any time, dynamic and rapid expansion of resources can be achieved through virtual machines.

[0047] like Figure 2B FIG. 1 is a flowchart of a resource scheduling method according to an exemplary embodiment of the present specification, comprising the following steps:

[0048] Step 202: Receive a running request of at least one container group.

[0049] Step 204: Determine whether the at least one container group can be successfully scheduled to run in the first type of working node in the working node cluster.

[0050] Step 206: In response to determining that there is at least one target container group that cannot be scheduled to run in the first type of working node, obtain a virtual machine corresponding to each of the target container groups, and add the obtained virtual machines to the working node cluster as a second type of working node, where the second type of working node is used to run the corresponding target container group.

[0051] The resource scheduling method of this embodiment can be applied to any computer device, including but not limited to a single server, a server group consisting of multiple servers, or a cloud consisting of a large number of hosts or servers based on cloud computing, etc. Optionally, when executing the resource scheduling method of this embodiment, the specific implementation can be executed by the master node of the container management platform to improve the existing scheduling function of the master node of the container management platform; in other examples, it can also be executed by a computer program connected to the master node of the container management platform to instruct the master node of the container management platform to perform resource scheduling based on the solution of this embodiment.

[0052] The operation request for at least one container group in step 202 may be received in a variety of different ways depending on the actual application scenario. For example, a user-oriented service may receive an instance creation request from a user, and the service may determine the number of Pods required to run the user based on the instance creation request. In some scenarios, the number of copies of each Pod may also be determined. Based on the determined one or more Pods and the number of copies of the Pod, an operation request for a set of container groups is generated and sent to the computer program for resource scheduling running this embodiment. In other scenarios, the computer program for resource scheduling running this embodiment may directly receive an operation request for at least one container group of the user, or receive an operation request for at least one container group sent by other objects, etc. This embodiment does not limit this.

[0053] Among them, for step 204, by carrying the resource requirements of the container group in the running request and based on the resource running status of each first-class working node in the first resource pool, it can be determined whether there are sufficient resources so that the container group in the container group set can be successfully scheduled to the first-class working node in the working node cluster for running.

[0054] Among them, for step 206, there can be many ways to implement obtaining the virtual machine, which can be obtained from a server that provides virtual machine services. For example, the solution in this embodiment can send a virtual machine acquisition request to the server that provides virtual machine services to obtain the virtual machine provided by the server; or you can build the virtual machine service yourself and schedule the required virtual machine when needed.

[0055] Optionally, the run request may include one or more container groups. In some examples, for a distributed multi-copy scenario, at least two container groups may be included. Existing cloud service solutions based on container management platforms fail to implement the ability to support distributed multi-copy. Distributed multi-copy means that the same data needs to be stored in multiple different storage nodes to achieve the purpose of data protection. In the related art, a request can be made to the Master of Kubernetes to create a Pod containing multiple copies. The Master will create multiple Pods according to the number of copies, and each Pod corresponds to a storage node for storing data, which is used to store a copy of the data. For the multiple Pods created, the Master will try to schedule multiple Pods to different Nodes based on the working conditions of the Nodes in the cluster, but it is not necessarily guaranteed. Therefore, the distributed multi-copy based on Kubernetes actually has certain risks. For example, due to the high load of the Node cluster, multiple Pods may be scheduled to the same Node by the Master. If the Node fails, all data may be lost; it may also be the case that although multiple Pods are distributed in multiple different Nodes, these multiple different Nodes are multiple virtual machines running on the same physical machine.

[0056] To improve reliability, the container group and its replicas are physically dispersed. Physical dispersion here means that the container group to be run and its replica container groups are not run on the same physical machine.

[0057] Based on this, the receiving a running request of at least one container group includes: receiving a running request of a container group set, wherein the container group set includes: at least two container groups; and determining whether the at least one container group can be successfully scheduled to run in the first type of working node in the working node cluster may include:

[0058] Determine whether the container groups in the container group set can be successfully scheduled to run in different first-type working nodes in the working node cluster;

[0059] In response to determining that there is at least one target container group in the container group set that cannot be scheduled to run in the first type of working node, obtaining a virtual machine corresponding to each of the target container groups includes:

[0060] In response to determining that there are at least two target container groups in the container group set that cannot be scheduled to run in the first-type working node, a virtual machine corresponding to each of the target container groups and running on a different computer device is obtained.

[0061] This embodiment does not limit the number of copies of the Pod, which can be two or more, and can be configured as needed in actual applications. Among them, the container group set contains n Pods, and these n Pods constitute multiple replica Pods, and the configurations of these n Pods can be the same. The master node of the container management platform can complete the creation, scheduling and automatic control tasks of a group of Pod copies throughout their life cycle. The master node usually schedules the Pod and its copies to available working nodes based on the load of each working node. For example, the container group set contains Pod_A, Pod_B and Pod_C, and these three Pods constitute a multiple replica Pod with three copies.

[0062] In this embodiment, based on the design of the aforementioned resource pool, during scheduling, physical scattered deployment can be further realized. First, it is determined whether each container group in the container group set can be successfully scheduled to run in different first-class working nodes in the working node cluster; if all first-class working nodes fail to realize the physical scattered deployment of all container groups, virtual machines will be expanded to schedule the remaining container groups.

[0063] The above embodiment may include multiple situations: all Pods may be physically dispersed and deployed on different first-class worker nodes; some Pods may be physically dispersed and deployed on different first-class worker nodes, while the remaining k Pods cannot be deployed on different first-class worker nodes; or all Pods cannot be physically dispersed and deployed on different first-class worker nodes.

[0064] Furthermore, according to the number of container groups that cannot be successfully scheduled on the first type of working node (referred to as target container groups in this embodiment for the sake of distinction), a corresponding number of virtual machines are obtained. Assuming there is only one target container group, one virtual machine is obtained; assuming there are k (k is an integer greater than 1) target container groups, k virtual machines are obtained, and these k virtual machines run on different computer devices, that is, these k virtual machines also support physical dispersion.

[0065] Based on this, through the above embodiments, the physically dispersed deployment of multiple Pods can be guaranteed, and the high availability requirements of the service can be guaranteed to the greatest extent.

[0066] For example, Figure 2A The four Pod collections shown in the figure each contain three Pods:

[0067] Set 1: SET1-A, SET1-B, and SET1-C; Set 2: SET2-A, SET2-B, and SET2-C; Set 3: SET3-A, SET3-B, and SET3-C; Set 4: SET4-A, SET4-B, and SET4-C. The three Pods in each set can run on different machines in pairs.

[0068] For example, the three Pods in Set 1 run on Node 11, Node 12, and Node 13. The same is true for the three Pods in Set 2 and the three Pods in Set 3.

[0069] Among the three Pods in Set 4, one Pod (SET4-A) runs on Node14, while the other two Pods (SET4-B and SET4-C) run on two virtual machines (Node21 and Node22), which run on different computer devices. Here, multiple copies in a set are allowed to be deployed on different media and can cross resource pools as an example. In practical applications, if multiple copies in a set are not allowed to cross resource pools, according to needs, when a container group in a set cannot be scheduled to the first resource pool, even if there are first-class working nodes in the first resource pool that can schedule some container groups, multiple second-class working nodes are introduced in the second resource pool, and all container groups in the set are scheduled to the second-class working nodes.

[0070] In some examples, the container management platform includes a master node for managing the worker node cluster, and determining whether the container group in the container group set can be successfully scheduled to run in different first-class worker nodes in the worker node cluster may include:

[0071] Sending a resource reservation request corresponding to the container group set to the master node, so that the master node attempts to allocate one or more reserved resources corresponding to the container group set and located at different first-category working nodes from the working node cluster based on the resource reservation request;

[0072] Obtain a resource reservation result of the master node, and determine, based on the resource reservation result, whether each container group in the container group set can be successfully scheduled to run in a first type of working node in the working node cluster.

[0073] In order to improve the success rate of resource scheduling, the master nodes of some container management platforms include a scheduling function for reserved resources. This embodiment can send a resource reservation request to the master node so that after the master node receives the resource reservation request, it reserves corresponding reserved resources for each container group in each first-class working node.

[0074] Still taking Kubernetes as an example, the master node can be configured with a scheduling function for implementing node reservation resources. The resource reservation request received by the master node can be a request containing one or more reservation objects. This function can forge a virtual Pod for each Reservation (reserved resource) object. The master node can schedule the virtual Pod to find a suitable working node. Once a suitable working node is found, the corresponding resources of the working node will be occupied. The occupied resources are reserved resources, and each reserved resource can be used to run the corresponding container group. When creating a Reservation, you can specify which Pods will use the reserved resources in the future, and you can specify a specific Pod, or Pods with certain labels. When these Pods are scheduled by the scheduler that implements reserved resources in the master node, the scheduler will find the Reservation object that can be used by the Pod, and can give priority to the use of the Reservation resources. In addition, you can record which Pod can use the Reservation in the Reservation object, or the annotation attribute of the Pod can record which Reservation it uses.

[0075] In some examples, in response to determining that there is at least one target container group in the container group set that has not been successfully scheduled to run in the first type of working node, obtaining a virtual machine corresponding to each of the target container groups may include:

[0076] In response to the number of successfully allocated reserved resources included in the resource reservation result being lower than the number of container groups in the container group set, the number of target container groups in the container group set that have not been successfully scheduled to run in the first type of working nodes is determined, and virtual machines corresponding to the number of target container groups are obtained.

[0077] The obtained virtual machine is added to the working node cluster as a second working node to serve as a reserved resource corresponding to the container group set.

[0078] In this embodiment, the resource reservation result may include the number of successfully allocated reserved resources. If the number of reserved resources is equal to the number of container groups in the container group set, it means that each container group in the container group set can be successfully scheduled to run in the first resource pool without expanding resources. If the number of reserved resources is lower than the number of container groups in the container group set, it is determined that resources need to be expanded. Based on the difference between the number of reserved resources and the number of container groups in the container group set, the number of target container groups in the container group set that have not been successfully scheduled to run in the first type of working nodes can be determined, and virtual machines corresponding to the number of target container groups can be obtained. The obtained virtual machines are then added to the working node cluster as the second working node to serve as the reserved resources corresponding to the container group set. In this way, this embodiment realizes automatic and rapid resource expansion in the process of reserving resources.

[0079] In some examples, after the step of adding the acquired virtual machine to the worker node cluster as a second type of worker node, the method may further include:

[0080] Send a scheduling request corresponding to the container group set to the master node, so that the master node schedules each container group in the container group set to run in the reserved resources corresponding to the container group set.

[0081] In this embodiment, after the reserved resources of the container group set are allocated, a scheduling request corresponding to the container group set may be sent to the master node, so that the master node creates each container group and schedules it to run in the corresponding reserved resources.

[0082] If the reserved resources cannot be successfully allocated to the container group set, it means that the resource scheduling service may have problems such as failures, and troubleshooting is required. In this embodiment, after determining that the reserved resources are successfully allocated, the scheduling request corresponding to the container group set is sent to the master node, which can ensure the stability of the service.

[0083] Most concepts in Kubernetes, such as Node or Pod, can be considered as a resource object. All resource objects in Kubernetes can be defined or described using files in YAML or JSON format. For example, you can submit a definition file of a worker node to the master node in Kubernetes so that the master node can obtain the information of the worker node defined in the file after receiving the definition file of the worker node. You can submit a definition file of a Pod to the master node in Kubernetes so that the master node can create a Pod and schedule it to run on a certain worker node after receiving the definition file of the Pod.

[0084] The container group definition file can be used to define one or more attributes of a Pod, such as name, label, number of replicas, CPU limit, memory limit or storage volume, etc. In actual applications, configuration can be performed as needed, and this embodiment does not limit this. If a container group set includes multiple replica Pods, only one container group definition file can be used to define it, and the number of replicas and the attribute configuration of the Pod can be specified in the file.

[0085] Among them, the working node definition file can be used to define one or more attributes of the Node, such as labels, address information, resource information, stains or comments, etc., which can be configured as needed in actual applications and are not limited in this embodiment.

[0086] Some of these properties are described below:

[0087] (1) Label: A label can be a key-value pair, where the key and value can be specified by the user. Labels can be attached to various resource objects, such as Node or Pod. A resource object can define any number of labels, and the same label can be added to any number of resource objects. Labels are usually determined when a resource object is defined, and can also be dynamically added or deleted after the object is created.

[0088] (2) podAntiAffinity (the mutual exclusion attribute between Pods). This attribute is one of the attributes of Pod. It is used to represent the mutual exclusion relationship between Pods at the Node level. It allows the Master to avoid placing the Pod on the node where other Pods with specific labels are located when scheduling the Pod to the working node.

[0089] Based on this, in this embodiment, one or more of the above attributes can be used to achieve physical separation between container groups in the container group set, and schedule each container group on a suitable working node. Specifically, in this embodiment, a set identifier can be generated for the container group set, and each container group in the container group set has the same set identifier; different container group sets have different set identifiers.

[0090] In some examples, sending the resource reservation request corresponding to the container group set to the master node may include:

[0091] The set identifier of the container group set is obtained, and a reserved resource object definition file corresponding to the container group set is generated and sent to the master node; wherein the reserved resource object definition file is used to define each reserved resource object having the same number as the container groups in the container group set, and the configuration of the label attribute and the container group mutual exclusivity attribute of each of the reserved resource objects includes the set identifier, so as to indicate to the master node that the reserved resources of each reserved resource object are mutually exclusive on the same working node, so that the master node allocates the reserved resources located at different first-class working nodes.

[0092] As described in the above-mentioned embodiment, it can be implemented based on the reserved resource function of the master node. Each reserved resource object can be a virtual Pod, and the reserved resource object definition file can be a Pod definition file of the virtual Pod. The reserved resource object definition file sent to the master node in this embodiment is similar to the Pod definition file.

[0093] Among them, the label attribute and podAntiAffinity attribute of the virtual Pod corresponding to each reserved resource object include the collection identifier. Based on this, for the Master, since the label attributes and podAntiAffinity attributes of these multiple virtual Pods are configured as the same collection identifier, it means that the resources reserved by these three virtual Pods are mutually exclusive at the Node level and will not be scheduled on the same Node. In this way, mutual exclusion between any two Pods in the container group set on the Node is also achieved.

[0094] Since the first type of worker node in the first resource pool is a physical machine, and different first type of worker nodes are different physical machines, physical dispersion can be achieved in the first resource pool. If there are at least two Pods in the container group that need to be scheduled to different second type of worker nodes in the second resource pool, as virtual machines of different second type of worker nodes run on different computer devices, physical dispersion can also be achieved in the second resource pool.

[0095] Among them, the configuration of the attributes described above, in addition to the configuration made to implement the solution of this embodiment, in actual applications, as needed, the configuration of the attributes may further include other configurations to implement other scheduling strategies, which is not limited in this embodiment.

[0096] In a Kubernetes cluster, when a new Node is added, you can install related services on the new Node, then configure the startup parameters, configure the address of the current Kubernetes cluster Master, and finally start these services. Kubernetes provides a default automatic registration mechanism, and the new Node will automatically join the existing Kubernetes working node cluster.

[0097] In some examples, adding the acquired virtual machine to the working node cluster as a second working node to serve as a reserved resource corresponding to the container group set may include:

[0098] Obtain the set identifier of the container group set, generate a working node definition file corresponding to the virtual machine and send it to the master node; wherein, in the working node definition file, the configuration of the label attribute of the working node corresponding to the virtual machine includes the set identifier, so that the master node configures the virtual machine as a second type of working node in the working cluster and as a reserved resource corresponding to the container group set.

[0099] In this embodiment, in order to enable the expanded virtual machine to be used to run the corresponding container group, the configuration of the label attribute of the working node corresponding to the virtual machine includes a set identifier, so that the working node can correspond to the container group set, and the subsequent master node can find the working node when scheduling the container group in the container group set.

[0100] In addition, the configuration of the taint attribute of the working node also includes the collection identifier, which can be combined with the tolerance attribute of the Pod to ensure that the working nodes composed of the expanded virtual machines can be scheduled to run the corresponding container group.

[0101] In some examples, the configuration of the owner attribute of each of the reserved resource objects in the reserved resource object definition file may include the set identifier; sending the scheduling request corresponding to the container group set to the master node so that the master node schedules each container group in the container group set to the reserved resource corresponding to the container group set for operation may include:

[0102] A container group definition file corresponding to the container group set is generated and sent to the master node; wherein the container group definition file includes the configuration of the label attribute of each container group in the container group set including the set identifier, so that the master node schedules each container group in the container group set to the reserved resources corresponding to the container group set for operation based on the configuration of the owner attribute of the reserved resource object and the configuration of the label attribute of each container group in the container group set.

[0103] In this embodiment, the owner attribute owners of Reservation can be configured to specify its owners as a collection identifier, so that the Reservation is specified as a container group with a label having a collection identifier, thereby realizing the association between Reservation and Pod. As an example, the owners attribute can include a sub-attribute label selector labelSelector, which can specify the matching label as a collection identifier, based on which the master node can select a container group with a label having a collection identifier during scheduling.

[0104] like Figure 2C Shown is a schematic diagram of a resource scheduling according to an exemplary embodiment of the present specification; the figure shows three nodes (Node01 to Node03) corresponding to three physical machines in the first resource pool, and three nodes (Node04 to Node06) corresponding to three virtual machines in the second resource pool.

[0105] The diagram also shows three container group sets:

[0106] Container group set 1 contains three Pods: Pod Set 1-01, Pod Set 1-02, and Pod Set 1-03;

[0107] Container group set 2 contains three Pods: Pod Set2-01, Pod Set2-02, and Pod Set2-03;

[0108] Container group set 3 contains three Pods: Pod Set3-01, Pod Set3-02, and Pod Set3-03.

[0109] The scheduling requirements of these three container group sets can be transmitted by the upper-layer service. For example, an upper-layer service receives a configuration request for a user's database instance. The service can determine the number of containers to be created based on the resource size of the database instance configured by the user; for example, it is determined that three containers need to be created, and in order to protect the data, each container needs a copy. Taking the number of copies as 2 as an example, three container group sets are required, and each container group set contains 3 Pods to implement the multi-copy function. The input parameters of the method of this embodiment can be the information of the three container group sets, and then the method of this embodiment can be used to schedule the three container group sets. In this embodiment, the scheduling process for the three container group sets is the same, and in actual applications, they can be executed separately (either in parallel or in serial). The resource scheduling embodiment of this embodiment implements the physical dispersion scheduling of the distributed multiple copies of the three container group sets.

[0110] Among them, the set identifiers of the three container group sets can be obtained; taking container group set 1 as an example, the set identifier SetID of the container group set is 1, that is, the three Pods have the same set identifier.

[0111] like Figure 2C As shown, this embodiment uses reserved resources to implement the scheduling of container group sets. Therefore, you can first request three reserved resources of the container group set from the Master. The implementation process of requesting reserved resources for the container group set is consistent with the implementation process of scheduling the container group set. Submit the reserved resource object (virtual Pod) definition file corresponding to the container group set to the Master. After successful reservation, send the container group definition file corresponding to the container group set; the content in the reserved resource object definition file can be the same as the container group definition file. For example, the virtual Pod defined in the reserved resource object definition file submitted to the Master of the container management platform, and the Pod defined in the container group definition file, can contain the following information:

[0112] The Replicas attribute indicates the number of replicas of the Pod, which is configured to 3;

[0113] The configurations of the three Pods are the same. The configuration of Pod attributes can include:

[0114] The label attribute of the Pod is configured as: SetID = 1;

[0115] Indicates that the Pod AntiAffinity attribute of the Pod is configured as SetID = 1;

[0116] Indicates the affinity attribute nodeAffinity between the Pod and the Node, which is configured as SetID = 1;

[0117] Indicates that the tolerance attribute of the Pod to the Node is configured as SetID=1.

[0118] In the reserved resource object definition file, the sub-attribute label selector labelSelector under the owners attribute of Reservation can specify the matching label as the set identifier, that is, matchLabels:SetID:"1".

[0119] Among them, NodeAffinity (node ​​affinity attribute) is an attribute of Pod, which enables Pod to be preferentially scheduled to run on certain Nodes (preferred selection or mandatory requirement); it is used to guide Master to select appropriate nodes according to the attributes of the working nodes when scheduling Pods to working nodes. Among them, NodeAffinity can be divided into two types: mandatory selection (indicating that Pod must meet node affinity, otherwise it will not be scheduled to the corresponding node) and priority selection (indicating that Pod will try its best to meet node affinity, but if it cannot be met, it can also be scheduled to the corresponding node). It can be configured as needed, and the NodeAffinity attribute of the Pod in this embodiment can be set to priority selection.

[0120] Node's taint and Pod's toleration (taint and tolerance); taint needs to be used in conjunction with toleration to allow Pods to avoid inappropriate Nodes. Among them, taint is a property of the Node, and tolerance is a property of the Pod. In contrast to NodeAffinity, taint enables working nodes to exclude a specific type of pod. A Node with a taint set will have a mutually exclusive relationship between the taint and the Pod, and the Pod will not be scheduled to the Node to a certain extent. However, toleration can be set on the Pod, which means that a Pod with tolerance set will tolerate the existence of taints and can be scheduled to a Node with taints.

[0121] Based on this, the Master can allocate corresponding reserved resources for running three Pods to the container group set; Figure 2C As shown in , for the three Pods of container group set 1, the reserved resources are distributed on three different Nodes in the first resource pool. Each allocated reserved resource will be associated with container group set 1, for example, the label of the reserved resource can be configured as SetID=1.

[0122] Based on the resource reservation result of the Master, since the reserved resources are successfully allocated in the first resource pool, the Pod definition file of the container group set 1 mentioned above can be submitted to the Master. After the Master creates 3 Pods based on the Pod definition file, they will be scheduled to the corresponding reserved resources. The scheduling process of the Master can first search for available working nodes from each first-class working node in the first resource pool. For example, when scheduling Pod Set1-03, the Master can find that there is a reserved resource in Node03 that is identified as Set1 for the container group set based on the label of PodSet1-03 and the owner attribute of Reservation. Therefore, Pod Set1-03 will be scheduled to the reserved resource first.

[0123] In this embodiment, in order to further ensure that the Reservation of the container group set is used to run each Pod in the container group set; although the owner attribute of the Reservation of the container group set is configured, and an association is established between the Reservation and the container groups in the container group set, in actual applications, although the master node will give priority to scheduling each container group in the container group to the associated Reservation, there may be situations where the corresponding scheduling cannot be achieved due to load imbalance and other reasons. Based on this, the Pod definition file of the container group set in this embodiment also configures the other three attributes of the Pod except the label in the same way as the attributes of the virtual Pod in the Reservation definition file, so as to further make the scheduling of the Pod subject to strong restrictions of these attributes, and further ensure that the Pod can be scheduled to the corresponding Reservation.

[0124] Since the three Pods in the container group set 1 have the same label (set identifier), and the Pod mutual exclusion attribute of the three Pods is also configured as the set identifier, the Pod mutual exclusion attribute is used to specify the mutual exclusion relationship between different Pods. Since the Pod mutual exclusion attribute of the three Pods is configured as the same set identifier, when the Master schedules work nodes for these three Pods separately, it will not schedule them to run on the same work node.

[0125] For the three Pods of container group set 2, the scheduling process is similar. The three Pods are scheduled in three different Nodes in the first resource pool.

[0126] The scheduling process is similar for the three Pods of container group set 3. During the processing of reserved resources, the Master finds that the three Pods cannot be successfully scheduled in the first resource pool. Based on the resource reservation result returned by the Master, it triggers the execution of obtaining three virtual machines running on different computer devices, and adds the obtained three virtual machines to the Master's working node cluster. Each virtual machine serves as a reserved resource corresponding to one of the Pods of container group set 3.

[0127] The specific implementation of adding the obtained virtual machine to the Master's working node cluster may be to submit a definition file of the virtual machine to the Master. The definition file of the virtual machine may include the following information to ensure that after the virtual machine is configured as a working node, only the scheduling of the Pod in the container group set 3 is allowed:

[0128] The label attribute of Node is configured as: SetID=3;

[0129] The taint attribute of Node is configured as: SetID=3.

[0130] After that, you can submit the Pod definition file of the container group set 3 mentioned above to the Master. The Pod definition file contains the following information:

[0131] The Replicas attribute indicates the number of replicas of the Pod, which is configured to 3;

[0132] The configurations of the three Pods are the same. The configuration of Pod attributes can include:

[0133] The label attribute of the Pod is configured as: SetID = 3;

[0134] Indicates that the Pod AntiAffinity attribute of the Pod is configured as SetID = 3;

[0135] The affinity attribute nodeAffinity, which indicates the affinity between the Pod and the Node, is configured as SetID=3;

[0136] Indicates that the tolerance attribute of the Pod to the Node is configured as SetID=3.

[0137] After the Master creates three Pods based on the Pod definition file, it will schedule them to the corresponding reserved resources. The Master's scheduling process may be to first search for reserved resources from each first-class worker node in the first resource pool;

[0138] For example, when scheduling Pod Set1-03, the Master can find that Node03 has a reserved resource for the container group set identified as Set1. Therefore, Pod Set1-03 will be scheduled to this reserved resource first.

[0139] In the above embodiment, for each Pod in the container group set, its label attribute and podAntiAffinity attribute are configured as the set identifier; based on this, for the Master, since the label attributes and podAntiAffinity attributes of these multiple Pods are configured as the same set identifier, it means that these three Pods are mutually exclusive at the Node level and will not be scheduled on the same Node, thus achieving mutual exclusion between any two Pods in the container group set on the Node.

[0140] Since the first type of worker node in the first resource pool is a physical machine, and different first type of worker nodes are different physical machines, physical dispersion can be achieved in the first resource pool. If there are at least two Pods in the container group that need to be scheduled to different second type of worker nodes in the second resource pool, as virtual machines of different second type of worker nodes run on different computer devices, physical dispersion can also be achieved in the second resource pool.

[0141] In order to schedule the Pod to the corresponding Node, and the second-class working node corresponding to the obtained virtual machine is only used to run the corresponding Pod, the NodeAffinity attribute and tolerance attribute of the Pod can also be configured as a collection identifier, and the label and tolerance of the corresponding second-class working node are also configured as a collection identifier; therefore, when the Master schedules the working node for the Pod, since the first-class working nodes in the first resource pool do not have the label attribute, the NodeAffinity field enables the Pod to be scheduled to the node with the label in the second resource pool with the collection identifier first. At the same time, the second-class Node in the second resource pool is configured with the taint field to prevent a single run request from involving multiple container group sets. For example, as shown in the figure, after adding multiple second-class Nodes corresponding to different container group sets, scheduling errors may occur. For example, if the second-class Node is not configured with the taint field, it may cause the Pod of container group set 1 to be scheduled to the second-class Node of container group set 2.

[0142] In some examples, when the Pod running on the second type of Node ends, the Node corresponding to the Pod can be cleared from the working node cluster, so that the scheduling of the second type of Node can be terminated.

[0143] In other examples, the load information of each first-class working node in the working cluster can also be obtained. If there are idle resources, the container group running in the second-class working node can be scheduled to run in the first-class working node.

[0144] As can be seen from the above embodiments, in this embodiment, there is not only a fixed pool composed of large physical machines, but also an elastic pool composed of virtual machines. The hybrid resource pool composed of the two can maximize the transparency and reliability of inventory. This embodiment can be applied to service requirements in distributed multi-copy scenarios. By using resource reservation, resources can be prepared for Pod in advance, and by using the physical scattering protocol proposed in this embodiment, the Pod in this scenario can be deployed in the form of physical scattering, which can maximize the high availability requirements of the service.

[0145] This embodiment implements distributed multi-copy high-security inventory scheduling. This solution can meet the high availability requirements of the service and ensure the success rate of scheduling through resource reservation and physical-level scattered deployment, and the solution has good scalability. This embodiment can adapt to resource reservation in multiple availability zones to further enhance the high availability of the service. This embodiment also supports resource requests that do not require multiple copies to be scattered.

[0146] Corresponding to the aforementioned embodiment of the resource scheduling method based on the container management platform, this specification also provides an embodiment of a resource scheduling device based on the container management platform and a computer device used therein.

[0147] The resource scheduling device embodiments of the container management platform in this specification can be applied to computer devices, such as servers or terminal devices. The device embodiments can be implemented through software, hardware, or a combination of software and hardware. Taking software implementation as an example, as a device in a logical sense, it is formed by the processor in which it is located reading the corresponding computer program instructions in the non-volatile memory into the memory and running them. From the hardware level, if Figure 3 As shown in the figure, it is a hardware structure diagram of the computer device where the resource scheduling device of the container management platform is located in this specification, except Figure 3 In addition to the processor 310, memory 330, network interface 320, and non-volatile memory 340 shown, the computer device where the resource scheduling device 331 based on the container management platform is located in the embodiment may also include other hardware according to the actual function of the computer device, which will not be described in detail.

[0148] like Figure 4 As shown, Figure 4This is a block diagram of a resource scheduling device based on a container management platform according to an exemplary embodiment of the present specification, wherein the container management platform is used to manage a work node cluster, wherein the work node cluster includes at least one first-class work node, wherein the first-class work node is a physical machine, and the physical machine is used to run at least one container group; the device includes:

[0149] A receiving module 41 is used to receive an operation request of at least one container group;

[0150] A determination module 42 is used to determine whether the at least one container group can be successfully scheduled to run in the first type of working node in the working node cluster;

[0151] The scheduling module 43 is used to obtain virtual machines corresponding to each of the target container groups in response to determining that there is at least one target container group that has not been successfully scheduled to run in the first type of working node, add the obtained virtual machines to the working node cluster as the second type of working node, and schedule the second type of working node to run the corresponding target container group.

[0152] In some examples, the receiving module 41 receives a running request of at least one container group, including:

[0153] Receive a request to run a container group set, wherein the container group set includes: at least two container groups

[0154] The determination module 42 determines whether the container group in the container group set can be successfully scheduled to run in the first type of working node in the working node cluster, including:

[0155] Determine whether the container groups in the container group set can be successfully scheduled to run in different first-type working nodes in the working node cluster;

[0156] In response to determining that there is at least one target container group in the container group set that cannot be scheduled to run in the first type of working node, obtaining a virtual machine corresponding to each of the target container groups includes:

[0157] In response to determining that there are at least two target container groups in the container group set that cannot be scheduled to run in the first-type working node, a virtual machine corresponding to each of the target container groups and running on a different computer device is obtained.

[0158] In some examples, the container management platform includes a master node for managing the worker node cluster, and the determination module 42 determines whether the container group in the container group set can be successfully scheduled to run in different first-class worker nodes in the worker node cluster, including:

[0159] Sending a resource reservation request corresponding to the container group set to the master node, so that the master node attempts to allocate one or more reserved resources corresponding to the container group set and located at different first-category working nodes from the working node cluster based on the resource reservation request;

[0160] Obtain a resource reservation result of the master node, and determine, based on the resource reservation result, whether each container group in the container group set can be successfully scheduled to run in a first type of working node in the working node cluster.

[0161] In some examples, the scheduling module 43, in response to determining that there is at least one target container group in the container group set that has not been successfully scheduled to run in the first type of working node, obtains a virtual machine corresponding to each of the target container groups, including:

[0162] In response to the number of successfully allocated reserved resources included in the resource reservation result being lower than the number of container groups in the container group set, determining the number of target container groups in the container group set that have not been successfully scheduled to run in the first type of working nodes, and acquiring virtual machines corresponding to the number of target container groups;

[0163] The obtained virtual machine is added to the working node cluster as a second working node to serve as a reserved resource corresponding to the container group set.

[0164] In some examples, the determining module 42 sends a resource reservation request corresponding to the container group set to the master node, including:

[0165] The set identifier of the container group set is obtained, and a reserved resource object definition file corresponding to the container group set is generated and sent to the master node; wherein the reserved resource object definition file is used to define each reserved resource object having the same number as the container groups in the container group set, and the configuration of the label attribute and the container group mutual exclusivity attribute of each of the reserved resource objects includes the set identifier, so as to indicate to the master node that the reserved resources of each reserved resource object are mutually exclusive on the same working node, so that the master node allocates the reserved resources located at different first-class working nodes.

[0166] In some examples, the determination module 42 adds the acquired virtual machine to the working node cluster as a second working node to serve as a reserved resource corresponding to the container group set, including:

[0167] Obtain the set identifier of the container group set, generate a working node definition file corresponding to the virtual machine and send it to the master node; wherein, in the working node definition file, the configuration of the label attribute of the working node corresponding to the virtual machine includes the set identifier, so that the master node configures the virtual machine as a second type of working node in the working cluster and as a reserved resource corresponding to the container group set.

[0168] In some examples, the determining module 42 is further configured to:

[0169] Send a scheduling request corresponding to the container group set to the master node, so that the master node schedules each container group in the container group set to run in the reserved resources corresponding to the container group set.

[0170] In some examples, the configuration of the owner attribute of each of the reserved resource objects in the reserved resource object definition file includes the set identifier;

[0171] The scheduling module 43 is also used for:

[0172] A container group definition file corresponding to the container group set is generated and sent to the master node; wherein the container group definition file includes the configuration of the label attribute of each container group in the container group set including the set identifier, so that the master node schedules each container group in the container group set to the reserved resources corresponding to the container group set for operation based on the configuration of the owner attribute of the reserved resource object and the configuration of the label attribute of each container group in the container group set.

[0173] The implementation process of the functions and effects of each module in the above-mentioned resource scheduling device based on the container management platform is specifically described in the implementation process of the corresponding steps in the above-mentioned resource scheduling method based on the container management platform, and will not be repeated here.

[0174] Correspondingly, an embodiment of the present specification also provides a container management system, the system includes a scheduling node and a working node cluster, the working node cluster includes at least one first-class working node, the first-class working node is a physical machine, and the physical machine is used to run at least one container group; the scheduling node is used to execute the steps of an embodiment of a resource scheduling method based on a container management platform.

[0175] Accordingly, an embodiment of the present specification further provides a computer program product, including a computer program, which implements the steps of the aforementioned resource scheduling method embodiment based on a container management platform when executed by a processor.

[0176] Correspondingly, an embodiment of the present specification also provides a computer device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the program, the steps of an embodiment of a resource scheduling method based on a container management platform are implemented.

[0177] Accordingly, an embodiment of the present specification further provides a computer-readable storage medium on which a computer program is stored. When the computer program is executed by a processor, the steps of an embodiment of a resource scheduling method based on a container management platform are implemented.

[0178] For the device embodiment, since it basically corresponds to the method embodiment, the relevant parts can refer to the partial description of the method embodiment. The device embodiment described above is only schematic, wherein the modules described as separate components may or may not be physically separated, and the components displayed as modules may or may not be physical modules, that is, they may be located in one place, or they may be distributed on multiple network modules. Some or all of the modules may be selected according to actual needs to achieve the purpose of the scheme of this specification. A person of ordinary skill in the art can understand and implement it without paying creative labor.

[0179] The above embodiments can be applied to one or more computer devices, where the computer device is a device that can automatically perform numerical calculations and / or information processing according to pre-set or stored instructions. The hardware of the computer device includes but is not limited to a microprocessor, an application specific integrated circuit (ASIC), a field-programmable gate array (FPGA), a digital signal processor (DSP), an embedded device, etc.

[0180] The computer device may be any electronic product that can perform human-computer interaction with a user, such as a personal computer, a tablet computer, a smart phone, a personal digital assistant (PDA), a game console, an interactive network television (IPTV), a smart wearable device, etc.

[0181] The computer device may also include a network device and / or a user device, wherein the network device includes, but is not limited to, a single network server, a server group consisting of multiple network servers, or a cloud consisting of a large number of hosts or network servers based on cloud computing.

[0182] The network where the computer device is located includes but is not limited to the Internet, a wide area network, a metropolitan area network, a local area network, a virtual private network (VPN), etc.

[0183] The above is a description of a specific embodiment of the specification. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recorded in the claims can be performed in an order different from that in the embodiments and still achieve the desired results. In addition, the processes depicted in the drawings do not necessarily require the specific order or continuous order shown to achieve the desired results. In some embodiments, multitasking and parallel processing are also possible or may be advantageous.

[0184] The step division of the above methods is only for clear description. When implemented, they can be combined into one step or some steps can be split and decomposed into multiple steps. As long as they include the same logical relationship, they are all within the protection scope of this patent; adding insignificant modifications to the algorithm or process or introducing insignificant designs without changing the core design of the algorithm and process are all within the protection scope of this application.

[0185] The description of "specific examples" or "some examples" means that the specific features, structures, materials or characteristics described in conjunction with the embodiment or example are included in at least one embodiment or example of this specification. In this specification, the schematic representation of the above terms does not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials or characteristics described can be combined in any one or more embodiments or examples in a suitable manner.

[0186] Those skilled in the art will readily appreciate other embodiments of the specification after considering the specification and practicing the invention claimed herein. The specification is intended to cover any variations, uses or adaptations of the specification that follow the general principles of the specification and include common knowledge or customary techniques in the art that are not claimed in the specification. The specification and examples are to be considered exemplary only, and the true scope and spirit of the specification are indicated by the following claims.

[0187] It should be understood that the present description is not limited to the precise structures that have been described above and shown in the drawings, and that various modifications and changes may be made without departing from the scope thereof. The scope of the present description is limited only by the appended claims.

[0188] The above description is only a preferred embodiment of this specification and is not intended to limit this specification. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of this specification should be included in the scope of protection of this specification.

Claims

1. A resource scheduling method based on a container management platform, wherein the container management platform is used to manage a work node cluster, wherein the work node cluster includes at least one first-class work node, wherein the first-class work node is a physical machine, and the physical machine is used to run at least one container group; include: Receive a running request of at least one container group; Determine whether the at least one container group can be successfully scheduled to run in a first type of working node in the working node cluster; In response to determining that there is at least one target container group that cannot be scheduled to run in the first type of worker node, a virtual machine corresponding to each of the target container groups is obtained, and the obtained virtual machine is added to the worker node cluster as a second type of worker node, where the second type of worker node is used to run the corresponding target container group.

2. The method according to claim 1, wherein receiving a running request of at least one container group, include: Receiving a request for running a container group set, wherein the container group set includes: at least two container groups; The determining whether the at least one container group can be successfully scheduled to run in the first type of working node in the working node cluster includes: Determine whether the container groups in the container group set can be successfully scheduled to run in different first-type working nodes in the working node cluster; In response to determining that there is at least one target container group in the container group set that cannot be scheduled to run in the first type of working node, obtaining a virtual machine corresponding to each of the target container groups includes: In response to determining that there are at least two target container groups in the container group set that cannot be scheduled to run in the first-type working node, a virtual machine corresponding to each of the target container groups and running on a different computer device is obtained.

3. The method according to claim 2, wherein the container management platform includes a master node for managing the worker node cluster, wherein determining whether the container group in the container group set can be successfully scheduled to run in different first-class worker nodes in the worker node cluster, include: Sending a resource reservation request corresponding to the container group set to the master node, so that the master node attempts to allocate one or more reserved resources corresponding to the container group set and located at different first-category working nodes from the working node cluster based on the resource reservation request; Obtain a resource reservation result of the master node, and determine, based on the resource reservation result, whether each container group in the container group set can be successfully scheduled to run in a first type of working node in the working node cluster.

4. The method according to claim 3, in response to determining that there is at least one target container group in the container group set that has not been successfully scheduled to run in the first type of working node, obtaining a virtual machine corresponding to each of the target container groups, include: In response to the number of successfully allocated reserved resources included in the resource reservation result being lower than the number of container groups in the container group set, determining the number of target container groups in the container group set that have not been successfully scheduled to run in the first type of working nodes, and acquiring virtual machines corresponding to the number of target container groups; The obtained virtual machine is added to the working node cluster as a second working node to serve as a reserved resource corresponding to the container group set.

5. The method according to claim 3, wherein the resource reservation request corresponding to the container group set is sent to the master node, include: The set identifier of the container group set is obtained, and a reserved resource object definition file corresponding to the container group set is generated and sent to the master node; wherein the reserved resource object definition file is used to define each reserved resource object having the same number as the container groups in the container group set, and the configuration of the label attribute and the container group mutual exclusivity attribute of each of the reserved resource objects includes the set identifier, so as to indicate to the master node that the reserved resources of each reserved resource object are mutually exclusive on the same working node, so that the master node allocates the reserved resources located at different first-class working nodes.

6. The method according to claim 4, wherein the obtained virtual machine is added to the working node cluster as a second working node to serve as a reserved resource corresponding to the container group set, include: Obtain the set identifier of the container group set, generate a working node definition file corresponding to the virtual machine and send it to the master node; wherein, in the working node definition file, the configuration of the label attribute of the working node corresponding to the virtual machine includes the set identifier, so that the master node configures the virtual machine as a second type of working node in the working cluster and as a reserved resource corresponding to the container group set.

7. The method according to claim 6, after the step of adding the acquired virtual machine to the working node cluster as the second type of working node, the method further include: Send a scheduling request corresponding to the container group set to the master node, so that the master node schedules each container group in the container group set to run in the reserved resources corresponding to the container group set.

8. The method according to claim 7, wherein the configuration of the owner attribute of each of the reserved resource objects in the reserved resource object definition file includes the set identifier; The scheduling request corresponding to the container group set is sent to the master node, so that the master node schedules each container group in the container group set to the reserved resources corresponding to the container group set for operation. include: A container group definition file corresponding to the container group set is generated and sent to the master node; wherein the configuration of the label attribute of each container group in the container group definition file includes the set identifier, so that the master node schedules each container group in the container group set to the reserved resources corresponding to the container group set for operation based on the configuration of the owner attribute of the reserved resource object and the configuration of the label attribute of each container group in the container group set.

9. A resource scheduling device based on a container management platform, wherein the container management platform is used to manage a work node cluster, wherein the work node cluster includes at least one first-class work node, wherein the first-class work node is a physical machine, and the physical machine is used to run at least one container group; the device include: A receiving module, configured to receive a running request of at least one container group; A determination module, used to determine whether the at least one container group can be successfully scheduled to run in a first type of working node in the working node cluster; A scheduling module, configured to, in response to determining that there is at least one target container group that cannot be scheduled to run in the first-type working node, obtain a virtual machine corresponding to each of the target container groups, and add the obtained virtual machines to the working node cluster as a second-type working node, where the second-type working node is used to run the corresponding target container group.

10. A container management system, the system comprising a scheduling node and a working node cluster, the working node cluster comprising at least one first-class working node, the first-class working node being a physical machine, the physical machine being used to run at least one container group; the scheduling node being used to execute the steps of any one of the methods described in claims 1 to 8.

11. A computer device comprising a memory, a processor and a computer program stored in the memory and executable on the processor, in, When the processor executes the computer program, the steps of the method according to any one of claims 1 to 8 are implemented.

12. A computer-readable storage medium having a computer program stored thereon, wherein the computer program, when executed by a processor, implements the steps of the method according to any one of claims 1 to 8.