A method and device for container resource management in a cluster

By using the seat mark bit decoupling scheduling module and container creation deployment module in the cluster, the coexistence problem caused by lag in container resource release is solved, and efficient resource scheduling and system throughput are achieved.

CN114780203BActive Publication Date: 2025-05-30HANGZHOU MAGIC SQUARE ARTIFICIAL INTELLIGENCE FOUNDATION RES CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210340621.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-03-31
Publication Date
2025-05-30
Estimated Expiration
2042-03-31

AI Technical Summary

Technical Problem

In the prior art, there is a lag time for the release of container resources, resulting in container coexistence problems. Newly launched containers need to compete with containers that have not been actually deleted, resulting in inefficient operation and may even cause serious problems such as insufficient resources.

Method used

By presetting seat marking bits in the cluster, decoupling the scheduling module from the container creation and deployment module to achieve real-time resource scheduling. After each round of scheduling, if the task running on the node needs to be interrupted and suspended, the task will be forced to be deleted, and the idle node will be entered as available nodes in the new round of scheduling.

Benefits of technology

Accelerate the implementation of resource scheduling of clusters, improve the efficiency of scheduling system, improve independence, eliminate the risks brought about by the coexistence of training containers, improve system throughput, and provide more refined control conditions for subsequent node resource management.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114780203B_ABST
    Figure CN114780203B_ABST
Patent Text Reader

Abstract

The present invention relates to a method for container resource management in a cluster, comprising the following steps: (1) presetting seats for each node in the cluster; (2) after each round of scheduling, if the task running on a node needs to be interrupted and suspended, forcibly deleting the task; (3) regarding the nodes where tasks are forcibly deleted by the scheduler and the tasks running are terminated as all idle available nodes to enter a new round of scheduling; (4) obtaining the existence status of containers inside a single node and updating the seat status in real time at intervals of a seat monitoring trigger time; (5) for the nodes that need to execute new training tasks in the new round of scheduling, applying for the resources of one seat when starting the container, and waiting for seat resources if there are no seat resources; (6) creating a new container to start a new training task, and changing the seat status. The present invention can improve the efficiency of the scheduling system, eliminate the risks brought by the coexistence of training containers, and improve the independence of the scheduling module at the same time.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computing clusters, and particularly to a method and device for container resource management in a cluster. Background Art

[0002] With the development of cloud computing technology, complex models need to be trained using large-scale computing clusters, and Kubernetes container technology is usually used for large-scale deployment of training environments. The process of model training needs to consider the issues of computing power allocation and task scheduling. The invention application (patent application number: 2021114594024) discloses a method for cluster management task scheduling. This method needs to trigger scheduling at regular intervals, allocate available nodes to tasks that have obtained running permissions, and for tasks that have not obtained running permissions: if they are in a non-running state, they are suspended and continue to wait; if they are already running, they need to be aborted and suspended, and the node resources they occupy also need to be released. However, in the prior art, for nodes that release resources by forced deletion, there is a certain lag time from the occurrence of the forced deletion behavior to the actual deletion of the container and the release of resources. During system resource scheduling, this lag time will cause container coexistence problems, and newly started containers need to compete for resources with containers that have not been actually deleted, resulting in low operating efficiency and even serious problems such as insufficient resources. If scheduling is performed after the containers are actually deleted, it will also lead to problems such as low efficiency of the scheduling system and reduced resource utilization. In this case, a more efficient and reasonable method for container resource management in a cluster becomes particularly important. Summary of the Invention

[0003] Aiming at the above-mentioned defects existing in the prior art, the present invention provides a method for container resource management in a cluster. By marking with seat flags, the scheduling module is decoupled from the container creation and deployment module, which can accelerate the realization of cluster resource scheduling, improve the efficiency of the scheduling system, improve the independence of the scheduling module, and eliminate the risks brought by the coexistence of training containers, so as to solve the problems existing in the prior scheduling technology. The application of this method also provides conditions for more refined control of subsequent node resource management.

[0004] The present invention achieves the above object through the following technical solutions: A method for container resource management in a cluster, comprising the following steps:

[0005] (1) Preset seats for each node in the cluster to mark the number of resources of the node that can be used to create new training containers;

[0006] (2) After each round of scheduling program scheduling, if the task running on the node needs to be interrupted and suspended, then perform forced deletion on the task;

[0007] (3) For the nodes that are forcibly deleted due to the execution of the scheduler task and the task running is terminated, all enter the new round of scheduling as idle available nodes;

[0008] (4) Every time the seat monitoring trigger time elapses, obtain the existence status of the containers inside a single node, and update the seat status in real time according to the existence status of the containers inside the node;

[0009] (5) For the nodes that need to execute new training tasks in the new round of scheduling, apply for the resources of one seat when starting the container on this node. If there are no seat resources, it is necessary to wait for the seat resources;

[0010] (6) After applying for the seat resources, create a new container and start a new training task, and the seat status changes.

[0011] Preferably, the seat resources are 0 or 1.

[0012] Preferably, the seat monitoring trigger time described in step (4) is set to 0.5 seconds to 1 second.

[0013] Preferably, in step (4), the existence status of the training containers inside a single node is directly read through the GPU low-level interface.

[0014] Preferably, if there are still containers in the node described in step (4), mark the seat as 0. If there are no containers, mark the seat as 1.

[0015] Preferably, if the seat of the node described in step (5) is marked as 1, create a new container to execute the new training task, and mark it as 0 after the seat is occupied; if the seat is marked as 0, continue to wait until the seat is marked as 1 and the resources are released.

[0016] A computer device applying the above method includes a memory, a processor, a bus, and a computer program stored on the memory and executable on the processor.

[0017] The beneficial effect of the present invention is that it provides a method for managing container resources in a cluster, which can accelerate the resource scheduling of the cluster. Through the marking of the seat flag bits, the scheduling module is decoupled from the container creation and deployment module. The scheduling module does not need to care whether the resources of each node are really released. As long as the node that terminates the operation or is forcibly deleted in the previous round of scheduling can enter the next round of scheduling as an idle available node. Before loading the new container, detect whether the node is really released again. For the nodes that are not released, control it to wait through the seat flag bits until the node resources are really released and start loading the new training task.

[0018] The above method can greatly improve the throughput of the system, enhance the efficiency of the scheduling system, and eliminate the risks brought by the coexistence of training containers. Meanwhile, the independence of the scheduling module is improved. The application of this method also provides conditions for more refined control of subsequent node resource management. Brief Description of the Drawings

[0019] Figure 1 is a schematic flowchart of the method of the present invention; Detailed Embodiments

[0020] The technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. The components of the embodiments of the present invention described and illustrated herein can be arranged and designed in various different configurations.

[0021] In a specific embodiment, the scheduling trigger time can be set to 1 second.

[0022] As Figure 1 shown, a method for managing container resources in a cluster is as follows: The following operation steps are cyclically triggered:

[0023] S101 Preset seats for each node in the cluster to mark the number of resources of the node that can be used to create new training containers.

[0024] The preset seats are used to mark the number of resources of the node that can be used to create new training containers.

[0025] In specific implementation, the cluster includes multiple nodes, and the nodes refer to computer servers, including one or more GPUs, CPUs, etc. When the scheduling module sends out the information to start a training task, the corresponding container will be started on the corresponding node, and the corresponding algorithm model and relevant parameters will be loaded into the GPU video memory and CPU memory according to the training task information to execute the training task.

[0026] In specific implementation, machine learning algorithm code is trained through a computer cluster, such as a neural network model training program (RNN), etc. After the user completes the code development of the training program through the client, it can be uploaded to the cluster training cloud service platform to create a training task, queue in the task list waiting for the allocation of computing nodes, realize task scheduling, deploy the environment and run the training task.

[0027] Through containerization technology, training environment containers and container images are created, and the training environment such as the operating system, driver, configuration file, project training framework, etc. is synchronized to the allocated node computers to complete the automatic deployment, planning, update and maintenance of the training environment.

[0028] In this embodiment, the seat resource is 0 or 1. Since there can be at most one training container on a node and one training task can be run. If the seat is 0, it means the seat resource is insufficient, there may be an old container on this node, and the seat resource of this node is occupied; if the seat is marked as 1, it means there is no container on this node, the seat resource of the node has been released, and there is one available seat, and a new training container can be created.

[0029] S102 After each round of scheduling by the scheduler, if the task running on the node needs to be interrupted and suspended, then forcibly delete this task.

[0030] The system platform preset scheduling module realizes the scheduling of node resources through the scheduler in the scheduling module. The scheduler refers to, through a certain allocation and scheduling method, realizing the specific allocation of training tasks to specific nodes for training and arranging and deploying. At the same time, there is only one training container on the same node and one training task can be run. Due to limited computing power resources, in order to realize the reasonable utilization of node resources, it is necessary for the scheduling module to arrange whether the training task can obtain permission to run, whether the training task needs to be suspended, and specifically allocate the training task to a specific node.

[0031] The scheduler performs resource allocation and scheduling for all tasks submitted by users. Except for the tasks that have been completed and terminated, including tasks in the waiting-to-run, running, and suspended states. After the task is submitted, it enters the waiting-to-run state and waits in the queue for the allocation of the scheduling module. For the tasks that obtain the running permission through the scheduling algorithm, they will obtain the corresponding allocated nodes and enter the running state. For the tasks that do not obtain the running permission, if they are in non-running states such as waiting-to-run and suspended, they will continue to maintain the waiting and suspended state; if the tasks are already in the running state, they need to be forcibly deleted to release the node resources, and the tasks enter the suspended state.

[0032] For a single node, if the task already running on this node is considered to be able to continue running after scheduling, then keep the running state unchanged; if after scheduling, the task running on this node needs to be interrupted and suspended, then forcibly delete this task, and the seat resource of the node will be released.

[0033] The scheduler runs once every interval of a trigger time. In cases such as new task creation and addition, task error, and exit, the scheduler is not immediately triggered, but waits until the interval scheduling trigger time and then runs this scheduler once again.

[0034] In a specific embodiment, the trigger time of the scheduler can be set to 1 second.

[0035] S103 For the nodes where the scheduler forcibly deletes tasks and the tasks run to termination, all are regarded as idle available nodes and enter a new round of scheduling.

[0036] The available nodes described above are the nodes among all cluster nodes that are in normal running status, not disabled, and not assigned. Among all the cluster node computer servers, some node servers can be manually disabled, and some node servers have been assigned by the scheduling module to execute tasks. The above statuses do not belong to available nodes and cannot be used for assigned scheduling and execution tasks.

[0037] In a specific embodiment, after one round of scheduling, the nodes that need to release resources are all regarded as idle available nodes and enter a new round of scheduling. Generally, there will be a certain lag time from when a forced deletion is executed until the containers on the node are truly released. The lag time usually varies from 1 to 50 seconds depending on the applications running in the containers. In the technical solution of this embodiment, for the nodes where forced deletion is executed due to the scheduling algorithm and the tasks are terminated, regardless of whether the containers are truly released, they are regarded as idle available nodes and directly enter a new round of scheduling.

[0038] S104 At every interval of a seat monitoring trigger time, obtain the existence status of the containers inside a single node, and update the seat status in real time according to the existence status of the containers inside the node.

[0039] In a specific embodiment, at every interval of a seat monitoring trigger time t, through the GPU low-level interface, detect whether there are still training containers in the corresponding node. The interface can be provided by docker ps or container runtimeclient. In a specific embodiment, each seat monitoring trigger time t can be adjusted according to specific circumstances and is usually set to 0.5s - 1s.

[0040] If there are still containers in this node, set the seat to 0. If there are no containers, set the seat to 1. On a node, after a forced deletion is executed, if the old training containers have not been deleted and the seat resources have not been released, set the seat to 0. If the old containers have been deleted and the seat resources have been released, the seat is 1, indicating that one seat resource can be obtained for creating a new training container.

[0041] S105 For the nodes that need to execute new training tasks in the new round of scheduling, when starting the containers on this node, apply for one seat resource. If there is no seat resource, it is necessary to wait for the seat resource.

[0042] If the seat of this node is 1, create a new container to execute the new training task, and the seat becomes 0 after being occupied. If the seat is 0, continue to wait for the seat resource until the seat flag becomes 1 and the resource is released.

[0043] In a specific implementation, after a new round of scheduling is completed, for the nodes assigned to execute the new training tasks, before starting a new container on a node, it is necessary to ensure that the old container has been deleted. This is achieved through a seat flag. When starting a container on a node, a seat resource needs to be applied for. Kubernetes will automatically check whether the conditions are met to ensure that the seat is 1 and there is no container on the node. At this time, a new container can be directly created to start the new training. After the seat is occupied, it becomes 0. If the seat flag is 0 and the seat condition is not met and the resources are not sufficient, the container on the node still exists, then it needs to be in the Pending state and wait for the old container to be deleted. After the resources are released and the seat is 1, a new container is created to start the training.

[0044] After obtaining the seat resource in S106, a new container is created to start a new training task, and the seat status changes.

[0045] For the case where the seat is 1, after obtaining the seat resource, a new training container can be started. At this time, the seat on this node is occupied and the seat status becomes 0.

[0046] In several embodiments provided by the present application, it should be understood that the disclosed devices and methods can also be implemented in other ways. The device embodiments described above are only illustrative. For example, the flowcharts and block diagrams in the accompanying drawings show the possible architectures, functions, and operations of devices, methods, and computer program products according to multiple embodiments of the present invention. In this regard, each block in the flowchart or block diagram may represent a module, a program segment, or a part of code, and the module, program segment, or part of code contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order from that marked in the accompanying drawings. For example, two consecutive blocks may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram and / or flowchart, and the combination of blocks in the block diagram and / or flowchart, can be implemented by a dedicated hardware-based system for performing the specified functions or actions, or can be implemented by a combination of dedicated hardware and computer instructions.

[0047] In addition, the functional modules in each embodiment of the present invention can be integrated together to form an independent part, or each module can exist separately, or two or more modules can be integrated to form an independent part.

[0048] When the above-mentioned functions are implemented in the form of software function modules and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present invention. The aforementioned storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), magnetic disks, or optical discs that can store program codes. It should be noted that in this article, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements, but also includes other elements not expressly listed, or also includes elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "including one..." does not exclude the existence of additional identical elements in the process, method, article or device including the said element.

[0049] The above are only the preferred embodiments of the present invention and are not used to limit the present invention. For those skilled in the art, the present invention can have various changes and modifications. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention. It should be noted that similar reference numerals and letters denote similar items in the following drawings. Therefore, once an item is defined in one drawing, it does not need to be further defined and explained in subsequent drawings.

[0050] The above is only the specific implementation manner of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention can easily think of changes or replacements, which should all be covered within the protection scope of the present invention. Therefore, the protection scope of the present invention should be subject to the protection scope of the claims.

Claims

1. A method for container resource management in a cluster, characterized in that, it includes the following steps: (1) Preset seats for each node in the cluster to mark the number of resources that the node can use to create new training containers; (2) After each round of scheduler scheduling, if the task running on the node needs to be interrupted and suspended, then forcibly delete the task; (3) For the nodes where the tasks are forcibly deleted due to scheduler tasks and the tasks run to termination, all of them enter a new round of scheduling as idle available nodes; (4) Every time an interval of seat monitoring trigger time elapses, obtain the existence status of containers inside a single node, and update the seat status in real time according to the existence status of containers inside the node; (5) For the nodes that need to execute new training tasks in the new round of scheduling, when starting a container on the node, apply for the resources of one seat. If there are no seat resources, then wait for seat resources; (6) After applying for seat resources, create a new container and start a new training task, and the seat status changes.

2. The method for container resource management in a cluster according to claim 1, characterized in that, the seat resources are 0 or 1.

3. The method for container resource management in a cluster according to claim 1, characterized in that, the seat monitoring trigger time described in step (4) is set to 0.5 seconds - 1 second.

4. The method for container resource management in a cluster according to claim 1, characterized in that, in step (4), the existence status of training containers inside a single node is directly read through the GPU low-level interface.

5. The method for container resource management in a cluster according to claim 1, characterized in that, if there are still containers existing in the node described in step (4), then mark the seat as 0. If there are no containers existing, then mark the seat as 1.

6. The method for container resource management in a cluster according to claim 1, characterized in that, if the seat of the node described in step (5) is marked as 1, then create a new container to execute a new training task, and mark it as 0 after the seat is occupied; if the seat is marked as 0, then continue to wait until the seat mark becomes 1 and the resources are released.

7. A computer device applying the method according to any one of claims 1 - 6, characterized in that, it includes a memory, a processor, a bus, and a computer program stored on the memory and executable on the processor.

Citation Information

Patent Citations

  • Task scheduling method and apparatus for container cloud

    CN107729126A

  • Training task resource scheduling method and device, equipment and medium

    CN113867959A