Container node scheduling method and device, equipment, storage medium and program product

By implementing a container node scheduling method based on resource information and Pod request attributes in the scheduler of Kubernetes cluster, the problem of low scheduling efficiency of Pod nodes is solved and a more efficient scheduling process is achieved.

CN120045319APending Publication Date: 2025-05-27CHINA MOBILE (SUZHOU) SOFTWARE TECH CO LTD +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510096662.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-21
Publication Date
2025-05-27

AI Technical Summary

Technical Problem

The node scheduling efficiency of Pods in Kubernetes cluster is low. In the prior art, Kube-scheduler needs to determine the node scheduling scheme of Pod through multiple blind guesses.

Method used

By implementing a container node scheduling method in the scheduler of the Kubernetes cluster, the method includes determining the preselected nodes based on the resource information of each node and the resource attributes of the Pod request, and scoring these nodes through the controller. Finally, the scheduler selects the scheduling node of the Pod based on the scoring result.

Benefits of technology

It improves the node scheduling efficiency of Pod, avoids the process of Kube-scheduler blind guessing multiple times, and ensures the effectiveness and efficiency of the scheduling results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120045319A_ABST
    Figure CN120045319A_ABST
Patent Text Reader

Abstract

The embodiment of the invention discloses a node scheduling method and device for a container, equipment, a storage medium and a program product, and the method comprises the steps: at least based on the information of resources in each node and the attributes which are required to be satisfied by the resources requested when Pod is created, determining a pre-selected node in each node; sending the identifiers of the pre-selected nodes to a controller, so that the controller determines a first adaptation degree score of each node in the pre-selected nodes and a scheduling node of the Pod; and receiving a first adaptation degree score sent by the controller, and selecting the scheduling node of the Pod from the pre-selected nodes based on a second adaptation degree score of each node in the pre-selected nodes determined by the scheduler and the scheduling node of the Pod and the first adaptation degree score.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application belongs to the field of cloud computing technology, and particularly relates to a method, device, equipment, storage medium and program product for node scheduling of containers. Background Art

[0002] In the related art, the Dynamic Resource Allocation (DRA) mechanism newly provided by Kubernetes can be adopted to implement node scheduling of a container group (Pod); exemplarily, the node scheduling of the Pod can be completed through the negotiation process between the scheduler (Kube-scheduler) of the Kubernetes cluster and the Resource Driver Controller. However, in the negotiation process between the Kube-scheduler and the Resource Driver Controller, the Kube-scheduler can only determine the node scheduling scheme of the Pod through multiple blind guesses, which reduces the scheduling efficiency. Summary of the Invention

[0003] To solve the problem of low efficiency in node scheduling of Pods existing in the related art, embodiments of this application propose a method, device, equipment, storage medium and program product for node scheduling of containers.

[0004] Embodiments of this application provide a method for node scheduling of containers, which is applied to the scheduler of a Kubernetes cluster. The method includes:

[0005] Determine preselected nodes among the various nodes based at least on information about resources in each node and attributes that the resources requested when creating the Pod need to satisfy;

[0006] Send the identifiers of the preselected nodes to the controller, so that the controller determines a first fitness score of each node among the preselected nodes with the scheduling node of the Pod;

[0007] Receive the first fitness score sent by the controller, and select the scheduling node of the Pod among the preselected nodes based on the second fitness score of each node among the preselected nodes determined by the scheduler with the scheduling node of the Pod and the first fitness score.

[0008] In some embodiments, determining preselected nodes among the various nodes at least based on the information of resources in each node and the attributes that the resources requested when creating a Pod need to satisfy includes: obtaining filtering rules for each type of resource preset; determining the preselected nodes among the various nodes based on the filtering rules for each type of resource, the information of resources in each node, and the attributes that the resources requested when creating a Pod need to satisfy.

[0009] In some embodiments, the filtering rules for each type of resource are judgment statements organized using logical operations.

[0010] In some embodiments, before determining preselected nodes among the various nodes at least based on the information of resources in each node and the attributes that the resources requested when creating a Pod need to satisfy, the method further includes: when obtaining a scheduling request for a Pod, creating a resource declaration, and determining at least two types of resources requested when creating the Pod according to the resource declaration.

[0011] An embodiment of the present application further provides another method for scheduling nodes of a container, which is applied to a controller, and the method includes:

[0012] Receiving the identifiers of the preselected nodes sent by the scheduler of the Kubernetes cluster, where the scheduler is used to determine the preselected nodes among the various nodes at least based on the information of resources in each node and the attributes that the resources requested when creating a Pod need to satisfy;

[0013] Determining a first fitness score of each node among the preselected nodes with the scheduling node of the Pod;

[0014] Sending the first fitness score to the scheduler, so that the scheduler selects the scheduling node of the Pod among the preselected nodes according to the second fitness score of each node among the preselected nodes determined by itself and the first fitness score.

[0015] In some embodiments, determining the first fitness score of each node among the preselected nodes with the scheduling node of the Pod includes: at least determining a resource allocation scheme of each node according to the topological relationship between different types of resources in each node; determining the first fitness score of each node with the scheduling node of the Pod according to the resource allocation scheme of each node.

[0016] In some embodiments, determining a first fitness score of each node and the scheduling node of the Pod according to the resource allocation scheme of each node includes: determining resource attribute data corresponding to the resource allocation scheme of each node, where the resource attribute data includes at least one of the following: the failure rate of each resource in each node, the minimum bandwidth between different resources in each node, and the matching degree between each resource in each node and the task characteristics of the Pod.

[0017] Determine a first fitness score of each node and the scheduling node of the Pod according to the resource attribute data corresponding to the resource allocation scheme of each node.

[0018] An embodiment of the present application further provides a node scheduling device for a container, which is applied to a scheduler of a Kubernetes cluster. The device includes:

[0019] A first processing module, configured to determine preselected nodes among the various nodes at least based on information about resources in each node and attributes that the resources requested when creating a Pod need to satisfy.

[0020] A second processing module, configured to send the identifiers of the preselected nodes to a controller, so that the controller determines a first fitness score of each node among the preselected nodes and the scheduling node of the Pod; receive the first fitness score sent by the controller, and select the scheduling node of the Pod among the preselected nodes based on the second fitness score of each node among the preselected nodes determined by the scheduler and the first fitness score.

[0021] An embodiment of the present application further provides another node scheduling device for a container, which is applied to a controller. The device includes:

[0022] A receiving module, configured to receive the identifiers of the preselected nodes sent by a scheduler of a Kubernetes cluster, where the scheduler is configured to determine preselected nodes among the various nodes at least based on information about resources in each node and attributes that the resources requested when creating a Pod need to satisfy.

[0023] A third processing module, configured to determine a first fitness score of each node among the preselected nodes and the scheduling node of the Pod; send the first fitness score to the scheduler, so that the scheduler selects the scheduling node of the Pod among the preselected nodes according to the second fitness score of each node among the preselected nodes determined by itself and the first fitness score.

[0024] An embodiment of the present application further provides an electronic device, which includes a processor and a memory for storing a computer program that can run on the processor; wherein, the processor is used to run the computer program to execute the node scheduling method of any one of the above containers.

[0025] An embodiment of the present application further provides a computer storage medium, on which a computer program is stored, and when the computer program is executed by a processor, it implements the node scheduling method of any one of the above containers.

[0026] An embodiment of the present application further provides a computer program product, including a computer program, and when the computer program is executed by a processor, it implements the node scheduling method of any one of the above containers.

[0027] It can be seen that in the embodiment of the present application, the scheduler of the Kubernetes cluster does not need to blindly guess multiple times to determine the nodes that the controller does not reject. The controller only scores the preselected nodes for the reference of the scheduler of the Kubernetes cluster. Finally, the scheduler of the Kubernetes cluster can determine the scheduling node of the Pod at one time according to the first fitness score and the second fitness score, improving the scheduling efficiency. Description of the Drawings

[0028] Figure 1 A schematic diagram for implementing the management and allocation of Kubernetes resources provided by the related art;

[0029] Figure 2 A schematic diagram for allocating resources to a container provided by the related art;

[0030] Figure 3 Another schematic diagram for allocating resources to a container provided by the related art;

[0031] Figure 4 Another schematic diagram for allocating resources to a container provided by the related art;

[0032] Figure 5 A flowchart of the node scheduling method of the container applied to the scheduler of the Kubernetes cluster in the embodiment of the present application;

[0033] Figure 6 A schematic diagram for selecting a scheduling node by a preselection rule and a preference rule provided by the embodiment of the present application;

[0034] Figure 7 A flowchart of the node scheduling method of the container applied to the controller in the embodiment of the present application;

[0035] Figure 8Schematic structural diagram of a node scheduling device for a container of a scheduler applied to a Kubernetes cluster in an embodiment of the present application;

[0036] Figure 9 Schematic structural diagram of a node scheduling device for a container of a controller in an embodiment of the present application;

[0037] Figure 10 Schematic compositional structure diagram of an electronic device provided in an embodiment of the present application. Detailed implementation manners

[0038] With the development of large models and cloud native technologies, a common technical solution for large model training and inference is as follows: Kubernetes is used for container orchestration, and intelligent computing resources such as Graphics Processing Unit (GPU), Remote Direct Memory Access (RDMA) devices, and Neural Processing Unit (NPU) are allocated to containers to train and infer models in the form of containers. For example, common tools such as kubeflow and argo are implemented based on this technical solution. When training a model in the form of a container, it is generally necessary to use a multi-instance method for distributed training to fully utilize the GPU resources of the entire cluster. For example, OpenAI uses a Kubernetes cluster with tens of thousands of nodes to train a Large Language Model (LLM).

[0039] In distributed training, the topology of the resources allocated to each subtask affects the overall speedup because data transfer is required between tasks, and the transfer bandwidth can become a bottleneck for training efficiency. Taking NVIDIA's collective communication library NCCL (NVIDIA Collective Communication Library) as an example, when planning the transfer path, NCCL preferentially selects GPUs connected by NVlink, which has higher bandwidth compared to GPUs connected by the Peripheral Component Interconnect express (PCIe). In addition, the GPU Direct RDMA technology provided by NVIDIA can achieve fast video memory exchange through RDMA technology, but it requires that the RDMA device and the GPU be connected only through PCIe to ensure the transfer bandwidth. Therefore, when using containers for distributed training, allocating resources such as GPUs, RDMA devices, and NPUs is only the basis. To obtain a higher speedup, the topological relationship between resources also needs to be considered during allocation. For example, when allocating multiple GPUs, try to allocate GPUs connected by NVlink or NVswitch; when allocating RDMA network cards, try to allocate RDMA network cards that are connected to the GPUs only through PCIe.

[0040] In related technologies, Kubernetes generally uses the Device Plugin mechanism for the management and allocation of GPU and RDMA resources. Refer to Figure 1 , the Device Plugin reports device information to the Kubelet. For example, the reported device information includes the device name, device Non Uniform Memory Access (NUMA) information, etc. The Kubelet reports the device information to the service provided by the Kubernetes cluster process (Kube-apiserver). The scheduler schedules to the corresponding node based on the device information in the kube-apiserver, and the kubelet on the node performs device allocation. After the kubelet allocates the devices, it sends an allocation request for the specified device to the Device Plugin. The Kubelet calls the Container Runtime Interface (CRI) based on the information returned by the Device Plugin, and the CRI starts and configures the container.

[0041] The Device Plugin can manage and allocate devices and is decoupled from Kubernetes, facilitating expansion and development. Common GPU and RDMA Device Plugins include the k8s DevicePlugin provided by NVIDIA to manage NVIDIA GPUs and the RDMA DevicePlugin to manage Mellanox RDMA.

[0042] In related technologies, the Container-Device-Interface (CDI) mechanism can also be used to allocate resources to containers. Refer to Figure 2 , Kubelet can call CRI to create a container; the CRI component reads the CDI json file at the specified path according to the information passed by Kubelet; CRI calls runc to create a container according to the information on the CDI json file.

[0043] Compared with the device plugin, CDI also supports the configuration of network devices, storage devices, etc., and has a wider range of usage scenarios. A common tool for using CDI to allocate GPUs can be the NVIDIA GPU container runtime nvidia-container-toolkit. Currently, as Containerd and the Device Plugin also start to support CDI, the CDI mechanism is gradually unified with the DevicePlugin mechanism.

[0044] The newly provided Dynamic Resource Allocation (DRA) mechanism in Kubernetes can also manage and allocate devices such as GPUs and RDMA. Refer to Figure 3 and Figure 4 , when creating a Pod, a Pod scheduling request will be generated and a corresponding ResourceClaim will be created. The ResourceClaim is used to represent the resources required by the Pod. kube-scheduler can perform node scheduling for the Pod by negotiating with the resource driver controller.

[0045] Exemplarily, the process of performing node scheduling for a Pod can include the following steps:

[0046] Step S1: kube-scheduler selects the optimal node according to the built-in scheduling rules and creates a PodScheduling resource, writing the selected optimal node to the SelectedNode of the PodScheduling resource. The scheduling rules of kube-scheduler include pre-selection rules and preference rules; the pre-selection rules represent the filtering (Filter) rules used to filter all nodes and select the nodes that meet the hard criteria. For example, through the pre-selection rules, nodes with sufficient CPU resources to run the Pod can be selected; the preference rules represent the rules for scoring and sorting the pre-selected nodes. For example, the nodes can be scored according to whether the images required by the Pod already exist on the nodes. After scoring and sorting the nodes, the node with the highest score can be selected.

[0047] Step S2: The resource driver controller reads the SelectedNode of the PodScheduling resource and queries the ResourceClaim corresponding to the Pod.

[0048] Step S3: If the SelectedNode does not meet the resource requirements (for example, the device resources of the SelectedNode are insufficient or the device model is not the required device model), the resource driver controller informs kube-scheduler through the DeallocationRequested message of the PodScheduling resource that the node needs to be re-selected, and writes the node that does not meet the resource requirements to the UnsuitableNodes field of the PodScheduling resource. Here, the resource requirements can be determined according to the content in the resource claim.

[0049] Step S4: If the SelectedNode meets the resource requirements, the resource driver controller then informs kube-scheduler through the Allocation message of the PodScheduling.

[0050] Step S5: When kube-scheduler determines that the SelectedNode does not meet the resource requirements, it returns to Step S1 and considers the UnsuitableNodes field of the PodScheduling resource when re-selecting the node.

[0051] Step S6: When kube-scheduler determines that SelectedNode meets the resource requirements, it executes the scheduling rules on SelectedNode again to ensure that SelectedNode still meets the resource requirements at this time (this step is to prevent the situation of SelectedNode from changing during the execution of steps S2 and S4 by the resource driver controller. For example, the CPU is occupied by other nodes, resulting in SelectedNode no longer meeting the resource requirements. If SelectedNode still meets the resource requirements at this time, then step S7 is executed; otherwise, it returns to step S1, and kube-scheduler selects a new node and continues to negotiate with the resource driver controller).

[0052] Step S7: kube-scheduler binds SelectedNode to the Pod to complete the node scheduling of the Pod.

[0053] Refer to Figure 4 , after the node scheduling of the Pod is completed, the kubelet of the corresponding node calls the Resource Driver Plugin to allocate devices and create containers.

[0054] Compared with the device plugin, CDI, and resource driver controller, the scheduling scheme adopting the DRA mechanism can participate in the scheduling of the cluster. Therefore, it can reasonably schedule containers according to device information. When using the device plugin or CDI for container scheduling, only kube-scheduler considers whether the number of devices meets the requirements. At the same time, in the scheduling schemes using DevicePlugin and CDI, which specific device to allocate is basically determined by kubelet, while in the scheduling scheme adopting the DRA mechanism in DRA, the allocation logic is handed over to the resource driver plugin, improving the flexibility of allocation.

[0055] In related technologies, nvidia-container-toolkit based on the CDI mechanism is only a device configuration tool for containers, lacking device management capabilities and having limited usage scenarios. Although k8s DevicePlugin and RDMADevicePlugin based on the Device Plugin mechanism can provide device configuration and management and are currently widely used solutions, they do not involve the scheduling level. Moreover, regarding which specific device to allocate, it can only be determined by kubelet. However, kubelet can only simply consider the NUMA topology and cannot further consider topology information such as nvlink and PCIe. Therefore, it is impossible to preferentially allocate GPUs connected by nvlink to containers to obtain higher training efficiency, nor can it allocate GPUs connected only by PCIe and RDMA to containers to achieve the GPU Direct RDMA function.

[0056] k8s-dra-driver based on the DRA mechanism can participate in the container scheduling decision and schedule the container to a node with a better topology (for example, when a container requests 2 GPUs, among nodes with the same remaining 2 GPU resources, a node with 2 GPUs connected by NVlink is better than a node with 2 GPUs connected by PCIe). Additionally, with the custom resource allocation logic of the same-resource driver plugin, it can preferentially select a group of GPUs with higher bandwidth on the same node for allocation. However, there are still the following two problems: First, it only manages a single resource and cannot meet scenarios that require coordinating multiple resources. For example, in the GPU Direct RDMA scenario, it is necessary to consider the topology affinity between the GPU and the RDMA device. Therefore, the DRA plugin for managing GPU resources needs to negotiate with the DRA plugin for managing RDMA resources, and the final resource allocation made by both should meet the topology affinity. However, there is no cooperation mechanism between one DRA plugin and another in the DRA mechanism. Second, the efficiency is too low. During the negotiation process between kube-scheduler and the resource driver controller, kube-scheduler can only make blind guesses. In the best case, kube-scheduler needs to calculate the scheduling rules 2 times. In non-optimal cases, kube-scheduler needs to try more times. For example, when SelectedNode does not meet the resource requirements, or when SelectedNode meets the resource requirements but due to changes in the node state, when kube-scheduler performs the second scheduling, it no longer meets the scheduling rules of kube-scheduler.

[0057] In view of the technical problems existing in the related art, the technical solution of the embodiment of the present application is proposed. The embodiment of the present application can solve the problem of low DRA scheduling efficiency, and at the same time retains the advantages that DRA can participate in scheduling and can customize resource allocation logic. By uniformly managing and topologically calculating various intelligent computing resources, the affinity between different intelligent computing resources is realized.

[0058] The following further details the embodiments of the present application in conjunction with the drawings and embodiments. It should be understood that the embodiments provided herein are only used to explain the embodiments of the present application and are not used to limit the embodiments of the present application. In addition, the embodiments provided below are partial embodiments for implementing the present application, rather than all embodiments for implementing the present application. Without conflict, the technical solutions described in the embodiments of the present application can be implemented in any combined manner.

[0059] It should be noted that in the embodiments of the present application, the terms "include", "comprise" or any other variant thereof are intended to cover non-exclusive inclusion, so that a method or device including a series of elements not only includes the elements clearly recited, but also includes other elements not explicitly listed, or further includes elements inherent to the implementation of the method or device. Without further limitation, an element defined by the statement "including one..." does not exclude the existence of other related elements in the method or device including the element (such as steps in the method or units in the device, and the unit can be part of a circuit, part of a processor, part of a program or software, etc.).

[0060] The node scheduling method for containers provided by the embodiments of the present application includes a series of steps, but the node scheduling method for containers provided by the embodiments of the present application is not limited to the recited steps. Similarly, the node scheduling device for containers provided by the embodiments of the present application includes a series of modules, but the device provided by the embodiments of the present application is not limited to including the explicitly recited modules, and may also include modules required for obtaining relevant information or processing based on the information.

[0061] Figure 5 It is a flowchart of the node scheduling method for containers of the scheduler applied to the Kubernetes cluster in the embodiments of the present application. As Figure 5 shown, the process includes:

[0062] Step 501: Determine the preselected nodes among the nodes based at least on the resource information in each node and the attributes that the resources requested when creating a container group Pod need to satisfy.

[0063] In the embodiments of the present application, each node may include at least one type of hardware resource. For example, the resources of each node include but are not limited to: GPU, CPU, RDMA device, NPU, etc.

[0064] Exemplarily, after the resource driver plugin of each node is started, it can scan the intelligent computing devices on this node to obtain device information, which is the information of the resources in the node. After obtaining the device information, the device information can be stored in kube-apiserver as a ResourcePerNode object. The information stored in ResourcePerNode includes but is not limited to device type, device model, device bandwidth, device characteristics, local device topology, whether the device instance is allocated, etc. The scheduler of the Kubernetes cluster can obtain the information of the resources in each node from kube-apiserver.

[0065] Exemplarily, the resource driver plugin can also query the monitoring system in the Kubernetes cluster, obtain the fault information of each device on this node, calculate the failure rate of each device on this machine, and record it in the ResourcePerNode object.

[0066] In some embodiments, the attributes that the resources requested when creating a container group Pod need to satisfy include but are not limited to: the number of required GPUs, the model of required GPUs, the size of GPU video memory, the characteristics of required GPUs, the number of required RDMA devices, the model of required RDMA devices, the characteristics of required RDMA devices, the number of required NPUs, the model of required NPUs, the characteristics of required NPUs, etc.

[0067] In practical applications, the resources requested when creating a container group Pod can be determined by creating a resource declaration; exemplarily, when obtaining the creation request of a Pod, a resource declaration can be created, and at least two resources requested when creating a container group Pod can be determined according to the resource declaration.

[0068] When a user creates a Pod, the scheduler of the Kubernetes cluster can obtain the scheduling request of the Pod. At this time, the scheduler of the Kubernetes cluster can create a corresponding resource declaration. Multiple resource categories can be set in the resource declaration, and at the same time, the attributes that each requested resource needs to satisfy are set through ParametersRef. It can be seen that since at least two resources requested when creating a container group Pod can be determined according to the resource declaration, therefore, the embodiments of the present application can more accurately and reasonably determine the preselected nodes on the basis of considering the attributes that at least one or two resources need to satisfy, which is beneficial to accurately selecting the scheduling node of the Pod in the preselected nodes subsequently.

[0069] In some embodiments, the process of determining preselected nodes among the various nodes based on at least the information of the resources in each node and the attributes that the resources requested when creating a Pod need to satisfy may include: obtaining the filtering rules for each type of resource preset; determining the preselected nodes among the various nodes based on the filtering rules for each type of resource, the information of the resources in each node, and the attributes that the resources requested when creating a container group Pod need to satisfy.

[0070] Here, the filtering rules for each type of resource may be judgment statements organized using logical operations. Exemplarily, the logical operations may be OR (or) operations, AND (and) operations, and the judgment statements include but are not limited to equal, include, intersection, exclude, greater, less, etc.

[0071] Exemplarily, a resource driver controller may create the filtering rules for each type of resource, and the filtering rules may use the YAML format.

[0072] Referring to Figure 6 , after receiving a scheduling request for a Pod, kube - scheduler no longer negotiates with the resource driver controller, but is assisted by the resource driver controller for scheduling. Exemplarily, kube - scheduler may obtain the ResourcePerNode object and the filtering rules created by the resource driver controller, and then determine the preselected nodes among the various nodes through the pre - selection rules. For example, for each type of resource in a node, kube - scheduler will additionally filter the various nodes through the resource declaration, the filtering rules corresponding to the resources, and the ResourcePerNode object of the node, so as to obtain the preselected nodes.

[0073] In an example, if the ParametersRef in the resource declaration sets the GPU video memory size to memoryValue, devices with a video memory value greater than memoryValue can be selected according to the ResourcePerNode object; if the resource declaration sets a specified function or a specified model of device, the corresponding device can be selected according to the ResourcePerNode object. After selecting devices that meet the above conditions, if the number of devices selected in the node is less than the number requested in the resource declaration, this node is eliminated; otherwise, if the number of devices selected in the node is greater than or equal to the number requested in the resource declaration, the node is determined as one of the preselected nodes.

[0074] It can be seen that in the embodiments of the present application, the preselected nodes can be more accurately determined among each node according to the filtering rules of each resource.

[0075] Step 502: Send the identifiers of the preselected nodes to the controller, so that the controller determines the first fitness score of each node among the preselected nodes with the scheduling node of the Pod.

[0076] The controller in this step can be the resource-driven controller described above. After sending the identifiers of the preselected nodes to the controller, the controller can score each node among the preselected nodes to obtain the first fitness score of each node among the preselected nodes with the scheduling node of the Pod; the controller can return the first fitness score to the scheduler of the Kubernetes cluster.

[0077] Step 503: Receive the first fitness score sent by the controller, and select the scheduling node of the Pod among the preselected nodes based on the second fitness score of each node among the preselected nodes determined by the scheduler and the first fitness score.

[0078] Refer to Figure 6 , kube-scheduler can select the scheduling node of the Pod among the preselected nodes according to the preference rules after determining the preselected nodes. Exemplarily, kube-scheduler can perform a weighted sum of the second fitness score determined by itself and the first fitness score sent by the resource-driven controller, so as to determine the comprehensive score of each node among the preselected nodes, and select the node with the highest comprehensive score among the preselected nodes as the scheduling node of the Pod.

[0079] In some embodiments, kube-scheduler can determine the second fitness score of each node with the scheduling node of the Pod according to the running state data of each node. Here, the running state data includes but is not limited to: the failure rate of each node, the matching degree of each node with the task characteristics of the Pod, etc. Exemplarily, kube-scheduler can calculate the failure rate of each node. The higher the failure rate, the lower the second fitness score of the node; kube-scheduler can also determine whether the mirror image required by the Pod already exists on the node. If the mirror image required by the Pod exists on the node, the second fitness score of the node is relatively high, otherwise, the second fitness score of the node is relatively low.

[0080] In some embodiments, kube-scheduler notifies the resource driver controller of the scheduling node of the Pod, and the resource driver controller can write the resource allocation scheme (i.e., the allocation result of the device) in the scheduling node of the Pod into the resource declaration. After completing the node scheduling of the Pod, the kubelet of the corresponding node calls the resource driver plugin to allocate resources (devices), and the resource driver plugin can perform the allocation operation according to the device allocation result written by the resource driver controller in the resource declaration.

[0081] In practical applications, steps 501 to 503 can be implemented based on a processor, and the above-mentioned processor can be at least one of an application specific integrated circuit (ASIC), a digital signal processor (DSP), a digital signal processing device (DSPD), a programmable logic device (PLD), a field programmable gate array (FPGA), a CPU, a controller, a microcontroller, and a microprocessor.

[0082] In the related art, the scheduling node can be determined through the mutual negotiation between kube-scheduler and the resource driver controller. In this process, both kube-scheduler and the resource driver controller have some rigid filtering rules and can veto scheduling to a certain node. Since kube-scheduler is the party that finally performs the scheduling operation (i.e., binds the Pod to the node), kube-scheduler needs to try to select a node that the resource driver controller does not reject and perform the scheduling operation, which is likely to result in multiple scheduling situations. In the embodiments of the present application, the scheduler of the Kubernetes cluster can select a node that meets the filtering rules at one time without negotiating with the controller, that is, determine the preselected node. The scheduler of the Kubernetes cluster does not need to determine the node that the controller does not reject through multiple blind guesses. The controller only scores the preselected nodes for reference by the scheduler of the Kubernetes cluster. Finally, the scheduler of the Kubernetes cluster can determine the scheduling node of the Pod at one time according to the first fitness score and the second fitness score, improving the scheduling efficiency.

[0083] Figure 7 It is a flowchart of the node scheduling method for the container applied to the controller in the embodiments of the present application, as Figure 7As shown in the figure, the process includes:

[0084] Step 701: Receive the identifiers of the preselected nodes sent by the scheduler of the Kubernetes cluster. The scheduler is used to determine the preselected nodes among all nodes at least based on the resource information in each node and the attributes that the resources requested when creating a Pod need to satisfy.

[0085] Step 702: Determine the first fitness score of each node among the preselected nodes with the scheduling node of the Pod.

[0086] Step 703: Send the first fitness score to the scheduler, so that the scheduler can select the scheduling node of the Pod among the preselected nodes according to the second fitness score of each node among the preselected nodes determined by itself and the first fitness score.

[0087] In practical applications, steps 701 to 703 can be implemented based on a processor, and the above processor can be at least one of ASIC, DSP, DSPD, PLD, FPGA, CPU, controller, microcontroller, and microprocessor.

[0088] It can be seen that in the embodiments of this application, the scheduler of the Kubernetes cluster does not need to determine the nodes that the controller does not reject through multiple blind guesses. The controller only scores the preselected nodes for the reference of the scheduler of the Kubernetes cluster. Finally, the scheduler of the Kubernetes cluster can determine the scheduling node of the Pod at one time according to the first fitness score and the second fitness score, improving the scheduling efficiency.

[0089] In some embodiments of this application, the process of determining the first fitness score of each node among the preselected nodes with the scheduling node of the Pod may include: determining the resource allocation scheme of each node at least according to the topological relationship between different types of resources in each node; and determining the first fitness score of each node with the scheduling node of the Pod according to the resource allocation scheme of each node.

[0090] In the scoring logic of the controller, the topological relationship between multiple resources will be considered. Exemplarily, in the implementation of the optimal allocation scheme of GPU and RDMA devices, select GPU and RDMA according to the following process:

[0091] a. Select GPU devices that meet the filtering rules as the GPUs to be allocated in the following order of b-g until the number of GPUs to be allocated meets the quantity requirements in the resource declaration.

[0092] b. When there is an NVSwitch device, select as many GPUs as possible from under one NVSwitch device as the GPUs to be allocated.

[0093] c. Select the GPUs connected to the currently to-be-allocated GPU device via NVLink and add them to the set of to-be-allocated GPUs.

[0094] d. Select the GPUs connected to the currently to-be-allocated GPU device via PCIe and add them to the set of to-be-allocated GPUs.

[0095] e. Select the GPUs under the same CPU as the currently to-be-allocated GPU device and add them to the set of to-be-allocated GPUs.

[0096] f. Select the GPUs under the same NUMA as the currently to-be-allocated GPU device and add them to the set of to-be-allocated GPUs.

[0097] g. Select other GPU devices on the node and add them to the set of to-be-allocated GPUs

[0098] h. Select RDMA devices based on the to-be-allocated GPU devices, and select the RDMA devices that meet the filtering rules as the to-be-allocated RDMA in the following order of i - m until the number of to-be-allocated RDMA meets the quantity requirements in the resource declaration.

[0099] i. Select the RDMA devices connected to the to-be-allocated GPU devices via NVLink and add them to the set of to-be-allocated RDMA.

[0100] j. Select the RDMA devices connected to the to-be-allocated GPU devices via PCIe and add them to the set of to-be-allocated RDMA.

[0101] k. Select the RDMA devices under the same CPU as the to-be-allocated GPU devices and add them to the set of to-be-allocated RDMA.

[0102] l. Select the RDMA devices under the same NUMA as the to-be-allocated GPU devices and add them to the set of to-be-allocated RDMA.

[0103] m. Select other RDMA devices on the node and add them to the set of to-be-allocated RDMA

[0104] Exemplarily, for the implementation of the optimal allocation scheme of NPU devices, select NPUs according to the following process:

[0105] a. Select the NPU devices that meet the filtering rules as the to-be-allocated NPUs in the following order of b - f until the number of to-be-allocated NPU devices meets the quantity requirements in the resource declaration.

[0106] b. Select the NPU devices connected to the currently to-be-allocated NPU device via HCCN (Huawei Cache Coherence Network) and add them to the set of to-be-allocated NPUs.

[0107] c. Select the NPU devices connected to the currently to-be-allocated NPU device via PCIe and add them to the set of to-be-allocated NPU devices.

[0108] d. Select the NPU devices under the same CPU as the currently to-be-allocated NPU device and add them to the set of to-be-allocated NPU devices.

[0109] e. Select the NPU devices under the same NUMA as the currently to-be-allocated NPU device and add them to the set of to-be-allocated NPU devices.

[0110] f. Select other NPU devices on the node and add them to the set of to-be-allocated NPU devices.

[0111] It can be seen that since the resource allocation scheme for each node can be determined according to the topological relationship between different types of resources in each node, it is beneficial to realize the reasonable allocation of node resources by coordinating multiple types of resources.

[0112] In some embodiments of the present application, determining the first fitness score between each node and the scheduling node of the Pod according to the resource allocation scheme of each node may include:

[0113] Determine the resource attribute data corresponding to the resource allocation scheme of each node, and the resource attribute data includes at least one of the following: the failure rate of each resource in each node, the minimum bandwidth between different resources in each node, and the matching degree between each resource in each node and the task characteristics of the Pod;

[0114] Determine the first fitness score between each node and the scheduling node of the Pod according to the resource attribute data corresponding to the resource allocation scheme of each node.

[0115] Exemplarily, after generating the optimal allocation scheme for GPUs, RDMA, or NPUs, score the nodes according to the set of to-be-allocated devices, and the scoring rules are as follows:

[0116] a. Calculate the minimum bandwidth between two devices in the set of to-be-allocated devices. The larger the minimum bandwidth, the higher the score, and vice versa.

[0117] b. Score according to task characteristic matching. When the Pod applying for resources is a training Pod, calculate the total video memory of the GPUs or NPUs in the set of to-be-allocated devices. The higher the total video memory, the higher the score. When the Pod applying for resources is an inference Pod, calculate the total floating-point operation ability (Floating Point Operations Per Second, FLOPS) of the GPUs or NPUs in the set of to-be-allocated devices. The higher the total floating-point operation ability, the higher the score.

[0118] c. Score according to the device failure rate recorded in ResourcesPerNode. Calculate the total failure rate of all devices in the set to be allocated (the sum of the failure rates of all devices). The higher the total failure rate, the lower the score.

[0119] After the controller determines the first fitness score, it can return the first fitness score to the scheduler of the Kubernetes cluster. At the same time, record the optimal allocation plan for each node in the resource declaration for the subsequent device allocation by the resource driver plugin.

[0120] In summary, the embodiments of the present application can achieve unified management and unified topology calculation of various intelligent computing resources. Different from the related technical solutions, the embodiments of the present application support Pods to apply for multiple resources through a single resource declaration. That is, in an instance of a resource declaration, a Pod can simultaneously apply for RDMA devices and GPU resources, which is mainly achieved by modifying the relationship between the resource declaration and the resource type from 1:1 to 1:many. At the preselected nodes during scheduling, nodes can be filtered according to different resources. At the selected nodes during scheduling, the topological relationship between multiple resources can be considered for scoring.

[0121] The embodiments of the present application improve the DRA mechanism in the related art, solving the problem that kube-scheduler needs to coordinate with the resource driver controller multiple times. In the embodiments of the present application, the resource driver controller only responsible for informing kube-scheduler how to filter nodes at the preselected nodes and how many points to add at the selected stage, and will not veto the scheduling result of kube-scheduler. At the same time, the embodiments of the present application support the unified management of multiple resources, and can perform unified topology calculation on various intelligent computing resources during the selected scoring.

[0122] The embodiments of the present application propose an improved DRA intelligent computing resource allocation process to improve the scheduling efficiency of DRA. The process includes: informing the preselection rules of kube-scheduler through filtering rules; providing additional selected scoring items for kube-scheduler through the resource driver controller. The embodiments of the present application can solve the problem of coordinated allocation between different types of resources and perform unified topology calculation on various intelligent computing resources. In the embodiments of the present application, multiple resource types can be applied for through resource declarations; nodes can be filtered according to the filtering rules of each resource; the resource driver controller generates an allocation set containing multiple resources according to certain rules and scores.

[0123] Compared with the k8s Device Plugin and RDMA Device Plugin implemented based on drive plugins, the embodiments of the present application have better scheduling capabilities and the ability to coordinate multiple types of resources for optimal allocation. Compared with the k8s-dra-driver based on the DRA mechanism, the embodiments of the present application do not require multiple negotiations and are more efficient; they can not only manage one type of resource, but also uniformly manage various intelligent computing resources, and coordinate during the allocation process of various resources, which is beneficial to improving the performance of intelligent computing.

[0124] Those skilled in the art can understand that in the above method of the specific implementation manner, the writing order of each step does not mean a strict execution order that constitutes any limitation on the implementation process, and the specific execution order of each step should be determined according to its function and possible internal logic.

[0125] Figure 8 It is a schematic structural diagram of a node scheduling device for a container of a scheduler applied to a Kubernetes cluster according to an embodiment of the present application, as Figure 8 shown, the device includes:

[0126] A first processing module 801, configured to determine preselected nodes among the various nodes at least based on the information of resources in each node and the attributes that the resources requested when creating a Pod need to satisfy.

[0127] A second processing module 802, configured to send the identifiers of the preselected nodes to a controller, so that the controller determines a first fitness score of each node among the preselected nodes with respect to the scheduling node of the Pod; receive the first fitness score sent by the controller, and based on the second fitness score of each node among the preselected nodes determined by the scheduler with respect to the scheduling node of the Pod and the first fitness score, select the scheduling node of the Pod among the preselected nodes.

[0128] In some embodiments, the first processing module 801 is configured to determine preselected nodes among the various nodes at least based on the information of resources in each node and the attributes that the resources requested when creating a Pod need to satisfy, including:

[0129] Obtain preset filtering rules for each type of resource;

[0130] Based on the filtering rules for each type of resource, the information of resources in each node, and the attributes that the resources requested when creating a Pod need to satisfy, determine the preselected nodes among the various nodes.

[0131] In some embodiments, the filtering rules for each type of resource are judgment statements organized using logical operations.

[0132] In some embodiments, the first processing module 801 is further configured to create a resource declaration when obtaining a scheduling request for a Pod, and determine at least two types of resources requested when creating the Pod according to the resource declaration, before determining preselected nodes among the various nodes, based at least on information about resources in each node and attributes that the resources requested when creating the Pod need to satisfy.

[0133] In practical applications, the first processing module 801 and the second processing module 802 may be implemented based on a processor and a communication device.

[0134] Figure 9 The following is a schematic structural diagram of a node scheduling device for a container applied to a controller according to an embodiment of the present application, as Figure 9 shown, the device includes:

[0135] A receiving module 901, configured to receive an identifier of a preselected node sent by a scheduler of a Kubernetes cluster, where the scheduler is configured to determine preselected nodes among the various nodes based at least on information about resources in each node and attributes that the resources requested when creating a container group Pod need to satisfy;

[0136] A third processing module 902, configured to determine a first fitness score of each node in the preselected nodes with respect to a scheduling node of the Pod; and send the first fitness score to the scheduler, so that the scheduler selects a scheduling node of the Pod from the preselected nodes according to a second fitness score of each node in the preselected nodes determined by itself and the first fitness score.

[0137] In some embodiments, the third processing module 902 is configured to determine a first fitness score of each node in the preselected nodes with respect to a scheduling node of the Pod, including:

[0138] Determining a resource allocation scheme for each node at least according to a topological relationship between different types of resources in each node;

[0139] Determining a first fitness score of each node with respect to a scheduling node of the Pod according to the resource allocation scheme of each node.

[0140] In some embodiments, the third processing module 902 is configured to determine a first fitness score of each node with respect to a scheduling node of the Pod according to the resource allocation scheme of each node, including:

[0141] Determine resource attribute data corresponding to the resource allocation scheme for each of the nodes, where the resource attribute data includes at least one of the following: the failure rate of each resource in each node, the minimum bandwidth between different resources in each node, and the degree of matching between each resource in each node and the task characteristics of the Pod;

[0142] Determine a first fitness score of each of the nodes and the scheduling node of the Pod according to the resource attribute data corresponding to the resource allocation scheme for each of the nodes.

[0143] In practical applications, the receiving module 901 and the third processing module 902 can be implemented based on a processor and a communication device.

[0144] It should be noted that the description of the above device embodiments is similar to the description of the above method embodiments and has beneficial effects similar to those of the method embodiments. For the technical details not disclosed in the device embodiments of the present application, please refer to the description of the method embodiments of the present application for understanding.

[0145] It should be noted that in the embodiments of the present application, if the above method is implemented in the form of software function modules and sold or used as an independent product, it can also be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the embodiments of the present application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a terminal, a server, etc.) to execute all or part of the methods described in the embodiments of the present application. The foregoing storage medium includes: various media such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a magnetic disk, or an optical disc that can store program codes. In this way, the embodiments of the present application are not limited to any specific combination of hardware and software.

[0146] Correspondingly, the embodiments of the present application further provide a computer program product, where the computer program product includes computer-executable instructions for implementing any node scheduling method for a container provided by the embodiments of the present application.

[0147] Correspondingly, the embodiments of the present application further provide a computer storage medium, where computer-executable instructions are stored on the computer storage medium for implementing any node scheduling method for a container provided by the above embodiments.

[0148] The embodiments of the present application further provide an electronic device. Figure 10 Shown in the following is a schematic structural diagram of an electronic device provided by an embodiment of the present application, as Figure 10 shown, the electronic device 100 may include:

[0149] A memory 1001 for storing executable instructions;

[0150] A processor 1002 for implementing the node scheduling method of any of the above containers when executing the executable instructions stored in the memory 1001.

[0151] The above-mentioned processor 1002 can be at least one of ASIC, DSP, DSPD, PLD, FPGA, CPU, controller, microcontroller, and microprocessor.

[0152] The above-mentioned computer-readable storage medium and memory 1002 can be a read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), ferromagnetic random access memory (FRAM), flash memory, magnetic surface memory, optical disc, or compact disc read-only memory (CD-ROM), etc.; or various terminals including one or any combination of the above memories, such as mobile phones, computers, tablet devices, personal digital assistants, etc.

[0153] In some embodiments, the functions or modules included in the device provided by the embodiments of the present application can be used to execute the methods described in the above method embodiments. The specific implementation can refer to the description of the above method embodiments. For the sake of brevity, it will not be repeated here.

[0154] The descriptions of the above embodiments tend to emphasize the differences between the embodiments. Their similarities or similarities can be referred to each other. For the sake of brevity, they will not be repeated in this article.

[0155] The methods disclosed in the method embodiments provided by the present application can be arbitrarily combined without conflict to obtain new method embodiments.

[0156] The features disclosed in the product embodiments provided by the present application can be arbitrarily combined without conflict to obtain new product embodiments.

[0157] The features disclosed in the method or device embodiments provided in this application can be arbitrarily combined without conflict to obtain new method embodiments or device embodiments.

[0158] Through the description of the above embodiments, those skilled in the art can clearly understand that the above embodiment methods can be implemented by means of software plus a necessary general hardware platform. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on such an understanding, the technical solution of this application, in essence, or the part that contributes to the prior art can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) and includes several instructions for causing a terminal (which can be a mobile phone, computer, server, air conditioner, or network device, etc.) to execute the methods described in the various embodiments of this application.

[0159] The embodiments of this application have been described above in conjunction with the accompanying drawings. However, this application is not limited to the above specific implementation manners. The above specific implementation manners are merely illustrative and not restrictive. Under the inspiration of this application, those of ordinary skill in the art can also make many forms without departing from the purpose of this application and the scope protected by the claims. All of these are within the protection scope of this application.

Claims

1. A node scheduling method for a container, characterized in that: Applied to a scheduler of a Kubernetes cluster, the method includes: Determine a preselected node from among the nodes based at least on information about resources in each node and attributes that the resources requested when creating a container group Pod need to satisfy; Sending the identifier of the pre-selected node to the controller, so that the controller determines a first fitness score of each of the pre-selected nodes and the scheduling node of the Pod; Receive a first fitness score sent by the controller, and select a scheduling node for the Pod from among the preselected nodes based on a second fitness score between each of the preselected nodes and the scheduling node for the Pod determined by the scheduler and the first fitness score.

2. The method according to claim 1, characterized in that The determining of the pre-selected nodes from among the nodes based at least on the information of the resources in the nodes and the attributes that the resources requested when creating the Pod need to satisfy includes: Get the preset filtering rules for each resource; Based on the filtering rules of each resource, the information of the resources in each node, and the attributes that the resources requested when creating a Pod need to satisfy, the pre-selected node is determined among the nodes.

3. The method according to claim 2, characterized in that The filtering rule for each resource is a judgment statement organized by logical operations.

4. The method according to claim 1, characterized in that: Based at least on the information of resources in each node and the attributes that the resources requested when creating the Pod need to satisfy, before determining the pre-selected nodes in the nodes, the method further includes: When obtaining a scheduling request for a Pod, a resource declaration is created, and at least two resources requested when creating the Pod are determined according to the resource declaration.

5. A node scheduling method for a container, characterized in that: Applied in a controller, the method comprises: Receiving an identifier of a pre-selected node sent by a scheduler of the Kubernetes cluster, wherein the scheduler is used to determine the pre-selected node among the nodes based on at least information about resources in each node and attributes that the resources requested when creating a container group Pod need to satisfy; Determine a first fitness score for each of the pre-selected nodes and the scheduling node of the Pod; The first fitness score is sent to the scheduler, so that the scheduler selects the scheduling node of the Pod from the pre-selected nodes according to the second fitness score between each node in the pre-selected nodes and the scheduling node of the Pod determined by the scheduler itself and the first fitness score.

6. The method according to claim 5, characterized in that Determining a first fitness score between each of the pre-selected nodes and the scheduling node of the Pod includes: Determining a resource allocation scheme for each node at least according to a topological relationship between different types of resources in each node; According to the resource allocation scheme of each node, a first fitness score between each node and the scheduling node of the Pod is determined.

7. The method according to claim 6, characterized in that Determining a first fitness score between each node and the scheduling node of the Pod according to the resource allocation scheme of each node includes: Determine resource attribute data corresponding to the resource allocation scheme of each node, the resource attribute data including at least one of the following: a failure rate of each resource in each node, a minimum bandwidth between different resources in each node, and a matching degree between each resource in each node and a task characteristic of the Pod; Determine a first fitness score between each node and a scheduling node of the Pod according to resource attribute data corresponding to the resource allocation scheme of each node.

8. A node scheduling device for a container, characterized in that: Applied to the scheduler of the Kubernetes cluster, the device includes: A first processing module is used to determine a pre-selected node from each node based on at least information about resources in each node and attributes that the resources requested when creating a container group Pod need to satisfy; The second processing module is used to send the identifier of the pre-selected node to the controller, so that the controller determines a first fitness score between each node in the pre-selected nodes and the scheduling node of the Pod; receive the first fitness score sent by the controller, and select the scheduling node of the Pod from the pre-selected nodes based on the second fitness score between each node in the pre-selected nodes and the scheduling node of the Pod determined by the scheduler and the first fitness score.

9. A node scheduling device for a container, characterized in that: Applied in a controller, the device comprises: A receiving module, configured to receive an identifier of a pre-selected node sent by a scheduler of a Kubernetes cluster, wherein the scheduler is configured to determine the pre-selected node from among the nodes based at least on information about resources in each node and properties that the resources requested when creating a container group Pod need to satisfy; The third processing module is used to determine a first fitness score between each of the pre-selected nodes and the scheduling node of the Pod; send the first fitness score to the scheduler, so that the scheduler selects the scheduling node of the Pod from the pre-selected nodes according to the second fitness score between each of the pre-selected nodes and the scheduling node of the Pod determined by itself and the first fitness score.

10. An electronic device, characterized in that: The electronic device comprises a processor and a memory for storing a computer program that can be run on the processor; wherein, The processor is configured to run the computer program to perform the method according to any one of claims 1 to 7.

11. A computer storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the method according to any one of claims 1 to 7 is implemented.

12. A computer program product, comprising a computer program, characterized in that The computer program implements the method according to any one of claims 1 to 7 when executed by a processor.