Inference service-oriented task migration method and device and storage medium

By using node labels and preset task migration rules to match task resources in the GPU cluster, the GPU fragmentation problem is solved, tasks are deployed on graphics card nodes with matching resources, and GPU resource utilization is improved.

CN120704857APending Publication Date: 2025-09-26ZHEJIANG DAHUA TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510592169.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-08
Publication Date
2025-09-26

AI Technical Summary

Technical Problem

Existing methods in GPU resource management fail to effectively deal with the fragmentation problem caused by dynamic changes in cluster resources, especially when the GPU allocation rate is high and they are unable to handle large-scale tasks.

Method used

Determine candidate graphics card nodes through node labels, match task resource requests with graphics card resources according to preset task migration rules, implement task migration to organize resource fragments, and use rolling migration to ensure service availability.

Benefits of technology

Effectively organize graphics card node resource fragments, ensure that tasks are deployed in nodes with matching resource amounts, improve GPU resource utilization, and reduce fragmentation problems.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120704857A_ABST
    Figure CN120704857A_ABST
Patent Text Reader

Abstract

The invention discloses a reasoning service-oriented task migration method and device and a storage medium. The method comprises the following steps: determining candidate graphics card nodes in graphics card nodes of a current graphics card cluster according to node labels; the node label is used for representing the graphics card resource quantity of the graphics card node; determining a to-be-migrated task in each candidate graphics card node according to a preset task migration rule; matching the resource application quantity of each to-be-migrated task with the graphics card resource quantity of each candidate graphics card node to obtain a target task and a target graphics card node which are successfully matched; and migrating the target task into the target graphics card node, thereby realizing resource fragmentation of the graphics card node through task migration.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of cloud native technology, and in particular to a task migration method, device, and storage medium for reasoning services. Background Art

[0002] With the rapid development of artificial intelligence technology and the continuous increase in the parameter size of neural network models, the demand for graphics processing units (GPUs) has increased dramatically, driving the large-scale construction of GPU computing clusters. Therefore, effective management of GPU resources is of great significance.

[0003] Kubernetes, as a container orchestration standard, effectively manages the software and hardware resources that computing relies on. Kubernetes-based cloud-native technologies can efficiently support GPU resource scheduling through containerized applications.

[0004] However, existing methods don't account for the dynamic changes in cluster resources during operation. As inference services are continuously deployed, upgraded, and adjusted, cluster GPU resources inevitably become fragmented. When the GPU allocation rate is high, even though the entire cluster has a large amount of remaining GPU resources, excessive fragmentation can make it impossible to handle large-scale tasks. Summary of the Invention

[0005] This application at least provides a task migration method, apparatus, device, and computer-readable storage medium for reasoning services.

[0006] The first aspect of the present application provides a task migration method for inference services, including: determining a candidate graphics card node among the graphics card nodes of the current graphics card cluster based on a node label; the node label is used to characterize the amount of graphics card resources possessed by the graphics card node; determining the tasks to be migrated in each candidate graphics card node according to a preset task migration rule; matching the resource application amount of each task to be migrated with the graphics card resource amount of each candidate graphics card node to obtain a successfully matched target task and target graphics card node; and migrating the target task to the target graphics card node.

[0007] In one embodiment, determining a candidate graphics card node from among the graphics card nodes in the current graphics card cluster based on the node label includes: obtaining a resource allocation rate of each labeled node in the current graphics card cluster; the labeled node is a graphics card node in the current graphics card cluster that has the node label; and determining a labeled node whose resource allocation rate is less than a resource allocation threshold as the candidate graphics card node.

[0008] In one embodiment, determining the labeled node whose resource allocation rate is less than the resource allocation threshold as the candidate graphics card node includes: sorting the labeled nodes whose resource allocation rate is less than the resource allocation threshold according to the resource allocation rate to obtain the candidate graphics card node.

[0009] In one embodiment, the method of determining the tasks to be migrated in each candidate graphics card node according to the preset task migration rules includes: traversing each candidate graphics card node, and performing rule comparison with the task attributes of the current node task in each candidate graphics card node according to the preset task migration rules; and determining the current node task that fails the comparison as the task to be migrated.

[0010] In one embodiment, the preset task migration rules are compared with the task attributes of the current node task in each candidate graphics card node, including: judging whether the current node task is at least one of a static task, a mirror task, and a daemon task based on the task attributes; if not, performing a consistency comparison based on the resource application amount of the current node task and the graphics card resource amount of the graphics card node where the current node task is located; in response to the resource application amount of the current node task being inconsistent with the graphics card resource amount of the graphics card node where the current node task is located, determining that the comparison between the preset task migration rules and the current node task has failed.

[0011] In one embodiment, the resource application amount of each task to be migrated is matched with the graphics card resource amount of each candidate graphics card node to obtain a successfully matched target task and target graphics card node, including: traversing each task to be migrated, and determining a matching graphics card node that matches the currently traversed task to be migrated from the candidate graphics card nodes based on the resource application amount and the graphics card resource amount; and determining a matching graphics card node with idle resources and the currently traversed task to be migrated as the target graphics card node and the target task.

[0012] In one embodiment, migrating the target task to the target graphics card node includes: performing task creation processing in the target graphics card node according to the target task to obtain a new task in the target graphics card node; and deleting the target task.

[0013] In one embodiment, before determining the candidate graphics card node from among the graphics card nodes in the current graphics card cluster based on the node label, the method further includes: determining the corresponding initial node label based on the resource requirements of the received initial scheduling task; determining each initial graphics card node in the current graphics card cluster that meets the resource requirements to obtain an initial node list; in response to the existence of an initial graphics card node corresponding to the initial node label in the initial node list, scheduling the initial scheduling task to the initial graphics card node corresponding to the initial node label; in response to the absence of an initial graphics card node corresponding to the initial node label in the initial node list, scheduling the initial scheduling task to an unlabeled node in the initial node list, and performing label setting processing on the unlabeled node based on the initial node label to obtain a labeled node.

[0014] The second aspect of the present application provides a task migration device for inference services, including: a candidate node determination module, used to determine a candidate graphics card node among the graphics card nodes of the current graphics card cluster according to the node label; the node label is used to characterize the amount of graphics card resources possessed by the graphics card node; a task to be migrated determination module, used to determine the tasks to be migrated in each candidate graphics card node according to a preset task migration rule; a task and node matching module, used to match the resource application amount of each task to be migrated with the graphics card resource amount of each candidate graphics card node, and obtain a successfully matched target task and target graphics card node; a task migration module, used to migrate the target task to the target graphics card node.

[0015] A third aspect of the present application provides an electronic device, including a memory and a processor, wherein the processor is used to execute program instructions stored in the memory to implement the above-mentioned task migration method for reasoning services.

[0016] In a fourth aspect, the present application provides a computer-readable storage medium having program instructions stored thereon, which implement the above-mentioned task migration method for reasoning services when the program instructions are executed by a processor.

[0017] The above scheme determines the candidate graphics card node among the graphics card nodes of the current graphics card cluster based on the node label; the node label represents the amount of graphics card resources possessed by the graphics card node; the preset task migration rule includes a method for determining whether there are unmatched tasks in the candidate graphics card node, so the tasks to be migrated in each candidate graphics card node can be determined according to the preset task migration rule; the resource application amount of each task to be migrated is matched with the graphics card resource amount of each candidate graphics card node respectively, so as to re-search each candidate graphics card node suitable for processing each task to be migrated, and obtain the target task and target graphics card node that are successfully matched; the target task is migrated to the target graphics card node, so that each task can be deployed in the graphics card node with matching resource volume through task migration, thereby realizing resource fragmentation of the graphics card node.

[0018] It is to be understood that the foregoing general description and the following detailed description are exemplary and explanatory only and are not restrictive of the present application. BRIEF DESCRIPTION OF THE DRAWINGS

[0019] The drawings herein are incorporated into and constitute a part of the specification. These drawings illustrate embodiments consistent with the present application and, together with the specification, are used to illustrate the technical solutions of the present application.

[0020] Figure 1 This is a flowchart of an exemplary embodiment of the inference service-oriented task migration method of the present application;

[0021] Figure 2 This is an exemplary architectural diagram of the inference service-oriented task migration method of this application;

[0022] Figure 3 This is a schematic diagram of an exemplary task migration scenario in the inference service-oriented task migration method of this application;

[0023] Figure 4 This is a flowchart of an exemplary card specification matching scheduling in the inference service-oriented task migration method of the present application;

[0024] Figure 5 This is an exemplary resource transfer diagram in the reasoning service-oriented task migration method of this application;

[0025] Figure 6 is a block diagram of a task migration device for reasoning services shown in an exemplary embodiment of the present application;

[0026] Figure 7 This is a structural diagram of an embodiment of an electronic device of the present application;

[0027] Figure 8 It is a structural diagram of an embodiment of a computer-readable storage medium of the present application. DETAILED DESCRIPTION

[0028] The following describes the embodiments of the present application in detail with reference to the accompanying drawings.

[0029] In the following description, for the purpose of explanation rather than limitation, specific details such as specific system structures, interfaces, and technologies are provided to facilitate a thorough understanding of the present application.

[0030] The term "and / or" in this article is simply a description of the association relationship of associated objects, indicating that three relationships can exist. For example, A and / or B can represent three situations: A exists alone, A and B exist at the same time, and B exists alone. In addition, the character " / " in this article generally indicates that the associated objects are in an "or" relationship. In addition, "many" in this article means two or more than two. In addition, the term "at least one" in this article means any combination of at least two of any one or more of a plurality of. For example, including at least one of A, B, and C can mean including any one or more elements selected from the set consisting of A, B, and C.

[0031] To facilitate understanding, the existing background technology involved in this application is now described as an example. With the rapid development of artificial intelligence technology, especially the rise of large-scale model technology, the parameter scale and multimodal capabilities of neural network models have continued to increase, and the demand for GPUs has increased dramatically, driving a large-scale construction wave of GPU computing clusters. Therefore, it is of great significance to effectively manage GPU resources, realize on-demand allocation of GPU resources, and improve the utilization rate of GPU resources.

[0032] Kubernetes, a container orchestration standard, effectively manages the software and hardware resources that computing tasks rely on. Kubernetes-based cloud-native technologies can efficiently support GPU resource scheduling through containerized applications.

[0033] However, existing scheduling algorithms match policy rules based on the current resource allocation in the cluster, failing to consider the dynamic changes in cluster resources during actual operation. As inference services are continuously deployed, upgraded, and adjusted, cluster GPU resources can become fragmented. Specifically, when the GPU allocation rate is high, while the entire GPU cluster may have a large amount of remaining GPU resources, multiple GPU nodes may only have a small number of remaining GPU cards, preventing the allocation of large resources.

[0034] Compared with training tasks, AI inference services also have some differences in operating characteristics. Training tasks are similar to batch tasks. After running for several hours or days, the resources are released when the task ends. Inference services are equivalent to online services, which usually run online for a long time and have requirements for service response latency. Inference services can have one or more pods (pod is the smallest unit that can be created (deployed) and managed in the k8s system), and inference services are constrained by neural network models. Under normal circumstances, the amount of GPU resources required to process pod tasks is generally 1 card (1 GPU card), 2 cards (2 GPU cards), 4 cards (4 GPU cards) and other resource specifications. Due to the differences in operating characteristics, different strategies are also required when dealing with GPU fragmentation problems in inference services to achieve optimal results.

[0035] See also Figure 1 , Figure 1 This is a flowchart of an exemplary embodiment of the task migration method for reasoning services of this application. Specifically, it may include the following steps:

[0036] Step S110 , determining a candidate graphics card node from among the graphics card nodes in the current graphics card cluster according to the node label; the node label is used to represent the amount of graphics card resources possessed by the graphics card node.

[0037] For example, reference may be made to Figure 2 As shown, Figure 2 It is an exemplary architectural diagram of the task migration method for reasoning services of the present application. In the present application, there may be one or more graphics card clusters (also referred to as graphics card node pools), and each graphics card cluster may have one or more graphics card nodes, and all or part of these graphics card nodes may be pre-set with node labels. And there may also be one or more candidate graphics card nodes determined, which is not limited here. Among them, the node label of each node may be set according to the amount of graphics card resources required by each node when executing the reasoning service, or it may be set manually. In the specific application process, the pod task may be deployed to the GPU node for operation through the k8s scheduler, and then for the pod task scheduled to the GPU node, the pod to be migrated can be determined therefrom, and the k8s rescheduler migrates the pod task to be migrated in the GPU node to achieve resource defragmentation.

[0038] It should be noted that, depending on the actual application scenario, the amount of graphics card resources can be expressed in different ways. For example, it can be expressed according to the number of graphics cards used by the graphics card node to execute each task (pod) (GPU card number specifications can be defined, generally 1 card, 2 cards, 4 cards, 8 cards, etc.); or it can be expressed according to the value of the graphics card memory used by the graphics card node to execute each task (for example, 12GB, 24GB, etc.). For ease of explanation, the following text mainly uses the number of graphics cards as an example to illustrate the amount of graphics card resources. For details, please refer to the existing technology for setting the GPU during model inference, which will not be repeated here.

[0039] Therefore, you can set corresponding node labels for the amount of graphics card resources (number of cards) when running inference services on different graphics card nodes. For example, a 1-card label indicates that each pod running the inference service on a node with this label will occupy one GPU card. Similarly, a 2-card label indicates that each pod running the inference service on this node will occupy two GPU cards. A 4-card label indicates that each pod running the inference service on this node will occupy four GPU cards.

[0040] Furthermore, graphics card nodes with the same node label can be grouped into a card specification resource type queue, thereby classifying each graphics card node according to its card number specifications. For example, based on the number of GPU cards required for inference services, you can pre-define the card number specification classification and the node labels corresponding to each node. The definition can refer to the following example:

[0041]

[0042] The following fields describe the main fields: metadata.name represents the name of the card-specification resource type queue (for example, one-gpu represents a 1-card queue, where the graphics card node uses one GPU card to run one pod). spec.nodeLabels refers to the node label corresponding to the card-specification resource type queue (for example, the node label corresponding to the 1-card GPU resource type queue has a value of 1, indicating one card). The above definitions can be used for each card number specification (2 cards, 4 cards, 8 cards, and so on).

[0043] Then, you can also define the mapping between the GPU node pool and the above card specification resource type queue. The definition can refer to the following example:

[0044]

[0045] The main fields are explained as follows: metadata.name: indicates the name of the GPU node pool. spec.resourceGroups.matchResourceKey refers to the key of the GPU resource defined in the task Pod. The k8s scheduler can select the corresponding card specification resource type queue based on the value corresponding to this key. spec.resourceGroups.flavors defines one or more card specification resource type queues mapped by this node pool (for example, the current h100 graphics card node pool defines 1 card, 2 cards, 4 cards, and 8 cards resource type queues). spec.resourceGroups.flavors.name.reservation is the number of reserved nodes for the corresponding card specification resource type queue. Reserved nodes can prevent overload operation.

[0046] It is understandable that the GPU node dynamically divides the GPU node pool into node queues (node ​​lists) with different card number specifications to support pod task scheduling, which is usually performed before pod migration. This is equivalent to scheduling each pod to the corresponding node first, and then analyzing whether the pod in the node needs to be migrated to reduce resource fragmentation. In summary, after the k8s scheduler pre-sets the GPU cluster and the GPU nodes therein according to the definition method of the above example, it can schedule task pods according to the GPU card specifications of each graphics card node, and the k8s rescheduler can also migrate task pods according to the GPU card specifications of each graphics card node.

[0047] It should also be noted that the current graphics card cluster may include labeled nodes with node labels and unlabeled nodes without node labels. During the execution of the present method, it is necessary to determine candidate graphics card nodes from the current graphics card cluster based on the node labels, that is, graphics card nodes that may require pod migration.

[0048] Step S120 : determining the tasks to be migrated in each candidate graphics card node according to a preset task migration rule.

[0049] Preset task migration rules are the pre-set rules used to determine which pods (tasks to be migrated) need to be migrated from each candidate graphics card node. These rules may include filtering pod attributes (such as whether the pod is a static pod or a mirror pod), filtering node configuration items, and filtering node resource types. These rules are not detailed here.

[0050] One or more tasks to be migrated may be determined from each candidate graphics card node according to preset task migration rules, which are not limited here.

[0051] Step S130 , matching the resource application amount of each task to be migrated with the graphics card resource amount of each candidate graphics card node to obtain a successfully matched target task and target graphics card node.

[0052] To illustrate the above steps, after determining the candidate graphics card nodes and their tasks to be migrated, the candidate graphics card nodes and the tasks to be migrated can be matched with their resource specifications. The successfully matched candidate graphics card nodes and tasks to be migrated are then determined as the target tasks and target graphics card nodes.

[0053] Exemplarily, for each pod task (task to be migrated), the GPU resource application amount of each pod can be read. For example, by reading the GPU resource application amount of the pod, it can be determined whether the pod is a 1-card task, a 2-card task, or a 4-card task, etc. If the pod is a 1-card task, it means that it needs to occupy 1 GPU card for deployment and operation. Therefore, the graphics card node suitable for running a 1-card pod is a 1-card node, and secondly, a 2-card node or a 4-card node can also be selected, but this will cause resource fragmentation in the 2-card or 4-card nodes, wasting graphics card resources. This is equivalent to 1-card node, 2-card node, and 4-card node all being migratable graphics card nodes for a 1-card pod task, but the 1-card node is the target graphics card node for the 1-card pod task.

[0054] Therefore, in this step, each task to be migrated is matched with each candidate graphics card node based on its resource request and the graphics card resource capacity of each candidate graphics card node, thereby obtaining successfully matched target tasks and target graphics card nodes (for example, matching 1-card pod to 1-card node, matching 2-card pod to 2-card node, etc.). For tasks to be migrated and candidate graphics card nodes that fail to match, this indicates that there are currently no suitable migration nodes for these tasks to be migrated, and these candidate graphics card nodes are currently unable to accommodate the tasks to be migrated. Therefore, the previous deployment status can be maintained without performing pod migration based on the tasks to be migrated and candidate graphics card nodes that failed to match.

[0055] Step S140: Migrate the target task to the target graphics card node.

[0056] In conjunction with the above steps, after determining the target task and target graphics card node that are successfully matched with each other, the target task can be migrated to the corresponding target graphics card node.

[0057] For example, a rolling migration approach can be used. First, a new pod task with the same name as the target pod task is added to the target graphics card node. After the new pod task is started on the target graphics card node, the target pod is deleted, thus achieving rolling pod migration. This ensures the continuous availability of the inference service during the migration process and meets requirements such as response latency.

[0058] It can be seen that the present application determines the candidate graphics card nodes based on the graphics card nodes in the current graphics card cluster; the preset task migration rules include a method for determining whether there are unmatched tasks in the candidate graphics card nodes, so the tasks to be migrated in each candidate graphics card node can be determined according to the preset task migration rules; the resource application amount of each task to be migrated is matched with the graphics card resource amount of each candidate graphics card node respectively, so as to re-search for each candidate graphics card node suitable for processing each task to be migrated, and obtain the target task and target graphics card node that are successfully matched; the target task is migrated to the target graphics card node, so that each task can be deployed in the graphics card node with matching resource volume through task migration, thereby realizing resource defragmentation of the graphics card node.

[0059] Based on the above embodiment, the present embodiment describes the steps of determining a candidate graphics card node from each graphics card node in the current graphics card cluster based on the node label. Specifically, the method of this embodiment includes the following steps:

[0060] Obtain the resource allocation rate of each labeled node in the current graphics card cluster; a labeled node is a graphics card node with a node label in the current graphics card cluster; and determine a labeled node with a resource allocation rate less than a resource allocation threshold as a candidate graphics card node.

[0061] In conjunction with the above embodiment, when screening candidate graphics card nodes from the current graphics card cluster, in addition to selecting based on whether the graphics card nodes have node labels, the graphics card nodes can also be screened based on their resource allocation rates.

[0062] For example, one or more resource allocation thresholds can be pre-set, and their specific values ​​depend on the actual application scenario and are not limited here. For example, the resource allocation threshold can be set to 100%, which is equivalent to if the resource allocation rate of a graphics card node in the current graphics card cluster is equal to 100%, indicating that the graphics card node will not generate resource fragmentation, and it will not be determined as a candidate graphics card node. Conversely, if the resource allocation rate of a graphics card node in the current graphics card cluster is less than 100%, indicating that the graphics card node may generate resource fragmentation, and if the graphics card node is a labeled node, it will be determined as a candidate graphics card node.

[0063] Based on the above embodiment, the present embodiment describes the steps of determining a labeled node with a resource allocation rate less than a resource allocation threshold as a candidate graphics card node. Specifically, the method of this embodiment includes the following steps:

[0064] According to the resource allocation rate, the labeled nodes with a resource allocation rate less than the resource allocation threshold are sorted to obtain candidate graphics card nodes.

[0065] In conjunction with the above embodiment, after selecting the labeled nodes whose resource allocation ratio is less than the resource allocation threshold, the labeled nodes whose resource allocation ratio is less than the resource allocation threshold may be sorted according to the resource allocation ratio.

[0066] Exemplarily, a node list may be generated based on the selected labeled nodes. Each labeled node in the node list is traversed, and the labeled nodes in the node list are sorted according to the resource allocation rate, so that each labeled node in the node list is sorted in descending order according to the resource allocation rate, and then each sorted labeled node in the node list is determined as a candidate graphics card node. It should be noted that the sorting method of the present application may be to sort and change the original node list to obtain a sorted node list, or to generate another sorted node list, which is not limited here.

[0067] Sorting the labeled nodes in the node list can prioritize GPU nodes with higher resource allocation rates. Generally, higher resource allocation rates increase the probability of resource fragmentation. Therefore, attempting to migrate pods sequentially for the sorted GPU nodes in the node list can more effectively address GPU resource fragmentation.

[0068] Based on the above embodiment, the present embodiment of the application describes the steps of determining the tasks to be migrated in each candidate graphics card node according to the preset task migration rules. Specifically, the method of this embodiment includes the following steps:

[0069] Traverse each candidate graphics card node and perform rule comparison with the task attributes of the current node task in each candidate graphics card node according to the preset task migration rules; determine the current node task that fails the comparison as the task to be migrated.

[0070] In conjunction with the above embodiment, the candidate graphics card nodes in this embodiment may be sorted according to resource allocation ratio or not, which is not limited here. Preferably, they may be sorted according to resource allocation ratio.

[0071] Then, each candidate graphics card node in the node list is traversed in order (traversing the nodes from high to low according to the resource allocation rate). For the candidate graphics card node currently traversed, the task attributes of the current node task (current pod) in the candidate graphics card node are compared with the preset task migration rules to determine whether there are pods (tasks to be migrated) that need to be migrated in each candidate graphics card node. Among them, pods that meet the preset task migration rules can remain in the original graphics card node and run normally, while pods that do not meet the preset task migration rules need to be migrated to other graphics card nodes for operation.

[0072] Based on the above embodiment, the present embodiment describes the steps of comparing the preset task migration rules with the task attributes of the current node task in each candidate graphics card node. Specifically, the method of this embodiment includes the following steps:

[0073] Determine whether the current node task is at least one of a static task, a mirror task, and a daemon task based on the task attributes; if not, perform a consistency comparison based on the resource application amount of the current node task and the graphics card resource amount of the graphics card node where the current node task is located; in response to the inconsistency between the resource application amount of the current node task and the graphics card resource amount of the graphics card node where the current node task is located, determine that the comparison between the preset task migration rule and the current node task has failed.

[0074] In combination with the above-mentioned embodiment, the preset task migration rules may include determining whether the pod is a task to be migrated based on its attributes. On this basis, it is also possible to determine whether the pod is a task to be migrated based on the configuration items corresponding to the candidate graphics card node. Then, it is also possible to determine whether it is a task to be migrated based on the consistency comparison results between the resource application amount of the pod and the graphics card resource amount of the graphics card node where it is located.

[0075] Pod task attributes can include static tasks, mirror tasks, and daemon tasks (pods generated by the DaemonSet controller, which ensures that each node in the Kubernetes cluster runs an identical copy of the pod). Pods with these attributes are generally not migrated. For more information, refer to the existing explanation of pods and are not detailed here.

[0076] The method of filtering tasks to be migrated based on configuration items may include but is not limited to: when include and exclude in the configuration item migratePodsOfWorkload have values, the rescheduler may select the pods to which include corresponds to the workload, and / or exclude the pods to which exclude corresponds to the workload according to the rules. Include is a pre-set workload of pods that need to be defragmented, and exclude is a workload of pods that need to be excluded (not defragmented). For the specific meaning of the parameters, please refer to the explanation of the configuration item in the prior art, which will not be elaborated here.

[0077] In addition, when PriorityThreshold is configured in the configuration item, pods can be filtered based on their priority and PriorityThreshold. For example, pods with a priority lower than PriorityThreshold can be selected or filtered (priorityThreshold.Value can be read first. If not set, the default value priorityThreshold.Name preset in the system can be read). This is not described in detail here.

[0078] If the attributes of the pod in the candidate graphics card node currently traversed include at least one of the above examples, the rule comparison is successful, indicating that the pod cannot be migrated. Conversely, if the attributes of the pod in the candidate graphics card node currently traversed do not include the above examples, the resource request amount of the pod (1 card, 2 cards, or 4 cards, etc.) can be obtained and compared with the graphics card resource amount of the graphics card node where the pod is located (for example, based on the value x in the node label queue.cube.xxx.com / gpu-classify: "x" corresponding to the graphics card node, the graphics card resource amount of the graphics card node can be determined to be 1 card, 2 cards, or 4 cards, etc.) for consistency comparison.

[0079] For example, if the resource request amount for the pod is 1 card, and the graphics card resource amount of the graphics card node where the pod is currently located is 1 card, then the resource request amount and the graphics card resource amount are consistent, and the comparison between the preset task migration rule and the current node task is determined to be successful. If the resource request amount for the pod is 1 card, and the graphics card resource amount of the graphics card node where the pod is currently located is 2 cards, then the resource request amount and the graphics card resource amount are inconsistent, and the comparison between the preset task migration rule and the current node task is determined to be a failure, which means that the pod is a task to be migrated.

[0080] Based on the above embodiment, the present embodiment describes the steps of matching the resource request amount of each task to be migrated with the graphics card resource amount of each candidate graphics card node to obtain a successfully matched target task and target graphics card node. Specifically, the method of this embodiment includes the following steps:

[0081] Traverse each task to be migrated, and determine the matching graphics card node that matches the currently traversed task to be migrated from the candidate graphics card nodes based on the resource application amount and graphics card resource amount; determine the matching graphics card node with idle resources and the currently traversed task to be migrated as the target graphics card node and target task.

[0082] With reference to the above embodiment, after each pod to be migrated is determined, a graphics card node suitable for re-migration and redeployment can be found for each pod to be migrated.

[0083] Exemplarily, matching can be performed based on the resource application amount of each pod to be migrated and the graphics card resource amount of each candidate graphics card node. Refer to the above example for explanation. For example, a pod with a resource application amount of 1 card matches a graphics card node with a graphics card resource amount of 1 card, and a pod with a resource application amount of 1 card does not match a graphics card node with a graphics card resource amount of 2 cards. Thus, a matching graphics card node that matches the currently traversed pod task to be migrated can be found. Among them, a matching graphics card node refers to a graphics card node whose resource amount specifications match those of the pod to be migrated. Therefore, in actual application scenarios, a matching graphics card node may support the transfer of the pod or not support the transfer of the pod. The target graphics card node refers to a node that can support the transfer of the pod. Therefore, it is specifically necessary to determine whether there are idle resources on the matching graphics card node. If there are idle resources, it indicates that the matching graphics card node can allocate GPU resources for the transfer of the pod (for example, a 1-card node can apply for 1-card GPU resources for the 1-card pod to be migrated); if there are no idle resources, it indicates that the matching graphics card node cannot allocate GPU resources for the transfer of the pod.

[0084] Therefore, after the above steps, the matching graphics card node with idle resources and the currently traversed task to be migrated are determined as the target graphics card node and target task. If the currently traversed task to be migrated does not have a corresponding matching graphics card node, or the matching graphics card node has no idle resources, it cannot be determined as the target graphics card node.

[0085] Optionally, during the implementation of this embodiment, the pods to be migrated that are screened out from each candidate graphics card node may be sorted first. The principle and process can be exemplified by the process of sorting the candidate graphics card nodes in the aforementioned embodiment. The larger the pod task, the more likely it is to cause resource fragmentation. Therefore, the pods to be migrated can be sorted from low to high according to the GPU application amount, and then the sorted pods to be migrated can be traversed in turn, so that the pods with small GPU application amounts are preferentially migrated to the appropriate graphics card node. By bidirectionally sorting the candidate graphics card nodes and the tasks to be migrated in the above example, the efficiency of resource fragmentation optimization can be improved.

[0086] Based on the above embodiment, the present embodiment of the application describes the steps of migrating the target task to the target graphics card node. Specifically, the method of this embodiment includes the following steps:

[0087] Perform task creation processing in the target graphics card node according to the target task to obtain the newly created task in the target graphics card node; and perform deletion processing on the target task.

[0088] In conjunction with the above embodiment, the method for migrating the target task to the target graphics card node in this embodiment can be based on a rolling migration method. According to the target pod task, pod creation processing is performed on the target graphics card node to expand a new pod (new pod) that is identical to the target pod in the target graphics card node. After the new pod is started in the target graphics card node, the original target pod can be deleted. This ensures the availability of the inference service during the migration process and meets requirements such as response delay.

[0089] In summary, please refer to Figure 3 As shown, Figure 3 This is a schematic diagram of an exemplary task migration scenario in the task migration method for inference services of the present application. The pod tasks to be migrated in each candidate graphics card node are traversed in sequence, thereby determining the target graphics card node and target pod tasks that match each other, and migrating the target pod tasks to the target GPU graphics card node. The graphics card node can then be reorganized into pods according to its card specifications (for example, 1 card, 2 cards, and 4 cards, etc.) to complete GPU defragmentation.

[0090] Based on the above embodiments, the embodiments of the present application should also explain that the overall steps provided in the above embodiments (re-migrating pod tasks from the graphics card node to the appropriate target graphics card node) can be performed irregularly or regularly. For example, regular defragmentation can be controlled according to a pre-configured periodic migration time. When the defragmentation is triggered regularly, steps including but not limited to S110 to S140 in the above embodiments can be executed.

[0091] Based on the above embodiment, the present embodiment describes the steps before determining a candidate graphics card node from each graphics card node in the current graphics card cluster based on the node label. Specifically, the method of this embodiment includes the following steps:

[0092] Determine the corresponding initial node label according to the resource requirements of the received initial scheduling task; determine the initial graphics card nodes that meet the resource requirements in the current graphics card cluster to obtain an initial node list; in response to the existence of an initial graphics card node corresponding to the initial node label in the initial node list, schedule the initial scheduling task to the initial graphics card node corresponding to the initial node label; in response to the absence of an initial graphics card node corresponding to the initial node label in the initial node list, schedule the initial scheduling task to an unlabeled node in the initial node list, and perform label setting processing on the unlabeled node according to the initial node label to obtain a labeled node.

[0093] In conjunction with the above embodiment, the above embodiment mainly illustrates a method for rescheduling pods to achieve pod migration to solve the resource fragmentation problem, which can be implemented by the Dscheduler (rescheduler) of k8s. In this embodiment, it is mainly explained that when the pod is initially scheduled and deployed on the GPU node, a pod scheduling method based on card specification classification is provided, which can be implemented by the scheduler (scheduler) of k8s. By classifying the card specification resource types mapped by the GPU node pool, the GPU nodes are dynamically classified into the corresponding card specification resource type queues during the scheduling phase, so that the GPU card allocation of each GPU node is regularized, which is conducive to reducing subsequent fragmentation. On this basis, the fragmentation method illustrated in the above example can be combined (equivalent to when receiving the pod task to be deployed, the method of this embodiment is used to schedule the pod to the appropriate graphics card node for deployment; then after triggering the defragmentation event, the pod tasks already deployed in the graphics card node can be analyzed to determine whether they need to be migrated to other graphics card nodes) to further optimize the fragmentation in the overall operating environment.

[0094] For example, it can be combined with Figure 4 As shown, Figure 4This is a flow chart of an exemplary card specification matching scheduling in the task migration method for reasoning services of the present application. Among them, when the initial scheduling task (initial pod) is received, the GPU resource application amount of the pod can be read, and the resource key is selected based on the spec.resourceGroups.matchResourceKey field defined by the scheduler. The kubernetes scheduler can perform pre-selected scheduling to filter out graphics card nodes that do not meet the preset requirements, such as filtering out nodes that cannot meet the resource requirements of the pod according to data such as CPU, memory, GPU, etc., thereby obtaining each initial graphics card node and forming an initial node list (pre-selected node list). The kubernetes scheduler reads the preset definition of the card specification resource type queue of the GPU node pool as in the aforementioned embodiment, reads the card rule resource type queue that is consistent with the GPU resource application amount of the pod (for example, a queue of 1 card pod corresponding to 1 card GPU node), and obtains the node label corresponding to the resource amount. Find out whether there is an initial graphics card node with such a node label from the initial node list.

[0095] If there is no initial graphics card node with such node label in the initial node list, or the GPU resources of such initial graphics card node have been allocated, you can select a node without node label and / or without such node label and GPU unassigned (for example, an unlabeled node) in the current graphics card cluster, schedule the pod to the node, and set the node label of the corresponding card specification resource type for the node to implement label setting processing, thereby obtaining a labeled node and dynamically classifying the GPU nodes in the current graphics card cluster.

[0096] If the initial graphics card node with this node label exists in the initial node list, you can filter out nodes with this node label, score each filtered node, and then select the node with the highest score to deploy the pod. The mathematical expression of the score of a single node can be as follows:

[0097]

[0098] Among them, Request refers to the resource request amount of the pod, which may include: GPURequest (GPU request amount), CPURequest (CPU request amount) and MemRequest (memory request amount); similarly, Allocated refers to the resource allocation amount of a single node, including: GPUAllocated (GPU allocation amount), CPUAllocated (CPU allocation amount) and MemAllocated (memory allocation amount); in addition, Total refers to the total resource amount of a single node, including: GPUTotal (GPU total amount), CPUTotal (CPU total amount) and MemTotal (memory total amount). Among them, weight1, weight2, and weight3 are pre-set weight parameters, and the specific values ​​and size relationships are not limited here. Optionally, since this application is for inference services, the weight1 weight of the GPU can be set to the highest.

[0099] In summary, after the received pod task is scheduled to the corresponding graphics card node for deployment, when the defragmentation method of the aforementioned embodiment is triggered, the graphics card node and the pod task in the graphics card node can be analyzed and judged. For details, please refer to the aforementioned embodiment, including but not limited to the method provided in steps S110 to S140, which will not be repeated here.

[0100] On the basis of the above embodiments, it should be noted that the embodiments of the present application can refer to Figure 5 As shown, Figure 5 This is an exemplary resource borrowing diagram in the task migration method for inference services of this application. When there are many inference services deployed, due to the tight allocation of GPU resources, a certain type of graphics card node may have an empty GPU idle node list, but graphics card nodes in other card specification resource type queues still have idle available GPU resources. In order to improve GPU utilization, when a certain type of GPU resources is insufficient, resources can be borrowed across card specification resource type queues. Figure 5 As shown in the figure, at a certain moment, a 2-card GPU service needs to be dispatched to a 2-card graphics card node for deployment. However, the GPU resources of the graphics card nodes in the 2-card resource queue have been fully allocated, and the corresponding GPU node idle list is empty. In other words, all GPU nodes have been assigned to the resource type queues of their respective card specifications. However, some nodes in the 4-card resource queue have idle GPU resources. To improve the overall GPU resource allocation rate, the scheduler allows the 2-card pod task to be dispatched to the graphics card node in the 4-card resource queue.

[0101] When overall GPU resource allocation is tight, task pods whose GPU card counts do not match the card specification resource type queues are allowed to borrow resources. This situation may disrupt the resource regularization of GPU card count specifications in the graphics card node, resulting in the risk of GPU fragmentation. Therefore, the resource defragmentation method provided in the above embodiment can be used to regularly defragment GPU resources and regularize task pods that do not match the GPU card count of the graphics card node.

[0102] It should be further explained that the execution subject of the task migration method for reasoning services may be a task migration device for reasoning services. For example, the task migration method for reasoning services may be executed by a terminal device, a server, or other processing device, wherein the terminal device may be a user equipment (UE), a computer, a mobile device, a user terminal, a terminal, a cellular phone, a cordless phone, a personal digital assistant (PDA), a handheld device, a computing device, an in-vehicle device, a wearable device, etc. In some possible implementations, the task migration method for reasoning services may be implemented by a processor calling computer-readable instructions stored in a memory.

[0103] Figure 6 FIG. 1 is a block diagram of a task migration device for reasoning services shown in an exemplary embodiment of the present application. Figure 6 As shown, the exemplary task migration apparatus 600 for reasoning services includes: a candidate node determination module 610, a task to be migrated determination module 620, a task and node matching module 630, and a task migration module 640. Specifically:

[0104] The candidate node determination module 610 is used to determine a candidate graphics card node from among the graphics card nodes in the current graphics card cluster according to the node label; the node label is used to represent the amount of graphics card resources possessed by the graphics card node.

[0105] The task to be migrated determining module 620 is configured to determine the tasks to be migrated in each candidate graphics card node according to a preset task migration rule.

[0106] The task and node matching module 630 is used to match the resource application amount of each task to be migrated with the graphics card resource amount of each candidate graphics card node, and obtain a successfully matched target task and target graphics card node.

[0107] The task migration module 640 is used to migrate the target task to the target graphics card node.

[0108] In this exemplary task migration device for inference services, candidate graphics card nodes are determined among the graphics card nodes of the current graphics card cluster based on node labels; the node labels represent the amount of graphics card resources possessed by the graphics card nodes; the preset task migration rules include a method for determining whether there are unmatched tasks in the candidate graphics card nodes, so the tasks to be migrated in each candidate graphics card node can be determined according to the preset task migration rules; the resource application amount of each task to be migrated is matched with the graphics card resource amount of each candidate graphics card node respectively, so as to re-search for each candidate graphics card node suitable for processing each task to be migrated, and obtain a successfully matched target task and target graphics card node; the target task is migrated to the target graphics card node, so that each task can be deployed in a graphics card node with matching resource volume through task migration, thereby realizing resource defragmentation of the graphics card node.

[0109] It should be noted that the apparatus provided in the above embodiments and the methods provided in the above embodiments are based on the same concept. The specific manner in which the various modules and units perform their operations has been described in detail in the method embodiments and will not be repeated here. In actual applications, the apparatus provided in the above embodiments can, as needed, allocate the above functions to different functional modules, i.e., divide the internal structure of the apparatus into different functional modules to perform all or part of the functions described above. This is not a limitation herein.

[0110] The functions of each module can be found in the embodiment of the task migration method for reasoning services, which will not be described in detail here.

[0111] See also Figure 7 , Figure 7 1 is a schematic diagram of the structure of an embodiment of an electronic device of the present application. Electronic device 100 includes memory 101 and processor 102. Processor 102 is configured to execute program instructions stored in memory 101 to implement the steps of any of the above-mentioned embodiments of the task migration method for reasoning services. In a specific implementation scenario, electronic device 100 may include, but is not limited to, a microcomputer and a server. In addition, electronic device 100 may also include mobile devices such as laptops and tablet computers, which are not limited here.

[0112] Specifically, the processor 102 is used to control itself and the memory 101 to implement the steps in any of the above-mentioned embodiments of the task migration method for reasoning services. The processor 102 can also be called a CPU (Central Processing Unit). The processor 102 may be an integrated circuit chip with signal processing capabilities. The processor 102 can also be a general-purpose processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA) or other programmable logic devices, discrete gates or transistor logic devices, discrete hardware components. The general-purpose processor can be a microprocessor or the processor can also be any conventional processor, etc. In addition, the processor 102 can be implemented by an integrated circuit chip.

[0113] In this exemplary electronic device, candidate graphics card nodes are determined among the graphics card nodes of the current graphics card cluster based on node labels; the node labels represent the amount of graphics card resources possessed by the graphics card nodes; the preset task migration rules include a method for determining whether there are unmatched tasks in the candidate graphics card nodes, so the tasks to be migrated in each candidate graphics card node can be determined based on the preset task migration rules; the resource application amount of each task to be migrated is matched with the graphics card resource amount of each candidate graphics card node respectively, so as to re-search for each candidate graphics card node suitable for processing each task to be migrated, and obtain a successfully matched target task and target graphics card node; the target task is migrated to the target graphics card node, so that each task can be deployed in a graphics card node with matching resource volume through task migration, thereby realizing resource defragmentation of the graphics card node.

[0114] See also Figure 8 , Figure 8 The computer-readable storage medium 110 stores program instructions 111 that can be executed by a processor, and the program instructions 111 are used to implement the steps of any of the above-mentioned task migration method embodiments for reasoning services.

[0115] In this exemplary storage medium, by running the program instructions in the storage medium, candidate graphics card nodes are determined in each graphics card node of the current graphics card cluster according to the node label; the node label represents the amount of graphics card resources possessed by the graphics card node; the preset task migration rules include a method for determining whether there are unmatched tasks in the candidate graphics card nodes, so the tasks to be migrated in each candidate graphics card node can be determined according to the preset task migration rules; the resource application amount of each task to be migrated is matched with the graphics card resource amount of each candidate graphics card node respectively, so as to re-search each candidate graphics card node suitable for processing each task to be migrated, and obtain a successfully matched target task and target graphics card node; the target task is migrated to the target graphics card node, so that each task can be deployed in a graphics card node with matching resource volume through task migration, thereby realizing resource defragmentation of the graphics card node.

[0116] In some embodiments, the functions or modules included in the device provided by the embodiments of the present disclosure can be used to execute the method described in the above method embodiments. The specific implementation can refer to the description of the above method embodiments. For the sake of brevity, it will not be repeated here.

[0117] The above description of the various embodiments tends to emphasize the differences between the various embodiments. The same or similar aspects can be referenced with each other and will not be repeated herein for the sake of brevity.

[0118] In the several embodiments provided in this application, it should be understood that the disclosed methods and devices can be implemented in other ways. For example, the device implementation methods described above are only schematic. For example, the division of modules or units is only a logical function division. There may be other division methods in actual implementation. For example, units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, and the indirect coupling or communication connection of devices or units can be electrical, mechanical or other forms.

[0119] In addition, the functional units in the various embodiments of the present application can be integrated into a processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The above-mentioned integrated unit can be implemented in the form of hardware or in the form of a software functional unit. If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application is essentially or the part that contributes to the prior art or all or part of the technical solution can be embodied in the form of a software product, which is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, server, or network device, etc.) or a processor to execute all or part of the steps of the various embodiments of the present application. The aforementioned storage medium includes various media that can store program code, such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk.

Claims

1. A task migration method for reasoning services, characterized in that: The method comprises: Determine a candidate graphics card node from among the graphics card nodes in the current graphics card cluster according to a node label, wherein the node label is used to characterize the amount of graphics card resources possessed by the graphics card node; Determine the tasks to be migrated in each candidate graphics card node according to the preset task migration rules; Match the resource application amount of each task to be migrated with the graphics card resource amount of each candidate graphics card node to obtain the successfully matched target task and target graphics card node; Migrate the target task to the target graphics card node.

2. The method according to claim 1, characterized in that The step of determining a candidate graphics card node from among the graphics card nodes in the current graphics card cluster according to the node label includes: Obtaining a resource allocation rate of each labeled node in the current graphics card cluster; the labeled node is a graphics card node in the current graphics card cluster that has the node label; The labeled node whose resource allocation rate is less than the resource allocation threshold is determined as the candidate graphics card node.

3. The method according to claim 2, characterized in that The step of determining the labeled node whose resource allocation rate is less than the resource allocation threshold as the candidate graphics card node includes: The labeled nodes whose resource allocation rates are lower than the resource allocation threshold are sorted according to the resource allocation rate to obtain the candidate graphics card nodes.

4. The method according to claim 1, wherein The step of determining the tasks to be migrated in each candidate graphics card node according to the preset task migration rules includes: Traverse each candidate graphics card node and compare the task attributes of the current node task in each candidate graphics card node with the preset task migration rules; The current node task that fails the comparison is determined as the task to be migrated.

5. The method according to claim 4, characterized in that The step of comparing the preset task migration rules with the task attributes of the current node task in each candidate graphics card node includes: Determine whether the current node task is at least one of a static task, a mirror task, and a daemon task according to the task attribute; If not, then performing a consistency comparison based on the resource application amount of the current node task and the graphics card resource amount of the graphics card node where the current node task is located; In response to the resource application amount of the current node task being inconsistent with the graphics card resource amount of the graphics card node where the current node task is located, it is determined that the comparison between the preset task migration rule and the current node task has failed.

6. The method according to claim 1, characterized in that The resource application amount of each task to be migrated is matched with the graphics card resource amount of each candidate graphics card node to obtain a successfully matched target task and target graphics card node, including: Traversing each task to be migrated, and determining a matching graphics card node that matches the currently traversed task to be migrated from the candidate graphics card nodes according to the resource application amount and the graphics card resource amount; A matching graphics card node with idle resources and a currently traversed task to be migrated are determined as the target graphics card node and the target task.

7. The method according to claim 1, characterized in that Migrating the target task to the target graphics card node includes: Performing task creation processing in the target graphics card node according to the target task to obtain a new task in the target graphics card node; The target task is deleted.

8. The method according to claim 1, characterized in that Before determining the candidate graphics card node from the graphics card nodes in the current graphics card cluster according to the node label, the method further includes: Determine the corresponding initial node label based on the resource requirements of the received initial scheduling task; Determine each initial graphics card node in the current graphics card cluster that meets the resource requirement, and obtain an initial node list; In response to the presence of an initial graphics card node corresponding to the initial node label in the initial node list, scheduling the initial scheduling task to the initial graphics card node corresponding to the initial node label; In response to the fact that the initial graphics card node corresponding to the initial node label does not exist in the initial node list, the initial scheduling task is scheduled to the unlabeled node in the initial node list, and the unlabeled node is label-set according to the initial node label to obtain a labeled node.

9. An electronic device, characterized in that: The method comprises a memory and a processor, wherein the processor is configured to execute program instructions stored in the memory to implement the method according to any one of claims 1 to 8.

10. A computer-readable storage medium having program instructions stored thereon, characterized in that: When the program instructions are executed by a processor, the method according to any one of claims 1 to 8 is implemented.