Resource management method and device of graphics processor, storage medium and computer equipment

By clustering tasks and fine-grained resource scheduling in Kubernetes system, the problem of GPU resources cannot be flexibly allocated is solved, and efficient GPU resource utilization and task processing efficiency are achieved.

CN120371492APending Publication Date: 2025-07-25MOMENTA (SUZHOU) TECHNOLOGY CO LTD
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202410095501.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-01-23
Publication Date
2025-07-25

AI Technical Summary

Technical Problem

The Kubernetes system cannot flexibly adapt to diversified needs in GPU resource allocation, resulting in the inability to effectively utilize GPU resources and video memory and low task processing efficiency.

Method used

By determining the target tasks that meet the clustering conditions in multiple tasks, performing classification processing, and binding the target tasks of the same category to the target computing nodes in the Kubernetes cluster, calling the graph processor resources of the target computing nodes to process the target tasks, realizing fine-grained scheduling of resources and clustering of tasks.

Benefits of technology

Effectively reduce resource fragmentation, improve GPU resource utilization and task throughput efficiency, broaden usage scenarios, and realize flexible allocation and efficient utilization of resources.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120371492A_ABST
    Figure CN120371492A_ABST
Patent Text Reader

Abstract

The invention discloses a resource management method and device of a graphics processor, a storage medium and computer equipment. Comprising the following steps: in response to execution instructions of a plurality of tasks, determining a target task meeting a clustering condition in the plurality of tasks; performing classification processing on the target task; the target tasks of the same type are bound to a target computing node in a Kubernetes cluster; and calling resources of the graphics processor of the target computing node to process the target task. Therefore, a group of related tasks are allocated to the same computing node, generation of resource fragments is effectively reduced, more complete machine resources are reserved to be provided for other tasks, the purposes of utilizing computing node resources as much as possible and reducing communication overhead are achieved, a user can flexibly use GPU resources on a Kubernetes platform conveniently, and the user experience is improved. The GPU resource utilization rate and the task throughput efficiency are greatly improved, the use scene is widened, and the method can be applied more flexibly.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of computer technologies, and in particular, to a method, apparatus, storage medium, and computer device for resource management of a graphics processing unit. Background Art

[0002] The K8s (Kubernetes, a container cluster management system) platform with containers as the application running carriers has been widely used in the fields of artificial intelligence (AI) and machine learning. In distributed machine learning tasks based on Kubernetes, for multi-modal training scenarios, the demand for GPU (Graphics Processing Unit) resources is often necessary and diverse. For example, in a training scenario, it often requires a large number of GPU cards of the whole machine or even multiple machines; in a single-machine Debug scenario, it may only require one CPU card.

[0003] In related technologies, the Kubernetes system provides the node orchestration ability of Nvidia GPUs. Usually, a GPU card is randomly or assigned to a container according to the specified label of the container, and does not support the shared use of a single GPU, nor does it address the reuse of GPUs and GPU video memories and the isolated allocation of video memories. However, technically, it cannot flexibly adapt to scenario applications. When GPU resources are in a long-term tight state and the demand is sufficient, GPUs and GPU video memories cannot be effectively utilized, resulting in the problem of low task processing efficiency. Summary of the Invention

[0004] In view of this, this application provides a method, apparatus, storage medium, and computer device for resource management of a graphics processing unit to solve the problems described in the background art.

[0005] According to one aspect of this application, a method for resource management of a graphics processing unit is provided, including:

[0006] In response to execution instructions of multiple tasks, determining target tasks that meet clustering conditions among the multiple tasks;

[0007] Performing classification processing on the target tasks;

[0008] Binding target tasks of the same category to target computing nodes in a Kubernetes cluster, where the Kubernetes cluster includes multiple computing nodes, each computing node includes at least one graphics processing unit, and the target computing nodes are one or more computing nodes in the multiple computing nodes that are in an available state;

[0009] Invoking the resources of the graphics processing unit of the target computing node to process the target tasks.

[0010] Further, binding target tasks of the same category to target computing nodes in the Kubernetes cluster specifically includes:

[0011] Obtain the total resource requirements of the target tasks of the same category and the resource supply of each available computing node in the Kubernetes cluster;

[0012] If the total resource requirements are less than or equal to the resource supply of any available computing node, bind the target tasks of the same category to any computing node;

[0013] If the total resource requirements are greater than the resource supply of each available computing node, determine the first computing node based on the priority of the available computing nodes in the Kubernetes cluster;

[0014] Bind the first task among the target tasks of the same category to the first computing node, where the total resource requirements of the first task are less than or equal to the resource supply of the first computing node;

[0015] Bind the second tasks other than the first task among the target tasks of the same category to the second computing node, where the second computing node is an available computing node in the Kubernetes cluster other than the first computing node.

[0016] Further, the resource management method of the graphics processor further includes:

[0017] Obtain the number of graphics processors included in the computing node, the resource amount of each graphics processor, and the resource call amount of each graphics processor;

[0018] Based on the resource amount and the resource call amount, calculate the resource available for calling of each graphics processor respectively;

[0019] Determine the sum of the resource available for calling of the graphics processors included in the computing node as the resource supply;

[0020] If the resource supply is greater than the preset resource amount, determine that the computing node is in an available state.

[0021] Further, calling the resources of the graphics processor of the target computing node to process the target task specifically includes:

[0022] If the resource requirement of the third task is less than or equal to the resource available for calling of the first graphics processor, call the resources of the first graphics processor to process the third task based on the resource requirement;

[0023] If the resource requirement of the third task is greater than the resource available for calling of each graphics processor of the target computing node, call the resources of the second graphics processor to process the third task based on the resource requirement;

[0024] Among them, the third task is one of the target tasks, and the first graphics processor is any graphics processor of the target computing node; the second graphics processor includes at least two graphics processors, and the sum of the resource requirements of the at least two graphics processors is greater than or equal to the resource requirement of the third task.

[0025] Furthermore, invoking the resources of the graphics processor of the target computing node to process the target task specifically includes:

[0026] If the third task runs in multiple container groups and the task information of the third task has the specified graphics processor information of the container group, based on the resource requirements of the container group, invoke the resources of the graphics processor corresponding to the specified graphics processor information to start the container group;

[0027] Among them, the third task is one of the target tasks.

[0028] Furthermore, classifying the target task specifically includes:

[0029] Obtain the task information of multiple tasks, where the task information includes at least one of the following: task type information, resource requirement information, task attribution information, and task attribute information;

[0030] Based on the task information, classify the target task to determine the category of the target task.

[0031] Furthermore, after invoking the resources of the graphics processor of the target computing node to process the target task, the resource management method of the graphics processor further includes:

[0032] If the target computing node fails to process the target task and the number of restart times is greater than the preset number, bind the target task to a new target computing node in the Kubernetes cluster.

[0033] According to another aspect of the present application, there is provided a resource management device for a graphics processor, including:

[0034] A screening module, configured to determine target tasks that meet the clustering conditions among multiple tasks in response to execution instructions of the multiple tasks;

[0035] A classification module, configured to classify the target tasks;

[0036] A scheduling module, configured to bind target tasks of the same category to target computing nodes in the Kubernetes cluster, where the Kubernetes cluster includes multiple computing nodes, each computing node includes at least one graphics processor, and the target computing node is one or more computing nodes in the multiple computing nodes that are in an available state;

[0037] A processing module, configured to call the resources of the graphics processor of the target computing node to process the target task.

[0038] Furthermore, a scheduling module, specifically configured to obtain the total resource requirements of the target tasks of the same category and the resource supply of each available computing node in the Kubernetes cluster; if the total resource requirements are less than or equal to the resource supply of any available computing node, bind the target tasks of the same category to any computing node; if the total resource requirements are greater than the resource supply of each available computing node, determine a first computing node based on the priority of the available computing nodes in the Kubernetes cluster; bind the first task in the target tasks of the same category to the first computing node, where the total resource requirements of the first task are less than or equal to the resource supply of the first computing node; bind the second tasks other than the first task in the target tasks of the same category to a second computing node, where the second computing node is an available computing node in the Kubernetes cluster other than the first computing node.

[0039] Furthermore, the resource management device of the graphics processor further includes:

[0040] A first obtaining module, configured to obtain the number of graphics processors included in the computing node, the resource amount of each graphics processor, and the resource call amount of each graphics processor.

[0041] A determining module, configured to calculate the available resource call amount of each graphics processor based on the resource amount and the resource call amount; and determine the sum of the available resource call amounts of the graphics processors included in the computing node as the resource supply; and if the resource supply is greater than the preset resource amount, determine that the computing node is in an available state.

[0042] Furthermore, the processing module is specifically configured to, if the resource requirement of the third task is less than or equal to the available resource call amount of the first graphics processor, call the resources of the first graphics processor to process the third task based on the resource requirement; if the resource requirement of the third task is greater than the available resource call amount of each graphics processor of the target computing node, call the resources of the second graphics processor to process the third task based on the resource requirement; where the third task is one of the target tasks, the first graphics processor is any graphics processor of the target computing node; the second graphics processor includes at least two graphics processors, and the sum of the resource requirements of the at least two graphics processors is greater than or equal to the resource requirement of the third task.

[0043] Further, the processing module is specifically configured to, if a third task runs on multiple container groups and the task information of the third task has the specified graphics processor information of the container group, call the resources of the graphics processor corresponding to the specified graphics processor information to start the container group based on the resource requirements of the container group; wherein, the third task is one of the target tasks.

[0044] Further, the graphics processor resource management device further includes:

[0045] A second acquisition module, configured to acquire the task information of multiple tasks, where the task information includes at least one of the following: task type information, resource requirement information, task attribution information, and task attribute information;

[0046] A classification module, specifically configured to perform classification processing on the target tasks based on the task information to determine the categories of the target tasks.

[0047] Further, the scheduling module is further configured to, if the target computing node fails to process the target task and the number of restart times is greater than a preset number of times, bind the target task to a new target computing node in the Kubernetes cluster.

[0048] According to another aspect of the present application, there is provided a computer-readable storage medium, on which a program or instruction is stored, and when the program or instruction is executed by a processor, the steps of the above-mentioned graphics processor resource management method are implemented.

[0049] According to another aspect of the present application, there is provided a computer device, including at least one processor, the processor is coupled with a memory, and the memory stores a computer program running on the processor, characterized in that when the processor executes the program, the steps of the above-mentioned graphics processor resource management method are implemented.

[0050] By means of the above technical solution, when the system needs to process multiple tasks, the target tasks that can be processed by clustering are screened out from the multiple tasks. The target tasks are divided through classification processing, and the target tasks of the same category are bound to the same target computing node in the Kubernetes cluster, so as to facilitate the system to call the resources of the graphics processor of the target computing node to process the target tasks of the same category. Thus, a group of related tasks are allocated to the same computing node, effectively reducing the generation of resource fragmentation, reserving more overall machine resources for other tasks, and further achieving the purpose of making the most of the computing node resources and reducing communication overhead, facilitating users to flexibly use GPU resources on the Kubernetes platform, greatly improving the GPU resource utilization rate and task throughput efficiency, broadening the usage scenarios, and enabling more flexible applications.

[0051] The above description is only an overview of the technical solution of the present application. In order to understand the technical means of the present application more clearly, it can be implemented according to the content of the specification. And in order to make the above and other purposes, features and advantages of the present application more obvious and understandable, the specific embodiments of the present application are specifically exemplified below. Description of the Drawings

[0052] The drawings described herein are used to provide a further understanding of the present application and form a part of the present application. The schematic embodiments of the present application and their descriptions are used to explain the present application and do not constitute an improper limitation to the present application. In the drawings:

[0053] Figure 1 A flowchart showing the resource management method of the graphics processor provided by the embodiment of the present application is shown;

[0054] Figure 2 A block diagram showing the structure of the resource management device of the graphics processor provided by the embodiment of the present application is shown. Detailed Embodiments

[0055] The present application will be described in detail below with reference to the drawings and in combination with embodiments. It should be noted that, without conflict, the embodiments in the present application and the features in the embodiments can be combined with each other.

[0056] The embodiments of the present application will be described in detail below. The examples of the embodiments are shown in the drawings, where the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions from beginning to end. The embodiments described below by referring to the drawings are exemplary and are only used to explain the present application and should not be construed as a limitation to the present application.

[0057] Those skilled in the art of the present technology can understand that, unless specifically stated otherwise, the singular forms "a", "an", "the" and "said" used herein may also include the plural forms. It should be further understood that the term "comprising" used in the specification of the present application means the presence of the described features, integers, steps, operations, elements and / or components, but does not exclude the presence or addition of one or more other features, integers, steps, operations, elements, components and / or their groups. It should be understood that when we say that an element is "connected" or "joined" to another element, it can be directly connected or joined to other elements, or there may also be intermediate elements. In addition, the "connection" or "joining" used herein may include wireless connection or wireless coupling. The phrase "and / or" used herein includes all or any unit and all combinations of one or more related listed items.

[0058] Now, exemplary embodiments according to the present application will be described in more detail with reference to the accompanying drawings. However, these exemplary embodiments can be implemented in many different forms and should not be construed as being limited only to the embodiments set forth herein. It should be understood that these embodiments are provided so that the disclosure of the present application is thorough and complete, and the concepts of these exemplary embodiments are fully conveyed to those of ordinary skill in the art.

[0059] In this embodiment, a resource management method for a graphics processor is provided. As Figure 1 shown, the method includes:

[0060] Step 101, in response to execution instructions of multiple tasks, determine target tasks among the multiple tasks that meet the clustering conditions.

[0061] Among them, the clustering conditions are used to filter tasks that can be processed by clustering in groups. For example, the clustering conditions can be set to use tasks without specified processing nodes as target tasks, or tasks with resource requirements lower than a preset value as target tasks, or tasks that do not need to occupy the entire machine's GPU card as target tasks, etc. The specific clustering conditions can be reasonably set according to actual needs, and the embodiments of the present application will not list them one by one.

[0062] It can be understood that the tasks in this embodiment refer to an abstract work unit, such as Deployment, Job, StatefulSet, which describes the application program to be run and its requirements. A task can be executed by one or more container groups (pod). Exemplarily, a task can be a detection task, a classification task, a training task, a segmentation task, etc. Multiple tasks can be tasks created by a user interface, or tasks created by other devices forwarded by a terminal. The present application does not specifically limit all possible types of tasks and the source methods of tasks.

[0063] Step 102, perform classification processing on the target tasks.

[0064] In a specific application scenario, Step 102, that is, performing classification processing on the target tasks, specifically includes the following steps:

[0065] Step 102-1, obtain the task information of the multiple tasks.

[0066] Among them, the task information includes at least one of the following: task type information (such as detection type, training type, etc.), resource requirement information (such as GPU resource requirement quantity, type of required GPU card, time for executing the task, etc.), task attribution information (such as user interface identifier for creating the task, terminal identifier for sending the task, etc.), and task attribute information (such as task identifier, task content, task priority, etc.).

[0067] For example, classify Task 1 and Task 2 created by User A into one category, or classify Task 3, Task 4, and Task 5 with GPU resource requirements less than 500MB into one category.

[0068] Step 102-2: Classify the target tasks based on the task information to determine the categories of the target tasks.

[0069] In this embodiment, different tasks are classified through the task information, and tasks with similar characteristics are divided into one category. This facilitates the centralized allocation of resources by the computing nodes to tasks in the same category, not only improving the processing efficiency and accuracy of processing these tasks, but also being able to reserve more overall machine resources for other tasks. At the same time, uniformly processing tasks in the same category makes it convenient for users to clearly understand the status and progress of the tasks, and thus better control the task progress.

[0070] Step 103: Bind the target tasks in the same category to the target computing nodes in the Kubernetes cluster.

[0071] Among them, the Kubernetes (K8s) cluster is a container cluster management platform for containers developed based on a cluster manager. It uniformly manages the underlying host, network, and storage resources based on container technology, etc. It provides functions such as application deployment, maintenance, and expansion mechanisms. Using K8s can conveniently manage containerized applications running across machines. The Kubernetes cluster includes multiple computing nodes, and the identifiers of different nodes are different. The computing nodes can be computing devices such as bare machines, physical machines, and virtual machines. Each computing node includes at least one graphics processor, and there is a topological relationship among at least one graphics processor. This topological relationship is used to represent the priority of at least one graphics processor being called. The target computing node is one or more computing nodes in the available state among the multiple computing nodes.

[0072] It can be understood that the target tasks in the same category can be one or multiple.

[0073] In this embodiment, when the system needs to process multiple tasks, target tasks that can be clustered together are screened out from the multiple tasks. The target tasks are divided through classification processing, and the target tasks in the same category are bound to the same target computing nodes in the Kubernetes cluster. In this way, the target tasks in the same category will share the same node memory, disk, and network resources, which can improve the data locality of the tasks, reduce data movement and communication overhead. Moreover, the target tasks in the same category can run independently in the same restricted environment, preventing them from interfering with other categories of tasks or computing nodes, achieving the effect of resource isolation, reducing potential failure points, and ensuring the correct execution order of the target tasks, thereby helping to improve the execution efficiency of the tasks.

[0074] In a specific application scenario, step 103, that is, binding target tasks of the same category to target computing nodes in a Kubernetes cluster, specifically includes the following steps:

[0075] Step 103-1: Obtain the total resource requirements of the target tasks of the same category and the resource supply of each available computing node in the Kubernetes cluster.

[0076] Among them, the resource supply refers to the amount of resources that a computing node can use for scheduling.

[0077] Step 103-2: If the total resource requirements are less than or equal to the resource supply of any available computing node, bind the target tasks of the same category to any computing node.

[0078] Among them, if there are multiple target tasks of the same category, the total resource requirements are the sum of the resource requirements of multiple target tasks of the same category.

[0079] In this embodiment, when the total resource requirements of the target tasks of the same category are less than or equal to the resource supply of any computing node, it indicates that the computing node has sufficient resources to process the target tasks of the same category. Then, bind the target tasks of the same category to the same computing node to call the resources of the same computing node to process the target tasks of the same category in parallel or serially. In this way, the target tasks of the same category will share the same node memory, disk, and network resources, thus simplifying the data transmission and communication between target tasks.

[0080] Step 103-3: If the total resource requirements are greater than the resource supply of each available computing node, determine the first computing node based on the priority of the available computing nodes in the Kubernetes cluster.

[0081] Step 103-4: Bind the first task among the target tasks of the same category to the first computing node.

[0082] Among them, the total resource requirements of the first task are less than or equal to the resource supply of the first computing node.

[0083] Step 103-5: Bind the second task other than the first task among the target tasks of the same category to the second computing node.

[0084] Among them, the second computing node is an available computing node in the Kubernetes cluster other than the first computing node.

[0085] In this embodiment, if the total resource demand of the target task is greater than the resource supply of each computing node in an available state, it means that a single computing node is difficult to support the processing requirements of all target tasks, and due to insufficient computing node resources, all target tasks may not be processed. At this time, the first computing node that may be called first among the computing nodes in an available state is determined according to the priority of the preset computing node. Then the first task whose total resource demand is less than or equal to the resource supply of the first computing node is bound to the first computing node, so as to process as many target tasks as possible through the resources of the first computing node. Then the remaining second task is bound to the second computing node to process the second task that the first computing node cannot process through other computing nodes. Thus, the resource limitation of the computing node is fully considered, and on the basis of realizing the executable allocation of the target task, it is ensured that the target tasks of the same category can be bound to the same computing node as much as possible, ensuring that the workload of the computing node is close to maximization, avoiding the failure of all target tasks of the entire category due to insufficient resources, and avoiding the delay caused by the task waiting for resources, making the maximum use of the computing power of the node, improving the task processing efficiency of the overall system, and enhancing the reliability and fault tolerance of the system.

[0086] It is understandable that the first task bound to the first computing node can be screened based on the task information. For example, the task priorities are sorted in order, and the target task that meets the resource supply of the first computing node and is at the front of the sorting order is taken as the first task; or the target task whose total resource demand is closest to the resource supply of the first computing node is taken as the first task.

[0087] For example, compute node 1 has a resource supply of 20, and compute node 2 has a resource supply of 15. Task A of the same category requires 10 resources, task B requires 8 resources, and task C requires 6 resources. The extended scheduler attempts to assign task A and task B to compute node 1 because their resource requirements are 10 and 8 respectively, totaling 18, which is less than the resource supply of 20 of compute node 1, and then assigns the remaining task C to compute node 2.

[0088] Furthermore, before step 103, the resource management method of the graphics processor also includes: obtaining the number of graphics processors included in the computing node, the resource amount of each graphics processor and the resource call amount of each graphics processor; based on the resource amount and the resource call amount, respectively calculating the resource callable amount of each graphics processor; determining the sum of the resource callable amounts of the graphics processors included in the computing node as the resource supply amount; if the resource supply amount is greater than the preset resource amount, determining that the computing node is in an available state.

[0089] Among them, the preset resource amount can be reasonably set according to the scheduling configuration requirements or task requirements.

[0090] In this embodiment, the number of graphics processors deployed on each computing node, the resource amount of each graphics processor, and the resource call amount of the resources already used by each graphics processor currently are determined. By calculating the difference between the resource amount and the resource call amount, the resource callable amount of the available resources of each graphics processor can be determined. According to the number of graphics processors deployed on a computing node, the sum of the available resources of the graphics processors of the computing node is used as the resource supply amount of the computing node. When the resource supply amount is greater than the preset resource amount, it indicates that the resources that can be called by the computing node can meet the resource requirements of at least one task, so it can be determined that the computing node is in an available state and can process the target task subsequently. Thus, taking the resource amount of the graphics processor as the minimum scheduling unit realizes a finer-grained division of the resources of the graphics processor, avoiding the problem that when a single graphics processor is used as the minimum scheduling unit, only one GPU card can be randomly or assigned to a container according to the specified label of the container, which helps to realize the scheduling of multiple tasks and multiple containers on a single card according to resource requirements.

[0091] Step 104, call the resources of the graphics processor of the target computing node to process the target task.

[0092] Through the resource management method of the graphics processor provided in this embodiment, when the system needs to process multiple tasks, the target tasks that can be processed by clustering are screened out from the multiple tasks. The target tasks are divided through classification processing, and the target tasks of the same category are bound to the same target computing node in the Kubernetes cluster, so that the system can call the resources of the graphics processor of the target computing node to process the target tasks of the same category. Thus, a group of related tasks are allocated to the same computing node, effectively reducing the generation of resource fragmentation, reserving more overall machine resources for other tasks, and further achieving the purpose of making the most use of the resources of the computing node and reducing communication overhead, facilitating users to flexibly use GPU resources on the Kubernetes platform, greatly improving the resource utilization rate and task throughput efficiency of the graphics processor, broadening the usage scenarios, and being able to be applied more flexibly.

[0093] It should be noted that the resources of the graphics processor include network, storage, and computing resources, etc., such as video memory resources.

[0094] In a specific application scenario, step 104, calling the resources of the graphics processor of the target computing node to process the target task, specifically includes the following steps:

[0095] Step 104-1, if the resource requirement amount of the third task is less than or equal to the resource callable amount of the first graphics processor, call the resources of the first graphics processor to process the third task based on the resource requirement amount.

[0096] Specifically, for example, if a computer is configured with two GPU cards with 16GiB of free video memory each, the total GPU-MEM that can be reported to the K8s controller is 2 × 16 × 1024MiB. GPU-NUM is 2. When a task A of 500MiB needs to be executed, either GPU card 1 or GPU card 2 can be selected to execute task A. Taking the execution of task A on GPU card 1 as an example, during the task execution, the available resource amount of GPU card 1 is 1024 - 500 = 524MiB, and the available resource amount of GPU card 2 remains 1024MiB. If a task B of 100MiB needs to be executed at this time, since the resource requirement of task B, 100 < 524 < 1024, either GPU card 1 or GPU card 2 can be selected to execute task B when resources are sufficient. To further improve resource utilization, in the case of resource shortage, the difference between the resource requirement and the available resource amount can be calculated, and GPU card 1 with the smallest difference is selected to execute task B.

[0097] Step 104-2, if the resource requirement of the third task is greater than the available resource amount of each graphics processor of the target computing node, the resources of the second graphics processor are called based on the resource requirement to process the third task.

[0098] Among them, the third task is one of the target tasks, the first graphics processor is any graphics processor of the target computing node; the second graphics processor includes at least two graphics processors, and the sum of the resource requirements of the at least two graphics processors is greater than or equal to the resource requirement of the third task.

[0099] In this embodiment, the available resource amount of each graphics processor in the target computing node is obtained. When the resource requirement amount of the third task is less than or equal to the available resource amount of any one graphics processor, that is, the existing resources of this graphics processor are sufficient to execute the third task. At this time, based on the resource requirement amount, the resources of the first graphics processor are called to process the third task. On the contrary, if the resource requirement amount of the third task is greater than the available resource amount of any one graphics processor, that is, none of the graphics processors in this computing node can independently process the third task. At this time, based on the resource requirement amount, the resources of at least two graphics processors (the second graphics processor) are called to process the third task. On the one hand, it fully considers the different possible resource requirements of different tasks for the graphics processor, takes the resource amount of the graphics processor as the minimum scheduling unit, expands the current minimum unit of the K8s graphics processor scheduling, realizes a finer-grained division of the graphics processor resources, and provides more scheduling dimensions for different types of target tasks. On the other hand, resource scheduling is performed according to the size of the resource amount, which can better match tasks and available resources, thereby maximizing the utilization of the image processor resources, better realizing task parallelism. Even if a certain task requires more resources, other tasks with smaller resource requirements can still be executed in parallel on the idle resource control of this graphics processor, so as to achieve the purpose of flexible resource configuration, maximize the utilization rate of the graphics processor and the task throughput efficiency, greatly enhance the flexibility and scalability of resource scheduling, and then improve the system adaptability and resource utilization rate.

[0100] It is worth mentioning that if the third task runs in multiple container groups, and the task information of the third task has the specified graphics processor information of the container group, the resources of the graphics processor corresponding to the specified graphics processor information are called based on the resource requirement amount of the container group to start the container group; where the third task is one of the target tasks.

[0101] In this embodiment, since a task may run in multiple container groups to improve the concurrency ability and overall performance of the task. When the third task runs in multiple container groups and the task information of the third task has the specified graphics processor information of the container group, that is, the third task has specified different graphics processors to start different container groups (pods). Then, the resource scheduling according to the size of the resource amount is abandoned, and the container group is bound and started according to the resources of the graphics processor corresponding to the specified graphics processor information. Thus, the on-demand allocation of resources is realized to ensure that the target task can be correctly executed and the expected results can be obtained, and the application of the system in special scenarios is expanded.

[0102] Furthermore, in some possible embodiments, after step 104, the method further includes: if the target computing node fails to process the target task and the restart times are greater than the preset times, bind the target task to a new target computing node in the Kubernetes cluster.

[0103] In this embodiment, a restart policy can be set in the configuration of the target task to specify the handling method when the target task fails. When the target computing node fails to process the task and the number of restarts exceeds the preset number, the scheduling function of Kubernetes can be used to specify a new target computing node. By specifying a suitable target computing node for the Pod of the target task and rescheduling and binding the task on the new target computing node, the correct execution of the target task and the expected result can be ensured.

[0104] Specifically, for example, an extended scheduler for a graphics processor is deployed in the Kubernetes cluster. When the system needs to process multiple tasks, the extended scheduler filters out the target tasks that can be clustered from multiple tasks. According to the resource requirements of the target tasks, the extended scheduler automatically adds Pod-level affinity information to the target tasks and uses PodAffinity to specify that multiple Pods corresponding to the target tasks with the same label or attribute are assigned to the same node. With the cooperation of this mechanism, the clustering scheduling effect of the Pods can be achieved, and resource fragmentation can be reduced as much as possible.

[0105] Furthermore, the extended scheduler is deployed with a GPU-Share-Sched-Extender component and a GPU-Share-Plugin component. At the same time, the video memory resource unit of the Extended Resource is redesigned and defined. The first ExtendedResource can be defined as GPU-MEMORY, that is, the video memory requirement; the second can be defined as GPU-NUM, that is, the number requirement of GPU cards, so as to use the video memory requirement and the number requirement of GPU cards as the resource scheduling unit. When resource scheduling is required, the GPU-Share-Plugin component reports the GPU video memory resource amount and GPU card number information of the computing node. The GPU-Share-Sched-Extender component uses the total Extended Resource of the GPU cards reported by the GPU-Share-Plugin as the total amount of GPU card video memory (node_gpu_mem_quota). By listing the GPU allocation situations in all active Pods, the used situation of the GPU card video memory is obtained

[0106] (node_gpu_mem_used), and the GPU-MEM request volume of the current Pod (pod_gpu_mem_require). Then, by comparing the difference between node_gpu_mem_quota and node_gpu_mem_used with pod_gpu_mem_require, it is possible to determine whether the current video memory meets the video memory resource requirements. When the video memory resource requirements are met, the Pod can be bound to the GPU of this node through the GPU-Share-Plugin. Assuming that the required video memory of the Pod is small, then it is possible to achieve the scheduling of multiple containers on a single card according to the video memory requirements. Thus, it expands the minimum unit of the current K8s for GPU scheduling, provides more scheduling dimensions for different types of GPU tasks, and maximizes the GPU utilization rate and task throughput efficiency.

[0107] It should be noted that the magnitudes of the sequence numbers of the steps in the above embodiments do not mean the order of execution. The execution order of each process should be determined according to its function and internal logic, and should not constitute any limitation to the implementation process of the embodiments of the present application.

[0108] Furthermore, as Figure 2 shown, as a specific implementation of the above method for resource management of a graphics processor, an embodiment of the present application provides a resource management device 200 for a graphics processor. The resource management device 200 for a graphics processor includes: a screening module 201, a classification module 202, a scheduling module 203, and a processing module 204.

[0109] Among them, the screening module 201 is configured to determine target tasks that meet the clustering conditions among multiple tasks in response to execution instructions of the multiple tasks;

[0110] The classification module 202 is configured to perform classification processing on the target tasks;

[0111] The scheduling module 203 is configured to bind the target tasks of the same category to target computing nodes in a Kubernetes cluster. Among them, the Kubernetes cluster includes multiple computing nodes, each computing node includes at least one graphics processor, and the target computing node is one or more computing nodes in the available state among the multiple computing nodes;

[0112] The processing module 204 is configured to call the resources of the graphics processor of the target computing node to process the target tasks.

[0113] In this embodiment, when the system needs to process multiple tasks, target tasks that can be clustered together are screened out from the multiple tasks. The target tasks are classified and processed, and the target tasks of the same category are bound to the same target computing node in the Kubernetes cluster, so that the system can call the resources of the graphics processor of the target computing node to process the target tasks of the same category. Thus, a group of related tasks are assigned to the same computing node, effectively reducing the generation of resource fragmentation, reserving more overall machine resources for other tasks, and further achieving the purpose of making the most use of the computing node resources and reducing communication overhead. This facilitates the flexible use of GPU resources by users on the Kubernetes platform, greatly improving the utilization rate of graphics processor resources and the task throughput efficiency, expanding the usage scenarios, and enabling more flexible applications.

[0114] Further, the scheduling module 203 is specifically configured to obtain the total resource requirements of the target tasks of the same category and the resource supply of each available computing node in the Kubernetes cluster; if the total resource requirements are less than or equal to the resource supply of any available computing node, bind the target tasks of the same category to any computing node; if the total resource requirements are greater than the resource supply of each available computing node, determine the first computing node based on the priority of the available computing nodes in the Kubernetes cluster; bind the first task in the target tasks of the same category to the first computing node, where the total resource requirements of the first task are less than or equal to the resource supply of the first computing node; and bind the second tasks other than the first task in the target tasks of the same category to the second computing node, where the second computing node is an available computing node in the Kubernetes cluster other than the first computing node.

[0115] Further, the resource management device 200 of the graphics processor further includes: a first acquisition module (not shown in the figure), where the first acquisition module is used to obtain the number of graphics processors included in the computing node, the resource amount of each graphics processor, and the resource call amount of each graphics processor; a determination module (not shown in the figure), where the determination module is used to calculate the available resource call amount of each graphics processor based on the resource amount and the resource call amount; and determine the sum of the available resource call amounts of the graphics processors included in the computing node as the resource supply; and if the resource supply is greater than the preset resource amount, determine that the computing node is in an available state.

[0116] Further, the processing module 204 is specifically configured to, if the resource requirement of the third task is less than or equal to the available resources of the first graphics processor, call the resources of the first graphics processor to process the third task based on the resource requirement; if the resource requirement of the third task is greater than the available resources of each graphics processor of the target computing node, call the resources of the second graphics processor to process the third task based on the resource requirement; wherein, the third task is one of the target tasks, and the first graphics processor is any graphics processor of the target computing node; the second graphics processor includes at least two graphics processors, and the sum of the resource requirements of the at least two graphics processors is greater than or equal to the resource requirement of the third task.

[0117] Further, the processing module 204 is specifically configured to, if the third task runs in multiple container groups and the task information of the third task has the specified graphics processor information of the container group, call the resources of the graphics processor corresponding to the specified graphics processor information based on the resource requirement of the container group to start the container group; wherein, the third task is one of the target tasks.

[0118] Further, the resource management device 200 of the graphics processor further includes: a second acquisition module (not shown in the figure), and the second acquisition module is configured to acquire the task information of multiple tasks, wherein the task information includes at least one of the following: task type information, resource requirement information, task attribution information, and task attribute information; a classification module 202, specifically configured to perform classification processing on the target tasks based on the task information to determine the category of the target tasks.

[0119] Further, the scheduling module 203 is further configured to, if the target computing node fails to process the target task and the restart times are greater than the preset times, bind the target task to a new target computing node in the Kubernetes cluster.

[0120] For the specific limitations of the resource management device of the graphics processor, reference may be made to the limitations of the resource management method of the graphics processor in the foregoing text, which will not be elaborated here. Each module in the above-mentioned resource management device of the graphics processor can be implemented in whole or in part by software, hardware, and their combination. The above-mentioned modules can be embedded in the processor of the computer device in the form of hardware or be independent of it, or can be stored in the memory of the computer device in the form of software, so as to facilitate the processor to call and execute the operations corresponding to the above-mentioned modules.

[0121] Based on the above method as Figure 1 shown, correspondingly, an embodiment of the present application further provides a readable storage medium, on which a computer program is stored, and when the program is executed by a processor, it implements the resource management method of the graphics processor as Figure 1 shown above.

[0122] Based on such understanding, the technical solution of the present application can be embodied in the form of a software product, which can be stored in a non-volatile storage medium (such as a CD-ROM, a USB flash drive, a portable hard drive, etc.), and includes a number of instructions for causing a computer device (such as a personal computer, a server, or a network device, etc.) to execute the methods described in various implementation scenarios of the present application.

[0123] The storage medium may further include an operating system and a network communication module. The operating system is a program for managing and storing the hardware and software resources of the computer device, and supports the operation of information processing programs and other software and / or programs. The network communication module is used to implement the communication between the components inside the storage medium, as well as the communication between the storage medium and other hardware and software in the entity device.

[0124] Based on the method as described above Figure 1 and Figure 2 the virtual device embodiment as described above, for the purpose of achieving the above object, the embodiments of the present application further provide a computer device, which may specifically be a personal computer, a server, a network device, etc. The computer device includes at least one processor; the processor is coupled to the memory, and the memory stores a computer program that runs on the processor; the processor executes the computer program to implement the resource management method of the graphics processor as described above Figure 1 as shown.

[0125] Optionally, the computer device may further include a user interface, a network interface, a camera, a radio frequency (RF) circuit, sensors, an audio circuit, a WI-FI module, etc. The user interface may include a display screen (Display) and an input unit such as a keyboard (Keyboard), etc. Optionally, the user interface may further include a USB interface, a card reader interface, etc. The network interface may optionally include a standard wired interface, a wireless interface (such as a Bluetooth interface, a WI-FI interface), etc.

[0126] Those skilled in the art can understand that the structure of a computer device provided in this embodiment does not constitute a limitation on the computer device, and may include more or fewer components, or combine certain components, or have different component arrangements.

[0127] Through the description of the above embodiments, those skilled in the art can clearly understand that this application can be implemented by means of software plus a necessary general hardware platform, or can also be implemented by hardware. In response to execution instructions of multiple tasks, determine target tasks that meet clustering conditions among the multiple tasks; classify the target tasks; bind the target tasks of the same category to target computing nodes in the Kubernetes cluster, where the Kubernetes cluster includes multiple computing nodes, each computing node includes at least one graphics processor, and the target computing node is one or more computing nodes in the available state among the multiple computing nodes; call the resources of the graphics processor of the target computing node to process the target tasks. In the embodiments of this application, when the system needs to process multiple tasks, filter out target tasks that can be clustered together from the multiple tasks. Divide the target tasks through classification processing, and bind the target tasks of the same category to the same target computing node in the Kubernetes cluster, so that the system can call the resources of the graphics processor of the target computing node to process the target tasks of the same category. Thus, a group of related tasks are allocated to the same computing node, effectively reducing the generation of resource fragmentation, reserving more overall machine resources for other tasks, and further achieving the purpose of making as much use of the computing node resources as possible and reducing communication overhead, facilitating users to flexibly use GPU resources on the Kubernetes platform, greatly improving the utilization rate of graphics processor resources and task throughput efficiency, broadening the usage scenarios, and being able to be applied more flexibly.

[0128] Those skilled in the art can understand that the drawings are only schematic diagrams of a preferred implementation scenario, and the modules or processes in the drawings are not necessarily essential for implementing this application. Those skilled in the art can understand that the modules in the device in the implementation scenario can be distributed in the device in the implementation scenario according to the description of the implementation scenario, or can be changed accordingly and located in one or more devices different from this implementation scenario. The modules in the above implementation scenario can be combined into one module, or further split into multiple sub-modules.

[0129] The above serial numbers of this application are only for description and do not represent the advantages or disadvantages of the implementation scenarios. The above disclosure is only several specific implementation scenarios of this application. However, this application is not limited thereto, and any changes that can be thought of by those skilled in the art should fall within the protection scope of this application.

Claims

1. A resource management method for a graphics processor, characterized in that The method includes: In response to execution instructions of multiple tasks, determining target tasks that meet clustering conditions among the multiple tasks; Performing classification processing on the target tasks; Binding the target tasks of the same category to target computing nodes in a Kubernetes cluster, where the Kubernetes cluster includes multiple computing nodes, each computing node includes at least one graphics processor, and the target computing nodes are one or more of the computing nodes in the available state among the multiple computing nodes; Invoking resources of the graphics processors of the target computing nodes to process the target tasks.

2. The resource management method of the graphics processor according to claim 1, characterized in that The binding of the target tasks of the same category to target computing nodes in the Kubernetes cluster specifically includes: Obtaining the total resource requirements of the target tasks of the same category and the resource supply amounts of each available computing node in the Kubernetes cluster; If the total resource requirements are less than or equal to the resource supply amount of any available computing node, binding the target tasks of the same category to the any available computing node; If the total resource requirements are greater than the resource supply amounts of each available computing node, determining a first computing node based on the priorities of the available computing nodes in the Kubernetes cluster; Binding a first task among the target tasks of the same category to the first computing node, where the total resource requirements of the first task are less than or equal to the resource supply amount of the first computing node; Binding a second task other than the first task among the target tasks of the same category to a second computing node, where the second computing node is an available computing node other than the first computing node in the Kubernetes cluster.

3. The resource management method of the graphics processor according to claim 1, wherein The method further includes: Obtaining the number of graphics processors included in the computing node, the resource amount of each graphics processor, and the resource invocation amount of each graphics processor; Based on the resource amount and the resource invocation amount, respectively calculating the resource invokable amounts of each graphics processor; Determining the sum of the resource invokable amounts of the graphics processors included in the computing node as the resource supply amount; If the resource supply amount is greater than a preset resource amount, determining that the computing node is in an available state.

4. The resource management method of the graphics processor according to claim 1, wherein The invoking of resources of the graphics processors of the target computing nodes to process the target tasks specifically includes: If the resource requirement amount of a third task is less than or equal to the resource invokable amount of a first graphics processor, invoking the resources of the first graphics processor to process the third task based on the resource requirement amount; If the resource requirement amount of the third task is greater than the resource invokable amounts of each graphics processor of the target computing node, invoking the resources of a second graphics processor to process the third task based on the resource requirement amount; Among them, the third task is one of the target tasks, and the first graphics processor is any one of the graphics processors of the target computing node; the at least two graphics processors, the second graphics processor includes at least two graphics processors, and the sum of the resource requirements of the at least two graphics processors is greater than or equal to the resource requirements of the third task.

5. The resource management method of the graphics processor according to claim 1, characterized in that The calling of the resources of the graphics processor of the target computing node to process the target task specifically includes: If the third task runs in multiple container groups and the task information of the third task has the specified graphics processor information of the container group, call the resources of the graphics processor corresponding to the specified graphics processor information based on the resource requirements of the container group to start the container group; Among them, the third task is one of the target tasks.

6. The resource management method of a graphics processor according to any one of claims 1 to 5, characterized in that The classification processing of the target task specifically includes: Obtain the task information of multiple tasks, where the task information includes at least one of the following: task type information, resource requirement information, task attribution information, and task attribute information; Based on the task information, perform classification processing on the target task to determine the category of the target task.

7. The resource management method of the graphics processor according to any one of claims 1 to 5, characterized in that, After the calling of the resources of the graphics processor of the target computing node to process the target task, the method further includes: If the target computing node fails to process the target task and the number of restart times is greater than the preset number of times, bind the target task to a new target computing node in the Kubernetes cluster.

8. A resource management device for a graphics processor, characterized in that The device includes: A screening module, configured to determine a target task that meets the clustering conditions among the multiple tasks in response to execution instructions of the multiple tasks; A classification module, configured to perform classification processing on the target task; A scheduling module, configured to bind the target tasks of the same category to a target computing node in the Kubernetes cluster, where the Kubernetes cluster includes multiple computing nodes, each computing node includes at least one graphics processor, and the target computing node is one or more of the multiple computing nodes in an available state; A processing module, configured to call the resources of the graphics processor of the target computing node to process the target task.

9. A computer-readable storage medium having a program or instructions stored thereon, characterized in that, When the program or instruction is executed by a processor, it implements the steps of the resource management method of the graphics processor according to any one of claims 1 to 7.

10. A computer device, comprising at least one processor, the processor being coupled to a memory, the memory storing a computer program that runs on the processor, characterized in that, When the processor executes the program, it implements the resource management method of the graphics processor according to any one of claims 1 to 7.

Citation Information

Cited By

  • Graphic processor resource management system, method and server

    CN120765447A

  • Graphic processor resource allocation method

    CN120892216A