GPU pass-through and resource hybrid scheduling method and system, and chip

By acquiring and processing GPU resource information, using mixed quantization calculation and full permutation combination functions for screening, GPU direct-through and resource hybrid scheduling are realized, solving the problem of low GPU resource utilization in traditional technology, and improving resource utilization and training efficiency.

WO2025124260A1PCT designated stage expired Publication Date: 2025-06-19CHINA TELECOM CLOUD TECH CO LTD

Patent Information

Application Number
PCT/CN2024/136816
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-12-15
Filing Date
2024-12-04
Publication Date
2025-06-19

AI Technical Summary

Technical Problem

Traditional technologies cannot perform efficient resource hybrid scheduling in the case of GPU direct access, resulting in low GPU resource utilization, resulting in waste of resources and inefficient training.

Method used

By obtaining the GPU resource information of each node, using mixed quantization calculation to obtain computing power data and video memory requirements information, combining the full arrangement combination function and pruning function for screening, obtaining the results of the preferred nodes, and realizing GPU direct-through and resource mixed scheduling.

Benefits of technology

It improves the utilization rate and utilization rate of GPU resources, saves manpower and material costs, has excellent scheduling performance and high accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2024136816_19062025_PF_FP_ABST
    Figure CN2024136816_19062025_PF_FP_ABST
Patent Text Reader

Abstract

The present application provides a GPU pass-through and resource hybrid scheduling method and system, and a chip. The method comprises: S1, acquiring resource information of GPUs of nodes, and storing the resource information in a scheduling buffer; S2, on the basis of hybrid quantization computation, acquiring computing power data and total display memory demand information of the GPUs of the nodes; S3, on the basis of the total display memory demand information of the GPUs of the nodes, performing first screening on the nodes to obtain a first screening result; S4, on the basis of a full permutation combination function and a pruning function, acquiring a simplified two-dimensional full permutation array; and S5, on the basis of the first screening result, performing a second screening operation on the nodes to obtain a preferred node result.
Need to check novelty before this filing date? Find Prior Art

Description

GPU direct and resource hybrid scheduling method, system and chip

[0001] Related applications

[0002] This application claims priority to Chinese patent application number 202311734812.4, filed on December 15, 2023, entitled “GPU direct pass-through and resource hybrid scheduling method, system and chip,” the entire text of which is hereby incorporated by reference. Technical Field

[0003] The present application relates to the technical field of cloud-native containers, and in particular to a GPU pass-through and resource hybrid scheduling method, system, and chip. Background Art

[0004] Graphics Processing Units (GPUs) have a large number of cores and high-speed memory, excel at parallel computing, and are suitable for scenarios such as deep learning training, inference, and high-performance scientific computing. Currently, many AI applications require the use of a large number of GPU resources to accelerate training. With more and more GPU resources and training tasks, the utilization and management of GPU resources has become difficult. For example, during daytime working hours, the demand for GPU resources is very tight, and it is often difficult to get a card. However, at night, a considerable number of GPU cards are idle, resulting in unnecessary waste of resources. Therefore, most companies and universities have adopted K8s (Kubernetes) as a unified orchestration and scheduling management platform to automate and unify the management of training resources, thereby improving the utilization of GPU resources and work efficiency.

[0005] Traditional technologies separate the two scenarios of GPU direct passthrough scheduling and virtualized scheduling, making the two scheduling scenarios unable to coexist. When GPU direct passthrough is enabled, the most efficient computing resources can be obtained, but virtualized scheduling cannot be enabled for the GPU. Often, a cluster's resources are fully occupied early, shutting out some requests that only require a small portion of computing power and video memory. When virtualized scheduling is enabled, direct passthrough of the entire card cannot be taken into account, resulting in some direct passthrough requests that require a large amount of computing resources only obtaining a portion of the computing power, resulting in inefficient training. This limits the scenarios in which GPUs can be used and causes unnecessary trouble. Traditional technologies are unable to perform efficiently with GPU direct passthrough. Therefore, how to perform GPU direct passthrough and mixed resource scheduling has become an urgent problem that needs to be solved.

[0006] Patent CN114816746A discloses a method for implementing mixed-type virtualization on GPUs with multiple virtualization types, including: monitoring, scheduling, and allocation of vGPU resources; vGPU resource monitoring is used to monitor GPU health and set monitoring indicators; scheduling is based on the monitoring indicators, obtaining useful data from them for scheduling to vGPU resources; scheduling includes filtering and evaluation: filtering selects the corresponding GPUs based on requirements, which is achieved through multiple filters; evaluation is the process of scoring the evaluation indicators among multiple resources that still meet the requirements after filtering, and finally selecting the highest-scoring resource for scheduling; allocation is achieved by adding multiple mdev devices of different vGPU types to the same PCI device directory to allocate mixed-type vGPUs. This method is inefficient for scheduling and cannot fully accommodate full-card passthrough.

[0007] Based on this, the present application provides a GPU direct and resource hybrid scheduling method, system and chip to improve traditional technology. Summary of the Invention

[0008] The purpose of this application is to provide a GPU direct and resource hybrid scheduling method, system and chip. By no longer distinguishing between GPU direct and virtualized scheduling scenarios, a new hybrid scheduling method can be used to expand the application scenarios of GPU scheduling, obtain the optimal node, further improve the utilization rate and utilization rate of GPU resources, save manpower and material costs, and have the advantages of better scheduling performance and high scheduling accuracy.

[0009] The purpose of this application is achieved by the following technical solutions:

[0010] In a first aspect, the present application provides a GPU direct and resource hybrid scheduling method, the method comprising:

[0011] S1: Obtain resource information of the GPU of each node and save it to the scheduling buffer area. The resource information includes total available resource information and remaining resource information;

[0012] S2: Based on hybrid quantization calculation, obtain the computing power data and total video memory requirement information of each node's GPU;

[0013] S3: Based on the total video memory requirement information of the GPU of each node, perform a first screening on the nodes to obtain a first screening result;

[0014] S4: Based on the full permutation combination function and the pruning function, a simplified two-dimensional full permutation array is obtained;

[0015] S5: Based on the first screening result, a second screening operation is performed on the node to obtain a preferred node result.

[0016] The beneficial effects of this technical solution are as follows: first, the resource information of the GPU of each node is obtained and saved in the scheduling cache; based on hybrid quantization calculation, the computing power data and total memory requirement information of the GPU of each node are obtained; secondly, based on the total memory requirement information of the GPU of each node, the nodes are first screened to obtain a first screening result; then, based on the full permutation combination function and pruning function, a simplified two-dimensional full permutation array is obtained; finally, based on the first screening result, a second screening operation is performed on the nodes to obtain the preferred node result. This method can expand the application scenarios of GPU scheduling by no longer distinguishing between GPU direct and virtualization scheduling scenarios, using a new hybrid scheduling method to obtain the optimal node, further improve the utilization rate and utilization of GPU resources, save manpower and material costs, and has the advantages of better scheduling performance and high scheduling accuracy.

[0017] In some embodiments, the obtaining and storing of resource information of the GPU of each node, wherein the resource information includes total available resource information and remaining resource information, includes:

[0018] Use the device-plugin component to obtain the attribute information of the GPU corresponding to each node, including the GPU information reporting time and GPU status information;

[0019] When the periodic scheduler starts scheduling, determining whether the GPU information reporting time is greater than the preset reporting time; if the GPU information reporting time is greater than the preset reporting time, setting the GPU to an unavailable state;

[0020] If the GPU information reporting time is not greater than the preset reporting time, determining whether the GPU status information is in a ready state; if the GPU status information is not in a ready state, setting the GPU to an unavailable state;

[0021] If the GPU state information is in a ready state, traversing the attribute information of the GPU to obtain the total available resource information and saving it to a scheduling buffer;

[0022] Based on the total available resource information, the remaining resource information is obtained and saved in a scheduling buffer area.

[0023] The beneficial effect of this technical solution is that, through the device-plugin component, the attribute information of the GPU corresponding to each node can be obtained in real time, including the GPU information reporting time and GPU status information. This helps to understand the real-time usage of GPU resources in the cluster and provides accurate data support for task scheduling. When the periodic scheduler starts scheduling, it can determine whether the GPU resources are available based on the preset reporting time. If the GPU information reporting time is greater than the preset reporting time, or the GPU status information is not in the ready state, the GPU is set to the unavailable state. This ensures that the scheduler does not schedule tasks to unavailable GPU resources, thereby improving task success rate. When the GPU status information is in the ready state, the GPU attribute information can be traversed to obtain the total available resource information and save it to the scheduling buffer. Then, based on the total available resource information, the remaining resource information can be obtained and saved to the scheduling buffer. This helps to better understand the remaining GPU resources in the cluster, thereby achieving optimal resource allocation during task scheduling and improving cluster resource utilization. The device-plugin component allows customized GPU resource management policies according to actual needs, such as setting different reporting times and preset states. This makes the technical solution highly flexible and adaptable to the GPU resource management needs of different scenarios. In summary, the technical solution of implementing GPU resource management using the device-plugin component has beneficial effects such as real-time monitoring, periodic scheduling, optimized resource utilization, and flexible adaptation to different scenarios, which helps to improve the utilization of GPU resources and task scheduling efficiency in Kubernetes clusters.

[0024] In some embodiments, obtaining the remaining resource information based on the total available resource information and saving it to a scheduling buffer area includes:

[0025] Get the GPU index and GPU usage corresponding to the application on each node;

[0026] Obtaining used resource information based on the GPU index and GPU usage corresponding to the application on each node;

[0027] Based on the total available resource information, the used resource information is removed from the scheduling buffer area to obtain the remaining resource information and save it in the scheduling buffer area.

[0028] The beneficial effect of this technical solution is that it can grasp the GPU usage of applications on each node in detail, including the GPU index used and the usage, which helps to more accurately understand the usage of GPU resources in the cluster. Based on the GPU index and GPU usage corresponding to the application on each node, the used resource information can be obtained. This helps to more accurately calculate the used resources when performing resource scheduling and avoid duplicate resource allocation. When obtaining the remaining resource information, the remaining resource information can be obtained and saved in the scheduling buffer by removing the used resource information in the scheduling buffer. This helps to better understand the remaining resources in the cluster, thereby achieving optimal resource allocation during task scheduling and improving cluster resource utilization. By refining resource monitoring and optimizing resource scheduling, it can be ensured that tasks can be executed efficiently in the cluster, improving task execution efficiency. In summary, the technical solution for obtaining the GPU index and GPU usage corresponding to the application on each node has the beneficial effects of refined management, accurate calculation of used resources, optimization of resource scheduling and improvement of task execution efficiency, which helps to improve the resource management efficiency and task execution efficiency of the cluster.

[0029] In some optional implementations, obtaining the computing power data and total video memory requirement information of the GPU of each node based on hybrid quantization calculation includes:

[0030] Determine whether the application on each node has GPU requirements;

[0031] If the application on a node has GPU requirements, the GPU index corresponding to the application on the node is calculated and saved; if the application on a node does not have GPU requirements, no operation is performed on the application on the node;

[0032] Determine whether the GPU requirement is a whole-card pass-through mode; if the GPU requirement is a whole-card pass-through mode, obtain computing power data and total video memory requirement information of the GPU of the node through a multiplication operation based on the GPU index;

[0033] If the GPU requirement is not the whole card pass-through mode, the computing power data and total video memory requirement information of the GPU of the node are obtained by an addition operation based on the GPU index.

[0034] The beneficial effect of this technical solution is that it determines whether the application on each node needs GPU resources. For applications with GPU requirements, the corresponding GPU index will be calculated and saved, while for applications without GPU requirements, no operation will be performed. In this way, resource management can be carried out according to the personalized needs of the application to improve resource utilization efficiency. For applications with GPU requirements, it will be further determined whether they need the whole-card pass-through mode. If the requirement is the whole-card pass-through mode, the computing power data and total video memory requirement information of the node's GPU will be obtained based on the GPU index through multiplication operations; if the requirement is not the whole-card pass-through mode, the corresponding information will be obtained through addition operations. In this way, the demand for GPU resources of applications on each node can be accurately calculated, providing accurate data support for resource scheduling. The technical solution can adapt to the GPU resource requirements in different scenarios. Whether it is the whole-card pass-through mode or other modes, the required resource information can be obtained through corresponding operations, and it has strong flexibility.

[0035] In some embodiments, the first screening of the nodes based on the total video memory requirement information of the GPUs of the nodes to obtain a first screening result includes:

[0036] Comparing the total graphics memory requirement information of the GPU of the node with the remaining resource information;

[0037] If the total memory requirement information of the GPU of the node is less than the remaining resource information, the node is considered to have insufficient resources and is filtered;

[0038] If the total graphics memory requirement information of the GPU of the node is not less than the remaining resource information, the node is saved to obtain the first screening result.

[0039] The beneficial effect of this technical solution is that if the node's total GPU memory requirement information is less than the remaining resource information, it means that the node's resources are insufficient to meet the new task requirements. In this case, the node can be filtered out to avoid scheduling tasks to nodes with insufficient resources, thereby improving the success rate of the task. If the node's total GPU memory requirement information is not less than the remaining resource information, it means that the node's resources can meet the new task requirements. In this case, the node will be saved and become the first screening result. In this way, it can ensure that nodes with sufficient resources are fully utilized and improve the resource utilization of the cluster. By comparing the node's total GPU memory requirement information and the remaining resource information, it can be quickly determined which nodes have sufficient resources to execute new tasks. This helps to improve task scheduling efficiency and reduce task failures caused by insufficient resources. It can adapt to the GPU resource requirements in different scenarios. Whether it is memory requirements or other requirements, it can be determined whether the node's resources are sufficient through comparison operations, which has strong flexibility.

[0040] In some embodiments, obtaining a simplified two-dimensional full permutation array based on the full permutation combination function and the pruning function includes:

[0041] Input the preset application information and get the index array with GPU;

[0042] Inputting the index array into a preset full permutation function to obtain a full permutation two-dimensional array;

[0043] Determine whether the computing power data of the GPU and the video memory data of the GPU are equal;

[0044] If the computing power data of the GPU and the video memory data of the GPU are equal, marking the index array of the GPU to obtain an equivalent array;

[0045] The fully permuted two-dimensional array and the equivalent array are input into a preset pruning function to obtain the simplified two-dimensional fully permuted array.

[0046] The beneficial effect of this technical solution is that by inputting preset application information, an index array with a GPU can be obtained according to the specific needs of the application, thereby realizing personalized resource management for different applications. By inputting the index array into the preset full permutation function, a fully permuted two-dimensional array can be obtained, so that when scheduling tasks, all possible GPU resource allocation schemes can be fully considered, thereby improving resource utilization and task scheduling efficiency. After determining whether the GPU computing power data and the GPU memory data are equal, if they are equal, the GPU index array is marked to obtain an equivalent array. Then, the fully permuted two-dimensional array and the equivalent array are input into the preset pruning function together to obtain a streamlined two-dimensional fully permuted array. In this way, precise pruning can be performed to remove GPU resource allocation schemes that do not meet the conditions, thereby improving task scheduling efficiency and resource utilization. The technical solution of inputting preset application information and obtaining an index array with a GPU, and then performing full permutation and pruning operations has the beneficial effects of personalized resource management, full permutation optimization, precise pruning, and adaptability to different scenarios, which helps to improve the resource management efficiency and task execution efficiency of the cluster.

[0047] In some embodiments, performing a second screening operation on the nodes based on the first screening result to obtain a preferred node result includes:

[0048] Based on the first screening result, performing comparative screening on the reduced two-dimensional full permutation array to obtain a comparative screening result;

[0049] Based on the comparison and screening results, the nodes are scored and optimized to obtain the preferred node results.

[0050] The beneficial effect of this technical solution is that: based on the first screening result, the simplified two-dimensional full permutation array is compared and screened to obtain a comparison screening result. In this way, nodes that meet the conditions can be accurately screened, improving task scheduling efficiency and resource utilization. Based on the comparison screening results, the nodes are scored and optimized to obtain the preferred node results. In this way, the best execution node can be selected for the task based on the node's resource status and task requirements, improving task execution efficiency and cluster resource utilization. By accurately screening and optimizing node selection, it can be ensured that tasks can be executed efficiently and smoothly in the cluster, improving the task execution effect.

[0051] In a second aspect, the present application provides an electronic device, comprising a memory and a processor, wherein the memory stores a computer program, and the processor is configured to implement the following steps when executing the computer program:

[0052] S1: Obtain resource information of the GPU of each node and save it to the scheduling buffer area. The resource information includes total available resource information and remaining resource information;

[0053] S2: Based on hybrid quantization calculation, obtain the computing power data and total video memory requirement information of each node's GPU;

[0054] S3: Based on the total video memory requirement information of the GPU of each node, perform a first screening on the nodes to obtain a first screening result;

[0055] S4: Based on the full permutation combination function and the pruning function, a simplified two-dimensional full permutation array is obtained;

[0056] S5: Based on the first screening result, a second screening operation is performed on the node to obtain a preferred node result.

[0057] In some embodiments, the processor is configured to obtain and save resource information of the GPU of each node in the following manner when executing the computer program, wherein the resource information includes total available resource information and remaining resource information:

[0058] Use the device-plugin component to obtain the attribute information of the GPU corresponding to each node, including the GPU information reporting time and GPU status information;

[0059] When the periodic scheduler starts scheduling, determining whether the GPU information reporting time is greater than the preset reporting time; if the GPU information reporting time is greater than the preset reporting time, setting the GPU to an unavailable state;

[0060] If the GPU information reporting time is not greater than the preset reporting time, determining whether the GPU status information is in a ready state; if the GPU status information is not in a ready state, setting the GPU to an unavailable state;

[0061] If the GPU state information is in a ready state, traversing the attribute information of the GPU to obtain the total available resource information and saving it to a scheduling buffer;

[0062] Based on the total available resource information, the remaining resource information is obtained and saved in a scheduling buffer area.

[0063] In some embodiments, the processor is configured to, when executing the computer program, obtain the remaining resource information based on the total available resource information and save it to the scheduling buffer in the following manner:

[0064] Get the GPU index and GPU usage corresponding to the application on each node;

[0065] Obtaining used resource information based on the GPU index and GPU usage corresponding to the application on each node;

[0066] Based on the total available resource information, the used resource information is removed from the scheduling buffer area to obtain the remaining resource information and save it in the scheduling buffer area.

[0067] In some embodiments, the processor is configured to obtain computing power data and total video memory requirement information of the GPU of each node based on hybrid quantization computing in the following manner when executing the computer program:

[0068] Determine whether the application on each node has GPU requirements;

[0069] If the application on a node has GPU requirements, the GPU index corresponding to the application on the node is calculated and saved; if the application on a node does not have GPU requirements, no operation is performed on the application on the node;

[0070] Determine whether the GPU requirement is a whole-card pass-through mode; if the GPU requirement is a whole-card pass-through mode, obtain computing power data and total video memory requirement information of the GPU of the node through a multiplication operation based on the GPU index;

[0071] If the GPU requirement is not the whole card pass-through mode, the computing power data and total video memory requirement information of the GPU of the node are obtained by an addition operation based on the GPU index.

[0072] In some embodiments, the processor is configured to, when executing the computer program, perform a first screening on the nodes based on the total video memory requirement information of the GPU of each node to obtain a first screening result in the following manner:

[0073] Comparing the total graphics memory requirement information of the GPU of the node with the remaining resource information;

[0074] If the total memory requirement information of the GPU of the node is less than the remaining resource information, the node is considered to have insufficient resources and is filtered;

[0075] If the total graphics memory requirement information of the GPU of the node is not less than the remaining resource information, the node is saved to obtain the first screening result.

[0076] In some embodiments, the processor is configured to obtain a reduced two-dimensional full permutation array based on a full permutation combination function and a pruning function in the following manner when executing the computer program:

[0077] Input the preset application information and get the index array with GPU;

[0078] Inputting the index array into a preset full permutation function to obtain a full permutation two-dimensional array;

[0079] Determine whether the computing power data of the GPU and the video memory data of the GPU are equal;

[0080] If the computing power data of the GPU and the video memory data of the GPU are equal, marking the index array of the GPU to obtain an equivalent array;

[0081] The fully permuted two-dimensional array and the equivalent array are input into a preset pruning function to obtain the simplified two-dimensional fully permuted array.

[0082] In some embodiments, the processor is configured to perform a second screening operation on the node based on the first screening result to obtain a preferred node result in the following manner when executing the computer program:

[0083] Based on the first screening result, performing comparative screening on the reduced two-dimensional full permutation array to obtain a comparative screening result;

[0084] Based on the comparison and screening results, the nodes are scored and optimized to obtain the preferred node results.

[0085] In a third aspect, the present application provides a scheduling system, wherein the electronic device includes:

[0086] The above electronic equipment.

[0087] In a fourth aspect, the present application provides a chip storing a computer program, which implements the method steps in any one of the embodiments of the first aspect when executed by a processor. BRIEF DESCRIPTION OF THE DRAWINGS

[0088] To more clearly illustrate the technical solutions in the embodiments of this application or related technologies, the following briefly introduces the drawings required for use in the embodiments or related technology descriptions. Obviously, the drawings described below are only some embodiments of this application. For those of ordinary skill in the art, other drawings can be derived from these drawings without inventive effort.

[0089] FIG1 shows a flow chart of a GPU direct and resource hybrid scheduling method provided in one embodiment of the present application.

[0090] FIG2 shows a schematic diagram of a process for obtaining resource information provided by an embodiment of the present application.

[0091] FIG3 shows a schematic diagram of a process for obtaining remaining resource information provided by an embodiment of the present application.

[0092] FIG4 shows a schematic diagram of a process for obtaining a simplified two-dimensional full permutation array according to an embodiment of the present application.

[0093] FIG5 shows a structural framework diagram of an electronic device provided in an embodiment of the present application.

[0094] FIG6 shows a schematic structural diagram of a scheduling system provided in an embodiment of the present application.

[0095] FIG7 shows a schematic structural diagram of a program product provided in an embodiment of the present application. DETAILED DESCRIPTION

[0096] The following is a further description of the embodiments of the present application in conjunction with the accompanying drawings and specific implementation methods. It should be noted that, under the premise of no conflict, the embodiments or technical features described below can be arbitrarily combined to form a new implementation method.

[0097] In the embodiments of the present application, "at least one" refers to one or more, and "more" refers to two or more. "And / or" describes the association relationship of associated objects, indicating that three relationships may exist. For example, A and / or B can represent: the existence of A alone, the existence of A and B at the same time, and the existence of B alone, where A and B can be singular or plural. The character " / " generally indicates that the objects associated before and after are in an "or" relationship. "At least one of the following" or similar expressions refers to any combination of these items, including any combination of single or plural items. For example, at least one of a, b or c can represent: a, b, c, a and b, a and c, b and c, a and b and c, where a, b and c can be single or multiple. It is worth noting that "at least one" can also be interpreted as "one or more items".

[0098] It should also be noted that in the embodiments of this application, words such as "exemplary" or "for example" are used to indicate examples, illustrations, or descriptions. Any implementation or design scheme described as "exemplary" or "for example" in the embodiments of this application should not be interpreted as being preferred or advantageous over other implementations or designs. Rather, the use of words such as "exemplary" or "for example" is intended to present the relevant concepts in a concrete manner.

[0099] Method Example

[0100] Referring to FIG1 , FIG1 shows a flow chart of a GPU direct-through and resource hybrid scheduling method provided in an embodiment of the present application.

[0101] This application provides a GPU direct and resource hybrid scheduling method, the method comprising:

[0102] S1: Obtain resource information of the GPU of each node and save it to the scheduling buffer area. The resource information includes total available resource information and remaining resource information;

[0103] S2: Based on hybrid quantization calculation, obtain the computing power data and total video memory requirement information of each node's GPU;

[0104] S3: Based on the total video memory requirement information of the GPU of each node, perform a first screening on the nodes to obtain a first screening result;

[0105] S4: Based on the full permutation combination function and the pruning function, a simplified two-dimensional full permutation array is obtained;

[0106] S5: Based on the first screening result, a second screening operation is performed on the node to obtain a preferred node result.

[0107] Therefore, first, the resource information of each node's GPU is obtained and saved in the scheduling cache; based on hybrid quantization calculations, the computing power data and total memory requirement information of each node's GPU are obtained; secondly, based on the total memory requirement information of each node's GPU, the nodes are first screened to obtain a first screening result; then, based on the full permutation combination function and pruning function, a simplified two-dimensional full permutation array is obtained; finally, based on the first screening result, a second screening operation is performed on the nodes to obtain the preferred node result. This method can expand the application scenarios of GPU scheduling by no longer distinguishing between GPU direct and virtualization scheduling scenarios, using a new hybrid scheduling method to obtain the optimal node, further improve the utilization rate and utilization of GPU resources, save manpower and material costs, and has the advantages of better scheduling performance and high scheduling accuracy.

[0108] Refer to Figure 2, which shows a schematic diagram of a process for obtaining resource information provided by an embodiment of the present application.

[0109] In some optional implementations, the obtaining and saving of resource information of the GPU of each node, wherein the resource information includes total available resource information and remaining resource information, includes:

[0110] Step S201: using the device-plugin component to obtain attribute information of the GPU corresponding to each node, wherein the attribute information of the GPU includes GPU information reporting time and GPU status information;

[0111] Step S202: When the periodic scheduler starts scheduling, it is determined whether the GPU information reporting time is greater than the preset reporting time; if the GPU information reporting time is greater than the preset reporting time, the GPU is set to an unavailable state;

[0112] Step S203: If the GPU information reporting time is not greater than the preset reporting time, determining whether the GPU status information is in a ready state; if the GPU status information is not in a ready state, setting the GPU to an unavailable state;

[0113] Step S204: If the GPU status information is in a ready state, traverse the attribute information of the GPU to obtain the total available resource information and save it in a scheduling buffer;

[0114] Step S205: The available total resource information, obtains the remaining resource information and saves it in the scheduling buffer area.

[0115] The device-plugin component enables real-time access to GPU attribute information for each node, including GPU reporting time and GPU status information. This helps understand the real-time usage of GPU resources in the cluster and provides accurate data support for task scheduling. When the periodic scheduler starts scheduling, it can determine the availability of GPU resources based on the preset reporting time. If the GPU reporting time exceeds the preset reporting time, or the GPU status information is not in the ready state, the GPU is set to unavailable. This ensures that the scheduler does not schedule tasks to unavailable GPU resources, improving task success rates. When the GPU status information is in the ready state, the GPU attribute information is traversed to obtain the total available resource information and save it to the scheduling buffer. Then, based on the total available resource information, the remaining resource information is obtained and saved to the scheduling buffer. This helps better understand the remaining GPU resources in the cluster, enabling optimal resource allocation during task scheduling and improving cluster resource utilization. The device-plugin component allows you to customize GPU resource management policies according to your needs, such as setting different reporting times and preset states. This makes the technical solution highly flexible and adaptable to the GPU resource management needs of different scenarios. In summary, the technical solution of implementing GPU resource management using the device-plugin component has beneficial effects such as real-time monitoring, periodic scheduling, optimized resource utilization, and flexible adaptation to different scenarios, which helps to improve the utilization of GPU resources and task scheduling efficiency in Kubernetes clusters.

[0116] In one specific embodiment, the total GPU resources and remaining resources of each node are counted and stored. This specifically includes selecting a node and obtaining GPU information reported by the node's device-plugin, including the model of each GPU card, the maximum number of virtualized vGPUs supported for partitioning, graphics card memory, graphics card computing power, graphics card health status, and graphics card information reporting time. When the scheduler begins a new scheduling cycle every 3 seconds, it compares the device-plugin reporting time with the current time. If the difference is greater than 60 seconds, all GPU cards on the node are considered unavailable, i.e., no graphics card resources are available on the node. If the graphics card reporting time is normal, it is then determined whether each GPU card on the node is in the Ready or Healthy state. If it is in the NotReady or Unhealthy state, the GPU card is considered unavailable. All GPU card information is then traversed and stored in the scheduling cache.

[0117] Refer to Figure 3, which shows a flow chart of a method for obtaining remaining resource information provided in an embodiment of the present application.

[0118] In some optional implementations, the acquiring the remaining resource information based on the available total resource information and saving it to the scheduling buffer area includes:

[0119] Step S301: Obtain the GPU index and GPU usage corresponding to the application on each node;

[0120] Step S302: Obtaining used resource information based on the GPU index and GPU usage corresponding to the application on each node;

[0121] Step S303: Based on the available total resource information, remove the used resource information in the scheduling buffer area to obtain the remaining resource information and save it in the scheduling buffer area.

[0122] This allows for a detailed understanding of the GPU usage of applications on each node, including the GPU index and usage, which helps to more accurately understand the usage of GPU resources in the cluster. Based on the GPU index and GPU usage corresponding to the application on each node, information about the used resources can be obtained. This helps to more accurately calculate the used resources when scheduling resources and avoid duplicate resource allocation. When obtaining remaining resource information, the remaining resource information can be obtained and saved in the scheduling buffer by removing the used resource information from the scheduling buffer. This helps to better understand the remaining resources in the cluster, thereby achieving optimal resource allocation during task scheduling and improving cluster resource utilization. By refining resource monitoring and optimizing resource scheduling, it is possible to ensure that tasks can be executed efficiently in the cluster, improving task execution efficiency. In summary, the technical solution for obtaining the GPU index and GPU usage corresponding to the application on each node has beneficial effects such as refined management, accurate calculation of used resources, optimized resource scheduling, and improved task execution efficiency, helping to improve the resource management efficiency and task execution efficiency of the cluster.

[0123] In a specific embodiment, the applications for which GPU usage is counted include applications that have been allocated GPU resources on the node and are not in the Succeed or Failed state. According to the index and usage of the GPU card declared in the application, the corresponding used resources are subtracted from the scheduling cache to obtain the remaining GPU resources of the node.

[0124] In some optional implementations, obtaining the computing power data and total video memory requirement information of the GPU of each node based on hybrid quantization calculation includes:

[0125] Determine whether the application on each node has GPU requirements;

[0126] If the application on a node has GPU requirements, the GPU index corresponding to the application on the node is calculated and saved; if the application on a node does not have GPU requirements, no operation is performed on the application on the node;

[0127] Determine whether the GPU requirement is a whole-card pass-through mode; if the GPU requirement is a whole-card pass-through mode, obtain computing power data and total video memory requirement information of the GPU of the node through a multiplication operation based on the GPU index;

[0128] If the GPU requirement is not the whole card pass-through mode, the computing power data and total video memory requirement information of the GPU of the node are obtained by an addition operation based on the GPU index.

[0129] Therefore, a judgment is made as to whether the application on each node requires GPU resources. For applications with GPU requirements, the corresponding GPU index will be calculated and saved, while for applications without GPU requirements, no operation is performed. In this way, resource management can be carried out according to the personalized needs of the application to improve resource utilization efficiency. For applications with GPU requirements, it will be further determined whether they need the whole-card pass-through mode. If the requirement is the whole-card pass-through mode, the computing power data and total video memory requirement information of the node's GPU will be obtained based on the GPU index through multiplication operations; if the requirement is not the whole-card pass-through mode, the corresponding information will be obtained through addition operations. In this way, the demand for GPU resources of applications on each node can be accurately calculated, providing accurate data support for resource scheduling. The technical solution can adapt to the GPU resource requirements in different scenarios. Whether it is the whole-card pass-through mode or other modes, the required resource information can be obtained through corresponding operations, which has strong flexibility.

[0130] In a specific embodiment, hybrid quantitative calculation and simple preliminary screening, the hybrid quantitative calculation method specifically includes: judging whether the declaration in the application has GPU resources and whether it is in a non-Succeeded and Failed state, and no GPU resources are allocated to the application's Annotation; if the above conditions are met, the index of the container with GPU requirements in the application is counted and saved; if there is no GPU requirement, all subsequent calculations are skipped; according to the above GPU container index, each container in the application is traversed; the container requirement is the whole card pass-through mode, and the total GPU computing power requirement of the container is the total computing power of the whole card 100 (the total computing power of a card can be regarded as 100) multiplied by the number N of applied GPU cards, and the container is obtained. The total GPU demand required by the container is 100*N; if the container is in virtualization mode, the total GPU computing power demand of the container is equal to the computing power V applied for by the container (the size range of V is (0,100]) multiplied by the number of applied GPU cards N, and the total video memory demand of the container is the video memory applied for by the container multiplied by the number of applied GPU cards, the result is V*N. Save the total GPU computing power and total video memory demand of the container; save the requirements of each container in the Map structure, where the key value of the Map structure is the container index, and the value value is the serialized GPU computing power and video memory requirements required by the container. At the same time, add the requirements of all containers in the application (including GPU direct and virtualization requirements) to obtain the total GPU computing power and total video memory of the application;

[0131] In some optional implementations, the first screening of the nodes based on the total graphics memory requirement information of the GPUs of the nodes to obtain a first screening result includes:

[0132] Comparing the total graphics memory requirement information of the GPU of the node with the remaining resource information;

[0133] If the total memory requirement information of the GPU of the node is less than the remaining resource information, the node is considered to have insufficient resources and is filtered;

[0134] If the total graphics memory requirement information of the GPU of the node is not less than the remaining resource information, the node is saved to obtain the first screening result.

[0135] Therefore, if a node's total GPU memory requirement is less than its remaining resource information, it means that the node's resources are insufficient to meet the new task's requirements. In this case, the node can be filtered out to avoid scheduling tasks to nodes with insufficient resources, thereby improving the task's success rate. If a node's total GPU memory requirement is not less than its remaining resource information, it indicates that the node's resources can meet the new task's requirements. In this case, the node is saved and becomes the first filtered result. This ensures that nodes with sufficient resources are fully utilized, improving cluster resource utilization. By comparing a node's total GPU memory requirement with its remaining resource information, it is possible to quickly determine which nodes have sufficient resources to execute the new task. This helps improve task scheduling efficiency and reduce task failures due to insufficient resources. The system can adapt to GPU resource requirements in different scenarios. Whether it is memory requirements or other requirements, the comparison operation can be used to determine whether a node has sufficient resources, providing strong flexibility.

[0136] In a specific embodiment, the total GPU computing power and total video memory required by the application are compared with the total resources remaining in each node in the scheduling cache. If the remaining GPU resources of the node are smaller than the resources required by the application, the node can be directly considered to have insufficient resources and filtered out. If the remaining resources of the node are sufficient, the screening continues.

[0137] Referring to FIG. 4 , FIG. 4 shows a schematic diagram of a process for obtaining a simplified two-dimensional full permutation array provided in an embodiment of the present application.

[0138] In some optional implementations, obtaining a simplified two-dimensional full permutation array based on the full permutation combination function and the pruning function includes:

[0139] Step S401: input preset application information and obtain an index array with GPU;

[0140] Step S402: inputting the index array into a preset full permutation function to obtain a full permutation two-dimensional array;

[0141] Step S403: Determine whether the GPU computing power data and the GPU memory data are equal;

[0142] Step S404: if the computing power data of the GPU and the video memory data of the GPU are equal, marking the index array of the GPU to obtain an equivalent array;

[0143] Step S405: inputting the full permutation two-dimensional array and the equivalent array into a preset pruning function to obtain the reduced two-dimensional full permutation array.

[0144] Thus, by inputting preset application information, an index array with a GPU can be obtained according to the specific needs of the application, thereby realizing personalized resource management for different applications. By inputting the index array into the preset full permutation function, a fully permuted two-dimensional array can be obtained. Therefore, when scheduling tasks, all possible GPU resource allocation schemes can be fully considered, thereby improving resource utilization and task scheduling efficiency. After determining whether the GPU computing power data and the GPU memory data are equal, if they are equal, the GPU index array is marked to obtain an equivalent array. Then, the fully permuted two-dimensional array and the equivalent array are input into the preset pruning function together to obtain a streamlined two-dimensional fully permuted array. In this way, precise pruning can be performed to remove GPU resource allocation schemes that do not meet the conditions, thereby improving task scheduling efficiency and resource utilization. The technical solution of inputting preset application information and obtaining an index array with a GPU, and then performing full permutation and pruning operations has the beneficial effects of personalized resource management, full permutation optimization, precise pruning, and adaptability to different scenarios, which helps to improve the resource management efficiency and task execution efficiency of the cluster.

[0145] In a specific embodiment, full permutation combination and pruning, the full permutation combination and pruning method specifically includes: inputting preset application information to obtain an index array of the container where the GPU requirement is located; inputting the index array as an input parameter into the full permutation function to obtain a full permutation two-dimensional array; judging the computing power and video memory applied for by each container, if a container has the same computing power and video memory applied for as other containers, it can be regarded as an "equivalent container", and marking each "equivalent container" in the container index array to obtain an "equivalent" array, inputting the full permutation two-dimensional array and the "equivalent" array into the pruning function together, removing the repeated permutations of the full permutation two-dimensional array, and obtaining a simplified two-dimensional full permutation array after pruning.

[0146] In some optional implementations, performing a second screening operation on the nodes based on the first screening result to obtain a preferred node result includes:

[0147] Based on the first screening result, performing comparative screening on the reduced two-dimensional full permutation array to obtain a comparative screening result;

[0148] Based on the comparison and screening results, the nodes are scored and optimized to obtain the preferred node results.

[0149] Based on the first screening results, the simplified two-dimensional full permutation array is then compared and screened to obtain a comparison screening result. This allows for precise screening of nodes that meet the requirements, improving task scheduling efficiency and resource utilization. Based on the comparison screening results, the nodes are scored and optimized to obtain a preferred node result. This allows for the selection of the optimal execution node for the task based on the node's resource status and task requirements, improving task execution efficiency and cluster resource utilization. By accurately screening and optimizing node selection, it is possible to ensure that tasks can be executed efficiently and smoothly within the cluster, improving task execution effectiveness.

[0150] In a specific embodiment, the comparison and screening specifically includes: taking out the i-th container index array A in the pruned simplified two-dimensional full-permutation array, initializing and setting the GPU usage mark bit M to 0, the container binding GPU card structure BindMap (the key value is the container index, and the value value is the index of the node matching GPU card), and the number of GPUs required by the application is recorded as X, and then a copy of the GPU margin R of the node resource is copied, and the data in the container index array A is taken out, which is the container index I; taking the margin of the n-th GPU in the GPU margin resources of the node and comparing it with the computing power and memory resources declared by the container. If the GPU margin resources are insufficient, jump out of this card and go to the next card; if the comparison If the remaining capacity is sufficient, the computing power and video memory requested by the container will be subtracted from the copied node resource R, the mark bit value M will be incremented by 1, and the key-value pair of container index I and GPU index n will be recorded in BindMap. Once the mark bit M is equal to the value of the required GPU number X, the scheduling resource is considered to be matched successfully, and BindMap is returned to exit the loop; all GPU cards are traversed, and if M is equal to X, BindMap is returned directly; otherwise, BindMap is cleared and the i+1th container index array B is entered until the last set of container index arrays is traversed; since the result is returned in the loop when the scheduling resource match is successful, the final result is a match failure, that is, the empty BindMap is directly returned.

[0151] The scoring priority step specifically includes using the BindMap obtained above to read the stored data. If the BindMap data length is exactly 1, it proves that only one node meets the conditions, so this node is directly selected for binding with the application. If the BindMap data length is greater than or equal to 2, multiple nodes meet the conditions. Considering different characteristics of the nodes, including the node's used resources, remaining resources, and the number of bound applications, as scoring criteria, the node with the highest score is listed (if multiple nodes have the same highest score, one of them is randomly selected) as the preferred node. The application is bound to the prioritized node, and the scheduling process for the application ends.

[0152] Device Example

[0153] An embodiment of the present application also provides an electronic device, which includes a memory and a processor, wherein the memory stores a computer program, and when the processor executes the computer program, the steps of any of the above methods are implemented. The specific implementation method is consistent with the implementation method and the technical effect achieved in the above method embodiment, and some contents will not be repeated here.

[0154] An embodiment of the present application further provides an electronic device, the electronic device comprising a memory and a processor, the memory storing a computer program, and the processor being configured to implement the following steps when executing the computer program:

[0155] S1: Obtain resource information of the GPU of each node and save it to the scheduling buffer area. The resource information includes total available resource information and remaining resource information;

[0156] S2: Based on hybrid quantization calculation, obtain the computing power data and total video memory requirement information of each node's GPU;

[0157] S3: Based on the total video memory requirement information of the GPU of each node, perform a first screening on the nodes to obtain a first screening result;

[0158] S4: Based on the full permutation combination function and the pruning function, a simplified two-dimensional full permutation array is obtained;

[0159] S5: Based on the first screening result, a second screening operation is performed on the node to obtain a preferred node result.

[0160] In some optional implementations, the processor is configured to, when executing the computer program, obtain and save resource information of the GPU of each node in the following manner, where the resource information includes total available resource information and remaining resource information:

[0161] Use the device-plugin component to obtain the attribute information of the GPU corresponding to each node, including the GPU information reporting time and GPU status information;

[0162] When the periodic scheduler starts scheduling, determining whether the GPU information reporting time is greater than the preset reporting time; if the GPU information reporting time is greater than the preset reporting time, setting the GPU to an unavailable state;

[0163] If the GPU information reporting time is not greater than the preset reporting time, determining whether the GPU status information is in a ready state; if the GPU status information is not in a ready state, setting the GPU to an unavailable state;

[0164] If the GPU state information is in a ready state, traversing the attribute information of the GPU to obtain the total available resource information and saving it to a scheduling buffer;

[0165] Based on the total available resource information, the remaining resource information is obtained and saved in a scheduling buffer area.

[0166] In some optional implementations, the processor is configured to, when executing the computer program, obtain the remaining resource information based on the total available resource information and save it to the scheduling buffer in the following manner:

[0167] Get the GPU index and GPU usage corresponding to the application on each node;

[0168] Obtaining used resource information based on the GPU index and GPU usage corresponding to the application on each node;

[0169] Based on the total available resource information, the used resource information is removed from the scheduling buffer area to obtain the remaining resource information and save it in the scheduling buffer area.

[0170] In some optional embodiments, the processor is configured to obtain computing power data and total video memory requirement information of the GPU of each node based on hybrid quantization computing in the following manner when executing the computer program:

[0171] Determine whether the application on each node has GPU requirements;

[0172] If the application on a node has GPU requirements, the GPU index corresponding to the application on the node is calculated and saved; if the application on a node does not have GPU requirements, no operation is performed on the application on the node;

[0173] Determine whether the GPU requirement is a whole-card pass-through mode; if the GPU requirement is a whole-card pass-through mode, obtain computing power data and total video memory requirement information of the GPU of the node through a multiplication operation based on the GPU index;

[0174] If the GPU requirement is not the whole card pass-through mode, the computing power data and total video memory requirement information of the GPU of the node are obtained by an addition operation based on the GPU index.

[0175] In some optional embodiments, the processor is configured to, when executing the computer program, perform a first screening on the nodes based on the total graphics memory requirement information of the GPUs of the nodes to obtain a first screening result in the following manner:

[0176] Comparing the total graphics memory requirement information of the GPU of the node with the remaining resource information;

[0177] If the total memory requirement information of the GPU of the node is less than the remaining resource information, the node is considered to have insufficient resources and is filtered;

[0178] If the total graphics memory requirement information of the GPU of the node is not less than the remaining resource information, the node is saved to obtain the first screening result.

[0179] In some optional embodiments, the processor is configured to obtain a reduced two-dimensional full permutation array based on a full permutation combination function and a pruning function in the following manner when executing the computer program:

[0180] Input the preset application information and get the index array with GPU;

[0181] Inputting the index array into a preset full permutation function to obtain a full permutation two-dimensional array;

[0182] Determine whether the computing power data of the GPU and the video memory data of the GPU are equal;

[0183] If the computing power data of the GPU and the video memory data of the GPU are equal, marking the index array of the GPU to obtain an equivalent array;

[0184] The fully permuted two-dimensional array and the equivalent array are input into a preset pruning function to obtain the simplified two-dimensional fully permuted array.

[0185] In some optional embodiments, the processor is configured to perform a second screening operation on the node to obtain a preferred node result based on the first screening result in the following manner when executing the computer program:

[0186] Based on the first screening result, performing comparative screening on the reduced two-dimensional full permutation array to obtain a comparative screening result;

[0187] Based on the comparison and screening results, the nodes are scored and optimized to obtain the preferred node results.

[0188] Referring to FIG. 5 , FIG. 5 shows a structural framework diagram of an electronic device provided in an embodiment of the present application.

[0189] The electronic device includes at least one memory 210, at least one processor 220, and a bus 230 connecting different platform systems.

[0190] The memory 210 may include a readable medium in the form of a volatile memory, such as a random access memory (RAM) 211 and / or a cache memory 212 , and may further include a read-only memory (ROM) 213 .

[0191] The memory 210 also stores a computer program, which can be executed by the processor 220, so that the processor 220 implements the steps of any of the above methods.

[0192] The memory 210 may also include a utility 214 having at least one program module 215, such program module 215 including but not limited to: an operating system, one or more application programs, other program modules, and program data, each of which or some combination may include an implementation of a network environment.

[0193] Accordingly, the processor 220 may execute the aforementioned computer program and the utility 214 .

[0194] The processor 220 may be implemented as one or more application specific integrated circuits (ASICs), DSPs, programmable logic devices (PLDs), complex programmable logic devices (CPLDs), field programmable gate arrays (FPGAs), or other electronic components.

[0195] Bus 230 may represent one or more of several types of bus structures, including a memory bus or memory controller, a peripheral bus, an accelerated graphics port, a processor, or a local bus using any of a variety of bus architectures.

[0196] The electronic device may also communicate with one or more external devices 240, such as a keyboard, pointing device, Bluetooth device, etc., and may also communicate with one or more devices capable of interacting with the electronic device, and / or any device that enables the electronic device to communicate with one or more other computing devices (e.g., a router, modem, etc.). Such communication may be performed via input / output interface 250. Furthermore, the electronic device may also communicate with one or more networks (e.g., a local area network (LAN), a wide area network (WAN), and / or a public network, such as the Internet) via network adapter 260. The network adapter 260 may communicate with other modules of the electronic device via bus 230. It should be understood that, although not shown in the figures, other hardware and / or software modules may be used in conjunction with the electronic device, including but not limited to microcode, device drivers, redundant processors, external disk drive arrays, RAID systems, tape drives, and data backup storage platforms.

[0197] System Example

[0198] Referring to FIG. 6 , FIG. 6 shows a schematic structural diagram of a scheduling system provided in an embodiment of the present application.

[0199] An embodiment of the present application also provides a scheduling system, which includes the above-mentioned electronic device.

[0200] Media Examples

[0201] An embodiment of the present application also provides a chip, which stores a computer program. When the computer program is executed by a processor, the steps of any of the above methods are implemented. The specific implementation method is consistent with the implementation method and the technical effect achieved in the above method embodiment, and some contents will not be repeated here.

[0202] Referring to FIG. 7 , FIG. 7 shows a schematic structural diagram of a program product provided in an embodiment of the present application.

[0203] The program product is used to implement any of the above methods. The program product can use a portable compact disc read-only memory (CD-ROM) and include program code, and can be run on a terminal device, such as a personal computer. However, the program product of the present application is not limited to this. In the embodiment of the present application, the readable storage medium can be any tangible medium that contains or stores a program, and the program can be used by or in combination with an instruction execution system, device or device. The program product can use any combination of one or more readable media. The readable medium can be a readable signal medium or a readable storage medium. The readable storage medium can be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or device, or any combination of the above. More specific examples of readable storage media (non-exhaustive list) include: an electrical connection with one or more wires, a portable disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above.

[0204] The chip may include a data signal propagated in the baseband or as part of a carrier wave, which carries a readable program code. This propagated data signal may take a variety of forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. The readable storage medium may also be any readable medium that can send, propagate, or transmit a program for use by or in conjunction with an instruction execution system, device, or component. The program code contained on the readable storage medium may be transmitted using any suitable medium, including but not limited to wireless, wired, optical cable, RF, etc., or any suitable combination of the above. The program code for performing the operations of this application may be written in any combination of one or more programming languages, including object-oriented programming languages ​​such as Java, C++, etc., as well as conventional procedural programming languages ​​such as C, Python, or similar programming languages. The program code may be executed entirely on the user computing device, partially on the user device, as a standalone software package, partially on the user computing device and partially on a remote computing device, or entirely on a remote computing device or server. Where a remote computing device is involved, the remote computing device may be connected to the user computing device through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computing device (e.g., through the Internet using an Internet service provider).

[0205] This application is explained from the perspectives of purpose of use, effectiveness, progress and novelty, and has met the functional enhancement and use requirements emphasized by the Patent Law. The above description and drawings of this application are only preferred embodiments of this application and are not intended to limit this application. Therefore, all similar or identical structures, devices, features, etc. of this application, that is, all equivalent replacements or modifications made in accordance with the scope of this application's patent application, shall fall within the scope of protection of this application's patent application. Therefore, the scope of protection of this application's patent shall be based on the attached claims.

Claims

1. A GPU direct and resource hybrid scheduling method, comprising: S1: Obtain resource information of the GPU of each node and save it in the scheduling buffer area, wherein the resource information includes total available resource information and remaining resource information; S2: Based on hybrid quantization computing, obtain the computing power data and total video memory requirement information of the GPU of each node; S3: Based on the total video memory requirement information of the GPU of each node, perform a first screening on the nodes to obtain a first screening result; S4: Based on the full permutation combination function and the pruning function, a simplified two-dimensional full permutation array is obtained; S5: Based on the first screening result, a second screening operation is performed on the node to obtain a preferred node result.

2. The GPU direct and resource hybrid scheduling method according to claim 1, wherein the resource information of the GPU of each node is obtained and saved, and the resource information includes total available resource information and remaining resource information, including: Use the device-plugin component to obtain the attribute information of the GPU corresponding to each node, where the attribute information of the GPU includes the GPU information reporting time and GPU status information; When the periodic scheduler starts scheduling, determining whether the GPU information reporting time is greater than the preset reporting time; if the GPU information reporting time is greater than the preset reporting time, setting the GPU to an unavailable state; If the GPU information reporting time is not greater than the preset reporting time, determining whether the GPU status information is in a ready state; if the GPU status information is not in a ready state, setting the GPU to an unavailable state; If the GPU state information is in a ready state, traverse the attribute information of the GPU to obtain the total available resource information and save it in a scheduling buffer; Based on the total available resource information, the remaining resource information is obtained and saved in a scheduling buffer area.

3. The GPU direct and resource hybrid scheduling method according to claim 2, wherein the obtaining the remaining resource information based on the available total resource information and saving it to the scheduling buffer area comprises: Get the GPU index and GPU usage corresponding to the application on each node; Obtaining used resource information based on the GPU index and GPU usage corresponding to the application on each node; Based on the total available resource information, the used resource information is removed from the scheduling buffer area to obtain the remaining resource information and save it in the scheduling buffer area.

4. The GPU direct access and resource hybrid scheduling method according to claim 1, wherein the step of obtaining the computing power data and total video memory requirement information of the GPU of each node based on hybrid quantization calculation includes: Determine whether the application on each node has GPU requirements; If the application on a certain node has GPU requirements, the GPU index corresponding to the application on the node is calculated and saved; if the application on a certain node does not have GPU requirements, no operation is performed on the application on the node; Determine whether the GPU requirement is a whole-card pass-through mode, and if the GPU requirement is a whole-card pass-through mode, obtain computing power data and total video memory requirement information of the GPU of the node through a multiplication operation based on the GPU index; If the GPU requirement is not the whole card pass-through mode, the computing power data and total video memory requirement information of the GPU of the node are obtained through an addition operation based on the GPU index.

5. The GPU direct and resource hybrid scheduling method according to claim 1, wherein the first screening of the nodes based on the total video memory requirement information of the GPU of each node to obtain a first screening result comprises: Compare the total video memory requirement information of the GPU of the node with the remaining resource information; If the total video memory requirement information of the GPU of the node is less than the remaining resource information, it is considered that the node has insufficient resources and the node is filtered; If the total video memory requirement information of the GPU of the node is not less than the remaining resource information, the node is saved to obtain the first screening result.

6. The GPU direct and resource hybrid scheduling method according to claim 4, wherein the obtaining of a simplified two-dimensional full permutation array based on a full permutation combination function and a pruning function comprises: Input the preset application information and get the index array with GPU; Inputting the index array into a preset full permutation function to obtain a full permutation two-dimensional array; Determine whether the computing power data of the GPU and the video memory data of the GPU are equal; If the computing power data of the GPU and the video memory data of the GPU are equal, marking the index array of the GPU to obtain an equivalent array; The fully permuted two-dimensional array and the equivalent array are input into a preset pruning function together to obtain the simplified two-dimensional fully permuted array.

7. The GPU direct and resource hybrid scheduling method according to claim 6, wherein based on the first screening result, Performing a second screening operation on the node to obtain a preferred node result includes: Based on the first screening result, performing comparative screening on the simplified two-dimensional full permutation array to obtain a comparative screening result; Based on the comparison and screening results, the nodes are scored and selected respectively to obtain the preferred node results.

8. An electronic device, comprising a memory and a processor, wherein the memory stores a computer program, wherein the processor is configured to implement the following steps when executing the computer program: S1: Obtain resource information of the GPU of each node and save it in the scheduling buffer area, wherein the resource information includes total available resource information and remaining resource information; S2: Based on hybrid quantization computing, obtain the computing power data and total video memory requirement information of the GPU of each node; S3: Based on the total video memory requirement information of the GPU of each node, perform a first screening on the nodes to obtain a first screening result; S4: Based on the full permutation combination function and the pruning function, a simplified two-dimensional full permutation array is obtained; S5: Based on the first screening result, a second screening operation is performed on the node to obtain a preferred node result.

9. A scheduling system, comprising: The electronic device as claimed in claim 8.

10. A chip storing a computer program, wherein the computer program implements the steps of the method according to any one of claims 1 to 7 when executed by a processor.

Citation Information

Patent Citations

  • Container scheduling method and device oriented to shared GPU cluster

    CN114968566A

  • Server GPU computing power distribution system and method and server

    CN115904699A

  • Virtual GPU (Graphics Processing Unit) allocation method and system under container cloud environment based on API (Application Program Interface) interception and forwarding

    CN116991553A

  • GPU straight-through and resource hybrid scheduling method, system and chip

    CN117873706A

  • Topology-aware provisioning of hardware accelerator resources in a distributed environment

    US20190312772A1

Cited By

  • GPU video memory fragment optimization scheduling method and system

    CN121501512A