GPU (Graphics Processing Unit) scheduling system and method for intelligently evaluating resource occupancy

Through linear regression method and improved approximation ideal solution sorting method, efficient GPU resource scheduling of machine learning model tasks is achieved, the problem of insufficient resource utilization in the existing technology is solved, and the computing efficiency and resource utilization are improved.

CN119988031AInactive Publication Date: 2025-05-13KUAIJI XINYUN (QINGDAO) TECHNOLOGY CO LTD
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202510199094.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-24
Publication Date
2025-05-13
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

The existing technology lacks real-time adjustment mechanism in GPU resource scheduling of machine learning model tasks, resulting in insufficient resource utilization and affecting computing efficiency.

Method used

The GPU prediction usage of the node is calculated by linear regression method, and the optimal sequence is generated using the improved approximation ideal solution sorting method to achieve efficient scheduling of GPU resources.

Benefits of technology

The resource utilization and performance of machine learning model tasks are improved, the problem of unbalanced use of GPU resources is avoided, the response time of task scheduling is reduced, and the performance of GPU cluster is optimized.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119988031A_ABST
    Figure CN119988031A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of GPU scheduling, in particular to a GPU scheduling system and method for intelligent assessment of resource occupancy, and the method comprises the steps: obtaining GPU resource occupancy of a plurality of nodes in a cluster through a Kubernetes interface when a workflow of a machine learning model task is executed to a plurality of set durations; calculating the GPU prediction usage amount of a machine learning model task for a plurality of nodes in a cluster according to the historical GPU resource occupation amount through a usage amount prediction model based on a linear regression method; generating an optimal sequence by using the GPU predicted usage amounts of the plurality of nodes through an improved approximation ideal solution sorting method; and GPU optimization scheduling of the nodes is carried out through Kubernetes according to the optimal sequence. According to the method, efficient scheduling of GPU resources used by the machine learning model task is realized, the problem of unbalanced use of the GPU resources is avoided, and the resource utilization rate and performance of the machine learning model task are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of GPU scheduling technology, and in particular to a GPU scheduling system and method for intelligent evaluation of resource occupancy. Background Art

[0002] In the field of data science and machine learning, machine learning models, or logistic regression models, AI models usually require developers to use a large number of different types of work task flows to carry out various data exploration preparation and training evaluation work. The characteristics of this workflow are: it contains a variable number of arbitrary tasks, and each task has different requirements for GPU resources. Specifically, the number of nodes and usage rate of GPU used by each workflow will change frequently. The current mainstream technical implementation method is to schedule the GPU nodes used by the workflow based on the Kubernetes containerized GPU scheduling operating environment. Usually, the GPU nodes corresponding to the workflow are configured in a static manner, that is, when the workflow starts, the system executes the resource application. If the application is passed, the GPU resources will start to be exclusive, and during the task execution, the GPU resources cannot be released or added. This static configuration method shows great deficiencies in the machine learning model development task scenario where the number of nodes and the resource usage of a single node will change dynamically. Due to the lack of a real-time adjustment mechanism for resource allocation for workflow task nodes, it often leads to insufficient resource utilization, affecting the overall computing efficiency, or causing the workflow task to fail to obtain more idle resources, greatly reducing the training efficiency of the machine learning model.

[0003] Specifically, to connect a customized GPU device to a Kubernetes container cluster, it is necessary to develop and install a device plug-in according to its defined framework. The plug-in can separate the vendor-specific code, but it does not allow resource sharing or partial allocation on the customized GPU device. When the Kubernetes container cannot fully utilize the customized device, it will inevitably lead to reduced resource utilization, especially when GPUs are introduced as a new type of coprocessor in high-performance machine learning model calculations, and resource utilization is obviously insufficient.

[0004] Therefore, how to predict the GPU resource usage of machine learning model tasks to implement the design of a Kubernetes-based resource scheduling engine is a technical problem that needs to be solved. Summary of the invention

[0005] To this end, the present invention provides a GPU scheduling system and method for intelligent evaluation of resource occupancy, which calculates the predicted GPU usage of a node through a linear regression method, and calculates the optimal sequence of GPUs through an improved approximate ideal solution sorting method, thereby achieving efficient scheduling of GPU resources used by machine learning model tasks, avoiding the problem of uneven GPU resource usage, and improving the resource utilization and performance of machine learning model tasks.

[0006] To achieve the above object, the present invention proposes a GPU scheduling method for intelligent evaluation of resource usage, comprising:

[0007] When the workflow of the machine learning model task is executed for multiple set durations, the GPU resource usage of multiple nodes in the cluster is obtained through the Kubernetes interface;

[0008] The historical resource usage of the multiple GPUs is calculated by a usage prediction model based on a linear regression method to predict the GPU usage of the machine learning model task for multiple nodes in the cluster;

[0009] Generate an optimal sequence of the predicted GPU usage of multiple nodes based on an improved approach to an ideal solution sorting method;

[0010] Kubernetes is used to optimize the GPU scheduling of nodes according to the optimal sequence.

[0011] Furthermore, the process of calculating the predicted GPU usage of multiple nodes in the cluster by the machine learning model task based on the usage prediction model based on the linear regression method is as follows:

[0012] Construct a time series of resource usage based on the historical resource usage of multiple GPUs;

[0013] Constructing a differential equation through the resource quantity time series, and solving the parameters to be estimated of the differential equation using a solution formula;

[0014] Inputting the resource occupation amounts of the GPUs of the most recently set duration into the differential equation generates the predicted usage amounts of the GPUs.

[0015] Furthermore, the solution formula is a least squares formula.

[0016] In the above scheme, the high computational efficiency of the prediction process of the linear regression method is used to infer and predict the resource usage of subsequent nodes that have not yet started execution in the machine learning model task workflow, thereby achieving efficient and rapid prediction of GPU predicted usage.

[0017] Furthermore, the process of generating the optimal sequence of the GPU predicted usage of multiple nodes by using the improved approach to ideal solution sorting method is as follows:

[0018] Calculate the historical predicted availability of multiple GPU resources by using the availability formula based on the amount of resources occupied by the multiple GPUs;

[0019] Calculate the historical occupancy rates of multiple GPU resources by using a standardized formula based on the predicted availability rates of multiple GPU resources;

[0020] Calculate the maximum and minimum GPU resource utilization rates of all nodes in the cluster within the same set duration;

[0021] Calculate the maximum node distance and the minimum node distance by using a distance formula for the multiple GPU resource occupancy rates, the maximum GPU resource occupancy rates and the minimum GPU resource occupancy rates;

[0022] Calculate the node evaluation score according to the maximum distance of the node and the minimum distance of the node;

[0023] The optimal sequence is generated according to the node evaluation scores of a plurality of nodes.

[0024] Furthermore, the process of calculating the predicted availability of the historical multiple GPU resources by using the availability formula for the amount of resources occupied by the multiple GPUs is as follows:

[0025] Install the DevicePlugin plug-in in Kubernetes to query the number of GPU cards and GPU card memory;

[0026] Calculate the maximum value of GPU resources according to the number of GPU cards and the memory of the GPU cards;

[0027] The predicted availability of the GPU resources is obtained by subtracting the amount of resources occupied by the GPU from the maximum value of the GPU resources.

[0028] Furthermore, the process of calculating the maximum node distance and the minimum node distance by using the distance formula for multiple GPU resource occupancy rates, the maximum GPU resource occupancy rates and the minimum GPU resource occupancy rates is as follows:

[0029] The GPU resource occupancy rate and the maximum value of the GPU resource occupancy rate are used to calculate the maximum distance of the pre-node through the Euclidean formula, and the GPU resource occupancy rate and the minimum value of the GPU resource occupancy rate are used to calculate the minimum distance of the pre-node through the Euclidean formula;

[0030] The node maximum distance is generated by weighted summing up the multiple pre-node maximum distances of all set time lengths of the same node, and the node minimum distance is generated by weighted summing up the multiple pre-node minimum distances of all set time lengths of the same node.

[0031] Furthermore, the process of calculating the node evaluation score according to the maximum node distance and the minimum node distance is as follows:

[0032] The ratio of the minimum distance of the node to the sum of the maximum distance of the node and the minimum distance of the node is used as the node evaluation score.

[0033] In the above scheme, the GPU node scheduling method based on the improved approximate ideal solution sorting method reduces the response time of task scheduling, improves the throughput and availability of the cluster, and reduces the load imbalance of the cluster, thereby optimizing the performance of the GPU cluster.

[0034] Furthermore, the process of optimizing the GPU scheduling of nodes through Kubernetes according to the optimal sequence is as follows:

[0035] According to the GPU resource demand of the machine learning model task, multiple optimal nodes in the optimal sequence are determined, and GPU resources of the multiple optimal nodes are scheduled through the AKCSS scheduling algorithm.

[0036] The present invention also provides a system applying the GPU scheduling method for intelligent evaluation of resource usage, comprising:

[0037] The query module is used to obtain the GPU resource usage of multiple nodes in the cluster through the Kubernetes interface when the workflow of the machine learning model task is executed for multiple set durations;

[0038] A prediction module, used to calculate the predicted GPU usage of multiple nodes in the cluster by the machine learning model task through the usage prediction model based on the linear regression method based on the historical usage of multiple GPU resources;

[0039] A preferred module, used to generate an optimal sequence of the predicted GPU usage of multiple nodes by using an improved approach to an ideal solution sorting method;

[0040] A scheduling module, used to optimize the GPU scheduling of nodes through Kubernetes according to the optimal sequence;

[0041] Wherein, the query module and the scheduling module are connected to the Kubernetes API server.

[0042] Furthermore, the query module queries the amount of resources occupied by the GPU by installing a DevicePlugin plug-in in Kubernetes.

[0043] In the above solution, a GPU resource dynamic scheduling engine based on Kubernetes is built. It combines real-time monitoring and prediction data to flexibly predict the resource allocation of the workflow, and dynamically sends it to the platform to configure the runtime environment of the working node. It realizes the configuration of the working node without shutting down, the dynamic loading of the configuration instruction file data, and the dynamic hot plugging of GPU resources through the DevicePlugin plug-in mechanism of Kubernetes.

[0044] Compared with the prior art, the present invention has the following beneficial effects:

[0045] 1. The predicted GPU usage of the node is calculated by linear regression method, and the optimal sequence of GPU is calculated by improving the approximate ideal solution sorting method, which realizes efficient scheduling of GPU resources for machine learning model tasks, avoids the problem of uneven GPU resource usage, and improves the resource utilization and performance of machine learning model tasks.

[0046] 2. Through the high computational efficiency of the linear regression prediction process, the resource usage of subsequent nodes that have not yet started to execute in the machine learning model task workflow is inferred and predicted, achieving efficient and fast prediction of GPU usage.

[0047] 3. The GPU node scheduling method based on the improved approximate ideal solution sorting method reduces the response time of task scheduling, improves the throughput and availability of the cluster, and reduces the load imbalance of the cluster, thereby optimizing the performance of the GPU cluster.

[0048] 4. We have realized the construction of a dynamic GPU resource scheduling engine based on Kubernetes. It combines real-time monitoring and prediction data to flexibly predict the resource allocation of workflows, and dynamically sends it to the platform to configure the runtime environment of the working nodes. This allows the configuration of working nodes without shutting down, dynamic loading of configuration instruction file data, and dynamic hot-swapping of GPU resources through the DevicePlugin plug-in mechanism of Kubernetes. BRIEF DESCRIPTION OF THE DRAWINGS

[0049] Figure 1 A schematic diagram of a flow chart of a GPU scheduling method for intelligent resource occupancy evaluation according to an embodiment of the present invention;

[0050] Figure 2 A schematic diagram of the structure of a GPU scheduling system for intelligent resource occupancy evaluation according to an embodiment of the present invention;

[0051] Figure 3 A schematic diagram of a dynamic resource scheduling engine flow of a GPU scheduling system and method for intelligent resource occupancy evaluation according to an embodiment of the present invention;

[0052] Figure 4 The present invention is a schematic diagram of a Kubernetes node allocation process of a GPU scheduling system and method for intelligent resource usage assessment according to an embodiment of the present invention. DETAILED DESCRIPTION

[0053] In order to make the objects and advantages of the present invention more clearly understood, the present invention is further described below in conjunction with embodiments; it should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.

[0054] The preferred embodiments of the present invention are described below with reference to the accompanying drawings. It should be understood by those skilled in the art that these embodiments are only used to explain the technical principles of the present invention and are not intended to limit the protection scope of the present invention.

[0055] It should be noted that, in the description of the present invention, terms such as "up", "down", "left", "right", "inside" and "outside" indicating directions or positional relationships are based on the directions or positional relationships shown in the drawings. This is merely for the convenience of description and does not indicate or imply that the device or element must have a specific orientation, be constructed and operated in a specific orientation. Therefore, it cannot be understood as a limitation on the present invention.

[0056] In addition, it should be noted that in the description of the present invention, unless otherwise clearly specified and limited, the terms "installed", "connected", and "connected" should be understood in a broad sense, for example, it can be a fixed connection, a detachable connection, or an integral connection; it can be a mechanical connection or an electrical connection; it can be a direct connection, or it can be indirectly connected through an intermediate medium, or it can be the internal communication of two components. For those skilled in the art, the specific meanings of the above terms in the present invention can be understood according to specific circumstances.

[0057] like Figures 1 to 4 As shown, the present invention provides a GPU scheduling system and method for intelligent evaluation of resource occupancy, which calculates the predicted GPU usage of a node by a linear regression method, and calculates the optimal sequence of GPUs by an improved approximate ideal solution sorting method, thereby achieving efficient scheduling of GPU resources used by machine learning model tasks, avoiding the problem of uneven GPU resource usage, and improving the resource utilization and performance of machine learning model tasks.

[0058] like Figures 1 to 4 As shown, this embodiment proposes a GPU scheduling method for intelligent evaluation of resource usage, including:

[0059] When the workflow of the machine learning model task is executed for multiple set durations, the GPU resource usage of multiple nodes in the cluster is obtained through the Kubernetes interface;

[0060] The historical resource usage of the multiple GPUs is calculated by a usage prediction model based on a linear regression method to predict the GPU usage of the machine learning model task for multiple nodes in the cluster;

[0061] Generate an optimal sequence of the predicted GPU usage of multiple nodes based on an improved approach to an ideal solution sorting method;

[0062] Kubernetes is used to optimize the GPU scheduling of nodes according to the optimal sequence.

[0063] It should be noted that the interface is an API interface. The machine learning model task is represented as a POD task of Kubernetes (K8s) in actual implementation. The POD task is the smallest deployment unit in Kubernetes for encapsulating one or more containers, storage resources, and network configuration and management options. Kubernetes provides query, management, and scheduling functions for POD tasks, and can obtain the GPU resource usage of machine learning model tasks at multiple set durations. Therefore, the method described in this implementation can be encapsulated in a dynamic resource scheduling engine, and the scheduling management of Kubernetes can be realized through the dynamic resource scheduling engine.

[0064] The set duration is preferably set to 0.1 seconds. The optimal sequence is updated every 0.1 seconds of execution of the workflow of the machine learning model task, thereby combining real-time monitoring and prediction data to flexibly predict the resource allocation of the workflow, and dynamically send it to the platform to configure the operating environment of the working node, thereby realizing the dynamic loading of configuration instruction file data without shutting down the configuration working node.

[0065] It is understandable that the method described in this implementation needs to improve the response speed compared to the existing KCSS algorithm to achieve a real-time resource monitoring module, so the method needs to be as simple and efficient as possible while ensuring the prediction accuracy. Therefore, this implementation uses a relatively simple linear regression method to calculate the predicted GPU usage, which provides decision support for the resource scheduling of the entire workflow task by continuously monitoring resources with a higher logical dimension and oriented to the workflow task flow compared to the traditional monitoring granularity of only a single resource object.

[0066] Furthermore, the process of calculating the predicted GPU usage of multiple nodes in the cluster by the machine learning model task based on the usage prediction model based on the linear regression method is as follows:

[0067] The resource usage time series of multiple historical GPUs is constructed; a differential equation is constructed through the resource time series, and the estimated parameters of the differential equation are solved by a solution formula; and the resource usage of multiple GPUs with the most recent set time is input into the differential equation to generate multiple predicted GPU usages.

[0068] Furthermore, the solution formula is a least squares formula.

[0069] Specifically, the usage prediction model based on linear regression is essentially to process the original data and establish a differential equation that reflects the regular changes between data, so as to predict the demand for GPU resources of machine learning tasks. The specific implementation process is as follows:

[0070] The time series of resource usage is represented by X ij ={X i1 ,X i2 ,...,X im}, where X ij The amount of GPU resources occupied by the i-th node at the j-th set duration.

[0071] The resource usage time series are superimposed in sequence to generate a superimposed sequence, which is represented by X i ={X i1 ,X i2 ,...,X im},in,

[0072] The differential equation is constructed by the superposition sequence, and the differential equation is specifically:

[0073]

[0074] In the formula, p and q are two parameters to be estimated, X i represents the i-th superposition sequence.

[0075] The least square method is used to solve the estimated parameters of the differential equation, specifically:

[0076] (p,q) T =(B T B) -1 B T X i

[0077] In the formula, p and q are two parameters to be estimated, X i represents the i-th superposition sequence, B is the characteristic matrix, and the process of calculating the characteristic matrix B by the least squares formula is:

[0078]

[0079] In the formula, B is the feature matrix, X in Represents the nth element of the i-th superposition sequence.

[0080] In the above scheme, the high computational efficiency of the prediction process of the linear regression method is used to infer and predict the resource usage of subsequent nodes that have not yet started execution in the machine learning model task workflow, thereby achieving efficient and rapid prediction of GPU predicted usage.

[0081] Furthermore, if Figure 1 As shown in FIG. 1 , the process of generating the optimal sequence of the GPU predicted usage of multiple nodes by using the improved approximate ideal solution sorting method is as follows:

[0082] The GPU resource predicted availability is calculated by using the availability formula for the amount of resources occupied by the GPU; the GPU resource predicted availability is calculated by using the standardized formula for the GPU resource occupancy; the maximum GPU resource occupancy and the minimum GPU resource occupancy of all nodes in the cluster are calculated in the same set time period; the GPU resource occupancy, the maximum GPU resource occupancy and the minimum GPU resource occupancy are calculated by using the distance formula for the maximum node distance and the minimum node distance; the node evaluation score is calculated according to the maximum node distance and the minimum node distance; and the optimal sequence is generated according to the node evaluation scores of multiple nodes.

[0083] Furthermore, the process of calculating the predicted availability of GPU resources by using the availability formula for the amount of GPU resources occupied is as follows:

[0084] The number of GPU cards and the memory of GPU cards are queried by installing the DevicePlugin plug-in in Kubernetes; the maximum value of GPU resources is calculated according to the number of GPU cards and the memory of GPU cards; and the predicted availability of GPU resources is obtained by subtracting the amount of resources occupied by the GPU from the maximum value of GPU resources.

[0085] Specifically, the process of querying the number of GPU cards and GPU card memory by installing the DevicePlugin plug-in in Kubernetes is as follows: DevicePlugin's GPU Share DevicePlugin uses the nvml library to query the number of GPU cards and GPU card memory, and reports to Kubernetes' Kubelet through its ListAndWatch function, and Kubelet reports to the KubernetesAPI server. Therefore, the query of GPU resource status can be realized through the KubernetesAPI server.

[0086] Specifically, the formula for calculating the predicted availability of the GPU resources is:

[0087] A ij =UX ij

[0088] In the formula, A ij is the predicted availability of GPU resources for the i-th node at the j-th set duration, U is the maximum value of GPU resources, X ij The amount of GPU resources occupied by the i-th node at the j-th set duration.

[0089] Specifically, the GPU card memory is accumulated and superimposed according to the number of GPU cards to calculate the maximum value of the GPU resources. It is understandable that the maximum value of the GPU resources is an estimated value.

[0090] Specifically, the formula for calculating the GPU resource occupancy rate is:

[0091]

[0092] In the formula, Z ij is the GPU resource occupancy rate of the i-th node at the j-th set duration, X ij The amount of GPU resources occupied by the i-th node at the j-th set duration.

[0093] Specifically, the formula for calculating the maximum GPU resource occupancy rate and the minimum GPU resource occupancy rate is:

[0094] Z jmax =max{Z 1j ,Z 2j ,…,Z ij}

[0095] Z jmin =min{Z 1j ,Z 2j ,…,Z ij}

[0096] In the formula, Z jmax , Z jmin Respectively represent the maximum distance of the pre-node and the minimum distance of the pre-node, Z 1j ,Z 2j ,…,Z ij Indicates the GPU resource usage of the 1st node to the i-th node.

[0097] Furthermore, the process of calculating the maximum node distance and the minimum node distance by using the distance formula for the GPU resource occupancy rate, the maximum GPU resource occupancy rate and the minimum GPU resource occupancy rate is as follows:

[0098] The GPU resource occupancy rate and the maximum value of the GPU resource occupancy rate are used to calculate the maximum distance of the pre-node through the Euclidean formula, and the GPU resource occupancy rate and the minimum value of the GPU resource occupancy rate are used to calculate the minimum distance of the pre-node through the Euclidean formula;

[0099] The node maximum distance is generated by weighted summing up the multiple pre-node maximum distances of all set time lengths of the same node, and the node minimum distance is generated by weighted summing up the multiple pre-node minimum distances of all set time lengths of the same node.

[0100] Specifically, the formula for calculating the maximum distance between nodes and the minimum distance between nodes is:

[0101]

[0102] Where, d imax , d imin is the maximum distance between nodes and the minimum distance between nodes, W j represents the weight of node i at the jth set time. Preferably, in order to reduce the impact of GPU resource occupancy fluctuation on distance calculation, the weight corresponding to the node obtained at a set time farther from the current moment is set larger. k is preferably 5, W1…W5 are 0.15, 0.15, 0.2, 0.2, 0.3 respectively, and Z jmax , Z jmin Respectively represent the maximum distance of the pre-node and the minimum distance of the pre-node, Z ij represents the GPU resource usage of the i-th node at the j-th set duration, Z jmax -Z ij , Z jmin -Z ij are the maximum distance of the prenode and the minimum distance of the prenode respectively.

[0103] Furthermore, the process of calculating the node evaluation score according to the maximum node distance and the minimum node distance is as follows:

[0104] The ratio of the minimum distance of the node to the sum of the maximum distance of the node and the minimum distance of the node is used as the node evaluation score.

[0105] Specifically, the formula for calculating the node evaluation score is:

[0106]

[0107] In the formula, S i The node evaluation score for the i-th node, d imax , d imin It is understood that the node evaluation score can accurately and quickly reflect the node resource availability and recommendation degree.

[0108] In the above scheme, the GPU node scheduling method based on the improved approximate ideal solution sorting method reduces the response time of task scheduling, improves the throughput and availability of the cluster, and reduces the load imbalance of the cluster, thereby optimizing the performance of the GPU cluster.

[0109] Furthermore, if Figures 2 to 4 As shown in the figure, the process of optimizing the GPU scheduling of nodes through Kubernetes according to the optimal sequence is:

[0110] Determine multiple optimal nodes in the optimal sequence according to the GPU resource demand of the machine learning model task, and schedule the GPU resources of the multiple optimal nodes through the AKCSS scheduling algorithm, that is, schedule the optimal nodes through Figure 3 Add the asynchronous binding shown in the figure to the Kubernetes API server, so that the Kubernetes API server can use Figure 4 The method shown is assigned to the corresponding machine learning model task (POD).

[0111] like Figure 2 As shown, this embodiment also provides a GPU scheduling system for intelligent evaluation of resource usage, including:

[0112] The query module is used to obtain the GPU resource usage of multiple nodes in the cluster through the Kubernetes interface when the workflow of the machine learning model task is executed for multiple set durations;

[0113] A prediction module, used to calculate the predicted GPU usage of multiple nodes in the cluster by the machine learning model task through the usage prediction model based on the linear regression method based on the historical usage of multiple GPU resources;

[0114] A preferred module, used to generate an optimal sequence of the predicted GPU usage of multiple nodes by using an improved approach to an ideal solution sorting method;

[0115] A scheduling module, used to optimize the GPU scheduling of nodes through Kubernetes according to the optimal sequence;

[0116] Wherein, the query module and the scheduling module are both connected to the Kubernetes API server.

[0117] It is understandable that in Figures 2 to 4In the structure shown, the query module, prediction module, optimization module and scheduling module are encapsulated as a dynamic resource scheduling engine. The dynamic resource scheduling engine can dynamically determine how many resources the workflow is expected to occupy based on monitoring and prediction data, create instructions for allocating resources, store them as yaml objects, and send them to the specified workflow runtime environment at the same time. Its specific implementation method is: develop a resource scheduling engine, take the global usage of the system GPU and the resource configuration application metadata of the current workflow as input, generate resource configuration update instructions for the task nodes in the workflow that have not yet started to execute the subsequent tasks in real time, store them in yaml form, and dynamically update the configuration instructions to the containerized program unit corresponding to the current workflow metadata configuration loading in the Kubernetes cluster (usually the background working node of the task scheduling cloud management platform), so as to realize the real-time update of the GPU resource configuration application to the workflow runtime environment.

[0118] Furthermore, the query module queries the amount of resources occupied by the GPU by installing the DevicePlugin plug-in in Kubernetes, so as to realize dynamic hot plugging of GPU resources.

[0119] It is understandable that the dynamic hot-plugging of the GPU resources adopts the GPU dynamic hot-plugging technology oriented to the logical workflow, and realizes the dynamic update of the memory-level resource configuration metadata based on the way that the container configuration instructions are dynamically sent to the management node.

[0120] The method and system described in this implementation dynamically updates the configuration instructions to the containerized program unit corresponding to the current workflow metadata configuration loading in the Kubernetes cluster (usually the background work node of the task scheduling cloud management platform) through the resource scheduling engine, so as to realize the real-time update of the GPU resource configuration application to the workflow runtime environment. Before the subsequent task nodes that have not yet started to execute in the workflow are started, the containerized orchestration resource configuration instructions corresponding to the nodes have been updated. When the container corresponding to the new task node is started, the container runtime resource application configuration interface of Kubernetes can be used to initiate a new resource configuration application to Kubernetes, and a containerized resource object that complies with the new resource configuration requirements can be created. Finally, when the container is started, the subsequent task nodes that have not yet started to execute in the workflow can obtain new GPU resources, either more than the original setting in the workflow orchestration stage to realize the automatic expansion of GPU resources for workflow tasks, or less than the original resource request to realize the automatic recovery of GPU resources.

[0121] In the above solution, a GPU resource dynamic scheduling engine based on Kubernetes is built. It combines real-time monitoring and prediction data to flexibly predict the resource allocation of the workflow, and dynamically sends it to the platform to configure the runtime environment of the working node. It realizes the configuration of the working node without shutting down, the dynamic loading of the configuration instruction file data, and the dynamic hot plugging of GPU resources through the DevicePlugin plug-in mechanism of Kubernetes.

[0122] In this embodiment, the GPU predicted usage of the node is calculated by the linear regression method, and the optimal sequence of the GPU is calculated by the improved approach to the ideal solution sorting method, so that the efficient scheduling of the use of GPU resources by the machine learning model task is realized, the problem of uneven use of GPU resources is avoided, and the resource utilization and performance of the machine learning model task are improved. Through the high calculation efficiency of the prediction process of the linear regression method, the resource usage of the subsequent nodes that have not yet started to be executed in the machine learning model task workflow is inferred and predicted, and the efficient and rapid prediction of the predicted usage of the GPU is realized. The GPU node scheduling method based on the improved approach to the ideal solution sorting method (TOPSIS) reduces the response time of task scheduling by optimizing the optimal sequence of nodes used by machine learning tasks, improves the throughput and availability of the cluster, and reduces the load imbalance of the cluster, thereby optimizing the performance of the GPU cluster. It realizes the construction of a GPU resource dynamic scheduling engine based on Kubernetes, which combines real-time monitoring and prediction data, flexibly predicts the resource allocation of the workflow, and dynamically sends it to the runtime environment of the platform configuration work node, realizes the configuration of the work node without stopping, the dynamic loading of the configuration instruction file data, and realizes the dynamic hot plug of GPU resources through the DevicePlugin plug-in mechanism of Kubernetes.

[0123] So far, the technical solutions of the present invention have been described in conjunction with the preferred embodiments shown in the accompanying drawings. However, it is easy for those skilled in the art to understand that the protection scope of the present invention is obviously not limited to these specific embodiments. Without departing from the principle of the present invention, those skilled in the art can make equivalent changes or substitutions to the relevant technical features, and the technical solutions after these changes or substitutions will fall within the protection scope of the present invention.

[0124] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. For those skilled in the art, the present invention may have various modifications and variations. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.

Claims

1. A GPU scheduling method for intelligent resource occupancy evaluation, characterized in that: include: When the workflow of the machine learning model task is executed for multiple set durations, the GPU resource usage of multiple nodes in the cluster is obtained through the Kubernetes interface; The historical resource usage of the multiple GPUs is calculated by a usage prediction model based on a linear regression method to predict the GPU usage of the machine learning model task for multiple nodes in the cluster; Generate an optimal sequence of the predicted GPU usage of multiple nodes based on an improved approach to an ideal solution sorting method; Kubernetes is used to optimize the GPU scheduling of nodes according to the optimal sequence.

2. The GPU scheduling method for intelligent resource occupancy evaluation according to claim 1, characterized in that: The process of calculating the predicted GPU usage of multiple nodes in the cluster by the machine learning model task through the usage prediction model based on the linear regression method is as follows: Construct a time series of resource usage based on the historical resource usage of multiple GPUs; Constructing a differential equation through the resource quantity time series, and solving the parameters to be estimated of the differential equation using a solution formula; Inputting the resource occupation amounts of the GPUs of the most recently set duration into the differential equation generates the predicted usage amounts of the GPUs.

3. The GPU scheduling method for intelligent resource occupancy evaluation according to claim 2, characterized in that: The solution formula is a least squares formula.

4. The GPU scheduling method for intelligent resource occupancy evaluation according to claim 1, characterized in that: The process of generating the optimal sequence of GPU predicted usage of multiple nodes by using the improved approach to ideal solution sorting method is as follows: Calculate the predicted availability of multiple GPU resources by using the availability formula based on the amount of resources occupied by the multiple GPUs; Calculate the GPU resource occupancy rate by using a standardized formula based on the predicted availability rates of the multiple GPU resources; Calculate the maximum and minimum GPU resource utilization rates of all nodes in the cluster within the same set duration; Calculate the maximum node distance and the minimum node distance by using a distance formula for the multiple GPU resource occupancy rates, the maximum GPU resource occupancy rates and the minimum GPU resource occupancy rates; Calculate the node evaluation score according to the maximum distance of the node and the minimum distance of the node; The optimal sequence is generated according to the node evaluation scores of a plurality of nodes.

5. The GPU scheduling method for intelligent resource occupancy evaluation according to claim 4, characterized in that: The process of calculating the predicted availability of GPU resources by using the availability formula for the amount of GPU resources occupied is as follows: Install the DevicePlugin plug-in in Kubernetes to query the number of GPU cards and GPU card memory; Calculate the maximum value of GPU resources according to the number of GPU cards and the memory of the GPU cards; The predicted availability of the GPU resources is obtained by subtracting the amount of resources occupied by the GPU from the maximum value of the GPU resources.

6. The GPU scheduling method for intelligent resource occupancy evaluation according to claim 4, characterized in that: The process of calculating the maximum node distance and the minimum node distance by using the distance formula to calculate the GPU resource occupancy rate, the maximum GPU resource occupancy rate and the minimum GPU resource occupancy rate is as follows: The GPU resource occupancy rate and the maximum value of the GPU resource occupancy rate are used to calculate the maximum distance of the pre-node through the Euclidean formula, and the GPU resource occupancy rate and the minimum value of the GPU resource occupancy rate are used to calculate the minimum distance of the pre-node through the Euclidean formula; The node maximum distance is generated by weighted summing up the multiple pre-node maximum distances of all set time lengths of the same node, and the node minimum distance is generated by weighted summing up the multiple pre-node minimum distances of all set time lengths of the same node.

7. The GPU scheduling method for intelligent resource occupancy evaluation according to claim 4, characterized in that: The process of calculating the node evaluation score based on the maximum node distance and the minimum node distance is: The ratio of the minimum distance of the node to the sum of the maximum distance of the node and the minimum distance of the node is used as the node evaluation score.

8. The GPU scheduling method for intelligent resource occupancy evaluation according to any one of claims 1 to 7, characterized in that: The process of optimizing GPU scheduling of nodes through Kubernetes according to the optimal sequence is as follows: According to the GPU resource demand of the machine learning model task, multiple optimal nodes in the optimal sequence are determined, and GPU resources of the multiple optimal nodes are scheduled through the AKCSS scheduling algorithm.

9. A GPU scheduling system for intelligent resource occupancy evaluation using the GPU scheduling method for intelligent resource occupancy evaluation according to any one of claims 1 to 8, characterized in that: include: The query module is used to obtain the GPU resource usage of multiple nodes in the cluster through the Kubernetes interface when the workflow of the machine learning model task is executed for multiple set durations; A prediction module, used to calculate the predicted GPU usage of multiple nodes in the cluster by the machine learning model task through a usage prediction model based on a linear regression method based on the historical usage of multiple GPU resources; A preferred module, used to generate an optimal sequence of the predicted GPU usage of multiple nodes by using an improved approach to an ideal solution sorting method; A scheduling module, used to optimize the GPU scheduling of nodes through Kubernetes according to the optimal sequence; Wherein, the query module and the scheduling module are both connected to the Kubernetes API server.

10. The GPU scheduling system for intelligent resource occupancy evaluation according to claim 9, characterized in that: The query module queries the amount of resources occupied by the GPU by installing the DevicePlugin plug-in in Kubernetes.

Citation Information

Patent Citations

  • Multi-factor strategy-based computing power resource optimal scheduling distribution method

    CN115550370A

  • Resource scheduling method and system for training tasks of deep recommendation system

    CN117492997A

  • Multi-tenant GPU cluster elastic quota scheduling method and system

    CN117707759A

  • Multidimensional resource scheduling method in kubernetes cluster architecture system

    US20210365290A1