GPU sharing scheduling based on interference detection
By starting the execution of workloads on multiple GPUs, extracting useful feature sets of utilization metrics using deep learning models, determining the workload type, and configuring shared execution of workloads based on interference detection and interference avoidance policies, the problem of low GPU resource utilization is solved, and more efficient resource sharing and task completion is achieved.
Patent Information
- Application Number
- CN202280101569.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2022-09-07
- Publication Date
- 2025-07-01
AI Technical Summary
In the prior art, GPU resource utilization is low and sharing capacity is limited, resulting in waste of resources and inefficiency, especially when a single GPU utilization is low in deep learning tasks.
By starting workload execution on multiple GPUs, extracting a useful feature set of utilization metrics using deep learning models, determining workload types, and configuring shared execution of workloads based on interference detection and interference avoidance policies, utilizing virtual GPU resources for optimized scheduling.
It improves the utilization rate of GPU resources, reduces interference between workloads, optimizes task completion time, and improves the overall efficiency of computing resources.
Smart Images

Figure CN120239853A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to computing resource sharing in a cloud-native environment, for example, AI-based scheduling for sharing graphics processing unit (GPU) resources. Background Art
[0002] Due to the integration of thousands of computing cores on a single chip, GPUs have powerful parallel processing capabilities. Therefore, GPUs can provide large-scale computing power support for deep learning (DL) tasks such as computer vision (CV), natural language processing (NLP), and high-performance computing (HPC). With the rapid development of the DL field in the past few years, different technologies for accessing and configuring GPU resources have emerged continuously. However, there are problems of low GPU resource utilization and limited GPU resource sharing ability in the prior art. Summary of the Invention
[0003] Various examples are now described to introduce some concepts in a simplified form, which will be further described in the detailed implementation. The summary of the invention is not intended to identify the key or essential features of the claimed subject matter, nor is it intended to limit the scope of the claimed subject matter.
[0004] According to a first aspect of the present invention, a computer-implemented method for AI-based workload scheduling is provided. The method includes: starting the execution of a first workload on one of a plurality of graphics processing units (GPUs); determining a utilization metric of the first workload, where the utilization metric is associated with the execution of the first workload on the GPU; using a transformation function of a deep learning (DL) model to extract a useful feature set of the utilization metric of the first workload, where the useful feature set includes a subset of the utilization metric; using the useful feature set to determine the workload type of the first workload; and configuring shared execution of the first workload and a second workload on a second GPU among the plurality of GPUs based on packing the first workload and the second workload, where the second workload is associated with the workload type of the first workload.
[0005] According to the first aspect, in the first implementation of the method, the DL model includes an AI-based encoder and an AI-based decoder. The method further includes: performing training of the DL model by using a first training dataset as the input of the AI-based encoder and using a second training dataset as the output of the AI-based decoder.
[0006] According to the first aspect or any of the foregoing implementations of the first aspect, in the second implementation of the method, the first training dataset is configured to include previous utilization metrics of a plurality of workloads executed before the execution of the first workload. The plurality of workloads includes the second workload.
[0007] According to the first aspect or any of the foregoing implementations of the first aspect, in the third implementation of the method, the second training dataset is configured as a plurality of joint completion times, where the plurality of joint completion times is associated with a corresponding plurality of joint executions, and the plurality of joint executions is associated with the plurality of workloads.
[0008] According to the first aspect or any of the foregoing implementations of the first aspect, in the fourth implementation of the method, one of the plurality of joint executions includes at least two workloads among the plurality of workloads executed on the same GPU among the plurality of GPUs.
[0009] According to the first aspect or any of the foregoing implementations of the first aspect, in the fifth implementation of the method, the conversion function is determined by using a subset of convolutional layers among a plurality of convolutional layers on the AI-based encoder of the DL model.
[0010] According to the first aspect or any of the foregoing implementations of the first aspect, in the sixth implementation of the method, the conversion function is applied to the utilization metrics of the plurality of workloads to obtain an additional useful feature set. The plurality of workloads is executed before the execution of the first workload, and the plurality of workloads includes the second workload.
[0011] According to the first aspect or any of the foregoing implementations of the first aspect, in the seventh implementation of the method, the workload type of the first workload is determined by using a comparison between the useful feature set and each additional useful feature set in the additional useful feature set. The second workload is selected based on the comparison.
[0012] According to the first aspect or any of the foregoing implementations of the first aspect, in the eighth implementation of the method, the selection of the second workload includes: when the difference between the useful feature set and an additional useful feature set in the additional useful feature sets does not exceed a threshold, select the second workload. The additional useful feature set is associated with the second workload.
[0013] According to the first aspect or any of the foregoing implementations of the first aspect, in the ninth implementation of the method, when the difference between the useful feature set and the additional useful feature set is not greater than the threshold, perform the configuration of the shared execution of the first workload and the second workload.
[0014] According to the first aspect or any of the foregoing implementations of the first aspect, in the tenth implementation of the method, configure a plurality of virtual GPUs (vGPUs) of the second GPU. Use the plurality of vGPUs of the second GPU to configure the shared execution of the first workload and the second workload.
[0015] According to the first aspect or any of the foregoing implementations of the first aspect, in the eleventh implementation of the method, the utilization rate metric includes at least one of the following: a GPU usage histogram of one or more containers associated with the execution of the first workload; a memory usage histogram of a computing node associated with the execution of the first workload; a GPU type associated with the GPU used for the execution of the first workload.
[0016] According to a second aspect of the present invention, there is provided a system for workload scheduling based on artificial intelligence (AI). The system includes a memory storing instructions and at least one processor communicating with the memory. The at least one processor is configured to perform operations including the following when executing the instructions: start the execution of a first workload on one GPU among a plurality of graphics processing units (GPUs); determine a utilization rate metric of the first workload, where the utilization rate metric is associated with the execution of the first workload on the GPU; extract a useful feature set of the utilization rate metric of the first workload using a conversion function of a deep learning (DL) model, where the useful feature set includes a subset of the utilization rate metric; determine a workload type of the first workload using the useful feature set; and configure the shared execution of the first workload and a second workload on a second GPU among the plurality of GPUs based on packing the first workload and the second workload, where the second workload is associated with the workload type of the first workload.
[0017] According to a second aspect, in a first implementation of the system, the DL model includes an AI-based encoder and an AI-based decoder. The method further includes: performing training of the DL model by using a first training dataset as an input to the AI-based encoder and using a second training dataset as an output of the AI-based decoder.
[0018] According to the second aspect or any of the foregoing implementations of the second aspect, in a second implementation of the system, the first training dataset is configured to include previous utilization metrics of a plurality of workloads executed prior to the execution of the first workload. The plurality of workloads includes the second workload.
[0019] According to the second aspect or any of the foregoing implementations of the second aspect, in a third implementation of the system, the second training dataset is configured as a plurality of joint completion times, where the plurality of joint completion times are associated with a corresponding plurality of joint executions, and the plurality of joint executions are associated with the plurality of workloads.
[0020] According to the second aspect or any of the foregoing implementations of the second aspect, in a fourth implementation of the system, one of the plurality of joint executions includes at least two of the plurality of workloads executed on the same GPU among the plurality of GPUs.
[0021] According to the second aspect or any of the foregoing implementations of the second aspect, in a fifth implementation of the system, the transformation function is determined using a subset of convolutional layers among a plurality of convolutional layers on the AI-based encoder of the DL model.
[0022] According to the second aspect or any of the foregoing implementations of the second aspect, in a sixth implementation of the system, the transformation function is applied to utilization metrics of a plurality of workloads to obtain an additional useful feature set. The plurality of workloads are executed prior to the execution of the first workload, and the plurality of workloads includes the second workload.
[0023] According to the second aspect or any of the foregoing implementations of the second aspect, in a seventh implementation of the system, the workload type of the first workload is determined by using a comparison of the useful feature set with each additional useful feature set in the additional useful feature set. The second workload is selected based on the comparison.
[0024] According to the second aspect or any of the foregoing implementations of the second aspect, in the eighth implementation of the system, the selection of the second workload includes: when the difference between the useful feature set and an additional useful feature set in the additional useful feature sets does not exceed a threshold, select the second workload. The additional useful feature set is associated with the second workload.
[0025] According to the second aspect or any of the foregoing implementations of the second aspect, in the ninth implementation of the system, when the difference between the useful feature set and the additional useful feature set is not greater than the threshold, perform the configuration of the shared execution of the first workload and the second workload.
[0026] According to the second aspect or any of the foregoing implementations of the second aspect, in the tenth implementation of the system, configure a plurality of virtual GPUs (vGPUs) of the second GPU. Use the plurality of vGPUs of the second GPU to configure the shared execution of the first workload and the second workload.
[0027] According to the second aspect or any of the foregoing implementations of the second aspect, in the eleventh implementation of the system, the utilization rate metrics include at least one of the following: a GPU usage histogram of one or more containers associated with the execution of the first workload; a memory usage histogram of a computing node associated with the execution of the first workload; a GPU type associated with the GPU used for the execution of the first workload.
[0028] According to a third aspect of the present invention, there is provided a non-transitory computer-readable medium storing instructions for workload scheduling based on artificial intelligence (AI). When the instructions are executed by one or more processors, the one or more processors are caused to perform operations. The operations include: starting the execution of a first workload on one GPU among a plurality of graphics processing units (GPUs); determining a utilization rate metric of the first workload, where the utilization rate metric is associated with the execution of the first workload on the GPU; using a conversion function of a deep learning (DL) model to extract a useful feature set of the utilization rate metric of the first workload, where the useful feature set includes a subset of the utilization rate metric; using the useful feature set to determine a workload type of the first workload; and based on packing the first workload and a second workload, configuring the shared execution of the first workload and the second workload on a second GPU among the plurality of GPUs, where the second workload is associated with the workload type of the first workload.
[0029] According to a third aspect, in a first implementation of the non-transitory computer-readable medium, the DL model includes an AI-based encoder and an AI-based decoder. The method further includes: performing training of the DL model by using a first training dataset as an input to the AI-based encoder and using a second training dataset as an output of the AI-based decoder.
[0030] According to the third aspect or any of the foregoing implementations of the third aspect, in a second implementation of the non-transitory computer-readable medium, the first training dataset is configured to include previous utilization metrics of a plurality of workloads executed before the execution of the first workload. The plurality of workloads includes the second workload.
[0031] According to the third aspect or any of the foregoing implementations of the third aspect, in a third implementation of the non-transitory computer-readable medium, the second training dataset is configured as a plurality of joint completion times, where the plurality of joint completion times are associated with a corresponding plurality of joint executions, and the plurality of joint executions are associated with the plurality of workloads.
[0032] According to the third aspect or any of the foregoing implementations of the third aspect, in a fourth implementation of the non-transitory computer-readable medium, one of the plurality of joint executions includes at least two workloads among the plurality of workloads executed on the same GPU among the plurality of GPUs.
[0033] According to the third aspect or any of the foregoing implementations of the third aspect, in a fifth implementation of the non-transitory computer-readable medium, a subset of convolutional layers among a plurality of convolutional layers on the AI-based encoder of the DL model is used to determine the transformation function.
[0034] According to the third aspect or any of the foregoing implementations of the third aspect, in a sixth implementation of the non-transitory computer-readable medium, the transformation function is applied to utilization metrics of a plurality of workloads to obtain an additional useful feature set. The plurality of workloads are executed before the execution of the first workload, and the plurality of workloads includes the second workload.
[0035] According to the third aspect or any of the foregoing implementations of the third aspect, in a seventh implementation of the non-transitory computer-readable medium, a comparison between the useful feature set and each additional useful feature set in the additional useful feature set is used to determine the workload type of the first workload. The second workload is selected based on the comparison.
[0036] According to the third aspect or any of the foregoing implementations of the third aspect, in the eighth implementation of the non-transitory computer-readable medium, the selection of the second workload includes: when the difference between the useful feature set and an additional useful feature set in the additional useful feature sets does not exceed a threshold, selecting the second workload. The additional useful feature set is associated with the second workload.
[0037] According to the third aspect or any of the foregoing implementations of the third aspect, in the ninth implementation of the non-transitory computer-readable medium, when the difference between the useful feature set and the additional useful feature set is not greater than the threshold, the configuration of the shared execution of the first workload and the second workload is performed.
[0038] According to the third aspect or any of the foregoing implementations of the third aspect, in the tenth implementation of the non-transitory computer-readable medium, a plurality of virtual GPUs (vGPUs) of the second GPU are configured. The shared execution of the first workload and the second workload is configured using the plurality of vGPUs of the second GPU.
[0039] According to the third aspect or any of the foregoing implementations of the third aspect, in the eleventh implementation of the non-transitory computer-readable medium, the utilization metric includes at least one of the following: a GPU usage histogram of one or more containers associated with the execution of the first workload; a memory usage histogram of a computing node associated with the execution of the first workload; a GPU type associated with the GPU used for the execution of the first workload.
[0040] According to a fourth aspect of the present invention, a system for workload scheduling based on artificial intelligence (AI) is provided. The system includes: a module for starting the execution of a first workload on one of a plurality of graphics processing units (GPUs); a module for determining a utilization metric of the first workload, where the utilization metric is associated with the execution of the first workload on the GPU; a module for extracting a useful feature set of the utilization metric of the first workload using a transformation function of a deep learning (DL) model, where the useful feature set includes a subset of the utilization metric; a module for determining a workload type of the first workload using the useful feature set; and a module for configuring the shared execution of the first workload and a second workload on a second GPU among the plurality of GPUs based on packing the first workload and the second workload, where the second workload is associated with the workload type of the first workload.
[0041] Any one of the above examples can be combined with any one or more of the other above examples to create new embodiments within the scope of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS
[0042] In the drawings, which are not necessarily drawn to scale, like numerals may describe like components in different views. The drawings generally illustrate, by way of example and not limitation, the various embodiments described herein.
[0043] Figure 1 is a high-level system diagram of a network architecture according to some exemplary embodiments, the network architecture having a workload management module that performs workload management functions.
[0044] Figure 2 is a block diagram of a computing node according to some exemplary embodiments, the computing node being used to implement Figure 1 the workload management module in
[0045] Figure 3 is a diagram of workload (or job) packing and the corresponding combined completion time for completing the workload according to some exemplary embodiments.
[0046] Figure 4 is according to some exemplary embodiments of Figure 1 the exemplary utilization metrics collected and used by the workload management module in
[0047] Figure 5 is according to some exemplary embodiments of the Figure 1 exemplary workloads that can be managed by the workload management module in
[0048] Figure 6 is a block diagram of an exemplary device plugin architecture used in conjunction with workload management according to some exemplary embodiments.
[0049] Figure 7 is a block diagram of an exemplary profiler architecture used in conjunction with workload management according to some exemplary embodiments.
[0050] Figure 8 is a block diagram of an exemplary workflow for scheduling workloads according to some exemplary embodiments.
[0051] Figure 9 is according to some exemplary embodiments of Figure 8 the exemplary pseudocode associated with the workflow in
[0052] Figure 10 is a diagram of an exemplary trained DL model used in conjunction with workload management according to some exemplary embodiments.
[0053] Figure 11 is a schematic diagram of the encoder network and the decoder network of the DL model in Figure 10 .
[0054] Figure 12 and Figure 13 is a schematic diagram of training the DL model in Figure 10 .
[0055] Figure 14 and Figure 15 is a schematic diagram of using the encoder network of the DL model in Figure 10 to generate a transformation function for workload scheduling.
[0056] Figure 16 is a block diagram of training a DL program using a deep learning (DL) training architecture (DLTA).
[0057] Figure 17 is a schematic diagram of generating a trained DL program using a neural network model trained within the DLTA.
[0058] Figure 18 is a flowchart of a method applicable to workload scheduling according to some exemplary embodiments.
[0059] Figure 19 is a block diagram of a representative software architecture according to some exemplary embodiments, which can be used in combination with various device hardware described herein.
[0060] Figure 20 is a block diagram of the circuitry of a device that implements an algorithm and executes a method according to some exemplary embodiments. DETAILED DESCRIPTION
[0061] First, it should be understood that although the following provides illustrative implementations of one or more embodiments, the systems and / or methods disclosed in connection with Figures 1 to 20 can be implemented using any number of techniques, whether currently known or not yet in existence. The present invention should in no way be limited to the illustrative implementations, drawings, and techniques described below, including the exemplary designs and implementations illustrated and described herein, but may be modified within the full scope of the appended claims and their equivalents.
[0062] In the following description, reference is made to the accompanying drawings, which form a part of this specification, and in which are shown, by way of illustration, specific embodiments in which the invention may be practiced. The embodiments are described in sufficient detail to enable those skilled in the art to practice the subject matter of the invention, it being understood that other embodiments may be used and that structural, logical, and electrical changes may be made without departing from the scope of the invention. Accordingly, the description of the exemplary embodiments that follow is not to be taken in a limiting sense, and the scope of the invention is defined by the appended claims.
[0063] As used herein, the term "network-based service infrastructure" includes a plurality of network devices (also referred to as hosts, nodes, or servers) that provide on-demand computing capabilities (e.g., via one or more virtual machines or other virtual resources running on the network devices) and storage capacity services to a population of end recipients (e.g., customers of the service infrastructure), where the end recipients are communicatively coupled to the network devices within the service infrastructure via a network. Customers of the service infrastructure can use one or more computing devices (or customer devices) to access and manage the services provided by the service infrastructure (e.g., workload scheduling services) via the network. The customer devices, the network, and the network-based service infrastructure may be collectively referred to as a "network architecture". Customers of the service infrastructure may also be referred to as "users".
[0064] As used herein, the term "resource usage" is synonymous with "computing resource usage" and refers to the computing resources being used by a virtual machine or container within a network-based service infrastructure. Computing resources may include one or more of the following resources of a host: central processing unit (CPU) resources, graphics processing unit (GPU) resources, memory resources, and other host resources. Additionally, the usage of computing resources can be monitored and can be dynamically changed and adjusted.
[0065] As used herein, the term "virtual machine" (VM) and the term "container" are used interchangeably to execute function code associated with the services provided by a network-based service architecture. More specifically, the function code and the function runtime can be hosted on a container or VM instantiated on a host device within the service architecture.
[0066] As used herein, the term "worker" (or "worker node") refers to a working machine that is part of a deep learning training architecture (DLTA) together with other workers. In some aspects, all working machines are coupled to each other (e.g., in a ring topology). Gradients can be exchanged between the working machines, and each working machine can perform its gradient averaging and gradient updates (e.g., gradient synchronization). As used herein, the terms "worker" and "working machine" can be used interchangeably.
[0067] As used herein, the terms "forward computation" and "backward computation" refer to computations performed in connection with the training of a neural network model (or another type of model). During forward and backward computations, the computations performed in the current iteration modify the weights based on the results of the previous iteration (e.g., based on the gradients generated at the end of the previous backward computation).
[0068] As used herein, the term "packaged workload" (or "packaging") means executing a workload while sharing GPU resources. For example, packaging workloads A and B means executing workload A using a first set of virtual GPU (vGPU) resources of a physical GPU, and after workload A has completed execution, executing workload B using a second (remaining) set of virtual resources of the physical GPU.
[0069] Scheduling workloads in a distributed computing cluster can consume a large amount of resources and rely on complex processing algorithms. For example, the Kubernetes (K8S) platform can be used to run and orchestrate containerized workloads (e.g., which can be regarded as an example of a cluster). The Kubernetes platform can include working machines (or nodes) that run containerized applications. In some aspects, scheduling in the Kubernetes platform refers to monitoring a set of containerized workloads (e.g., pods), analyzing their resource (e.g., CPU, memory, network) requests, and determining the optimal nodes for pod placement and running in combination with some high-level goal (e.g., shortest job completion time, highest resource utilization, etc.).
[0070] Due to the integration of thousands of computing cores on a single chip, a graphics processing unit (GPU) has powerful parallel processing capabilities. In this regard, the GPU can provide large-scale computing power support for different DL tasks. In some aspects, a device plugin mechanism can be configured in the Kubernetes platform so that GPU-related workloads can access the physical GPU cards installed in the nodes as extended hardware resources. Although the K8S cluster can access and manage GPUs through device plugins, most cluster management still has defects such as low GPU resource utilization.
[0071] Insufficient GPU resource utilization in the K8S platform may be caused by the enforcement of exclusive GPU usage, which prevents GPUs from being shared across pods. In other words, the default scheduling in K8S only supports addition and subtraction at the whole GPU card granularity, and does not support arbitrary sharding of individual GPUs. Since the GPU usage of each DL application is not affected by other applications, using integer GPU granularity is a design suitable for AI-based (e.g., DL) jobs. However, such processing may lead to serious underutilization of resources, especially in model development and inference scenarios where the utilization rate of a single GPU is low. In this regard, allowing more services to share a single physical GPU card can significantly improve resource utilization in the cluster.
[0072] In addition, even if a physical GPU can be virtualized into sharded virtual GPUs (also known as vGPUs) and the vGPUs between pods can be isolated, only local policies are available for scheduling these fractional GPUs. For example, simple bin-packing or spread methods. In the bin-packing scheduling policy, workloads are placed on nodes to reserve the least amount of unused vGPU resources, thus helping to optimize resource utilization. The spread scheduling policy evenly distributes workloads across the cluster, thus helping to maximize availability. However, both of these options ignore workload characteristics and do not consider any potential interference between the workloads packed into a single physical GPU.
[0073] Table 1 below shows the interference when two jobs are packed together. As shown in Table 1, when Job (or workload) A and Job B share a single physical GPU, their joint completion time (JCT) increases compared to when using the GPU exclusively. Although the increase in JCT is expected to some extent, it can be seen from Table 1 that the interference is workload-specific. For example, due to certain characteristics, Job C may have less impact on Job A than on Job B. In other words, in terms of job completion time, there is a "best partner" that can be packed with Job A.
[0074] Table 1
[0075]
[0076] The disclosed workload scheduling techniques can be used to expose a single physical GPU as shareable by containers executing workloads. By using the disclosed scheduling techniques, GPU-related workloads can be scheduled more efficiently by monitoring the actual utilization of resources and reducing interference between workloads sharing the GPU. The disclosed workload scheduling techniques can also be used to find the "best partner" (also referred to as the optimal partner or optimal workload partner) for each workload to share the GPU based on interference detection and interference avoidance.
[0077] In some aspects, the disclosed techniques can use a deep learning-based model to detect workload interference without any manual feature engineering. To train the deep learning-based model, a profiler module is designed to collect utilization metrics from different types of AI workloads. Additionally, the disclosed techniques use a scheduling pipeline to implement a scheduling algorithm with an online learning mode. In some aspects, the online learning mode can be used to handle different types of AI workloads.
[0078] Compared with existing solutions that use single-level scheduling (e.g., scheduling based on cloud service orchestration schemes that create warm containers, where each container uses / reserves actual host resources), the disclosed techniques use a workload management module that is configured with an interference detection-based scheduling algorithm in a GPU sharing scenario. When the GPU is shared by individual jobs, traditional scheduling techniques only use whole-card GPU scheduling or retain simple scheduling policies without considering any interference. Additionally, the disclosed scheduling techniques can be trained based on utilization metrics defined and collected by the profiler module (which can be part of the workload management module). In some aspects, the disclosed scheduling algorithm is trained using data collected by the profiler module from at least 1000 simulated workloads.
[0079] The workload management module also includes a device plugin and a scheduler module. The device plugin can be used to virtualize GPU resources into multiple vGPUs. The scheduler module can be configured with a scheduling pipeline that performs the following functions: (a) using the profiler module to perform a dry-run process for a new workload to collect metrics; (b) using a DL model to determine the category of the new workload; (c) assigning the new workload to its optimal workload partner; and (d) performing incremental DL model training and configuration. The following is combined with Figures 1 to 20 Provide additional description of the workload management module (including the device plugin, profiler module, and scheduler module).
[0080] Figure 1 is a high-level system diagram of a network architecture according to some exemplary embodiments, which has a workload management module that performs workload management functions. Refer to Figure 1, the network architecture 100 may include multiple devices (e.g., user devices) 102A, …, 102N (collectively referred to as devices 102), which are communicatively coupled to a network-based service infrastructure 114 via a network 112. Devices 102A, …, 102N are associated with corresponding users 106A, …, 106N, and may interact with the network-based service infrastructure 114 using a network access client (e.g., one of network access clients 104A, …, 104N). Network access clients 104A, …, 104N may be implemented as Web clients or application (app) clients.
[0081] Users 106A, …, 106N may be collectively referred to as “a user 106” or collectively as “users 106”. Each user 106 may be a human user (e.g., a person), a machine user (e.g., a computer configured by a software program to interact with devices 102 and the network-based service infrastructure 114), or any suitable combination thereof (e.g., a person assisted by a machine or a machine supervised by a person). Users 106 are not part of the network architecture 100, but each user 106 is associated with one or more devices 102 and may be a user of a device 102 (e.g., user 106A may be the owner of device 102A, and user 106N may be the owner of device 102N). For example, device 102A may be a desktop computer, an in-vehicle computer, a tablet computer, a navigation device, a portable media device, or a smartphone belonging to user 106A. Users 106A, …, 106N may use devices 102A, …, 102N to access services (e.g., workload scheduling services) provided by a workload management module of the network-based service infrastructure 114. In this regard, users 106 may also be referred to as “customers 106” or “tenants 106” of the network-based service infrastructure 114. For example, the workload scheduling service may include configuring GPU resources (e.g., virtual GPUs or other computing resources such as memory resources, CPU resources, etc.) and scheduling one or more workloads 108, …, 110 provided by any device 102 to execute on the configured GPU resources (e.g., packing at least two workloads to execute on the same vGPU), thereby improving resource utilization and reducing interference between workloads.
[0082] The network-based service infrastructure 114 can include multiple computing devices 116, 118, ……, 120 (which can also be referred to as nodes). For example, the computing device 118 can be configured as a master node, and the computing devices 116 and 120 can be configured as worker nodes. In some aspects, the computing devices 116, 118, ……, 120 are configured as part of a Kubernetes infrastructure, where the worker nodes (e.g., worker nodes 116 and 120) are used to schedule and execute workloads (e.g., workloads 108, ……, 110 configured through one or more of devices 102 and network 112) using one or more virtual containers. The computing devices 116, 118, and 120 include corresponding GPU resources 126, 130, and 136, which can be used to share the execution of workloads according to the disclosed techniques.
[0083] In some embodiments, the network-based service infrastructure 114 includes a workload management module 115 for performing the disclosed workload management functions. For example, the workload management module 115 can include a scheduler module 128 configured at the master node 118, at least one profiler module (e.g., profiler modules 122 and 132 configured at the corresponding worker nodes 116 and 120), and at least one device plugin (e.g., device plugins 124 and 134 configured at the corresponding worker nodes 116 and 120). Although Figure 1 a particular embodiment is shown where the scheduler module 128, profiler modules 122 and 132, and device plugins 124 and 134 are configured at different computing nodes, the present invention is not limited thereto, and other implementations of the workload management module 115 are possible (e.g., all components of the workload management module are implemented in a single computing device among the computing devices 116, ……, 120). In conjunction with Figure 2 a more detailed description of the overall architecture of the workload management module and its components is provided. In conjunction with Figure 6 a more detailed description of the device plugin is provided. In conjunction with Figure 7 a more detailed description of the profiler module is provided. In conjunction with Figure 8 、 Figure 9 and Figure 18 a more detailed description of the overall workflow performed by the components of the workload management module is provided.
[0084] Figure 1Any of the devices shown can be implemented in a general-purpose computer that is modified (e.g., configured or programmed) by software to be a special-purpose computer to perform the functions described herein for a machine, database, or device. As used herein, a "database" is a data storage resource that stores data that is structured as text files, tables, spreadsheets, relational databases (e.g., object-relational databases, NoSQL databases, network or graph databases), triple stores, hierarchical data stores, or any suitable combination thereof. Additionally, data accessed (or stored) via an application programming interface (API) or a remote procedure call (RPC) can be considered to be accessed from (or stored in) a database. Additionally, Figure 1 Any two or more of the devices or databases shown can be combined into a single machine, database, or device, and the functions described herein for any single machine, database, or device can be subdivided among multiple machines, databases, or devices.
[0085] Network 112 can be any network capable of communicating between machines, databases, and devices (e.g., devices 102A, …, 102N and devices 116, 118, …, 120 within the network-based service infrastructure 114). Thus, network 112 can be a wired network, a wireless network (e.g., a mobile or cellular network), or any suitable combination thereof. Network 112 can include one or more portions that make up a private network, a public network (e.g., the Internet), or any suitable combination thereof.
[0086] Figure 2 is a block diagram 200 of computing nodes 116, 118, and 120 according to some exemplary embodiments, the computing nodes 116, 118, and 120 being for implementing Figure 1 the workload management module 115 in. Referring to Figure 2 , the master node 118 includes a scheduler module 128 and other components 202. The worker node 116 includes a device plugin 124 and a profiler module 122. Similarly, the worker node 120 includes a device plugin 134 and a profiler module 132.
[0087] Device plugins 134 and 124 can be used to implement access to the sharded GPUs, as well as enforce limits and isolation between containers. For example, device plugin 134 exposes the physical GPU 210 as a virtualized GPU (vGPU) 214, and device plugin 124 exposes the physical GPU 208 as a vGPU 212.
[0088] The profiler module 122 includes suitable circuitry, logic, interfaces, and / or code and is used to collect utilization metrics, extract representations of the utilization metrics, and aggregate these representations into an input for the scheduling algorithm of the scheduler module 128. The profiler module 132 can perform similar functions as the profiler module 122. In some aspects, the profiler module 122 can be associated with a storage server (e.g., the Prometheus server 206 of the Kubernetes architecture), which can be used to store metrics and metadata from the worker node 116 and metrics and metadata from the profiler modules in other worker nodes 218 (e.g., metadata from the profiler module 132 of the worker node 120).
[0089] The scheduler module 128 includes suitable circuitry, logic, interfaces, and / or code and uses a deep learning (DL) model to learn workload-specific scheduling policies without manual input and (e.g., via scheduling decision 216) allocate workloads to compute nodes for execution while sharing GPU resources and minimizing the overall job completion time (JCT). For example, the scheduler module 128 uses the AI-driven analyzer 204 to determine the optimal placement of workloads based on the scheduling algorithm. In some aspects, the AI-driven analyzer 204 is deployed on the same node as the Prometheus server for easy communication (e.g., as Figure 2 shown).
[0090] In some aspects, the scheduling algorithm used by the scheduler module 128 is based on interference detection during workload sharing of the GPU. In contrast, existing workload schedulers only use whole-card GPUs or retain simple scheduling policies without considering any interference in the case of GPU sharing.
[0091] Figure 3 is a schematic diagram of Table 300 of workload (or job) packing and the corresponding combined completion time for completing the workload according to some exemplary embodiments. As Figure 3 shown, packing two workloads into one physical GPU (e.g., packing job A with job B) causes interference and lengthens the JCT of the workload. Figure 3It is also illustrated that the interference can be workload - specific, and workloads can be selected (e.g., using the disclosed techniques) to share GPU resources to minimize interference and reduce JCT. More specifically, the workload management module 115 can use a profiler module (e.g., profiler module 122) to collect utilization metrics and utilize a DL model (e.g., used by the AI - driven analyzer 204) to predict the optimal workload to be packed with the selected workload. In some aspects, the "optimal workload" can indicate the workload that causes the least interference when packed with the selected workload and / or the workload that causes the least JCT for each packed workload.
[0092] Figure 4 is a schematic diagram of Table 400 of exemplary utilization metrics collected and used by the workload management module in Figure 1 . Referring to Figure 4 , Table 400 shows exemplary cluster utilization metrics that can be collected by the profiler module and used by an AI - based scheduling algorithm (e.g., the DL model using the AI - driven analyzer 204). In some aspects, the utilization metrics listed in Table 400 can be used to train the DL model.
[0093] Figure 4 Some exemplary utilization metrics shown include histo_gpu_usage (indicating the GPU usage histogram of a pod), histo_mem_usage (indicating the memory usage histogram in a node), type_GPU (indicating the GPU type, e.g., Tesla V100, GeForce RTX 2080, Ti, RTX 8000, Titan X, etc.). Table 2 below provides additional utilization metrics.
[0094] Figure 5 is a schematic diagram of an exemplary workload 500 that can be managed by the workload management module in Figure 1 . In some aspects, a DL model (e.g., as used by the AI - driven analyzer 204) can be trained using the utilization metrics collected by the profiler module, while the training workloads (e.g., including the workloads shown in Figure 5 ) are executed by the worker nodes of the network - based service infrastructure 114.
[0095] Figure 6 is a block diagram of an exemplary device plugin architecture 600 used in conjunction with workload management. Referring to Figure 6, the device plugin architecture 600 includes a Kubernetes Application Programming Interface (K8S API) server 602, a K8S scheduler 604, a K8S kubelet 606, docker (or container management tool) 624, a device plugin 608, containers 626, a host system 634, physical GPU resources 614, a GPU user space driver 616, and a GPU kernel space driver 622. The K8S API server 602, the K8S scheduler 604, the K8S kubelet 606, and docker 624 can be modules configured as part of a Kubernetes-based network architecture.
[0096] The device plugin 608 includes a K8S device plugin with a GPU manager 610 and a virtual GPU registration server 612. Containers 626 can be used to configure a vGPU library 632 and execute workloads 628 and 630. The GPU user space driver 616 includes a GPU driver API 618 (also known as Compute Unified Device Architecture (CUDA)) and a GPU monitoring API 620 (also known as NVIDIA Management Library API (NVML API)).
[0097] In a Kubernetes infrastructure, processing can be based on the assumption that all K8S devices / modules on a node are the same and that GPU usage at the container level is exclusive. When workload scheduling uses GPUs on the same node, existing APIs may not be able to express GPU requirements (e.g., sharing the same GPU between containers) or GPU hardware characteristics (memory, computing power, etc.) in the K8S pod specification.
[0098] In some embodiments, the device plugin 608 uses a Kubernetes extension mechanism to enable Kubernetes-managed containers to access GPUs. Compared with the NVIDIA device plugin, the disclosed device plugin 608 can provide sliced GPUs (also known as vGPUs), where vGPU usage limit enforcement and vGPU isolation between containers are achieved. These features can be implemented through the K8S LD_PRELOAD mechanism.
[0099] In some aspects, the LD_PRELOAD mechanism is a technique that affects the linking and symbol (function) resolution of shared libraries at runtime. In short, a library is a collection of compiled functions that can be used directly without being rewritten. This can be achieved by including library code (e.g., static libraries) in a program or by dynamically linking at runtime (e.g., shared libraries). In some aspects, the shared libraries used to build a program may require runtime linker / loader support. Thus, the required symbols can be loaded and prepared before the program is executed. In some aspects, the LD_PRELOAD mechanism can be used in the program execution preparation phase. In some aspects, the Linux system programs ld.so and ld-linux.so (dynamic linkers / loaders) can use LD_PRELOAD to load the specified shared libraries. The dynamic loader can first load the shared libraries in LD_PRELOAD before any other libraries. Thus, once a user wants to share GPU memory and computing resources between multiple isolated containers, it is necessary to intercept a special library, namely the GPU driver API 618, through the LD_PRELOAD mechanism. In some aspects, intercepting at this level is to support CUDA-based applications and have a stable public API.
[0100] As Figure 6 shown, the overall architecture of the device plugin 608 includes three components: the K8S device plugin 610, the vGPU registration server 612, and the vGPU library 632 (also called libintercept.so in Figure 6 ).
[0101] The K8S device plugin 610 is used to notify the K8S kubelet 606 about the GPU. The K8S device plugin 610 runs on the host and is responsible for creating vGPUs using the physical GPU resources 614 and communicating with the K8S kubelet 606 through a remote procedure call API (e.g., gRPC) service.
[0102] The K8S device plugin 610 registers with the K8S kubelet 606 through registration, request, and allocation calls 636 to notify the kubelet of the existence of the K8S device plugin 610. When a user requires GPU devices in a container specification, the kubelet arbitrarily selects the corresponding number of devices from the device list sent by the K8S device plugin.
[0103] After successful registration, the kubelet sends a ListAndWatch request 638 to the GPU manager of the K8S device plugin to query device information. The GPU manager returns a list of devices managed by the GPU manager to the kubelet. What is sent to the kubelet is not the physical GPUs, but a list of vGPUs. In some aspects, the physical GPU resources 614 are virtualized in two resource dimensions: memory and computing resources.
[0104] The vGPU registration server 612 is configured to run on the host to issue container configurations and monitor the containers allocated with vGPUs. When a container requests GPU resources, the server sends the configuration of the container (e.g., the required GPU resources) and the name of the container to the vGPU manager of the K8S device plugin 610.
[0105] The vGPU library 632 runs in the container 626 to manage GPU resources. When a first GPU application is executed in the container 626, the vGPU library 632 can be started. After startup, the vGPU library 632 registers with the vGPU manager. It intercepts the memory-related APIs and computing-related APIs in the CUDA library through the LD_LIBRARY_PATH mechanism. In some aspects, LD_LIBRARY_PATH is an environment variable in the Linux system that affects the runtime linking of programs and allows certain directories to be loaded before the standard set of directories.
[0106] In some embodiments, the device plugin architecture 600 can perform the following processing flow. The GPU manager of the K8S device plugin 610 registers with the kubelet 606 through vGPUs and then processes the ListAndWatch request 638. Once a GPU request is received, the kubelet sends the request to the GPU manager. The GPU manager sends a scheduling request to the scheduler, and the scheduler returns a response with the allocated GPUs. The GPU manager sends a response to the vGPU registration server 612. The GPU manager returns the environment variables of the container, the mount information (e.g., the host file system 634 mounted on the container 626), and the device information to the kubelet. The kubelet creates and initializes the container 626. Before the container is executed, the GPU driver API (e.g., CUDA API) is intercepted by the LD_LIBRARY mechanism, which allows some directories to be loaded first. Deploy vGPUs for the container 626. The vGPU server 612 manages vGPU resources and cleans up the container when the container is deactivated.
[0107] Figure 7 is a block diagram of an exemplary profiler architecture 700 used in conjunction with workload management according to some exemplary embodiments. Refer to Figure 7, the profiler architecture 700 includes compute nodes that implement components of the workload management module, such as the scheduler module 710, the profiler modules 718 and 730, and the device plugins 712 and 736. More specifically, the profiler architecture includes a master node 702 with an API server 708 and a scheduler module 710, and worker nodes 704, …, 706. The worker node 704 includes a GPU 728, a profiler module 718, a device plugin 712, a GPU driver API 714, and a container runtime 716. The profiler module 718 includes an AI-driven analyzer (which can be a component of the scheduler module 710), a K8S Prometheus server 722, a CPU metric collector module 724, and a GPU metric collector module 726. The worker node 706 includes a GPU 742, a profiler module 730, a device plugin 736, a GPU driver API 738, and a container runtime 740. The profiler module 730 includes a CPU metric collector module 732 and a GPU metric collector module 734.
[0108] The scheduler module 710, the AI-driven analyzer module 720, the profiler modules 718 and 730, and the device plugins 712 and 736 are functionally similar to the corresponding modules discussed in Figures 1 to 6 discussion.
[0109] The profiler modules 718 and 730 can be used to collect and analyze GPU metrics at various levels, such as pod level, node level, job / workload level, GPU level, CPU level, memory level, and network traffic. Table 2 below provides exemplary metrics that can be defined, collected, and stored by the profiler modules 718 and 730.
[0110] Table 2
[0111]
[0112]
[0113]
[0114] The GPU metric collector modules 726 and 734 can include NVIDIA's Data Center GPU Manager (DCGM), which is used to collect the disclosed utilization metrics (such as the metrics listed in Tables 1 and 2) and store them in the server 722. The CPU metric collector modules 724 and 734 are used to collect CPU metrics and store them in the server 722. In some aspects, each cluster can use a single Prometheus server.
[0115] The AI-driven analyzer module 720 is used to read metrics from the Prometheus server 722, determine the workload type using a DL model (e.g., classify the workload as "visible" or "invisible"), update the utilization metrics stored in the server 722 using the metrics from the dry-run process (e.g., as discussed in conjunction with Figure 8 ), and resume the scheduling decision.
[0116] In some aspects, analysis functions (e.g., loop pattern detection and trend prediction) are built into the profiler (e.g., via the AI-driven analyzer module 720) to predict the workload type, utilization, etc. The analysis results are written back to the server 722 as part of the object annotations. Additionally, the profiler modules 718 and 730 can generate short-term trial workloads (i.e., dry runs) with different device placements (allocating different types and quantities of GPUs) and track the execution efficiency. Thus, the proposed scheduling algorithm can perform dynamic optimization by considering the results of the trial workloads.
[0117] Figure 8 is a block diagram of an exemplary workflow 800 for scheduling workloads according to some exemplary embodiments. Figure 8 The exemplary operations (or steps) in
[0118] can be performed by components of the workload management module disclosed herein, e.g., the scheduler module 804, the profiler module 808, and the AI-driven analyzer module 810 with a trained DL model 812. i where the metadata of a single job on a single GPU (the metadata of workload i is denoted as F i ), and the JCT is denoted as T ij ); and (b) two jobs packed onto a single GPU (the JCT of the packed jobs i and j is denoted as T
[0119] Operations 1, 2, 3: To predict the optimal workload partner for packing with the new workload 802, each upcoming workload will first be assigned to a single idle GPU and executed within a predefined time (also known as a dry run). For example, a dry run 806 of the new workload 802 is performed. The scheduler 804 can arbitrarily assign the workload to a single GPU and have that GPU run the first iteration. The profiler module 808 estimates the utilization metrics of the workload 802 and represents it as F n , which is the input to the AI-driven analyzer module 810 with the DL model 812.
[0120] Operation 4: Classification of the workload 802 can be performed. The DL model 812 in the AI-driven analyzer module 810 can classify the workload 802 into at least ten types (or classes) based on the utilization metrics of the DL model 812: nine visible (or previously known) workload types and one invisible (or previously unknown) workload type. At operation 814, if the workload 802 belongs to a visible type, processing continues at operation 5b. At operation 814, if the workload 802 belongs to an invisible type, processing continues at operation 5a, where the dry run is maintained until completion.
[0121] Operations 5b and 6b are associated with the elimination function. If the workload 802 is classified as one of the visible types (e.g., one of the nine visible types), its dry run terminates at operation 5b (also known as operation 820). The workload 802 is reallocated (e.g., by the scheduler module 804) to other GPUs according to the packing table, where the optimal workload partner can be selected (e.g., at operation 816) to share the GPU and thus minimize the JCT of the workload. In some aspects, the packing table is constructed during an offline training phase at operation 0.
[0122] Operations 5a and 6a are associated with the online learning function of the DL model 812. If the workload 802 is classified into the invisible class of workloads, its dry run process continues (at operation 818) until the workload is completed. The profiler module 808 records the corresponding utilization metrics and metadata of the executed workload. At operation 822, the AI-driven analyzer module 810 (e.g., the packing table used by the analyzer) is updated. In some aspects, this update includes two sub-operations: (a) incrementing the number of visible classes by 1; (b) updating the packing table with the metadata of the workload.
[0123] Figure 9 is a schematic diagram of an exemplary pseudocode 900 associated with the Figure 8 workflow in accordance with some exemplary embodiments.
[0124] Figure 10 is a schematic diagram of an exemplary trained DL model 1000 for use in conjunction with workload management according to some exemplary embodiments. The DL model 1000 may be the same as the DL model 812 discussed in conjunction with Figure 8 discussion.
[0125] In some embodiments, the DL model 1000 includes a neural network encoder 1004 and a neural network decoder 1008. The DL model 1000 can be used to predict whether an incoming (or new) workload (e.g., workload 802) belongs to a type (or class) among a plurality of visible types (or classes) detected during offline simulation and training of the DL model 1000.
[0126] In some aspects, the encoder 1004 can be used for dimensionality reduction. For example, the input 1002 can be configured as input X = [m_1, m_2] and can include utilization metrics collected from a plurality (e.g., approximately 1000) of workloads with a dimensionality of 200. The encoder output 1006 can be designated as output Z = E(x), where E is a transformation function that reduces the dimensionality to 10 (e.g., nine visible classes and one invisible class). The encoder output 1006 is also the input to the decoder 1008, and the decoder 1008 generates output 1010. The output 1010 can be designated as where D is a second transformation function based on the output Z of the encoder 1004.
[0127] Figure 11 is according to some exemplary embodiments of Figure 10 schematic diagram 1100 of the encoder network and decoder network of the DL model in Figure 11 More specifically, Figure 10 shows a detailed schematic diagram of the convolutional layers and dimensionalities associated with the encoder 1004 and decoder 1008 used in the DL model 1000 in
[0128] Figure 12 and Figure 13 is a schematic diagram of the training of the DL model according to some exemplary embodiments of Figure 10 in Figure 12 , schematic diagram 1200 is the supervised training phase of the DL model 1000. More specifically, the matrix table 1202 can be used as the input to the encoder 1204, while the packing table 1208 can be used as the output of the decoder 1206. The matrix table 1202 includes the matrix m of the utilization metrics of workloads i and j i m j . The packing table 1208 includes the JCT t when workloads i and j are packed together to share GPU resources ij .
[0129] Figure 13 Schematic diagram 1300 showing a more detailed view of training the DL model 1000. Refer to Figure 13 , the training of the DL model uses utilization metrics m1 and m2 as the input 1302 to the first convolutional layer 1304 of the encoder 1204. Before generating the concatenated layer 1312 using the encoder output of the convolutional layer 1310 corresponding to the utilization metrics m1 and m2, additional convolutional layers 1306, 1308, and 1310 can be applied. A regression layer 1314 can be applied to the output of the concatenated layer 1312 to generate JCT t 12 as the output of the DL model 1000.
[0130] Figure 14 and Figure 15 is a schematic diagram of using the encoder network of the DL model in Figure 10 to generate a transfer function for workload scheduling according to some exemplary embodiments.
[0131] Refer to Figure 14 , schematic diagram 1400 shows the matrix table 1202, including the matrix m of the utilization metrics of workloads i and j i m j which matrix m i m j can be provided as input to the encoder 1204 of the DL model 1000. The output of the encoder 1204 is a useful feature set F 1402 corresponding to the transfer function E(m). In some aspects, E is a transfer function that reduces the dimension to 10. In some embodiments, the useful feature set F1402 is a subset of the utilization metrics of each workload in the workload. In this regard, the term "useful feature set" is used in combination with the disclosed technology to indicate a subset of the utilization metrics generated as the output of the DL model encoder that uses the matrix of workload utilization metrics as input.
[0132] The following is an example of how the AI-driven analyzer module of the scheduler module disclosed herein uses the useful feature set. For each new workload n, the scheduler module assigns it to a single idle GPU and executes the workload within a predefined time (e.g., as discussed in conjunction with Figure 8 ). The utilization metrics from the workload n are collected and represented as m n . The workload type can be determined by classifying the workload as at least one workload type of the prior workloads with the utilization metric m1. The workload management module 115 can perform the following functions to determine the workload type:
[0133] (a) If |F(m n) - If |F(m1)|2 <= θ, then workload n belongs to type 1 associated with a prior workload having a utilization metric m1 (where θ is a predefined threshold and F is a set of useful features determined for the new workload n and the prior workload 1). In some aspects, when the difference between the useful feature set of the prior workload and the useful feature set of the new workload is no greater than the threshold θ, the workload management module 115 may configure the shared execution of the workloads (e.g., the prior workload and the new workload).
[0134] (b) If |F(m n ) - F(m1)|2 > θ, then workload n is an invisible type. In some aspects, when the difference between the useful feature set of the prior workload and the useful feature set of the new workload is greater than the threshold θ, the workload management module 115 may avoid configuring the shared execution of the workloads (e.g., the prior workload and the new workload).
[0135] After a dry run (e.g., dry run 806), if workload n belongs to any known type (or class), then the optimal workload partner for packing with workload n can be looked up in the packing table (to minimize interference and JCT, etc.). If workload n belongs to an invisible type (e.g., does not match a known type), then the scheduler may avoid configuring the shared execution of the workload. In this case, the dry run can continue until workload n is completed and its utilization metric is profiled (e.g., as a new type that can be used to match subsequent workloads). In this regard, after the dry run of workload n is completed, the packing table is updated with the new workload type.
[0136] In some embodiments, the disclosed AI-based scheduling algorithm considers the interference patterns of multiple AI jobs and schedules the jobs to share the GPU with the least interference to the jobs. In some aspects, the interference patterns between individual AI jobs are obtained offline through AI training. In other words, these patterns are detected based on AI training of the data obtained from running the jobs that share these GPUs.
[0137] Figure 15 FIG. 1500 shows a schematic diagram of obtaining a transfer function E(m) 1508 using the convolutional layer of the encoder of the DL model 1000. More specifically, the utilization metric 1502 is provided as an input to the encoder of the DL model. In some aspects, a subset of the convolutional layers of the encoder is used to generate the transfer function E(m) 1508. In Figure 15 the example shown, the transfer function E(m) 1508 is generated by freezing the weights of the convolutional layers 1504 and 1506, and this transfer function 1508 is the output of the second convolutional layer 1506. Other configurations that utilize a subset of the encoder convolutional layers to generate the transfer function can also be used.
[0138] Figure 16 is a block diagram 1600 of training a deep learning (DL) model 1608 using a DL training architecture (DLTA) 1604 according to some exemplary embodiments. In some exemplary embodiments, machine-learning programs (MLPs) (including deep learning programs) are also collectively referred to as machine learning algorithms or tools for performing operations associated with relevant data or other artificial intelligence (AI)-based functions.
[0139] As Figure 16 shown, deep learning program training 1606 can be performed within the DLTA 1604 based on training data 1602 (which may include utilization metrics or other predefined metrics corresponding to a predefined output). During deep learning program training 1606, the features of the training data 1602 can be evaluated to further train the DL model 1608. The DL program training 1606 results in a trained DL model 1608, which can include one or more classifiers 1614 that can be used to provide an evaluation 1612 based on new data 1610. The trained DL model 1608 can be the same as the DL model 812 used by the AI-driven analyzer module 810.
[0140] Deep learning is part of machine learning, which is a field of study that enables computers to learn without being explicitly programmed. Machine learning explores the research and construction of algorithms (also referred to as tools in this article) that can learn from existing data, correlate data, and make predictions on new data. Such machine learning tools operate by constructing models based on exemplary training data (e.g., training data 1602) to make data-driven predictions or decisions, represented as an output or evaluation 1612. Although exemplary embodiments of some machine learning tools (e.g., deep learning training architectures) are provided, the principles provided herein can be applied to other machine learning tools.
[0141] In some exemplary embodiments, different machine learning tools can be used. For example, during program training 1606 (e.g., for correlating training data 1602), tools such as Logistic Regression (LR), Naive-Bayes, Random Forest (RF), neural network (NN), matrix factorization, and Support Vector Machine (SVM) can be used.
[0142] Two common types of problems in machine learning are classification problems and regression problems. Classification problems (also known as categorization problems) aim to classify items into one of several categorical values (e.g., is this object an apple or an orange?). Regression algorithms aim to quantify some item (e.g., by providing a real value). In some embodiments, DLTA1604 can use a machine learning algorithm that utilizes training data 1602 to find correlations between identified features that affect the result.
[0143] The machine learning algorithm utilizes the features of the training data 1602 to analyze new data 1610 (e.g., utilization metrics of a new workload) to generate an assessment 1612 (e.g., an assessment of the workload type performed at operation 814 in Figure 8 ). These features include individual measurable attributes of the phenomena that are observed and used to train the ML program. The concept of features is related to the concept of explanatory variables used in statistical techniques such as linear regression. In pattern recognition, classification, and regression, selecting informative, distinguishable, and independent features is very important for the efficient operation of the MLP. Features can be of different types, e.g., numerical features, strings, and graphics. In some aspects, the training data can be of different types, where the features are numbers for use by a computing device.
[0144] The machine learning algorithm utilizes the training data 1602 to find correlations between identified features that affect the result of the assessment 1612. In some exemplary embodiments, the training data 1602 includes labeled data, which is known data for one or more identified features and one or more results. Through the training data 1602 (which can include identified features), the DL program training 1606 within the DLTA 1604 is used to train the DL model. The result of the training is the trained DL model 1608. When the DL model 1608 is used to perform an assessment, new data 1610 is provided as input to the trained DL model 1608, and the DL model 1608 generates an assessment 1612 as output.
[0145] Figure 17 is a schematic diagram 1700 of generating a trained DL model 1706 using a neural network model 1704 trained within the DLTA1604 according to some exemplary embodiments. Referring to Figure 17 , the source data 1702 can be analyzed by the neural network model 1704 (or another type of machine learning algorithm or technique) to generate a trained DL model 1706 (which can be the same as the trained DL model 1608). The source data 1702 can include a training data set, e.g., the training data 1602, including data identified by one or more features. As used herein, the terms "neural network" and "neural network model" are used interchangeably.
[0146] Machine learning techniques train a model to make accurate predictions based on data fed into the model (e.g., what the user said in a given utterance; whether a noun is a person, place, or thing; what the weather will be tomorrow). During the learning phase, the model is developed against a training dataset of inputs to optimize the model to correctly predict the output for a given input. Generally, the learning phase can be supervised, semi-supervised, or unsupervised; the level of indicating the "correct" output corresponding to the training input gradually decreases. During the supervised learning phase, all outputs are provided to the model, and the model is guided to develop general rules or algorithms that map the input to the output. In contrast, during the unsupervised learning phase, no expected output is provided for the input, so that the model can develop its own rules to discover relationships in the training dataset. During the semi-supervised learning phase, an incomplete-label training set is provided, where for the training dataset, some outputs are known and some outputs are unknown.
[0147] The model can run several rounds against the training dataset, where the training dataset is repeatedly fed into the model to refine the results of the model (i.e., the entire dataset is processed during one round). During an iteration, the model (e.g., a neural network model or other type of machine learning model) runs against a small batch (or a portion) of the entire dataset. During the supervised learning phase, the model is developed to predict the output for a given set of inputs (e.g., source data 1702) and is evaluated over several rounds to more reliably provide an output that specifies the given input corresponding to the maximum number of inputs of the training dataset. In another example, during the unsupervised learning phase, the model is developed to cluster the data into n groups and the dataset is evaluated over several rounds to determine the consistency of the dataset in placing a given input into a given group and the reliability of the dataset in producing n desired clusters in each round.
[0148] After running one round, the model is evaluated and the values of its variables (e.g., weights, biases, or other parameters) are adjusted to attempt to iteratively refine the model better. As used herein, the term "weights" is used to refer to the parameters used by a machine learning model. During backpropagation, the model can output gradients that can be used to update the weights associated with the forward computation.
[0149] In different aspects, the evaluation is biased towards false negatives, towards false positives, or towards the overall accuracy of the model. These values can be adjusted in several ways depending on the machine learning technique used. For example, in genetic or evolutionary algorithms, the values of the models that are most successful in predicting the desired output are used to develop the values of the models to be used in subsequent rounds (which may include random transformations / mutations) to provide additional data points. Those of ordinary skill in the art are familiar with several other machine learning algorithms applicable to the present invention, including linear regression, random forests, decision tree learning, neural networks, deep neural networks, etc.
[0150] Each model develops rules or algorithms over several rounds by varying the values of one or more variables that affect the input to more closely map to the desired outcome. However, since the training dataset may vary, and vary quite significantly, it may not be possible to achieve optimal accuracy and precision. Thus, the several rounds that make up the learning phase can be set to a given number of trials or a fixed time / computation budget, or can be terminated before reaching that number / budget if the accuracy of a given model is high enough or low enough or has reached an accuracy plateau. For example, if the training phase is designed to run for n rounds and produce a model with an accuracy of at least 95%, and the model is produced before the nth round, then the learning phase can end early and use the model produced that meets the final target accuracy threshold. Similarly, if a given model is not accurate enough to meet a random chance threshold (e.g., the model has an accuracy of only 55% in determining the true / false output for a given input), then the learning phase of that model may end early while other models in the learning phase may continue training. Similarly, when a given model continues to provide similar accuracy over several rounds or its results are wavering (a performance plateau has been reached), the learning phase of the given model may be terminated before reaching the number of rounds / computation budget.
[0151] Once the learning phase is complete, the model is finalized. In some exemplary embodiments, the finalized model is evaluated according to test criteria. In the first example, a test dataset including the known output of its input is fed into the finalized model to determine the accuracy of the model when processing data on which it has not been trained. In the second example, the false positive rate or false negative rate can be used to evaluate the finalized model. In the third example, the partitioning between data clusters in each model is used to select the model that produces the clearest boundaries for the data clusters.
[0152] In some exemplary embodiments, the DL model 1706 is trained by a neural network model 1704 (e.g., a deep learning, deep convolutional, or recurrent neural network) that includes a series of "neurons", such as long short-term memory (LSTM) nodes arranged in a network. A neuron is an architectural element for data processing and artificial intelligence (especially machine learning), including a memory that can determine when to "remember" and when to "forget" the values stored in the memory based on the weights of the inputs provided to a given neuron. Each neuron used herein is configured to receive a predefined number of inputs from other neurons in the network in order to provide relational and sub-relational outputs for the content of the frame being analyzed. Individual neurons can be linked together and / or organized into a tree structure in various configurations of the neural network to provide interactive and relational learning modeling to understand how each frame in a discourse relates to one another.
[0153] For example, an LSTM, as a neuron, includes a number of gates to process an input vector (e.g., a phoneme of a discourse), a storage unit, and an output vector (e.g., a context representation). The input gate and the output gate control the information flowing into and out of the storage unit, respectively, while the forget gate selectively removes information from the storage unit based on the inputs from earlier linked units in the neural network. The weight vectors and bias vectors of the various gates are adjusted throughout the training phase, and once the training phase is complete, these weights and biases will ultimately be determined for normal operation. Those skilled in the art will understand that neurons and neural networks can be constructed by programming (e.g., via software instructions) or by dedicated hardware that links each neuron to form a neural network.
[0154] Neural networks utilize features to analyze data to generate an assessment (e.g., to identify speech units). A feature is an individual measurable property of an observed phenomenon. The concept of a feature is related to the concept of an explanatory variable used in statistical techniques such as linear regression. Additionally, a deep feature represents the output of a node in a hidden layer of a deep neural network.
[0155] A neural network, sometimes referred to as an artificial neural network or neural network model (e.g., neural network model 1704), can include a computing system based on the biological neural network of an animal brain. Such systems improve performance (referred to as learning) step by step to perform tasks, usually without task-specific programming. For example, in image recognition, a neural network can be taught to identify images containing objects by analyzing exemplary images that have been labeled with object names, and after learning the objects and names, the analysis results can be used to identify objects in unlabeled images. A neural network is based on a collection of connected units called neurons, where each connection (called a synapse) between neurons can transmit a unidirectional signal, and the activation intensity of the unidirectional signal varies with the connection strength. A receiving neuron can activate the signal and propagate the signal to downstream neurons connected to it, usually based on whether the combined incoming signals from potentially multiple transmitting neurons have sufficient strength, where strength is a parameter.
[0156] A deep neural network (DNN) is a stacked neural network composed of multiple layers. These layers consist of nodes, which are the locations where computations occur, roughly mimicking neurons in the human brain that trigger when exposed to sufficient stimuli. The nodes combine the input from the data with a set of coefficients or weights that can amplify or attenuate the input, thereby assigning importance to the input for the task the algorithm is trying to learn. The sum of these input weight products is calculated and passed through a node activation function to determine whether and to what extent the signal further propagates through the network to affect the result. DNNs use a cascade of multiple non-linear processing units for feature extraction and transformation. Each successive layer uses the output of the previous layer as input. High-level features are derived from low-level features to form a hierarchical representation. The layers after the input layer can be convolutional layers that produce feature maps that filter the results of the input and are used by the next convolutional layer.
[0157] In the training of a DNN architecture, regression (structured as a set of statistical processes for estimating relationships between variables) can include minimizing a cost function. The cost function can be implemented as a function that returns a number representing how well the neural network performs in mapping training examples to the correct output. During training, if the cost function value is not within a predetermined range, based on known training images, backpropagation is used, where backpropagation is a common method for training artificial neural networks and is used together with optimization methods such as the stochastic gradient descent (SGD) method.
[0158] Using backpropagation can include propagation and weight updates. When an input is provided to a neural network, the input propagates forward layer by layer through the neural network until it reaches the output layer. Then, a cost function is used to compare the output of the neural network with the desired output, and an error value is calculated for each node in the output layer. The error value propagates backward from the output until each node has an associated error value that roughly represents its contribution to the original output. Backpropagation can use these error values to calculate the gradient of the cost function associated with the weights in the neural network. The calculated gradient is fed into a selected optimization method to update the weights, thereby attempting to minimize the cost function.
[0159] Although the training architecture 1604 is referred to as a deep learning training architecture using a neural network model (the program being trained is referred to as a trained deep learning model, e.g., DL model 1608 or DL model 1706), the present invention is not limited thereto, and other types of machine learning training architectures can also be used for model training using the techniques disclosed herein.
[0160] Figure 18 is a flowchart of a method 1800 applicable to workload scheduling according to some exemplary embodiments. Method 1800 includes operations 1802, 1804, 1806, 1808, and 1810. By way of example and not limitation, method 1800 is performed by one or more components of workload management module 115 (also referred to as Figure 19 workload management module 1960 of Figure 20 or 2060 of
[0161] Operation 1802: Initiate the execution of a first workload on one of a plurality of GPUs. For example, scheduler module 804 schedules the execution of workload 802 on at least one vGPU associated with one of the physical GPUs 728 of worker node 704.
[0162] Operation 1804: Determine the utilization metric of the first workload. For example, profiler module 808 (which may be the same as profiler module 718) determines the utilization metric associated with the idle execution of workload 802 on at least one vGPU.
[0163] Operation 1806: Use the transformation function of the DL model to extract a useful feature set of the utilization metric of the first workload. For example, as discussed in connection with Figures 12 to 16 the encoder network of DL model 812 is used to determine the transformation function E(m). Then, the transformation function and the obtained utilization metric are used to extract the useful feature set F. In some aspects, the useful feature set is a subset of the utilization metric.
[0164] Operation 1808: Determine the workload type of the first workload using the useful feature set. For example, as discussed in conjunction with Figure 14 the workload type can be determined by classifying the workload into at least one workload type of the prior workloads based on a comparison of the corresponding useful feature sets.
[0165] Operation 1810: Configure shared execution of the first workload and the second workload on a second GPU among multiple GPUs based on packing the first workload and the second workload. The second workload is associated with the determined workload type of the first workload. For example, if the difference between the useful feature set of the first workload and the prior workload is less than a threshold amount, the first workload can be indicated as being of the same type as the prior workload. Then, a packing table can be referred to and used to select an optimal workload partner to pack and share GPU resources. An optimal workload partner of the same type as the first workload type can be selected, and workload interference and JCT can be minimized as much as possible.
[0166] Figure 19 is a block diagram of a representative software architecture 1900 according to some exemplary embodiments, which can be used in combination with various device hardware described herein. Figure 19 is only a non-limiting example of the software architecture 1902, and it should be understood that many other architectures can be implemented to facilitate the functions described herein. The software architecture 1902 can be executed on Figure 20 the hardware of the computing device 2000 in, which includes a processor 2005, a memory 2010, storage elements 2015 and 2020, and I / O components (or interfaces) 2025 and 2030. A representative hardware layer 1904 is shown, for example, which can represent Figure 20 the computing device 2000 in. The representative hardware layer 1904 includes one or more processing units 1906 with associated executable instructions 1908. The executable instructions 1908 represent the executable instructions of the software architecture 1902, including Figures 1 to 18 the implementation of methods, modules, etc. in. The hardware layer 1904 also includes a memory and / or storage module 1910, which also has executable instructions 1908. The hardware layer 1904 may also include other hardware 1912, which represents any other hardware of the hardware layer 1904, for example, other hardware shown as part of the computing device 2000.
[0167] In Figure 19In the exemplary architecture, the software architecture 1902 can be conceptually visualized as a stack of layers, where each layer provides a specific function. For example, the software architecture 1902 can include layers such as an operating system 1914, libraries 1916, frameworks / middleware 1918, applications 1920, and a presentation layer 1944. In operation, the applications 1920 and / or other components within the layer can make application programming interface (API) calls 1924 through the software stack and receive responses, return values, etc., shown as messages 1926, in response to the API calls 1924. Figure 19 The layers shown are representative in nature, and not all software architectures 1902 have all layers. For example, some mobile or specialized operating systems may not provide frameworks / middleware 1918, while other operating systems may provide such layers. Other software architectures may include additional or different layers.
[0168] The operating system 1914 can manage hardware resources and provide common services. For example, the operating system 1914 can include a kernel 1928, services 1930, drivers 1932, and a workload management module 1960. The workload management module 1960 can include a scheduler module 1962 (with an AI-driven analyzer module for DL models), a profiler module 1964, and a device plugin 1966. The kernel 1928 can act as an abstraction layer between the hardware and other software layers. For example, the kernel 1928 can be responsible for memory management, processor management (e.g., scheduling), component management, networking, security settings, etc. The services 1930 can provide other common services to other software layers. The drivers 1932 can be responsible for controlling or interacting with the underlying hardware. For example, depending on the hardware configuration, the drivers 1932 can include a display driver, a camera driver, a driver, a flash driver, a serial communication driver (e.g., a Universal Serial Bus (USB) driver), a driver, an audio driver, a power management driver, etc.
[0169] In some aspects, the workload management module 1960, the scheduler module 1962, the profiler module 1964, and the device plugin 1966 can be the same as (and perform the same functions as) the corresponding similarly named modules discussed in conjunction with Figures 1 to 18 the corresponding similar names.
[0170] Library 1916 can provide a common infrastructure that can be used by application 1920 and / or other components and / or layers. Library 1916 generally provides functions that make it easier for other software modules to perform tasks than directly interacting with the functions of the underlying operating system 1914 (e.g., kernel 1928, services 1930, drivers 1932, and / or modules 1960 to 1966). Library 1916 can include system library 1934 (e.g., C standard library), and this system library 1934 can provide functions such as memory allocation functions, string operation functions, mathematical functions, etc. In addition, library 1916 can include API library 1936, for example, media library (e.g., a library that supports the rendering and operation of various media formats (e.g., MPEG4, H.264, MP3, AAC, AMR, JPG, PNG)), graphics library (e.g., the OpenGL framework that can be used to render 2D and 3D graphic content on a display), database (e.g., SQLite that can provide various relational database functions), web library (e.g., WebKit that can provide web browsing functions), etc. Library 1916 can also include a variety of other libraries 1938 to provide many other APIs to application 1920 and other software components / modules.
[0171] Framework / middleware 1918 (sometimes also called middleware) can provide an advanced common infrastructure that can be used by application 1920 and / or other software components / modules. For example, framework / middleware 1918 can provide various graphical user interface (GUI) functions, advanced resource management, advanced location services, etc. Framework / middleware 1918 can provide various other APIs that can be used by application 1920 and / or other software components / modules, and some of these APIs can be specific to the operating system 1914 or platform.
[0172] Application 1920 includes built-in applications 1940 and / or third-party applications 1942. Examples of representative built-in applications 1940 can include but are not limited to contact applications, browser applications, reader applications, location applications, media applications, messaging applications, and / or game applications. Third-party applications 1942 can include any built-in applications 1940 as well as various other applications. In a specific example, third-party applications 1942 (e.g., applications developed by entities other than the vendor of a specific platform using Android TM or iOS TM software development kit (SDK)) can be applications developed on iOS TM 、Android TM 、 Mobile software running on a mobile operating system such as Phone or other mobile operating systems. In this example, the third-party application 1942 can invoke the API call 1924 provided by the mobile operating system (e.g., the operating system 1914) to facilitate the implementation of the functions described herein.
[0173] The application 1920 can create a user interface by utilizing built-in operating system functions (e.g., the kernel 1928, services 1930, drivers 1932, and / or modules 1960 to 1964), libraries (e.g., the system library 1934, API library 1936, and other libraries 1938), and the framework / middleware 1918 to interact with the system user. Alternatively or additionally, in some systems, interaction with the user can be through a presentation layer (e.g., the presentation layer 1944). In these systems, the "logic" of the application / module can be separated from aspects of the application / module that interact with the user.
[0174] Certain software architectures use virtual machines. In Figure 19 the example, the virtual machine 1948 illustrates this. The virtual machine creates a software environment in which the application / module can execute as if these application / module were executing on a hardware machine (e.g., Figure 20 the computing device 2000 in). The virtual machine 1948 is hosted by the host operating system ( Figure 19 the operating system 1914 in), and typically (but not always) has a virtual machine monitor 1946 that manages the operation of the virtual machine 1948 and its interaction with the host operating system (i.e., the operating system 1914). The software architecture 1902 executes within the virtual machine 1948 such as the operating system 1950, libraries 1952, framework / middleware 1954, applications 1956, and / or presentation layer 1958. The software architecture layers executed within the virtual machine 1948 can be the same as or different from the corresponding layers described above.
[0175] Figure 20 is a block diagram of a circuit of a device that implements an algorithm and executes a method according to some exemplary embodiments. All components are not required to be used in various embodiments. For example, client, server, and cloud-based network devices can each use different sets of components, or, in the case of a server, use a larger storage device.
[0176] An exemplary computing device in the form of a computer 2000 (also referred to as computing device 2000, computer system 2000, or computer 2000) can include a processor 2005, a memory 2010, a removable storage element 2015, a non-removable storage element 2020, an input interface 2025, an output interface 2030, and a communication interface 2035, all of which are connected via a bus 2040. Although the exemplary computing device is shown and described as a computer 2000, in different embodiments, the computing device can take different forms.
[0177] The memory 2010 can include volatile memory 2045 and non-volatile memory 2050, and can store a program 2055. The computer 2000 can include or have access to a computing environment that includes various computer-readable media such as volatile memory 2045, non-volatile memory 2050, removable storage element 2015, and non-removable storage element 2020. Computer storage includes random-access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM), and electrically erasable programmable read-only memory (EEPROM), flash memory, or other storage technologies, compact disc read-only memory (CDROM), digital versatile disc (DVD), or other optical disc storage, magnetic tape cartridges, tapes, magnetic disk storage, or other magnetic storage devices, or any other medium capable of storing computer-readable instructions.
[0178] Computer-readable instructions stored on a computer-readable medium (e.g., program 2055 stored in memory 2010) can be executed by a processor 2005 of computer 2000. Hard disk drives, CD-ROMs, and RAM are some examples of articles that include non-transitory computer-readable media (e.g., storage devices). The terms "computer-readable medium" and "storage device" do not include carrier waves because carrier waves are too transitory. "Computer-readable non-transitory medium" includes all types of computer-readable media, including magnetic storage media, optical storage media, flash media, and solid-state storage media. It should be understood that software can be installed on a computer and sold with the computer. Alternatively, software can be obtained and loaded onto the computer, including obtaining software through a physical medium or a distribution system, including, for example, obtaining software from a server owned by the software creator or from a server not owned by but used by the software creator. For example, software can be stored on a server for distribution over the Internet. As used herein, the terms "computer-readable medium" and "machine-readable medium" can be used interchangeably.
[0179] Program 2055 can utilize the modules discussed herein, e.g., workload management module 2060, which can be the same as (and perform the same functions as) the workload management module discussed in connection with Figures 1 to 19 the workload management module discussed.
[0180] Any one or more of the modules described herein can be implemented using hardware (e.g., a processor of a machine, an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or any suitable combination thereof). Additionally, any two or more of these modules can be combined into a single module, and the functionality of a single module described herein can be subdivided among multiple modules. Further, according to various exemplary embodiments, modules described herein as being implemented within a single machine, database, or device can be distributed among multiple machines, databases, or devices.
[0181] In some aspects, one or more modules included in workload management module 2060 can be integrated into a single module to perform the corresponding functions of the integrated module.
[0182] Although some embodiments have been described in detail above, other modifications are possible. For example, the logical flows depicted in the figures do not require the particular order or sequence shown to achieve the desired result. Other steps can be provided from the described flows, or steps can be reduced, and other components can be added to or removed from the described system. Other embodiments are within the scope of the appended claims.
[0183] It should also be understood that software including one or more computer-executable instructions can be installed in and sold with one or more computing devices in accordance with the present invention, and the one or more computer-executable instructions facilitate the processing and operations as described above for any one or all of the steps of the present invention. Alternatively, the software can be obtained and loaded into one or more computing devices, including obtaining the software through a physical medium or a distribution system, including, for example, obtaining the software from a server owned by the software creator or from a server not owned by but used by the software creator. For example, the software can be stored in a server for distribution via the Internet.
[0184] Furthermore, those skilled in the art should understand that the present invention is not limited in its application to the details of construction and arrangement of components set forth in the specification or illustrated in the drawings. The embodiments herein support other embodiments and support practicing or implementing in various ways. In addition, it should be understood that the wording and terminology used herein are for description and should not be regarded as restrictive. The terms "including", "comprising" or "having" and their variants used herein are intended to include the items listed thereafter and their equivalents as well as other items. Unless otherwise defined, the terms "connected", "coupled" and "installed" and their variants are used broadly and cover direct and indirect connections, couplings and installations. Additionally, the terms "connected" and "coupled" and their variants are not limited to physical or mechanical connections or couplings. Moreover, terms such as "above", "below", "bottom" and "top" are relative and are used to assist in the description but are not restrictive.
[0185] The components of the illustrative devices, systems and methods used in accordance with the illustrated embodiments can be implemented, at least in part, in digital electronic circuitry, analog electronic circuitry, or in computer hardware, firmware, software, or combinations thereof. For example, these components can be implemented as a computer program product (e.g., a computer program, program code, or computer instructions) tangibly embodied in an information carrier or a machine-readable storage device for execution by a data processing apparatus (e.g., a programmable processor, a computer, or multiple computers), or for controlling the operation of the data processing apparatus.
[0186] A computer program may be written in any form of programming language, including compiled or interpreted languages, and may be deployed in any form, including as a stand-alone program or as a module, component, subroutine, or other unit suitable for use in a computing environment. The computer program may be deployed to execute on one computer or multiple computers at one site, or may be distributed across multiple sites and interconnected by a communication network. Additionally, as will be readily appreciated by those skilled in the art, the functional programs, code, and code segments for implementing the techniques described herein are within the scope of the claims. The method steps associated with the illustrative embodiments may be performed by one or more programmable processors executing a computer program, code, or instructions to perform functions (e.g., by operating on input data and / or generating output). For example, the method steps may also be performed by special purpose logic circuitry, such as a field programmable gate array (FPGA) or an application specific integrated circuit (ASIC), and the apparatus for performing these methods may be implemented as such special purpose logic circuitry.
[0187] The various illustrative logical blocks, modules, and circuits described in connection with the embodiments disclosed herein may be implemented or performed using a general purpose processor, a digital signal processor (DSP), an ASIC, an FPGA or other programmable logic device, discrete gate or transistor logic, discrete hardware components, or any combination thereof designed to perform the functions described herein. A general purpose processor may be a microprocessor, or, alternatively, the processor may be any conventional processor, controller, microcontroller, or state machine. The processor may also be implemented as a combination of computing devices, e.g., a combination of a DSP and a microprocessor, multiple microprocessors, one or more microprocessors in conjunction with a DSP core, or any other similar configuration.
[0188] Processors suitable for executing computer programs include, for example, both general and special purpose microprocessors, as well as any one or more processors in any type of digital computer. Typically, a processor receives instructions and data from a read only memory or a random access memory or both. The essential elements of a computer are a processor for executing the instructions and one or more storage devices for storing the instructions and data. Typically, a computer also includes one or more mass storage devices for storing data (e.g., magnetic disks, magneto-optical disks, or optical disks), or is operatively coupled to one or more mass storage devices for storing data to receive data from or transfer data to these mass storage devices. Information carriers suitable for embodying computer program instructions and data include various forms of non-volatile memory, for example, including semiconductor storage devices such as electrically programmable read-only memory or ROM (electrically programmable read-only memory, EPROM), electrically erasable programmable ROM (electrically erasable programmable ROM, EEPROM), flash memory devices, or data storage disks (e.g., magnetic disks, internal hard disks, or removable disks, magneto-optical disks, CD-ROM and DVD-ROM disks). The processor and the memory may be supplemented by, or incorporated in, special purpose logic circuitry.
[0189] Those skilled in the art will appreciate that information and signals can be represented using any of a variety of different technologies. For example, the data, instructions, commands, information, signals, bits, symbols, and chips described above can be represented by voltage, current, electromagnetic waves, magnetic fields or particles, optical fields or particles, or any combination thereof.
[0190] As used herein, "machine-readable medium" (or "computer-readable medium") refers to a device that can store instructions and data temporarily or permanently, and may include, but is not limited to, random access memory (RAM), read-only memory (ROM), buffer memory, flash memory, optical media, magnetic media, cache memory, other types of storage elements (e.g., erasable programmable read-only memory (EEPROM)), and / or any suitable combination thereof. The term "machine-readable medium" should be understood to include a single medium or multiple media that can store processor instructions (e.g., a centralized or distributed database, or associated cache and server). The term "machine-readable medium" should also be understood to include any medium (or combination of media) that can store instructions executable by one or more processors 2005 such that when the instructions are executed by one or more processors 2005, cause one or more processors 2005 to perform any one or more of the methods described herein. Thus, "machine-readable medium" refers to a single storage device or apparatus, as well as a "cloud-based" storage system or storage network that includes multiple storage devices or apparatuses. As used herein, the term "machine-readable medium" does not include a signal per se.
[0191] Additionally, without departing from the scope of the present invention, the various technologies, systems, subsystems, and methods described and illustrated as discrete or separate in the various embodiments may be combined or integrated with other systems, modules, technologies, or methods. Other items shown or described as being coupled, directly coupled, or communicating with each other may be indirectly coupled or communicating electrically, mechanically, or otherwise through some interface, device, or intermediate component. Those skilled in the art can identify other examples of changes, substitutions, and alterations and make changes, substitutions, and alterations without departing from the scope of the present invention.
[0192] Although the present invention has been described with reference to specific features and embodiments, it is apparent that various modifications and combinations thereof can be made without departing from the scope of the present invention. For example, other components may be added to or removed from the described system. Accordingly, the specification and drawings are to be regarded only as illustrative of the invention as defined by the appended claims, and it is contemplated that any modifications, variations, combinations, or equivalents that fall within the scope of the present invention will be covered. Other aspects may be within the scope of the appended claims.
Claims
1. A computer-implemented method for artificial intelligence (AI)-based workload scheduling, characterized in that, The method includes: Starting the execution of a first workload on one of a plurality of Graphics Processing Units (GPUs); Determining a utilization metric of the first workload, wherein the utilization metric is associated with the execution of the first workload on the GPU; Extracting a useful feature set of the utilization metric of the first workload using a transformation function of a Deep Learning (DL) model, wherein the useful feature set includes a subset of the utilization metric; Determining a workload type of the first workload using the useful feature set; and Configuring shared execution of the first workload and a second workload on a second GPU among the plurality of GPUs based on packing the first workload and the second workload, wherein the second workload is associated with the workload type of the first workload.
2. The computer-implemented method according to claim 1, wherein The DL model includes an AI-based encoder and an AI-based decoder, and the method further includes: Performing training of the DL model by using a first training dataset as an input to the AI-based encoder and using a second training dataset as an output of the AI-based decoder.
3. The computer-implemented method according to claim 2, wherein It further includes: Configuring the first training dataset to include previous utilization metrics of a plurality of workloads executed before the execution of the first workload, wherein the plurality of workloads includes the second workload.
4. The computer-implemented method according to any one of claims 2 and 3, characterized in that, It further includes: Configuring the second training dataset as a plurality of joint completion times, wherein the plurality of joint completion times are associated with corresponding multiple joint executions, and the multiple joint executions are associated with the plurality of workloads.
5. The computer-implemented method according to claim 4, wherein One of the multiple joint executions includes at least two workloads among the plurality of workloads executed on the same GPU among the plurality of GPUs.
6. The computer-implemented method according to any one of claims 1 to 5, characterized in that, It further includes: Determining the transformation function using a subset of convolutional layers among a plurality of convolutional layers on the AI-based encoder of the DL model.
7. The computer-implemented method according to any one of claims 1 to 6, characterized in that, It further includes: Applying the transformation function to utilization metrics of a plurality of workloads to obtain: an additional useful feature set, the plurality of workloads executed before the execution of the first workload, and the plurality of workloads including the second workload.
8. The computer-implemented method according to claim 7, wherein It further includes: Determining the workload type of the first workload by using a comparison between the useful feature set and each additional useful feature set in the additional useful feature set; and And Selecting the second workload based on the comparison.
9. The computer-implemented method according to claim 8, wherein The selection of the second workload includes: Selecting the second workload when a difference between the useful feature set and an additional useful feature set in the additional useful feature set does not exceed a threshold, wherein the additional useful feature set is associated with the second workload.
10. The computer-implemented method according to claim 9, wherein It further includes: Performing the configuration of the shared execution of the first workload and the second workload when the difference between the useful feature set and the additional useful feature set is not greater than the threshold.
11. The computer-implemented method according to any one of claims 1 to 10, characterized in that, It further includes: Configuring a plurality of virtual GPUs (vGPUs) of the second GPU; And Configure the shared execution of the first workload and the second workload using the plurality of vGPUs of the second GPU.
12. The computer-implemented method according to any one of claims 1 to 11, wherein The utilization metric includes at least one of the following: A GPU usage histogram of one or more containers associated with the execution of the first workload; A memory usage histogram of a compute node associated with the execution of the first workload; and The GPU type associated with the GPU used for the execution of the first workload.
13. A system for artificial intelligence (AI)-based workload scheduling, characterized in that, The system includes: At least one processor that communicates with the memory, the at least one processor being configured to perform operations including the following when executing the instructions: Initiate the execution of a first workload on a graphics processing unit (GPU) among a plurality of GPUs; Determine a utilization metric for the first workload, where the utilization metric is associated with the execution of the first workload on the GPU; Extract a useful feature set of the utilization metric of the first workload using a transformation function of a deep learning (DL) model, where the useful feature set includes a subset of the utilization metric; Determine the workload type of the first workload using the useful feature set; and Based on packing the first workload with a second workload, configure the shared execution of the first workload and the second workload on a second GPU among the plurality of GPUs, where the second workload is associated with the workload type of the first workload.
14. The system according to claim 13, wherein The DL model includes an AI-based encoder and an AI-based decoder, and the operations further include: Perform training of the DL model using a first training dataset as an input to the AI-based encoder and a second training dataset as an output of the AI-based decoder.
15. The system according to claim 14, wherein, The operations further include: Configure the first training dataset to include previous utilization metrics of a plurality of workloads executed before the execution of the first workload, where the plurality of workloads includes the second workload.
16. The system according to any one of claims 14 and 15, characterized in that, The operations further include: Configure the second training dataset as a plurality of joint completion times, where the plurality of joint completion times are associated with corresponding plurality of joint executions, the plurality of joint executions being associated with the plurality of workloads.
17. The system according to claim 16, wherein One of the plurality of joint executions includes at least two workloads among the plurality of workloads executed on the same GPU among the plurality of GPUs.
18. The system according to any one of claims 13 to 17, characterized in that The operations further include: Determine the transformation function using a subset of convolutional layers among a plurality of convolutional layers on the AI-based encoder of the DL model.
19. The system according to any one of claims 13 to 18, characterized in that, The operations further include: Apply the transformation function to utilization metrics of a plurality of workloads to obtain: an additional useful feature set, the plurality of workloads executed before the execution of the first workload, and the plurality of workloads including the second workload.
20. The system according to claim 19, wherein The operations further include: Using a comparison of the useful feature set with each additional useful feature set in the additional useful feature sets to determine the workload type of the first workload; and Selecting the second workload based on the comparison.
21. The system according to claim 20, wherein The selection of the second workload includes: When the difference between the useful feature set and an additional useful feature set in the additional useful feature sets does not exceed a threshold, selecting the second workload, where the additional useful feature set is associated with the second workload.
22. The system according to claim 21, wherein The operation further includes: When the difference between the useful feature set and the additional useful feature set is not greater than the threshold, performing the configuration of the shared execution of the first workload and the second workload.
23. The system according to any one of claims 13 to 22, characterized in that, The operation further includes: Configuring a plurality of virtual GPUs (vGPUs) of the second GPU; and Using the plurality of vGPUs of the second GPU to configure the shared execution of the first workload and the second workload.
24. The system according to any one of claims 13 to 23, characterized in that, The utilization metric includes at least one of the following: A GPU usage histogram of one or more containers associated with the execution of the first workload; A memory usage histogram of a compute node associated with the execution of the first workload; and The GPU type associated with the GPU used for the execution of the first workload.
25. A non-transitory computer-readable medium stores computer instructions for artificial intelligence (AI)-based workload scheduling, characterized in that, When the instructions are executed by one or more processors, causing the one or more processors to perform operations including the following: Initiating the execution of a first workload on one GPU among a plurality of graphics processing units (GPUs); Determining a utilization metric of the first workload, where the utilization metric is associated with the execution of the first workload on the GPU; Extracting a useful feature set of the utilization metric of the first workload using a transformation function of a deep learning (DL) model, where the useful feature set includes a subset of the utilization metric; Using the useful feature set to determine the workload type of the first workload; and Based on packing the first workload and a second workload, configuring the shared execution of the first workload and the second workload on a second GPU among the plurality of GPUs, where the second workload is associated with the workload type of the first workload.
26. The non-transitory computer-readable medium according to claim 25, wherein The DL model includes an AI-based encoder and an AI-based decoder, and the operation further includes: Performing the training of the DL model by using a first training dataset as the input of the AI-based encoder and using a second training dataset as the output of the AI-based decoder.
27. The non-transitory computer-readable medium according to claim 26, wherein The operation further includes: Configuring the first training dataset to include previous utilization metrics of a plurality of workloads executed before the execution of the first workload, where the plurality of workloads includes the second workload.
28. The non-transitory computer-readable medium according to any one of claims 26 and 27, characterized in that, The operation further includes: Configure the second training data set as a plurality of joint completion times, wherein the plurality of joint completion times are associated with a corresponding plurality of joint executions, and the plurality of joint executions are associated with the plurality of workloads.
29. The non-transitory computer-readable medium according to claim 28, wherein, One of the plurality of joint executions includes at least two of the plurality of workloads executed on the same GPU among the plurality of GPUs.
30. The non-transitory computer-readable medium according to any one of claims 25 to 29, characterized in that, The operations further include: Determine the transformation function using a subset of convolutional layers among a plurality of convolutional layers on an AI-based encoder of the DL model.
31. The non-transitory computer-readable medium according to any one of claims 25 to 30, characterized in that, The operations further include: Apply the transformation function to utilization metrics of a plurality of workloads to obtain: an additional useful feature set, the plurality of workloads executed before the execution of the first workload, and the plurality of workloads including the second workload.
32. The non-transitory computer-readable medium according to claim 31, wherein The operations further include: Use a comparison of the useful feature set with each additional useful feature set in the additional useful feature set to determine the workload type of the first workload; and Select the second workload based on the comparison.
33. The non-transitory computer-readable medium according to claim 32, wherein, The operations for the selection of the second workload include: When the difference between the useful feature set and an additional useful feature set in the additional useful feature set does not exceed a threshold, select the second workload, wherein the additional useful feature set is associated with the second workload.
34. The non-transitory computer-readable medium according to claim 33, wherein The operations further include: When the difference between the useful feature set and the additional useful feature set is not greater than the threshold, perform the configuration of the shared execution of the first workload and the second workload.
35. The non-transitory computer-readable medium according to any one of claims 25 to 34, wherein The operations further include: Configure a plurality of virtual GPUs (vGPUs) of the second GPU; and Configure the shared execution of the first workload and the second workload using the plurality of vGPUs of the second GPU.
36. The non-transitory computer-readable medium according to any one of claims 25 to 35, wherein The utilization metrics include at least one of the following: A GPU usage histogram of one or more containers associated with the execution of the first workload; A memory usage histogram of a computing node associated with the execution of the first workload; and The GPU type associated with the GPU used for the execution of the first workload.
37. An apparatus for artificial intelligence (AI)-based workload scheduling, characterized in that, The apparatus includes: A module for starting the execution of a first workload on one GPU among a plurality of graphics processing units (GPUs); A module for determining utilization metrics of the first workload, wherein the utilization metrics are associated with the execution of the first workload on the GPU; A module for extracting a useful feature set of the utilization metrics of the first workload using a transformation function of a deep learning (DL) model, wherein the useful feature set includes a subset of the utilization metrics; A module for determining the workload type of the first workload using the useful feature set; and A module for configuring shared execution of the first workload and the second workload on a second GPU among the multiple GPUs based on packing the first workload and the second workload, wherein the second workload is associated with the workload type of the first workload.