Multilevel Scheduling for Improved Quality of Service
A scheduler module in GPUs enforces job limits with slack time to manage resource allocation, addressing disproportionate consumption by ill-behaved virtual functions and enhancing quality of service for all virtual machines.
Patent Information
- Application Number
- JP2025534228
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2022-12-21
- Filing Date
- 2023-12-20
- Publication Date
- 2025-12-25
AI Technical Summary
Virtual functions in processing systems, such as GPUs, often consume disproportionate shares of hardware resources due to inconsistent job submission cadences and sizes, affecting the quality of service for other virtual machines.
Implementing a scheduler module that enforces job limits by allocating time partitions with slack time to accommodate variance in job submission cadence and size, prioritizing well-behaved virtual functions over ill-behaved ones to ensure fair resource allocation.
This approach mitigates the impact of ill-behaved virtual functions on well-behaved ones, improving overall quality of service by preventing excessive resource consumption and ensuring fair access to hardware resources.
Smart Images

Figure 2025542143000001_ABST
Abstract
Description
[Background technology]
[0001] Conventional processing units, such as graphics processing units (GPUs), support virtualization, allowing multiple virtual machines (VMs) to use the GPU's hardware resources. Some VMs implement operating systems that allow the VM to emulate a physical machine. Other VMs are designed to execute code in a platform-independent environment. A hypervisor creates and runs VMs, which are also called guest machines or guests. The virtual environment implemented on the GPU provides virtual functions to other virtual components implemented on the physical machine. A single physical function implemented on the GPU is used to support one or more virtual functions (VFs). The physical function allocates the virtual functions to different VMs on the physical machine based on time slices or time partitions. For example, the physical function allocates a first virtual function to a first VM in a first time interval and a second virtual function to a second VM in a second subsequent time interval. The single root input / output virtualization (SR-IOV) specification allows multiple VMs to share a GPU interface to a single bus, such as a peripheral component interconnect express (PCIe) bus. Components access virtual functions by sending requests over the bus. Summary of the Invention [Means for solving the problem]
[0002] To prevent the virtual functions from consuming a disproportionate share of the hardware resources, the processing system enforces job limits on the virtual functions to facilitate an expected quality of service for each of the virtual functions assigned to virtual machines executing on the processing system. In a first embodiment, a method includes allocating a plurality of time partitions within a scheduling period to a plurality of virtual functions for execution of jobs on a parallel processor, the scheduling period including slack time to allow variance in at least one of a cadence at which each virtual function submits jobs for execution and a size of the jobs submitted for execution. The method further includes preventing execution of a job that exceeds the assigned time partition for any virtual function of the plurality of virtual functions.
[0003] In some embodiments, preventing the execution includes scheduling the first virtual function to execute the first plurality of jobs after the second virtual function in response to the submission of the first plurality of jobs exceeding an expected cadence and the submission of the second plurality of jobs by the second virtual function not exceeding the expected cadence.
[0004] The method may further include scheduling the first virtual function to execute the first plurality of jobs after the second virtual function in response to the first plurality of jobs exceeding an expected job size and the second plurality of jobs submitted by the second virtual function not exceeding an expected job size.
[0005] In some embodiments, the method includes allocating a job size credit to the first virtual function when a size of a job submitted by the first virtual function is smaller than an expected job size. The job size credit may be used by the first virtual function in a scheduling period immediately subsequent to the scheduling period in which the job size credit is allocated.
[0006] The method may further include maintaining a first level list and a second level list, the first level list including a first virtual function that submits a first plurality of jobs that are within an expected cadence and take less time to execute than an expected job size, and the second level list including a second virtual function not included in the first level list. The method also includes scheduling the first plurality of jobs for the virtual functions in the first level list before scheduling the second plurality of jobs for the virtual functions in the second level list.
[0007] In some embodiments, the method further includes bypassing scheduling a second plurality of jobs for a virtual function in the second level list within a current scheduling period. In some embodiments, the scheduling period is a first scheduling period of the multiple scheduling periods, and the method further includes scheduling the first virtual function after the second virtual function if a first number of the first plurality of jobs submitted by the first virtual function within the multiple scheduling periods is greater than a second number of the second plurality of jobs submitted by the second virtual function within the multiple scheduling periods.
[0008] The method may also include allowing a virtual function that has remaining time within its assigned time partition and has not completed a job within the scheduling period to submit a job when the parallel processor is idle.
[0009] In another embodiment, a processing system includes a parallel processor and a scheduler module circuit configured to allocate a plurality of time partitions within a scheduling period to a plurality of virtual functions for execution of jobs on the parallel processor, the scheduling period including slack time to allow for variance in at least one of a cadence at which each virtual function submits jobs for execution and a size of the jobs submitted for execution, and the scheduler module circuit is configured to prevent execution of the job beyond the allocated time partition for a first virtual function of the plurality of virtual functions.
[0010] The scheduler module circuitry may also be configured to schedule the first virtual function after the second virtual function in response to the submission of the first plurality of jobs by the first virtual function exceeding an expected cadence and the submission of the second plurality of jobs by the second virtual function not exceeding an expected cadence. In some embodiments, the scheduler module circuitry is further configured to schedule the first virtual function after the second virtual function in response to the first plurality of jobs submitted by the first virtual function exceeding an expected job size and the second plurality of jobs submitted by the second virtual function not exceeding an expected job size.
[0011] The scheduler module circuitry can also be configured to allocate a job size credit to a virtual function when the size of a job submitted by the virtual function is smaller than the expected job size. The job size credit can be used by the virtual function in a scheduling period immediately subsequent to the scheduling period in which the job size credit is allocated.
[0012] In some embodiments, the scheduler module circuitry is further configured to maintain a first level list and a second level list, the first level list including a first virtual function that submits a first plurality of jobs that are within an expected cadence and take less time to execute than an expected job size, and the second level list including a second virtual function that is not included in the first level list. The scheduler module circuitry may also be configured to schedule the first plurality of jobs for the virtual functions in the first level list before scheduling the second plurality of jobs for the virtual functions in the second level list.
[0013] The scheduler module circuitry may be further configured to bypass scheduling the second plurality of jobs for the second virtual function within the current scheduling period. In some embodiments, the scheduling period is a first scheduling period of the multiple scheduling periods, and the scheduler module circuitry is further configured to schedule the first virtual function after the second virtual function if a first number of the first plurality of jobs submitted by the first virtual function within the multiple scheduling periods is greater than a second number of the second plurality of jobs submitted by the second virtual function within the multiple scheduling periods.
[0014] In some embodiments, the scheduler module circuitry is further configured to allow a virtual function that has remaining time within its assigned time partition and has not completed a job within the scheduling period to submit a job when the parallel processor is idle.
[0015] In another embodiment, a server includes parallel processors configured to execute jobs submitted by a plurality of virtual functions and a scheduler module circuit configured to assign a time partition to each of the plurality of virtual functions within a scheduling period based on expected job size and cadence and slack time to allow for distribution of submitted job sizes and cadences, and to prevent execution on the parallel processors of a job that exceeds the time partition assigned to any virtual function of the plurality of virtual functions.
[0016] Additionally, the scheduler module circuitry may be further configured to schedule the first virtual function before the second virtual function in response to the first plurality of jobs submitted by the first virtual function having a frequency below an expected cadence and the second plurality of jobs submitted by the second virtual function exceeding the expected cadence.
[0017] In some embodiments, the scheduler module circuitry is further configured to schedule the first virtual function before the second virtual function in response to the first plurality of jobs submitted by the first virtual function not exceeding an expected job size and the second plurality of jobs submitted by the second virtual function exceeding an expected job size. The scheduler module circuitry may be further configured to maintain a first level list and a second level list. The first level list includes the first virtual function submitting the first plurality of jobs that are within an expected cadence and take less time to execute than the expected job size, and the second level list includes the second virtual function not included in the first level list. The scheduler module circuitry is further configured to schedule the first plurality of jobs for the first virtual function in the first level list before scheduling the second plurality of jobs for the second virtual function in the second level list.
[0018] The present disclosure may be better understood, and its numerous features and advantages made apparent to those skilled in the art by referencing the accompanying drawings, in which: The use of the same reference numbers in different drawings indicates similar or identical items. [Brief explanation of the drawings]
[0019] [Figure 1] FIG. 1 is an exemplary block diagram of a processing system configured to enforce job limits for virtual functions, according to some embodiments. [Figure 2] FIG. 2 is an exemplary block diagram of a mapping of virtual functions to virtual machines implemented in a processing unit, according to some embodiments. [Figure 3] FIG. 1 illustrates an exemplary block diagram of a scheduler module according to some embodiments. [Figure 4] FIG. 1 is an exemplary block diagram of time partitioning to support fair access to virtual machines associated with virtual functions in a processing unit, according to some embodiments. [Figure 5] FIG. 10 is an exemplary diagram illustrating the cadence and job size of various jobs submitted for execution by multiple virtual functions, according to some embodiments. [Figure 6] FIG. 1 is a flow diagram of an exemplary method for scheduling first and second virtual functions, according to some embodiments. DETAILED DESCRIPTION OF THE INVENTION
[0020] The hardware resources of a GPU are partitioned according to SR-IOV using a physical function (PF) and one or more virtual functions (VF). Each virtual function is associated with a single physical function. In a native (host OS) environment, the physical functions are used by native user-mode and kernel-mode drivers, and all virtual functions are disabled. All GPU registers are assigned to the physical functions via trusted access. In a virtual environment, the physical functions are used by the hypervisor (host VM), and the GPU exposes a specific number of virtual functions according to the PCIe SR-IOV standard, such as one virtual function per guest VM. Each virtual function is assigned to a guest VM by the hypervisor.
[0021] Typically, the central processing unit (CPU) is partitioned across virtual functions so that each virtual function has a dedicated CPU. The CPU prepares jobs for the virtual function and submits them to the GPU. Each virtual function can receive remote user input, prepare job submissions based on the remote user input, and also submit jobs orthogonal to the user input. The CPU can submit jobs to the GPU for a virtual function at any time. However, job execution on the GPU occurs during the time partition assigned to the virtual function. In many cases, when the job submitted by a virtual function is for streaming video, the virtual function submits jobs that run at a regular cadence to achieve a target frames-per-second (FPS) rate. Consistent submission of single or multiple units of work that collectively correspond to a similar-sized "job" results in a regular cadence for job execution.
[0022] However, because jobs are submitted based on CPU ready timing, jobs may not align with the GPU time partitions assigned to virtual functions. Furthermore, jobs submitted by multiple virtual functions may vary in size such that each job takes longer to complete execution than its assigned time partition. Furthermore, in some instances, virtual functions may behave in a greedy or malicious manner by submitting an excessive number of jobs (or jobs that collectively take longer to execute) within their assigned time partitions, thereby not submitting jobs within the expected cadence.
[0023] When a virtual function behaves in a greedy or malicious manner, it consumes a disproportionate share of hardware resources, which adversely affects the quality of service of other virtual functions assigned to the VM. The impact on a virtual function, in some embodiments, is based on both throughput and latency. Throughput refers to the job execution rate, which is possibly related to one or more of the coding resolution and frame rate of desktop or game frames, the video decoding rate, or the rendering rate experienced by each virtual function. Latency refers to the time between submission of a job to a GPU until the job completes execution on the GPU so that the job results can be consumed within the expected latency. For example, in the context of video encoding, latency refers to the time it takes for a job to complete so that the encoded frames can be streamed.
[0024] 1-5 disclose embodiments of a processing unit, such as a graphics processing unit (GPU) of a processing system or server, configured to enforce job limits on virtual functions to facilitate expected quality of service for each of the virtual functions assigned to VMs executing on the processing unit. A scheduler module circuit defines a scheduling period as the sum of a time partition for each virtual function and an additional period, referred to herein as “slack time.” The slack time allows for occasional variance in the cadence at which each virtual function submits jobs for execution by the GPU. For example, if a frame takes longer than expected to render, the virtual function may miss an opportunity to submit the frame during its assigned time slice and submit the frame along with subsequent frames during the next assigned time slice, resulting in a first time slice in which the virtual function does not submit a job and a second time slice in which the virtual function submits a batch of two jobs. Such occasional variance in job submission cadence is expected and does not indicate greedy or malicious behavior on the part of the virtual function. The GPU defines the time partition for each virtual function as the expected job size of the virtual function multiplied by N, where N is a number greater than or equal to 1 and depends on the GPU's tolerance for variance in job submission behavior.
[0025] The GPU monitors jobs submitted for execution by the virtual functions within a scheduling period to determine whether the jobs are being submitted within an expected cadence or frequency. The GPU further monitors whether submitted jobs take longer to execute than the expected job size. Individual virtual functions are designated as either well-behaved or ill-behaved depending on whether the individual virtual functions submit jobs at the expected cadence and depending on whether the submitted jobs take longer to execute than the expected job size. The GPU scheduler schedules well-behaved virtual functions before ill-behaved virtual functions to prevent the ill-behaved virtual functions from consuming a disproportionate share of hardware resources, thereby mitigating the impact of the ill-behaved virtual functions on the quality of service of the well-behaved virtual functions.
[0026] FIG. 1 is a block diagram of a processing system 100 configured to implement job limits for virtual functions in accordance with some embodiments. The techniques described herein may be utilized, in various embodiments, in any of a variety of parallel processors, such as vector processors, graphics processing units (GPUs), general-purpose GPUs (GPGPUs), non-scalar processors, highly parallel processors, artificial intelligence (AI) processors, inference engines, machine learning processors, and other multi-threaded processing units. FIG. 1 illustrates an example of a parallel processor, and in particular, a graphics processing unit (GPU) 115 (e.g., a virtual GPU), in accordance with some embodiments. However, references herein to a GPU will be understood to include any of a variety of parallel processors, unless otherwise specified.
[0027] Processing system 100 includes or has access to memory 105 or other storage components implemented using a non-transitory computer-readable storage medium, such as dynamic random-access memory (DRAM). However, memory 105 may also be implemented using other types of memory, including static random access memory (SRAM), non-volatile RAM, etc. Processing system 100 also includes a bus 110 for supporting communication between entities implemented in processing system 100, such as memory 105. In the illustrated embodiment, bus 110 is configured as a PCIe bus. Some embodiments of processing system 100 include other buses, bridges, switches, routers, etc., which are not shown in FIG. 1 for clarity.
[0028] Processing system 100 also includes a central processing unit (CPU) 150 connected to bus 110 and in communication with GPU 115 and memory 105 via bus 110. In the illustrated embodiment, CPU 150 implements multiple processing elements (also referred to as processor cores) 155 configured to execute instructions simultaneously or in parallel. CPU 150 executes instructions, such as program code 160, stored in memory 105, and CPU 150 stores information, such as results of executed instructions, in memory 105. CPU 150 initiates graphics operations by issuing draw calls to GPU 115.
[0029] Input / Output (I / O) engine 165 handles input or output operations associated with display 120 and other elements of processing system 100, such as a keyboard, mouse, printer, external disk, network, etc. I / O engine 165 is coupled to bus 110 such that I / O engine 165 communicates with memory 105, GPU 115, or CPU 150. In the illustrated embodiment, I / O engine 165 is configured to read information stored on external storage component 170, which is implemented using a non-transitory computer-readable storage medium such as a flash drive. I / O engine 165 can also write information, such as results of processing by GPU 115 or CPU 150, to external storage component 170. Display 120 can be remotely connected to the VM via a network connection using an appropriate protocol.
[0030] Processing system 100 includes one or more graphics processing units (GPUs) 115 configured to render images for presentation on display 120. For example, GPU 115 may render objects to generate pixel values that are provided to display 120, which uses the pixel values to display images representing the rendered objects. GPU 115 includes a GPU core 125 that is comprised of a set of computational units, a set of fixed function units, or a combination thereof for simultaneously or in parallel executing instructions. GPU core 125 may include tens, hundreds, or thousands of computational units or fixed function units for executing instructions.
[0031] GPU 115 includes internal (or on-chip) memory 130, which may include a local data store and caches, registers, or buffers utilized by shader engine 125. Internal memory 130 stores data structures that describe tasks to be performed by one or more of the compute units or fixed function units within GPU core 125. The compute units or fixed function units within GPU core 125 may also access information in (external) memory 105. In the illustrated embodiment, GPU 115 communicates with memory 105 via bus 110. However, some embodiments of GPU 115 communicate with memory 105 via a direct connection or through other buses, bridges, switches, routers, etc. GPU 115 executes instructions stored in memory 105, and GPU 115 stores information, such as results of executed instructions, in memory 105. For example, memory 105 may store copies 135 of instructions from program code executed by GPU 115, such as program code representing shaders, virtual functions, or other code executed by one or more compute units or fixed function units implemented in GPU core 125.
[0032] GPU 115 includes encoder 140, which is used to encode information for transmission over bus 110. Encoder 140 also provides security features to support secure communication over bus 110. In some embodiments, encoder 140 encodes pixel values for transmission to display 120, which implements a decoder to decode the pixel values and reconstruct an image for presentation. Display 120 can be remotely connected to a VM via a network connection. Some embodiments of encoder 140 encode and encrypt information generated by virtual functions implemented on GPU 115 for communication over bus 110.
[0033] Some embodiments of GPU 115 act as a physical function supporting one or more virtual functions shared over bus 110. For example, GPU 115 may use a dedicated portion of bus 110 to securely share several VMs using the SR-IOV standard defined for PCIe buses. GPU 115 includes bus interface 145 that provides an interface between GPU 115 and bus 110, for example, according to the SR-IOV standard. Bus interface 145 provides functionality including doorbell detection, register redirection, frame buffer apertures, doorbell write redirection, and other functionality as described below.
[0034] As described in more detail below, at least one of the virtual functions supported and enabled by GPU 115 may each submit a job for execution that exceeds the expected use of the virtual functions. As an example, a job submitted for execution by a virtual function is a video encoding job. In some embodiments, the video encoding job encodes video information for a gaming application being executed by a VM, such as a gaming application executed by a cloud gaming platform. The game may be expected to run at a maximum frames per second (fps) and a maximum resolution. To prevent a job submitted for execution by at least one of the virtual functions from exceeding the expected use of the virtual functions, GPU 115 includes a scheduler that prevents jobs for virtual functions that have exceeded their expected use or are expected to exceed their expected use from executing before executing jobs for virtual functions that have not exceeded their expected use or are expected not to exceed their expected use. In some embodiments, the virtual functions may each further include a scheduler that restricts job submission.
[0035] FIG. 2 is a block diagram of a mapping 200 of virtual functions to VMs implemented on a processing unit, according to some embodiments. Mapping 200 represents a mapping of virtual functions to VMs implemented in some embodiments of GPU 115 shown in FIG. 1. Host machine 201 includes a physical function 205, such as GPU 115, that is partitioned into virtual functions 210, 211, 212, and 213 during initialization of physical function 205. In some embodiments, each of virtual functions 210-213 includes an application (e.g., a video game), an application programming interface (API), a user mode driver (UMD), and a kernel mode driver (KMD). Host machine 201 implements a host operating system or hypervisor 215 for physical function 205. Hypervisor 215 launches one or more VMs 220, 221, 222, and 223 to execute on physical resources, such as GPU 115, that support physical function 205. In some embodiments, VMs 220-223 include a GPU virtualization driver (GPUV) that can receive configuration information (e.g., from a server administrator) and pass the configuration information to a virtual GPU component and virtual video core or virtual engine (e.g., video core next (VCN)) assigned to each of virtual functions 210-213.
[0036] VMs 220-223 are assigned to corresponding virtual functions 210-213. In the illustrated embodiment, VM 220 is assigned to virtual function 210, VM 221 is assigned to virtual function 211, VM 222 is assigned to virtual function 212, and VM 223 is assigned to virtual function 213. Virtual functions 210-213 submit jobs to GPU 115, which provides GPU functionality to corresponding VMs 220-223. Thus, the virtualized GPU 115 is shared across many VMs 220-223. In some embodiments, time slicing, also known as time partitioning, and context switching are used to provide fair access to the GPU 115 by the virtual functions 210-213, such that each of the virtual functions 210-213 is assigned a respective time partition for execution of multiple jobs by the GPU 115.
[0037] VMs 220-223 each further include a scheduler module circuit 230 that manages virtual function access to GPU 115. In some embodiments, GPU 115 includes scheduler module circuit 230. In some embodiments, scheduler module circuit 230 is an SR-IOV multimedia GPU scheduler, a hardware and / or firmware scheduler. In some embodiments, scheduler module circuit 230 is implemented in various forms, such as a processor, a field programmable gate array (FPGA), or other forms of circuitry. This GPU-side enforcement by scheduler module circuit 230 for misbehaving virtual functions is difficult to address and provides security for such scheduling. Scheduler module circuit 230 defines a period, or scheduling period, during which jobs submitted by VMs 220-223 may be executed by GPU 115, respectively.
[0038] In some embodiments, the scheduler module circuitry 230 assigns a time slice to each of the virtual functions 210-213 that is tolerant to variance in submission behavior. In such embodiments, the scheduler module circuitry 230 can assign a time partition to each of the virtual functions 210-213 to be equal to the expected job size for the particular virtual function multiplied by n, where n is 1, 2, etc., depending on the desired tolerance to variance in the submission of jobs to the multiple virtual functions 210-213. For example, four virtual functions 210-213 may submit jobs for execution by the GPU 115 having 1080p resolution at 60 fps, with each job expected to take less than 3 ms. If n=2, the time partitioning assigned to each of the virtual functions 210-213 is equal to 6 ms; if n=2, then 2 jobs * 3 ms / job = 6 ms. In some embodiments, scheduler module circuitry 230 supports configurable tolerances for various metrics (e.g., expected job size + / - tolerance). In some embodiments, at least one of virtual functions 210-213 supports multiple streams (e.g., a single virtual function submits one 1080p 60fps stream + one 720p 30fps stream for execution).
[0039] Scheduler module circuit 230 includes a job cadence monitor circuit, referred to as job cadence monitor 310, and / or a job size monitor circuit, referred to as job size monitor 320, shown in FIG. 3. Job cadence monitor 310 monitors jobs submitted by virtual functions 210-213 within a scheduling period to determine whether jobs are being executed by GPU 115 at an expected cadence with time periods between jobs submitted by a particular virtual function that are approximately equal for each job by that virtual function. In some embodiments, because the time between job submissions is subject to variation due to submission jitter, job cadence monitor 310 monitors whether the expected number of jobs are received within each scheduling period. Job size monitor 320 further monitors the size of jobs submitted by virtual functions 210-213 for execution by GPU 115 within each scheduling period. Based on monitoring by job cadence monitor 310 and / or job size monitor 320, scheduler module circuitry 230 determines which of virtual functions 210-213 are over-utilizing the bandwidth and time partition of an assigned virtual engine, such as GPU 115, and which of virtual functions 210-213 are not "ill-behaved" but are instead "well-behaved" in that they are not "more than they pay for" or are not over-utilizing the bandwidth and time partition of their assigned virtual engine. Scheduler module circuitry 230 schedules virtual functions 210-213 based at least in part on the determination of whether the virtual functions are ill-behaved or well-behaved.
[0040] FIG. 4 is a block diagram of time partitioning 400 supporting fair access to virtual machines associated with virtual functions within GPU 115, according to some embodiments. Time partitioning 400 is implemented in some embodiments of GPU 115 shown in FIG. 1. Time partitioning 400 is used to provide fair access to some embodiments of virtual machines 210-213 shown in FIG. 3. In FIG. 3, time increases from left to right. A first time partition 405 is assigned to a first virtual function, such as virtual machine 220 assigned to virtual function 210. State information for virtual function 210 is stored by a bus interface, such as bus interface 145 shown in FIG. 1. Upon completion of first time partition 405, the processing unit performs a context switch 406, which includes saving the current context and state information of the first virtual function to memory. Context switch 406 also includes retrieving the context and state information of a second virtual function from memory and loading the information into memory or registers within the processing unit. The second time partition 407 is assigned to the second virtual function, so that the second virtual function has full access to the processing unit's resources for the duration of the second time partition 407. The scheduler module circuit 230 defines a scheduling period as the sum of these per-virtual function time partitions plus an additional period, which is the period during which an expected job size executes for virtual function time N, where N is a number greater than or equal to 1 and dependent on the GPU's tolerance for variance in job submission behavior.
[0041] 5 is a diagram illustrating the cadence and job size of various jobs submitted by multiple virtual functions for execution by GPU 115. Multiple time partitions are assigned to virtual functions 210-213. These time partitions are assigned to virtual function 210 for submitting jobs 511, 521, 531, and 541; to virtual function 211 for submitting jobs 512, 522, 532, and 542; to virtual function 212 for submitting jobs 513, 514, 523, 524, 533, 534, 543, and 544; and to virtual function 213 for submitting jobs 515, 525, 535, and 545. Job cadence monitor 310 determines that multiple jobs 511, 521, 531, and 541 submitted for execution by virtual function 210 are within the expected cadence, and multiple jobs 512, 522, 532, and 542 submitted for execution by virtual function 211 are within the expected cadence. As shown, jobs 511, 521, 531, 541, 512, 522, 532, and 542 submitted by virtual functions 210 and 211, respectively, start at approximately the same time within each of scheduling periods 510-540. Thus, virtual functions 210 and 211 each submit a single job at the expected cadence within each of the first, second, third, and fourth scheduling periods 510-540. As used herein, jobs "submitted" by virtual functions 210-213 include not only jobs that are actually executed by GPU 115, but also jobs that attempt to use a disproportionate share of the bandwidth available by GPU 115. In some embodiments, virtual functions 210-213 submit jobs at substantially the same expected cadence. In other embodiments, virtual functions 210-213 submit jobs at different cadences (e.g., some virtual functions submit jobs at 30 fps and some virtual functions submit jobs at 60 fps).
[0042] In contrast to the jobs submitted by virtual functions 210 and 211, virtual function 212 is shown as submitting jobs more frequently than virtual functions 210 and 211. Virtual function 212 is shown as submitting jobs 513, 514, 523, 524, 533, 534, 543, and 544. Thus, virtual function 212 submits eight jobs for execution within four scheduling periods 510, 520, 530, and 540. At four expected cadences, virtual function 212 submits more than its fair share of jobs for execution. Job cadence monitor 310 determines that virtual function 212 is submitting more jobs for execution than the expected number of jobs for execution of virtual function 212 and identifies virtual function 212 as misbehaving. Also, in contrast to the jobs submitted by virtual functions 210 and 211, virtual function 213 is shown submitting jobs that are larger than the expected job size. Virtual function 213 submits four jobs 515, 525, 535, and 545 within four scheduling periods 510, 520, 530, and 540, but the size of the jobs submitted by virtual function 213 is larger than the jobs submitted by virtual functions 210 and 211 and is larger than the expected job size. Job size monitor 320 determines that virtual function 213 is submitting jobs for execution that are larger than the expected job size and identifies virtual function 213 as misbehaving.
[0043] In some embodiments, the jobs submitted by the virtual functions 210-213 are video encoding jobs, such as video games. Each of the virtual machines 220-223 is capable of independently executing the video games. Based on the configuration of each of the video games, each of the video games is expected to run at 1080p resolution with balanced encoding preset to 60 frames per second (fps). If all of the jobs submitted by the virtual functions 210-213 run at 1080p resolution at 60 fps, the scheduler module circuit 230 identifies all of the virtual functions 210-213 as well-behaved virtual functions. However, in some cases, such as when a malicious virtual function utilizes open source code (e.g., OpenGL) being executed by the virtual machine, the virtual functions consistently exceed the expected job submission cadence and / or job size. As an example, jobs may be submitted more frequently than the expected fps for execution by the virtual functions 210-213 (e.g., 120 fps vs. the expected 60 fps). The job cadence monitor 310 monitors the cadence of jobs submitted at a rate of 120 fps and identifies virtual functions 210-213 submitting jobs as misbehaving. As another example, a virtual function that consistently submits jobs larger than the expected video resolution of 1080p, such as 4k resolution, is attempting to use an unfair share of the physical resources of the GPU 115. The job size monitor 320 identifies such virtual functions 210-213 as misbehaving. A virtual function determined to be misbehaving adversely affects the execution of jobs submitted by virtual functions determined to be well-behaving without intervention by the GPU 115.
[0044] When the scheduler module circuit 230 determines whether any virtual functions are misbehaving and any virtual functions are well-behaving, the scheduler module circuit 230 implements multi-level scheduling to minimize the impact of the misbehaving virtual functions on the well-behaving virtual functions, thereby improving the quality of service (QoS) of the well-behaving virtual functions. In some embodiments, the scheduler module circuit 230 maintains two lists: a first list for well-behaving virtual functions and a second list for poorly behaving virtual functions. The scheduler module circuit 230 schedules well-behaving virtual functions from the first list when the virtual engine is idle. If there are no pending jobs for the well-behaving virtual functions, the scheduler module circuit 230 schedules the poorly behaving virtual functions. This allows for configurable tolerance toward poorly behaving virtual functions to accommodate more graceful handling of exception situations in which at least one well-behaving virtual function may behave poorly due to an exception. In some embodiments, scheduler module circuitry 230 maintains more than two lists to categorize virtual functions more finely than well-behaved and poorly-behaved, e.g., in some embodiments, scheduler module circuitry 230 maintains lists of very well-behaved virtual functions, very poorly-behaved virtual functions, occasionally poorly-behaved virtual functions, occasionally well-behaved virtual functions, etc.
[0045] In some embodiments, if a particular virtual function is on the misbehaving list, the particular virtual function is not scheduled at all in a period as a penalty. For example, if a virtual function overused its share in a past scheduling period, the virtual function must wait until the overusage is deducted from a subsequent scheduling period and is eventually granted a time share again in a future period in which the virtual function can be rescheduled. In some embodiments, the classification of a virtual function as misbehaving resets after the virtual function is penalized to provide tolerance for exceptional circumstances or changes in use cases at runtime. The timing and conditions for the classification reset are configurable. If the virtual function continues to be classified as misbehaving for a longer period, the virtual function may eventually be prohibited from submitting jobs. In such a case, scheduler module circuit 230 bypasses submitting jobs for the misbehaving virtual function during the current scheduling period.
[0046] A scheduling period is the sum of the per-virtual-function time partitions assigned to each virtual function. As shown in FIG. 5, a first scheduling period 510 is shown, during which further individual jobs submitted by virtual machines 220-223 are expected to be executed. First scheduling period 510 is followed by scheduling period 520, which is followed by scheduling period 530, which is followed by scheduling period 540. While four scheduling periods are shown, this is for ease of explanation, and it should be understood that the defined scheduling periods continue to repeat as virtual machines 220-223 continue to submit further jobs for execution.
[0047] In some embodiments, the scheduler module circuit 230 adds additional time, or slack time, to the scheduling period. The purpose of the slack time in the period is to accommodate expected variations in submission behavior due to different virtual functions that are unavoidable in the use case. The slack time extends the size of the scheduling period so that jobs that run near the end of the scheduling period have time to execute past the scheduling period. The scheduler module circuit 230 thereby prevents jobs that do not finish execution by the end of the scheduling period from being inappropriately classified as misbehaving, thereby providing flexibility to the scheduling period. As shown in FIG. 5 , the first scheduling period 510 includes additional slack time such that additional time 519 remains after job 511 ends and before the end of the first scheduling period 510. Because job 511 does not always start at the same time and thereby finish within the scheduling period, the additional time 519 provides flexibility to refrain from classifying job 511 as misbehaving in such instances. For example, if job 511 is submitted near the end of a scheduling period, a portion of job 511 may execute during the scheduling period, and the remainder of job 511 may execute during the next scheduling period. In such a case, job 511 uses only a portion of its allotted time in the scheduling period and has excess time that can be carried forward to the next (i.e., subsequent) scheduling period. In some embodiments, excess time is not carried forward beyond the next scheduling period to prevent virtual functions from accumulating a large excess.
[0048] Similarly, the second scheduling period 520 includes slack time such that after job 541 finishes, additional time 529 remains before the end of the second scheduling period 520; the third scheduling period 530 includes slack time such that after job 521 finishes, additional time 539 remains before the end of the third scheduling period 530; and the fourth scheduling period 540 includes slack time such that after job 531 finishes, additional time 549 remains before the end of the fourth scheduling period 540.
[0049] As an example, the scheduler module circuit 230 defines the slack time as 33.3 ms (two periods of 60 fps, i.e., 2*16.67 ms) - 24 ms (the expected usage portion of 33.3 ms, which is 4 VFs*3 ms*2 jobs / VF within 33.3 ms). In some cases, a virtual function may not submit a job on time within 16.67 ms (e.g., job preparation is delayed) and instead submit two jobs in the next 16.67 ms (e.g., the virtual function is trying to catch up to still achieve 60 fps). The greater the expected variance between jobs, the larger n and / or the larger the slack time. In some embodiments, the scheduler module circuit 230 supports dynamic reconfiguration of per-VF behavior + algorithm parameters (e.g., slack, etc.).
[0050] In some embodiments, scheduler module circuitry 230 can issue a single job size credit to at least one of virtual functions 210-213. If any of virtual functions 210-213 has more than an unused job size (submitted jobs smaller than the expected job size) in any of scheduling periods 510-540, that particular one of virtual functions 210-213 gets a single job size credit for the scheduling period immediately following the job size credit allocation. For example, if virtual function 210 had more than an unused job size in the current scheduling period 510, the virtual function gets a single job size credit for the next scheduling period 520 relative to the current scheduling period 510. Thus, if a job of a particular virtual function is delayed, that particular virtual function is permitted to submit two jobs in the next scheduling period. In some embodiments, only a single job size credit is given to a particular virtual function, so as not to cause undue disturbance to other virtual functions. In some embodiments, this job size credit is configurable as to how much the job size credit is carried forward (number of scheduling periods, eg, 0 (job size credit is disabled), 1, 2, etc.).
[0051] In some embodiments, after completing an expected number of jobs within a particular scheduling period, a well-behaved virtual function that has time remaining within that particular scheduling period is granted a one-time exception to execute additional jobs submitted within that particular scheduling period if the GPU is idle. This corresponds to a scenario in which a well-behaved virtual function that completed one or more jobs within a scheduling period still has time remaining within the scheduling period that could be used to complete jobs if granted the exception. In some embodiments, this one-time adaptation is not allowed in the next scheduling period or the next X scheduling periods, where X is configurable, and if repeated, the virtual function is determined to be ill-behaved by the scheduler module circuitry 230. In some embodiments, the exception may be adjustable or disabled, for example, depending on GPU utilization within the scheduling period or recent history.
[0052] 6 is a flow diagram of a method 600 for scheduling first and second virtual functions of a plurality of virtual functions, according to some embodiments. In some embodiments, method 600 is implemented by a processing system, such as processing system 100 of FIG. 1. At block 604, a plurality of time partitions are assigned (allocated) to a plurality of virtual functions for executing a plurality of jobs.
[0053] At block 606, the job cadence monitor 310 determines whether the first plurality of jobs submitted by the first virtual function is within the expected cadence and the second plurality of jobs submitted by the second virtual function exceeds the expected cadence. If at block 606 the job cadence monitor 310 determines that the plurality of jobs 513, 514, 523, 524, 533, 534, 543, 544 submitted by the virtual function 211 exceeds the expected cadence and the job cadence monitor 310 determines that the plurality of jobs 511, 521, 531, 541 submitted by the virtual function 212 is within the expected cadence and the plurality of jobs 512, 522, 532, 542 submitted by the virtual function 210 is within the expected cadence, method flow proceeds to block 610. If, at block 606 , the job cadence monitor 310 determines that any of the jobs submitted by a particular virtual function are not exceeding the expected cadence, then the method flow proceeds from block 606 to block 608 .
[0054] In block 608, job size monitor 320 determines whether a first plurality of jobs submitted by a first virtual function is taking longer (or longer) to execute than the expected job size, and whether a second plurality of jobs submitted by a second virtual function is taking longer (or longer) to execute than the expected job size. Referring to Figures 3 and 5, job size monitor 320 determines that a plurality of jobs 515, 525, 535, 545 submitted by virtual function 213 are taking longer to execute than the expected job size, and that a plurality of jobs 511, 521, 531, 541 submitted by virtual function 210 and a plurality of jobs 512, 522, 532, 542 submitted by virtual function 211 are taking shorter to execute than the expected job size. Although virtual function 212 is shown submitting jobs 513, 514, 523, 524, 533, 534, 543, 544 that exceed the expected cadence, and virtual function 213 is shown submitting jobs 515, 525, 535, 545 that take longer to execute than the expected job size, such is shown for ease of explanation. In some embodiments, a virtual function (not shown) can submit jobs that are a combination of exceeding the expected cadence (e.g., higher than expected fps) and taking longer to execute than the expected job size (e.g., higher than expected resolution) and is scheduled accordingly, having been determined to be a misbehaving virtual function.
[0055] Note that because block 610 has already determined that virtual function 212 is a misbehaving virtual function, job size monitor 320 does not need to determine whether jobs 513, 514, 523, 524, 533, 534, 543, and 544 submitted by virtual function 212 are taking longer to execute than the expected job size, and appropriate corrective action is taken by block 606 for virtual function 212. If job size monitor 320 determines that any of the jobs submitted by a particular virtual function are taking longer to execute than the expected job size, block 606 proceeds to block 610. Otherwise, if job size monitor 320 determines that any of the jobs submitted by a particular virtual function are not taking longer to execute than the expected job size, block 608 proceeds to block 606, and method 600 continues to monitor the expected cadence and expected job size for virtual functions 210-213 by blocks 606 and 608, respectively.
[0056] Although not shown, method 600 can terminate for any one of the virtual functions and continue for the remaining virtual functions. For example, in the context of cloud gaming, a game ceasing execution on a cloud gaming server similarly causes virtual functions 210-213 to stop submitting jobs to GPU 115. Thus, method 600 halts that game but continues to execute any remaining games and any newly added games executed by GPU 115 that still receive jobs from the remaining ones of virtual functions 210-213, thereby continuing to determine whether any virtual functions are submitting jobs beyond the expected cadence and taking longer to execute than the expected job size. Method 600 can further include any of the functionality described above for scheduler module circuit 230, job cadence monitor 310, and / or job size monitor 320.
[0057] In some embodiments, the above-described apparatus and techniques are implemented in a system that includes one or more integrated circuit (IC) devices (also referred to as integrated circuit packages or microchips), such as the processing system 100 described above with reference to FIGS. 1-6 . Electronic design automation (EDA) and computer-aided design (CAD) software tools can be used in the design and manufacture of these IC devices. These design tools are typically represented as one or more software programs. The one or more software programs include code executable by a computer system for operating the computer system to operate on code representing the circuits of one or more IC devices to perform at least a portion of a process for designing or adapting a manufacturing system for producing the circuits. This code may include instructions, data, or a combination of instructions and data. The software instructions representing the design or manufacturing tools are typically stored on a computer-readable storage medium accessible to the computing system. Similarly, code representing one or more stages of the design or manufacture of the IC devices is stored on and accessed from the same or a different computer-readable storage medium.
[0058] A computer-readable storage medium includes any non-transitory storage medium or combination of non-transitory storage media that can be accessed by a computer system during use to provide instructions and / or data to the computer system. Such storage media may include, but are not limited to, optical media (e.g., compact discs (CDs), digital versatile discs (DVDs), Blu-ray discs), magnetic media (e.g., floppy disks, magnetic tape, magnetic hard drives), volatile memory (e.g., random access memory (RAM) or cache), non-volatile memory (e.g., read-only memory (ROM) or flash memory), or micro-electromechanical systems (MEMS)-based storage media. The computer-readable storage medium (e.g., system RAM or ROM) may be internal to the computing system, the computer-readable storage medium (e.g., a magnetic hard drive) may be permanently attached to the computing system, the computer-readable storage medium (e.g., an optical disk or Universal Serial Bus (USB)-based flash memory) may be removably attached to the computing system, or the computer-readable storage medium (e.g., network-accessible storage (NAS)) may be coupled to the computer system via a wired or wireless network.
[0059] In some embodiments, certain aspects of the techniques described above are implemented by one or more processors of a processing system executing software. The software includes one or more sets of executable instructions stored or otherwise tangibly embodied in a non-transitory computer-readable storage medium. The software may include instructions and specific data that, when executed by one or more processors, operate the one or more processors to perform one or more aspects of the techniques described above. The non-transitory computer-readable storage medium may include, for example, a magnetic or optical disk storage device, a solid-state storage device such as flash memory, a cache, a random access memory (RAM), or other non-volatile memory device(s). The executable instructions stored on the non-transitory computer-readable storage medium may be implemented as source code, assembly language code, object code, or other form of instructions that can be interpreted or otherwise executed by one or more processors.
[0060] In addition to the above, it should be noted that not all activities or elements described in the summary description are required, that some of the particular activities or devices may not be required, that one or more additional activities may be performed, and that one or more additional elements may be included. Furthermore, the order in which the activities are listed is not necessarily the order in which they are performed. Also, the concepts have been described with reference to specific embodiments. However, those skilled in the art will recognize that various modifications and variations can be made without departing from the scope of the invention as set forth in the claims. Accordingly, the specification and drawings should be regarded in an illustrative rather than a restrictive sense, and all such modifications are intended to be included within the scope of the present invention.
[0061] Benefits, other advantages, and solutions to problems have been described above with regard to specific embodiments. However, the benefits, advantages, solutions to problems, and features from which any benefit, advantage, or solution may arise or be manifested are not construed as critical, essential, or essential features of any or all claims. Moreover, the specific embodiments described above are illustrative only, since the disclosed invention may be modified and practiced in different, but similar manners apparent to those skilled in the art having the benefit of the teachings herein. No limitations are intended to the details of construction or design herein shown, other than as described in the appended claims. It is therefore apparent that the specific embodiments described above may be altered or modified, and that all such variations are considered within the scope of the disclosed invention. Accordingly, the protection sought herein is set forth in the appended claims.
Claims
1. 1. A method comprising: allocating a plurality of time partitions within a scheduling period to a plurality of virtual functions for execution of jobs on a parallel processor, the scheduling period including slack time to allow for variance in at least one of a cadence at which each virtual function submits jobs for execution and a size of jobs submitted for execution; preventing execution of a job that exceeds a time partition assigned to any of the plurality of virtual functions; method.
2. Preventing said execution includes: scheduling a first virtual function to execute the first plurality of jobs after a second virtual function in response to submission of a first plurality of jobs exceeding an expected cadence and submission of a second plurality of jobs by the second virtual function not exceeding the expected cadence; 10. The method of claim 1.
3. scheduling a first virtual function to execute the first plurality of jobs after a second virtual function in response to a first plurality of jobs exceeding an expected job size and a second plurality of jobs submitted by a second virtual function not exceeding the expected job size; 10. The method of claim 1.
4. assigning a job size credit to a first virtual function when a size of a job submitted by the first virtual function is smaller than an expected job size; the job size credit may be used by the first virtual function in a scheduling period immediately subsequent to the scheduling period in which the job size credit is allocated; 10. The method of claim 1.
5. maintaining a first level list and a second level list, the first level list including a first virtual function that submits a first plurality of jobs that are within an expected cadence and take less time to execute than an expected job size, and the second level list including a second virtual function that is not included in the first level list; scheduling the first plurality of jobs of virtual functions in the first level list before scheduling a second plurality of jobs of virtual functions in the second level list.
10. The method of claim 1.
6. bypassing scheduling the second plurality of jobs of virtual functions in the second level list within a current scheduling period. The method of claim 5.
7. the scheduling period is a first scheduling period of a plurality of scheduling periods; scheduling the first virtual function after a second virtual function when a first number of a first plurality of jobs submitted by the first virtual function within the plurality of scheduling periods is greater than a second number of a second plurality of jobs submitted by the second virtual function within the plurality of scheduling periods; 10. The method of claim 1.
8. enabling a virtual function that has time remaining within an assigned time partition and has not completed a job within the scheduling period to submit the job when the parallel processor is idle. The method of any one of claims 1 to 7.
9. 1. A processing system comprising: a parallel processor; a scheduler module circuit; The scheduler module circuit includes: allocating a plurality of time partitions within a scheduling period to a plurality of virtual functions for execution of jobs on the parallel processor, the scheduling period including slack time to allow for variance in at least one of a cadence at which each virtual function submits jobs for execution and a size of jobs submitted for execution; Preventing execution of a job that exceeds a time partition assigned to a first virtual function of the plurality of virtual functions; configured to: Processing system.
10. The scheduler module circuit includes: configured to schedule the first virtual function after the second virtual function in response to submission of a first plurality of jobs by the first virtual function exceeding an expected cadence and submission of a second plurality of jobs by a second virtual function not exceeding the expected cadence; The processing system of claim 9.
11. The scheduler module circuit includes: configured to schedule the first virtual function after the second virtual function in response to a first plurality of jobs submitted by the first virtual function exceeding an expected job size and a second plurality of jobs submitted by a second virtual function not exceeding the expected job size. The processing system of claim 9.
12. The scheduler module circuit includes: configured to allocate a job size credit to a virtual function when a size of a job submitted by the virtual function is smaller than an expected job size; the job size credit may be used by the virtual function in a scheduling period immediately subsequent to the scheduling period in which the job size credit is allocated; The processing system of claim 9.
13. The scheduler module circuit includes: maintaining a first level list and a second level list, the first level list including the first virtual functions that submit a first plurality of jobs that are within an expected cadence and take less time to execute than an expected job size, and the second level list including a second virtual function that is not included in the first level list; scheduling the first plurality of jobs for virtual functions in the first level list before scheduling a second plurality of jobs for virtual functions in the second level list; configured to: The processing system of claim 9.
14. The scheduler module circuit includes: configured to bypass scheduling the second plurality of jobs of the second virtual function within a current scheduling period. The processing system of claim 13.
15. the scheduling period is a first scheduling period of a plurality of scheduling periods; The scheduler module circuit includes: configured to schedule the first virtual function after the second virtual function when a first number of the first plurality of jobs submitted by the first virtual function within the plurality of scheduling periods is greater than a second number of a second plurality of jobs submitted by the second virtual function within the plurality of scheduling periods; The processing system of claim 9.
16. The scheduler module circuit includes: configured to allow a virtual function that has remaining time within an assigned time partition and has not completed a job within the scheduling period to submit the job when the parallel processor is idle. The processing system according to any one of claims 9 to 15.
17. a server, a parallel processor configured to execute jobs submitted by a plurality of virtual functions; a scheduler module circuit; The scheduler module circuit includes: assigning a time partition to each of the plurality of virtual functions based on expected job size and cadence within a scheduling period and slack time to allow for distribution of submitted job sizes and cadences; Preventing execution of a job in the parallel processor that exceeds a time partition assigned to any one of the plurality of virtual functions; configured to: server.
18. The scheduler module circuit includes:
18. The server of claim 17, further configured to schedule a first virtual function before a second virtual function in response to a first plurality of jobs submitted by a first virtual function having a frequency below an expected cadence and a second plurality of jobs submitted by a second virtual function exceeding the expected cadence.
19. The scheduler module circuit includes: configured to schedule a first virtual function before a second virtual function in response to a first plurality of jobs submitted by the first virtual function not exceeding an expected job size and a second plurality of jobs submitted by the second virtual function exceeding the expected job size; 18. The server of claim 17.
20. The scheduler module circuit includes: maintaining a first level list and a second level list, the first level list including a first virtual function that submits a first plurality of jobs that are within an expected cadence and take less time to execute than an expected job size, and the second level list including a second virtual function that is not included in the first level list; scheduling the first plurality of jobs of the first virtual function in the first level list before scheduling a second plurality of jobs of the second virtual function in the second level list; configured to:
18. The server of claim 17.