Multi-level scheduling for improved quality of service

By introducing scheduling cycles and slack time in the virtualized environment, monitoring job submissions, the scheduler module circuit prevents the virtual function from disproportionately consuming hardware resources, solving the problem of degradation in service quality of virtual functions, and achieving fair hardware resource allocation and improvement of service quality.

CN120584337APending Publication Date: 2025-09-02ADVANCED MICRO DEVICES INC +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202380084995.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2022-12-21
Filing Date
2023-12-20
Publication Date
2025-09-02

AI Technical Summary

Technical Problem

In a virtualized environment, virtual functions may consume hardware resources disproportionately, resulting in a degradation in service quality of other virtual functions, especially if job submissions are not at the expected pace and size.

Method used

By allocating time partitions during the scheduling cycle and introducing slack time, monitoring the submission rhythm and size of the job, the scheduler module circuit prevents virtual functions from exceeding their expected use, maintains the service quality of good-performing virtual functions, and prioritizes good-performing virtual functions through multi-level scheduling mechanisms.

Benefits of technology

It effectively prevents the disproportionate consumption of hardware resources by virtual functions, improves the service quality of virtual functions, and ensures that each virtual function obtains fair hardware resource allocation within expectations.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120584337A_ABST
    Figure CN120584337A_ABST
Patent Text Reader

Abstract

The parallel processor (115) is configured to implement job restrictions on the virtual functions (210, 211, 212, 213) to facilitate an expected quality of service for each of the virtual functions assigned to a virtual machine (220, 221, 222, 223) executing at the processing unit. A scheduler (230) schedules a well-functioning virtual function prior to a poorly-functioning virtual function to prevent the poorly-functioning virtual function from consuming a disproportionate hardware resource share, thereby mitigating the impact of the poorly-functioning virtual function on the quality of service of the well-functioning virtual function.
Need to check novelty before this filing date? Find Prior Art

Description

Background Art

[0001] Conventional processing units such as graphics processing units (GPUs) support virtualization, which allows multiple virtual machines (VMs) to use the GPU's hardware resources. Some VM implementations allow the VM to emulate the operating system of a physical machine. Other VMs are designed to execute code in a platform-independent environment. A hypervisor creates and runs a VM, which is also called a client or guest. The virtual environment implemented on the GPU also provides virtual functions for other virtual components implemented on the physical machine. A single physical function implemented in the GPU is used to support one or more virtual functions (VFs). The physical function assigns virtual functions to different VMs on the physical machine based on time slicing or time partitioning. For example, the physical function assigns a first virtual function to a first VM in a first time interval and assigns a second virtual function to a second VM in a subsequent second time interval. The Single Root Input / Output Virtualization (SR-IOV) specification allows multiple VMs to share a GPU interface on a single bus, such as a Peripheral Component Interconnect Express (PCIe) bus. Components access virtual functions by transmitting requests over the bus. Summary of the Invention

[0002] To prevent virtual functions from consuming a disproportionate share of hardware resources, a processing system enforces job limits on virtual functions to promote a desired quality of service for each of the virtual functions assigned to a virtual machine executing on the processing system. In a first embodiment, a method includes allocating a plurality of time partitions within a scheduling cycle for executing jobs on a parallel processor to a plurality of virtual functions, wherein the scheduling cycle includes slack time to allow for variation in at least one of a cadence of jobs submitted for execution by each virtual function and a size of the jobs submitted for execution. The method also includes preventing execution of jobs that exceed the time partitions allocated to the virtual functions of the plurality of virtual functions.

[0003] In some implementations, preventing execution includes scheduling the first virtual function to execute the first plurality of jobs after the second virtual function in response to submission of the first plurality of jobs exceeding an expected cadence and submission of the second plurality of jobs by the second virtual function not exceeding the expected cadence.

[0004] The method may further include scheduling the first virtual function to execute the first plurality of jobs after the second virtual function in response to the first plurality of jobs exceeding an expected job size and a second plurality of jobs submitted by the second virtual function not exceeding the expected job size.

[0005] In some implementations, the method further includes assigning a job size credit to the first virtual function if the size of the job submitted by the first virtual function is less than the expected job size. The job size credit is usable by the first virtual function in a subsequent scheduling cycle immediately following the scheduling cycle during which the job size credit was assigned.

[0006] The method may further include maintaining a first-level list and a second-level list. The first-level list includes a first virtual function that submitted a first plurality of jobs that were within an expected cadence and did not take longer to execute than an expected job size, and the second-level list includes a second virtual function that was not included in the first-level list. The method may further include scheduling the first plurality of jobs for the virtual function in the first-level list before scheduling the second plurality of jobs for the virtual function in the second-level list.

[0007] In some implementations, the method further includes: bypassing scheduling the second plurality of jobs of the virtual function in the second-level list during a current scheduling period. In some implementations, the scheduling period is a first scheduling period of a plurality of scheduling periods, and the method further includes: scheduling the first virtual function after the second virtual function if a first number of the first plurality of jobs submitted by the first virtual function during the plurality of scheduling periods is greater than a second number of the second plurality of jobs submitted by the second virtual function during the plurality of scheduling periods.

[0008] The method may further include: when the parallel processor is idle, allowing a virtual function that has remaining time in the allocated time partition and has not completed a job within the scheduling period to submit the job.

[0009] In another specific implementation, a processing system includes: a parallel processor; and a scheduler module circuit configured to allocate a plurality of time partitions within a scheduling period to a plurality of virtual functions for executing jobs on the parallel processor. The scheduling period includes slack time to allow for variations in at least one of a cadence of jobs submitted for execution by each virtual function and a size of the jobs submitted for execution. The scheduler module circuit is further configured to prevent execution of a job that exceeds the time partition allocated to a first virtual function among the plurality of virtual functions.

[0010] The scheduler module circuitry may also be configured to: in response to the submission of the first plurality of jobs by the first virtual function exceeding an expected cadence and the submission of the second plurality of jobs by the second virtual function not exceeding the expected cadence, schedule the first virtual function after the second virtual function. In some implementations, the scheduler module circuitry may be further configured to: in response to the first plurality of jobs submitted by the first virtual function exceeding an expected job size and the second plurality of jobs submitted by the second virtual function not exceeding the expected job size, schedule the first virtual function after the second virtual function.

[0011] The scheduler module circuitry may be further configured to assign a job size credit to the virtual function if the size of the job submitted by the virtual function is less than the expected job size. The job size credit is usable by the virtual function in a subsequent scheduling cycle immediately following the scheduling cycle during which the job size credit was assigned.

[0012] In some specific implementations, the scheduler module circuitry is further configured to maintain a first-level list and a second-level list. The first-level list includes the first virtual function that submitted a first plurality of jobs, the first plurality of jobs being executed within an expected cadence and not taking longer than an expected job size, and the second-level list includes a second virtual function that is not included in the first-level list. The scheduler module circuitry is further configured to schedule the first plurality of jobs for the virtual function in the first-level list before scheduling the second plurality of jobs for the virtual function in the second-level list.

[0013] The scheduler module circuitry may also be configured to bypass scheduling the second plurality of jobs of the second virtual function during a current scheduling cycle. In some implementations, the scheduling cycle is a first scheduling cycle among a plurality of scheduling cycles, and the scheduler module circuitry may be further configured to schedule the first virtual function after the second virtual function if a first number of the first plurality of jobs submitted by the first virtual function during the plurality of scheduling cycles is greater than a second number of the second plurality of jobs submitted by the second virtual function during the plurality of scheduling cycles.

[0014] In some specific implementations, the scheduler module circuit is further configured to: when the parallel processor is idle, allow a virtual function that has remaining time in the allocated time partition and has not completed the job within the scheduling cycle to submit the job.

[0015] In another embodiment, a server includes a parallel processor configured to execute jobs submitted by a plurality of virtual functions, and a scheduler module circuit configured to: assign a time partition to each of the plurality of virtual functions within a scheduling period based on an expected job size and cadence and slack time to allow for variations in the size and cadence of the submitted jobs; and prevent execution of a job on the parallel processor that exceeds the time partition assigned to the virtual function in the plurality of virtual functions.

[0016] The scheduler module circuitry may be further configured to schedule the first virtual function before the second virtual function in response to a first plurality of jobs submitted by the first virtual function having a frequency below an expected cadence and a second plurality of jobs submitted by the second virtual function exceeding the expected cadence.

[0017] In some specific implementations, the scheduler module circuit is further configured to: in response to a first plurality of jobs submitted by a first virtual function not exceeding an expected job size and a second plurality of jobs submitted by a second virtual function exceeding the expected job size, schedule the first virtual function before the second virtual function. The scheduler module circuit may also be configured to: maintain a first-level list and a second-level list. The first-level list includes the first virtual function that submitted the first plurality of jobs, the first plurality of jobs being within the expected cadence and not taking longer to execute than the expected job size, and the second-level list includes the second virtual function that is not included in the first-level list. The scheduler module circuit is further configured to: schedule the first plurality of jobs of the first virtual function in the first-level list before scheduling the second plurality of jobs of the second virtual function in the second-level list. BRIEF DESCRIPTION OF THE DRAWINGS

[0018] The present disclosure may be better understood by reference to the accompanying drawings, and its numerous features and advantages will be apparent to those skilled in the art. The use of the same reference numerals in different drawings indicates similar or identical items.

[0019] Figure 1 is an exemplary block diagram of a processing system configured to enforce job limits on virtual functions, according to some embodiments.

[0020] Figure 2 is an exemplary block diagram of the mapping of virtual functions to virtual machines implemented in a processing unit according to some embodiments.

[0021] Figure 3 is an exemplary block diagram of a scheduler module according to some embodiments.

[0022] Figure 4is an exemplary block diagram of time partitioning to support fair access to virtual machines associated with virtual functions in a processing unit, according to some embodiments.

[0023] Figure 5 is an exemplary diagram illustrating the cadence and job sizes of various jobs submitted for execution by multiple virtual functions according to some embodiments.

[0024] Figure 6 is a flow chart of an exemplary method of scheduling a first virtual function and a second virtual function according to some embodiments. DETAILED DESCRIPTION

[0025] The GPU's hardware resources are partitioned using physical functions (PFs) and one or more virtual functions (VFs) according to SR-IOV. Each virtual function is associated with a single physical function. In a native (host OS) environment, the physical function is used by native user mode, and the kernel mode driver and all virtual functions are disabled. All GPU registers are assigned to the physical function via trusted access. In a virtual environment, the physical function is used by the hypervisor (host VM), and the GPU exposes a certain number of virtual functions according to the PCIe SR-IOV standard, such as one virtual function per guest VM. Each virtual function is assigned to the guest VM by the hypervisor.

[0026] Typically, a central processing unit (CPU) is partitioned across virtual functions so that each virtual function has a dedicated CPU. The CPU prepares and submits jobs for the virtual functions to the GPU. Each virtual function receives remote user input and prepares job submissions based on the remote user input, and may also submit jobs unrelated to the user input. The CPU can submit jobs for a virtual function to the GPU at any time; however, execution of the job on the GPU occurs during the time partition assigned to the virtual function. In the case where many of the jobs submitted by the virtual functions are for streaming video, the virtual functions will submit jobs to be executed at a fixed cadence to achieve a target frames per second (FPS) rate. Continuously submitting single or multiple work units that collectively correspond to "jobs" of similar size results in the jobs being executed at a fixed cadence.

[0027] However, because jobs are submitted based on CPU readiness timing, they may not align with the GPU time partition assigned to a virtual function. Additionally, the sizes of jobs submitted by multiple virtual functions may vary, causing each job to take longer to complete execution than the assigned time partition. Furthermore, in some instances, virtual functions may act in a greedy or malicious manner by submitting an excessive number of jobs (or jobs that collectively take longer to execute) within the assigned time partition, thereby not submitting jobs within the expected cadence.

[0028] When a virtual function behaves in a greedy or malicious manner, it consumes a disproportionate share of hardware resources, which can negatively impact the quality of service of other virtual functions assigned to the VM. In some embodiments, the impact on the virtual function is based on both throughput and latency. Throughput refers to the rate at which a job is executed, which in some cases is related to one or more of the encoding resolution and frame rate, the video decoding rate, or the rendering rate of the desktop or game frames experienced by each virtual function. Latency refers to the time between when a job is submitted to the GPU and when the job completes execution at the GPU, so that the results of the job can be consumed within the expected latency. For example, in the context of video encoding, latency refers to the time it takes for a job to complete so that the encoded frames can be streamed.

[0029] Figures 1 to 5 Embodiments of a processing unit (such as a graphics processing unit (GPU)) of a processing system or server are disclosed. The processing unit is configured to enforce job limits on virtual functions (VFs) to promote an expected quality of service for each VF assigned to a VM executing on the processing unit. Scheduler module circuitry defines a scheduling period as the sum of a per-VF time partition plus an additional time period referred to herein as "slack time." The slack time allows for occasional variations in the cadence of jobs submitted by each VF for execution by the GPU. For example, if a frame takes longer than expected to render, the VF may miss an opportunity to submit the frame during its allocated time slice and may submit the frame along with subsequent frames during the next allocated time slice, resulting in a first time slice in which the VF submits no jobs and a second time slice in which the VF submits a batch of two jobs. Such occasional variations in the cadence of job submissions are expected and do not indicate greedy or malicious behavior on the part of the VFs. The GPU defines a per-VF time partition as the VF's expected job size multiplied by N, where N is a number equal to or greater than 1 and depends on the GPU's tolerance for variations in job submission behavior.

[0030] The GPU monitors jobs submitted for execution by virtual functions during a scheduling cycle to determine whether the jobs are submitted within an expected cadence or frequency. The GPU also monitors whether the submitted jobs take longer to execute than the expected job size. Based on whether the respective virtual functions submit jobs at the expected cadence and based on whether the submitted jobs take longer to execute than the expected job size, the respective virtual functions are designated as well-performing or underperforming. The GPU scheduler schedules well-performing virtual functions ahead of underperforming virtual functions to prevent the underperforming virtual functions from consuming a disproportionate share of hardware resources, thereby mitigating the impact of the underperforming virtual functions on the quality of service of the well-performing virtual functions.

[0031] Figure 1 is a block diagram of a processing system 100 configured to enforce job limits on virtual functions, according to some embodiments. In various embodiments, the techniques described herein are applied to any of a variety of parallel processors, such as vector processors, graphics processing units (GPUs), general-purpose GPUs (GPGPUs), non-scalar processors, highly parallel processors, artificial intelligence (AI) processors, inference engines, machine learning processors, other multi-threaded processing units, and the like. Figure 1 The example of a parallel processor according to some embodiments is illustrated, and in particular is an example of a graphics processing unit (GPU) 115 (e.g., a virtual GPU). However, unless otherwise specified, references herein to a GPU will be understood to include any of a variety of parallel processors.

[0032] The processing system 100 includes or has access to a memory 105 or other storage component implemented using non-transitory computer-readable media, such as dynamic random access memory (DRAM). However, the memory 105 may also be implemented using other types of memory, including static random access memory (SRAM), non-volatile RAM, etc. The processing system 100 also includes a bus 110 to support communication between entities implemented in the processing system 100, such as the memory 105. In the illustrated embodiment, the bus 110 is configured as a PCIe bus. Some embodiments of the processing system 100 include other buses, bridges, switches, routers, etc., which are not listed here for clarity. Figure 1 Shown in.

[0033] The processing system 100 also includes a central processing unit (CPU) 150 that is connected to the bus 110 and communicates with the GPU 115 and the memory 105 via the bus 110. In the illustrated embodiment, the CPU 150 implements a plurality of processing elements (also referred to as processor cores) 155 that are configured to execute instructions concurrently or in parallel. The CPU 150 executes instructions, such as program code 160 stored in the memory 105, and the CPU 150 stores information, such as the results of the executed instructions, in the memory 105. The CPU 150 initiates graphics processing by issuing draw calls to the GPU 115.

[0034] Input / output (I / O) engine 165 handles input or output operations associated with display 120 and other elements of processing system 100 (such as a keyboard, mouse, printer, external disk, network, etc.). I / O engine 165 is coupled to bus 110, allowing I / O engine 165 to communicate with memory 105, GPU 115, or CPU 150. In the illustrated embodiment, I / O engine 165 is configured to read information stored on external storage component 170, which is implemented using a non-transitory computer-readable medium such as a flash drive. I / O engine 165 can also write information (such as the results of processing performed by GPU 115 or CPU 150) to external storage component 170. Display 120 can be remotely connected to the VM via a network connection using an appropriate protocol.

[0035] The processing system 100 includes one or more graphics processing units (GPUs) 115 configured to render images for presentation on a display 120. For example, the GPU 115 may render an object to generate pixel values ​​that are provided to the display 120, which uses the pixel values ​​to display an image representing the rendered object. The GPU 115 includes a GPU core 125, which is composed of a set of computational units, a set of fixed-function units, or a combination thereof for executing instructions concurrently or in parallel. The GPU core 125 may include dozens, hundreds, or even thousands of computational units or fixed-function units for executing instructions.

[0036] GPU 115 includes internal (or on-chip) memory 130, which includes a frame buffer and a local data store (LDS), as well as caches, registers, or other buffers utilized by the compute units in GPU core 125. Internal memory 130 stores data structures that describe tasks executed on one or more of the compute units or fixed-function units in GPU core 125. The compute units or fixed-function units in GPU core 125 can also access information in (external) memory 105. In the illustrated embodiment, GPU 115 communicates with memory 105 via bus 110. However, some embodiments of GPU 115 communicate with memory 105 via a direct connection or via other buses, bridges, switches, routers, etc. GPU 115 executes instructions stored in memory 105, and GPU 115 stores information (such as the results of the executed instructions) in memory 105. For example, memory 105 may store a copy 135 of instructions from program code to be executed by GPU 115 , such as program code representing shaders, virtual functions, or other code executed by one or more of the compute units or fixed function units implemented in GPU core 125 .

[0037] GPU 115 includes an encoder 140 for encoding information for transmission over bus 110. Encoder 140 also provides security functions to support secure communication over bus 110. In some embodiments, encoder 140 encodes pixel values ​​for transmission to display 120, which implements a decoder to decode the pixel values ​​to reconstruct an image for presentation. Display 120 can be remotely connected to the VM via a network connection. Some embodiments of encoder 140 encode and encrypt information generated by virtual functions implemented on GPU 115 for communication over bus 110.

[0038] Some embodiments of GPU 115 operate as a physical function supporting one or more virtual functions shared via bus 110. For example, GPU 115 can use a dedicated portion of bus 110 to securely share multiple VMs using the SR-IOV standard defined for the PCIe bus. GPU 115 includes a bus interface 145 that provides an interface between GPU 115 and bus 110, for example, in accordance with the SR-IOV standard. Bus interface 145 provides functionality including doorbell detection, register redirection, frame buffer aperture, doorbell write redirection, and other functionality as discussed below.

[0039] As discussed in more detail below, at least one of the multiple virtual functions supported and enabled by the GPU 115 may submit jobs for execution that exceed their respective expected usage of the multiple virtual functions. For example, the job submitted for execution by the virtual function is a video encoding job. In some embodiments, the video encoding job encodes video information of a game application executed by the VM (such as a game application executed by a cloud gaming platform). The game can be expected to execute at a maximum frames per second (fps) and a maximum resolution. In order to prevent the job submitted for execution by at least one of the multiple virtual functions from exceeding its expected usage of the multiple virtual functions, the GPU 115 includes a scheduler that prevents the execution of jobs for the virtual function that exceed its expected usage or are expected to exceed its expected usage before executing jobs for the virtual function that have not yet exceeded its expected usage or are expected to not exceed its expected usage. In some embodiments, the virtual functions may also each include a scheduler that limits job submissions.

[0040] Figure 2 is a block diagram of a mapping 200 of virtual functions to VMs implemented in a processing unit according to some embodiments. The mapping 200 represents the mapping of virtual functions to VMs implemented in a processing unit. Figure 1115. A host 201 includes a physical function 205 (such as GPU 115), which is partitioned into virtual functions 210, 211, 212, 213 during initialization of the physical function 205. In some embodiments, each of the virtual functions 210-213 includes an application (e.g., a video game), an application programming interface (API), a user-mode driver (UMD), and a kernel-mode driver (KMD). The host 201 implements a host operating system or hypervisor 215 for the physical function 205. The hypervisor 215 launches one or more VMs 220, 221, 222, 223 for execution on the physical resources (such as GPU 115) supporting the physical function 205. In some embodiments, VMs 220-223 include a GPU virtualization driver (GPUV) that can receive configuration information assigned to each of the virtual functions 210-213 (e.g., from a server administrator and pass the configuration information to the virtual GPU component and the virtual video core or virtual engine (e.g., a Video Core Next Generation (VCN)).

[0041] VMs 220-223 are assigned to corresponding virtual functions 210-213. In the illustrated embodiment, VM 220 is assigned to virtual function 210, VM 221 is assigned to virtual function 211, VM 222 is assigned to virtual function 212, and VM 223 is assigned to virtual function 213. Virtual functions 210-213 submit jobs to GPU 115, which provides GPU functionality to the corresponding VMs 220-223. Thus, the virtualized GPU 115 is shared across many VMs 220-223. In some embodiments, time slicing (also known as time partitioning) and context switching are used to provide fair access to GPU 115 for virtual functions 210-213, so that each of virtual functions 210-213 is assigned a corresponding time partition for GPU 115 to execute multiple jobs.

[0042] VMs 220-223 also each include a scheduler module circuit 230 that manages virtual function access to GPU 115. In some embodiments, GPU 115 includes scheduler module circuit 230. In some embodiments, scheduler module circuit 230 is an SR-IOV multimedia GPU scheduler, a hardware and / or firmware scheduler. In some embodiments, scheduler module circuit 230 is implemented in various forms, such as a processor, a field programmable gate array (FPGA), or other forms of circuitry. This GPU-side implementation of scheduler module circuit 230 for underperforming virtual functions is difficult to bypass, thereby providing security for such scheduling. Scheduler module circuit 230 defines a time period or scheduling period during which jobs submitted by VMs 220-223 can be executed by GPU 115, respectively.

[0043] In some embodiments, the scheduler module circuitry 230 assigns a time slice to each of the virtual functions 210-213 that has a tolerance for variations in submission behavior. In such embodiments, the scheduler module circuitry 230 may assign a time partition to each of the virtual functions 210-213 that is equal to the expected job size of the particular virtual function multiplied by n, where n is 1, 2, etc., depending on the desired tolerance for variations in submitting jobs to the plurality of virtual functions 210-213. For example, four (4) virtual functions 210-213 may submit jobs for execution by the GPU 115 at 60fps and 1080p resolution, where each job is expected to be <3ms. If n=2, then the time partition assigned to each of the virtual functions 210-213 is equal to 6ms, and when n is 2, 2 jobs*3ms / job=6ms. In some embodiments, the scheduler module circuitry 230 supports configurable tolerances for various metrics (e.g., expected job size + / - tolerance). In some embodiments, at least one of the virtual functions 210-213 supports multiple streams (eg, a single virtual function is submitted to execute 1 1080p 60fps stream + 1 720p 30fps stream).

[0044] The scheduler module circuit 230 includes a job pacing monitor circuit (referred to as job pacing monitor 310) and / or a job size monitor circuit (referred to as job size monitor 320), such as Figure 3As shown. Job cadence monitor 310 monitors jobs submitted by virtual functions 210-213 during a scheduling cycle to determine whether the jobs are executed by GPU 115 at an expected cadence, with the expected cadence having a time period between jobs submitted by the virtual function that is approximately equal between the various jobs of a particular virtual function. In some embodiments, job cadence monitor 310 monitors whether an expected number of jobs are received during each scheduling cycle because the time between job submissions can vary due to submission jitter. Job size monitor 320 also monitors the size of jobs submitted by virtual functions 210-213 for execution by GPU 115 during each scheduling cycle. Based on the monitoring performed by the job pacing monitor 310 and / or the job size monitor 320, the scheduler module circuitry 230 determines which of the virtual functions 210-213 are overutilizing the bandwidth (such as the GPU 115) and time partition of the assigned virtual engine, and which of the virtual functions 210-213 are not "underperforming," but are "well-performing" because they do not "over-utilize" or over-utilize the bandwidth and time partition of the assigned virtual engine. The scheduler module circuitry 230 schedules the virtual functions 210-213 based at least in part on determining whether the virtual functions are determined to be underperforming or well-performing.

[0045] Figure 4 is a block diagram of time partitioning 400 that supports fair access to virtual machines associated with virtual functions in GPU 115, according to some embodiments. Time partitioning 400 is Figure 1 The time partition 400 is used to provide the GPU 115 of the embodiment shown. Figure 3 Fair access to some embodiments of virtual machines 210-213 is shown. Figure 3 , time increases from left to right. A first time partition 405 is assigned to a first virtual function, such as virtual machine 220 assigned to virtual function 210. State information of virtual function 210 is communicated via a bus interface such as Figure 1145). Once the first time partition 405 is complete, the processing unit performs a context switch 406, which includes saving the current context and state information of the first virtual function to memory. The context switch 406 also includes retrieving the context and state information of the second virtual function from memory and loading the information into memory or registers in the processing unit. A second time partition 407 is assigned to the second virtual function so that the second virtual function has full access to the resources of the processing unit for the duration of the second time partition 407. The scheduler module circuit 230 defines the scheduling period as the sum of these per-virtual function time partitions plus an additional time period. These per-virtual function time partitions are periods in which the expected job size executed for the virtual function is multiplied by N, where N is a number equal to or greater than 1 and depends on the GPU's tolerance for variations in job submission behavior.

[0046] Figure 5 1 is a diagram illustrating the cadence and job sizes of various jobs submitted by multiple virtual functions for execution by GPU 115. Multiple time partitions are assigned to virtual functions 210-213. These time partitions are assigned to virtual function 210 to submit jobs 511, 521, 531, 541, to virtual function 211 to submit jobs 512, 522, 532, 542, to virtual function 212 to submit jobs 513, 514, 523, 524, 533, 534, 543, 544, and to virtual function 213 to submit jobs 515, 525, 535, 545. Job cadence monitor 310 determines that the multiple jobs 511, 521, 531, 541 submitted for execution by virtual function 210 are within the expected cadence, and that the multiple jobs 512, 522, 532, 542 submitted for execution by virtual function 211 are within the expected cadence. As shown, jobs 511, 521, 531, 541, 512, 522, 532, 542 submitted by virtual functions 210, 211, respectively, start at substantially the same time within each of the scheduling cycles 510-540. Thus, each of virtual functions 210, 211 submits a single job at an expected cadence within each of the first, second, third, and fourth scheduling cycles 510-540. As used herein, jobs "submitted" by virtual functions 210-213 include not only jobs actually executed by GPU 115, but also jobs that attempt to use a disproportionate share of the bandwidth available to GPU 115. In some embodiments, virtual functions 210-213 submit jobs at substantially the same expected cadence. In other embodiments, virtual functions 210-213 submit jobs at different cadences (e.g., some virtual functions submit jobs at 30 fps, and some virtual functions submit jobs at 60 fps).

[0047] 21 1 . Compared to the jobs submitted by virtual functions 210, 211, virtual function 212 is shown as submitting jobs more frequently than virtual functions 210, 211. Virtual function 212 is shown as submitting jobs 513, 514, 523, 524, 533, 534, 543, 544. Thus, virtual function 212 submits eight (8) jobs for execution within four (4) scheduling periods 510, 520, 530, 540. With an expected cadence of four (4), virtual function 212 submits more than its fair share of jobs for execution. Job cadence monitor 310 determines that virtual function 212 is submitting more jobs for execution than the expected number of jobs for execution for virtual function 212 and identifies virtual function 212 as underperforming. Also, compared to the jobs submitted by virtual functions 210, 211, virtual function 213 is shown as submitting jobs that are larger than the expected job size. Although the virtual function 213 submitted four (4) jobs 515, 525, 535, 545 within four (4) scheduling cycles 510, 520, 530, 540, the size of the jobs submitted by the virtual function 213 is larger than the jobs submitted by the virtual functions 210, 211 and is larger than the expected job size. The job size monitor 320 determines that the virtual function 213 is submitting jobs for execution that are larger than the expected job size and identifies the virtual function 213 as underperforming.

[0048] In some embodiments, the job submitted by the virtual function 210-213 is a video encoding job, such as for a video game. Each virtual machine in the virtual machine 220-223 can execute a video game independently. Based on the configuration of each video game in the video game, it is expected that each video game in the video game will be executed with a 1080p resolution, wherein balanced encoding is preset to 60 frames per second (fps). If all jobs in the job submitted by the virtual function 210-213 are executed with 60fps and 1080p resolution, the scheduler module circuit 230 will identify all virtual functions in the virtual function 210-213 as well-performing virtual functions. However, in some cases, such as when a malicious virtual function utilizes the open source code (e.g., OpenGL) being executed by the virtual machine, the virtual function continues to exceed the expected job submission rhythm and / or job size. For example, the virtual function 210-213 can submit jobs for execution more frequently than the expected fps (e.g., 120fps compared to the expected 60fps). The job cadence monitor 310 monitors the cadence of jobs submitted at a rate of 120 fps and identifies the virtual functions 210-213 that submitted the jobs as underperforming virtual functions. For another example, a virtual function that continuously submits jobs with an expected video resolution greater than 1080p (such as 4k resolution) is attempting to use an unfair share of the physical resources of the GPU 115. The job size monitor 320 identifies such virtual functions 210-213 as underperforming virtual functions. A virtual function determined to be underperforming can negatively impact the execution of jobs submitted by virtual functions determined to be well-performing without intervention by the GPU 115.

[0049] Once the scheduler module circuitry 230 determines which of the virtual functions are underperforming and which are performing well, the scheduler module circuitry 230 implements multi-level scheduling to minimize the impact of the underperforming virtual functions on the well-performing virtual functions, thereby improving the quality of service (QoS) of the well-performing virtual functions. In some embodiments, the scheduler module circuitry 230 maintains two lists: a first list for well-performing virtual functions and a second list for underperforming virtual functions. When the virtual engine is idle, the scheduler module circuitry 230 schedules the well-performing virtual functions from the first list. If there are no pending jobs for the well-performing virtual functions, the scheduler module circuitry 230 schedules the underperforming virtual functions. This allows for configurable laxity for underperforming virtual functions to accommodate more graceful handling of abnormal situations where at least one well-performing virtual function may be underperforming due to an anomaly. In some embodiments, the scheduler module circuitry 230 maintains more than two lists to categorize virtual functions more finely than just well-performing and underperforming. For example, in some embodiments, the scheduler module circuitry 230 maintains a list of, for example, very well-performing virtual functions, very poorly-performing virtual functions, occasionally poorly-performing virtual functions, occasionally well-performing virtual functions, and the like.

[0050] In some embodiments, when a particular virtual function is on the underperforming list, as a penalty, the particular virtual function will not be scheduled at all for a time period. For example, if a virtual function overused its share in a past scheduling cycle, the virtual function must wait until the overuse is deducted from the subsequent scheduling cycle and is eventually granted a time share again in a future cycle during which the virtual function can be rescheduled. In some embodiments, the classification of a virtual function as underperforming is reset after the virtual function has been penalized to provide tolerance for anomalies or changes in use cases at runtime. The timing and conditions for the classification reset are configurable. If a virtual function continues to be classified as underperforming for an extended period of time, the virtual function may eventually be prohibited from submitting jobs. In such cases, the scheduler module circuit 230 bypasses submitting jobs for the underperforming virtual function during the current scheduling cycle.

[0051] The scheduling period is the sum of the per-virtual function time partitions assigned to each virtual function. Figure 5As shown, a first scheduling period 510 is shown in which further jobs submitted by virtual machines 220-223 are expected to be executed. First scheduling period 510 is immediately followed by scheduling period 520, which is immediately followed by scheduling period 530, and which is immediately followed by scheduling period 540. Although four (4) scheduling periods are shown, this is shown for ease of explanation, and it should be understood that the defined scheduling periods continue to repeat as long as virtual machines 220-223 continue to submit further jobs for execution.

[0052] In some embodiments, the scheduler module circuitry 230 adds additional time or slack time to the scheduling period. The purpose of the slack time within a period is to incorporate expected variations in the submission behavior of different virtual functions, which may be unavoidable for a use case. The slack time extends the size of the scheduling period so that jobs that are executed near the end of the scheduling period have time to execute after the scheduling period. The scheduler module circuitry 230 thereby prevents jobs that have not completed execution at the end of the scheduling period from being incorrectly classified as poorly performing, thereby providing flexibility to the scheduling period. Figure 5 As shown, the first scheduling period 510 includes additional slack time such that additional time 519 remains after the job 511 ends and before the first scheduling period 510 ends. Since the job 511 may not always start and end at the same time within the scheduling period, the additional time 519 provides flexibility for such instances to avoid classifying the job 511 as underperforming. For example, if the job 511 is submitted near the end of the scheduling period, a portion of the job 511 may be executed during the scheduling period, and the remainder of the job 511 may be executed during the next scheduling period. In such cases, the job 511 uses only a portion of its allocated time in the scheduling period and has surplus time that can be carried forward to the next (i.e., subsequent) scheduling period. In some embodiments, the surplus time is not further carried forward to the next scheduling period to prevent the virtual function from accumulating a large surplus.

[0053] Similarly, the second scheduling cycle 520 includes slack time such that additional time 529 remains after job 521 ends and before the second scheduling cycle 520 ends; the third scheduling cycle 530 includes slack time such that additional time 539 remains after job 531 ends and before the third scheduling cycle 530 ends; and the fourth scheduling cycle 540 includes slack time such that additional time 549 remains after job 541 ends and before the fourth scheduling cycle 540 ends.

[0054] For example, the scheduler module circuit 230 limits the slack time to 33.3ms (two cycles of 60fps, i.e., 2*16.67ms) - 24ms (the expected usage portion of 33.3ms, i.e., 4 VFs*3ms*2 jobs per VF within 33.3ms). In some cases, a virtual function may not submit a job on time within 16.67ms (e.g., job preparation is delayed), but instead submits 2 jobs within the next 16.67ms (e.g., the virtual function tries to catch up in order to still achieve 60fps). As the expected variation between jobs increases, n increases and / or the slack time increases. In some embodiments, the scheduler module circuit 230 supports dynamic reconfiguration of per-VF behavior + algorithm parameters (e.g., slack, etc.).

[0055] In some embodiments, the scheduler module circuitry 230 may issue a single job size credit to at least one of the virtual functions 210-213. If any of the virtual functions 210-213 has an unused job size (submits a job smaller than the expected job size) during one of the scheduling cycles 510-540, then that particular one of the virtual functions 210-213 receives a single job size credit during the scheduling cycle immediately preceding the assigned job size credit. For example, if the virtual function 210 has an unused job size during the current scheduling cycle 510, then that virtual function will receive a single job size credit during the next scheduling cycle 520 relative to the current scheduling cycle 510. Thus, if a job of a particular virtual function is delayed, that particular virtual function is permitted to submit two jobs during the next scheduling cycle. In some embodiments, a particular virtual function is given only a single job size credit so as not to cause excessive interference to other virtual functions. In some embodiments, the job size credit is configurable based on a job size credit carryover range (a number of scheduling periods, eg, 0 (job size credit disabled), 1, 2, etc.).

[0056] In some embodiments, if the GPU is idle, then after completing the expected number of jobs in a particular scheduling cycle, a well-behaved virtual function with time remaining in that particular scheduling cycle is given a one-time exception to run additional jobs submitted in that particular scheduling cycle. This accommodates scenarios where a well-behaved virtual function that has completed one or more jobs in a scheduling cycle still has time remaining in that scheduling cycle that could be used to complete the job if granted an exception. In some embodiments, this one-time adaptation is not allowed in the next scheduling cycle or in the next X scheduling cycles, where X is configurable, and if repeated, the virtual function is determined by the scheduler module circuitry 230 to be underperforming. In some embodiments, this exception is adjustable or can be disabled, depending on, for example, GPU utilization within the scheduling cycle or recent history.

[0057] Figure 6 is a flow chart of a method 600 for scheduling a first virtual function and a second virtual function in a plurality of virtual functions according to some embodiments. In some embodiments, the method 600 is performed by a processing system such as Figure 1 At block 604, a plurality of time partitions are assigned to a plurality of virtual functions to execute a plurality of jobs.

[0058] At block 606, the job cadence monitor 310 determines whether the first plurality of jobs submitted by the first virtual function are within the expected cadence and whether the second plurality of jobs submitted by the second virtual function are exceeding the expected cadence. If, at block 606, the job cadence monitor 310 determines that the plurality of jobs 513, 514, 523, 524, 533, 534, 543, 544 submitted by the virtual function 212 are exceeding the expected cadence, and the job cadence monitor 310 determines that the plurality of jobs 511, 521, 531, 541 submitted by the virtual function 210 are within the expected cadence, and the plurality of jobs 512, 522, 532, 542 submitted by the virtual function 211 are within the expected cadence, then the method flow proceeds to block 610. If, at block 606, the job cadence monitor 310 determines that none of the jobs submitted by the particular virtual function are exceeding the expected cadence, then the method flow proceeds from block 606 to block 608.

[0059] At block 608, the job size monitor 320 determines whether the first plurality of jobs submitted by the first virtual function do not take (or do take) longer to execute than the expected job size, and whether the second plurality of jobs submitted by the second virtual function do take (or do not take) longer to execute than the expected job size. Figure 3 and Figure 5, the job size monitor 320 determines that the plurality of jobs 515, 525, 535, 545 submitted by the virtual function 213 take longer to execute than the expected job size, and the plurality of jobs 511, 521, 531, 541 submitted by the virtual function 210 and the plurality of jobs 512, 522, 532, 542 submitted by the virtual function 211 do not take longer to execute than the expected job size. Although the virtual function 212 is shown as submitting jobs 513, 514, 523, 524, 533, 534, 543, 544 that exceed the expected cadence, and the virtual function 213 is shown as submitting jobs 515, 525, 535, 545 that take longer to execute than the expected job size, this is shown for ease of explanation. In some embodiments, a virtual function (not shown) may submit a job that is a combination of exceeding an expected cadence (e.g., greater fps than expected) and taking longer than expected job size (e.g., higher resolution than expected) and is scheduled based on the virtual function being determined to be underperforming.

[0060] Note that the job size monitor 320 does not need to determine whether the jobs 513, 514, 523, 524, 533, 534, 543, 544 submitted by the virtual function 212 take longer to execute than the expected job size because block 606 has already determined that the virtual function 212 is an underperforming virtual function and block 610 has taken appropriate corrective action with respect to the virtual function 212. If the job size monitor 320 determines that any of the jobs submitted by the particular virtual function take longer to execute than the expected job size, block 606 proceeds to block 610. Otherwise, if the job size monitor 320 determines that none of the jobs submitted by the particular virtual function take longer to execute than the expected job size, block 608 proceeds to block 606, causing the method 600 to continue monitoring the expected cadence and expected job size of the virtual functions 210-213 through blocks 606 and 608, respectively.

[0061] Although not shown, the method 600 may end for one of the plurality of virtual functions and continue for the remaining virtual functions. For example, in the context of cloud gaming, stopping a game executing on the cloud gaming server will also cause the virtual functions 210-213 to stop submitting jobs to the GPU 115. Thus, the method 600 will stop for that game, but continue for any remaining games executed by the GPU 115 that are still receiving jobs from the remaining virtual functions 210-213, as well as any newly added games, thereby continuing to determine whether any virtual function is submitting jobs that exceed the expected cadence and take longer to execute than the expected job size. The method 600 may also include any of the functions described above for the scheduler module circuit 230, the job cadence monitor 310, and / or the job size monitor 320.

[0062] In some embodiments, the apparatus and techniques described above are implemented in a system (such as the one described above with reference to FIG) that includes one or more integrated circuit (IC) devices (also referred to as integrated circuit packages or microchips). Figures 1 to 6 The processing system 100 described herein is implemented in a computer system. Electronic design automation (EDA) and computer-aided design (CAD) software tools can be used in the design and manufacture of these IC devices. These design tools are typically represented by one or more software programs. One or more software programs include code that can be executed by a computer system to manipulate the computer system to operate on the code representing the circuit of one or more IC devices in order to perform at least a portion of the process for designing or adjusting the manufacturing system to manufacture the circuit. The code may include instructions, data, or a combination of instructions and data. The software instructions representing the design tool or manufacturing tool are typically stored in a computer-readable storage medium accessible to the computing system. Similarly, the code representing one or more stages of the design or manufacture of the IC device can be stored in and accessed from the same computer-readable storage medium or different computer-readable storage media.

[0063] Computer-readable storage media may include any non-transitory storage medium or combination of non-transitory storage media that can be accessed by a computer system during use to provide instructions and / or data to the computer system. Such storage media may include, but are not limited to, optical media (e.g., compact discs (CDs), digital versatile discs (DVDs), Blu-ray discs), magnetic media (e.g., floppy disks, magnetic tapes, or magnetic hard drives), volatile memory (e.g., random access memory (RAM) or cache), non-volatile memory (e.g., read-only memory (ROM) or flash memory), or microelectromechanical system (MEMS)-based storage media. Computer-readable storage media may be embedded in a computing system (e.g., system RAM or ROM), fixedly attached to a computing system (e.g., a magnetic hard drive), removably attached to a computing system (e.g., an optical disc or flash memory based on a universal serial bus (USB)), or coupled to a computer system via a wired or wireless network (e.g., a network accessible storage device (NAS)).

[0064] In some embodiments, certain aspects of the technology described above can be implemented by one or more processors of a processing system that executes software. The software includes one or more sets of executable instructions that are stored or otherwise tangibly embodied on a non-transitory computer-readable storage medium. The software may include instructions and certain data that, when executed by one or more processors, manipulate the one or more processors to perform one or more aspects of the technology described above. The non-transitory computer-readable storage medium may include, for example, a magnetic or optical disk storage device, a solid-state storage device such as a flash memory, a cache, a random access memory (RAM), or other one or more non-volatile memory devices. The executable instructions stored on the non-transitory computer-readable storage medium may be source code, assembly language code, object code, or other instruction formats that are interpreted or otherwise executed by one or more processors.

[0065] It should be noted that not all activities or elements described above in the general description are required, a particular activity or part of the device may not be required, and one or more additional activities may be performed, or elements may be included in addition to those described. Furthermore, the order in which the activities are listed is not necessarily the order in which they are performed. In addition, these concepts have been described with reference to specific embodiments. However, it is understood by those skilled in the art that various modifications and changes may be made without departing from the scope of the present disclosure as set forth in the following claims. Accordingly, the specification and drawings are to be regarded as illustrative rather than restrictive, and all such modifications are intended to be included within the scope of the present disclosure.

[0066] Benefits, other advantages and solutions to problems have been described above with respect to specific embodiments. However, benefits, advantages, solutions to problems, and any features that may cause any benefit, advantage, or solution to appear or become more pronounced should not be construed as key, required, or essential features of any or all of the claims. Furthermore, the specific embodiments disclosed above are merely illustrative, as the disclosed subject matter may be modified and practiced in different but equivalent manners that will be apparent to those skilled in the art having the benefit of the teachings herein. No limitation is intended to the details of construction or design shown herein, except as described in the claims below. It is therefore apparent that the specific embodiments disclosed above may be changed or modified, and all such variations are considered to be within the scope of the disclosed subject matter. Accordingly, the protection sought herein is as set forth in the claims below.

Claims

1. A method comprising: allocating a plurality of time partitions within a scheduling period to a plurality of virtual functions for executing jobs at the parallel processors, wherein the scheduling period includes slack time to allow for variation in at least one of a pace of jobs submitted for execution by each virtual function and a size of the jobs submitted for execution; as well as Execution of a job that exceeds a time partition allocated for a virtual function of the plurality of virtual functions is prevented.

2. The method of claim 1 , wherein preventing execution comprises: In response to submission of the first plurality of jobs exceeding an expected cadence and submission of a second plurality of jobs by a second virtual function not exceeding the expected cadence, scheduling the first virtual function after the second virtual function to execute the first plurality of jobs.

3. The method according to claim 1, further comprising: In response to the first plurality of jobs exceeding an expected job size and a second plurality of jobs submitted by a second virtual function not exceeding the expected job size, the first virtual function is scheduled to execute the first plurality of jobs after the second virtual function.

4. The method according to claim 1, further comprising: In a case where the size of a job submitted by a first virtual function is less than an expected job size, a job size credit is assigned to the first virtual function, wherein the job size credit is usable by the first virtual function in a subsequent scheduling period immediately following the scheduling period during which the job size credit was assigned.

5. The method according to claim 1, further comprising: maintaining a first level list and a second level list, the first level list including first virtual functions that submitted a first plurality of jobs that were within an expected cadence and did not take longer than an expected job size to execute, and the second level list including second virtual functions that were not included in the first level list; as well as The first plurality of jobs for the virtual functions in the first level list are scheduled before the second plurality of jobs for the virtual functions in the second level list are scheduled.

6. The method according to claim 5, further comprising: In a current scheduling cycle, the second plurality of jobs of the virtual function in the second-level list are bypassed for scheduling.

7. The method according to claim 1, wherein the scheduling period is a first scheduling period among a plurality of scheduling periods, the method further comprising: The first virtual function is scheduled after the second virtual function if a first number of jobs submitted by the first virtual function during the plurality of scheduling periods is greater than a second number of jobs submitted by the second virtual function during the plurality of scheduling periods.

8. The method according to any one of claims 1 to 7, further comprising: In the case that the parallel processor is idle, a virtual function having remaining time in the allocated time partition and not completing a job within the scheduling period is allowed to submit the job.

9. A processing system, comprising: Parallel processors; and A scheduler module circuit, wherein the scheduler module circuit is configured to: allocating a plurality of time partitions within a scheduling period to a plurality of virtual functions for executing jobs at the parallel processor, wherein the scheduling period includes slack time to allow for variation in at least one of a pace of jobs submitted for execution by each virtual function and a size of the jobs submitted for execution; as well as Execution of a job that exceeds a time partition allocated for a first virtual function of the plurality of virtual functions is prevented.

10. The processing system of claim 9, wherein the scheduler module circuit is further configured to: In response to submission of a first plurality of jobs by the first virtual function exceeding an expected cadence and submission of a second plurality of jobs by a second virtual function not exceeding the expected cadence, scheduling the first virtual function after the second virtual function.

11. The processing system of claim 9, wherein the scheduler module circuit is further configured to: In response to a first plurality of jobs submitted by the first virtual function exceeding an expected job size and a second plurality of jobs submitted by a second virtual function not exceeding the expected job size, scheduling the first virtual function after the second virtual function.

12. The processing system of claim 9, wherein the scheduler module circuit is further configured to: In a case where the size of a job submitted by a virtual function is less than an expected job size, a job size credit is assigned to the virtual function, wherein the job size credit is usable by the virtual function in a subsequent scheduling cycle immediately following the scheduling cycle during which the job size credit was assigned.

13. The processing system of claim 9, wherein the scheduler module circuit is further configured to: maintaining a first level list and a second level list, the first level list including the first virtual functions that submitted a first plurality of jobs that were within an expected cadence and did not take longer than an expected job size to execute, and the second level list including second virtual functions that were not included in the first level list; The first plurality of jobs for the virtual functions in the first level list are scheduled before the second plurality of jobs for the virtual functions in the second level list are scheduled.

14. The processing system of claim 13 , wherein the scheduler module circuit is further configured to: During a current scheduling cycle, scheduling of the second plurality of jobs of the second virtual function is bypassed.

15. The processing system of claim 9, wherein the scheduling period is a first scheduling period among a plurality of scheduling periods, the scheduler module circuit being further configured to: The first virtual function is scheduled after the second virtual function if a first number of jobs submitted by the first virtual function during the plurality of scheduling periods is greater than a second number of jobs submitted by a second virtual function during the plurality of scheduling periods.

16. The processing system of any one of claims 9 to 15, wherein the scheduler module circuit is further configured to: In the case that the parallel processor is idle, a virtual function having remaining time in the allocated time partition and not completing a job within the scheduling period is allowed to submit the job.

17. A server, comprising: a parallel processor configured to execute jobs submitted by a plurality of virtual functions; and A scheduler module circuit, wherein the scheduler module circuit is configured to: allocating a time partition to each of the plurality of virtual functions within a scheduling period based on expected job size and cadence and slack time to allow for variations in submitted job size and cadence; and Jobs that exceed a time partition allocated for a virtual function of the plurality of virtual functions are prevented from executing at the parallel processor.

18. The server of claim 17, wherein the scheduler module circuit is further configured to: In response to a first plurality of jobs submitted by a first virtual function having a frequency below an expected cadence and a second plurality of jobs submitted by a second virtual function exceeding the expected cadence, scheduling the first virtual function before the second virtual function.

19. The server of claim 17, wherein the scheduler module circuit is further configured to: In response to a first plurality of jobs submitted by a first virtual function not exceeding an expected job size and a second plurality of jobs submitted by a second virtual function exceeding the expected job size, scheduling the first virtual function before the second virtual function.

20. The server of claim 17, wherein the scheduler module circuit is further configured to: maintaining a first level list and a second level list, the first level list including first virtual functions that submitted a first plurality of jobs that were within an expected cadence and did not take longer than an expected job size to execute, and the second level list including second virtual functions that were not included in the first level list; The first plurality of jobs for the first virtual function in the first level list are scheduled before the second plurality of jobs for the second virtual function in the second level list.