Method, system, and medium for configuring a processor to function as multiple independent processors

By partitioning the PPU into isolated logical partitions, the problems of uneven resource utilization and interference in multi-tenant environments in multi-CPU process processing are solved, and more efficient resource utilization and stable multi-tenant support are achieved.

CN112445610BActive Publication Date: 2025-07-29NVIDIA CORP
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202010213365.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2019-09-05
Filing Date
2020-03-24
Publication Date
2025-07-29
Estimated Expiration
2040-04-28

AI Technical Summary

Technical Problem

Traditional GPUs use unevenly when processing multiple CPU processes, resulting in some processes being stagnant and resources being idle. There is interference between different processing subcontexts in a multi-tenant environment, affecting GPU performance and multi-tenant support capabilities.

Method used

The parallel processing unit (PPU) is partitioned into isolated logical partitions, each partition contains a subset of hardware resources, and multiple engines are generated within each partition to support multiple independent processing contexts to ensure isolation and uniform utilization of resources.

Benefits of technology

It realizes resource isolation between multiple processing contexts of PPU, improves the resource utilization rate of GPU and the stability of multi-tenant environment, and is suitable for cloud-based deployment.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN112445610B_ABST
    Figure CN112445610B_ABST
Patent Text Reader

Abstract

A parallel processing unit (PPU) can be divided into multiple partitions. Each partition is configured to perform operations similar to the operations of the entire PPU. A given partition includes a subset of the computing and storage resources associated with the entire PPU. Software executed on a CPU will partition the PPU for an administrator user. A guest user is assigned to a partition, and the guest user can perform processing tasks within that partition isolated from any other guest users assigned to any other partition. Because the PPU can be divided into isolated partitions, multiple CPU processes can effectively utilize the PPU resources.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Various embodiments generally relate to parallel processing architectures, and more particularly, to techniques for configuring a processor to function as multiple independent processors. Background Art

[0002] A traditional central processing unit (CPU) typically includes a relatively small number of processing cores that can execute a relatively small number of CPU processes. In contrast, a traditional graphics processing unit (GPU) typically includes hundreds of processing cores that can execute hundreds of threads in parallel with each other. Thus, given the large amount of processing resources that can be deployed when using a traditional GPU, a traditional GPU can generally execute certain processing tasks faster and more efficiently than a traditional CPU.

[0003] In some embodiments, a CPU process executing on a CPU can offload a given processing task to a GPU in order to have the processing task executed more quickly. To do so, the CPU process generates a processing context on the GPU that specifies the target state of the various GPU resources that are to be implemented to execute the processing task. Those GPU resources can include processing, graphics, and memory resources, among others. The CPU process then launches a thread group on the GPU according to the processing context, and the thread group utilizes the various GPU resources to execute the processing task. In many of these types of implementations, the GPU is configured according to only one processing context at a time. However, in some cases, the CPU needs to offload more than one CPU process to the GPU within the same time interval. In such cases, the CPU can dynamically change the processing context implemented on the GPU at different points in time in order to serially service these CPU processes during a given time interval. However, a drawback of this approach is that the processing tasks offloaded by certain CPU processes do not fully utilize the resources of the GPU. Thus, when one or more processing tasks associated with these CPU processes are serially executed on the GPU, certain GPU resources may be left idle, which reduces overall GPU performance and utilization.

[0004] One method for simultaneously executing multiple CPU processes on a GPU is to generate multiple different processing sub-contexts within a given "parent" processing context and assign each different processing sub-context to a different CPU process. Then, the multiple CPU processes can simultaneously launch different thread groups on the GPU, where each thread group utilizes specific GPU resources configured according to a specific processing sub-context. By this method, the GPU can be utilized more effectively because more than one CPU process can offload processing tasks to the GPU at the same point in time, thereby potentially avoiding situations where certain GPU resources are left idle.

[0005] One problem with the above method is that CPU processes associated with different processing sub - contexts may unfairly consume GPU resources, which should be more evenly allocated or distributed across different processing sub - contexts. For example, a first CPU process may start a first thread group in a first processing sub - context that executes a large number of read requests and consumes a large amount of available GPU memory bandwidth. A second CPU process may then start a second thread group in a second processing sub - context that also executes a large number of read requests. However, since the first thread group has already consumed much of the available GPU memory bandwidth, the second thread set may experience high latency, which may cause the second CPU process to stall.

[0006] Another problem with the above method is that because processing sub - contexts share a parent context, any error that occurs when a thread associated with one processing sub - context executes may interfere with the execution of other threads associated with another processing sub - context that shares the same parent context. For example, a first CPU process may start a first thread group associated with a first processing sub - context to perform a first processing task. A second CPU process may start a second thread group associated with a second processing sub - context, and the second thread group may then malfunction and fail. To recover from the failure, the GPU will have to reset the parent context, which will automatically reset the first processing sub - context and the second processing sub - context. In this case, even though the failure was caused by the second thread group and not the first thread group, the execution of the first thread group will be interrupted.

[0007] As previously mentioned, there is a need in the art for more efficient techniques for configuring a GPU to perform processing tasks associated with multiple contexts. Summary of the Invention

[0008] Various embodiments include a computer - implemented method that includes: partitioning a set of hardware resources included in a processor to generate a first logical partition that includes a first subset of the hardware resources; and generating a plurality of engines within the first logical partition, wherein each engine included in the plurality of engines is assigned a different portion of the first subset of the hardware resources and executes in functional isolation from all other engines included in the plurality of engines.

[0009] A technical advantage of the disclosed technology over the prior art is that, using the disclosed technology, a parallel processing unit (PPU) (e.g., a GPU) can support multiple contexts simultaneously and be functionally isolated from each other. Thus, multiple CPU processes can effectively utilize PPU resources by executing multiple different contexts simultaneously without interfering with each other. Brief Description of the Drawings

[0010] In order to understand the manner in which the above-described features of the various embodiments can be detailed, the inventive concept briefly outlined above can be described in more specific terms by reference to the various embodiments, some of which are illustrated in the accompanying drawings. However, it should be noted that the drawings only show typical embodiments of the inventive concept and should not be regarded as limiting the scope in any way, and there are other equivalent embodiments.

[0011] Figure 1 is a block diagram of a computer system configured to implement one or more aspects of the various embodiments;

[0012] Figure 2 is included in accordance with the various embodiments Figure 1 in the parallel processing subsystem of the parallel processing unit (PPU) block diagram;

[0013] Figure 3 is included in accordance with the various embodiments Figure 2 in the parallel processing unit of the general processing cluster block diagram;

[0014] Figure 4 is included in accordance with the various embodiments Figure 2 in the PPU of the partitioning unit block diagram;

[0015] Figure 5 is included in accordance with the various embodiments Figure 2 in the PPU of the various PPU resources block diagram;

[0016] Figure 6 is in accordance with the various embodiments Figure 1 of the hypervisor how the PPU resources are logically grouped into PPU partition sets example;

[0017] Figure 7 shows in accordance with the various embodiments Figure 1 of the hypervisor how to configure the PPU partition sets to implement one or more simultaneous multi-context (SMC) engines;

[0018] Figure 8A is in accordance with the various embodiments Figure 7 of the DRAM more detailed illustration;

[0019] Figure 8B shows in accordance with the various embodiments of how to address Figure 8B each DRAM portion;

[0020] Figure 9 is in accordance with the various embodiments Figure 1 of the hypervisor how to partition and configure the PPU data flow diagram;

[0021] Figure 10is a flowchart of method steps for partitioning and configuring a PPU on behalf of one or more users according to various embodiments;

[0022] Figure 11 illustrates a partition configuration table according to which a hypervisor can configure one or more PPU partitions according to various embodiments; Figure 1 of the hypervisor can configure one or more PPU partitions;

[0023] Figure 12 illustrates how a hypervisor according to various embodiments partitions a PPU to generate one or more PPU partitions; Figure 1 of the hypervisor partitions the PPU to generate one or more PPU partitions;

[0024] Figure 13 illustrates how a hypervisor according to various embodiments allocates various PPU resources during partitioning; Figure 1 of the hypervisor allocates various PPU resources during partitioning;

[0025] Figure 14A illustrates how multiple guest operating systems running multiple VMs simultaneously start multiple processing contexts within one or more PPU partitions according to various embodiments;

[0026] Figure 14B illustrates how a host operating system simultaneously starts multiple processing environments within one or more PPU partitions according to various embodiments;

[0027] Figure 15 illustrates how a hypervisor according to various embodiments assigns virtual address space identifiers to different SMC engines; Figure 1 of the hypervisor assigns virtual address space identifiers to different SMC engines;

[0028] Figure 16 illustrates how a memory management unit converts local virtual address space identifiers during fault mitigation according to various embodiments;

[0029] Figure 17 illustrates how a hypervisor according to various embodiments implements soft ownership when migrating a processing context between SMC engines on different PPUs; Figure 1 of the hypervisor implements soft ownership;

[0030] Figure 18 is a flowchart of method steps for configuring computing resources within a PPU to simultaneously support operations associated with multiple processing contexts according to various embodiments;

[0031] Figure 19 illustrates a set of boundary options according to which a hypervisor can generate one or more PPU memory partitions according to various embodiments; Figure 1 of the hypervisor can generate one or more PPU memory partitions;

[0032] Figure 20 illustrates according to various embodiments;Figure 1 An example of how the management program partitions the PPU memory to generate one or more PPU memory partitions;

[0033] Figure 21 Illustrates how, according to various embodiments, Figure 16 the memory management unit provides access to different PPU memory partitions;

[0034] Figure 22 Illustrates how, according to various embodiments, Figure 16 the memory management unit performs various address translations;

[0035] Figure 23 Illustrates how, according to various embodiments, Figure 16 the memory management unit simultaneously provides support operations associated with multiple processing contexts;

[0036] Figure 24 Is a flowchart of method steps for configuring memory resources within a PPU to simultaneously support operations associated with multiple processing contexts according to various embodiments;

[0037] Figure 25 Is, according to various embodiments, an illustration of a set of timelines of VM-level time slices associated with Figure 2 the PPU;

[0038] Figure 26 Is, according to various other embodiments, an illustration of another set of timelines of VM-level time slices associated with Figure 2 the PPU;

[0039] Figure 27 Is, according to various embodiments, an illustration of a timeline of SMC-level time slices associated with Figure 2 the PPU;

[0040] Figure 28 Illustrates how, according to various embodiments, a VM migrates from one PPU to another PPU;

[0041] Figure 29 Is, according to various embodiments, an illustration of a set of timelines of fine-grained VM migration associated with Figure 2 the PPU;

[0042] Figure 30A - 30B Illustrates how, according to various embodiments, a method step flowchart for time slicing a VM in Figure 2 the PPU;

[0043] Figure 31 Is a memory map according to various embodiments, which illustrates how the BAR0 address space is mapped to Figure 2The privileged register space within the PPU;

[0044] Figure 32 is a flowchart of method steps for addressing a privileged register address space in a Figure 2 PPU according to various embodiments;

[0045] Figure 33 is for Figure 2 a block diagram of a performance monitoring system of a PPU according to various embodiments;

[0046] Figure 34A - 34B shows various configurations of a Figure 33 performance multiplexer unit according to various embodiments;

[0047] Figure 35 is for Figure 2 a block diagram of a performance monitor aggregation system of a PPU according to various embodiments;

[0048] Figure 36 shows the format of a trigger packet associated with a Figure 35 performance monitor aggregation system according to various embodiments;

[0049] Figure 37 is a flowchart of method steps for monitoring the Figure 2 performance of a PPU according to various embodiments;

[0050] Figure 38 is for Figure 2 a block diagram of a power and clock frequency management system of a PPU according to various embodiments; and

[0051] Figure 39 is a flowchart of method steps for managing the Figure 2 power consumption of a PPU 200 according to various embodiments. DETAILED DESCRIPTION

[0052] In the following description, numerous specific details are set forth to provide a more thorough understanding of the various embodiments. However, it will be apparent to one of ordinary skill in the art that the inventive concept may be practiced without one or more of these specific details.

[0053] As described above, traditional GPUs can generally execute certain processing tasks faster than traditional CPUs. In some configurations, a CPU process executing on a CPU can offload a given processing task to a GPU for faster execution of the processing task. In this way, the CPU process generates a processing context on the GPU that specifies the target state of various GPU resources and then launches a thread group on the GPU to execute the processing task.

[0054] In some cases, multiple CPU processes may need to offload processing tasks to a GPU during the same time interval. However, a GPU can only be configured according to one processing context at a time. In such cases, the CPU can dynamically change the processing context of the GPU at different time points to continuously serve multiple CPU processes across the time interval. However, some CPU processes may not fully utilize the GPU resources when executing processing tasks, sometimes leaving various GPU resources idle. To address this issue, the CPU can generate multiple processing sub-contexts within a "parent" processing context and assign these processing sub-contexts to different CPU processes. Then, these CPU processes can simultaneously launch different thread groups on the GPU, and each thread group can utilize specific GPU resources configured according to a specific processing sub-context. This method can be implemented to more effectively utilize GPU resources. However, this method has several drawbacks.

[0055] First, the CPU processes associated with different processing sub-contexts consume the GPU resources that should be fairly shared among different processing sub-contexts unfairly, resulting in a situation where one CPU process may stall the progress of another CPU process. Second, since the processing sub-contexts share the parent processing context, any error that occurs during the execution of a thread associated with one processing sub-context may disrupt the execution of threads associated with other processing sub-contexts included in the same parent processing context. In some cases, a failure that occurs in one processing sub-context may cause all other processing sub-contexts in the same parent processing context to be reset and restarted.

[0056] Generally, the above-mentioned drawbacks associated with processing sub-contexts limit the extent to which traditional GPUs can support multi-tenancy. As referred to herein, "multi-tenancy" refers to a GPU configuration in which multiple users or "tenants" use GPU resources to perform processing operations simultaneously or during overlapping time intervals. Generally, traditional GPUs provide support for multi-tenancy by allowing different tenants to execute different processing tasks using different processing sub-contexts within a given parent processing context. However, processing sub-contexts are not isolated computing environments because, for the various reasons mentioned above, the processing tasks executed in different processing sub-contexts may interfere with each other. Therefore, any given tenant occupying a given GPU will have a negative impact on the quality of service provided by the GPU to other tenants. These factors may reduce the attractiveness of cloud-based GPU deployments, in which multiple users may access the same GPU simultaneously.

[0057] To address these issues, various embodiments include a parallel processing unit (PPU) that can be partitioned. Each partition is configured to concurrently execute processing tasks associated with multiple processing environments. A given partition includes a logical grouping or "slice" of one or more GPU resources. Each slice provides sufficient computing, graphics, and memory resources to emulate the operation of the entire PPU. A hypervisor executing on a CPU performs various techniques on behalf of an administrator user to partition the PPU. A guest user is assigned to a partition and can then execute processing tasks isolated from any other guest users assigned to any other partition within that partition.

[0058] Relative to the prior art, one technical advantage of the disclosed technology is that, with the disclosed technology, the PPU can support multiple processing contexts simultaneously and functionally isolated from each other. Thus, multiple CPU processes can effectively utilize the PPU resources via multiple different processing contexts without interfering with each other. Another technical advantage of the disclosed technology is that, since the PPU can be partitioned into isolated computing environments using the disclosed technology, the PPU can support a more robust form of multi-tenancy relative to prior art methods that rely on processing sub-contexts to provide multi-tenancy functionality. Thus, when the disclosed technology is implemented, the PPU becomes more suitable for cloud-based deployments in which access to different partitions within the same PPU can be provided to different and potentially competing entities. These technical advantages represent one or more technological advancements over prior art methods.

[0059] System Overview

[0060] Figure 1 is a block diagram of a computer system configured to implement one or more aspects of the present invention. As shown, computer system 100 includes a central processing unit (CPU) 110, a system memory 120, and a parallel processing subsystem 130 coupled together via a memory bridge 132. The parallel processing subsystem 130 is coupled to the memory bridge 132 via a communication path 134. One or more display devices 136 may be coupled to the parallel processing subsystem 130. Computer system 100 also includes a system disk 140, one or more additional cards 150, and a network adapter 160. The system disk 140 is coupled to an I / O bridge 142. The I / O bridge 142 is coupled to the memory bridge 132 via a communication path 144 and is also coupled to an input device 146. The one or more additional cards 150 and the network adapter 160 are coupled together via a switch 148, which in turn is coupled to the I / O bridge 142.

[0061] Memory bridge 132 is a hardware unit that facilitates communication between CPU 110, system memory 120, and parallel processing subsystem 130, as well as other components of computer system 100. For example, memory bridge 132 may be a northbridge chip. Communication path 134 is a high-speed and / or high-bandwidth data connection that facilitates low-latency communication between parallel processing subsystem 130 and memory bridge 132 across one or more independent channels. For example, communication path 134 may be a Peripheral Component Interconnect Express (PCIe) link, Accelerated Graphics Port (AGP), HyperTransport, or any other technically feasible communication bus type.

[0062] I / O bridge 142 is a hardware unit that facilitates input and / or output operations performed using system disk 140, input devices 146, one or more add-in cards 150, network adapter 160, and various other components of computer system 100. For example, I / O bridge 143 may be a southbridge chip. Communication path 144 is a high-speed and / or high-bandwidth data connection that facilitates low-latency communication between memory bridge 132 and I / O bridge 142. For example, communication path 144 may be a PCIe link, AGP, HyperTransport, or any other technically feasible communication bus type. Using the configuration shown, any component coupled to memory bridge 132 or I / O bridge 142 can communicate with any other component coupled to memory bridge 132 or I / O bridge 142.

[0063] The CPU 110 is a processor configured to coordinate the overall operation of the computer system 100. As such, the CPU 110 executes instructions to issue commands to various other components included in the computer system 100. The CPU 110 is also configured to execute instructions to process data generated and / or stored by any other components included in the computer system 100, including the system memory 120 and the system disk 140. The system memory 120 and the system disk 140 are memory devices that include computer-readable media configured to store data and software applications. The system memory 120 includes device drivers 122 and a hypervisor 124, the operation of which is described in more detail below. The parallel processing subsystem 130 includes one or more parallel processing units (PPUs) that are configured to perform multiple operations simultaneously in a highly parallel processing architecture. Each PPU includes one or more compute engines that perform general-purpose computing operations in parallel and / or one or more graphics engines that perform graphics-oriented operations in parallel. A given PPU can be configured to generate pixels for display via a display device 136. An exemplary PPU is described below in conjunction with Figures 2 - 4 Describe in more detail.

[0064] Device driver 122 is a software application that, when executed by CPU 110, operates as an interface between CPU 110 and parallel processing subsystem 130. In particular, device driver 122 allows CPU 110 to offload various processing operations to parallel processing subsystem 130 for highly parallel execution, including general computing operations as well as graphics processing operations. Hypervisor 124 is a software application that, when executed by CPU 110, partitions the various computing, graphics, and memory resources included in parallel processing subsystem 130 to provide separate users with independent use of these resources, which will be described in more detail below in conjunction with Figures 5 - 10 More detailed description.

[0065] In various embodiments, some or all of the components of computer system 100 may be implemented in a cloud-based environment that is potentially distributed across a vast geographical area. For example, the various components of computer system 100 may be deployed across geographically distinct data centers. In such embodiments, the various components of computer system 100 may communicate with each other via one or more networks, including any number of local intranets and / or the Internet. In various other embodiments, certain components of computer system 100 may be implemented via one or more virtualization devices. For example, CPU 110 may be implemented as a virtual instance of a hardware CPU. In some embodiments, part or all of parallel processing subsystem 130 may be integrated with one or more other components of computer system 100 to form a single chip, such as a system-on-chip (SoC).

[0066] Those skilled in the art will appreciate that the architecture of computer system 100 is flexible enough to be implemented across a vast range of potential scenarios and use cases. For example, computer system 100 may be implemented in a cloud computing center to expose general computing capabilities and / or general graphics processing capabilities to one or more users. Alternatively, computer system 100 may be deployed in an automotive implementation to perform data processing operations associated with vehicle navigation. Those skilled in the art will further appreciate that, without departing from the overall scope and spirit of the present embodiments, the various components of computer system 100 and the connection topology between these components may be modified in any technically feasible manner.

[0067] Figure 2 Is included in accordance with various embodiments in Figure 1Block diagram of the PPU in the parallel processing subsystem. As shown, PPU 200 includes an I / O unit 210, a host interface 220, a system (sys) pipeline 230, an array of processing clusters 240, a crossbar switch unit 250, and a memory interface 260. PPU 200 is coupled to PPU memory 270. Each of the components shown can be implemented by any technically feasible type of hardware and / or any technically feasible combination of hardware and software.

[0068] The I / O unit 210 is coupled to the Figure 1 CPU 110 via communication path 134 and memory bridge 132. The I / O unit 210 is also coupled to the host interface 220 and the crossbar switch unit 250. The host interface 220 is coupled to one or more physical copy engines (PCEs) 222, which in turn are coupled to one or more PCE counters 224. The host interface 220 is also coupled to the system pipeline 230. A given system pipeline 230 includes a front end 232, task / work units 234, and a performance monitor (PM) 236, and is coupled to the array of processing clusters 240. The array of processing clusters 240 includes general processing clusters (GPCs) 242(0) to 242(A), where A is a positive integer. The array of processing clusters 240 is coupled to the crossbar switch unit 250. The crossbar switch unit 250 is coupled to the memory interface 260. The memory interface 260 includes partition units 262(0) to 262(B), where B is a positive integer value. Each partition unit 262 can be separately connected to the crossbar switch unit 250. The PPU memory 270 includes dynamic random access memories (DRAMs) 272(0) to 272(C), where C is a positive integer value. To facilitate simultaneous operation on multiple processing contexts, the individual units within PPU 200 are replicated as follows: (a) the host interface 220 includes PBDMAs 520(0) to 520(7); (b) the system pipeline 230 includes system pipelines 230(0) to 230(7) such that the task / work units 234 correspond to SKEDs 500(0) to 500(7); and the task / work units 234 correspond to CWDs 560(0) to 560(7).

[0069] In operation, the I / O unit 210 obtains various types of command data from the CPU 110 and distributes the command data to the relevant components of the PPU 200 for execution. In particular, the I / O unit 210 obtains command data associated with a processing task from the CPU 110 and routes the command data to the host interface 220. The I / O unit 210 also obtains command data associated with a memory access operation from the CPU 110 and routes the command data to the crossbar unit 250. Command data related to a processing task typically includes one or more pointers to task metadata (TMD) that is stored in a command queue in the PPU memory 270 or elsewhere within the computer system 100. A given TMD is an encoded processing task that describes the index of the data to be processed, the operations to be performed on the data, the status parameters associated with those operations, the execution priority, and other processing task-oriented information.

[0070] The host interface 220 receives command data related to a processing task from the I / O unit 210 and then distributes the command data to the system pipeline 230 via one or more command streams. In some configurations, the host interface 220 generates different command streams for each different system pipeline 230, where a given command stream includes a pointer to the TMD related to the corresponding system pipeline 230.

[0071] A given system pipeline 230 performs various preprocessing operations on the received command data to facilitate the execution of the corresponding processing task on the GPCs 242 within the processing cluster array 240. When command data associated with one or more processing tasks is received, the front end 232 in a given system pipeline 230 obtains the associated processing tasks and relays the processing tasks to the task / work units 234. The task / work units 234 configure one or more GPCs 242 into an operational state suitable for executing the processing task and then transmit the processing tasks to those GPCs 242 for execution. Each system pipeline 230 can offload copy tasks to one or more PCEs 222 that perform dedicated copy operations. The PCE counter 224 tracks the usage of the PCEs 222 to balance the copy operation workload between different system pipelines 230. The PM 236 monitors the overall performance and / or resource consumption of the corresponding system pipeline 230 and can limit the various operations performed by that system pipeline 230 to maintain balanced resource consumption across all system pipelines 230.

[0072] Each GPC 242 includes multiple parallel processing cores that are capable of executing a large number of threads simultaneously and with any degree of independence and / or isolation from other GPCs 242. For example, a given GPC 242 can execute hundreds or thousands of concurrent threads in conjunction with or in isolation from any other GPC 242. Concurrent thread groups executing on a GPC 242 can execute separate instances of the same program or separate instances of different programs. In some configurations, GPCs 242 are shared among all system pipelines 230, while in other configurations, different groups of GPCs 242 are assigned to operate with specific system pipelines 230. Each GPC 242 receives processing tasks from one or more system pipelines 230 and, in response, launches one or more thread groups to perform those processing tasks and generate output data. After completing a given processing task, the given GPC 242 transmits the output data to another GPC 242 for further processing, or to the crossbar switch unit 250 for appropriate routing. Figure 3 An exemplary GPC is described in more detail.

[0073] The crossbar unit 250 is a switching mechanism that routes various types of data between the I / O unit 210, the processing cluster array 240, and the memory interface 260. As described above, the I / O unit 210 sends command data related to memory access operations to the crossbar unit 250. In response, the crossbar unit 250 submits the associated memory access operation to the memory interface 260 for processing. In some cases, the crossbar unit 250 also routes read data returned from the memory interface 260 back to the component that requested the read data. As described above, the crossbar unit 250 also receives output data from the GPCs 242 and can then route this output data to the I / O unit 210 for transmission to the CPU 110, or route this data to the memory interface 260 for storage and / or processing. The crossbar unit 250 is generally configured to route data between the GPCs 242 and from any GPC 242 to any partition unit 262. In various embodiments, the crossbar unit 250 can implement virtual channels to separate traffic flows between the GPCs 242 and the partition units 262. In various embodiments, crossbar unit 250 may allow for unshared paths between a set of GPCs 242 and a set of partition units 262 .

[0074] The memory interface 260 implements partition units 262 to provide high-bandwidth memory access to the DRAM 272 in the PPU memory 270. Each partition unit 262 can perform memory access operations in parallel with different DRAM 272s, thereby effectively utilizing the available memory bandwidth of the PPU memory 270. A given partition unit 262 also provides cache support via one or more internal caches. The exemplary partition unit 262 is described in more detail below in conjunction with Figure 4 and will be described in more detail below.

[0075] Generally, the PPU memory 270, and in particular the DRAM 272, can be configured to store any technically feasible data associated with general computing applications and / or graphics processing applications. For example, the DRAM 272 can store a large data value matrix associated with a neural network in a general computing application, or alternatively, one or more frame buffers including multiple render targets in a graphics processing application. In various embodiments, the DRAM 272 can be implemented via any technically feasible memory device.

[0076] The architecture described above allows the PPU 200 to perform a variety of processing operations in a fast manner and asynchronously with respect to the operation of the CPU 110. In particular, the parallel architecture of the PPU 200 allows a large number of operations to be performed in parallel and have any degree of independence from each other and from the operations performed on the CPU 110, thereby accelerating the overall performance of these operations.

[0077] In one embodiment, the PPU 200 can be configured to perform general computing operations to accelerate computations involving large data sets. Such data sets may involve financial time series, dynamic simulation data, real-time sensor readings, neural network weight matrices and / or tensors, and machine learning parameters, among others. In another embodiment, the PPU 200 can be configured to act as a graphics processing unit (GPU) that implements one or more graphics rendering pipelines to generate pixel data based on graphics commands generated by the CPU 110. The PPU 200 can then output the pixel data as one or more frames via the display device 136. The PPU memory 170 can be configured to act as a graphics memory that stores one or more frame buffers and / or one or more render targets in the manner described above. In yet another embodiment, the PPU 200 can be configured to perform general computing operations and graphics processing operations simultaneously. In such a configuration, one or more system pipelines 230 can be configured to implement general computing operations via one or more GPCs 242, and one or more other system pipelines 230 can be configured to implement one or more graphics processing pipelines via one or more GPCs 242.

[0078] For any of the above-described architectures, the device driver 122 and the hypervisor 124 interoperate to subdivide the various computing, graphics, and memory resources included in the PPU 200 into separate "PPU partitions". Alternatively, there may be multiple device drivers 122, each associated with a "PPU partition". Preferably, the device driver executes on a set of cores in the CPU 110. A given PPU partition operates in a manner generally similar to the PPU 200 as a whole. In particular, each PPU partition may be configured to perform general computing operations, graphics processing operations, or both types of operations in relative isolation from other PPU partitions. Additionally, a given PPU partition may be configured to implement multiple processing contexts while executing one or more virtual machines (VMs) on the computing, graphics, and memory resources allocated to the given PPU partition. The logical grouping of PPU resources into PPU partitions is described in more detail below in conjunction with Figure 5 - FIG. 8. The techniques for partitioning and configuring PPU resources are described in more detail below in conjunction with Figures 9 - 10 more detail.

[0079] Figure 3 is a block diagram of a GPC included in a Figure 2 PPU according to various embodiments of the present invention. As shown, the GPC 242 is coupled to a memory management unit (MMU) 300 and includes a pipeline manager 310, a work distribution crossbar 320, one or more texture processing clusters (TPCs) 330, one or more texture units 340, a level 1.5 (L1.5) cache 350, a PM 360, and a raster pre-operation processor (preROP) 370. The pipeline manager 310 is coupled to the work distribution crossbar 320 and the TPC 330. Each TPC 330 includes one or more streaming multiprocessors (SMs) 332 and is coupled to the texture unit 340, the MMU 300, the L1.5 cache 350, the PM 360, and the preROP 370. The texture unit 340 and the L1.5 cache 350 are also coupled to the MMU 300 and to each other. The preROP 370 is coupled to the work distribution crossbar 320. Each of the illustrated components may be implemented by any technically feasible type of hardware and / or any technically feasible combination of hardware and software.

[0080] The GPC 242 is configured with a highly parallel architecture that supports the parallel execution of a large number of threads. As mentioned herein, a "thread" is an instance of a particular program that executes on a particular set of input data to perform various types of operations (including general computing operations and graphics processing operations). In one embodiment, the GPC 242 may implement single instruction multiple data (SIMD) techniques to support the parallel execution of a large number of threads without relying on multiple independent instruction units.

[0081] In another embodiment, the GPC 242 can implement single instruction multiple thread (SIMT) technology to support parallel execution of a large number of generally synchronous threads via a common instruction unit that issues instructions to one or more processing engines. Those skilled in the art will understand that SIMT execution allows different threads to more easily follow divergent execution paths through a given program, which is different from SIMD execution in which all threads generally follow non-divergent execution paths through a given program. Those skilled in the art will recognize that SIMD technology represents a functional subset of SIMT technology.

[0082] The GPC 242 can execute a large number of parallel threads via the SM 332s included in the TPC 330. Each SM 332 includes a set of functional units (not shown), including one or more execution units and / or one or more load / store units, which are configured to execute instructions associated with the received processing tasks. A given functional unit can execute instructions in a pipelined manner, which means that instructions can be issued to the functional unit before the execution of a previous instruction is completed. In various embodiments, the functional units within the SM 332 can be configured to perform a variety of different operations, which include integer and floating-point arithmetic (e.g., addition and multiplication, etc.), comparison operations, boolean operations (e.g., AND, OR, and XOR, etc.), shift operations, and calculations of various algebraic functions (e.g., planar interpolation and trigonometric functions, exponential functions, and logarithmic functions, etc.). Each functional unit can store intermediate data in a level 1 (L1) cache resident in the SM 332.

[0083] Through the above functional units, the SM 332 is configured to process one or more "thread groups" (also referred to as "warps"), which execute the same program on different input data simultaneously. Each thread in a thread group generally executes via a different functional unit, although in some cases not all functional units execute threads. For example, if the number of threads included in a thread group is less than the number of functional units, the unused functional units may remain idle during the processing of the thread group. In other cases, multiple threads in a thread group execute via the same functional unit at different times. For example, if the number of threads included in a thread group is greater than the number of functional units, one or more functional units can execute different threads in consecutive clock cycles.

[0084] In one embodiment, a set of related thread groups can be active simultaneously in different execution phases within an SM 332. A set of related thread groups is referred to herein as a "cooperative thread array" (CTA) or "thread array". Threads within the same CTA or threads in different CTAs can typically share intermediate data and / or output data with each other via one or more L1 caches (including those of the SM 332, the L1.5 cache 350, one or more L2 caches shared between SM 332s), or via any shared memory, global memory, or other type of memory residing on any memory device included in the computer system 100. In one embodiment, the L1.5 cache 350 can be configured to cache instructions to be executed by threads on the SM 332.

[0085] Each thread within a given thread group or CTA is typically assigned a unique thread identifier (thread ID) that can be accessed during execution. The thread ID assigned to a given thread can be defined as a one-dimensional or multi-dimensional numerical value. The execution and processing behavior of a given thread can vary depending on the thread ID. For example, a thread can determine which portion of an input data set to process and / or which portion of an output data set to write based on the thread ID.

[0086] In one embodiment, a per-thread instruction sequence can include at least one instruction that defines the cooperative behavior between a given thread and one or more other threads. For example, a per-thread instruction sequence can include an instruction that, when executed, suspends the given thread in a particular execution state until some or all of the other threads reach a corresponding execution state. In another example, a per-thread instruction sequence can include an instruction that, when executed, causes the given thread to store data in shared memory that can be accessed by some or all of the other threads. In yet another example, a per-thread instruction sequence can include an instruction that, when executed, causes the given thread to automatically read and update data stored in shared memory that can be accessed by some or all of the other threads, depending on the thread IDs of those threads. In yet another example, a per-thread instruction sequence can include an instruction that, when executed, causes the given thread to calculate an address in shared memory based on the corresponding thread ID in order to read data from that shared memory. Using the above synchronization techniques, a first thread can write data to a given location in shared memory, and a second thread can read that data from the shared memory in a predictable manner. Thus, threads can be configured to implement a variety of data sharing patterns within a given thread group or within a given CTA or across threads in different thread groups or different CTAs. In various embodiments, software applications written in the Compute Unified Device Architecture (CUDA) programming language describe the behavior and operations of threads executing on the GPC 242, including any of the above behaviors and operations.

[0087] In operation, the pipeline manager 310 generally coordinates the parallel execution of processing tasks in the GPC 242. The pipeline manager 310 receives processing tasks from the task / work unit 234 and assigns those processing tasks to the TPC 330 for execution by the SM 332. A given processing task is generally associated with one or more CTAs that can be executed on one or more SM 332s in one or more TPC 330s. In one embodiment, a given task / work unit 234 can assign one or more processing tasks to the GPC 242 by launching one or more CTAs for one or more specific TPC 330s. The pipeline manager 310 can receive the launched CTAs from the task / work unit 234 and transmit the CTAs to the associated TPC 330 for execution by one or more SM 332s included in the TPC 330. During or after the execution of a given processing task, each SM 332 generates output data and transmits the output data to various locations based on the current configuration and / or the nature of the current processing task.

[0088] In a configuration related to general computing or graphics processing, the SM 332 can transmit the output data to the work assignment crossbar 320, and the work assignment crossbar 320 then routes the output data to one or more GPC 242s for additional processing or routes the output data to the crossbar unit 250 for further routing. The crossbar unit 250 can route the output data to the L2 cache, the PPU memory 270, or the system memory 120 included in a given partition unit 262 and other destinations. The pipeline manager 310 generally coordinates the routing of the output data performed by the work output work assignment crossbar 320 based on the processing task associated with the output data.

[0089] In a configuration specific to graphics processing, the SM 332 can transmit the output data to the texture unit 340 and / or the preROP 370. In some embodiments, the preROP 370 can implement some or all of the raster operations specified in the 3D graphics API, in which case the preROP 370 implements some or all of the operations performed by the ROP 410. The texture unit 340 generally performs texture mapping operations, including, for example, determining texture sample positions, reading texture data, and filtering texture data. The preROP 370 generally performs raster-oriented operations, including, for example, organizing pixel color data and performing color blending optimizations. The preROP 370 can also perform address translation and direct the output data received from the SM 332 to one or more raster operation processor (ROP) units in the partition unit 262.

[0090] In any of the above configurations, one or more PMs 360 monitor the performance of the various components of the GPC 242 to provide performance data to the user, and / or balance the utilization of computing, graphics, and / or memory resources across thread groups, and / or balance the utilization of these resources with the resources of other GPCs 242. Additionally, in any of the above configurations, the SM 332 and other components in the GPC 242 can perform memory access operations with the memory interface 260 via the MMU 300. The MMU 300 generally writes output data to and reads input data from various storage spaces on behalf of the GPC 242 and the components included therein. The MMU 300 is configured to map virtual addresses to physical addresses via a set of page table entries (PTEs) and one or more optional translation lookaside buffers (TLBs). The MMU 300 can cache various data in the L1.5 cache 350, including the read data returned from the memory interface 260. In the illustrated embodiment, the MMU 300 is externally coupled to the GPC 242 and can potentially be shared with other GPCs 242. In other embodiments, the GPC 242 can include a dedicated instance of the MMU 300 that provides access to one or more partition units 262 included in the memory interface 260.

[0091] Figure 4 is a block diagram of the partition unit 262 included in the Figure 2 PPU 200 according to various embodiments. As shown, the partition unit 262 includes an L2 cache 400, a frame buffer (FB) DRAM interface 410, a raster operation processor (ROP) 420, and one or more PMs 430. The L2 cache 400 is coupled between the FB DRAM interface 410, the ROP 420, and the PM 430.

[0092] The L2 cache 400 is a read / write cache that performs load and store operations received from the crossbar unit 250 and the ROP 420. The L2 cache 400 outputs read misses and urgent write-back requests to the FB DRAM interface 410 for processing. The L2 cache 400 also sends dirty updates to the FB DRAM interface 410 for opportunistic processing. In some embodiments, during operation, the PM 430 monitors the utilization of the L2 cache 400 to fairly allocate memory access bandwidth among different GPCs 242 and other components of the PPU 200. The FB DRAM interface 410 directly interfaces with a specific DRAM 272 to perform memory access operations, including writing data to the DRAM 272 and reading data from the DRAM 272. In some embodiments, this set of DRAMs 272 is distributed across multiple DRAM chips, with a portion of the multiple DRAM chips corresponding to each DRAM 272.

[0093] In a configuration related to graphics processing, the ROP 420 performs raster operations to generate graphics data. For example, the ROP 420 can perform stencil operations, z-test operations, blending operations, and compression and / or decompression operations on z or color data. The ROP 420 can be configured to generate various types of graphics data, including pixel data, graphics objects, fragment data, and the like. The ROP 420 can also distribute graphics processing tasks to other compute units. In one embodiment, each GPC 242 includes a dedicated ROP 420 that performs raster operations on behalf of the corresponding GPC 242.

[0094] Those skilled in the art will understand that Figures 1 - 4 The architecture described in this document in no way limits the scope of the present embodiments, and the techniques disclosed herein may be implemented on any suitably configured processing unit, including but not limited to one or more CPUs, one or more multi-core CPUs, one or more PPUs 200, one or more GPCs 242, one or more GPUs or other special-purpose processing units, etc., without departing from the scope and spirit of the embodiments of the invention.

[0095] Logical grouping of hardware resources

[0096] Figure 5 According to various embodiments Figure 2 1 is a block diagram of various PPU resources included in a PPU. As shown, the PPU resources 500 include system pipelines 230(0) to 230(7), a control crossbar and SMC arbiter 510, a privileged register interface (PRI) hub 512, GPCs 242, a crossbar unit 250, and an L2 cache 400. The L2 cache 400 is depicted here as a collection of "L2 cache slices," each slice corresponding to a different region of DRAM 262. The system pipelines 230, GPCs 242, and the PRI hub 512 are coupled together via the control crossbar and SMC arbiter 510. The GPCs 242 and the various slices of the L2 cache 400 are coupled together via the crossbar unit 250. In the example discussed herein, the PPU resources 500 include eight system pipelines 230, eight GPCs 242, and a specific number of other components. However, those skilled in the art will appreciate that the PPU resources 500 may include any technically feasible number of these components.

[0097] Each system pipeline 230 generally includes PBDMAs 520 and 522, a Front - End Context Switch (FECS) 530, a Compute (COMP) Front - End (FE) 540, a Scheduler (SKED) 550, and a CUDA Work Dispatcher (CWD) 560. The PBDMAs 520 and 522 are hardware memory controllers that manage the communication between the device driver 122 and the PPU 200. The FECS 530 is a hardware unit that manages context switching. The Compute FE 540 is a hardware unit that prepares compute tasks for execution. The SKED 550 is a hardware unit that schedules processing tasks for execution. The CWD 560 is a hardware unit configured to queue and dispatch one or more thread grids to one or more GPCs 242 to execute one or more processing tasks. In one embodiment, a given processing task can be specified in a CUDA program. Through the above components, the system pipeline 230 can be configured to perform and / or manage general - purpose computing operations.

[0098] The system pipeline 230(0) further includes a Graphics Front - End (FE) unit 542 (shown as GFX FE 542), a State Change Controller SCC 552, and a Primitive Distributor Stage A / Stage B unit (PDA / PDB) 562. The Graphics FE 542 is a hardware unit that prepares graphics processing tasks for execution. The SCC 552 is a hardware unit that manages the work parallelization with different API states (e.g., shader programs, constants used by shaders, and how to sample textures) to maintain an orderly application of the API state even if primitives are not processed in order. The PDA / PDB 562 is a hardware unit that distributes primitives (e.g., triangles, lines, points, quadrilaterals, meshes, etc.) to the GPCs 242. Through these additional components, the system pipeline 230(0) can be further configured to perform graphics processing operations. In various embodiments, some or all of the system pipeline 230 can be configured to include components similar to those of the system pipeline 230(0), and thus be capable of performing general - purpose computing operations or graphics processing operations. Optionally, in various other embodiments, some or all of the system pipeline 230 can be configured to include components similar to those of the system pipeline 230(1) to 230(7), and thus be capable of performing only general - purpose computing operations. Generally, Figure 2 the front - end 232 can be configured to include the Compute FE 540, the Graphics FE 542, or both the Compute FE 540 and the Graphics FE 542. Thus, for general considerations, the front - end 232 is hereinafter referred to with reference to one or both of the Compute FE 540 and the Graphics FE 542.

[0099] The control crossbar and SMC arbiter 510 facilitate communication between the system pipeline 230 and the GPC 242. In some configurations, one or more specific GPCs 242 are programmably assigned to perform processing tasks on behalf of a specific system pipeline 230. In such configurations, the control crossbar and SMC arbiter 510 are configured to route data between any given GPC 242 and the corresponding system pipeline 230. The PRI hub 512 provides access to a set of privileged registers through the CPU 110 and / or PPU 200 units to control the configuration of the PPU 200. The register address space of the PPU 200 can be configured through the PRI registers, and in this way, the PRI hub 212 is used to configure the mapping of the PRI register addresses between the general PRI address space and the PRI address spaces defined separately for each system pipeline 230. This PRI address space configuration provides the function of broadcasting from the SMC engine to multiple PRI registers, which will be described in conjunction with Figure 7 below. The GPC 242 writes data to and reads data from the L2 cache 400 via the crossbar unit 250 in the foregoing manner. In some configurations, each GPC 242 is assigned a separate set of L2 slices derived from the L2 cache 400, and any given GPC 242 can perform write / read operations on the corresponding set of L2 slices.

[0100] Any of the PPU resources 500 discussed above can be logically grouped or partitioned into one or more PPU partitions, and each partition operates in the same manner as the PPU 200 as a whole. Specifically, a given PPU partition can be configured with sufficient computing, graphics, and memory resources to perform any technically feasible operation that can be performed by the PPU 200. An example of how the PPU resources 500 are logically grouped into partitions is described in detail below in conjunction with Figure 6 below.

[0101] Figure 6 is an example of how a hypervisor according to various embodiments Figure 1 logically groups PPU resources into a set of PPU partitions. As shown, the PPU partition 600 includes one or more PPU slices 610. Specifically, the PPU partition 600(0) includes PPU slices 610(0) to 610(3), the PPU partition 600(4) includes PPU slices 610(4) and 610(5), the PPU partition 600(6) includes PPU slice 610(6), and the PPU partition 600(7) includes PPU slice 610(7). In the example discussed herein, the PPU partition 600 includes the specific number of PPU slices 610 shown. However, in other configurations, the PPU partition 600 can include other numbers of PPU slices 610.

[0102] Each PPU slice 610 includes various resources derived from a system pipeline 230, including PBDMAs 520 and 522, FECS 530, front end 232, SKED 550, and CWD 560. Each PPU slice 610 also includes a GPC 242, an L2 slice set 620, and a corresponding portion of the DRAM 272 (not shown here). The various resources included within a given PPU slice 610 confer sufficient functionality such that any given PPU slice 610 can perform at least some of the general computing and / or graphics processing operations that the PPU 200 is capable of performing.

[0103] For example, a PPU slice 610 can receive processing tasks via the front end 232 and then schedule those processing tasks for execution via the SKED 550. Then, the CWD 560 can issue a thread grid to execute those processing tasks on the GPC 242. The GPC 242 can execute multiple thread groups in parallel in the manner described above in connection with Figure 3 The PBDMAs 520 and 522 can perform memory access operations on behalf of the various components included within the PPU slice 610. In certain embodiments, the PBDMAs 520 and 522 fetch commands from memory and send the commands to the FE 232 for processing. As needed, the various components of the PPU slice 610 can write data to and read data from the corresponding L2 cache slice set 620. The components of the PPU slice 610 can also interface with external components included within the PPU 200 as needed, including the I / O unit 210 and / or the PCE 222, etc. The FECS 530 can perform context switching operations when time slicing one or more VMs across the various resources included within the PPU slice 610.

[0104] In the illustrated embodiment, each PPU slice 610 includes resources derived from the system pipeline 230 that are configured to coordinate general computing operations. Thus, the PPU slice 610 is configured to perform only general processing tasks. However, in other embodiments, each PPU slice 610 can further include resources derived from the system pipeline 230 that are configured to coordinate graphics processing operations, such as the system pipeline 230(0). In these embodiments, the PPU slice 610 can be configured to additionally perform graphics processing tasks.

[0105] Typically, each PPU partition 600 is a hard partition of resources that provides a dedicated parallel computing environment isolated from other PPU partitions 600 for one or more users. A given PPU partition 600 includes one or more dedicated PPU slices 610 as shown, which together provide the various general computing, graphics processing, and memory resources needed to at least somewhat mimic the overall functionality of the PPU 200. Thus, a given user can perform parallel processing operations within a given PPU partition 600 in a manner similar to that of a similar user performing those same parallel processing operations on the PPU 200 when the PPU 200 is not partitioned. Each PPU partition 600 is fault-insensitive to other PPU 600s, and each PPU partition can be reset independently of other PPU partitions 600 and without interrupting the operation of other PPU partitions 600. As described in more detail below, various resources not specifically shown here are fairly distributed across different PPU partitions 600 in proportion to the sizes of those different PPU partitions 600.

[0106] In the example configuration of the PPU partition 600 discussed herein, PPU partition 600(0) is assigned four of the eight PPU slices 610 and is thus provided with half of the PPU resources 500, including various types of bandwidth such as memory bandwidth. Thus, PPU partition 610(0) will be constrained to consume half of the available system storage bandwidth, half of the available PPU storage bandwidth, half of the available PCE 212 bandwidth, etc. Similarly, PPU partition 600(4) is assigned two of the eight PPU slices 610 and is thus provided with one-quarter of the PPU resources 500. Thus, PPU partition 610(4) will be limited to consuming one-quarter of the available system storage bandwidth, one-quarter of the available PPU storage bandwidth, one-quarter of the available PCE 212 bandwidth, etc. Other PPU partitions 600(6) and 600(7) will be constrained in a similar manner. Those skilled in the art will understand how to implement the above exemplary partitioning and associated resource provisioning using any other technically feasible configuration of the PPU partition 600.

[0107] In some embodiments, each PPU partition 600 executes a context for a virtual machine (VM). In one embodiment, the PPU 200 can implement various performance monitors and throttling counters that record the amount of local and / or system-wide resources being consumed by each PPU partition 600 in order to maintain proportional resource consumption across all PPU partitions 600. Allocating an appropriate portion of the PPU storage bandwidth to the PPU partition 600 can be achieved by allocating the same portion of the L2 slices 400 to the PPU partition 600.

[0108] In general, the PPU partitions 600 can be configured to operate in a functionally isolated manner relative to each other. As referred to herein, the term "functionally isolated" as applied to a PPU partition set 600 generally means that any PPU partition 600 can perform one or more operations independently of, without interfering with, and without being interfered with by any operations performed by any other PPU partition 600 in the PPU partition set 600.

[0109] A given PPU partition 600 can be configured to execute processing tasks associated with multiple processing contexts simultaneously. The term "processing context" or "context" generally refers to the state of hardware, software, and / or memory resources during the execution of one or more threads, and generally corresponds to a process on the CPU 110. The multiple processing contexts associated with a given PPU partition 600 can be different processing contexts or different instances of the same processing context. When configured in this manner, the specific PPU resources assigned to a given PPU partition 600 are logically grouped into separate "SMC engines" that execute separate processing tasks associated with separate processing contexts, as described below in conjunction with Figure 7 Thus, a given processing context may include hardware settings, per-thread instructions, and / or register contents associated with a thread executed in the SMC engine 700 .

[0110] Figure 7 Shown according to various embodiments Figure 1 2. An example of how a hypervisor may configure a set of PPU partitions to implement one or more simultaneous multi-context (SMC) engines. As shown, PPU partition 600 includes one or more SMC engines 700. In particular, PPU partition 600(0) includes SMC engines 700(0) and 700(2), PPU partition 600(4) includes SMC engine 700(4), PPU partition 600(6) includes SMC engine 700(6), and PPU partition 600(7) includes SMC engine 700(7). Each SMC engine 700 may be configured to execute one or more processing contexts and / or to execute one or more processing tasks associated with a given processing context, in a manner similar to that of the PPU 200 as a whole.

[0111] A given SMC engine 700 generally includes computational and memory resources associated with at least one PPU slice 610. For example, SMC engines 700(6) and 700(7) include computational and memory resources associated with PPU slices 610(6) and 610(7), respectively. Each SMC engine 700 also includes a set of virtual engine identifiers (VEIDs) 702 that locally reference one or more sub-contexts, where the VEID is associated with a virtual address space identifier for selecting a virtual address space and may be the same as it, and the pages of the virtual address space are described by page tables managed by MMU 1600. A given SMC engine 700 may also include computational and memory resources associated with multiple PPU slices 610. For example, SMC engine 700(0) includes computational resources associated with PPU slices 610(0) and 610(1), but does not utilize system pipeline 230(1). SMC engine 700(0) includes and utilizes the L2 slices in four PPU slices 610(0), 610(1), 610(2), and 610(3). In some embodiments, SMC engines 700 within the same PPU partition 600 share the L2 slices within PPU partition 600. In this configuration, system pipeline 230(1) of the shown PPU partition 600(1) is not used because SMC engines generally run one processing environment at a time, and one processing environment only requires one system pipeline 230. SMC engine 700(2) is configured in a similar manner to SMC engine 700(0). The memory resources contained in any particular PPU partition 600 may be shown as PPU memory partition 710, and this memory resource may be allocated to and / or distributed among any one or more SMC engines 700 within that particular PPU partition 600.

[0112] A given PPU memory partition 710 includes the set of L2 slices included in PPU partition 600 and the corresponding portion of DRAM 272. Generally, multiple SMC engines 700 share one PPU memory partition 710 if those SMC engines 700 are included in the same PPU partition 600. The allocation for each SMC engine 700 is provided to the contexts running on those SMC engines 700, and the allocation within PPU memory partition 710 is implemented based on pages.

[0113] Each SMC engine 700 can be configured to independently execute processing tasks associated with one processing context at any given time. Thus, the PPU partition 600(0) having two SMC engines 700(0) and 700(2) can be configured to simultaneously execute processing tasks associated with two separate processing contexts at any given time. On the other hand, each of the PPU partitions 600(4), 600(6), and 600(7) respectively including one SMC engine 700(4), 700(6), and 700(7) can be configured to execute processing tasks associated with one processing context at a time. In some embodiments, contexts running on SMC engines 700 in different PPU partitions 600 can share data by sharing one or more pages in one or two PPU partitions 600.

[0114] Any given SMC engine 700 can be further configured to time slice different processing contexts over different time intervals. Thus, each SMC engine 700 can independently support the execution of processing tasks associated with multiple processing contexts, although not necessarily simultaneously. For example, the SMC engine 700(6) can time slice four different processing contexts over four different time intervals, thereby allowing the processing tasks associated with these four processing contexts to be executed within the PPU partition 600(6). In some embodiments, the VM is time sliced over one or more PPU partitions 600. For example, the PPU partition 600(0) can time slice between two VMs, where each VM simultaneously executes two processing contexts, one processing context on each of the SMC engines 700(0) and 700(1). In these embodiments, preferably, all processing contexts are switched out of the first VM before context switching to the second VM in the processing context, which is advantageous when the processing contexts running on the PPU partition 600(0) share the L2 slice 400 within the PPU partition 600(0).

[0115] In one embodiment, a given VM can be associated with a GPU function ID (GFID). A given GFID can contain one or more bits that correspond to physical functions (PFs) associated with the hardware in which the VM executes. The given GFID can also include a set of bits that correspond to virtual functions (VFs) uniquely assigned to the VM. Among other uses, the given GFID can be used to route errors to a location corresponding to the guest operating system of the VM.

[0116] The SMC engines 700 within different PPU partitions 600 generally operate isolated from each other because, as previously described, each PPU partition 600 is a hard partition of the PPU resources 500. Multiple SMC engines 700 within the same PPU partition 600 can generally operate independently of each other, and in particular can context-switch independently of each other. For example, the SMC engine 700(0) within the PPU partition 600(0) can context-switch independently and asynchronously relative to the SMC engine 700(2). In some embodiments, multiple SMC engines 700 within the same PPU partition 600 can context-switch synchronously to support certain operating modes, such as time-slicing between two VMs.

[0117] Generally, Figure 1 the device driver 122 and the hypervisor 124 interoperate in the manner described so far to partition the PPU 200. In addition, the device driver 122 and the hypervisor 124 interoperate to configure each PPU partition 600 into one or more SMC engines 700. In this way, the device driver 122 and the hypervisor 124 configure the DRAM 272 and / or the L2 cache 400 so as to partition the set of L2 slices into groups each of which is an SMC memory partition 710, as described in more detail below in connection with FIG. 8. In some embodiments, the hypervisor 124 responds to control by a system administrator to allow the system administrator to create a configuration of PPU partitions. These PPU partitions 600 are switched to the guest OS 916 of a VM, and the guest OS 916 then sends a request to the hypervisor 124 to configure the associated PPU partition 600 as one or more SMC engines 700. In some embodiments, because sufficient isolation is added to prevent one guest OS from affecting the PPU partition 600 of another guest OS, the guest OS can directly configure the SMC engines 700 within the PPU partition 600.

[0118] Figure 8A is a more detailed illustration of Figure 7 the DRAM according to various embodiments. As shown, the DRAM 272 can be accessed via the L2 slices 800, which include Figure 7 each of the DRAMs 272(0) to 272(7). Each L2 slice 800 corresponds to a different part of the L2 cache 400 and is configured to access a corresponding subset of the locations within the DRAM 272. Generally, the partitioning of the DRAM 272 corresponds to the original 2D address space, which is organized similarly to the DRAM 272 shown herein.

[0119] Also as shown, DRAM 272 is divided into a top portion 810, a partitionable portion 820, and a bottom portion 830. The top portion 810 and the bottom portion 830 are storage partitions derived from the top and bottom portions of all DRAMs 272(0) through 272(7), respectively. The device driver 122, the hypervisor 124, and other system-level entities can access the top portion 810 and / or the bottom portion 830, which in some embodiments are not accessible to the PPU partitions 600. On the other hand, the partitionable portion 820 is generally designated for use by the general PPU partitions 600 and is specifically used by the SMC engine 700. In some embodiments, secure data resides in the top portion 810 or the bottom portion 830 and is accessible to all PPU partitions 600. In some embodiments, the top portion 810 or the bottom portion 830 is used for hypervisor data that is not accessible to VMs.

[0120] In the exemplary memory partitioning shown, the partitionable portion 820 includes a DRAM portion 822(0) corresponding to the PPU memory partition 710(0) within the PPU partition 600(0), a DRAM portion 822(4) corresponding to the PPU memory partition 710(4) within the PPU partition 600(2), a DRAM portion 822(6) corresponding to the PPU memory partition 710(6) within the PPU partition 600(6), and a DRAM portion 822(7) corresponding to the PPU memory partition 710(7) within the PPU partition 600(7). Each DRAM portion 822 corresponds to an intermediate portion of the address corresponding to a set of L2 cache slices 800. A given DRAM portion 822 can be further subdivided to provide separate sets of L2 cache slices for different VMs that execute processing tasks associated with different processing contexts. For example, the DRAM portion 822(4) can be subdivided into two or more regions to support two or more VMs that execute processing tasks associated with two or more processing contexts. Once a DRAM portion is configured and in use, it is generally used by one VM running on the PPU partition 600 at a time.

[0121] In operation, device driver 122 and hypervisor 124 perform memory access operations in a relatively balanced manner within top portion 810 and / or bottom portion 830 via the top and bottom portions of the address ranges corresponding to all L2 cache slices 800, thereby proportionally penalizing the memory bandwidth on each L2 slice 800. In some embodiments, SMC engine 700 performs memory access operations to system memory 120 via L2 cache slices 800, and its throughput is controlled by throttling counter 840. Each throttling counter 840 monitors the memory bandwidth consumed when SMC engine 700 accesses system memory 120 via the L2 cache slice 800 associated with the corresponding PPU memory partition 710, in order to provide proportional memory bandwidth to each PPU partition 600. As discussed, access to various system-wide resources is provided to PPU partitions in proportion to the configuration of those PPU partitions 600. In the example shown, PPU partition 600(0) is allocated half of PPU resources 500 and thus is allocated half of the partitionable portion 820 (shown as DRAM portion 822(0)), and correspondingly, half of the available memory bandwidth is allocated to system memory 120. The partitioning of DRAM 272 is described in more detail below in conjunction with Figures 19 - 24 the partitioning of DRAM 272.

[0122] Figure 8B illustrates how various Figure 8B DRAM portions are addressed according to various embodiments. As shown, one-dimensional (1D) system physical address (SPA) space 850 includes top address 852 corresponding to top portion 810, partitionable address 854 divided into address regions 856 and corresponding to DRAM portions 822, and bottom address 858 corresponding to bottom portion 830. Top address 852 is obfuscated (i.e., pseudo-random interleaving based on the SPA address) across all L2 slices 800 and corresponds to the top portion of those L2 slices. Bottom address 858 is obfuscated across all L2 slices 800 and corresponds to the bottom portion of those L2 slices. Generally, only system-level entities (e.g., hypervisor 124) and / or any entity operating via a physical function (PF) can access top address 852 and bottom address 858. Partitionable address 854 is allocated to PPU partitions 600. In particular, address region 856(0) is allocated to PPU partition 600(0), address region 856(4) is allocated to PPU partition 600(4), address region 856(6) is allocated to PPU partition 600(6), and address region 856(7) is allocated to PPU partition 600(7). Address regions 856 can only be accessed by one or more SMC engines 700 executing within the corresponding PPU partition 600.

[0123] Overall reference Figures 5 - 8B , the above method of partitioning the PPU resources 500 supports a variety of usage scenarios, including single-tenant and multi-tenant usage scenarios. In a single-tenant usage scenario, the PPU 200 can be partitioned to provide different users associated with a single tenant with independent access to the PPU resources. For example, different users associated with a given tenant can perform different predetermined workloads on different PPU partitions 600. In a single-tenant usage scenario, access to the entire PPU resource 500 can be provided to a single entity. In a multi-tenant usage scenario, the PPU 200 can be partitioned to provide independent access to the PPU resources to one or more users associated with one or more different tenants. In a multi-tenant usage scenario, multiple entities can be provided with access to different PPU partitions 600, which include different parts of the PPU resources 500.

[0124] In any usage scenario, the device driver 122 and the hypervisor 124 interoperate to perform a two-step process that first involves dividing the PPU 200 into PPU partitions 600 and secondly involves configuring these PPU partitions 600 as SMC engines 700. Figures 9 - 10 Provide a more detailed explanation.

[0125] Techniques for configuring logical groupings of hardware resources

[0126] Figure 9 is a diagram showing various embodiments of the present invention. Figure 1 Flowchart of how the hypervisor partitions and configures PPUs. As shown, the hypervisor environment 900 includes a guest environment 910 and a host environment 920 separated from each other by a hypervisor trust boundary 930. The guest environment 910 includes a system management interface (SMI) 912, a kernel driver 914, and a guest operating system (OS) 916. The host environment 920 includes an SMI 922, a virtual GPU (vGPU) plug-in 924, a host OS 926, and a kernel driver 928. Modules included in the guest environment 910 that reside above the hypervisor trust boundary 930 generally execute at a lower privilege level than modules included in the host environment 920 that reside below the hypervisor trust boundary 930, including duplicate instances of the same module, such as SMI 912 and SMI 922. The hypervisor 124 executes with a kernel-level set of privileges and can grant appropriate privileges to any of the modules shown. In some embodiments, there is a one-to-one correspondence between VMs and guest environments 910 ; and when multiple virtual machines are not context-switched to PPU partition 600 , there is typically a one-to-one correspondence between guest environments and PPU partitions.

[0127] In operation, an administrator user of the PPU 200 interacts with the PPU 200 via the host environment 920 and the host OS 926 to configure the PPU partition 600. In particular, the administrator user provides a partition input 904 to the SMI 922. In response, the SMI 922 issues a "create partition" command to the core driver 928, indicating the target configuration of the PPU partition 600. The core driver 928 transmits the "create partition" command to the host interface 220 in the PPU 200 to partition various PPU resources 500. In this way, the administrator user can initialize the PPU 200 to have a specific configuration with the PPU partition 600. Typically, the administrator user has unrestricted access to the PPU 200. For example, the administrator user can be a system administrator of the data center where the PPU 200 is located. The administrator user can be a system administrator of the data center where multiple PPU 200s reside. The administrator user partitions the PPU 200 in the described manner to prepare individual PPU partitions 600 to be independently configured and used by respective guest users, as described in more detail below.

[0128] A guest user of the PPU 200 interacts with a specific "guest" PPU partition 600 via a VM executed within the guest environment 910 to configure the SMC engine 700 within that guest PPU partition 600. Specifically, the guest user provides a configuration input 902 to the SMI 912. Then the SMI 912 issues a "configure partition" command to the core driver 914, indicating the target configuration of the SMC engine 700. The core driver 914 sends the "configure partition" command across the hypervisor trust boundary 930 to the vGPU plugin 924 via the guest OS 916. The vGPU plugin 924 makes various VM calls to the core driver 928. The core driver 928 sends the "configure partition" command to the host interface 220 in the PPU 200 to configure various resources of the guest PPU partition 600. In this way, the guest user can configure a given PPU partition 600 to have a specific configuration with the SMC engine 700. Typically, the guest user can only access a portion of the PPU resources 500 associated with the guest PPU partition 600. For example, the guest user can be a customer of the data center where the PPU 200 is located who has purchased access to a portion of the PPU 200. In one embodiment, the guest OS 916 can be configured with sufficient security measures to allow each guest OS 916 to configure the corresponding PPU partition 600 without the involvement of the host environment 920 and / or the hypervisor 124.

[0129] Figure 10 is a flowchart of method steps for partitioning and configuring a PPU on behalf of one or more users according to various embodiments. Although combined with Figures 1 to 9The system describes method steps, but those skilled in the art will understand that any system configured to execute the method steps in any order falls within the scope of this embodiment.

[0130] As shown, the method begins at step 1000, where the hypervisor 124 receives a partition input 904 via the host environment 920. The host environment 920 executes at an elevated privilege level, thus allowing an administrator user to directly interact with the PPU 200. The partition input 904 specifies the target configuration of the PPU partition 600, including the desired size and layout of the PPU partition 600. In one embodiment, the partition input can be received from an administrator user via the host environment.

[0131] At step 1004, the hypervisor 124 generates one or more PPU partitions 600 within the PPU 200 based on the partition input 904 received at step 1002. In particular, the hypervisor 124 implements the SMI 922 to issue a "create partition" command to the core driver 928. In response, the core driver 928 interacts with the host interface 220 of the PPU 200 to create one or more PPU partitions 600 with the desired configuration.

[0132] At step 1006, the hypervisor 124 allocates memory resources to the one or more PPU partitions 600 generated at step 1004. In particular, via the "create partition" command discussed above, the hypervisor 124 subdivides the DRAM 272 in the manner described above to allocate different regions of the DRAM 272 and the corresponding L2 cache slices 800 to the one or more PPU partitions 600. In one embodiment, the hypervisor 124 can also configure the address mapping unit to perform partition-specific scrambling operations to provide access to those different regions of the DRAM 272 via the L2 cache slices 800. This particular embodiment is described in more detail below in connection with Figure 8A the above. Figure 22 is described in more detail.

[0133] In step 1008, the hypervisor 124 allocates PPU computing and / or graphics resources on one or more PPU partitions 600. To do so, the hypervisor 124 allocates one or more system pipelines 230 and one or more GPCs 242 to one or more PPU partitions 600 via a "create partition" command. In one embodiment, the hypervisor 124 may implement steps 1006 and 1008 by logically allocating one or more PPU slices 610 to one or more PPU partitions 600, thereby allocating memory resources and computing / graphics resources together. When steps 1002, 1004, 1006, and 1008 of method 1000 are complete, the PPU 200 is partitioned, and then the guest user may configure one or more PPU partitions 600, as described below.

[0134] In step 1010, the hypervisor 124 receives a configuration input 902 associated with the first PPU partition via the guest environment 910. The guest environment 910 executes at a reduced privilege level, thereby allowing the guest user to interact only with the first PPU partition 600. The configuration input 902 specifies a target configuration of the SMC engine 700 within the first PPU partition 600, including the desired size and arrangement of the SMC engine 700. In one embodiment, the configuration input may be received from the guest user via the guest environment.

[0135] In step 1012, the hypervisor 124 generates one or more SMC engines 700 within the first PPU partition 600 based on the configuration input 902 received in step 1010 via a "configure partition" command. A given SMC engine 700 may include computing and / or graphics resources derived from one system pipeline 230 and one or more GPCs 242 derived from one or more PPU slices 610. A given SMC engine 700 may access at least a portion of the PPU memory partition 710 contained within the PPU partition 600 in which the SMC engine 700 resides, where the PPU memory partition 710 includes one or more sets of L2 cache slices and a corresponding portion of the DRAM 272.

[0136] In step 1014, the hypervisor 124 distributes the memory resources allocated to the first PPU partition 600 across one or more SMC engines 700 generated in step 1012. Generally, one or more SMC engines 700 share the first PPU memory partition 710 if these SMC engines 700 are included in the same PPU partition 600. The allocation of each SMC engine 700 is provided to the contexts running on these SMC engines 700, and the allocation in the PPU memory partition 710 is implemented based on pages. In some embodiments, the guest OS 916 of the VM performs step 1014 by distributing the memory resources to the contexts running on the SMC engines 700, and the SMC engines 700 are part of the PPU partition 600 of the guest environment 910. In other embodiments, the hypervisor 124 performs the memory resource allocation 1014 for multiple VMs using the PPU partition 600, and each VM also distributes the memory resources 1024 to the contexts running on the SMC engines 700.

[0137] In step 1016, the hypervisor 124 distributes the computing and / or graphics resources allocated to the first PPU partition 600 across one or more SMC engines 700 generated in step 1012. Through the "configure partition" command, the hypervisor 124 allocates the system pipeline 230 included in the guest PPU partition 600 to each SMC engine 700. The hypervisor 124 also allocates one or more GPCs 242 to each SMC engine 700. When the steps 1010, 1012, 1014, and 1016 of the method 1000 are completed, the first PPU partition 600 is configured, and then the guest user can initiate processing operations on one or more SMC engines 700 within the PPU partition 600.

[0138] In step 1018, the hypervisor 124 time-slices one or more VMs across one or more SMC engines 700 configured within the first PPU partition 600. The time-sliced VMs can operate independently of other VMs executing within a first PPU partition 600 and operate isolated from other VMs executing within other PPU partitions 600. In one embodiment, one or more VMs can be time-sliced simultaneously across one or more SMC engines 700. In this way, the disclosed techniques allow the partitioned PPU to support the parallel execution of processing tasks associated with multiple different processing contexts.

[0139] In some embodiments, the techniques disclosed herein operate in a non-virtualized system. Those skilled in the art will recognize that a single OS usage model on the PPU 200 or a group of PPU 200s can use all of the mechanisms described in connection with VMs. In some embodiments, a container corresponds to the description of a VM, which means that a container on a single OS can implement the processing isolation provided to the VMs described herein.

[0140] Partitioning computing resources to support multiple contexts simultaneously

[0141] In various embodiments, when the hypervisor 124 partitions the 200 PPUs on behalf of an administrator user in the manner described above, the hypervisor 124 receives input from the administrator user that indicates various boundaries between the PPU partitions 600. Based on this input, the hypervisor 124 logically groups the PPU slices 610 into PPU partitions 600, allocates various hardware resources to each PPU partition 600, and coordinates various other operations to support the simultaneous implementation of multiple processing contexts within a given PPU partition 600. The hypervisor 124 also performs additional techniques to support the migration of processing contexts between PPU partitions 600 configured on different PPUs 200. These various techniques are described in more detail below in connection with Figures 11 - 18 These various techniques are described in more detail below.

[0142] Figure 11 An embodiment of a partition configuration table according to various embodiments is shown, according to which Figure 1 the hypervisor 124 can configure one or more PPU partitions. As shown, the partition configuration table 1100 includes partition options 0 through 14. The partition options 0 - 14 are depicted above the PPU slices 610. Each of the partition options 0 - 14 spans a different grouping of the PPU slices 610 and in this way represents different possible partitions of the PPU 600. Specifically, partition option 0 spans PPU slices 610(0) through 610(7) and thus represents a PPU partition 600 that includes all eight PPU slices 610. Similarly, partition option 1 spans PPU slices 610(0) through 610(3) and thus represents a PPU partition 600 that includes only the first four PPU slices 610, similar to Figures 6 - 7 the PPU partition 600(0) shown in. Partition option 2 spans PPU slices 610(4) through 610(7) and thus represents a PPU partition 600 that includes only the last four PPU slices 610. Partition options 3, 4, 5, and 6 span different groupings of two adjacent PPU slices 610, while partition options 7, 8, 9, 10, 11, 12, 13, and 14 span only a respective single PPU slice 610.

[0143] The partition configuration table 1100 also includes boundary options 1110 representing different possible positions of the partition boundaries. Specifically, boundary options 1110(1) and 1110(9) represent the boundaries of partition option 0. Boundary options 1110(1) and 110(5) represent the boundaries of partition option 1, while boundary options 1110(5) and 1110(9) represent the boundaries of partition option 2. Boundary options 1110(1), 1110(3), 1110(5), 1110(7), and 1110(9) represent the boundaries associated with partition options 3, 4, 5, and 6. Boundary options 1110(1) through 1110(9) represent the boundaries associated with partition options 7 through 14. It should be understood that those skilled in the art can create many different schemes to achieve the same functionality as the configuration table 1100, which may be a set of enable bits, a list of predefined selections, or any other form that allows control of how the PPU slices 610 are divided into PPU partitions 600. Additionally, those skilled in the art will understand that the partition configuration table 1100 can include any technically feasible numbers or entries other than Figure 11 those shown.

[0144] During partitioning, the hypervisor 124 or device driver 122 running at the hypervisor level receives partition input from an administrator user that indicates a specific partition option according to which the PPU 200 should be partitioned. The hypervisor 124 or device driver 122 then activates a specific set of boundary options 1110 that logically isolate one or more groups of PPU slices 610 from each other to achieve the desired partitioning, as described in more detail below in conjunction with Figure 12 the example of.

[0145] Figure 12 illustrates how the hypervisor 124 or device driver 122 of Figure 1 partition the PPU to generate one or more PPU partitions according to various embodiments. As shown, during partitioning, the hypervisor 124 or device driver 122 running at the hypervisor level receives input from an administrator user that indicates that the PPU 200 should be partitioned according to partition options 1, 5, 13, and 14 (emphasized for clarity). In response, the hypervisor 124 or device driver 122 activates boundary options 1110(1), 1110(5), 1110(7), 1110(8), and 1110(9) and deactivates the other boundary options in order to generate PPU partitions 600(0), 600(4), 600(6), and 600(7). The exemplary configuration of the PPU partitions 600 is also shown in Figures 6 - 7 .

[0146] Generally referring to Figures 11 - 12, the hypervisor 124 or the device driver 122 implements the above technique by mapping each selection of the partitioning option to a specific binary value, which is then used to enable and disable the boundary option 1110. The binary value associated with a given partitioning option is referred to herein as a "swizzle identifier" (swizID). The various swizIDs implemented by the hypervisor 124 are listed in Table 1 below:

[0147]

[0148]

[0149] Table 1

[0150] The hypervisor 124 activates or deactivates the boundary option 1110 for a given partitioning option based on the swizID associated with the given partitioning option. For example, the hypervisor 124 can activate the boundary options 1110(1) and 1110(3) to configure the PPU 200 according to the partitioning option 3 based on the corresponding swizID 10000001011. Bits 1 and 3 of this swizID activate the boundary options 1110(1) and 1110(3) respectively, and bits 2 and 4 - 9 deactivate the remaining boundary options. Bits 0 and 10 of all swizIDs are set to one (1) to activate the boundary within the L2 cache 400, as described in more detail below in conjunction with Figures 19 - 20 The hypervisor 124 collects various swizIDs for different selected configuration options and computes an OR operation across all the collected swizIDs to generate a configuration swizID that defines the configuration of the PPU partition 600. The configuration swizID indicates all the boundary options 1110 that should be activated and deactivated to achieve the desired configuration of the PPU partition 600.

[0151] Those skilled in the art will recognize that certain combinations of partitioning options are not feasible. For example, partitioning options 0 and 1 cannot be implemented in combination with each other because partitioning options 0 and 1 overlap. During partitioning, the hypervisor 124 corrects these combinations by automatically detecting infeasible combinations of partitioning options and by modifying one or more partitioning options and / or the corresponding swizIDs or omitting one or more partitioning options and / or the corresponding swizIDs.

[0152] In addition, the hypervisor 124 can dynamically detect hardware faults that render certain partitioning options infeasible. For example, assume that PPU slice 610(0) includes a non-functional GPC 242 that was scratched and fused during manufacturing by the floor. In this case, the PPU partition 600 that includes only PPU slice 610(0) will lack sufficient computing resources to operate and will thus be difficult to implement. In such a case, the hypervisor 124 will not allow the selection of partitioning option 7 and / or the use of the corresponding swizID because any PPU partition 600 configured according to that partitioning option will be unable to perform computational operations and will thus not function properly.

[0153] In some cases, the hypervisor 124 can allow certain configuration options that include a certain amount of non-functional hardware, as long as the PPU partition 600 configured according to such a configuration option can still function to some extent. In the above example, the hypervisor 124 can allow the selection of configuration option 3 as long as PPU slice 610(1) includes a functional GPC 242. Any PPU partition 600 configured according to configuration option 3 will still function, but it will only include half of the computing resources compared to a similar PPU partition 600 that does not include any non-functional hardware.

[0154] After partitioning the PPU 200 in the above-described manner, the hypervisor 124 allocates various hardware resources to the resulting PPU partitions 600. Some of these resources are statically allocated to individual PPU slices 610 and provide dedicated support for specific operations, while other resources are shared within the same PPU partition 600 or between different PPU slices 610 in different PPU partitions 600, as described in more detail below Figure 13 and described in more detail below.

[0155] Figure 13 illustrates how the Figure 1 hypervisor allocates various PPU resources during partitioning according to various embodiments. As shown, PCEs 222(0) to 222(7) are coupled to PPU slices 610(0) to 610(7). In this example, the number of PCEs 222 included in the PPU 200 is equal to the number of PPU slices 610. Thus, the hypervisor 124 can statically allocate each PCE 222 to a different PPU slice 610 and configure those PCEs 222 to perform copy operations on behalf of the corresponding PPU slice 610 in a dedicated manner.

[0156] Other hardware resources included in the PPU 200 cannot be statically allocated in the above - described manner because these resources may be relatively scarce. In the example shown, the PPU 200 includes only two decoders 1300 that need to be allocated across eight PPU slices 610. Thus, the hypervisor 124 dynamically allocates decoder 1300(0) to PPU slices 610(0) through 610(3) included in PPU partition 600(0). The hypervisor 124 also dynamically allocates decoder 1300(1) to PPU slices 610(4) and 610(5) included in PPU partition 600(4), PPU slices 610(6) included in PPU partition 600(6), and PPU slice 600(7) included in PPU partition 600(7).

[0157] In the configuration shown, decoder 1300(0) is dynamically allocated to perform decoding operations for PPU partition 600(0) in an exclusive manner, but decoder 1300(1) is shared among PPU partitions 600(4), 600(6), and 600(7). In various embodiments, one or more performance monitors may manage the use of the hardware resources shared in the described manner to balance the resource usage among different PPU slices 610. The hypervisor 124 performs the above techniques to allocate any technically feasible resources of the PPU 200 to the PPU partitions 600.

[0158] When partitioning has been performed and the various resources of the PPU 200 have been statically or dynamically allocated to the respective PPU slices 610, the hypervisor 124 is ready to allow the VMs to begin executing processing tasks within those PPU partitions 600. Thus, a VM can simultaneously start multiple processing contexts within a given PPU partition 600 that are isolated from other processing contexts associated with other PPU partitions 600, as described above and as will be described in detail below Figures 14A - 14B and as described in detail below. Unused PPU slices 610 can be repartitioned into other PPU partitions 600, while other PPU slices are used within the active PPU partitions 600.

[0159] Figure 14AAccording to various embodiments, it is shown how multiple guest operating systems 916 running multiple VMs can simultaneously start multiple processing contexts within one or more PPU partitions. As shown, the guest operating systems 916 include various processing contexts 1400 associated with different PPU partitions 600. Processing contexts 1400(0) and 1400(1) are associated with PPU partition 600(0) and can be started on SMC engine 700(0) or SMC engine 700(1). In some embodiments, once a processing context is assigned to an SMC engine 700, it remains on that SMC engine 700 until completion. Processing context 1400(4) is associated with PPU partition 600(4) and can be started on SMC engine 700(4). Processing contexts 1400(6) and 1400(6) are associated with PPU partitions 600(6) and 600(7) respectively, and can be started on SMC engines 700(6) and 700(7) respectively.

[0160] As previously described in connection with Figure 7 each SMC engine 700 can perform the processing tasks associated with a given processing context 1400 independently of other SMC engines 700 that execute the processing tasks associated with any given processing context 1400. The processing tasks performed by a given SMC engine 700 in conjunction with a given processing context 1400 are scheduled independently of the other processing tasks performed by other SMC engines 700 in conjunction with any other processing context 1400. Additionally, as described in more detail below in connection with Figures 15 - 16 the SMC engines 700 can experience failures and / or errors independently of each other and can be reset without interrupting the operation of other SMC engines 700.

[0161] In addition, each SMC engine 700 can be configured to perform processing tasks associated with one or more processing sub-contexts 1410, which are included in and / or derived from a single parent processing context 1400. As shown, a given processing context 1400(0) includes one or more processing sub-contexts 1410(0), and a given processing context 1400(1) includes one or more processing sub-contexts 1410(1). The hypervisor 124 configures the processing sub-contexts 1410 and the corresponding device drivers. The processing sub-contexts 1410 associated with a given parent processing context 1400 are launched on the same SMC engine 700 that launched the parent processing context 1400. Thus, in the example shown, the processing sub-context 1410(0) is launched on SMC engine 700(0), and the processing sub-context 1410(1) is launched on SMC engine 700(1). In one embodiment, each guest OS 916 can configure its respective PPU partition 600 independently of the hypervisor 124 and without disturbing the configuration of other PPU partitions 600.

[0162] In some embodiments that do not use virtualization, the hypervisor 124 and the guest OS 916 may not be present, and the host OS 926 may configure and launch the processing context 1400 and the processing sub-context 1410, as described in more detail below in conjunction with Figure 14B More detailed description.

[0163] Figure 14B Illustrated is how a host OS can launch multiple processing contexts simultaneously within one or more PPU partitions according to various embodiments. As shown, the host OS 926 includes a processing context 1400 and processing sub-contexts 1410. In the embodiment shown, the host OS 926 is configured to launch the processing context 1400 and the processing sub-contexts 1410 on the SMC engine 700 without involving the hypervisor or other virtualization software. The embodiment shown can be implemented in a "bare metal" scenario.

[0164] Generally referring to Figures 14A - 14B, processing tasks associated with processing sub-contexts 1410 in the same parent processing context 1400 are generally not scheduled independently of each other and generally share the resources of the corresponding SMC engine 700. Additionally, in some cases, a processing sub-context 1410 initiated in a given SMC engine 700 may result in a failure and / or error that causes the SMC engine 700 to reset any associated processing context 1400 and / or the processing sub-context 1410 to be restarted. A local virtual address space identifier is assigned to the processing context 1400 and / or the processing sub-context 1410, which is derived from the global virtual address space identifier 1510 associated with the PPU 200 as a whole, as described in more detail below in conjunction with Figure 15 described in more detail.

[0165] In some embodiments, there is no virtualization and thus no hypervisor, but it will be clear to those skilled in the art that the single OS usage model on the PPU 200 or a group of PPU 200s can use all of the mechanisms belonging to the VM as described. In some embodiments, a container corresponds to the description of a VM, which means that a container on a single OS can obtain the processing isolation provided to the VM in the present description. Examples of containers are LXC (Linux Containers) and Docker containers, which are well known in the computer industry. For example, each Docker container can correspond to a PPU partition 600, and thus the present invention provides isolation between multiple Docker containers running under one OS.

[0166] Figure 15 illustrates how a hypervisor according to various embodiments Figure 1 assigns virtual address space identifiers to different SMC engines. As shown, the virtual address space identifier 1500 includes separate virtual address ranges for each SMC engine 700. Each virtual address range starts with zero (0) to maintain consistency between the SMC engines 700, but each virtual address range corresponds to a different portion of the global virtual address space identifier 1510. For example, the virtual address space identifier 0 - 15 assigned to SMC engine 700(0) corresponds to the global virtual address space identifier 0 - 15, but the virtual address space identifier 0 - 15 assigned to SMC engine 700(1) corresponds to the global virtual address space identifier 16 - 31. In one embodiment, the global set of global virtual address space identifiers 1510 can be a virtual address space or a physical address space. In some embodiments, there are also virtual address space identifiers for each PPU partition so that the guest OS of a VM has a set of virtual address space identifiers starting from zero for all SMC engines 700 it owns.

[0167] The hypervisor 124 allocates a certain range of virtual address space identifiers to a given SMC engine 700 based on the number of PPU slices 610 from which the SMC engine 700 has been allocated resources. In the example shown, the hypervisor 124 allocates virtual address space identifiers 0 - 15 to SMC engine 700(0), virtual address space identifiers 0 - 15 to SMC engine 700(1), and virtual address space identifiers 0 - 15 to SMC engine 700(4). The hypervisor 124 allocates 16 virtual address space identifiers to SMC engines 700(0), 700(1), and 700(4) because these SMC engines extract resources from two PPU slices 610, as Figure 7 shown. In contrast, the hypervisor 124 allocates virtual address space identifiers 0 - 7 to SMC engines 700(6) and 700(7) because these SMC engines 700 extract resources from one PPU slice 610 respectively. The hypervisor 124 can further subdivide the virtual address space identifiers allocated to a given SMC engine 700 to support multiple processing contexts 1400. For example, the hypervisor 124 can subdivide the virtual address space identifiers 0 - 15 allocated to SMC engine 700(0) into two ranges 0 - 7 and 0 - 7, each of which can be allocated to a different processing context 1400. This example shows how the global virtual address space identifiers 1510 are distributed proportionally in the composition of 0 - 15, 16 - 31, 32 - 47, 48 - 55, and 55 - 63. In some embodiments, the virtual address space identifiers are unique, so the above example would have virtual space identifiers 0 - 15, 16 - 31, 32 - 47, 48 - 55, and 55 - 63 instead of 0 - 15, 0 - 15, 0 - 15, 0 - 7, and 0 - 7, as Figure 15 shown. In certain embodiments, the allocation of the global virtual address space identifiers 1510 is not proportional to the number of PPU slices 610, and the hypervisor can freely allocate any subset of the global virtual address space identifiers 1510 to the PPU partitions 600 or SMC engines 700.

[0168] The hypervisor 124 allocates virtual address space identifiers in the described manner to allow different SMC engines 700 to execute processing tasks associated with any given processing context 1400 without remapping the virtual addresses specified by those processing tasks. Thus, the hypervisor 124 can dynamically migrate processing contexts 1400 between SMC engines 700 without significant changes to those processing contexts. During the execution of various processing tasks associated with a given processing context 1400, any given SMC engine 700 sometimes encounters a failure and is configured to report these failures using locally allocated virtual addresses, as described below in connection with Figure 16As will be described in more detail. After the migration occurs, the migrated processing context still uses the same virtual address space identifier 1500, but these identifiers may correspond to different global virtual address space identifiers 1510.

[0169] Figure 16 Illustrates how a memory management unit converts a local virtual address space identifier 1500 to a global virtual address space identifier 1510 when mitigating a fault, according to various embodiments. As shown, during execution, as previously discussed, the SMC engine 700 may experience faults and / or errors and crash independently of each other. In the example shown, the SMC engine 700(1) encounters an error and causes a local fault identifier 1610 to be output to the memory management unit (MMU) 1600. The access by the SMC engine 700 to an unmapped page causes the MMU to generate a fault and also causes a local fault identifier.

[0170] The MMU 1600 maintains a mapping between the local virtual address space identifier 1500 and the global virtual address space identifier 1510. Based on this mapping, the MMU 1600 generates a global fault identifier 1620 and sends the global fault identifier 1620 to the guest OS 916(0). In response to receiving the global fault identifier 1620, the guest OS 916(0) can reset the SMC engine 700(1) without interrupting the operation of any other SMC engine 700 and then restart the processing context 1400(1). By this method, each SMC engine 700 runs with a different set of virtual address space identifiers, which start from zero and span a potentially similar range, but correspond to different parts of the global memory. Thus, the global virtual address space identifier 1510 can be partitioned among the SMC engines 700, but retains the appearance of a dedicated address space. In some embodiments, for the entire PPU partition 600, the fault identifier 1620 may start from zero. In other embodiments, the fault identifier 1620 may be the identifier of the SMC engine 700 and the virtual address space identifier 1500.

[0171] In one embodiment, the global fault identifier 1620 can be reported to the hypervisor 124, and the hypervisor 124 can perform various operations to resolve the associated fault. In another embodiment, some types of faults can be reported to the associated guest OS 916, while other types of faults (e.g., hard errors occurring within the top portion 810 or bottom portion 830 of the DRAM 272) can be reported to the hypervisor 124. In response to such a fault, the hypervisor 124 can reset some or all of the SMC engines 700. In various other embodiments, a given global fault identifier 1620 can be virtualized and thus not directly correspond to a true global identifier. In operation, the MMU 1600 can route faults to the appropriate VM based on the GFID associated with those VMs. The GFID was discussed above in connection with Figure 7 and is discussed below in connection with

[0172] Generally referring to Figures 15 - 16 , the hypervisor 124 can implement techniques similar to the above-described techniques to assign identifiers to the various hardware resources associated with each PPU partition 600 and / or each SMC engine 700. For example, the hypervisor 124 can assign a local GPC identifier (GPC ID) within a local GPC ID range starting from zero (0) to each GPC 242 included in a given PPU partition 600. Each local GPC ID will correspond to a different global GPC ID. This method can be implemented with any PPU resource in order to maintain a set of identifiers that is internally consistent within any given PPU partition 600 and / or SMC engine 700. As described above, this method facilitates the migration of the processing context 1400 between the SMC engines 700 and further allows the processing context 1400 to migrate between different PPUs 200.

[0173] When the hypervisor 124 migrates the processing context 1400 between different SMC engines 700 residing on different PPUs 200, the hypervisor 124 performs a technique herein referred to as "soft floor sweeping" in order to configure the target PPU 200 with hardware resources similar to those of the source PPU 200. As described in more detail below in connection with Figure 17 and is shown in

[0174] Figure 17 which illustrates, according to various embodiments, when migrating the processing context between SMC engines on different PPUs Figure 1How the hypervisor implements soft floor cleaning. As shown, the computing environment 1700(0) includes an instance of the hypervisor 124(0) and the PPU partition 600(0). The PPU partition 600(0) is configured with the SMC engine 700(0). The SMC engine 700(0) executes processing tasks associated with the processing context 1710. Resources 1720(0) and 1720(1) are allocated to the SMC engine 700(0), but the resource 1720(1) is not operative. Thus, during manufacturing, the resource 1720(1) is fused ("floor cleaned"). The resource 1720 can be any computing, graphics, or memory resource described so far. For example, a given resource 1720 can be a GPC 242, a TPC 330 within the GPC 242, an SM 332 within the TPC 330, a GFX FE 542, or an L2 cache slice 800, etc.

[0175] In various cases, the hypervisor 124(0) can determine that the processing context 1710 should be migrated from the computing context 1700(0) to the computing context 1700(1). For example, the computing context 1700(0) can be scheduled for planned downtime, and in order to maintain continuous service, the hypervisor 124(0) determines that when the computing context 1700(0) is unavailable, the processing context 1710 should be migrated at least temporarily to a different computing environment.

[0176] In such a case, the hypervisor 124(0) interacts with the corresponding hypervisor 124(1) executing in the computing environment 1700(1) to configure the PPU partition 600(1) to provide the same or similar resources as the PPU partition 600(0). As shown, the PPU partition 600(1) includes resources 1720(2) and 1720(3), but in order to mimic the amount of resources provided by the PPU partition 600(0), 1720(3) is made unavailable. Thus, the processing context 1710 can be migrated from the SMC engine 700(0) within the PPU partition 600(0) to the SMC engine 700(1) within the PPU partition 600(1) with no significant change in quality of service. This method helps maintain the appearance of any given PPU partition 600 operating in a manner similar to the PPU 200 by providing access to a consistent set of resources while also allowing the migration of the processing context between different hardware. The hypervisor 124 can also implement the above method to migrate the SMC engine 700 between partitions 600 within the same PPU 200. In one embodiment, the hypervisors 124(0) and 124(1) can execute as a unified software entity that manages the operation of multiple PPUs 200 in different computing environments 1700.

[0177] Generally referring to Figures 11 - 17, the hypervisor 124 implements the above techniques to partition PPU resources in a manner that supports the concurrent execution of processing tasks associated with multiple processing contexts. The following describes these techniques in greater detail in conjunction with Figure 18 More specifically.

[0178] Figure 18 is a flowchart of method steps for configuring computing resources within a PPU to concurrently support operations associated with multiple processing contexts according to various embodiments. Although the method steps are described in the context of Figures 1 - 17 a system, those skilled in the art will understand that any system configured to execute the method steps in any order falls within the scope of this embodiment.

[0179] As shown, method 1800 begins at step 1802, where Figure 1 the hypervisor 124 of

[0180] evaluates PPU 200 to determine a set of available hardware resources. Certain hardware resources may sometimes not be properly fabricated during the manufacture of a given PPU 200 and may be non-functional. In fact, these non-functional hardware resources have been fused and are not used. However, other hardware resources within a given PPU 200 are functional and thus the entire PPU 200 can still operate, albeit with reduced performance. Salvaging partially functional PPUs and other types of units in the described manner is known in the art as "floor sweeping." Figure 11 As described above, a given swizID defines a set of hardware boundaries that can be enabled and disabled to isolate different groups of PPU slices 610 within PPU 200 to form PPU partitions 600. In the case where certain hardware resources are unavailable, the hypervisor 124 determines that certain swizIDs correspond to infeasible partition configurations and should thus be made unavailable.

[0181] At step 1806, the hypervisor 124 generates a set of swizIDs based on the partition input. For example, the hypervisor 124 may receive input from an administrator user indicating a set of partition options and then map those partition options to a corresponding set of swizIDs derived from the set of available swizIDs determined at step 1804. Optionally, the hypervisor 124 may directly receive the set of swizIDs from the administrator user and then modify any swizIDs not included in the set of available swizIDs.

[0182] In step 1808, the hypervisor 124 configures a set of boundaries among the hardware resources based on the swizID group generated in step 1806. In this case, the hypervisor 124 computes a logical OR across the swizID group to generate a configured swizID (or “local” swizID) that indicates which boundary options should be activated as boundaries and which boundary options should be disabled. The exemplary set of partitioning options and corresponding boundary options were described above in conjunction with Figure 12 a described set of exemplary partitioning options and corresponding boundary options.

[0183] In step 1810, the guest OS 916 starts a set of processing contexts in the PPU partition 600 assigned to the guest user, at least in part, based on one or more swizIDs. The hypervisor 124 assigns a set of virtual address space identifiers 1500 to a portion of the global virtual address space identifier 1510 corresponding to Figure 15 the PPU partition 600. The hypervisor 124 or the SMC engine 700 within the PPU partition 600 can subdivide the set of virtual address space identifiers 1500 into different ranges, which are in turn assigned to different processing contexts. This approach allows each processing context to operate using a consistent set of virtual address spaces across all SMC engines 700, thereby allowing for easier migration of processing contexts.

[0184] In step 1812, the hypervisor 124 or the corresponding guest OS 916 resets a subset of the processing contexts started in step 1810 in response to one or more faults. These faults can occur at the execution unit level, the SMC engine level, the VM level, etc. Importantly, faults that occur during the execution of a processing task associated with one processing context typically do not affect the execution of processing tasks associated with other processing contexts. This fault isolation between processing contexts specifically addresses the problems found in prior art approaches that rely on processing sub-contexts. Optionally, between steps 1810 and 1812, a debugger can be invoked to control the SMC engine 700 that encountered the fault.

[0185] In step 1814, the hypervisor 124 configures a migration target based on the available hardware resources associated with the PPU 200. The migration target can be another PPU 200, but in some cases, the migration target 200 is another SMC 700 within a given PPU partition 600 or another PPU partition 600 within the PPU 200. When configuring the migration target, the hypervisor 124 can perform a technique herein referred to as “soft floor sweeping” to cause the migration target to provide hardware resources similar to those used by a set of processing contexts..

[0186] In step 1816, the hypervisor 124 migrates a set of processing contexts to a migration target. Processing tasks associated with those processing contexts can continue with few or no interruptions and can continue using similar available hardware resources. Thus, these techniques allow for providing balanced quality of service in situations where processing contexts need to move between different PPU partitions 600 or different PPUs 200.

[0187] Generally referring to Figures 11 - 18 , the hypervisor 124, guest OS 916, and / or host OS 926 perform the disclosed techniques to divide the various computing resources associated with the PPU 200 into isolated and independent PPU partitions 600 in which different processing contexts can be simultaneously activated. Thus, the resources of the PPU 200 can be utilized more efficiently compared to traditional methods that may support one processing context at a time with incomplete utilization of the PPU. Different PPU partitions 600 can also be accessed and configured independently by multiple different tenants. Thus, the disclosed techniques provide reliable support for multi-tenancy and can therefore meet the consumer demand for an effective cloud-based parallel processing platform.

[0188] Partitioning memory resources to support multiple contexts simultaneously

[0189] In addition to partitioning the computing resources associated with the PPU 200 to support multiple processing contexts simultaneously, the hypervisor 124 also partitions the memory resources associated with the PPU 200 to support multiple contexts simultaneously, thus providing reliable support for multi-tenancy. The hypervisor 124 implements various techniques when partitioning the memory resources associated with the PPU 200, which are described in more detail below in conjunction with Figures 19 - 24 more detail.

[0190] Figure 19 Illustrates a set of boundary options according to various embodiments, according to which Figure 1 the hypervisor can generate one or more PPU memory partitions. As shown, the DRAM 272 includes a set of boundary options 1900 that can be activated during partitioning to divide the L2 cache into individual sections and partitions.

[0191] Specifically, the boundary options 1900(0), 1900(1), 1900(9), and 1900(10) divide the DRAM 272 into Figure 8AThe top portion 810, the partitionable portion 820, and the bottom portion 830. The boundary option 1900(0) forms the lower boundary of the bottom portion 830, and the boundary option 1900(1) forms the upper boundary of the bottom portion 830. The boundary option 1900(1) also forms the lower boundary of the partitionable portion 820 and the left boundary of the partitionable portion 820. The boundary option 1900(9) forms the right boundary of the partitionable portion 820 and the upper boundary of the partitionable portion 820. The boundary option 1900(9) also forms the lower boundary of the top portion 810, and the boundary option 1900(10) forms the upper boundary of the top portion 810. The boundary options 1900(1) to 1900(8) further subdivide the partitionable portion 820 into various memory partitions, which will be described in more detail in conjunction with Figure 20 as follows.

[0192] Also as shown in the figure, the total size of the DRAM 272 is M, the total size of the top portion 810 is T, the total size of the partitionable portion 820 is P, and the total size of the bottom portion 830 is B. Further, the portion of a given cache slice corresponding to the partitionable portion 820 is given by F, and the portion of a given cache slice corresponding to the bottom portion is given by W. F and W are configurable parameters that can be set by the hypervisor 124, and in some embodiments, they can completely constrain the values of T, P, and B relative to M.

[0193] During configuration, the hypervisor 124 configures the DRAM 272 into the top portion 810, the partitionable portion 820, and the bottom portion 830 based on M, F, and W. In this way, the hypervisor 124 determines the values of T, P, and B based on M, F, and W. The hypervisor 124 also activates a specific boundary option 1900 based on a configuration swizID generated via interaction with an administrator user, as described above in conjunction with Figures 11 - 12 stated. The exemplary activation of the boundary option 1900 will be described in conjunction with Figure 20 below.

[0194] Figure 20 illustrates according to various embodiments Figure 1An example of how a hypervisor partitions the PPU memory to generate one or more PPU memory partitions. As shown, boundary options 1900(0), 1900(1), 1900(9), and 1900(10) are activated, forming the top portion 810, the partitionable portion 820, and the bottom portion 830 of DRAM 272. Boundary options 1900(1), 1900(5), 1900(7), and 1900(8) are also activated, forming PPU memory partitions 710(0), 710(4), 710(6), and 710(7) corresponding to DRAM partitions 822(0), 822(4), 822(6), and 822(7) within the partitionable portion 820, respectively. Boundary options 1900(2), 1900(3), 1900(4), and 1900(6) are not activated and are thus omitted. The hypervisor 124 configures DRAM 272 in the manner shown based on a configuration swizID equal to "11110100011".

[0195] The boundary options 1900 associated with DRAM 272 logically correspond to Figures 11 - 12 the boundary options 1110 shown in Figures 11 - 12 As discussed above in connection with Figure 20 each bit of a given configuration swizID indicates whether the corresponding boundary option 1110 should be activated or deactivated to group the PPU slices 610 together. In a similar manner, as

[0196] shown, each bit in the exemplary configuration swizID "11110100011" indicates whether the corresponding boundary option 1900 associated with DRAM 272 should be activated or deactivated.

[0196] By default, bits 0 and 10 of the exemplary swizID are set to 1 to activate boundary options 1900(0) and 1900(10). Bits 1 and 9 of the exemplary swizID are set to 1 to activate boundary options 1900(1) and 1900(9) and establish the partitionable portion 820. Bits 5, 7, and 8 of the exemplary swizID are set to 1 to activate boundary options 1900(5), 1900(7), and 1900(8) and divide the partitionable portion 820 into DRAM portions 822 associated with the PPU memory partitions 710. The other bits of the swizID are set to zero to deactivate the corresponding boundary options. The partitioning of DRAM 272 shown here corresponds to Figure 12 the exemplary configuration of PPU slices 610 shown in Figures 21 - 23 Once partitioned in this manner by the hypervisor 124, the SMC engine 700 executing within the PPU partition 600 can perform memory access operations via the L2 cache slice 800 in the manner described below in connection with Figures 21 - 23

[0197] Figure 21 illustrates how a memory management unit according to various embodiments Figure 16 provides access to different PPU memory partitions. As shown, Figure 16 the MMU 1600 is coupled between the DRAM 272 and the 1D SPA space 850. The 1D SPA space 850 is divided into a top address 852 corresponding to the top portion 810, a partitionable address 854 corresponding to the partitionable portion 820, and a bottom address 858 corresponding to the bottom portion 830, as also Figure 8B shown. During partitioning, the hypervisor 124 generates the 1D SPA space 850 based on the configuration of the DRAM 272.

[0198] The MMU 1600 includes an address mapping unit (AMAP) 2110 that is configured to map the top address 852, the partitionable address 854, and the bottom address 858 to the original addresses associated with the top portion 810, the partitionable portion 820, and the bottom portion 830, respectively. In this way, the MMU 1600 services memory access requests received from the hypervisor 124 that target the top portion 810 and / or the bottom 830 portion, and memory access requests received from the SMC engine 700 that target the partitionable portion 820, as described in more detail below in conjunction with Figure 22 is described.

[0199] Figure 22 illustrates how a memory management unit according to various embodiments Figure 16 performs various address translations. As shown, the partitionable address 854 includes an address region 856(0) that includes addresses corresponding to the PPU memory partition 710(0), as described above in conjunction with Figure 8B is discussed. The MMU 1600 converts the physical addresses included in the address region 856(0) to the original addresses associated with the DRAM portion 822(0) via the AMAP 2110. The AMAP 2110 is configured to rearrange the addresses from the address region 856(0) on the L2 cache slice 800(0) included in the PPU memory partition 710(0) to avoid situations where striding causes the same L2 cache slice 800(0) to be accessed repeatedly (also known as "camping").

[0200] In one embodiment, the AMAP 2110 can implement a "memory access" swizID that identifies the memory interleaving factor for a given memory region. A given memory access swizID determines a set of L2 cache slices that are cross-interleaved for various types of memory access, including video memory, system memory, and peer memory access. Different PPU partitions 600 typically implement different and non-overlapping memory regions 822 within the partitionable section 829 to minimize interference between concurrently executing jobs. The hypervisor 124 can use a memory access swizID of zero to balance memory access operations across the L2 cache slices, which will typically access either the top section 810 or the bottom section 830.

[0201] A given memory access swizID can be a "local" swizID that is calculated based on the system physical address and is used to interleave or scramble memory access requests across the relevant L2 slices and the corresponding portions of the DRAM. A given local swizID associated with a given PPU partition 600 can correspond to the swizID used to configure that PPU partition. By this method, the AMAP 2110 can scramble addresses within the boundaries of a given PPU memory partition according to the swizID used to activate those boundaries. The scrambled addresses based on the memory access swizID allow the MMU 1600 to interleave the DRAM 272 such that each PPU partition 600 views its partitionable section 820 as contiguous within the linear system physical address space 850. This method can maintain the isolation between PPU partitions 600 and the integrity associated with those PPU partitions.

[0202] A given memory access swizID can alternatively be a "remote" swizID provided by the device driver 122 or the hypervisor 124 and is used to interleave memory access requests across the L2 slices for system memory access operations. For processing operations that occur within a given PPU partition 600, the local swizID and the remote swizID can be the same. Different PPU partitions 600 typically have different remote swizIDs to allow system memory access operations to occur only through the L2 slices 800 that belong to the PPU partition 600.

[0203] The MMU 1600 also provides support for translating virtual addresses associated with the virtual address space identifier 1500 into system physical addresses in the 1D system physical address space 850. For example, assume Figure 7The SMC engine 700(0) executes using the PPU memory partition 710(0) and the corresponding DRAM portion 822(0), and in doing so, causes a memory fault. The MMU 1600 issues a fault with a local fault identifier 1610. The MMU 1600, in turn, converts the local fault identifier 1610 to a global fault identifier 1620. Faults and errors may be reported to the virtual function in accordance with the SR-IOV public specification.

[0204] The MMU 1600 also facilitates subdividing the address regions 856 and PPU memory partitions 710 to provide support for multiple SMC engines 700, multiple VMs, and / or multiple processing contexts 1400 executing within a given PPU partition 600, as described below in conjunction with Figure 23 Described in more detail.

[0205] Figure 23 shows various embodiments Figure 16 How the memory management unit of the SMC engine 700(0) provides support for operations associated with multiple processing contexts simultaneously. As shown, the address region 856(0) contains multiple virtual memory pages 2310 of different sizes. For the SMC engine 700(0), the virtual memory space identifier 1500 is mapped to the global virtual address space identifier 1510, which selects a page table for the specific virtual address space used by the processing context on the SMC engine 700(0). The page specified by page table A selects page 2310(A) within the DRAM portion 822(0). At the same time, the SMC engine 700(2) can use the page specified by page table B, which also selects page 2310(B) within the DRAM portion 822(0). Through the page-based virtual memory management scheme, pages in the DRAM portion 822(0) can be assigned to different subcontexts or different processing contexts. Note that a processing context can use multiple virtual address space identifiers 1500 because it can execute many subcontexts.

[0206] Subdividing the address region 856(0) and the DRAM portion 822(0) corresponding to the PPU memory partition 710(0) in the manner shown provides dedicated memory resources within the PPU memory partition 822(0) for different SMC engines 700 executing within the corresponding PPU partition 600. Thus, multiple SMC engines 700 in different PPU 600 partitions can simultaneously execute processing tasks within different processing contexts without interfering with each other in terms of bandwidth.

[0207] The above-described page-based approach can also be applied to a single SMC engine 700 that executes multiple processing sub-contexts, where each processing sub-context requires a dedicated portion of the PPU memory partition 710(0). Similarly, the above method can be applied to different VMs executing on one or more SMC engines 700 and requiring dedicated portions of the PPU memory partition 710(0).

[0208] Generally referring to Figures 19 - 23 , the disclosed techniques allow a given PPU partition 600 to be configured in the manner described above in combination with Figures 11 - 12 to simultaneously and securely start multiple processing contexts. In particular, as described, partitioning the L2 cache fairly allocates the DRAM portion 822 to different PPU partitions 600. Additionally, the various address translations implemented via the MMU 1600 and AMAP 2110 effectively and fairly utilize the memory bandwidth, thereby providing a consistent quality of service to all tenants of the PPU 200. The techniques described in combination with Figure 24 are also described in more detail below in combination with Figures 19 - 23 .

[0209] Figure 24 is a flowchart of method steps for configuring memory resources within a PPU to simultaneously support operations associated with multiple processing contexts according to various embodiments. Although the method steps are described in the context of the Figures 1 - 23 system, those skilled in the art will understand that any system configured to execute the method steps in any order falls within the scope of this embodiment. [[ID=...]]

[0210] As shown, method 2400 begins at step 2402, where Figure 1 the hypervisor 124 of

[0211] determines a set of memory configuration parameters for partitioning the DRAM 272. The set of memory configuration parameters can indicate any technically feasible set of parameters that describe any attribute of the DRAM 272, including the total size (M) of the DRAM 272, the size T of the top portion 810, the size P of the partitionable portion 820, the size B of the bottom portion 830, the number of L2 cache slices 800 (e.g., 96), the size F of each cache slice portion corresponding to the partitionable portion 820, and / or the size W of each cache slice portion corresponding to the bottom portion 830. In one embodiment, a set of configuration parameters only needs to include the parameters M, F, and W.

[0211] At step 2404, the hypervisor 124 activates a first set of boundary options based on the set of memory configuration parameters determined in step 2402 to divide the DRAM 272 into multiple portions. In particular, the hypervisor 124 activates Figure 19The boundary options 1900(0), 1900(1), 1900(9), and 1900(10) shown are used to divide the DRAM 272 into a top portion 810, a partitionable portion 820, and a bottom portion 820. In one embodiment, the hypervisor 124 can modify the position of a given boundary option to adjust the size of the corresponding portion of the DRAM 272.

[0212] In step 2406, the hypervisor 124 determines a configuration swizID based on the partitioning input. The partitioning input can be obtained via step 1002 of method 1000 described above in connection with Figure 10 The partitioning input indicates a set of target PPU partitions of the PPU 200. The hypervisor 124 can determine the configuration swizID of the set of target partitions through the techniques described above in connection with Figures 11 - 12 In one embodiment, the configuration swizID can be obtained before step 2404, and then the first set of boundary options can be activated based on the configuration swizID.

[0213] In step 2408, the hypervisor activates a second set of boundary options based on the configuration swizID to generate one or more PPU memory partitions 710 within the partitionable portion 820 of the DRAM 272. The second set of boundary options can include Figure 19 Any of the boundary options 1900(2) to 1900(8) shown in. These different boundary options can divide the partitionable portion 820 into multiple DRAM portions 822 corresponding to the PPU memory partitions 710, which are equal to or less than the number of PPU slices 610. In various embodiments, steps 2404 and 2408 of method 2400 can be performed in combination with each other based on a configuration swizID obtained via or generated based on an administrator user input.

[0214] In step 2410, the hypervisor 124 determines a set of partitionable addresses 854 based on a set of memory configuration parameters and / or the configuration swizID. In this way, the hypervisor 124 divides the 1D SPA space 850 into a top address 852, partitionable addresses 854, and a bottom address 858, as Figure 21 shown. The top address 852 can be converted to an original address associated with the top portion 810, the partitionable addresses 854 can be converted to original addresses associated with the partitionable portion 820, and the bottom address 858 can be converted to an original address associated with the bottom portion 830.

[0215] In step 2412, the MMU 1600 services a memory access request by obfuscating the allocatable address 854 on the L2 cache slice 800 within the PPU memory partition 710 corresponding to the DRAM section 822. The MMU 1600 obfuscates the allocatable address via the AMAP 2110 based on the memory access swizID (or "remote" swizID) associated with the PPU memory partition 710. In one embodiment, the memory access swizID is derived from the swizID according to which the PPU memory partition 710 is configured. Obfuscating the partitionable address 854 in this way can reduce repeated access to the respective L2 cache slices 800 (also referred to as "camping").

[0216] In step 2414, the MMU 1600 receives a local fault identifier associated with a memory fault and converts the local fault identifier to a global fault identifier. In this way, the MMU 1600 can convert the virtual address associated with the local fault identifier to a global address associated with the global fault identifier. For example, when a given SMC engine 700 encounters an error during a memory read operation or a memory write operation on the PPU memory partition 710, it may result in a memory fault. Implementing the fault ID in the local virtual address space allows the SMC engine 700 to operate in a similar address space, thus allowing for a simpler migration of the processing context between the PPU partitions 600, as described above in conjunction with Figure 17 what has been described. Converting those fault IDs to the global fault identifier 1620 allows the hypervisor 124 to resolve the faults corresponding to the fault IDs from the perspective of the global virtual address space identifier 1510.

[0217] Generally referring to Figures 19 - 24 what has been described above in conjunction with Figures 11 - 18 the techniques for partitioning PPU memory resources disclosed herein complete the techniques for partitioning PPU computing resources described above in conjunction with

[0218] Time slicing multiple VMs and processing contexts

[0219] As further discussed herein, Figure 2The PPU 200 supports two levels of partitioning. In the first partitioning level, referred to herein as the "PPU partition", the PPU resources 500 of the PPU 200 are divided into PPU partitions 600, also referred to herein as "partitioned PPUs". In some embodiments, one or both of the PP memory 270 and the DRAM 272 may be divided into SMC memory partitions 710. At this partitioning level, each PPU partition 600 executes one VM at any given time. In the second partitioning level, referred to herein as the "SMC partition", each PPU partition 600 is further divided into SMC engines 700. Each PPU partition 600 includes one or more SMC engines 700. At this partitioning level, each SMC engine 700 executes one processing context for one VM at any given time.

[0220] Over time, the SMC engine 700 switches from executing a specific processing context for a VM to executing a different processing context for the same VM or a different processing context for a different VM. Since the execution time of the SMC engine 700 is "sliced" among multiple processing contexts corresponding to one or more VMs, this process is referred to herein as "time slicing".

[0221] Each SMC engine 700 time slices between the processing contexts listed on the run list, as managed by the PBDMAs 520 and 522 of the SMC engine 700. Generally, when switching between VMs, the run lists on all affected SMC engines 700 are replaced to time slice different sets of processing contexts. If multiple SMC engines 700 are active, the run lists are replaced simultaneously. This type of scheduling via run list replacement is referred to herein as "software scheduling". Switching between processing contexts within the same VM may similarly involve replacing the run list and is very similar to switching VMs, except that the VM does not change due to the context switch. As a result, this type of context switch does not require additional hardware support. In some embodiments, a VM may have many different processing contexts of various sizes. In such embodiments, software scheduling can consider how to pack these processing contexts into the SMC engines 700 for correct and efficient execution. Additionally, software scheduling can reconfigure the number of PPU partitions 600 and the number of SMC engines within each PPU partition 600 to correctly and efficiently execute the processing contexts for the VMs.

[0222] To enable the PPU resource 500 to support the above two partitioning levels, the PPU 200 correspondingly supports two time-slice levels. Corresponding to the PPU partitioning, the PPU 200 performs VM-level time slicing, where each PPU partition 600 time-slices among multiple virtual machines. Corresponding to the SMC partitioning, the PPU 200 performs SMC-level time slicing, where each VM 600 time-slices among multiple processing contexts. In various embodiments, the two time-slice levels maintain the same number of TPCs in each GPC 242 over time. In various embodiments, time slicing can involve changing the number of TPCs in one or more GPCs 242 over time. Now, the two levels of time slicing are described.

[0223] Figure 25 is the timeline set 2500, which shows, according to various embodiments, the VM-level time slicing associated with the Figure 2 PPU partition 600 of the PPU 200. As shown, the timeline set 2500 includes, but is not limited to, four PPU partition timelines 2502(0), 2502(4), 2502(6), and 2502(7). In some embodiments, the PPU partition timelines 2502(0), 2502(4), 2502(6), and 2502(7) can respectively correspond to Figure 6 the PPU partitions 600(0), 600(4), 600(6), and 600(7). In such embodiments, the PPU partition timeline 2502(0) can correspond to PPU slices 610(0)-610(3), and the PPU partition timeline 2502(4) can correspond to PPU slices 610(4)-610(5). Similarly, the PPU partition timeline 2502(6) can correspond to PPU slice 610(6), and the PPU partition timeline 2502(7) can correspond to PPU slice 610(7). As a result, the PPU partition 600(0) can execute up to four processing contexts simultaneously, the PPU partition 600(4) can execute up to two processing contexts simultaneously, and each of the PPU partitions 600(6) and 600(7) can execute one processing context at a time.

[0224] As shown in PPU partition timeline 2502(0), PPU partition 600(0) time-slices between two VMs, referred to as VMA and VM B. Timeline 2502(0) shows the time-slicing between VM A and VM B, where the processing context associated with VM A is shown in the form of context 2510(Ax-y), and the processing context associated with VM B is shown in the form of context 2510(Bx-y). From time t0 to time t1, PPU partition 600(0) executes the processing context of VM A. The first SMC engine 700(0) of PPU partition 600(0) sequentially executes processing context 2510(A0-1), processing context 2510(A1-1), and processing context 2510(A2-1). At the same time, the second SMC engine 700(2) of PPU partition 600(0) sequentially executes processing context 2510(A3-1), processing context 2510(A4-1), and processing context 2510(A5-1). At time t1, PPU partition 600(0) stops executing the processing context associated with VM A and switches to the processing context associated with VM B. Processing context 2510(A2-1) and processing context 2510(A5-1) are context-switched out of the respective SMC engines 700(0) and 700(2). PPU partition 600(0) is reconfigured from having two SMC engines 700(0) (which includes two GPCs 230(0) and 230(1)) and 700(2) (which includes two GPCs 230(2) and 230(3)) to having one SMC engine 700(0) (which includes four GPCs 230(0), 230(1), 230(2), and 230(3)). Once the reconfiguration is complete, PPU partition 600(0) starts executing the processing context 2510(B0-1) of VM B.

[0225] From time t1 to time t4, the PPU partition 600(0) executes the processing context of VM B. The first SMC engine 700(0) of the PPU partition 600(0) sequentially executes the processing context 2510(B0-1), the processing context 2510(B1-1), and the processing context 2510(B2-1). At time t4, the PPU partition 600(0) stops executing the processing context associated with VM B and switches to the processing context associated with VM A. The processing context 2510(B2-1) is context-switched out of the corresponding SMC engine 700(0). The PPU partition 600(0) is reconfigured from having one SMC engine 700(0) (which includes four GPCs 230(0), 230(1), 230(2), and 230(3)) to having two SMC engines 700(0) (which includes two GPCs 230(0) and 230(1)) and 700(2) (which includes two GPCs 230(2) and 230(3)). Once the reconfiguration is complete, the PPU partition 600(0) starts executing the processing contexts 2510(A2-2) and 2510(A5-2) of VM A. Starting from time t4, the PPU partition 600(0) executes the processing context of VM A again. The first SMC engine 700(0) of the PPU partition 600(0) sequentially executes the processing context 2510(A2-2) and the processing context 2510(A0-2). Meanwhile, the second SMC engine 700(2) of the PPU partition 600(0) sequentially executes the processing context 2510(A5-2) and the processing context 2510(A3-)2).

[0226] As shown in PPU partition timeline 2502(4), PPU partition 600(4) time slices between two VMs, referred to as VMC and VM D. Timeline 2502(4) shows the time slicing between VM C and VM D, where the processing context associated with VM C is shown in the form of context 2510(Cx-y), and the processing context associated with VM D is shown in the form of context 2510(Dx-y). From time t0 to time t2, PPU partition 600(4) is in an idle state and does not execute any processing context. From time t2 to time t5, PPU partition 600(4) executes the processing context of VM C. The first SMC engine 700(4) of PPU partition 600(4) sequentially executes processing context 2510(C0-1) and processing context 2510(C1-1). At time t5, PPU partition 600(4) stops executing processing context 2510(C1-1) and reconfigures to start executing the processing context of VM D. From time t5 onwards, PPU partition 600(4) executes the processing context of VM D. The first SMC engine 700(4) of PPU partition 600(4) executes processing context 2510(D1-1).

[0227] As shown in PPU partition timeline 2502(6), PPU partition 600(6) time slices between two VMs, referred to as VME and VM F. Timeline 2502(6) shows the time slicing between VM E and VM F, where the processing context associated with VM E is shown in the form of context 2510(Ex-y), and the processing context associated with VM F is shown in the form of context 2510(Fx-y). From time t0 to time t3, PPU partition 600(6) is in an idle state and does not execute any processing context. From time t3 to time t6, PPU partition 600(6) executes the processing context of VM E. The first SMC engine 700(6) of PPU partition 600(6) sequentially executes processing context 2510(E0-1) and processing context 2510(E1-1). At time t6, PPU partition 600(6) stops executing processing context 2510(E1-1) and reconfigures to start executing the processing context of VM F. From time t6 onwards, PPU partition 600(4) executes the processing context of VM F. The first SMC engine 700(6) of PPU partition 600(6) executes processing context 2510(F0-1).

[0228] As shown in PPU partition timeline 2502(7), PPU partition 600(7) performs time slicing within a single VM, referred to as VMG. Timeline 2502(6) shows the time slicing of VM G, where the processing context associated with VM G is shown in the form of context 2510(Gx - y). From time t0 to time t5, PPU partition 600(7) is in an idle state and does not execute any processing context. Starting from time t5, PPU partition 600(7) executes processing context G. The first SMC engine 700(7) of PPU partition 600(7) executes processing context 2510(G0 - 1).

[0229] In this way, each of PPU partitions 600(0), 600(4), 600(6), and 600(7) performs time slicing between processing contexts corresponding to one or more VMs. Each of PPU partitions 600(0), 600(4), 600(6), and 600(7) transitions independently from one processing context to another. For example, PPU partition 600(0) can switch from executing one processing context for a particular VM to another processing context for the same or a different VM, regardless of whether any one or more of PPU partitions 600(4), 600(6), and 600(7) switch processing contexts. During the shown time period, each of PPU partitions 600(0), 600(4), 600(6), and 600(7) maintains a constant number of PPU slices 610. In some embodiments, the number of PPU slices 610 for each PPU partition 600 can change, as now described.

[0230] Figure 26 is another set of timelines 2600 according to various other embodiments, which shows VM - level time slicing associated with Figure 2 the PPU 200. The functions of the processing contexts shown in the set of timelines 2600 are substantially the same as Figure 25 the set of timelines 2500, except as further described below. As shown, the set of timelines 2600 includes, but is not limited to, four PPU partition timelines 2602(0), 2602(4), 2602(6), and 2602(7). In some embodiments, PPU partition timelines 2602(0), 2602(4), 2602(6), and 2602(7) can respectively correspond to Figure 6 PPU partitions 600(0), 600(4), 600(6), and 600(7).

[0231] As shown by PPU partition timeline 2602(0), PPU partition 600(0) time slices between two VMs, referred to as VMA and VM B. Timeline 2602(0) shows the time slicing between VM A and VM B, where the processing context associated with VM A is shown in the form of context 2610(Ax-y), and the processing context associated with VM B is shown in the form of context 2610(Bx-y). From time t0 to time t1, PPU partition 600(0) executes the processing context of VM B. The first SMC engine 700(0) of PPU partition 600(0) sequentially executes processing context 2610(B1-1) and processing context 2610(B2-1). At time t1, PPU partition 600(0) stops executing the processing context associated with VM B and switches to the processing context associated with VM A. Processing context 2610(B2-1) is context switched out of the corresponding SMC engine 700(0). PPU partition 600(0) is reconfigured from having one SMC engine 700(0) (which includes four GPCs 230(0), 230(1), 230(2) and 230(3)) to having two SMC engines 700(0) (which includes two GPCs 230(0) and 230(1)) and 700(2) (which includes two GPCs 230(2) and 230(3)). Once the reconfiguration is complete, PPU partition 600(0) starts executing the processing context of VM A. From time t1 to time t3, PPU partition 600(0) executes the processing context of A. The first SMC engine 700(0) of PPU partition 600(0) sequentially executes processing context 2610(A2-1), processing context 2610(A0-1) and processing context 2610(A1-1). At the same time, the second SMC engine 700(2) of PPU partition 600(0) sequentially executes processing context 2610(A5-1), processing context 2610(A3-1) and processing context 2610(A4-1). At time t3, PPU partition 600(0) stops executing the processing context associated with VM A and switches to the processing context associated with VM B. Processing context 2610(A1-1) and processing context 2610(A4-1) are context switched out of the corresponding SMC engines 700(0) and 700(2). PPU partition 600(0) is reconfigured from having two SMC engines 700(0) (which includes two GPCs 230(0) and 230(1)) and 700(2) (which includes two GPCs 230(2) and 230(3)) to having one SMC engine 700(0) (which includes four GPCs 230(0), 230(1), 230(2) and 230(3)).Once the reconfiguration is complete, PPU partition 600(0) starts executing the processing environment of VM B. Starting from time t3, PPU partition 600(0) executes the processing context of VM B again. The first SMC engine 700(0) of PPU partition 600(0) sequentially executes processing context 2610(B2-2) and processing context 2610(B0-1).

[0232] As shown in PPU partition timeline 2602(4), from time t0 to time t2, the first SMC engine 700(4) of PPU partition 600(4) sequentially executes processing context 2610(D0-1) and processing context 2610(D1-1). Then, the first SMC engine 700(4) of PPU partition 600(4) becomes idle. As shown in PPU partition timeline 2602(6), from time t0 to time t2, the first SMC engine 700(6) of PPU partition 600(6) sequentially executes processing context 2610(F0-1) and processing context 2610(F1-1). Then, the first SMC engine 700(6) of PPU partition 600(6) becomes idle. As shown in PPU partition timeline 2602(7), PPU partition 600(7) is idle from time t0 to time t2.

[0233] At time t2, PPU partitions 600(4), 600(6), and 600(7) are merged to form a single PPU partition 600(4) with four SMC engines 700(4)-700(7). As shown in PPU partition timeline 2604(4), the merged PPU partition 600(4) executes the processing context of VM H. Starting from time t2, the first SMC engine 700(4) of PPU partition 600(4) sequentially executes processing context 2610(H0-1), processing context 2610(H1-1), processing context 2610(H2-1), and processing context 2610(H0-2).

[0234] In this way, PPU partitions 600 can be merged and / or sliced into partitions of different sizes during time slicing. PPU partitions 600 can be merged and / or sliced independently of each other. In Figure 25 and 26 , each VM executes on a fixed number of SMC engines 700, resulting in a given VM executing a constant number of processing contexts simultaneously. As shown, VM A executes simultaneously on two SMC engines 700, while the remaining VMs execute on one SMC engine 700 at a time. In some embodiments, a particular VM can change the number of SMC engines 700 on which it executes the VM, as now described.

[0235] Figure 27is the timeline 2700 according to various embodiments, which shows the SMC-level time slices associated with the Figure 2 PPU 200. The processing contexts shown in the timeline 2700 are respectively associated with the Figure 25 and 26 timeline sets 2500 and 2600, except as further described below. As shown, the timeline 2700 represents a single PPU partition timeline. In some embodiments, the PPU partition timeline represented by the timeline 2700 may correspond to any PPU partition 600 of the Figure 6 that includes at least two SMC engines 700.

[0236] As shown in the timeline 2700, the time slices of the PPU partition 600 within a single VM are referred to as VM A. The timeline 2700 illustrates the time slices of VM A, where the processing context associated with VM A is shown in the form of context 2710(Ax-y). Starting from time t0, VM A executes on two SMC engines 700. The first SMC engine 700 included in the PPU partition 600 executes the processing context 2710(A0-1), and then idles until time t1. Meanwhile, the second SMC engine 700 included in the PPU partition 600 executes the processing context 2710(A1-1), and then idles until time t1. The duration between time t0 and time t1 is long enough to ensure that there is sufficient time for the processing contexts 2710(A0-1) and 2710(A1-1) to complete execution and for the SMC engines to enter the idle state. In some embodiments, the processing context 2710(A0-1) and the processing context 2710(A1-1) may perform the same task simultaneously on two separate SMC engines 700, thereby providing spatial redundancy. In such embodiments, the processing context 2710(A0-1) and the processing context 2710(A1-1) may perform tasks on redundant SMC engines 700 having the same configuration as each other, and then compare the accuracy and determinacy of the results.

[0237] Between time t1 and time t2, the run lists for the processing contexts 2710(A0-1) and 2710(A1-1) are removed from the PPU partition 600. The PPU partition 600 is reconfigured from two SMC engines 700 to one SMC engine 700, which includes all the resources of the two SMC engines 700. Then, the PPU partition 600 uses the new run list to execute the processing contexts 2710(B2-1), 2710(B3-1), and 2710(B4-1).

[0238] Between time t2 and time t3, VM B executes on an SMC engine 700. The SMC engine 700 included in the PPU partition 600 sequentially executes processing context 2710 (B2-1), processing context 2710 (B3-1), and processing context 2710 (B4-1). In some embodiments, the processing contexts 2710 (B2-1), 2710 (B3-1), and 2710 (B4-1) may execute performance-intensive tasks, which may benefit from execution on a single SMC engine 700 that has more computing resources than the SMC engines that execute processing contexts 2710 (A0-1) and 2710 (A1-1). The SMC engine 700 then idles until time t3. In some embodiments, the SMC engine 700 executes offline scheduling tasks during this idle period.

[0239] Between time t3 and time t4, the run lists for processing contexts 2710 (B2-1), 2710 (B3-1), and 2710 (B4-1) are removed from the PPU partition 600. The PPU partition 600 is reconfigured from one SMC engine 700 to two SMC engines 700. Each of the two SMC engines 700 contains a portion of the resources included in one SMC engine 700. The PPU partition 600 then uses the new run list to execute processing contexts 2710 (A0-2) and 2710 (A1-2).

[0240] Starting from time t4, VM A executes again on two SMC engines 700. The first SMC engine 700 executes processing context 2710 (A0-2), and then idles until time t5. Meanwhile, the second SMC engine 700 executes processing context 2710 (A1-2), and then idles until time t5. The duration between time t4 and time t5 is long enough to ensure that there is sufficient time for processing contexts 2710 (A0-2) and 2710 (A1-2) to complete execution and for the SMC engines to enter the idle state. In some embodiments, processing context 2710 (A0-2) and processing context 2710 (A1-2) may execute the same task simultaneously on two separate SMC engines 700, thereby providing spatial redundancy. In such embodiments, processing context 2710 (A0-2) and processing context 2710 (A1-2) may execute tasks on redundant SMC engines 700 that have the same configuration as each other, and then compare the accuracy and determinacy of the results.

[0241] Between time t5 and time t6, the run lists for processing contexts 2710(A0-2) and 2710(A1-2) were removed from PPU partition 600. PPU partition 600 was reconfigured from two SMC engines 700 to one SMC engine 700, which includes all the resources of the two SMC engines 700. Then PPU partition 600 used the new run list to execute processing context 2710(B3-2). Starting at time t6, VM B executed again on one SMC engine 700. SMC engine 700 executed processing context 2710(B3-2).

[0242] In some embodiments, PPU partition 600 can be quickly reconfigured between executing on one SMC engine and executing on two SMC engines, a process referred to herein as "quick reconfiguration". Quick reconfiguration increases the utilization of the resources of PPU200 while providing a mechanism for multiple processing contexts to execute on a single PPU partition 600 in different modes. One or both of the core driver 914 and the hardware microcode within PPU 200 include various optimizations for implementing quick reconfiguration. These optimizations are now described.

[0243] During reconfiguration, certain resources in PPU 200, such as FECS 530 and GPC 242 environment contexts, are not reset unless an error occurs in the resources. As a result, loading the microcode into these resources during reconfiguration can be divided into multiple phases. In particular, the microcode loading sequence for FECS 530 and GPC 242 can be divided into a LOAD phase and an INIT phase. The LOAD phase is executed in parallel for all available FECS 530 and GPC 242 environment contexts within PPU partition 600, thereby reducing the time required to load the microcode into these resources. The INIT phase is executed during reconfiguration, thereby initializing the context switch of FECS 530 and GPC 242 in parallel with reconfiguring PPU partition 600. During the initialization phase, PPU 200 ensures that the LOAD phase for all FECS 530 and GPC 242 environment contexts has been completed. As a result, the time to load and initialize the resources of PPU partition 600 is reduced. Additionally, PPU 200 stores a cache of standardized processing context images for each possible configuration of PPU partition 600, which is referred to herein as the "golden processing context image". The appropriate golden processing context image is retrieved and loaded during the LOAD and INIT phases, thereby further reducing the time to load and initialize the resources of PPU partition 600.

[0244] As described herein, a particular VM may change over time the number of SMC engines 700 on which the VM executes. In one particular example, VM A includes various tasks associated with an autonomous vehicle. Some tasks of the autonomous vehicle are more critical than others. For example, tasks associated with autonomous driving (such as detecting traffic signals and avoiding collisions) will be considered more critical than tasks associated with the vehicle's entertainment system. These more critical tasks may need to comply with certain regulations or industry standards. One such standard assigns classification levels known as Automotive Safety Integrity Levels (ASILs). To increase the level of integrity, ASIL includes four levels, namely ASIL-A, ASIL-B, ASIL-C, and ASIL-D. Tasks such as detecting traffic signals and avoiding collisions will be classified as ASIL-D. Less important tasks can be classified into lower ASIL levels. Tasks not related to safety (such as tasks associated with the vehicle's entertainment system) can be classified as QM, which indicates that only standard quality management specifications apply.

[0245] In this regard, processing contexts 2710(A0-1) and 2710(A1-1) may include two instances of the same ASIL-D level task that are executed simultaneously on two different SMC engines 700 of the PPU partition 600. After processing contexts 2710(A0-1) and 2710(A1-1) complete execution, the results of processing contexts 2710(A0-1) and 2710(A1-1) are compared with each other. If processing context 2710(A0-1) and processing context 2710(A1-1) generate the same result, the result has been verified and the vehicle will continue to drive based on the result. On the other hand, a failure in one or more components associated with processing context 2710(A0-1) or processing context 2710(A1-1) may cause the affected processing context to generate an incorrect result. Therefore, if processing context 2710(A0-1) and processing context 2710(A1-1) generate different results, the result is invalid and the vehicle will perform an appropriate evasive action, such as slowly moving from the traffic flow to the closest position.

[0246] After completion of execution of processing contexts 2710(A0-1) and 2710(A1-1), PPU partition 600 is reconfigured to include only one SMC engine 700. The SMC engine 700 in turn executes QM-level processing contexts 2710(B2-1), 2710(B3-1) and 2710(B4-1). These processing contexts include less critical tasks such as tasks associated with the vehicle's entertainment system. After completion of processing contexts 2710(B2-1), 2710(B3-1) and 2710(B4-1), PPU partition 600 is reconfigured to include two SMC engines 700. The SMC engines 700 concurrently execute processing contexts 2710(A0-2) and 2710(A1-2), which are two instances of the same ASIL-D level task. After completion of execution of processing contexts 2710(A0-2) and 2710(A1-2), PPU partition 600 is again reconfigured to include only one SMC engine 700 and executes QM-level processing context 2710(B3-2).

[0247] In this way, PPU partition 600 is dynamically reconfigured between multiple SMC engines 700 that execute ASIL-D level tasks and a single SMC engine 700 that executes QM-level tasks. The duration between consecutive ASIL-D processing contexts (e.g., the duration between time t0 and time t4) is referred to as a "frame", where the portion of the frame between time t0 and time t1 is allocated for execution of ASIL-D tasks.

[0248] Figure 28 Shows how a VM migrates from one PPU 200(1) to another PPU 200(2) according to various embodiments. As shown in FIG. 2800, PPU 200(1) executes four VMs 2810A, 2810B, 2810C and 2810D. Each of these VMs 2810A, 2810B, 2810C and 2810D executes on a different SMC engine 700 included in PPU 200(1). Similarly, PPU 200(2) executes four VMs 2810E, 2810F, 2810G and 2810H. Each of these VMs 2810E, 2810F, 2810G and 2810H executes on a different SMC engine 700 included in PPU 200(1).

[0249] Over time, for various reasons, a VM may migrate from one PPU 200 to another PPU 200, including but not limited to preparing for system maintenance, consolidating VMs onto fewer PPU 200s to improve utilization or save power, and gaining efficiency by migrating to a different data center. In a first example, when the system on which a VM is currently executing is to be shut down for system maintenance, the VM can be forced to migrate to a different PPU 200. In a second example, the processing context of one or more VMs may be idle for an indeterminate period of time. If all processing contexts in one or more VMs are idle, the VMs can migrate from one PPU 200 to another PPU 200 to improve PPU 200 utilization or reduce power consumption. In a third example, VMs associated with a particular user or user group can be migrated from a geographically distant data center to a closer data center to improve communication latency. More generally, a VM can be migrated to a different PPU 200.

[0250] More generally, a VM can migrate from one PPU 200 to another PPU 200 at any time when the context has been removed from hardware execution via context save. The associated operating system, such as Figure 9 the guest operating system 916, can preempt the context and force a context save at any time, not just when the corresponding VM is idle. The context can be preempted when all work for the corresponding VM has been completed. Additionally, the context can be preempted by forcing the context to stop submitting further work and exhausting the currently ongoing work, even if the VM has other work to perform. In either case, once the ongoing work of the context has been exhausted and the context has been saved, the VM can be migrated from one PPU200 to another PPU 200. During VM migration, the VM may pause execution for a period of about a few milliseconds

[0251] In some embodiments, a VM may be migrated only to a PPU partition 600 in another PPU 200 that has the same configuration as the PPU partition 600 currently executing the VM. For example, a VM may be restricted to migrating only to a PPU partition 600 in another PPU 200 that has the same number of GPCs 242 as the PPU partition 600 currently executing the VM. As shown in FIG. 2802, four VMs are in an idle state. PPU 200(1) executes two VMs 2810A and 2810C. The other two VMs 2810B and 2810D that were previously executed on PPU 200(1) are in an idle state. Similarly, PPU 200(2) executes two VMs 2810E and 2810H. The other two VMs 2810F and 2810G that were previously executed on PPU 200(1) are in an idle state. As a result, each of PPU 200(1) and 200(2) is underutilized. In this case, the currently executing VMs may be migrated to better utilize the available PPU resources. In one example, the VMs executing on PPU 200(1) may consume half of the available hardware resources on PPU 200(1). Similarly, the VMs executing on PPU 200(2) may consume half of the available hardware resources on PPU 200(2). As a result, each of PPU 200(1) and PPU 200(2) will operate at approximately 50% capacity. If all the VMs executing on PPU 200(2) are migrated to PPU 200(1), then PPU 200(1) will operate at approximately 100% capacity. PPU 200(2) will operate at 0% capacity. As a result, to reduce power consumption, the power voltage to PPU 200(2) may be reduced. As shown in FIG. 2804, VMs 2810E and 2810H have been migrated from PPU 200(2) to PPU 200(1). As a result, PPU200(1) executes four VMs 2810A, 2810E, 2810C, and 2810H. Thus, PPU 200(1) is more fully utilized. After the VM migration, PPU 200(2) no longer executes any VMs. As a result, PPU 200(2) may be powered off to reduce power consumption. If additional VMs subsequently start executing, then PPU 200(2) may be powered on to execute the additional VMs.

[0252] Figure 29 is a set of timelines 2900 according to various embodiments, which shows fine-grained VM migrations associated with Figure 2 the PPU 200. The functions of the set of timelines 2900 are respectively related to Figure 25 and 26 the timelines 2500 and 2600 of Figure 27is substantially the same as the timeline 2700, except as further described below. As shown, the set of timelines 2900 includes, but is not limited to, four PPU partition timelines 2902(0), 2902(1), 2902(2), and 2902(3). In some embodiments, the PPU partition timelines 2902(0), 2902(1), 2902(2), and 2902(3) may correspond to Figure 6 the four PPU partitions 600. Each PPU partition 600 includes an SMC engine 700. Thus, each PPU partition 600 can execute one VM at a time. As shown by the PPU partition timelines 2902(0), 2902(1), 2902(2), and 2902(3), each PPU partition 600 time-slices among five VMs, which are referred to as VM A through VM F. As a result, each of the five VMs migrates among the four PPU partitions 600.

[0253] During [[ID= the period shown, VM A executes processing context 2910(A0) on the first PPU partition, as shown by the PPU partition timeline 2902(3). Then, VM A migrates to the second PPU partition and executes processing context 2910(A1) on the PPU partition timeline 2902(2). Subsequently, VM A migrates to the third and fourth PPU partitions in turn and executes processing contexts 2910(A2) and 2910(A3) on the PPU partition timelines 2902(1) and 2902(0), respectively. Then, VM A migrates back to the first PPU partition and executes processing context 2910(A4) on the PPU partition timeline 2902(3). Finally, VM A migrates to the second PPU partition again and executes processing context 2910(A5) on the PPU partition timeline 2902(2).

[0254] In a similar manner, VM B, which executes processing contexts 2910(B0) through 2910(B4), executes on the first PPU partition 600 and then migrates among the other three PPU partitions, as shown by the PPU partition timelines 2902(0), 2902(1), 2902(2), and 2902(3). The remaining three VMs similarly migrate among the four PPU partitions, where VM C executes processing contexts 2910(C0) through 2910(C5), VM D executes processing contexts 2910(D0) through 2910(D5), and VM E executes processing contexts 2910(E0) through 2910(E5).

[0255] In this manner, five VMs are migrated across four PPU partitions 600, with each VM accessing substantially the same amount of PPU resources. Overall, each of the five VMs is able to execute approximately 80% of the time, with four PPU partitions 600 divided by five VMs equaling 4 / 5, or 80%. Consequently, fine-grained VM migration provides load balancing across a group of VMs regardless of the number of VMs relative to the number of PPU partitions 600.

[0256] ​ Various embodiments are shown for ​ Flowchart of the method steps for time slicing a VM in the PPU 200. ​ The method steps are described with respect to a system, but one of ordinary skill in the art will understand that any system configured to perform the method steps, in any order, is within the scope of the present disclosure.

[0257] As shown, method 3000 begins at step 3002, where PPU 200 determines that at least one VM is about to switch from a first set of one or more processing contexts to a second set of one or more processing contexts. A VM may switch one or more processing contexts for any technically feasible reason, including but not limited to, the VM completing execution of all tasks, the VM entering an idle state, the VM having executed for a maximum allotted time, or the VM having generated an error.

[0258] At step 3004, the PPU 200 determines whether the PPU 200 needs to perform an intra-PPU partition change to accommodate the new one or more processing contexts. An intra-PPU partition change occurs when the PPU partition 600 maintains the same number of PPU slices 610 during a context switch but changes the number of active SMC engines 700 within the PPU partition 600. If the PPU 200 does not need to perform an intra-PPU partition change to accommodate the new one or more processing contexts, the method 3000 proceeds to step 3008. However, if the PPU 200 needs to perform an intra-PPU partition change to accommodate the new one or more processing contexts, the method 3000 proceeds to step 3006, where the PPU 200 reconfigures the PPU partition 600 to maintain the same number of PPU slices 610 while changing the number of active SMC engines 700.

[0259] In step 3008, the PPU 200 determines whether the PPU 200 needs to perform an inter-PPU partition change to accommodate the new one or more processing contexts. An inter-PPU partition change occurs when the PPU partition 600 changes the number of PPU slices 610 by merging or splitting one or more PPU partitions 600 during a context switch. As a result, the PPU 200 also changes the number of PPU partitions 600 within 700. Depending on the new processing context, the PPU 200 may or may not change the number of active SMC engines 700 in each PPU partition 600. If the PPU 200 does not need to perform an inter-PPU partition change to accommodate the new one or more processing contexts, the method 3000 proceeds to step 3012. However, if the PPU 200 needs to perform an inter-PPU partition change to accommodate the new one or more processing contexts, the method 3000 proceeds to step 3010, where the PPU 200 reconfigures the PPU partition 600 to change the number of PPU slices 610 included in the PPU partition 600. To change the number of PPU slices 610 in the PPU partition 600, the PPU 200 merges two or more PPU partitions 600 into a single PPU partition 600. Additionally or alternatively, the PPU 200 can split the PPU partition 600 into two or more PPU partitions 600.

[0260] In step 3012, the PPU 200 determines whether all VMs executed on the given PPU 200 are idle. If one or more VMs on the given PPU 200 are not idle (active), the method proceeds to step 3018. On the other hand, if all VMs executed on the given PPU 200 are idle, the method 3000 proceeds to step 3014, where the PPU 200 determines whether one or more other PPU 200s have resources available for executing the idle VMs. If the resources are not available on one or more other PPU 200s, the method proceeds to step 3018. On the other hand, if the resources are available on one or more other PPU 200s, the method proceeds to step 3016, where the PPU 200 migrates the idle VMs to one or more other PPU 200s.

[0261] In step 3018, after performing the intra-PPU partition change, the inter-PPU partition change, and / or the VM migration, the PPU 200 starts to execute the new processing context. Then, the method 3000 terminates. In various embodiments, the PPU 200 determines the need to perform an intra-PPU change independently of determining the need to perform an inter-PPU change. Similarly, in various embodiments, the PPU 200 determines the need to perform an inter-PPU change independently of determining the need to perform an intra-PPU change.

[0262] Privileged Register Address Mapping

[0263] As further described herein ​ The PRI hub 512 and the internal PRI bus (not shown) enable any unit in the CPU 110 and / or PPU 200 to read and write to privileged registers, also known as "PRI registers": which are distributed throughout the PPU 200. Thus, the PRI hub 212 is configured to map PRI bus addresses between a common address space that covers all PRI bus registers and an address space defined separately for each system pipeline 230. When communicating via a PCIe link (typically from the CPU), the PRI registers are accessed through the PCIe address space, which is referred to herein as the "Base Address Register 0" space, or more simply, the "BAR0" address space, which is typically used for devices connected to the PCIe bus. Typically, since a large number of devices must all fit within the BAR0 address space, the addressable memory range of the BAR0 address space of the PPU 200 is limited to 16 megabytes (MB). The 16MB address range is sufficient to access the privileged registers of a single SMC engine 700. However, to support multiple SMC engines 700, the address range may exceed 16MB. Thus, the PRI hub 512 provides two addressing modes to support execution with multiple SMC engines 300. The first addressing mode, referred to herein as the "traditional mode", is applicable to operations involving a single SMC engine 700. The second addressing mode, referred to herein as the "SMC Engine Addressing Mode", is applicable to operations involving multiple SMC engines 700. The addressing modes are now described.

[0264] ​ is a memory map according to various embodiments, which shows how the BAR0 address space 3110 maps to ​The privileged register address space 3120 within the PPU 200. The BAR0 address space 3110 includes, but is not limited to, a first address space 3112, a graphics register (GFX REG) address space 3114, and a second address space 3116. The privileged register address space 3120 includes, but is not limited to, a first address space 3122, a legacy graphics register address space 3124, a second address space 3126, and an SMC graphics register address space 3128(0)-3128(7). The BAR0 address space 3110 supports two addressing modes, a legacy addressing mode and an SMC addressing mode. Generally, the goals of the two modes are: (1) the legacy mode, in which the entire PPU 200 is treated as one engine with a set of PRI registers; (2) the SMC mode, in which each PPU partition 600 is addressed as if each PPU partition 600 is a separate engine and as if each PPU partition 600 itself is the entire PPU 200. When processing the entire PPU in the legacy mode and when processing only on one PPU partition 600, the SMC mode allows the driver software 122 to be the same. That is, the driver can be written once and used in both the legacy mode and the SMC mode scenarios..

[0265] In the legacy addressing mode, the PPU 200 executes tasks as a single cluster of hardware resources rather than as separate PPU partitions 600 with separate SMC engines 700. In the legacy mode, a memory read or write to a memory address is directed to the first address space 3112 or the second address space 3116 of the BAR0 address space 3110, accessing the corresponding memory address in the first address space 3122 or the second address space 3126 of the privileged register address space 3120, respectively. Similarly, a memory read or write to a memory address is directed to the graphics register address space 3114 of the BAR0 address space 3110, accessing the corresponding memory address within the legacy graphics register address space 3124 of the privileged register address space 3120. The legacy graphics register address space 3124 includes the address ranges of various components within the PPU 200, including, but not limited to, the compute FE 540, the graphics FE 542, the SKED 550, the CWD 560, and the PDA / PDB 562. In addition, the legacy graphics register address space 3124 includes the address range for each GPC 242. The GPC 242 can be individually addressed via a dedicated address range within the legacy graphics register address space 3124, minus any GPC 242 that has been removed due to floor sweeping. Additionally or alternatively, the legacy graphics register address space 3124 includes an address range for broadcasting data to all GPC 242 simultaneously. These GPC broadcast address spaces can be useful when all GPC 242 are configured identically.

[0266] In the SMC addressing mode, the PPU 200 executes tasks as a separate PPU partition 600 with a separate SMC engine 700. As in the traditional mode, a memory read or write to a memory address is directed to either the first address space 3112 or the second address space 3116 of the BAR0 address space 3110, accessing the corresponding memory address in either the first address space 3122 or the second address space 3126 of the privileged register address space 3120, respectively. In the SMC mode, the SMC graphics register address spaces 3128(0)-3128(7) are provided to access the respective components in each of the SMC engines 700(0)-700(7). These components include, but are not limited to, the compute FEs 540(0)-540(7), the graphics FEs 542(0)-542(7), the SKEDs 550(0)-550(7), the CWDs 560(0)-560(7), and the PDAs / PDBs 562(0)-562(7). The GPCs 242 of a particular corresponding SMC engine 700 are individually addressable through a dedicated address range within the traditional graphics register address space 3124, minus any GPCs 242 that have been removed due to floor sweeping. Additionally or alternatively, the traditional graphics register address space 3124 includes an address range for simultaneously broadcasting data to all of the GPCs 242 of a particular corresponding SMC engine 700. The BAR0 address space 3110 provides two mechanisms for accessing the SMC graphics register address spaces 3128(0)-3128(7).

[0267] In the first mechanism, the graphics register address space 3114 of the BAR0 address space 3110 is mapped to one of the SMC graphics register address spaces 3128(0)-3128(7) in the privileged register address space 3120. A specific address within the BAR0 address space 3110 accesses the SMC window register. The SMC window register includes two fields. These two fields include an SMC enable field and an SMC index field. The SMC enable field contains a binary logical value that is FALSE or TRUE. If the SMC enable field is FALSE, the BAR0 address space 3110 accesses the privileged register address space 3120 in a traditional addressing mode, as described herein. If the SMC enable field is TRUE, the BAR0 address space 3110 accesses the privileged register address space 3120 in an SMC addressing mode based on the value of the SMC index field. The value of the SMC index field specifies which SMC engine 700 is currently mapped to the BAR0 address space 3110. For example, if the value of the SMC index field is 0, the SMC graphics register address space 3128(0) of the privileged register address space 3120 will be mapped to the graphics register address space 3114 of the BAR0 address space 3110. Similarly, if the value of the SMC index field is 1, the SMC graphics register address space 3128(1) of the privileged register address space 3120 will be mapped to the register address space 3114 of the graphics BAR0 address space 3110, and so on. A memory read or write to a memory address directed to the graphics register address space 3114 of the BAR0 address space 3110 accesses the corresponding memory address within the SMC graphics register address space 3128 specified by the SMC index field. Through this first mechanism, the SMC graphics register address space 3128 of the SMC engine 700 specified by the SMC index field can be accessed, while access to the remaining SMC graphics register address spaces 3128 is prohibited.

[0268] In a second mechanism, certain privileged components (e.g., the hypervisor 124) can access a separate address of any one of the SMC graphics register address spaces 3128(0) - 3128(7). This second mechanism accesses the SMC graphics register address spaces 3128(0) - 3128(7) through two specific addresses in the BAR0 address space 3110. One of these two addresses accesses the SMC address register. The other of these two addresses accesses the SMC data register. A specific memory address at any location in the SMC graphics register address spaces 3128(0) - 3128(7) can be accessed in two steps. In the first step, an address corresponding to an address in the SMC graphics register address spaces 3128(0) - 3128(7) is written to the SMC address register. In the second step, the SMC data register is read or written with a data value. Reading or writing the SMC data register in the BAR0 address space 3110 causes a corresponding read or write to the privileged register address space 3120 at the memory address specified in the SMC address register. The SMC address register is then dereferenced, enabling the SMC address register and the SMC data register for subsequent transactions. Reading or writing the SMC address register does not cause a read or write to the privileged register address space 3120.

[0269] ​ is a flowchart of method steps for addressing a privileged register address space in a ​ PPU 200 according to various embodiments. Although the method steps are described in connection with a ​ system, those of ordinary skill in the art will understand that any system configured to execute the method steps in any order is within the scope of the present disclosure.

[0270] As shown, method 3200 begins at step 3202, where the PPU 200 detects a memory access to the privileged register address space 3120. More specifically, the PPU 200 detects a memory access to the BAR0 address space 3110. At step 3204, the PPU 200 determines whether the memory access is to the graphics register address space 3114. If the memory access is not to the graphics register address space 3114, method 3200 proceeds to step 3214, where the PPU 200 generates a memory transaction to the address specified by the memory access. Then, method 3200 terminates.

[0271] Return to step 3204. If the memory access is to the graphics register address space 3114, method 3200 proceeds to step 3206, where PPU 200 determines whether the memory access is in the legacy mode. If the memory access is in the legacy mode, method 3200 advances to step 3214, where PPU 200 generates a memory transaction to the address specified by the memory access. Then, method 3200 terminates. On the other hand, if the memory access is not in the legacy mode, method 3200 proceeds to step 3208, where PPU 200 determines whether the memory access is in the window mode.

[0272] If the memory access is in the window mode, the method proceeds to step 3212, where PPU 200 generates a memory transaction based on the value of the SMC index field specified in the SMC window register. The value of the SMC index field specifies which SMC engine 700 is currently mapped to the BAR0 address space 3110. For example, if the value of the SMC index field is 0, the SMC graphics register address space 3128(0) of the privileged register address space 3120 will be mapped to the graphics register address space 3114 of the BAR0 address space 3110. Similarly, if the value of the SMC index field is 1, the SMC graphics register address space 3128(1) of the privileged register address space 3120 will be mapped to the graphics register address space 3114 of the BAR0 address space 3110, and so on. Direct a memory read or write to the memory address to the graphics register address space 3114 of the BAR0 address space 3110, accessing the corresponding memory address within the SMC graphics register address space 3128 specified by the SMC index field. Through this first mechanism, the SMC graphics register address space 3128 of the SMC engine 700 specified by the SMC index field can be accessed, while access to the remaining SMC graphics register address space 3128 is prohibited. Then, method 3200 terminates.

[0273] Return to step 3208. If the memory access is not in window mode, the method proceeds to step 3210, where the PPU 200 generates a memory transaction based on the values of the SMC address register and the SMC data register. More specifically, the PPU 200 accesses the SMC graphics register address spaces 3128(0)-3128(7) through two specific addresses within the BAR0 address space 3110. One of these two addresses accesses the SMC address register. The other of these two addresses accesses the SMC data register. A specific memory address at any location within the SMC graphics register address spaces 3128(0)-3128(7) can be accessed in two steps. In the first step, an address corresponding to the address within the SMC graphics register address spaces 3128(0)-3128(7) is written to the SMC address register. In the second step, the SMC data register is read or written with a data value. Reading or writing the SMC data register within the BAR0 address space 3110 causes a corresponding read or write to the privileged register address space 3120 at the memory address specified in the SMC address register. The SMC address register is then dereferenced, enabling the SMC address register and the SMC data register for subsequent transactions. Reading or writing the SMC address register does not cause a read or write to the privileged register address space 3120. Then, method 3200 terminates.

[0274] Performance monitoring using multiple SMC engines

[0275] As further discussed herein, a performance monitor (PM), such as ​ PM 236 of ​ PM 360 of ​ PM 430 of, monitors the overall performance and / or resource consumption of the corresponding components included in the PPU 200. The performance monitor (PM) is included in a performance monitoring system that provides performance monitoring and performance analysis across multiple SMC engines 700. The performance monitoring system profiles multiple VMs and the processing contexts executing within the VMs simultaneously or substantially simultaneously. The performance monitoring system isolates multiple virtual machines and the multiple processing contexts executing within the VMs from each other with respect to how performance data is generated and captured to prevent leakage of performance data between VMs. The PM 232 and associated counters in the performance data monitoring system track the performance data for the attributes of a specific SMC engine 700. In the case of shared resources and units where the attributes are not traceable to a specific SMC engine 700, a device with a higher privileged entity (such as the hypervisor 124) collects the performance data for the shared resources and units. When a VM migrates to other partition PPUs 600 and / or other PPUs 200, the performance monitoring system profiles the compute engine and the graphics engine, as well as profiles the VM. The performance monitoring system is now described.

[0276] ​ According to various embodiments, ​ 200 , a block diagram of a performance monitoring system 3300 for a PPU 200. As shown, the performance monitoring system 3300 includes, but is not limited to, a performance monitor 3310, a selection multiplexer 3320, a monitoring bus 3330, and a performance multiplexer unit 3340. The performance monitor 3310 and the selection multiplexer 3320 together constitute a performance monitor module (PMM). Each GPC 242, each partition unit 262, and each system pipeline 230 includes at least one PMM. The performance multiplexer unit 3340 is included in each unit being monitored. The logic within the performance multiplexer unit 3340 is included in one or more of the FE 540, SKED 550, CWD 560, and / or other suitable functional units. As further described herein, all components of the various performance monitoring systems 3300 are integrated with the performance monitor aggregator ( ​ Except as further described below, the functionality of the performance monitoring system 3300 is similar to that of the ​ PM 236, ​ PM 360 and PM 430 are basically the same.

[0277] In operation, performance multiplexer units 3340(0)-3340(P) enable programmable selection of groups of signals within PPU 200 that can be monitored by corresponding performance monitors 3340. Each performance multiplexer 3340 can select a group of signals to be sent to monitor bus 3330. A subset of signals from monitor bus 3330 is selected for monitoring by selection multiplexer 3320. Signals from monitor bus 3330 that are not selected by selection multiplexer 3320 are not monitored. Signals from within PPU 200 are connected to performance multiplexer units 3340(0)-3340(P) in groups so that one group is selected at a time for monitoring. Performance multiplexer units 3340 multiplex signals so that signals in a particular signal group are selected as a group. The select inputs of the multiplexers included in performance multiplexer units 3340 are programmed by one or more registers included in privileged register address space 3120. As a result, the specific signals transmitted by performance multiplexer unit 3340 to monitoring bus 3330 are programmable.

[0278] The monitor bus 3330 receives a set of signals from the performance multiplexer units 3340(0)-3340(P). Each signal sent to the monitor bus 3330 is connected as an input to each selection multiplexer 3320.

[0279] The select multiplexer 3320 includes a set of individual multiplexers 3322(0)-3322(M) and 3324(0)-3324(N). The input sides of each of the multiplexers 3322(0)-3322(M) and 3324(0)-3324(N) receive all signals from the monitoring bus 3330 and select one signal for transmission. The select inputs of the multiplexers 3322(0)-3322(M) and 3324(0)-3324(N) are programmed by one or more registers included in the privileged register address space 3120. As a result, a specific set of signals transmitted by the select multiplexer 3320 is programmable. The select multiplexer 3320 transmits the selected signals to the performance monitor 3310.

[0280] Due to the composition of the programmable performance multiplexer unit 3340 and the programmable select multiplexer 3320, the specific PPU signals monitored by the performance monitor 3310 are programmable.

[0281] The performance monitor 3310 includes a performance counter array 3312, a shadow counter array 3314, and a trigger function table 3316. The performance monitor 3310 receives the signals transmitted by the select multiplexer 3320. More specifically, the shadow counter array 3314 receives the signals transmitted by the multiplexers 3322(0)-3322(M). Similarly, the trigger function table 3316 receives the signals transmitted by the multiplexers 3324(0)-3324(N). As further described, the counters within the shadow counter array 3314 are updated based on the signals received from the multiplexers 3322(0)-3322(M) and various trigger conditions. Generally, the shadow counter array 3314 includes a set of one or more signal counters, where each counter increments whenever the signal received from the corresponding multiplexer 3322 is in a specific logic state. Based on certain signals in the form of trigger conditions, the values in the shadow counter array 3314 are transferred to the performance counter array 3312. The performance counter array 3312 includes a set of one or more signal counters corresponding to the signal counters included in the shadow counter array.

[0282] In one operating mode, after being transferred to the performance counter array 3312, the counters in the shadow counter array 3314 are reset to zero so that the shadow counter values stored in the shadow counter array 3314 always correspond to the activity since the previous trigger.

[0283] The performance monitor 3310 can be configured according to various counting modes that define the number of counters included in the performance counter array 3312 and the shadow counter array 3314. The counting modes also define how and when to trigger the performance counter array 3312 and the shadow counter array 3314, and how to transfer data from the performance counter array 3312 to other devices within the PPU 200. These various counting modes can be divided into two main performance monitoring modes - non-stream performance monitoring and stream performance monitoring.

[0284] In the non-stream performance monitoring mode, the trigger function table 3316 is programmed to receive signals from the multiplexers 3324(0)-3324(N) according to certain specified logical signal expression combinations. When the conditions of one or more of these logical signal expressions are satisfied, the trigger function table 3316 sends a signal in the form of a logic trigger 3350 to the performance counter array 3312. In response to receiving the logic trigger 3350, the performance counter array 3312 samples and stores the current value in the shadow counter array 3314. Then, the value in the performance counter array 3312 is read through the privileged register address space 3120.

[0285] In the stream performance monitoring mode, the performance monitor aggregator (PMA) sends a signal in the form of a PMA trigger 3352 to the performance counter array 3312. In response to receiving the PMA trigger 3352, the performance counter array 3312 samples and stores the current value in the shadow counter array 3314. The performance monitor 3310 generates a performance monitor (PMM) record, which can include but is not limited to the value in the performance counter array 3312 when the PMA trigger 3352 is received from the PMA, the count of the total number of PMA triggers to which the performance monitor 3310 responds, the SMC engine ID, and the PMMID that uniquely identifies the PMM that generates the record in the system. Then, these PMM records are sent to the PMM router associated with one or more performance monitors 3310. Then, the PMM router sends the PMM records to the PMA. In some embodiments, the PMM ID for each performance monitor 3310 can be programmed via one or more registers included in the privileged register address space 3120.

[0286] Generally, a specific performance monitor 3310 in a specific performance monitoring system 3300 in the PPU 200 is in the same clock frequency domain as the signal monitored by that specific performance monitor 3310. However, a specific performance monitor 3310 can be in the same clock frequency domain or a different clock frequency domain relative to another performance monitor in the PPU 200.

[0287] Now, the various configurations of the performance multiplexer unit 3340 are described.

[0288] ​ illustrates various configurations of a performance multiplexer unit 3340 in accordance with various embodiments ​ of the performance multiplexer unit 3340.

[0289] As ​ shown, a first configuration of the performance multiplexer unit 3340(0) includes, but is not limited to, signal groups A 3420(0)-3420(P), signal groups B 3430(0)-3430(Q), and multiplexers 3412(0) and 3412(1). In operation, multiplexer 3412(0) selects one of the signal groups A 3420(0)-3420(P), where each of the signal groups A 3420(0)-3420(P) is a subgroup within a larger signal group C. Multiplexer 3412(0) selects one of the signal groups A 3420(0)-3420(P) and transmits the selected signal group to the monitoring bus 3330. Similarly, multiplexer 3412(1) selects one of the signal groups B 3430(0)-3430(Q), where each of the signal groups B 3430(0)-3430(Q) is a subgroup within a larger signal group D. Multiplexer 3412(1) selects one of the signal groups B 3430(0)-3430(Q) and transmits the selected subgroup to the monitoring bus 3330. The selection inputs of the multiplexers 3412 included in the performance multiplexer unit 3340(0) are programmed via one or more registers included in the privileged register address space 3120. As a result, the set of signals transmitted by the performance multiplexer unit 3340(0) is programmable.

[0290] As ​As shown, the second configuration of the performance multiplexer unit 3340(1) includes, but is not limited to, signal groups C3440(0)-3440(R) and multiplexer 3412(2). In operation, multiplexer 3412(2) selects one of signal groups C 3440(0)-3440(R), where each of signal groups C 3440(0)-3440(R) is a subgroup within a larger signal group E. Multiplexer 3412(2) selects one of signal groups C3440(0)-3440(R) and transmits the selected signal group to the monitoring bus 3330. In the configuration of performance multiplexer unit 3340(1), several signals are transmitted to multiple signal groups. In particular, signal C1 3450 is sent to signal group C 3440(0) and signal group C 3440(1). Similarly, signal C2 3452 is sent to signal group C 3440(1) and signal group C 3440(2). The selection inputs of multiplexer 3412(2) included in performance multiplexer unit 3340(1) are programmed via one or more registers included in the privileged register address space 3120. As a result, the set of signals transmitted by performance multiplexer unit 3340(1) is programmable. The configuration of performance multiplexer unit 3340(1) may be useful in facilitating the visibility of certain signal groups in a single pass of the performance monitoring system 3300 by making signals available in multiple signal groups 3440.

[0291] ​ is a block diagram of a performance monitor aggregation system 3500 for a PPU 200 according to various embodiments. As shown, performance monitor aggregation system 3500 includes, but is not limited to, GPCs 242(0)-242(M), partition units 262(0)-262(N), crossbar unit 250, control crossbar and SMC arbiter 510, PM management system 3530, and performance analysis system 3540. ​

[0292] ​ In operation, GPCs 242(0)-242(M) perform various processing tasks for one or more system pipelines 230. Each GPC 242 includes multiple parallel processing cores capable of executing a large number of threads simultaneously and has an arbitrary degree of independence and / or isolation from other GPCs 242. Each GPC 242(0)-242(M) includes one or more PMs 360(0)-360(M) and GPC PMM routers 3514(0)-3514(M). The functions of PMs 360(0)-360(M) are substantially similar to ​Performance Monitor 3310. PM 360(0)-360(M) generates PMM records that include performance data of the corresponding GPC 242(0)-242(M). PM 360(0)-360(M) sends these PMM records to the corresponding GPC PMM routers 3514(0)-3514(M) and receives data from them. GPC PMM routers 3514(0)-3514(M) transmit the PMM records to the PM management system 3530 through the crossbar unit 250.

[0293] Partition units 262(0)-262(N) provide high-bandwidth memory access to the DRAM within the PPU memory ( ​ not shown). Each partition unit 262 performs memory access operations using different DRAMs in parallel with each other, thus effectively utilizing the available memory bandwidth of the PPU memory. Each of the partition units 262(0)-262(N) includes one or more PM430(0)-430(N) and partition unit (PU) PMM routers 3524(0)-3524(N). The functions of PM 430(0)-430(N) are substantially similar to ​ those of the performance monitor 3310. PM 430(0)-430(N) generates PMM records that include performance data of the corresponding partition units 262(0)-262(N). PM 430(0)-430(N) sends these PMM records to the corresponding PU PMM routers 3524(0)-3524(N) and receives data from them. PU PMM routers 3524(0)-3524(N) in turn transmit the PMM records to the PM management system 3530 through the control crossbar and SMC arbiter 510.

[0294] The PM management system 3530 controls the collection of PMM records and stores the PMM records for reporting purposes. The PM management system 3530 includes, but is not limited to, a system performance monitor 3532, a system (SYS) PMM router 3534, a performance monitor aggregator (PMA) 3536, a high-speed hub (HSHUB), and transmission logic 3539.

[0295] The function of the system PM 3532 is substantially similar to ​ that of the performance monitor 3310. The system PM 3532 generates PMM records that contain performance data for system-wide components not included in a specific GPC 242 or partition unit 262. The system PM 3532 sends these PMM records to the system PMM router 3534 and receives data from it. The system PMM router 3534 in turn sends the PMM records to the PMA 3536.

[0296] PMA 3536 generates triggers for various performance monitors including PM 360(0)-360(M), PM 430(0)-430(N), and system PM 3532. PMA 3536 generates these triggers through two techniques. In the first technique, when the system pipeline 230 receives a command from the host interface 220, PMA 3536 generates a trigger in response to a signal sent by each system pipeline 230. In the second technique, PMA 3536 generates a trigger by periodically sending a programmatically controlled trigger pulse to the PM. Generally, the performance monitor aggregation system 3500 includes at least one programmable trigger pulse generator corresponding to each system pipeline 230 in addition to another trigger pulse generator independent of any system pipeline 230. PMA 3536 sends the trigger signal to the GPC PMM routers 3514(0)-3514(M) and the PU PMs 3524(0)-3524(N) by controlling the crossbar switch and the SMC arbiter 510. PMA 3536 sends the trigger directly to the system PMM router 3534 through a communication link inside the PM management system 3530. The PMM router sends the PMA trigger to the corresponding PM. In some embodiments, these triggers take the form of trigger packets combined with ​ described below.

[0297] ​ FIG. shows the format of trigger packets associated with the ​ performance monitor aggregation system 3500 according to various embodiments. The purpose of the trigger packet is to convey information about the trigger source to the performance monitor 3310. Each of the performance monitors 3310 uses this information to determine whether to respond to a specific trigger. In this regard, the trigger packet contains information that can be used by each performance monitor 3310 to determine whether to respond to a specific trigger packet. Each performance monitor 3310 associated with a specific SMC engine 700 is programmed with the SMC engine ID corresponding to that SMC engine 700. Such a performance monitor 3310 responds to each SMC trigger packet including the same SMC engine ID. Each performance monitor 3310 not associated with a specific SMC engine 700 or programmed with an invalid SMC engine ID does not respond to each SMC trigger packet. Instead, such a performance monitor 3310 responds to a shared trigger packet.

[0298] Figure 3600 shows the general format of a trigger packet. As shown, Figure 3600 includes a packet type 3602 indicating that the packet is a PM trigger, a trigger type 3604, and a trigger payload 3606. The PM trigger type 3604 is an enumerated value that identifies the type of trigger format. For example, to identify three different trigger packet types, the PM trigger type 3604 can be a 2-bit value. The trigger payload 3606 includes data that varies based on the PM trigger type 3604. Now, three different types of trigger packets are described, where the three types of trigger packets correspond to three categories of performance monitoring data (legacy data, per-SMC data, and shared data).

[0299] Figure 3610 shows the format of a legacy trigger packet. The legacy trigger packet includes a packet type 3602 indicating that the packet is a PM trigger and a trigger type 3614 indicating that the trigger packet is a legacy trigger packet. The trigger payload 3606 of the legacy trigger packet includes an unused field 3616.

[0300] Figure 3620 illustrates the format of a per-SMC trigger packet. The per-SMC trigger packet includes a packet type 3602 indicating that the packet is a PM trigger and a trigger type 3624 indicating that the trigger packet is a per-SMC trigger packet. The trigger payload 3606 of the per-SMC trigger packet includes an SMC engine ID field. The SMC engine ID field 3626 identifies the specific SMC engine 700 to which the trigger applies.

[0301] Figure 3630 shows the format of a shared trigger packet. The shared trigger packet includes a packet type 3602 indicating that the packet is a PM trigger and a trigger type 3634 indicating that the trigger packet is a shared trigger packet. The trigger payload 3606 of the shared trigger packet includes an unused field.

[0302] The type of trigger packet sent by the PMA 3536 is determined by the trigger source and one or more registers included in the privilege register address space 3120 corresponding to each trigger source, which are programmed to indicate the type of trigger packet that the PMA should send for that source. In one mode of operation, the PMA is programmed such that trigger packets generated in response to a source associated with an SMC engine are per-SMC trigger packets with the SMC engine ID set to the corresponding SMC engine, and trigger packets generated in response to a source not associated with an SMC engine are shared trigger packets. Trigger packets can be generated at any technically feasible rate, with at most one trigger packet per computational cycle.

[0303] In response to receiving a trigger packet from PMA 3536, each PM checks the trigger type and trigger payload to determine whether the PM should respond to the trigger. Each PM unconditionally responds to a legacy trigger packet. In the case of each SMC trigger packet, the PM responds only if the SMC engine ID included in the trigger payload matches the SMC engine ID programmed for the PM via the register privilege register address space 3120. An invalid SMC engine ID programmed in this register ensures that the PM does not respond to each SMC trigger packet. In the case of a shared trigger packet, the PM responds only if the PM has been programmed via a register in the privilege register address space 3120 to respond. In one mode of operation, all PMs that are uniquely assigned to an SMC engine as monitoring units are programmed to respond to each SMC trigger packet with the corresponding SMC engine ID payload, while all other PMs are programmed to respond to shared trigger packets but not to each SMC trigger packet.

[0304] In the case where the PM determines to ensure a response to the trigger, the PM samples a counter included in the corresponding PM. Then, the responding PM transmits a PMM record that includes the sampled counter value, the total number of responding triggers, the SMC engine ID assigned to the PM, and the PMM ID that uniquely identifies the PM in the system. The PMM router, PMA 3536, and / or the performance analysis system 3540 use the PMM ID to identify which PM sent the corresponding PMM record. Then, the PMM router sends the record to PMA 3536. More specifically, the GPC PMM routers 3514(0)-3514(M) send the PMM record to PMA 3536 via the crossbar unit 250 and the high-speed hub 3538. The PU PMM routers 3524(0)-3524(N) send the PMM record to PMA 3536 via the control crossbar and the SMC arbiter 510. The system PMM router 3534 directly transmits the PMM record to the PM management system 3530 via a communication link. In this way, PMA 3536 receives PMM records from all relevant PMs in the PPU 200.

[0305] In some embodiments, when a transmission trigger occurs, PMA 3536 additionally generates a PMA record. Generally, the purpose of a PMA record is to record the timestamp at which a particular PMA trigger occurred and to associate the corresponding PMM record from the performance monitor 3310 in response to that PMA trigger at that timestamp. The PMA record includes, but is not limited to, the timestamp, the SMC engine ID associated with the source of the PMA trigger, the total number of triggers generated by the source with the same SMC engine ID, and associated metadata. When the performance monitor 3310 receives a trigger, the performance monitor 3310 also generates a PMM record with a trigger count. Subsequently, when the PMA record and the PMM record are parsed, the PMM record with a particular trigger count can be associated with the PMA record with the same trigger count. In this way, the timestamp corresponding to the PMM record is established based on the timestamp of the associated PMA record. As a result, the behavior of the PPU 200 reflected by the PMM record is precisely associated with the time range defined by two adjacent PMA triggers from the same source.

[0306] When receiving the PMM record and generating the PMA record, PMA 3536 stores the PMM record and the PMA record in the data store in the PPU memory in the form of a record buffer via the high-speed hub 3538. The high-speed hub 3538 sends the PMM record and the PMA record to the partitioning units 262(0)-262(N). Then, the partitioning units store the PMM record and the PMA record in the record buffer in the PPU memory. Additionally or alternatively, the high-speed hub 3538 transmits the PMM record and the PMA record to the performance analysis system 3540 via the transmission logic 3539. In some embodiments, the high-speed hub 3538, the transmission logic 3539, and the performance analysis system 3540 may communicate with each other via a PCIe link. The user can view the PMM record and the PMA record on the performance analysis system 3540 to characterize the behavior of the PPU 200 as reflected in the PMM record. The performance analysis system 3540 collects the PMM record and the PMA record with the same trigger count. Then, the performance analysis system 3540 uses the timestamp from the PMA record and the performance data from the PMM record with the same trigger count to determine the timestamp associated with the performance data. In some embodiments, the performance analysis system 3540 may access the performance record buffer as virtual memory. As a result of placing the performance record buffer in different virtual address spaces, the performance monitoring data of different SMC engines 700 can be isolated from each other, as will now be described.

[0307] PMA 3536 provides isolation of performance monitoring data among several SMC engines 700. In particular, PMA 3536 classifies PMM records and PMA records into categories based on the operating mode. When the PPU 200 operates in the legacy mode, the PPU 200 executes tasks as a single cluster of hardware resources rather than as separate PPU partitions 600 with separate SMC engines 700. In the legacy mode, PMA 3536 stores PMM records and PMA records in a single category as a single set of performance monitoring data. When the PPU 200 operates in the SMC mode, the PPU 200 executes tasks as separate PPU partitions 600 with separate SMC engines 700. In the SMC mode, PMA 3536 classifies and stores PMM records and PMA records in different categories using the SMC engine ID fields of the PMM records and PMA records. Records with each SMC engine ID are stored in a separate data store in the form of a record buffer, which can be accessed from a different virtual address space that matches the corresponding SMC engine 700. As described above, PMM records and PMA records that cannot be traced to a specific SMC engine 700 contain an invalid SMC engine ID. PMA stores these records in a separate data store in the form of an SMC record buffer in the virtual address space, which can be accessed by any authorized entity that has sufficient privilege to access data for all SMC engines 700. Such authorized entities include, but are not limited to, the hypervisor 124 in a virtual environment and the root user or the operating system kernel in a non-virtual environment. Each SMC engine 700 can access some or all of the performance monitoring data in the non-SMC PMA record buffer by requesting data from an authorized entity.

[0308] In some embodiments, PMA 3536 is configured to generate triggers corresponding to each SMC engine 700 in accordance with context switch events of the same SMC engine 700. In such embodiments, PM is configured such that the counters in the shadow counter array 3314 are reset to zero after each trigger so that, while time slicing is enabled, data transmitted from PM to PMA 3536 for each SMC engine ID can be attributed to a single context or VM.

[0309] ​ is a flowchart of method steps for monitoring ​ the performance of the PPU 200 in accordance with various embodiments. Although the method steps are described in connection with ​ a system, those of ordinary skill in the art will understand that any system configured to execute the method steps in any order is within the scope of the present disclosure.

[0310] As shown in the figure, method 3700 begins at step 3702, where PMA 3536 generates and sends a trigger to PM 3310. In addition, PMA 3536 generates a corresponding PMA record, which includes a timestamp and, optionally, also includes the SMC engine ID corresponding to the source of the trigger. PM 3310 receives the trigger to sample performance data.

[0311] At step 3704, in response to receiving the trigger, PM 3310 determines whether to guarantee a response to the trigger. Each PM 3310 examines the trigger type and trigger payload from the trigger packet to determine whether PM 3310 should respond to the trigger. Each PM3310 unconditionally responds to traditional trigger packets. In the case of each SMC trigger packet, PM 3310 responds only when the SMC engine ID contained in the trigger payload matches the SMC engine ID assigned to PM 3310 by programming in the register in the privileged register address space 3120. Programming an invalid SMC engine ID in this register ensures that PM 3310 does not respond to each SMC trigger packet. In the case of a shared trigger packet, PM 3310 responds only when PM 3310 has been programmed to respond through the register in the privileged register address space 3120. In one operating mode, all PM 3310s that are uniquely assigned to SMC engine 700 as monitoring units are programmed to respond to each SMC trigger packet with the corresponding SMC engine ID payload. All other PM 3310s are programmed to respond to shared trigger packets but not to each SMC trigger packet.

[0312] If a response is required, the performance counter array 3312 samples and stores the current value in the shadow counter array 3314. In non-stream mode, other components can read the values in the performance counter array through the privileged register interface hub 512.

[0313] In streaming mode, the method proceeds to step 3706, where PM 3310 sends the sampled performance data, and PMA 3536 receives the sampled performance data from PM 3310. PM 3310 generates a PMM record, which includes the values in the performance counter array 3312 when the PMA trigger 3352 is received from PMA3536. These PMM records are then sent to the PMM router associated with the specific performance monitor 3310. The PMM router in turn sends the PMM records to PMA 3536.

[0314] In step 3708, PMA 3536 classifies PMM records and PMA records into categories based on the operating mode. When PPU 200 operates in the traditional mode, PMA 3536 classifies the PMM records and PMA records in a single category as a single set of performance monitoring data. When PPU 200 operates in the SMC mode, PMA 3536 uses the SMC engine ID fields of the PMM records and PMA records to classify the PMM records and PMA records into different categories. As described above, PMM records and PMA records that cannot be traced to a specific SMC engine 700 contain invalid SMC engine IDs. PMA 3536 classifies these PMM records and PMA records into a separate category.

[0315] In step 3710, PMA 3536 stores the PMM records and / or PMA records in a PMA record buffer in the PPU memory. When PPU 200 operates in the traditional mode, PMA 3536 stores the PMM records and PMA records in a single category as a single set of performance monitoring data. When PPU 200 operates in the SMC mode, PMA 3536 stores the PMM records and PMA records associated with each SMC engine ID in separate data stores that can be accessed from different virtual address spaces that match the corresponding SMC engines 700. PMA 3536 stores the PMM records and PMA records that cannot be traced to a specific SMC engine 700 in a separate data store in the virtual address space in the form of a non-SMC PMA record buffer, and this virtual address space can only be accessed by any authorized entity that has sufficient privileges to access the data of all SMC engines 700. Such authorized entities include, but are not limited to, the hypervisor 124 in a virtualized environment and the root user or the operating system kernel in a non-virtualized environment. Each SMC engine 700 can access some or all of the performance monitoring data in the non-SMC PMA record buffer by requesting data from an authorized entity.

[0316] More specifically, PMA 3536 streams the PMM records and PMA records to the high-speed hub 3538. The high-speed hub 3538 sends the PMM records and PMA transmissions to the partitioning units 262(0)-262(N). Then, the partitioning units store the PMM records and PMA records in the PMA record buffer in the PPU memory. The PMA record buffer for each record is selected based on the SMC engine ID field of the record so that each record ultimately resides in the PPU memory that can be accessed in the virtual address space that matches the SMC engine corresponding to the PMA record buffer. The PMM records and PMA records that cannot be traced to a specific SMC engine 700 in the separate data store can only be accessed by authorized entities.

[0317] In step 3712, PMA 3536 sends PMA records and / or PMM records to performance analysis system 3540 via high-speed hub 3538 and transmission logic 3539. Additionally or alternatively, performance analysis system 3540 accesses the PMA records and / or PMM records via one or more virtual addresses in the virtual address space. Generally, performance analysis system 3540 includes a software application executed on CPU 110 and / or any other technically feasible processor. Performance analysis system 3540 directly accesses virtual memory to access the PMA records and / or PMM records. The virtual memory may be associated with PPU 200 and / or CPU 110. A user may view the PMA records and / or PMM records on performance analysis system 3540 to characterize the behavior of PPU 200 as reflected in the PMA records and / or PMM records. Then, method 3700 terminates.

[0318] Power and Clock Frequency Management of the SMC Engine

[0319] Complex systems, such as ​ the PPU 200 may consume a large amount of power. More specifically, certain components within PPU 200 may have different power consumption levels from each other at different points in time. In one example, components in one PPU partition 600 may perform computational and / or graphics-intensive tasks, thereby increasing power consumption relative to other PPU partitions 600. In another example, due to leakage currents and related factors, even when in an idle state, PPU partition 600 may consume power. Additionally, the increased power consumption may result in a higher operating temperature, which may in turn lead to a performance degradation. As a result, PPU 200 includes power and clock frequency management that takes into account how power consumption within one PPU partition 600 may negatively impact the performance of other PPU partitions 600.

[0320] ​ is for a ​ block diagram of power and clock frequency management system 3800 of PPU 200 according to various embodiments. Power and clock frequency management system 3800 includes, but is not limited to, circuit sub-parts 3810(0)-3810(N), power gating controller 3820, and clock frequency controller 3830.

[0321] Each of circuit sub-parts 3810(0)-3810(N) includes any set of components included in PPU 200 at any level of granularity. In this regard, each of circuit sub-parts 3810(0)-3810(N) may include, but is not limited to, system pipeline 230, PPU partition 600, PPU slice 610, SMC engine 700, or any technically feasible subset thereof.

[0322] In operation, the power gating controller 3820 monitors the active state of the circuit sub - parts 3810(0) - 3810(N). If the power gating controller 3820 determines that a particular circuit sub - part (such as circuit sub - part 3810(2)) is in an idle state, the power gating controller 3820 reduces the supply voltage of the circuit sub - part 3810(2) to a voltage that is less than the operating voltage but maintains the data stored in the memory. Optionally, the power gating controller 3820 can remove power from the circuit sub - part 3810(2), thus turning off the circuit sub - part 3810(2). Subsequently, if the circuit sub - part 3810(2) is needed to perform certain tasks, the power gating controller 3820 increases the power supply voltage of the circuit sub - part 3180(2) to a voltage suitable for operation.

[0323] The clock frequency controller 3830 monitors the power consumption of the circuit sub - parts 3810(0) - 3810(N). If the clock frequency controller 3830 determines that a particular circuit sub - part (e.g., circuit sub - part 3810(3)) consumes more power relative to other circuit sub - parts 3810, the clock frequency controller 3830 reduces the frequency of the clock signal associated with the circuit sub - part 3810(3). As a result, the power consumed by the circuit sub - part 3810(3) is reduced. Subsequently, if the clock frequency controller 3830 determines that the circuit sub - part 3810(3) consumes less power relative to other circuit sub - parts 3810, the clock frequency controller 3830 increases the frequency of the clock signal associated with the circuit sub - part 3810(3), thereby increasing the performance of the circuit sub - part 3810(3).

[0324] In this way, the power gating controller 3820 and the clock frequency controller 3830 reduce the overall power consumption of the PPU 200 and reduce the negative impact that one PPU partition 600 has on another PPU partition 600 due to temperature effects.

[0325] ​ is a flowchart of method steps for managing ​ the power consumption of the PPU 200 according to various embodiments. Although the method steps are described in connection with ​ a system, those of ordinary skill in the art will understand that any system configured to perform the method steps in any order is within the scope of this disclosure.

[0326] As shown, method 3900 begins at step 3902, where the power and clock frequency management system 3800 of PPU 200 monitors the active state of the VMs executing on the respective circuit sub - parts 3810 of PPU 200. At step 3904, the power and clock frequency management system 3800 determines whether any of the circuit sub - parts 3810 is in an idle state. If none of the circuit sub - parts 3810 is in an idle state, the method proceeds to step 3908. However, if one or more of the circuit sub - parts 3810 are in an idle state, the method proceeds to step 3906, where the power and clock frequency management system 3800 reduces the power supply voltage of the idle circuit sub - parts 3810. In particular, the power gating controller 3820 in the power and clock frequency management system 3800 reduces the supply voltage of the circuit sub - part 3810(2) to a voltage that is less than the operating voltage but maintains the data stored in the memory. Optionally, the power gating controller 3820 can remove power from the circuit sub - part 3810(2), thus turning off the circuit sub - part 3810(2).

[0327] At step 3908, the power and clock frequency management system 3800 monitors the power consumption of each SMC engine 700 in the PPU. At step 3910, the power and clock frequency management system 3800 determines whether one or more of the SMC engines 700 are consuming too much power relative to the other MC engines 700. If none of the SMC engines 700 is consuming too much power, the method proceeds to step 3902 to continue monitoring. However, if one or more of the SMC engines 700 are consuming too much power, the method proceeds to step 3912, where the clock frequency controller 3830 within the power and clock frequency management system 3800 reduces the clock frequency of one or more of the circuit sub - parts 3810 associated with the SMC engine 700 that is consuming too much power. Then the method proceeds to step 3902 to continue monitoring.

[0328] In summary, various embodiments include a parallel processing unit (PPU) that can be partitioned. Each partition is configured to simultaneously execute processing tasks associated with multiple processing contexts. A given partition includes one or more logical groupings or "slices" of GPU resources. Each slice provides sufficient computing, graphics, and memory resources to emulate the operation of the entire PPU. A hypervisor executing on a CPU performs various techniques on behalf of an administrator user to partition the PPU. Guest users are assigned to partitions and can then execute processing tasks within that partition that are isolated from any other guest users assigned to any other partition.

[0329] Relative to the prior art, one technical advantage of the disclosed technology is that, using the disclosed technology, the PPU can support multiple processing contexts simultaneously and be functionally isolated from each other. Thus, multiple CPU processes can effectively utilize the PPU resources via multiple different processing contexts without interfering with each other. Another technical advantage of the disclosed technology is that, since the PPU can be partitioned into isolated computing environments using the disclosed technology, the PPU can support a more reliable form of multi-tenancy relative to prior art methods that rely on processing sub-contexts to provide multi-tenancy functionality. Thus, when the disclosed technology is implemented, the PPU becomes more suitable for cloud-based deployments, where access to different partitions within the same PPU can be provided to different and potentially competing entities. These technical advantages represent one or more technological advancements over prior art methods.

[0330] 1. In some embodiments, a computer-implemented method includes: configuring one or more partitions of a first processor such that a first set of engines is included within the one or more partitions, where each engine included in the first set of engines runs within a given partition of the one or more partitions, and each engine included in the first set of engines is assigned a different set of hardware resources included in the given partition; and reconfiguring the one or more partitions in response to an execution state associated with at least one engine in the first set of engines such that a second set of engines is included within the one or more partitions, where each engine included in the second set of engines runs within a given partition of the one or more partitions, and each engine included in the second set of engines is assigned a different set of hardware resources included in the given partition.

[0331] 2. The computer-implemented method according to claim 1, wherein a first engine and a second engine are included in the first set of engines and the second set of engines, and the method further includes: before reconfiguring the one or more partitions: executing a first processing context of a plurality of processing contexts on the first engine, and executing a second processing context of the plurality of processing contexts on the second engine; and in response to reconfiguring the one or more partitions: executing a third processing context of the plurality of processing contexts on the first engine, and continuing to execute the second processing context on the second engine.

[0332] 3. The computer-implemented method according to claim 1 or 2, wherein the first set of engines includes a first engine and a second engine, and the second set of engines includes the first engine but not the second engine, and wherein reconfiguring the one or more partitions includes allocating at least a portion of the set of hardware resources assigned to the second engine to the first engine.

[0333] 4. The computer-implemented method according to claims 1-3, wherein the first engine and the second engine are included in a first partition included in the one or more partitions.

[0334] 5. The computer-implemented method according to claims 1-4, wherein the first set of engines includes a first engine, and the second set of engines includes the first engine and a second engine, and wherein reconfiguring the one or more partitions includes allocating a portion of a set of hardware resources assigned to the first engine to the second engine.

[0335] 6. The computer-implemented method according to claims 1-5, further comprising: executing a first instance of a first processing context in a plurality of processing contexts on a first engine included in the first set of engines; and executing a second instance of the first processing context on a second engine included in the first set of engines.

[0336] 7. The computer-implemented method according to claims 1-6, wherein the first instance of the first processing context and the second instance of the first processing context comply with the Automotive Safety Integrity Level D (ASIL-D) standard.

[0337] 8. The computer-implemented method according to claims 1-7, further comprising: detecting a memory access operation that accesses a first memory address associated with the first processor; determining, based on an index value, that the memory access operation is applicable to a first engine included in at least one of the first set of engines and the second set of engines; calculating a second memory address based on the first memory address and the index value, wherein the second memory address accesses a register associated with the first engine; and performing the memory access operation.

[0338] 9. The computer-implemented method according to claims 1-8, further comprising: detecting a memory write access operation that stores a first value at a first memory address associated with the first processor; detecting a memory read access operation for a second memory address associated with the first processor; performing a memory read access operation to read a second value at a third memory address based on the first value, wherein the third memory address is associated with a first engine included in at least one of the first set of engines and the second set of engines; and storing the second value at the second memory address.

[0339] 10. The computer-implemented method according to claims 1-9, further comprising: detecting a memory write access operation for storing a first value at a first memory address associated with the first processor; detecting a memory write access operation for storing a second value at a second memory address associated with the first processor; and performing a memory write access operation to write the second value to a third memory address based on the first value, wherein the third memory address is associated with a first engine included in at least one of the first set of engines and the second set of engines.

[0340] 11. In some embodiments, a non-transitory computer-readable medium stores program instructions that, when executed by a processor, cause the processor to perform the following steps: configuring one or more partitions of a first processor such that a first set of engines is included within the one or more partitions, wherein each engine included in the first set of engines operates within a given partition of the one or more partitions; and reconfiguring the one or more partitions in response to an execution state associated with at least one engine in the first set of engines such that a second set of engines is included within the one or more partitions, wherein each engine included in the second set of engines operates within a given partition of the one or more partitions.

[0341] 12. The non-transitory computer-readable medium according to claim 11, wherein a first engine and a second engine are included in the first set of engines and the second set of engines, and further comprising: before reconfiguring the one or more partitions: executing a first processing context among a plurality of processing contexts on the first engine, and executing a second processing context among the plurality of processing contexts on the second engine; and in response to reconfiguring the one or more partitions: executing a third processing context among the plurality of processing contexts on the first engine, and continuing to execute the second processing context on the second engine.

[0342] 13. The non-transitory computer-readable medium according to claim 11 or 12, wherein the first set of engines includes a first engine and a second engine, and the second set of engines includes the first engine and does not include the second engine, and wherein reconfiguring the one or more partitions includes allocating at least a portion of the hardware resources assigned to the second engine to the first engine.

[0343] 14. The non-transitory computer-readable medium according to claims 11-13, wherein the first engine and the second engine are included in a first partition included in the one or more partitions.

[0344] 15. The non-transitory computer-readable medium according to claims 11-14, wherein the first set of engines includes a first engine, and the second set of engines includes the first engine and a second engine, and wherein reconfiguring the one or more partitions includes allocating a portion of the hardware resources allocated to the first engine to the second engine.

[0345] 16. The non-transitory computer-readable medium according to claims 11-15, further comprising: executing a first instance of a first processing context among a plurality of processing contexts on a first engine included in the second set of engines; and executing a second instance of the first processing context on a third engine included in the second set of engines.

[0346] 17. The non-transitory computer-readable medium according to claims 11-16, wherein reconfiguring the one or more partitions includes: merging the hardware resources of a first partition included in the one or more partitions with the hardware resources of a second partition included in the one or more partitions to generate a first modified partition.

[0347] 18. The non-transitory computer-readable medium according to claims 11-17, wherein reconfiguring the one or more partitions includes: allocating a portion of the hardware resources of a first partition included in the one or more partitions to a second partition included in the one or more partitions.

[0348] 19. The non-transitory computer-readable medium according to claims 11-18, further comprising: executing a first processing context associated with a first virtual machine on a first engine included in the second set of engines; terminating a second processing context associated with a second virtual machine executed on a second engine included in the second set of engines; determining that the first virtual machine is in an idle state; determining that a second processor has resources available for executing the first processing context associated with the first virtual machine; and migrating the first virtual machine executed on the first processor to the second processor.

[0349] 20. In some embodiments, a system includes: a memory that stores a software application; and a processor that, when executing the software application, is configured to perform the following steps: configure one or more partitions of the processor such that the one or more partitions include a first set of engines, wherein each engine included in the first set of engines executes a different processing context included in a first set of processing contexts within a given partition of the one or more partitions, and each engine included in the first set of engines is assigned a different set of hardware resources included in the given partition; and in response to an execution state associated with at least one engine in the first set of engines, reconfigure the one or more partitions such that the one or more partitions include a second set of engines, wherein each engine included in the second set of engines executes a different processing context included in the first set of processing contexts within a given partition of the one or more partitions, and each engine included in the second set of engines is assigned a different set of hardware resources included in the given partition.

[0350] In any way, any and all combinations of any claim element recited in any claim and / or any element described in this application fall within the intended scope of this embodiment and protection.

[0351] Descriptions of various embodiments have been given for illustrative purposes, but are not intended to be exhaustive or limited to the disclosed embodiments. Many modifications and variations will be apparent to those of ordinary skill in the art without departing from the scope and spirit of the described embodiments.

[0352] Aspects of the embodiments of the present invention can be implemented as a system, method, or computer program product. Accordingly, aspects of the present disclosure can take the form of an entirely hardware embodiment, an entirely software embodiment (including firmware, resident software, microcode, etc.), or an embodiment combining software and hardware aspects, which are generally referred to herein as "modules", "systems", or "computers". Additionally, any hardware and / or software technology, process, function, component, engine, module, or system described in this disclosure can be implemented as a circuit or a collection of circuits. Furthermore, aspects of the present disclosure can take the form of a computer program product embodied in one or more computer-readable media having computer-readable program code embodied thereon.

[0353] Any combination of one or more computer-readable media can be utilized. The computer-readable media can be a computer-readable signal medium or a computer-readable storage medium. A computer-readable storage medium can be, for example, but not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any suitable combination of the foregoing. More specific examples (a non-exhaustive list) of the computer-readable storage medium would include the following: an electrical connection having one or more wires, a portable computer floppy disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any other suitable combination of the foregoing. In the context of this document, a computer-readable storage medium can be any tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device.

[0354] Aspects of the present disclosure have been described above with reference to the flowcharts and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the present disclosure. It will be understood that each block of the flowcharts and / or block diagrams, and combinations of blocks in the flowcharts and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general purpose computer, a special purpose computer, or other programmable data processing apparatus to produce a machine. When the instructions are executed by a processor of a computer or other programmable data processing apparatus, the functions / acts specified in the blocks of the flowcharts and / or block diagrams can be implemented. Such a processor can be, but is not limited to, a general purpose processor, a special purpose processor, an application specific processor, or a field programmable gate array.

[0355] The flowcharts and block diagrams in the figures illustrate the architecture, functionality, and operation of systems, methods, and computer program products possible according to various embodiments of the present invention. In this regard, each block in the flowcharts or block diagrams can represent a module, segment, or portion of code, which includes one or more executable instructions for implementing one or more specified logical functions. It should also be noted that in some alternative implementations, the functions noted in the blocks may occur out of the order noted in the figures. For example, depending on the functionality involved, two blocks shown in succession may in fact be executed substantially concurrently, or sometimes the blocks may be executed in the reverse order. It should also be noted that each block of the block diagrams and / or flowchart illustrations, and combinations of blocks in the block diagrams and / or flowchart illustrations, can be implemented by a system based on dedicated hardware for performing the specified functions or acts, or by a combination of dedicated hardware and computer instructions.

[0356] Although the foregoing is directed to embodiments of the present disclosure, other and further embodiments of the present disclosure may be devised without departing from the basic scope thereof, and the scope of the present disclosure is determined by the appended claims.

Claims

1. A computer-implemented method, comprising: Configuring a first processor into a first set of partitions, where each partition in the first set of partitions represents a part of the first processor; Configuring a first partition included in the first set of partitions to include a first set of engines, where each processing context in a first set of processing contexts within the first partition can be assigned to a different engine in the first set of engines, and each engine included in the first set of engines runs within the first partition and is assigned a first set of hardware resources included in the first partition; And In response to an execution state associated with at least one engine in the first set of engines: Reconfiguring the first processor into a second set of partitions, where the second set of partitions is different from the first set of partitions; And Reconfiguring a second partition included in the second set of partitions to include a second set of engines, where each processing context in a second set of processing contexts within the second partition can be assigned to a different engine in the second set of engines, and each engine included in the second set of engines runs within the second partition and is assigned a second set of hardware resources included in the second partition; Where: Each engine included in the first set of engines includes a logical grouping of a subset of the first set of hardware resources, and Each engine included in the second set of engines includes a logical grouping of a subset of the second set of hardware resources.

2. The computer-implemented method according to claim 1, where a first engine and a second engine are included in the first set of engines and the second set of engines, and the method further comprises: Before reconfiguring the second partition: Executing a first processing context among a plurality of processing contexts on the first engine, and Executing a second processing context among the plurality of processing contexts on the second engine; And In response to reconfiguring the second partition: Executing a third processing context among the plurality of processing contexts on the first engine, and Continuing to execute the second processing context on the second engine.

3. The computer-implemented method according to claim 1, where the first set of engines includes a first engine and a second engine, and the second set of engines includes the first engine but does not include the second engine, and where reconfiguring the second partition includes allocating at least a part of the first set of hardware resources assigned to the second engine to the first engine.

4. The computer-implemented method according to claim 3, where the first engine and the second engine are included in the first partition.

5. The computer-implemented method according to claim 1, where the first set of engines includes a first engine, and the second set of engines includes the first engine and a second engine, and where reconfiguring the second partition includes allocating a part of the first set of hardware resources assigned to the first engine to the second engine.

6. The computer-implemented method according to claim 1, further comprising: Execute a first instance of a first processing context among a plurality of processing contexts on a first engine included in the first set of engines; And Execute a second instance of the first processing context on a second engine included in the first set of engines.

7. The computer-implemented method according to claim 6, wherein the first instance of the first processing context and the second instance of the first processing context comply with the Automotive Safety Integrity Level D (ASIL-D) standard.

8. The computer-implemented method according to claim 1, further comprising: Detect a memory access operation that accesses a first memory address associated with the first processor; Based on an index value, determine that the memory access operation is applicable to a first engine included in at least one of the first set of engines and the second set of engines; Calculate a second memory address based on the first memory address and the index value, wherein the second memory address accesses a register associated with the first engine; And Execute the memory access operation.

9. The computer-implemented method according to claim 1, further comprising: Detect a first memory write access operation that stores a first value at a first memory address associated with the first processor; Detect a first memory read access operation for a second memory address associated with the first processor; Execute a second memory read access operation to read a second value at a third memory address, wherein the third memory address is based on the first value and is associated with a first engine included in at least one of the first set of engines and the second set of engines; And Execute a second memory write access operation to store the second value at the second memory address.

10. The computer-implemented method according to claim 1, further comprising: Detect a memory write access operation that stores a first value at a first memory address associated with the first processor; Detect a memory write access operation that stores a second value at a second memory address associated with the first processor; And Execute a memory write access operation to write the second value to a third memory address based on the first value, wherein the third memory address is associated with a first engine included in at least one of the first set of engines and the second set of engines.

11. A non-transitory computer-readable medium storing program instructions that, when executed by a processor, cause the processor to perform the following steps: Configure a first processor into a first set of partitions, wherein each partition included in the first set of partitions represents a part of the first processor; Configure a first partition included in the first set of partitions to include a first set of engines, wherein each processing context in the first set of processing contexts included in the first partition can be assigned to a different engine in the first set of engines, and each engine included in the first set of engines runs within the first partition and is assigned a first set of hardware resources included in the first partition; And In response to an execution state associated with at least one engine in the first set of engines: Reconfigure the first processor into a second set of partitions, where the second set of partitions is different from the first set of partitions; And Reconfigure a second partition included in the second set of partitions to include a second set of engines, where each processing context in the second set of processing contexts included within the second partition can be assigned to a different engine in the second set of engines, and each engine included in the second set of engines runs within the second partition and is assigned a second set of hardware resources included in the second partition; Where: Each engine included in the first set of engines includes a logical grouping of a subset of the first set of hardware resources, and Each engine included in the second set of engines includes a logical grouping of a subset of the second set of hardware resources.

12. The non-transitory computer-readable medium according to claim 11, wherein a first engine and a second engine are included in the first set of engines and the second set of engines, and further comprising: Before reconfiguring the second partition: Execute a first processing context among a plurality of processing contexts on the first engine, and Execute a second processing context among the plurality of processing contexts on the second engine; And In response to reconfiguring the second partition: Execute a third processing context among the plurality of processing contexts on the first engine, and Continue to execute the second processing context on the second engine.

13. The non-transitory computer-readable medium according to claim 11, wherein the first set of engines includes a first engine and a second engine, and the second set of engines includes the first engine and does not include the second engine, and wherein reconfiguring the second partition includes allocating at least a portion of the hardware resources assigned to the second engine to the first engine.

14. The non-transitory computer-readable medium according to claim 13, wherein the first engine and the second engine are included in the first partition.

15. The non-transitory computer-readable medium according to claim 11, wherein the first set of engines includes a first engine, and the second set of engines includes the first engine and a second engine, and wherein reconfiguring the second partition includes allocating a portion of the hardware resources assigned to the first engine to the second engine.

16. The non-transitory computer-readable medium according to claim 11, further comprising: Execute a first instance of a first processing context among a plurality of processing contexts on a first engine included in the second set of engines; And Execute a second instance of the first processing context on a third engine included in the second set of engines.

17. The non-transitory computer-readable medium according to claim 11, wherein reconfiguring the second partition comprises: Merge the hardware resources of the first partition with the hardware resources of the second partition to generate a first modified partition.

18. The non-transitory computer-readable medium according to claim 11, wherein reconfiguring the second partition comprises: Allocate a portion of the hardware resources of the first partition to the second partition.

19. The non-transitory computer-readable medium according to claim 11, further comprising: Execute a first processing context associated with a first virtual machine on a first engine included in the second set of engines; Terminate a second processing context associated with a second virtual machine executed on a second engine included in the second set of engines; Determine that the first virtual machine is in an idle state; Determine that a second processor has resources available for executing the first processing context associated with the first virtual machine; And Migrate the first virtual machine executed on the first processor to the second processor.

20. A system, comprising: A memory that stores software applications; And A processor that, when executing the software applications, is configured to perform the following steps: Configure a first processor into a first set of partitions, where each partition in the first set of partitions represents a portion of the first processor; Configure a first partition included in the first set of partitions to include a first set of engines, where each processing context in a first set of processing contexts included within the first partition can be assigned to a different engine in the first set of engines, and each engine included in the first set of engines runs within the first partition and is assigned a first set of hardware resources included in the first partition; And In response to an execution state associated with at least one engine in the first set of engines: Reconfigure the first processor into a second set of partitions, where the second set of partitions is different from the first set of partitions; And Reconfigure a second partition included in the second set of partitions to include a second set of engines, where each processing context in a second set of processing contexts included within the second partition can be assigned to a different engine in the second set of engines, and each engine included in the second set of engines runs within the second partition and is assigned a second set of hardware resources included in the second partition; Wherein: Each engine included in the first set of engines includes a logical grouping of a subset of the first set of hardware resources, and Each engine included in the second set of engines includes a logical grouping of a subset of the second set of hardware resources.

Citation Information

Patent Citations

  • Graphics engine partitioning mechanism

    US20180308198A1

  • Configuring sets of processor cores for processing instructions

    US7734895B1

  • System utilization through dedicated uncapped partitions

    US8302102B2