Techniques for configuring a processor to function as multiple independent processors
By partitioning the GPU and configuring multiple isolated engines, the problem of uneven distribution of GPU resources is solved, efficient parallel execution of multiple contexts is achieved, and the performance and utilization of GPU are improved.
Patent Information
- Application Number
- CN202010214099.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2019-09-05
- Filing Date
- 2020-03-24
- Publication Date
- 2025-05-06
- Estimated Expiration
- 2040-03-24
AI Technical Summary
When multiple CPU processes are executed simultaneously on the GPU, the prior art results in uneven distribution of GPU resources, which in turn leads to performance degradation and resource idleness.
By partitioning the hardware resources of the GPU, multiple logical partitions are generated, multiple engines are configured in each partition, and each engine is functionally isolated to support parallel execution of multiple contexts.
Multi-context support for GPUs is implemented in functions, avoiding the problem of uneven resource allocation and improving the overall performance and utilization of GPUs.
Smart Images

Figure CN112445611B_ABST
Abstract
Description
Technical Field
[0001] Various embodiments relate generally to parallel processing architectures and, more particularly, to techniques for configuring a processor to function as multiple independent processors. Background Art
[0002] A conventional central processing unit (CPU) typically includes a relatively small number of processing cores that can execute a relatively small number of CPU processes. In contrast, a conventional graphics processing unit (GPU) typically includes hundreds of processing cores that can execute hundreds of threads in parallel with each other. Therefore, given the large number of processing resources that can be deployed when using a conventional GPU, a conventional GPU can typically perform certain processing tasks faster and more efficiently than a conventional CPU.
[0003] In some embodiments, a CPU process executed on a CPU can offload a given processing task to a GPU so that the processing task can be executed faster. In doing so, the CPU process generates a processing context on the GPU, which specifies the target state of various GPU resources to be implemented to perform the processing task. Those GPU resources may include processing, graphics, and memory resources, etc. Then, the CPU process starts a thread group on the GPU according to the processing context, and the thread group uses various GPU resources to perform the processing task. In many of these types of implementations, the GPU is configured according to only one processing environment at a time. However, in some cases, the CPU needs to offload more than one CPU process to the GPU within the same time interval. In this case, the CPU can dynamically change the processing context implemented on the GPU at different time points to provide serial services for these CPU processes during a certain time interval. However, one disadvantage of this approach is that the processing tasks offloaded by some CPU processes cannot fully utilize the resources of the GPU. Therefore, when one or more processing tasks associated with these CPU processes are executed serially on the GPU, some GPU resources may be idle, which reduces overall GPU performance and utilization.
[0004] One approach to executing multiple CPU processes simultaneously on the GPU is to generate multiple different processing subcontexts within a given "parent" processing context, and assign each different processing subcontext to a different CPU process. Multiple CPU processes can then simultaneously launch different thread groups on the GPU, where each thread group utilizes specific GPU resources configured according to a specific processing subcontext. With this approach, the GPU can be utilized more efficiently because more than one CPU process can offload processing tasks to the GPU at the same time, potentially avoiding situations where some GPU resources are left idle.
[0005] One problem with the above approach is that CPU processes associated with different processing subcontexts may unfairly consume GPU resources that should be more evenly allocated or distributed among the different processing subcontexts. For example, a first CPU process may launch a first thread group in a first processing subcontext that performs a large number of read requests and consumes a large amount of available GPU memory bandwidth. A second CPU process may then launch a second thread group in a second processing subcontext that also performs a large number of read requests. However, because the first thread group has consumed a lot of available GPU memory bandwidth, the second thread set may experience high latency, which may cause the second CPU process to stall.
[0006] Another problem with the above approach is that because the processing subcontexts share the parent context, any errors that occur while threads associated with one processing subcontext are executing may interfere with the execution of other threads associated with another processing subcontext that shares the same parent context. For example, a first CPU process may launch a first thread group associated with a first processing subcontext to perform a first processing task. A second CPU process may launch a second thread group associated with a second processing subcontext, and the second thread group may subsequently malfunction and fail. In order to recover from the failure, the GPU will have to reset the parent context, which will automatically reset the first processing subcontext and the second processing subcontext. In this case, the execution of the first thread group will be interrupted even if the failure was caused by the second thread group rather than the first thread group.
[0007] As previously stated, there is a need in the art for more efficient techniques for configuring a GPU to perform processing tasks associated with multiple contexts. Summary of the invention
[0008] Various embodiments include a computer-implemented method comprising: partitioning a set of hardware resources included in a processor to generate a first logical partition including a first subset of the hardware resources; and generating a plurality of engines within the first logical partition, wherein each engine in the plurality of engines is allocated a different portion of the first subset of the hardware resources and executes in functional isolation from all other engines included in the plurality of engines.
[0009] One technical advantage of the disclosed technology over the prior art is that, using the disclosed technology, a parallel processing unit (PPU) (e.g., GPU) can support multiple contexts simultaneously and be functionally isolated from each other. Therefore, multiple CPU processes can effectively utilize PPU resources by executing multiple different contexts simultaneously without interfering with each other. BRIEF DESCRIPTION OF THE DRAWINGS
[0010] In order that the manner in which the above-mentioned features of various embodiments can be understood in detail, the inventive concept briefly summarized above can be described in more detail by reference to various embodiments (some of which are shown in the accompanying drawings). However, it should be noted that the accompanying drawings only illustrate typical embodiments of the inventive concept and therefore should not be considered to limit the scope in any way, and there are other equivalent embodiments.
[0011] Figure 1 is a block diagram of a computer system configured to implement one or more aspects of various embodiments;
[0012] Figure 2 According to various embodiments, the Figure 1 A block diagram of a parallel processing unit (PPU) in a parallel processing subsystem of FIG.
[0013] Figure 3 According to various embodiments, the Figure 2 A block diagram of a general processing cluster in a parallel processing unit of;
[0014] Figure 4 According to various embodiments, the Figure 2 A block diagram of a partition unit in a PPU;
[0015] Figure 5 According to various embodiments, the Figure 2 A block diagram of various PPU resources in a PPU;
[0016] Figure 6 According to various embodiments Figure 1 An example of how the hypervisor can logically group PPU resources into PPU partition sets;
[0017] Figure 7 It shows the various embodiments Figure 1 How the hypervisor configures a PPU partition set to implement one or more simultaneous multi-context (SMC) engines;
[0018] Fig. 8A According to various embodiments Figure 7 A more detailed illustration of the DRAM;
[0019] Figure 8B shows how to address according to various embodiments Figure 8B The various DRAM parts;
[0020] Fig. 9 According to various embodiments Figure 1 Data flow diagram of how the hypervisor partitions and configures the PPU;
[0021] Fig.10is a flow chart of method steps for partitioning and configuring a PPU on behalf of one or more users according to various embodiments;
[0022] Fig.11 According to various embodiments, Figure 1 The hypervisor may configure a partition configuration table of one or more PPU partitions;
[0023] Fig.12 It shows the various embodiments Figure 1 How the hypervisor partitions the PPU to generate one or more PPU partitions;
[0024] Fig.13 It shows the various embodiments Figure 1 How the hypervisor allocates various PPU resources during partitioning;
[0025] Fig.14A shows how multiple guest OSes running multiple VMs can simultaneously launch multiple processing contexts within one or more PPU partitions according to various embodiments;
[0026] Fig. 14B shows how a host OS can simultaneously launch multiple processing environments within one or more PPU partitions according to various embodiments;
[0027] Fig.15 It shows the various embodiments Figure 1 How the hypervisor assigns virtual address space identifiers to different SMC engines;
[0028] Fig.16 illustrates how a memory management unit translates local virtual address space identifiers when mitigating faults according to various embodiments;
[0029] Fig.17 FIG. 2 shows a flow chart of the process flow when migrating processing context between SMC engines on different PPUs according to various embodiments. Figure 1 How to achieve soft voice in management procedures;
[0030] Fig.18 is a flow chart of method steps for configuring computing resources within a PPU to simultaneously support operations associated with multiple processing contexts according to various embodiments;
[0031] Fig.19 shows a boundary option set according to various embodiments, according to the boundary option set Figure 1 The hypervisor may generate one or more PPU memory partitions;
[0032] Fig. 20 It shows the various embodiments Figure 1 An example of how a hypervisor may partition a PPU memory to generate one or more PPU memory partitions;
[0033] Fig.21 It shows the various embodiments Fig.16 How the memory management unit provides access to different PPU memory partitions;
[0034] Fig. 22 It shows the various embodiments Fig.16 How the memory management unit performs various address translations;
[0035] Fig.23 It shows the various embodiments Fig.16 How the memory management unit of a processor provides support operations associated with multiple processing contexts simultaneously;
[0036] Fig.24 is a flow chart of method steps for configuring memory resources within a PPU to simultaneously support operations associated with multiple processing contexts according to various embodiments;
[0037] Fig.25 is a diagram showing various embodiments of the present invention. Figure 2 The set of timelines for VM-level time slices associated with the PPU;
[0038] Fig.26 According to various other embodiments, Figure 2 Another timeline set of VM-level time slices associated with the PPU;
[0039] Fig. 27 is a diagram showing various embodiments of the present invention. Figure 2 The timeline of the SMC-level time slice associated with the PPU;
[0040] Fig.28 shows how a VM can be migrated from one PPU to another PPU according to various embodiments;
[0041] Fig.29 is an illustration according to various embodiments and Figure 2 A set of timelines for fine-grained VM migration associated with a PPU;
[0042] Figures 30A-30B A method for performing Figure 2 A flowchart of the method steps for time slicing a VM in a PPU;
[0043] Fig.31 is a memory map according to various embodiments, which shows how the BAR0 address space is mapped to Figure 2The privileged register space within the PPU;
[0044] Fig.32 According to various embodiments, Figure 2 A flowchart of the method steps for addressing the privileged register address space in the PPU;
[0045] Fig.33 According to various embodiments, Figure 2 A block diagram of a performance monitoring system for a PPU;
[0046] Figures 34A-34B It shows the various embodiments Fig.33 performance of various configurations of multiplexer units;
[0047] Fig.35 According to various embodiments, Figure 2 A block diagram of the PPU performance monitor aggregation system;
[0048] Fig.36 According to various embodiments, Fig.35 The format of a trigger packet associated with a performance monitor aggregation system;
[0049] Fig.37 According to various embodiments, a method for monitoring Figure 2 A flow chart of steps of a method for determining the performance of a PPU;
[0050] Fig.38 According to various embodiments, Figure 2 A block diagram of a power and clock frequency management system of a PPU; and
[0051] Fig.39 is used to manage Figure 2 Flow chart of method steps for calculating power consumption of PPU 200. DETAILED DESCRIPTION
[0052] In the following description, numerous specific details are set forth to provide a more thorough understanding of various embodiments. However, it will be apparent to one skilled in the art that the inventive concept may be practiced without one or more of these specific details.
[0053] As described above, a conventional GPU can generally perform certain processing tasks faster than a conventional CPU. In some configurations, a CPU process executing on a CPU can offload a given processing task to a GPU in order to perform the processing task faster. To do so, the CPU process generates a processing context on the GPU that specifies the target state of various GPU resources, and then launches a thread group on the GPU to perform the processing task.
[0054] In some cases, multiple CPU processes may be required to offload processing tasks to the GPU during the same time interval. However, the GPU can only be configured according to one processing context at a time. In this case, the CPU can dynamically change the processing context of the GPU at different points in time to continuously serve multiple CPU processes across time intervals. However, some CPU processes may not fully utilize GPU resources when performing processing tasks, sometimes leaving various GPU resources idle. To solve this problem, the CPU can generate multiple processing subcontexts in a "parent" processing context and assign these processing subcontexts to different CPU processes. These CPU processes can then launch different thread groups on the GPU at the same time, and each thread group can utilize specific GPU resources configured according to a specific processing subcontext. This method can be implemented to more efficiently utilize GPU resources. However, this method has several disadvantages.
[0055] First, CPU processes associated with different processing subcontexts can unfairly consume GPU resources that should be shared fairly between different processing subcontexts, leading to a situation where one CPU process can stall the progress of another CPU process. Second, because processing subcontexts share a parent processing context, any error that occurs during the execution of a thread associated with one processing subcontext can disrupt the execution of threads associated with other processing subcontexts included in the same parent processing context. In some cases, a failure that occurs in one processing subcontext can cause all other processing subcontexts in the same parent processing context to be reset and restarted.
[0056] In general, the above disadvantages associated with processing subcontexts limit the extent to which traditional GPUs can support multi-tenancy. As referred to herein, "multi-tenancy" refers to a GPU configuration in which multiple users or "tenants" use GPU resources to perform processing operations simultaneously or during overlapping time intervals. In general, traditional GPUs provide support for multi-tenancy by allowing different tenants to perform different processing tasks using different processing subcontexts within a given parent processing context. However, processing subcontexts are not isolated computing environments because processing tasks performed in different processing subcontexts may interfere with each other for the various reasons described above. Therefore, any given tenant occupying a given GPU will have a negative impact on the quality of service provided by the GPU to other tenants. These factors may reduce the attractiveness of cloud-based GPU deployments in which multiple users may access the same GPU simultaneously.
[0057] To address these issues, various embodiments include parallel processing units (PPUs) that can be divided into partitions. Each partition is configured to simultaneously execute processing tasks associated with multiple processing environments. A given partition includes a logical grouping or "slice" of one or more GPU resources. Each slice provides sufficient computing, graphics, and memory resources to simulate the operation of an entire PPU. A hypervisor executing on the CPU performs various techniques to partition the PPU on behalf of an administrator user. A guest user is assigned to a partition and can then perform processing tasks within that partition that are isolated from any other guest users assigned to any other partition.
[0058] One technical advantage of the disclosed technology relative to the prior art is that, using the disclosed technology, the PPU can support multiple processing contexts simultaneously and functionally isolated from each other. Therefore, multiple CPU processes can effectively utilize PPU resources via multiple different processing contexts without interfering with each other. Another technical advantage of the disclosed technology is that because the PPU can be partitioned into isolated computing environments using the disclosed technology, the PPU can support a more robust form of multi-tenancy relative to the prior art methods that rely on processing sub-contexts to provide multi-tenant functionality. Therefore, when the disclosed technology is implemented, the PPU becomes more suitable for cloud-based deployments in which different and potentially competing entities can be provided with access to different partitions within the same PPU. These technical advantages represent one or more technical advances relative to the prior art methods.
[0059] System Overview
[0060] Figure 1 1 is a block diagram of a computer system configured to implement one or more aspects of the present invention. As shown, computer system 100 includes a central processing unit (CPU) 110, a system memory 120, and a parallel processing subsystem 130 coupled together via a memory bridge 132. Parallel processing subsystem 130 is coupled to memory bridge 132 via a communication path 134. One or more display devices 136 can be coupled to parallel processing subsystem 130. Computer system 100 also includes a system disk 140, one or more add-in cards 150, and a network adapter 160. System disk 140 is coupled to I / O bridge 142. I / O bridge 142 is coupled to memory bridge 132 via communication path 144, and is also coupled to input device 146. One or more add-in cards 150 and network adapter 160 are coupled together via switch 148, which is in turn coupled to I / O bridge 142.
[0061] Memory bridge 132 is a hardware unit that facilitates communication between CPU 110, system memory 120, and parallel processing subsystem 130, as well as other components of computer system 100. For example, memory bridge 132 may be a north bridge chip. Communication path 134 is a high-speed and / or high-bandwidth data connection that facilitates low-latency communication between parallel processing subsystem 130 and memory bridge 132 across one or more independent channels. For example, communication path 134 may be a peripheral component interconnect express (PCIe) link, accelerated graphics port (AGP), HyperTransport, or any other technically feasible communication bus type.
[0062] The I / O bridge 142 is a hardware unit that facilitates input and / or output operations performed using the system disk 140, input devices 146, one or more add-in cards 150, network adapter 160, and various other components of the computer system 100. For example, the I / O bridge 143 can be a south bridge chip. The communication path 144 is a high-speed and / or high-bandwidth data connection that facilitates low-latency communication between the memory bridge 132 and the I / O bridge 142. For example, the communication path 142 can be a PCIe link, AGP, HyperTransport, or any other technically feasible communication bus type. With the configuration shown, any component coupled to the memory bridge 132 or the I / O bridge 142 can communicate with any other component coupled to the memory bridge 132 or the I / O bridge 142.
[0063] The CPU 110 is a processor configured to coordinate the overall operation of the computer system 100. As such, the CPU 110 executes instructions to issue commands to various other components included in the computer system 100. The CPU 110 is also configured to execute instructions in order to process data generated and / or stored by any other components included in the computer system 100, including the system memory 120 and the system disk 140. The system memory 120 and the system disk 140 are memory devices that include computer-readable media configured to store data and software applications. The system memory 120 includes a device driver 122 and a hypervisor 124, the operation of which is described in more detail below. The parallel processing subsystem 130 includes one or more parallel processing units (PPUs) that are configured to perform multiple operations simultaneously in a highly parallel processing architecture. Each PPU includes one or more computing engines that perform general computing operations in parallel and / or one or more graphics engines that perform graphics-oriented operations in parallel. A given PPU may be configured to generate pixels for display via a display device 136. An exemplary PPU is described below in conjunction with Figure 2-Figure 4 Describe in more detail.
[0064] Device driver 122 is a software application that, when executed by CPU 110, operates to interface between CPU 110 and PP subsystem 130. In particular, device driver 122 allows CPU 110 to offload various processing operations to PP subsystem 130 for highly parallel execution, including general purpose computing operations as well as graphics processing operations. Hypervisor 124 is a software application that, when executed by CPU 110, partitions the various computing, graphics, and memory resources included in PP subsystem 130 to provide independent use of those resources to separate users, as will be discussed below in conjunction with the present disclosure. Figure 5-Figure 10 Describe in more detail.
[0065] In various embodiments, some or all of the components of computer system 100 can be implemented in a cloud-based environment that is potentially distributed across a wide geographic area. For example, the various components of computer system 100 can be deployed across geographically different data centers. In such an embodiment, the various components of computer system 100 can communicate with each other through one or more networks (including any number of local intranets and / or the Internet). In various other embodiments, some components of computer system 100 can be implemented via one or more virtualized devices. For example, CPU 110 can be implemented as a virtual instance of a hardware CPU. In some embodiments, some or all of parallel processing subsystem 130 can be integrated with one or more other components of computer system 100 to form a single chip, such as a system on chip (SoC).
[0066] Those skilled in the art will appreciate that the architecture of the computer system 100 is flexible enough to be implemented across a wide range of potential scenarios and use cases. For example, the computer system 100 can be implemented in a cloud computing center to expose general computing capabilities and / or general graphics processing capabilities to one or more users. Alternatively, the computer system 100 can be deployed in an automotive implementation to perform data processing operations associated with vehicle navigation. Those skilled in the art will further appreciate that the various components of the computer system 100 and the connection topology between these components can be modified in any technically feasible manner without departing from the overall scope and spirit of the present embodiment.
[0067] Figure 2 According to various embodiments, the Figure 12 is a block diagram of a PPU in a parallel processing subsystem of FIG. As shown, PPU 200 includes I / O unit 210, host interface 220, system (sys) pipeline 230, processing cluster array 240, crossbar switch unit 250, and memory interface 260. PPU 200 is coupled to PPU memory 270. Each component shown can be implemented by any technically feasible hardware type and / or any technically feasible combination of hardware and software.
[0068] I / O unit 210 is coupled to memory bridge 132 via communication path 134. Figure 1 The I / O unit 210 is also coupled to a host interface 220 and a crossbar unit 250. The host interface 220 is coupled to one or more physical copy engines (PCEs) 222, which in turn are coupled to one or more PCE counters 224. The host interface 220 is also coupled to a system pipeline 230. A given system pipeline 230 includes a front end 232, a task / work unit 234, and a performance monitor (PM) 236, and is coupled to a processing cluster array 240. The processing cluster array 240 includes general processing clusters (GPCs) 242 (0) to 242 (A), where A is a positive integer. The processing cluster array 240 is coupled to the crossbar unit 250. The crossbar unit 250 is coupled to a memory interface 260. The memory interface 260 includes partition units 262 (0) to 262 (B), where B is a positive integer value. Each partition unit 262 can be connected to the crossbar unit 250, respectively. The PPU memory 270 includes dynamic random access memory (DRAM) 272(0) to 272(C), where C is a positive integer value. To facilitate simultaneous operation on multiple processing contexts, various units within the PPU 200 are replicated as follows: (a) the host interface 220 includes PBDMA 520(0) to 520(7); (b) the system pipeline 230 includes system pipeline 230(0) to 230(7), such that the task / work unit 234 corresponds to SKED 500(0) to SKED 500(7); and the task / work unit 234 corresponds to CWD 560(0) to 560(7).
[0069] In operation, the I / O unit 210 obtains various types of command data from the CPU 110 and distributes the command data to the relevant components of the PPU 200 for execution. In particular, the I / O unit 210 obtains command data associated with a processing task from the CPU 110 and routes the command data to the host interface 220. The I / O unit 210 also obtains command data associated with a memory access operation from the CPU 110 and routes the command data to the crossbar unit 250. The command data associated with a processing task typically includes one or more pointers to task metadata (TMD) stored in a command queue in the PPU memory 270 or elsewhere within the computer system 100. A given TMD is an encoded processing task that describes the index of the data to be processed, the operations to be performed on the data, the state parameters associated with those operations, the execution priority, and other processing task-oriented information.
[0070] The host interface 220 receives command data related to the processing tasks from the I / O unit 210 and then distributes the command data to the system pipelines 230 via one or more command streams. In some configurations, the host interface 210 generates a different command stream for each different system pipeline 230, where a given command stream includes a pointer to a TMD related to the corresponding system pipeline 230.
[0071] A given system pipeline 230 performs various pre-processing operations with the received command data to facilitate execution of corresponding processing tasks on the GPCs 242 within the processing cluster array 240. Upon receiving command data associated with one or more processing tasks, the front end 232 in a given system pipeline 230 obtains the associated processing tasks and relays these processing tasks to the task / work unit 234. The task / work unit 234 configures one or more GPCs 242 to an operating state suitable for executing the processing tasks, and then transmits the processing tasks to those GPCs 242 for execution. Each system pipeline 230 can offload copy tasks to one or more PCEs 222 that perform dedicated copy operations. The PCE counters 224 track the use of the PCEs 222 in order to balance the copy operation workload between different system pipelines 230. The PM 236 monitors the overall performance and / or resource consumption of the corresponding system pipeline 230, and can limit various operations performed by the system pipeline 230 to maintain balanced resource consumption across all system pipelines 230.
[0072] Each GPC 242 includes multiple parallel processing cores that are capable of executing a large number of threads simultaneously and with any degree of independence and / or isolation from other GPCs 242. For example, a given GPC 242 may execute hundreds or thousands of concurrent threads in conjunction or isolation with any other GPC 242. Concurrent groups of threads executing on a GPC 242 may execute separate instances of the same program or separate instances of different programs. In some configurations, GPCs 242 are shared among all system pipelines 230, while in other configurations, different groups of GPCs 242 are assigned to operate with specific system pipelines 230. Each GPC 242 receives processing tasks from one or more system pipelines 230 and, in response, launches one or more thread groups to perform those processing tasks and generate output data. After completing a given processing task, the given GPC 242 transmits the output data to another GPC 242 for further processing, or to the crossbar switch unit 250 for appropriate routing. The following is a description of the system pipelines 230 used in conjunction with the system pipelines 230. Figure 3 An exemplary GPC is described in more detail.
[0073] The crossbar unit 250 is a switching mechanism that routes various types of data between the I / O unit 210, the processing cluster array 240, and the memory interface 260. As described above, the I / O unit 210 sends command data related to memory access operations to the crossbar unit 250. In response, the crossbar unit 250 submits the associated memory access operation to the memory interface 260 for processing. In some cases, the crossbar unit 250 also routes the read data returned from the memory interface 260 back to the component requesting the read data. As described above, the crossbar unit 250 also receives output data from the GPC 242, which can then be routed to the I / O unit 210 for transmission to the CPU 110, or the data can be routed to the memory interface 260 for storage and / or processing. The crossbar unit 250 is generally configured to route data between the GPCs 242 and from any GPC 242 to any partition unit 262. In various embodiments, the crossbar unit 250 can implement virtual channels to separate traffic flows between the GPCs 242 and the partition units 262. In various embodiments, crossbar unit 250 may allow for non-shared paths between a set of GPCs 242 and a set of partition units 262 .
[0074] The memory interface 260 implements partition units 262 to provide high bandwidth memory access to the DRAM 272 in the PPU memory 270. Each partition unit 262 can perform memory access operations in parallel with different DRAMs 272, thereby efficiently utilizing the available memory bandwidth of the PPU memory 270. A given partition unit 262 also provides cache support via one or more internal caches. Figure 4 An exemplary partitioning unit 262 is described in further detail.
[0075] In general, PPU memory 270, and in particular DRAM 272, may be configured to store any technically feasible data associated with general computing applications and / or graphics processing applications. For example, DRAM 272 may store large matrices of data values associated with a neural network in a general computing application, or alternatively, store one or more frame buffers including multiple render targets in a graphics processing application. In various embodiments, DRAM 272 may be implemented via any technically feasible memory device.
[0076] The architecture set forth above allows PPU 200 to perform a wide variety of processing operations in a rapid manner and asynchronously with respect to the operations of CPU 110. In particular, the parallel architecture of PPU 200 allows a large number of operations to be performed in parallel and with any degree of independence from each other and from operations performed on CPU 110, thereby speeding up the overall performance of those operations.
[0077] In one embodiment, the PPU 200 may be configured to perform general computing operations to speed up calculations involving large data sets. Such data sets may involve financial time series, dynamic simulation data, real-time sensor readings, neural network weight matrices and / or tensors, and machine learning parameters, etc. In another embodiment, the PPU 200 may be configured to be used as a graphics processing unit (GPU), which implements one or more graphics rendering pipelines to generate pixel data based on graphics commands generated by the CPU 110. The PPU 200 may then output the pixel data as one or more frames via the display device 136. The PPU memory 170 may be configured to be used as a graphics memory, which stores one or more frame buffers and / or one or more rendering targets in a manner as described above. In yet another embodiment, the PPU 200 may be configured to perform general computing operations and graphics processing operations simultaneously. In such a configuration, one or more system pipelines 230 may be configured to implement general computing operations via one or more GPCs 242, and one or more other system pipelines 230 may be configured to implement one or more graphics processing pipelines via one or more GPCs 242.
[0078] For any of the above structures, the device driver 122 and the hypervisor 124 interoperate to subdivide the various compute, graphics, and memory resources contained in the PPU 200 into separate "PPU partitions." Alternatively, there may be multiple device drivers 122, each associated with a "PPU partition." Preferably, the device drivers execute on a set of cores in the CPU 110. A given PPU partition as a whole operates in a manner substantially similar to the PPU 200. In particular, each PPU partition may be configured to perform general-purpose compute operations, graphics processing operations, or both types of operations in relative isolation from other PPU partitions. In addition, a given PPU partition may be configured to simultaneously implement multiple processing contexts while simultaneously executing one or more virtual machines (VMs) on the compute, graphics, and memory resources allocated to the given PPU partition. The following is a description of the device drivers 122 and the PPU partitions described below. Figure 5 - Figure 8 describes in more detail the logical grouping of PPU resources into PPU partitions. Figure 9-10 Techniques for partitioning and configuring PPU resources are described in more detail.
[0079] Figure 3 According to various embodiments of the present invention, Figure 2 2 is a block diagram of a GPC in a PPU of FIG. As shown, a GPC 242 is coupled to a memory management unit (MMU) 300 and includes a pipeline manager 310, a work distribution crossbar 320, one or more texture processing clusters (TPCs) 330, one or more texture units 340, a level 1.5 (L1.5) cache 350, a PM 360, and a pre-raster operation processor (preROP) 370. The pipeline manager 310 is coupled to the work distribution crossbar 320 and the TPC 330. Each TPC 330 includes one or more streaming multiprocessors (SMs) 332 and is coupled to the texture unit 340, the MMU 300, the L1.5 cache 350, the PM 360, and the preROP 370. The texture unit 340 and the L1.5 cache 350 are also coupled to the MMU 300 and to each other. The PreROP 370 is coupled to the work distribution crossbar 320. Each component shown may be implemented by any technically feasible type of hardware and / or any technically feasible combination of hardware and software.
[0080] GPC 242 is configured with a highly parallel architecture that supports the parallel execution of a large number of threads. As referred to herein, a "thread" is an instance of a specific program that is executed on a specific set of input data to perform various types of operations, including general computing operations and graphics processing operations. In one embodiment, GPC 242 may implement single instruction multiple data (SIMD) technology to support the parallel execution of a large number of threads without having to rely on multiple independent instruction units.
[0081] In another embodiment, GPC 242 may implement single instruction multiple thread (SIMT) technology to support parallel execution of a large number of generally synchronized threads via a common instruction unit that issues instructions to one or more processing engines. Those skilled in the art will appreciate that SIMT execution allows different threads to more easily follow divergent execution paths through a given program, as opposed to SIMD execution in which all threads generally follow non-divergent execution paths through a given program. Those skilled in the art will recognize that SIMD technology represents a functional subset of SIMT technology.
[0082] GPC 242 can execute a large number of parallel threads via SM 332 included in TPC 330. Each SM 332 includes a set of functional units (not shown), including one or more execution units and / or one or more load storage units, which are configured to execute instructions associated with received processing tasks. A given functional unit can execute instructions in a pipeline manner, which means that instructions can be issued to the functional unit before the execution of the previous instruction is completed. In various embodiments, the functional units within SM332 can be configured to perform a variety of different operations, including integer and floating point arithmetic (e.g., addition and multiplication, etc.), comparison operations, Boolean operations (e.g., AND, OR, and XOR, etc.), displacements, and calculations of various algebraic functions (e.g., plane interpolation and trigonometric functions, exponential functions, and logarithmic functions, etc.). Each functional unit can store intermediate data in a level 1 (L1) cache residing in SM 332.
[0083] Through the above-mentioned functional units, SM 332 is configured to process one or more "thread groups" (also called "warps") that execute the same program simultaneously on different input data. Each thread in a thread group is generally executed via a different functional unit, although in some cases not all functional units execute threads. For example, if the number of threads included in a thread group is less than the number of functional units, unused functional units may remain idle during the processing of the thread group. In other cases, multiple threads in a thread group are executed via the same functional unit at different times. For example, if the number of threads included in a thread group is greater than the number of functional units, one or more functional units may execute different threads in consecutive clock cycles.
[0084] In one embodiment, a group of related thread groups may be simultaneously active in different execution stages within an SM 332. A group of related thread groups is referred to herein as a "cooperative thread array" (CTA) or "thread array". Threads in the same CTA or threads in different CTAs may typically share intermediate data and / or output data with each other through one or more L1 caches (including those of the SM 332, the L1.5 cache 350, one or more L2 caches shared between SMs 332), or through any shared memory, global memory, or other type of memory residing on any memory device included in the computer system 100. In one embodiment, the L1.5 cache 350 may be configured to cache instructions to be executed by threads on the SM 332.
[0085] Each thread in a given thread group or CTA is typically assigned a unique thread identifier (thread ID) that can be accessed during execution. The thread ID assigned to a given thread can be defined as a one-dimensional or multi-dimensional numeric value. The execution and processing behavior of a given thread may vary depending on the thread ID. For example, a thread may determine which portion of an input data set to process and / or which portion of an output data set to write based on the thread ID.
[0086] In one embodiment, each thread instruction sequence may include at least one instruction that defines the cooperative behavior between a given thread and one or more other threads. For example, each thread instruction sequence may include an instruction that, when executed, suspends a given thread in a specific execution state until the other part or all threads reach the corresponding execution state. In another example, each thread instruction sequence may include an instruction that, when executed, causes a given thread to store data in a shared memory accessible to other parts or all threads. In another example, each thread instruction sequence may include an instruction that, when executed, causes a given thread to automatically read and update data stored in a shared memory accessible to other parts or all threads, depending on the thread IDs of those threads. In another example, each thread instruction sequence may include an instruction that, when executed, causes a given thread to calculate an address in a shared memory based on a corresponding thread ID so as to read data from the shared memory. Using the above-mentioned synchronization technology, a first thread may write data to a given location in a shared memory, and a second thread may read the data from the shared memory in a predictable manner. Thus, threads may be configured to implement a variety of data sharing patterns within a given thread group or within a given CTA or across threads of different thread groups or different CTAs. In various embodiments, a software application written in the Compute Unified Device Architecture (CUDA) programming language describes the behavior and operations of threads executing on GPC 242, including any of the behaviors and operations described above.
[0087] In operation, the pipeline manager 310 generally coordinates the parallel execution of processing tasks in the GPCs 242. The pipeline manager 310 receives processing tasks from the task / work units 234 and assigns those processing tasks to the TPCs 330 for execution by the SMs 332. A given processing task is generally associated with one or more CTAs that may be executed on one or more SMs 332 in one or more TPCs 330. In one embodiment, a given task / work unit 234 may assign one or more processing tasks to the GPCs 242 by launching one or more CTAs for one or more specific TPCs 330. The pipeline manager 310 may receive launched CTAs from the task / work units 234 and transmit the CTAs to the associated TPCs 330 for execution by one or more SMs 332 contained in the TPCs 330. During or after execution of a given processing task, each SM 332 generates output data and transmits the output data to various locations depending on the current configuration and / or the nature of the current processing task.
[0088] In configurations related to general computing or graphics processing, the SMs 332 may transmit output data to the work distribution crossbar 320, and the work distribution crossbar 320 may then route the output data to one or more GPCs 242 for additional processing or route the output data to the crossbar unit 250 for further routing. The crossbar unit 250 may route the output data to an L2 cache included in a given partition unit 262, the PPU memory 270, or the system memory 120, among other destinations. The pipeline manager 310 typically coordinates the routing of the output data performed by the work output crossbar 320 based on the processing tasks associated with the output data.
[0089] In a configuration specific to graphics processing, SM 332 may transmit output data to texture unit 340 and / or preROP 370. In some embodiments, preROP 370 may implement some or all of the raster operations specified in a 3D graphics API, in which case preROP 370 implements some or all of the operations performed by ROP 410. Texture unit 340 typically performs texture mapping operations, including, for example, determining texture sample locations, reading texture data, and filtering texture data. PreROP 370 typically performs raster-oriented operations, including, for example, organizing pixel color data and performing color blending optimizations. PreROP 370 may also perform address translation and direct output data received from SM 332 to one or more raster operations processor (ROP) units in partition unit 262.
[0090] In any of the above configurations, one or more PMs 360 monitor the performance of various components of GPC 242 to provide performance data to a user, and / or balance the utilization of computing, graphics and / or memory resources across thread groups, and / or balance the utilization of these resources with the resources of other GPCs 242. In addition, in any of the above configurations, SM 332 and other components in GPC 242 can perform memory access operations with memory interface 260 through MMU 300. MMU 300 generally writes output data to various storage spaces and / or reads input data from various storage spaces on behalf of GPC 242 and the components included therein. MMU 300 is configured to map virtual addresses to physical addresses via a set of page table entries (PTEs) and one or more optional address translation lookaside buffers (TLBs). MMU 300 can cache various data in L1.5 cache 350, including read data returned from memory interface 260. In the illustrated embodiment, MMU 300 is externally coupled to GPC 242 and may potentially be shared with other GPCs 242. In other embodiments, GPC 242 may include a dedicated instance of MMU 300 that provides access to one or more partition units 262 included in memory interface 260.
[0091] Figure 4 According to various embodiments, the Figure 2 2 is a block diagram of a partition unit 262 in a PPU 200 of FIG. As shown, the partition unit 262 includes an L2 cache 400, a frame buffer (FB) DRAM interface 410, a raster operation processor (ROP) 420, and one or more PMs 430. The L2 cache 400 is coupled between the FB DRAM interface 410, the ROP 420, and the PM 430.
[0092] The L2 cache 400 is a read / write cache that performs load and store operations received from the crossbar unit 250 and the ROP 420. The L2 cache 400 outputs read misses and urgent writeback requests to the FB DRAM interface 410 for processing. The L2 cache 400 also sends dirty updates to the FB DRAM interface 410 for opportunistic processing. In some embodiments, during operation, the PM 430 monitors the utilization of the L2 cache 400 to fairly distribute memory access bandwidth between different GPCs 242 and other components of the PPU 200. The FB DRAM interface 410 directly interfaces with a specific DRAM 272 to perform memory access operations, including writing data to the DRAM 272 and reading data from the DRAM 272. In some embodiments, the group of DRAMs 272 is divided into multiple DRAM chips, where a portion of the multiple DRAM chips corresponds to each DRAM 272.
[0093] In a configuration related to graphics processing, ROP 420 performs raster operations to generate graphics data. For example, ROP 420 can perform stencil operations, z-test operations, blending operations, and compression and / or decompression operations on z or color data, etc. ROP 420 can be configured to generate various types of graphics data, including pixel data, graphics objects, fragment data, etc. ROP 420 can also assign graphics processing tasks to other computing units. In one embodiment, each GPC 242 includes a dedicated ROP 420 that performs raster operations on behalf of the corresponding GPC 242.
[0094] Those skilled in the art will understand that Figure 1-Figure 4 The architecture described in the specification in no way limits the scope of the present embodiments, and the techniques disclosed herein may be implemented on any appropriately configured processing unit, including but not limited to, one or more CPUs, one or more multi-core CPUs, one or more PPUs 200, one or more GPCs 242, one or more GPUs or other special-purpose processing units, etc., without departing from the scope and spirit of the embodiments of the present invention.
[0095] Logical grouping of hardware resources
[0096] Figure 5 According to various embodiments Figure 2 2 is a block diagram of various PPU resources included in a PPU of FIG. As shown, PPU resources 500 include system pipelines 230 (0) to 230 (7), a control crossbar and SMC arbiter 510, a privileged register interface (PRI) hub 512, GPCs 242, a crossbar unit 250, and an L2 cache 400. The L2 cache 400 is depicted here as a collection of "L2 cache slices", each slice corresponding to a different area of DRAM 262. The system pipelines 230, GPCs 242, and PRI hub 512 are coupled together via the control crossbar and SMC arbiter 510. The GPCs 242 and the various slices of the L2 cache 400 are coupled together via the crossbar unit 250. In the example discussed herein, the PPU resources 500 include eight system pipelines 230, eight GPCs 242, and a specific number of other components. However, those skilled in the art will appreciate that the PPU resources 500 may include any technically feasible number of these components.
[0097] Each system pipeline 230 typically includes PBDMA 520 and 522, a front-end context switch (FECS) 530, a compute (COMP) front end (FE) 540, a scheduler (SKED) 550, and a CUDA work distributor (CWD) 560. PBDMA 520 and 522 are hardware memory controllers that manage communications between device drivers 122 and PPU 200. FECS 530 is a hardware unit that manages context switching. Compute FE 540 is a hardware unit that prepares processing computing tasks for execution. SKED 550 is a hardware unit that schedules processing tasks for execution. CWD 560 is a hardware unit configured to queue and dispatch one or more thread grids to one or more GPCs 242 to execute one or more processing tasks. In one embodiment, a given processing task can be specified in a CUDA program. Through the above components, the system pipeline 230 can be configured to perform and / or manage general computing operations.
[0098] The system pipeline 230(0) also includes a graphics front end (FE) unit 542 (shown as GFX FE 542), a state change controller SCC 552, and a primitive distributor stage A / stage B unit (PDA / PDB) 562. The graphics FE 542 is a hardware unit that prepares graphics processing tasks for execution. The SCC 552 is a hardware unit that manages the parallelization of work with different API states (e.g., shader programs, constants used by shaders, and how to sample textures) to maintain an orderly application of API states, even if the primitives are not processed in order. The PDA / PDB 562 is a hardware unit that distributes primitives (e.g., triangles, lines, points, quads, meshes, etc.) to the GPCs 242. With these additional components, the system pipeline 230(0) can be further configured to perform graphics processing operations. In various embodiments, some or all of the system pipeline 230 can be configured to include components similar to the system pipeline 230(0) and thus be able to perform general computing operations or graphics processing operations. Alternatively, in various other embodiments, some or all of system pipeline 230 may be configured to include components similar to system pipelines 230(1) to 220(7), and thus be capable of performing only general purpose computing operations. Figure 2 The front end 232 may be configured to include a compute FE 540, a graphics FE 542, or both a compute FE 540 and a graphics FE 542. Therefore, for general considerations, references to the front end 232 are made hereinafter with reference to one or both of the compute FE 540 and the graphics FE 542.
[0099] The control crossbar switch and SMC arbiter 510 facilitates communication between the system pipelines 230 and the GPCs 242. In certain configurations, one or more specific GPCs 242 are programmably assigned to perform processing tasks on behalf of a specific system pipeline 230. In such a configuration, the control crossbar switch and SMC arbiter 510 is configured to route data between any given GPC 242 and the corresponding system pipeline 220. The PRI hub 512 provides access to a set of privileged registers through the CPU 110 and / or PPU 200 unit to control the configuration of the PPU 200. The register address space of the PPU 200 can be configured through PRI registers, and as such, the PRI hub 212 is used to configure the mapping of PRI register addresses between a general PRI address space and a PRI address space defined separately for each system pipeline 230. This PRI address space configuration provides the ability to broadcast from the SMC engine to multiple PRI registers, as described below in conjunction with Figure 7 The GPC 242 writes data to and reads data from the L2 cache 400 via the crossbar unit 250 in the manner described above. In some configurations, each GPC 242 is assigned a separate set of L2 slices derived from the L2 cache 400, and any given GPC 242 can perform write / read operations on a corresponding set of L2 slices.
[0100] Any of the PPU resources 500 discussed above may be logically grouped or divided into one or more PPU partitions, with each partition operating in the same manner as the PPU 200 as a whole. Specifically, a given PPU partition may be configured with sufficient compute, graphics, and memory resources to perform any technically feasible operation that may be performed by the PPU 200. Examples of how the PPU resources 500 may be logically grouped into partitions are described below in conjunction with Figure 6 Detailed explanation is given.
[0101] Figure 6 According to various embodiments Figure 1 6. An example of how a hypervisor can logically group PPU resources into sets of PPU partitions is shown. As shown, a PPU partition 600 includes one or more PPU slices 610. Specifically, PPU partition 600(0) includes PPU slices 610(0) through 610(3), PPU partition 600(4) includes PPU slices 610(4) and 610(5), PPU partition 600(6) includes PPU slice 610(6), and PPU partition 600(7) includes PPU slice 610(7). In the examples discussed herein, the PPU partition 600 includes the specific number of PPU slices 610 shown. However, in other configurations, the PPU partition 600 may include other numbers of PPU slices 610.
[0102] Each PPU slice 610 includes various resources derived from one system pipeline 230, including PBDMA 520 and 522, FECS 530, front end 232, SKED 550, and CWD 560. Each PPU slice 610 also includes GPC 242, L2 slice set 620, and a corresponding portion (not shown here) of DRAM 272. The various resources included within a given PPU slice 610 impart sufficient functionality so that any given PPU slice 610 can perform at least some of the general computing and / or graphics processing operations that the PPU 200 is capable of performing.
[0103] For example, PPU slice 610 may receive processing tasks via front end 232 and then schedule those processing tasks for execution via SKED 550. CWD 560 may then issue a thread grid to execute those processing tasks on GPC 242. GPC 242 may be configured to perform the above operations in conjunction with the GPC 242. Figure 3 Multiple thread groups are executed in parallel in the manner described. PBDMA 520 and 522 can perform memory access operations on behalf of various components included in the PPU slice 610. In some embodiments, PBDMA 520 and 522 obtain commands from the memory and send the commands to FE 232 for processing. As needed, various components of the PPU slice 610 can write data to and read data from the corresponding L2 cache slice set 620. The components of the PPU slice 610 can also interface with external components included in the PPU 200, including the I / O unit 210 and / or the PCE 222, etc. as needed. When one or more VMs are time sliced on various resources included in the PPU slice 610, the FECS 530 can perform context switching operations.
[0104] In the illustrated embodiment, each PPU slice 610 includes resources derived from the system pipeline 230 that are configured to coordinate general computing operations. Thus, the PPU slice 610 is configured to perform only general processing tasks. However, in other embodiments, each PPU slice 610 may further include resources derived from the system pipeline 230 that are configured to coordinate graphics processing operations, such as the system pipeline 230(0). In these embodiments, the PPU slice 610 may be configured to additionally perform graphics processing tasks.
[0105] Typically, each PPU partition 600 is a hard partition of resources that provides a dedicated parallel computing environment isolated from other PPU partitions 600 for one or more users. A given PPU partition 600 includes one or more dedicated PPU slices 610 as shown, which collectively provide various general computing, graphics processing, and memory resources required to mimic the overall functionality of the PPU 200 as a whole, at least to some extent. Thus, a given user can perform parallel processing operations within a given PPU partition 600 in a manner similar to that of a similar user who performs those same parallel processing operations on the PPU 200 when the PPU 200 is not partitioned. Each PPU partition 600 is insensitive to failures of other PPUs 600, and each PPU partition can be reset independently of other PPU partitions 600 and without interrupting the operation of other PPU partitions 600. As described in more detail below, various resources not specifically shown here are fairly distributed across different PPU partitions 600 in proportion to the size of those different PPU partitions 600.
[0106] In the example configuration of PPU partitions 600 discussed herein, PPU partition 600(0) is assigned four of the eight PPU slices 610 and is therefore provided half of the PPU resources 500, including various types of bandwidth, such as memory bandwidth. Thus, PPU partition 610(0) will be constrained to consume half of the available system memory bandwidth, half of the available PPU memory bandwidth, half of the available PCE 212 bandwidth, etc. Similarly, PPU partition 600(4) is assigned two of the eight PPU slices 610 and is therefore provided one-quarter of the PPU resources 500. Thus, PPU partition 610(4) will be constrained to consume one-quarter of the available system memory bandwidth, one-quarter of the available PPU memory bandwidth, one-quarter of the available PCE 212 bandwidth, etc. The other PPU partitions 600(6) and 600(7) will be constrained in a similar manner. One skilled in the art will appreciate how to implement the above-described example partitions and associated resource provisioning using any other technically feasible configuration of PPU partitions 600.
[0107] In some embodiments, each PPU partition 600 is a virtual machine (VM) execution context. In one embodiment, the PPU 200 may implement various performance monitors and throttling counters that record the amount of local and / or system-wide resources being consumed by each PPU partition 600 in order to maintain proportional resource consumption across all PPU partitions 600. Allocating an appropriate portion of the PPU memory bandwidth to a PPU partition 600 may be achieved by allocating the same portion of the L2 slice 400 to the PPU partition 600.
[0108] In general, the PPU partitions 600 may be configured to operate in a functionally isolated manner relative to each other. As referred to herein, the term "functionally isolated" as applied to a PPU partition set 600 generally means that any PPU partition 600 may perform one or more operations independently of any operations performed by any other PPU partition 600 in the PPU partition set 600, without interfering with any operations performed by any other PPU partition 600 in any PPU partition set 600, and without being interfered with by any operations performed by any other PPU partition 600 in the PPU partition set 600.
[0109] A given PPU partition 600 may be configured to simultaneously execute processing tasks associated with multiple processing contexts. The term "processing context" or "context" generally refers to the state of hardware, software, and / or memory resources during the execution of one or more threads, and generally corresponds to a process on the CPU 110. The multiple processing contexts associated with a given PPU partition 600 may be different processing contexts or different instances of the same processing context. When configured in this manner, specific PPU resources assigned to a given PPU partition 600 are logically grouped into separate "SMC engines" that execute separate processing tasks associated with separate processing contexts, as described below in conjunction with Figure 7 Thus, a given processing context may include hardware settings executed in the SMC engine 700, per-thread instructions, and / or register contents associated with a thread.
[0110] Figure 7 It shows the various embodiments Figure 1 200 . An example of how a hypervisor may configure a set of PPU partitions to implement one or more simultaneous multi-context (SMC) engines. As shown, PPU partition 600 includes one or more SMC engines 700. In particular, PPU partition 600(0) includes SMC engines 700(0) and 700(2), PPU partition 600(4) includes SMC engine 700(4), PPU partition 600(6) includes SMC engine 700(6), and PPU partition 600(7) includes SMC engine 700(7). Each SMC engine 700 may be configured to execute one or more processing contexts and / or to execute one or more processing tasks associated with a given processing context, in a manner similar to PPU 200 as a whole.
[0111] A given SMC engine 700 typically includes compute and memory resources associated with at least one PPU slice 610. For example, SMC engines 700(6) and 700(7) include compute and memory resources associated with PPU slices 610(6) and 610(7), respectively. Each SMC engine 700 also includes a set of virtual engine identifiers (VEIDs) 702 that locally reference one or more subcontexts, where the VEID is associated with, and may be the same as, a virtual address space identifier used to select a virtual address space, where pages of the virtual address space are described by page tables managed by MMU1600. A given SMC engine 700 may also include compute and memory resources associated with multiple PPU slices 610. For example, SMC engine 700(0) includes compute resources associated with PPU slices 610(0) and 610(1), but does not utilize system pipeline 230(1). The SMC engine 700(0) includes and utilizes L2 slices from four PPU slices 610(0), 610(1), 610(2), and 610(3). In some embodiments, SMC engines 700 within the same PPU partition 600 share L2 slices within the PPU partition 600. In this configuration, the system pipeline 230(1) of the illustrated PPU partition 600(1) is not used because the SMC engine typically runs one processing context at a time and only one system pipeline 230 is required for one processing context. The SMC engine 700(2) is configured in a similar manner to the SMC engine 700(0). The memory resources contained in any particular PPU partition 600 may be illustrated as PPU memory partitions 710, which may be allocated to and / or distributed among any one or more SMC engines 700 in that particular PPU partition 700.
[0112] A given PPU memory partition 710 includes the set of L2 slices included in the PPU partition 600 and the corresponding portion of the DRAM 272. Typically, multiple SMC engines 700 share one PPU memory partition 710 if those SMC engines 700 are included in the same PPU partition 600. Allocations for each SMC engine 700 are provided to the contexts running on those SMC engines 700, and allocations within the PPU memory partition 710 are implemented on a page basis.
[0113] Each SMC engine 700 can be configured to independently execute processing tasks associated with one processing context at any given time. Thus, a PPU partition 600(0) having two SMC engines 700(0) and 700(2) can be configured to simultaneously execute processing tasks associated with two separate processing contexts at any given time. On the other hand, PPU partitions 600(4), 600(6), and 600(7), each of which includes one SMC engine 700(4), 700(6), and 700(7), can be configured to execute processing tasks associated with one processing context at a time. In some embodiments, contexts running on SMC engines 700 in different PPU partitions 600 can share data by sharing one or more pages in one or both PPU partitions 600.
[0114] Any given SMC engine 700 may further be configured to time slice different processing contexts at different time intervals. Thus, each SMC engine 700 may independently support the execution of processing tasks associated with multiple processing contexts, albeit not necessarily simultaneously. For example, SMC engine 700(6) may time slice four different processing contexts at four different time intervals, thereby allowing processing tasks associated with the four processing contexts to be executed within PPU partition 600(6). In some embodiments, VMs are time sliced across one or more PPU partitions 600. For example, PPU partition 600(0) may be time sliced between two VMs, each of which concurrently executes two processing contexts, one processing context per SMC engine 700(0) and 700(1). In these embodiments, preferably, all processing contexts are context switched out of a first VM before context switching in a processing context from a second VM, which is advantageous when the processing contexts running on PPU partition 600(0) share an L2 slice 400 within PPU partition 600(0).
[0115] In one embodiment, a given VM may be associated with a GPU function ID (GFID). A given GFID may include one or more bits that correspond to a physical function (PF) associated with the hardware in which the VM executes. A given GFID may also include a set of bits that correspond to a virtual function (VF) uniquely assigned to the VM. Among other uses, a given GFID may be used to route errors to a location corresponding to the guest operating system of the VM.
[0116] The SMC engines 700 within different PPU partitions 600 generally operate in isolation from each other because, as described above, each PPU partition 600 is a hard partition of the PPU resources 500. Multiple SMC engines 700 within the same PPU partition 600 can generally operate independently of each other, and in particular can context switch independently of each other. For example, the SMC engine 700(0) within the PPU partition 600(0) can context switch independently and asynchronously relative to the SMC engine 700(2). In some embodiments, multiple SMC engines 700 within the same PPU partition 600 can synchronize context switching to support certain operating modes, such as time slicing between two VMs.
[0117] generally, Figure 1 The device driver 122 and the hypervisor 124 interoperate in the manner described so far to partition the PPU 200. In addition, the device driver 122 and the hypervisor 124 interoperate to configure each PPU partition 600 as one or more SMC engines 700. In this way, the device driver 122 and the hypervisor 124 configure the DRAM 272 and / or the L2 cache 400 so as to partition the L2 slice set into groups that are each SMC memory partitions 710, as described in more detail below in conjunction with FIG. 8. In some embodiments, the hypervisor 124 is responsive to the control of the system administrator to allow the system administrator to create a configuration of the PPU partitions. These PPU partitions 600 are switched to the guest OS 916 of the VM, and the guest OS 916 then sends a request to the hypervisor 124 to configure the associated PPU partition 600 as one or more SMC engines 700. In some embodiments, because sufficient isolation is added to prevent one guest OS from affecting the PPU partition 700 of another guest OS, the guest OS can directly configure the SMC engine 700 within the PPU partition 700.
[0118] Fig. 8A According to various embodiments Figure 7 As shown, DRAM 272 is accessible through L2 slice 800, which includes Figure 7 Each L2 slice 800 corresponds to a different portion of the L2 cache 400 and is configured to access a corresponding subset of locations in the DRAM 272. In general, the partitions of the DRAM 272 correspond to a raw 2D address space organized similarly to the DRAM 272 shown herein.
[0119] As also shown, DRAM 272 is divided into a top portion 810, a partitionable portion 820, and a bottom portion 830. Top portion 810 and bottom portion 830 are storage partitions derived from the top portion and bottom portion, respectively, of all DRAMs 272(0) to 272(7). Device drivers 122, hypervisor 124, and other system-level entities can access top portion 810 and / or bottom portion 830, which in some embodiments are not accessible to PPU partitions 600. On the other hand, partitionable portion 820 is generally designated for use by general PPU partitions 600, and in particular by SMC engine 700. In some embodiments, secure data resides in top portion 810 or bottom portion 830 and is accessible to all PPU partitions 600. In some embodiments, top portion 810 or bottom portion 830 is used for hypervisor data that is not accessible to VMs.
[0120] In the exemplary memory partitions shown, the partitionable portion 820 includes a DRAM portion 822(0) corresponding to the PPU memory partition 710(0) within the PPU partition 600(0), a DRAM portion 822(4) corresponding to the PPU memory partition 710(4) within the PPU partition 600(2), a DRAM portion 822(6) corresponding to the PPU memory partition 710(6) within the PPU partition 600(6), and a DRAM portion 822(7) corresponding to the PPU memory partition 710(7) within the PPU partition 600(7). Each DRAM portion 822 corresponds to a middle portion of addresses corresponding to a set of L2 cache slices 800. A given DRAM portion 822 may be further subdivided to provide separate sets of L2 cache slices for different VMs executing processing tasks associated with different processing contexts. For example, the DRAM portion 822(4) may be subdivided into two or more regions to support two or more VMs executing processing tasks associated with two or more processing contexts. Once a DRAM portion is configured and used, it is typically used by one VM running on PPU partition 600 at a time.
[0121] In operation, the device driver 122 and the hypervisor 124 perform memory access operations within the top portion 810 and / or the bottom portion 830 via the top portion and the bottom portion of the address range corresponding to all L2 cache slices 800 in a relatively balanced manner, thereby proportionally penalizing the memory bandwidth on each L2 slice 800. In some embodiments, the SMC engine 700 performs memory access operations to the system memory 120 via the L2 cache slices 800, the throughput of which is controlled by the throttle counters 840. Each throttle counter 840 monitors the memory bandwidth consumed when the SMC engine 700 accesses the system memory 120 via the L2 cache slice 800 associated with the corresponding PPU memory partition 710, so as to provide proportional memory bandwidth to each PPU partition 600. As discussed, the PPU partitions are provided with access to various system-wide resources in proportion to the configuration of those PPU partitions 600. In the example shown, PPU partition 600(0) is allocated half of PPU resources 500, and therefore half of partitionable portion 820 (shown as DRAM portion 822(0)), and accordingly, half of the available memory bandwidth is allocated to system memory 120. Figure 19-Figure 24 The partitioning of DRAM 272 is described in more detail.
[0122] Figure 8B shows how to address according to various embodiments Figure 8B As shown, a one-dimensional (1D) system physical address (SPA) space 850 includes a top address 852 corresponding to the top portion 810, a partitionable address 854 divided into address regions 856 and corresponding to the DRAM portion 822, and a bottom address 858 corresponding to the bottom portion 830. The top address 852 is fuzzified across all L2 slices 800 (i.e., based on a pseudo-random interleaving of SPA addresses) and corresponds to the top portion of those L2 slices. The bottom address 858 is fuzzified across all L2 slices 800 and corresponds to the bottom portion of those L2 slices. Typically, only system-level entities (e.g., the hypervisor 124) and / or any entity operating through a physical function (PF) can access the top address 852 and the bottom address 858. The partitionable address 854 is assigned to the PPU partition 600. Specifically, address region 856(0) is assigned to PPU partition 600(0), address region 856(4) is assigned to PPU partition 600(4), address region 856(6) is assigned to PPU partition 600(6), and address region 856(7) is assigned to PPU partition 600(7). Address regions 856 are accessible only by one or more SMC engines 700 executing in the corresponding PPU partition 600.
[0123] Overall reference Figure 5-8B The above method of partitioning the PPU resources 500 supports a variety of usage scenarios, including single-tenant and multi-tenant usage scenarios. In a single-tenant usage scenario, the PPU 200 can be partitioned to provide independent access to the PPU resources for different users associated with a single tenant. For example, different users associated with a given tenant can perform different predetermined workloads on different PPU partitions 600. In a single-tenant usage scenario, access to the entire PPU resource 500 can be provided to a single entity. In a multi-tenant usage scenario, the PPU 200 can be partitioned to provide independent access to the PPU resources for one or more users associated with one or more different tenants. In a multi-tenant usage scenario, multiple entities can be provided with access to different PPU partitions 600, which include different parts of the PPU resources 500.
[0124] In any usage scenario, the device driver 122 and the hypervisor 124 interoperate to perform a two-step process that first involves dividing the PPU 200 into PPU partitions 600, and secondly involves configuring these PPU partitions 600 as SMC engines 700. Figure 9-10 A more detailed explanation is given.
[0125] Techniques for configuring logical groupings of hardware resources
[0126] Fig. 9 is a diagram showing various embodiments of the present invention. Figure 1 Flowchart of how the hypervisor partitions and configures the PPU. As shown, the hypervisor environment 900 includes a guest environment 910 and a host environment 920 separated from each other by a hypervisor trust boundary 930. The guest environment 910 includes a system management interface (SMI) 912, a core driver 914, and a guest operating system (OS) 916. The host environment 920 includes an SMI 922, a virtual GPU (vGPU) plug-in 924, a host OS 926, and a core driver 928. Modules included in the guest environment 910 residing above the hypervisor trust boundary 930 generally execute at a lower privilege level than modules included in the host environment 920 residing below the hypervisor trust boundary 930, including duplicate instances of the same module, such as SMI 912 and SMI 922. The hypervisor 124 executes with a kernel-level permission set and can grant appropriate permissions to any of the modules shown. In some embodiments, there is a one-to-one correspondence between VMs and guest environments 910 ; and when multiple virtual machines are not context-switched to PPU partition 600 , there is usually a one-to-one correspondence between guest environments and PPU partitions.
[0127] In operation, an administrator user of the PPU 200 interacts with the PPU 200 through the host environment 920 and the host OS 926 to configure the PPU partitions 600. In particular, the administrator user provides the partition input 904 to the SMI 922. In response, the SMI 922 issues a "create partition" command to the core driver 928, indicating the target configuration of the PPU partitions 600. The core driver 928 transmits the "create partition" command to the host interface 220 in the PPU 200 to partition the various PPU resources 500. In this way, the administrator user can initialize the PPU 200 to have a specific configuration of the PPU partitions 600. Generally, the administrator user has unrestricted access to the PPU 200. For example, the administrator user may be a system administrator of the data center where the PPU 200 is located. The administrator user may be a system administrator of the data center where multiple PPUs 200 reside. The administrator user partitions the PPU 200 in the described manner to prepare individual PPU partitions 600 to be independently configured and used by individual guest users, as described in more detail below.
[0128] A guest user of the PPU 200 interacts with a specific "guest" PPU partition 600 via a VM executing within a guest environment 910 in order to configure the SMC engine 700 within that guest PPU partition 600. Specifically, the guest user provides configuration input 902 to the SMI 912. The SMI 912 then issues a "configure partition" command to the core driver 914, indicating the target configuration of the SMC engine 700. The core driver 914 sends the "configure partition" command to the vGPU plug-in 924 across the hypervisor trust boundary 930 through the guest OS 916. The vGPU plug-in 924 issues various VM calls to the core driver 928. The core driver 928 sends the "configure partition" command to the host interface 220 in the PPU 200 to configure various resources of the guest PPU partition 600. In this way, a guest user can configure a given PPU partition 600 to have a specific configuration of the SMC engine 700. Typically, a guest user can only access a portion of the PPU resources 500 associated with a guest PPU partition 600. For example, a guest user may be a customer of a data center where a PPU 200 resides who has purchased access to a portion of a PPU 200. In one embodiment, the guest OS 916 may be configured with sufficient safeguards to allow each guest OS 916 to configure a corresponding PPU partition 600 without the involvement of the host environment 920 and / or the hypervisor 124.
[0129] Fig.10 is a flow chart of method steps for partitioning and configuring a PPU on behalf of one or more users according to various embodiments. Figures 1 to 9The method steps are described with reference to a system, but those skilled in the art will understand that any system configured to perform the method steps in any order falls within the scope of the present embodiment.
[0130] As shown, the method begins at step 1000, where the hypervisor 124 receives a partition input 904 via a host environment 920. The host environment 920 executes at an elevated privilege level, allowing an administrator user to interact directly with the PPU 200. The partition input 904 specifies a target configuration for the PPU partitions 600, including the desired size and arrangement of the PPU partitions 600. In one embodiment, the partition input may be received from an administrator user via the host environment.
[0131] At step 1004, the hypervisor 124 generates one or more PPU partitions 600 within the PPU 200 based on the partition input 904 received at step 1002. In particular, the hypervisor 124 implements the SMI 922 to issue a "create partition" command to the core driver 928. In response, the core driver 928 interacts with the host interface 220 of the PPU 200 to create one or more PPU partitions 600 having the desired configuration.
[0132] At step 1006, the hypervisor 124 allocates memory resources on the one or more PPU partitions 600 generated at step 1004. In particular, through the "create partition" command discussed above, the hypervisor 124 allocates memory resources on the one or more PPU partitions 600 generated at step 1004. Fig. 8A DRAM 272 is subdivided in the above manner to allocate different regions of DRAM 272 and corresponding L2 cache slices 800 to one or more PPU partitions 600. In one embodiment, hypervisor 124 may also configure the address mapping unit to perform partition-specific confusion operations to provide access to those different regions of DRAM 272 via L2 cache slices 800. Fig. 22 This specific embodiment is described in more detail.
[0133] At step 1008, the hypervisor 124 allocates PPU compute and / or graphics resources across one or more PPU partitions 600. In doing so, the hypervisor 124 allocates one or more system pipelines 230 and one or more GPCs 242 to one or more PPU partitions 600 via a "create partition" command. In one embodiment, the hypervisor 124 may implement steps 1006 and 1008 by logically assigning one or more PPU slices 610 to one or more PPU partitions 600, thereby allocating memory resources and compute / graphics resources together. When steps 1002, 1004, 1006, and 1008 of method 1000 are completed, the PPU 200 is partitioned, and a guest user may then configure one or more PPU partitions 600, as described below.
[0134] At step 1010, the hypervisor 124 receives configuration input 902 associated with the first PPU partition via the guest environment 910. The guest environment 910 executes at a reduced privilege level, allowing a guest user to interact only with the first PPU partition 600. The configuration input 902 specifies a target configuration for the SMC engine 700 within the first PPU partition 600, including a desired size and arrangement of the SMC engine 700. In one embodiment, the configuration input may be received from a guest user via the guest environment.
[0135] At step 1012, the hypervisor 124 generates one or more SMC engines 700 within the first PPU partition 600 via a “configure partition” command based on the configuration input 902 received at step 1010. A given SMC engine 700 may include compute and / or graphics resources derived from a system pipeline 230 and one or more GPCs 242 derived from one or more PPU slices 610. A given SMC engine 700 may access at least a portion of a PPU memory partition 710 included in the PPU partition 600 in which the SMC engine 700 resides, wherein the PPU memory partition 710 includes one or more sets of L2 cache slices and corresponding portions of the DRAM 272.
[0136] At step 1014, the hypervisor 124 distributes the memory resources allocated to the first PPU partition 600 across the one or more SMC engines 700 generated at step 1012. Generally speaking, the one or more SMC engines 700 share the first PPU memory partition 710 if the SMC engines 700 are included in the same PPU partition 600. The allocations of each SMC engine 700 are provided to the contexts running on the SMC engines 700, and the allocations in the PPU memory partition 710 are implemented on a page basis. In some embodiments, the guest OS 916 of the VM performs step 1014 by distributing memory resources to the contexts running on the SMC engine 700, which is part of the PPU partition 600 belonging to the guest environment 910. In other embodiments, the hypervisor 124 performs memory resource allocation 1014 for multiple VMs using the PPU partition 600, and each VM also distributes memory resources 1024 to the contexts running on the SMC engine 700.
[0137] At step 1016, the hypervisor 124 distributes the computing and / or graphics resources allocated to the first PPU partition 600 across the one or more SMC engines 700 generated at step 1012. Through the "configure partition" command, the hypervisor 124 allocates the system pipelines 230 contained in the guest PPU partition 600 to each SMC engine 700. The hypervisor 124 also allocates one or more GPCs 242 to each SMC engine 700. When steps 1010, 1012, 1014, and 1016 of method 1000 are completed, the first PPU partition 600 is configured, and then the guest user can initiate processing operations on one or more SMC engines 700 within the PPU partition 600.
[0138] At step 1018, the hypervisor 124 causes the first PPU partition 600 to time slice one or more VMs across one or more SMC engines 700 configured within the first PPU partition 600. The time sliced VMs may operate independently of other VMs executing within one first PPU partition 600, and operate in isolation from other VMs executing within other PPU partitions 600. In one embodiment, one or more VMs may be time sliced simultaneously across one or more SMC engines 700. In this manner, the disclosed techniques allow a partitioned PPU to support parallel execution of processing tasks associated with multiple different processing contexts.
[0139] In some embodiments, the technology disclosed herein operates in a non-virtualized system. Those skilled in the art will recognize that a single OS usage model on a PPU 200 or a group of PPUs 200 can use all of the mechanisms described in conjunction with a VM. In some embodiments, a container corresponds to a description of a VM, which means that a container on a single OS can achieve the processing isolation provided to a VM as described herein.
[0140] Partitioning computing resources to support multiple contexts simultaneously
[0141] In various embodiments, when the hypervisor 124 partitions the PPU 200 on behalf of the administrator user in the manner described above, the hypervisor 124 receives input from the administrator user indicating various boundaries between the PPU partitions 600. Based on the input, the hypervisor 124 logically groups the PPU slices 610 into PPU partitions 600, allocates various hardware resources to each PPU partition 600, and coordinates various other operations to support the simultaneous implementation of multiple processing contexts within a given PPU partition 600. The hypervisor 124 also performs additional techniques to support the migration of processing contexts between PPU partitions 600 configured on different PPUs 200. Figure 11-Figure 18 These various techniques are described in more detail.
[0142] Fig.11 An embodiment of a partition configuration table according to various embodiments is shown. Figure 1 The hypervisor 124 may configure one or more PPU partitions. As shown, the partition configuration table 1100 includes partition options 0 through 14. Partition options 0-14 are depicted above the PPU slices 610. Each of the partition options 0-14 spans a different grouping of the PPU slices 610, and in this manner represents a different possible partitioning of the PPU 600. Specifically, partition option 0 spans PPU slices 610(0) through 610(7), and thus represents a PPU partition 600 that includes all eight PPU slices 610. Similarly, partition option 1 spans PPU slices 610(0) through 610(3), and thus represents a PPU partition 600 that includes only the first four PPU slices 610, similar to the first four PPU slices 610. Figure 6-Figure 7 610(0). Partition option 2 spans PPU slices 610(4) to 610(7), and thus represents a PPU partition 600 that includes only the last four PPU slices 610. Partition options 3, 4, 5, and 6 span different groups of two adjacent PPU slices 610, while partition options 7, 8, 9, 10, 11, 12, 13, and 14 span only one corresponding PPU slice 610.
[0143] The partition configuration table 1100 also includes boundary options 1110 that represent different possible locations for partition boundaries. Specifically, boundary options 1110(1) and 1110(9) represent boundaries for partition option 0. Boundary options 1110(1) and 1110(5) represent boundaries for partition option 1, while boundary options 1110(5) and 1110(9) represent boundaries for partition option 2. Boundary options 1110(1), 1110(3), 1110(5), 1110(7), and 1110(9) represent boundaries associated with partition options 3, 4, 5, and 6. Boundary options 1110(1) through 1110(9) represent boundaries associated with partition options 7 through 14. It should be understood that one skilled in the art can create many different schemes to implement the same functionality as the configuration table 1100, which could be a set of enable bits, a list of predefined selections, or any other form that allows control over how the PPU slice 610 is divided into the PPU partitions 600. In addition, those skilled in the art will appreciate that the partition configuration table 1100 may include Fig.11 Any technically feasible figures or entries other than those shown.
[0144] During partitioning, the hypervisor 124 or device driver 122 running at the hypervisor level receives partition input from an administrator user indicating specific partitioning options according to which the PPU 200 should be partitioned. The hypervisor 124 or device driver 122 then activates a specific set of boundary options 1110 that logically isolates one or more groups of PPU slices 610 from each other to achieve the desired partitioning, as described below in conjunction with Fig.12 The examples are described in more detail.
[0145] Fig.12 It shows the various embodiments Figure 1 110(0), 600(4), 600(6), and 600(7). Figure 6-Figure 7 This exemplary configuration of PPU partition 600 is also shown in FIG.
[0146] Overall reference Figure 11-Figure 12, the hypervisor 124 or device driver 122 implements the above technique by mapping each selection of partition options to a specific binary value, which is then used to enable and disable the boundary option 1110. The binary value associated with a given partition option is referred to herein as a "swiz identifier" (swizID). The various swizIDs implemented by the hypervisor 124 are listed in Table 1 below:
[0147] Partitioning Options SwizID 0 11000000011 1 10000100011 2 11000100001 3 10000001011 4 10000101001 5 10010100001 6 11010000001 7 10000000111 8 10000001101 9 10000011001 10 10000110001 11 10001100001 12 10011000001 13 10110000001 14 11100000001
[0148] Table 1
[0149] The hypervisor 124 activates or deactivates boundary options 1110 for a given partition option based on the swizID associated with the given partition option. For example, the hypervisor 124 may activate boundary options 1110(1) and 1110(3) to configure the PPU 200 according to partition option 3 based on the corresponding swizID 10000001011. Bits 1 and 3 of this swizID activate boundary options 1110(1) and 1110(3), respectively, and bits 2 and 4-9 deactivate the remaining boundary options. Bits 0 and 10 of all swizIDs are set to one (1) to activate boundaries within the L2 cache 400, as described below in conjunction with Figure 19-20 The hypervisor 124 collects various swizIDs for different selected configuration options and computes an OR operation across all collected swizIDs to generate a configuration swizID that defines the configuration of the PPU partition 600. The configuration swizID indicates all boundary options 1110 that should be activated and deactivated to achieve the desired configuration of the PPU partition 600.
[0150] Those skilled in the art will recognize that some combinations of partition options are not feasible. For example, partition options 0 and 1 cannot be implemented in conjunction with each other because partition options 0 and 1 overlap each other. During partitioning, the hypervisor 124 automatically detects infeasible combinations of partition options and corrects these combinations by modifying one or more partition options and / or corresponding swizIDs or omitting one or more partition options and / or corresponding swizIDs.
[0151] In addition, the hypervisor 124 can dynamically detect hardware failures that render certain partitioning options infeasible. For example, assume that the PPU slice 610(0) includes a non-functional GPC 242 that was scratched by a floor during manufacturing and fused. In this case, a PPU partition 600 that includes only the PPU slice 610(0) will lack sufficient computing resources to run and will therefore be difficult to implement. In this case, the hypervisor 124 will not allow the selection of partition option 7 and / or the use of the corresponding swizID because any PPU partition 600 configured according to that partition option will be unable to perform computing operations and will therefore be unable to operate normally.
[0152] In some cases, the hypervisor 124 may allow certain configuration options that include a certain amount of non-functional hardware, as long as a PPU partition 600 configured according to such configuration options can still function to some extent. In the above example, the hypervisor 124 may allow configuration option 3 to be selected as long as the PPU slice 610(1) includes a functional GPC 242. Any PPU partition 600 configured according to configuration option 3 will still function, but it will include only half the computing resources compared to a similar PPU partition 600 that does not include any non-functional hardware.
[0153] After partitioning the PPU 200 in the manner described above, the hypervisor 124 allocates various hardware resources to the resulting PPU partitions 600. Some of these resources are statically allocated to individual PPU slices 610 and provide dedicated support for specific operations, while other resources are shared between different PPU slices 610 within the same PPU partition 600 or within different PPU partitions 600, as described below in conjunction with Fig.13 Described in more detail.
[0154] Fig.13 It shows the various embodiments Figure 1 10. The hypervisor 124 may allocate various PPU resources during partitioning. As shown, PCEs 222(0) to 222(7) are coupled to PPU slices 610(0) to 610(7). In this example, the PPU 200 includes a number of PCEs 222 equal to the number of PPU slices 610. Thus, the hypervisor 124 may statically allocate each PCE 222 to a different PPU slice 610 and configure those PCEs 222 to perform replication operations on behalf of the corresponding PPU slices 610 in a dedicated manner.
[0155] Other hardware resources included in the PPU 200 cannot be statically allocated in the manner described above because these resources may be relatively scarce. In the example shown, the PPU 200 includes only two decoders 1300 that need to be allocated across eight PPU slices 610. Therefore, the hypervisor 124 dynamically allocates the decoder 1300(0) to the PPU slices 610(0) to 610(3) included in the PPU partition 600(0). The hypervisor 124 also dynamically allocates the decoder 1300(1) to the PPU slices 610(4) and 610(5) included in the PPU partition 600(4), the PPU slice 610(6) included in the PPU partition 600(6), and the PPU slice 600(7) included in the PPU partition 600(7).
[0156] In the illustrated configuration, decoder 1300(0) is dynamically assigned to perform decoding operations in a dedicated manner for PPU partition 600(0), but decoder 1300(1) is shared among PPU partitions 600(4), 600(6), and 600(7). In various embodiments, one or more performance monitors may manage the use of hardware resources shared in the manner described to load balance resource usage among different PPU slices 610. Hypervisor 124 performs the above techniques to allocate any technically feasible resources of PPU 200 to PPU partitions 600.
[0157] When partitioning has been performed and the various resources of the PPU 200 have been statically or dynamically allocated to the various PPU slices 610, the hypervisor 124 is ready to allow the VMs to begin executing processing tasks within those PPU partitions 600. In this way, a VM can simultaneously launch multiple processing contexts within a given PPU partition 600 that are isolated from other processing contexts associated with other PPU partitions 600, as described above and as described below in conjunction with Figure 14A-14B The unused PPU slices 610 may be repartitioned into other PPU partitions 600 , while the other PPU slices are used in the active PPU partitions 600 .
[0158] Fig.14AAccording to various embodiments, multiple guest OS 916 running multiple VMs can simultaneously start multiple processing contexts within one or more PPU partitions. As shown, the guest OS 916 includes various processing contexts 1400 associated with different PPU partitions 600. Processing contexts 1400(0) and 1400(1) are associated with PPU partition 600(0) and can be started on SMC engine 700(0) or SMC engine 700(1). In some embodiments, once a processing context is assigned to an SMC engine 700, it remains on that SMC engine 700 until completion. Processing context 1400(4) is associated with PPU partition 600(4) and can be started on SMC engine 700(4). Processing contexts 1400(6) and 1400(6) are associated with PPU partitions 600(6) and 600(7), respectively, and can be started on SMC engines 700(6) and 700(7), respectively.
[0159] As previously combined Figure 7 As described above, each SMC engine 700 may perform processing tasks associated with a given processing context 1400 independently of other SMC engines 700 that perform processing tasks associated with any given processing context 1400. The processing tasks performed by a given SMC engine 700 in conjunction with a given processing context 1400 are scheduled independently of other processing tasks performed by other SMC engines 700 in conjunction with any other processing context 1400. In addition, as described below in conjunction with Figure 15-16 As described in more detail, SMC engines 700 may experience failures and / or errors independently of one another and may be reset without interrupting the operation of other SMC engines 700 .
[0160] In addition, each SMC engine 700 can be configured to perform processing tasks associated with one or more processing subcontexts 1410, which are contained in / or derived from a single parent processing context 1400. As shown, a given processing context 1400(0) includes one or more processing subcontexts 1410(0), and a given processing context 1400(1) includes one or more processing subcontexts 1410(1). The hypervisor 124 configures the processing subcontexts 1410 and corresponding device drivers. The processing subcontexts 1410 associated with a given parent processing context 1400 are started on the same SMC engine 700 that started the parent processing context 1400. Therefore, in the example shown, the processing subcontext 1410(0) is started on the SMC engine 700(0), and the processing subcontext 1410(1) is started on the SMC engine 700(1). In one embodiment, each guest OS 916 can configure its own PPU partition 600 independently of the hypervisor 124 and can not interfere with the configuration of other PPU partitions 600 .
[0161] In some embodiments that do not use virtualization, the hypervisor 124 and the guest OS 916 may not be present, and the host OS 926 may configure and start the processing context 1400 and the processing subcontext 1410 as described below in conjunction with Fig. 14B Described in more detail.
[0162] Fig. 14B 926 includes a processing context 1400 and a processing subcontext 1410. In the illustrated embodiment, the host OS 926 is configured to launch the processing context 1400 and the processing subcontext 1410 on the SMC engine 700 without involving a hypervisor or other virtualization software. The illustrated embodiment can be implemented in a "bare metal" scenario.
[0163] Overall reference Figure 14A-14B, processing tasks associated with a processing subcontext 1410 in the same parent processing context 1400 are typically not scheduled independently of one another, and typically share resources of the corresponding SMC engine 700. In addition, in some cases, a processing subcontext 1410 launched in a given SMC engine 700 may result in a failure and / or error that causes the SMC engine 700 to reset any related processing contexts 1400 and / or processing subcontexts 1410 to be restarted. A local virtual address space identifier is assigned to a processing context 1400 and / or a processing subcontext 1410, which is derived from a global virtual address space identifier 1510 associated with the PPU 200 as a whole, as described below in conjunction with Fig.15 Described in more detail.
[0164] In some embodiments, there is no virtualization, and therefore no hypervisor, but it is clear to those skilled in the art that a single OS usage model on a PPU 200 or a group of PPUs 200 can use all of the mechanisms described as belonging to a VM. In some embodiments, a container corresponds to a description of a VM, which means that a container on a single OS can obtain the processing isolation provided to a VM described herein. Examples of containers are LXC (LinuX containers) and Docker containers, which are well known in the computer industry. For example, each Docker container can correspond to a PPU partition 600, so the present invention provides isolation between multiple Docker containers running under one OS.
[0165] Fig.15 It shows the various embodiments Figure 1 15. The hypervisor of FIG. 1500 may be configured to assign virtual address space identifiers to different SMC engines. As shown, the virtual address space identifier 1500 includes a separate virtual address range for each SMC engine 700. Each virtual address range starts with zero (0) to maintain consistency between the SMC engines 700, but each virtual address range corresponds to a different portion of the global virtual address space identifier 1510. For example, the virtual address space identifiers 0-15 assigned to SMC engine 700 (0) correspond to global virtual address space identifiers 0-15, but the virtual address space identifiers 0-15 assigned to SMC engine 700 (1) correspond to global virtual address space identifiers 16-31. In one embodiment, the global set of virtual address spaces 1510 may be a virtual address space or a physical address space. In some embodiments, there is also a virtual address space identifier for each PPU partition so that the guest OS of the VM has a set of virtual address space identifiers starting at zero for all SMC engines 700 it owns.
[0166] The hypervisor 124 assigns a range of virtual address space identifiers to a given SMC engine 700 based on the number of PPU slices 610 from which the SMC engine 700 has allocated resources. In the example shown, the hypervisor 124 assigns virtual address space identifiers 0-15 to SMC engine 700(0), virtual address space identifiers 0-15 to SMC engine 700(1), and virtual address space identifiers 0-15 to SMC engine 700(4). The hypervisor 124 assigns 16 virtual address space identifiers to SMC engines 700(0), 700(1), and 700(4) because these SMC engines draw resources from two PPU slices 610, as shown in FIG. Figure 7 As shown. In contrast, the hypervisor 124 assigns virtual address space identifiers 0-7 to SMC engines 700(6) and 700(7) because these SMC engines 700 each draw resources from one PPU slice 610. The hypervisor 124 can further subdivide the virtual address space identifiers assigned to a given SMC engine 700 to support multiple processing contexts 1400. For example, the hypervisor 124 can subdivide the virtual address space identifiers 0-15 assigned to SMC engine 700(0) into two ranges 0-7 and 0-7, each of which can be assigned to a different processing context 1400. This example shows how the global virtual address space identifiers 1510 are distributed in proportion to the components 0-15, 16-31, 32-47, 48-55, and 55-63. In some embodiments, the virtual address space identifiers are unique, so the above example would have virtual space identifiers 0-15, 16-31, 32-47, 48-55, and 55-63, rather than 0-15, 0-15, 0-15, 0-7, and 0-7, as shown in FIG. Fig.15 In some embodiments, the allocation of global virtual address space identifiers 1510 is not proportional to the number of PPU slices 610 , and the hypervisor is free to allocate any subset of the global virtual address space identifiers 1510 to the PPU partitions 600 or the SMC engine 700 .
[0167] The hypervisor 124 allocates virtual address space identifiers in the manner described to allow different SMC engines 700 to perform processing tasks associated with any given processing context 1400 without remapping the virtual addresses specified by those processing tasks. Thus, the hypervisor 124 can dynamically migrate processing contexts 1400 between SMC engines 700 without making significant changes to those processing contexts. During the execution of the various processing tasks associated with a given processing context 1400, any given SMC engine 700 may sometimes encounter faults and is configured to report those faults using locally assigned virtual addresses, as described below in conjunction with Fig.16After the migration occurs, the migrated processing context still uses the same virtual address space identifiers 1500 , but these identifiers may correspond to different global virtual address space identifiers 1510 .
[0168] Fig.16 1 shows how the memory management unit converts a local virtual address space identifier 1500 to a global virtual address space identifier 1510 when mitigating faults according to various embodiments. As shown, as previously discussed, during execution, the SMC engine 700 can experience faults and / or errors and crash independently of each other. In the example shown, the SMC engine 700 (1) encounters an error and causes a local fault identifier 1610 to be output to the memory management unit (MMU) 1600. Accesses by the SMC engine 700 to unmapped pages cause the MMU to generate a fault and also cause a local fault identifier.
[0169] The MMU 1600 maintains a mapping between the local virtual address space identifier 1500 and the global virtual address space identifier 1510. Based on the mapping, the MMU 1600 generates a global fault identifier 1620 and sends the global fault identifier 1620 to the guest OS 916 (0). In response to receiving the global fault identifier 1620, the guest OS 916 (0) can reset the SMC engine 700 (1) without interrupting the operation of any other SMC engine 700, and then restart the processing context 1400 (1). In this way, each SMC engine 700 runs with a different set of virtual address space identifiers, which start at zero and span a range that may be similar, but correspond to different parts of the global memory. Therefore, the global virtual address space identifier 1510 can be divided between the SMC engines 700, but retain the appearance of a dedicated address space. In some embodiments, the fault identifier 1620 can be zero-based for the entire PPU partition 600. In other embodiments, the fault identifier 1620 can be an identifier of the SMC engine 700 and the virtual address space identifier 1500.
[0170] In one embodiment, the global fault identifier 1620 may be reported to the hypervisor 124, and the hypervisor 124 may perform various operations to resolve the associated fault. In another embodiment, some types of faults may be reported to the associated guest OS 916, while other types of faults (e.g., hard errors occurring within the top portion 810 or bottom portion 830 of the DRAM 272) may be reported to the hypervisor 124. In response to such a fault, the hypervisor 124 may reset part or all of the SMC engine 700. In various other embodiments, a given global fault identifier 1620 may be virtualized and therefore not directly correspond to a real global identifier. In operation, the MMU 1600 may route faults to the appropriate VMs based on the GFIDs associated with those VMs. In conjunction with the above Figure 7 GFID was discussed.
[0171] Overall reference Figure 15-16 , the hypervisor 124 may implement techniques similar to those described above to assign identifiers to various hardware resources associated with each PPU partition 600 and / or each SMC engine 700. For example, the hypervisor 124 may assign a local GPC identifier (GPC ID) from a range of local GPC IDs starting at zero (0) to each GPC 242 included in a given PPU partition 600. Each local GPC ID will correspond to a different global GPC ID. The method may be implemented with any PPU resource in order to maintain a set of identifiers that are internally consistent in any given PPU partition 600 and / or SMC engine 700. As described above, the method facilitates migrating processing contexts 1400 between SMC engines 700, and further allows processing contexts 1400 to be migrated between different PPUs 200.
[0172] When the hypervisor 124 migrates the processing context 1400 between different SMC engines 700 residing on different PPUs 200, the hypervisor 124 performs a technique referred to herein as "soft floor cleaning" to configure the target PPU 200 with similar hardware resources as the source PPU 200. Fig.17 Described in more detail.
[0173] Fig.17 FIG. 2 shows how, according to various embodiments, processing contexts are migrated between SMC engines on different PPUs. Figure 1. As shown, computing environment 1700(0) includes an instance of hypervisor 124(0) and PPU partition 600(0). PPU partition 600(0) is configured with SMC engine 700(0). SMC engine 700(0) performs processing tasks associated with processing context 1710. Resources 1720(0) and 1720(1) are allocated to SMC engine 700(0), but resource 1720(1) does not work. Thus, during manufacturing, resource 1720(1) is blown ("floor sweeping"). Resource 1720 can be any computing, graphics, or memory resource described so far. For example, a given resource 1720 can be a GPC 242, a TPC 330 within a GPC 242, an SM 332 within a TPC 330, a GFX FE 542, or an L2 cache slice 800, etc.
[0174] In various circumstances, hypervisor 124(0) may determine that processing context 1710 should be migrated from computing context 1700(0) to computing context 1700(1). For example, computing context 1700(0) may be scheduled for planned downtime, and in order to maintain continuous service, hypervisor 124(0) determines that processing context 1710 should be at least temporarily migrated to a different computing environment while computing context 1700(0) is unavailable.
[0175] In such a case, hypervisor 124(0) interacts with a corresponding hypervisor 124(1) executing in computing environment 1700(1) to configure PPU partition 600(1) to provide the same or similar resources as PPU partition 600(0). As shown, PPU partition 600(1) includes resources 1720(2) and 1720(3), but 1720(3) is made unavailable in order to mimic the amount of resources provided by PPU partition 600(0). In this way, processing context 1710 can be migrated from SMC engine 700(0) within PPU partition 600(0) to SMC engine 700(1) within PPU partition 600(1) without noticeable change in quality of service. This approach helps maintain the appearance that any given PPU partition 600 is operating in a manner similar to PPU 200 by providing access to a consistent set of resources while also allowing processing contexts to be migrated between different hardware. Hypervisor 124 may also implement the above method to migrate SMC engine 700 between partitions 600 within the same PPU 200. In one embodiment, hypervisors 124(0) and 124(1) may execute as a unified software entity that manages the operation of multiple PPUs 200 in different computing environments 1700.
[0176] Overall reference Figure 11-Figure 17, the hypervisor 124 implements the above-described techniques to partition PPU resources in a manner that supports the simultaneous execution of processing tasks associated with multiple processing contexts. Fig.18 These techniques are described in more detail.
[0177] Fig.18 is a flow chart of method steps for configuring computing resources within a PPU to simultaneously support operations associated with multiple processing contexts according to various embodiments. Figure 1-Figure 17 The method steps are described with reference to a system, but those skilled in the art will understand that any system configured to perform the method steps in any order falls within the scope of the present embodiment.
[0178] As shown, method 1800 begins at step 1802, wherein Figure 1 The hypervisor 124 evaluates the PPU 200 to determine a set of available hardware resources. Certain hardware resources are sometimes not manufactured correctly during the manufacture of a given PPU 200 and may be non-functional. In effect, these non-functional hardware resources have been fused and are not used. However, other hardware resources within the given PPU 200 are functional and thus the entire PPU 200 can still operate, albeit at a lower performance. Salvaging partially functional PPUs and other types of units in the manner described is known in the art as "floor sweeping."
[0179] At step 1804, the hypervisor 124 determines a set of available swizIDs based on the available hardware resources determined at step 1802. Fig.11 As described, a given swizID defines a set of hardware boundaries that can be enabled and disabled to isolate different groups of PPU slices 610 in PPU 200 to form PPU partitions 600. In the event that certain hardware resources are unavailable, hypervisor 124 determines that certain swizIDs correspond to infeasible partition configurations and should therefore be disabled.
[0180] At step 1806, the hypervisor 124 generates a set of swizIDs based on the partition input. For example, the hypervisor 124 may receive input from an administrator user indicating a set of partition options and then map those partition options to corresponding swizID sets derived from the set of available swizIDs determined at step 1804. Alternatively, the hypervisor 124 may receive the set of swizIDs directly from the administrator user and then modify any swizIDs that are not included in the set of available swizIDs.
[0181] In step 1808, the hypervisor 124 configures a set of boundaries between the hardware resources based on the swizID group generated in step 1806. In this case, the hypervisor 124 calculates a logical OR across the swizID group to generate a configuration swizID (or "local" swizID) that indicates which boundary options should be activated as boundaries and which boundary options should be disabled. Fig.12 An exemplary set of partitioning options and corresponding boundary options are described.
[0182] At step 1810, the guest OS 916 starts a set of processing contexts in the PPU partition 600 assigned to the guest user based at least in part on the one or more swizIDs. The hypervisor 124 assigns a set of virtual address space identifiers 1500 to the PPU partitions 600 corresponding to the guest user. Fig.15 The PPU partition 600 that is part of the global virtual address space identifier 1510. The hypervisor 124 or the SMC engine 700 within the PPU partition 600 can subdivide the set of virtual address space identifiers 1500 into different ranges, which are in turn assigned to different processing contexts. This approach allows each processing context to operate using a consistent set of virtual address spaces across all SMC engines 700, thereby allowing easier migration of processing contexts.
[0183] In step 1812, the hypervisor 124 or corresponding guest OS 916 resets the subset of the processing context started in step 1810 in response to one or more faults. These faults may occur at the execution unit level, the SMC engine level, or the VM level, etc. Importantly, the fault generated during the execution of the processing task associated with a processing context will not usually affect the execution of the processing task associated with other processing contexts. This fault isolation between processing contexts specifically solves the problems found in the prior art methods that rely on processing subcontexts. Optionally, between steps 1810 and 1812, a debugger can be called to control the SMC engine 700 that encounters a fault.
[0184] At step 1814, the hypervisor 124 configures a migration target based on the available hardware resources associated with the PPU 200. The migration target may be another PPU 200, but in some cases the migration target 200 is another SMC 700 within a given PPU partition 600 or another PPU partition 600 in a PPU 200. In configuring the migration target, the hypervisor 124 may perform a technique referred to herein as "soft floor sweeping" to cause the migration target to provide hardware resources similar to those used by a set of processing contexts.
[0185] At step 1816, the hypervisor 124 migrates a set of processing contexts to the migration target. The processing tasks associated with those processing contexts can continue with little or no interruption and can continue using similar available hardware resources. Thus, these techniques allow balanced quality of service to be provided in situations where processing contexts need to be moved between different PPU partitions 600 or different PPUs 200.
[0186] Overall reference Figure 11-Figure 18 , the hypervisor 124, the guest OS 916, and / or the host OS 926 perform the disclosed techniques to partition various computing resources associated with the PPU 200 into isolated and independent PPU partitions 600, in which different processing contexts can be activated simultaneously. Therefore, the resources of the PPU 200 can be more efficiently utilized compared to the traditional method of supporting one processing context at a time when the PPU may not be fully utilized. Different PPU partitions 600 can also be accessed and configured independently of each other by multiple different tenants. Therefore, the disclosed techniques provide reliable support for multi-tenancy, and thus can meet consumer demand for an effective cloud-based parallel processing platform.
[0187] Partitioning memory resources to support multiple contexts simultaneously
[0188] In addition to partitioning the compute resources associated with the PPU 200 to support multiple processing contexts simultaneously, the hypervisor 124 also partitions the memory resources associated with the PPU 200 to support multiple contexts simultaneously, thereby providing robust support for multi-tenancy. The hypervisor 124 implements various techniques when partitioning the memory resources associated with the PPU 200, which are described below in conjunction with Figure 19-Figure 24 Describe in more detail.
[0189] Fig.19 shows a set of boundary options according to various embodiments, according to which boundary options, Figure 1 The hypervisor may generate one or more PPU memory partitions. As shown, DRAM 272 includes a boundary option set 1900 that may be activated during partitioning to divide the L2 cache into various portions and partitions.
[0190] In particular, boundary options 1900(0), 1900(1), 1900(9), and 1900(10) divide DRAM 272 into Fig. 8AThe top portion 810, the partitionable portion 820, and the bottom portion 830 are formed by border option 1900(0) forming a lower border of the bottom portion 830, and border option 1900(1) forming an upper border of the bottom portion 830. Border option 1900(1) also forms a lower border of the partitionable portion 820 and a left border of the partitionable portion 820. Border option 1900(9) forms a right border of the partitionable portion 820 and an upper border of the partitionable portion 820. Border option 1900(9) also forms a lower border of the top portion 810, and border option 1900(10) forms an upper border of the top portion 810. Boundary options 1900(1) to 1900(8) further subdivide the partitionable portion 820 into various memory partitions, which will be described below in conjunction with the description of the border options. Fig. 20 Describe in more detail.
[0191] As also shown, the total size of DRAM 272 is M, the total size of top portion 810 is T, the total size of partitionable portion 820 is P, and the total size of bottom portion 830 is B. Further, the portion of a given cache tile corresponding to partitionable portion 820 is given by F, and the portion of a given cache tile corresponding to the bottom portion is given by W. F and W are configurable parameters that may be set by hypervisor 124, and in some embodiments may fully constrain the values of T, P, and B relative to M.
[0192] During configuration, the hypervisor 124 configures the DRAM 272 into the top portion 810, the partitionable portion 820, and the bottom portion 830 based on M, F, and W. Thus, the hypervisor 124 determines the values of T, P, and B based on M, F, and W. The hypervisor 124 also activates a specific boundary option 1900 based on a configuration swizID generated via interaction with an administrator user, as described above in conjunction with Figure 11-Figure 12 The following is combined Fig. 20 An exemplary activation of the boundary option 1900 is depicted.
[0193] Fig. 20 It shows the various embodiments Figure 1270 . An example of how a hypervisor partitions PPU memory to generate one or more PPU memory partitions is shown. As shown, boundary options 1900(0), 1900(1), 1900(9), and 1900(10) are activated, thereby forming a top portion 810, a partitionable portion 820, and a bottom portion 830 of DRAM 272. Boundary options 1900(1), 1900(5), 1900(7), and 1900(8) are also activated, thereby forming PPU memory partitions 710(0), 710(4), 710(6), and 710(7) corresponding to DRAM partitions 822(0), 822(4), 822(6), and 822(7) within partitionable portion 820, respectively. Boundary options 1900(2), 1900(3), 1900(4), and 1900(6) are not activated and have therefore been omitted. Hypervisor 124 configures DRAM 272 in the manner shown based on a configuration swizID equal to "11110100011".
[0194] The boundary options 1900 associated with DRAM 272 logically correspond to Figure 11-Figure 12 The boundary options 1110 shown in FIG. Figure 11-Figure 12 As discussed, each bit of a given configuration swizID indicates whether the corresponding boundary option 1110 should be activated or deactivated to group PPU slices 610 together. Fig. 20 As shown, each bit in the exemplary configuration swizID "11110100011" indicates whether the corresponding boundary option 1900 associated with the DRAM 272 should be activated or deactivated.
[0195] By default, bits 0 and 10 of the exemplary swizID are set to 1 to activate boundary options 1900(0) and 1900(10). Bits 1 and 9 of the exemplary swizID are set to 1 to activate boundary options 1900(1) and 1900(9) and establish partitionable portion 820. Bits 5, 7, and 8 of the exemplary swizID are set to 1 to activate boundary options 1900(5), 1900(7), and 1900(8) and partition partitionable portion 820 into DRAM portions 822 associated with PPU memory partition 710. The other bits of the swizID are set to zero to deactivate the corresponding boundary options. The partitions of DRAM 272 shown here correspond to Fig.12 6. Once partitioned in this manner by the hypervisor 124, the SMC engine 700 executing within the PPU partition 600 may be combined with the L2 cache slice 800 as follows: Figure 21-23 Memory access operations are performed in the described manner.
[0196] Fig.21 It shows the various embodiments Fig.16 How the memory management unit provides access to different PPU memory partitions. As shown in the figure, Fig.16 The MMU 1600 is coupled between the DRAM 272 and the 1D SPA space 850. The 1D SPA space 850 is divided into a top address 852 corresponding to the top portion 810, a partitionable address 854 corresponding to the partitionable portion 820, and a bottom address 856 corresponding to the bottom portion 830, as shown in FIG. Figure 8B During partitioning, the hypervisor 124 generates a 1D SPA space 850 based on the configuration of the DRAM 272 .
[0197] The MMU 1600 includes an address mapping unit (AMAP) 2110 configured to map the top address 852, the partitionable address 854, and the bottom address 858 to original addresses associated with the top portion 810, the partitionable portion 820, and the bottom portion 830, respectively. In this manner, the MMU 1600 services memory access requests received from the hypervisor 124 that target the top portion 810 and / or the bottom portion 830, as well as memory access requests received from the SMC engine 700 that target the partitionable portion 820, as described below in conjunction with Fig. 22 Described in more detail.
[0198] Fig. 22 It shows the various embodiments Fig.16 As shown, the partitionable address 854 includes an address region 856 (0), which includes an address corresponding to the PPU memory partition 710 (0), as described above in conjunction with Figure 8B The MMU 1600, via the AMAP 2110, converts the physical addresses included in the address region 856(0) to the original addresses associated with the DRAM portion 822(0). The AMAP 2110 is configured to reorder addresses from the address region 856(0) across the L2 cache slices 800(0) included in the PPU memory partition 710(0) to avoid situations where striding causes the same L2 cache slice 800(0) to be repeatedly accessed (also known as "camping").
[0199] In one embodiment, the AMAP 2110 may implement a "memory access" swizID that identifies the memory interleaving factor for a given memory region. A given memory access swizID determines a set of L2 cache slices that are interleaved for various types of memory accesses, including video memory, system memory, and peer memory accesses. Different PPU partitions 600 typically implement different and non-overlapping memory regions 822 within the partitionable portion 829 to minimize interference between concurrently executing jobs. The hypervisor 124 may use a memory access swizID of zero to balance memory access operations across L2 cache slices, which will typically access either the top portion 810 or the bottom portion 830.
[0200] A given memory access swizID may be a "local" swizID that is calculated based on a system physical address and used to interleave or obfuscate memory access requests across the associated L2 slices and corresponding portions of DRAM. A given local swizID associated with a given PPU partition 600 may correspond to the swizID used to configure that PPU partition. In this way, the AMAP 2110 may obfuscate addresses within the boundaries of a given PPU memory partition based on the swizID used to activate those boundaries. Obfuscated addresses based on memory access swizIDs allow the MMU 1600 to interleave the DRAM 272 so that each PPU partition 600 treats its partitionable portion 820 as contiguous in the linear system physical address space 850. This approach may maintain isolation between PPU partitions 600 and integrity associated with those PPU partitions 600.
[0201] A given memory access swizID may alternatively be a "remote" swizID provided by the device driver 122 or hypervisor 124 and used to interleave memory access requests across the L2 slices for system memory access operations. For processing operations occurring within a given PPU partition 600, the local swizID and the remote swizID may be the same. Different PPU partitions 600 typically have different remote swizIDs to allow system memory access operations to pass only through the L2 slices 800 that belong to the PPU partition 600.
[0202] MMU 1600 also provides support for translating a virtual address associated with virtual address space identifier 1500 into a system physical address in 1D system physical address space 850. For example, assuming Figure 7The SMC engine 700(0) of FIG. 160 executes using the PPU memory partition 710(0) and the corresponding DRAM portion 822(0), and in doing so causes a memory fault. The MMU 1600 issues a fault with a local fault identifier 1610. The MMU 1600 in turn translates the local fault identifier 1610 into a global fault identifier 1620. Faults and errors may be reported to the virtual function in accordance with the SR-IOV public specification.
[0203] The MMU 1600 also facilitates subdividing the address regions 856 and PPU memory partitions 710 to provide support for multiple SMC engines 700, multiple VMs, and / or multiple processing contexts 1400 executing within a given PPU partition 600, as described below in conjunction with Fig.23 Described in more detail.
[0204] Fig.23 It shows the various embodiments Fig.16 822(0). As shown, address region 856(0) contains multiple virtual memory pages 2310 of different sizes. For SMC engine 700(0), virtual memory space identifier 1500 is mapped to a global virtual memory space identifier 1510, which selects a page table for a particular virtual address space used by a processing context on SMC engine 700(0). The page specified by page table A selects page 2310(A) within DRAM portion 822(0). At the same time, SMC engine 700(2) can use the page specified by page table B, which also selects page 2310(B) within DRAM portion 822(0). Through a page-based virtual memory management scheme, pages in DRAM portion 822(0) can be assigned to different subcontexts or different processing contexts. Note that a processing context can use multiple virtual address space identifiers 1500 because it can execute many subcontexts.
[0205] Subdividing the address region 856(0) and the DRAM portion 822(0) corresponding to the PPU memory partition 710(0) in the manner shown provides dedicated memory resources within the PPU memory partition 822(0) for different SMC engines 700 executing within the corresponding PPU partition 600. Thus, multiple SMC engines 700 in different PPU 600 partitions can simultaneously execute processing tasks within different processing contexts without interfering with each other in terms of bandwidth.
[0206] The page-based approach described above can also be applied to a single SMC engine 700 executing multiple processing subcontexts, each of which requires a dedicated portion of the PPU memory partition 710(0). Likewise, the above approach can be applied to different VMs executing on one or more SMC engines 700 and requiring a dedicated portion of the PPU memory partition 710(0).
[0207] Overall reference Figure 19-Figure 23 The disclosed technology allows the above combination Figure 11-Figure 12 The described manner configures a given PPU partition 600 to safely launch multiple processing contexts simultaneously. In particular, as described, the partitioning of the L2 cache fairly allocates the DRAM portions 822 to different PPU partitions 600. In addition, the various address translations implemented via the MMU 1600 and AMAP 2110 efficiently and fairly utilize memory bandwidth, thereby providing consistent quality of service to all tenants of the PPU 200. Fig.24 Describing the combination in more detail Figure 19-Figure 23 Describe the technology.
[0208] Fig.24 is a flow chart of method steps for configuring memory resources within a PPU to simultaneously support operations associated with multiple processing contexts according to various embodiments. Figure 1-Figure 23 The method steps are described with reference to a system, but those skilled in the art will understand that any system configured to perform the method steps in any order falls within the scope of the present embodiment.
[0209] As shown, method 2400 begins at step 2402, wherein Figure 1 The hypervisor 124 determines a set of memory configuration parameters for partitioning the DRAM 272. The set of memory configuration parameters may indicate any technically feasible set of parameters that describe any attributes of the DRAM 272, including the total size (M) of the DRAM 272, the size T of the top portion 810, the size P of the partitionable portion 820, the size B of the bottom portion 830, the number of L2 cache slices 800 (e.g., 96), the size F of each cache slice portion corresponding to the partitionable portion 820, and / or the size W of each cache slice portion corresponding to the bottom portion 830. In one embodiment, the set of configuration parameters need only include the parameters M, F, and W.
[0210] At step 2404, the hypervisor 124 activates a first set of boundary options to divide the DRAM 272 into multiple portions based on the set of memory configuration parameters determined at step 2402. In particular, the hypervisor 124 activates Fig.19Boundary options 1900(0), 1900(1), 1900(9), and 1900(10) are shown to divide DRAM 272 into a top portion 810, a partitionable portion 820, and a bottom portion 820. In one embodiment, hypervisor 124 may modify the position of a given boundary option to adjust the size of the corresponding portion of DRAM 272.
[0211] At step 2406, the hypervisor 124 determines the configuration swizID based on the partition input. Fig.10 Step 1002 of the described method 1000 obtains a partition input. The partition input indicates a set of target PPU partitions of the PPU 200. The hypervisor 124 may obtain a partition input by combining the above Figure 11-Figure 12 The described technology determines the configuration swizID of the target partition set. In one embodiment, the configuration swizID can be obtained before step 2404, and then the first set of boundary options can be activated based on the configuration swizID.
[0212] In step 2408, the hypervisor activates a second set of boundary options based on the configuration swizID to generate one or more PPU memory partitions 710 within the partitionable portion 820 of DRAM 272. The second set of boundary options may include Fig.19 1900(2) through 1900(8) shown in the method 2400. These different boundary options may subdivide the partitionable portion 820 into a number of DRAM portions 822 corresponding to the PPU memory partitions 710 that are equal to or less than the number of PPU slices 610. In various embodiments, steps 2404 and 2408 of method 2400 may be performed in conjunction with one another based on a configuration swizID obtained via administrator user input or generated based on administrator user input.
[0213] At step 2410, the hypervisor 124 determines a set of partitionable addresses 854 based on a set of memory configuration parameters and / or a configuration swizID. Thus, the hypervisor 124 divides the 1D SPA space 850 into a top address 852, a partitionable address 854, and a bottom address 858, as shown in FIG. Fig.21 The top address 852 may be converted to the original address associated with the top portion 810 , the partitionable address 854 may be converted to the original address associated with the partitionable portion 820 , and the bottom address 856 may be converted to the original address associated with the bottom portion 830 .
[0214] In step 2412, the MMU 1600 services the memory access request by obfuscating the allocatable address 854 on the L2 cache slice 800 within the PPU memory partition 710 corresponding to the DRAM portion 822. The MMU 1600 obfuscates the allocatable address via the AMAP 2110 based on the memory access swizID (or "remote" swizID) associated with the PPU memory partition 710. In one embodiment, the memory access swizID is derived from the swizID according to which the PPU memory partition 710 is configured. Obfuscating the partitionable address 854 in this manner can reduce duplicate accesses to individual L2 cache slices 800 (also referred to as "camping").
[0215] At step 2414, the MMU 1600 receives a local fault identifier associated with the memory fault and converts the local fault identifier to a global fault identifier. In this way, the MMU 1600 can convert the virtual address associated with the local fault identifier into a global address associated with the global fault identifier. For example, a memory fault may result when a given SMC engine 700 encounters an error during a memory read operation or a memory write operation performed with a PPU memory partition 710. Implementing the fault ID in the local virtual address space allows the SMC engine 700 to operate with a similar address space, thereby allowing for simpler migration of processing contexts between PPU partitions 600, as described above in conjunction with Fig.17 Converting those fault IDs to global fault identifiers 1620 allows the hypervisor 124 to resolve faults corresponding to the fault IDs from the global virtual address space identifier 1510 perspective.
[0216] Overall reference Figure 19-Figure 24 The disclosed technology for partitioning PPU memory resources completes the above combination Figure 11-Figure 18 Techniques for partitioning PPU computing resources are described. Through these techniques, the hypervisor 124 can configure and partition the PPU 200 to support multiple processing contexts simultaneously. With this capability, the PPU 200 can safely perform various different processing tasks on behalf of multiple CPU processes without allowing these CPU processes to interfere with each other, thereby more efficiently utilizing PPU resources. In addition, the disclosed techniques can be applied to provide reliable support for multi-tenancy in cloud-based PPU deployments, thereby meeting product needs that have not been met by previous approaches historically.
[0217] Time slicing across multiple VMs and processing contexts
[0218] As discussed further in this article, Figure 2The PPU 200 supports two levels of partitioning. In a first partitioning level, referred to herein as "PPU partitioning," the PPU resources 500 of the PPU 200 are divided into PPU partitions 600, also referred to herein as "partitioned PPUs." In some embodiments, one or both of the PP memory 270 and the DRAM 272 may be divided into SMC memory partitions 710. With this partitioning level, each PPU partition 600 executes one VM at any given time. In a second partitioning level, referred to herein as "SMC partitioning," each PPU partition 600 is further divided into SMC engines 700. Each PPU partition 600 includes one or more SMC engines 700. In this partitioning level, each SMC engine 700 executes one processing context for one VM at any given time.
[0219] Over time, the SMC engine 700 switches from executing a particular processing context for a VM to executing a different processing context for the same VM or executing a different processing context for a different VM. Because the execution time of the SMC engine 700 is "sliced" among multiple processing contexts corresponding to one or more VMs, this process is referred to herein as "time slicing."
[0220] Each SMC engine 700 time slices between the processing contexts listed on the run list, as managed by the PBDMA 520 and 522 of the SMC engine 700. Typically, when switching between VMs, the run lists on all affected SMC engines 700 are replaced so that different groups of processing contexts are time sliced. If multiple SMC engines 700 are active, the run lists are replaced at the same time. This type of scheduling by run list replacement is referred to as "software scheduling" in this article. Switching between processing contexts within the same VM may also similarly involve replacing the run list, and is very similar to switching VMs, except that the VM does not change due to context switching. As a result, this type of context switching does not require additional hardware support. In some embodiments, a VM may have many different processing contexts of various sizes. In such an embodiment, software scheduling can consider how to pack these processing contexts into the SMC engine 700 for correct and efficient execution. In addition, software scheduling can reconfigure the number of PPU partitions 600 and the number of SMC engines within each PPU partition 600 to correctly and efficiently execute the processing context for the VM.
[0221] In order for the PPU resource 500 to support the above two partitioning levels, the PPU 200 supports two time slicing levels accordingly. Corresponding to PPU partitioning, the PPU 200 performs VM-level time slicing, where each PPU partition 600 performs time slicing between multiple virtual machines. Corresponding to SMC partitioning, the PPU 200 performs SMC-level time slicing, where each VM 600 performs time slicing between multiple processing contexts. In various embodiments, the two time slicing levels maintain the same number of TPCs in each GPC 242 over time. In various embodiments, time slicing may involve changing the number of TPCs in one or more GPCs 242 over time. The two levels of time slicing are now described.
[0222] Fig.25 is a timeline set 2500 illustrating the Figure 2 2500. As shown, the timeline set 2500 includes, but is not limited to, four PPU partition timelines 2502(0), 2502(4), 2502(6), and 2502(7). In some embodiments, the PPU partition timelines 2502(0), 2502(4), 2502(6), and 2502(7) may correspond to, respectively. Figure 6 2502(7) may correspond to PPU slices 610(0)-610(3). Similarly, PPU partition timeline 2502(6) may correspond to PPU slice 610(6), and PPU partition timeline 2502(7) may correspond to PPU slice 610(7). As a result, PPU partition 600(0) may execute up to four processing contexts simultaneously, PPU partition 600(4) may execute up to two processing contexts simultaneously, and each of PPU partitions 600(6) and 600(7) may execute one processing context at a time.
[0223] As shown in PPU partition timeline 2502(0), PPU partition 600(0) time slices between two VMs, referred to as VM A and VM B. Timeline 2502(0) shows a time slice between VM A and VM B, where a processing context associated with VM A is shown in the form of context 2510(Ax-y) and a processing context associated with VM B is shown in the form of context 2510(Bx-y). From time t0 to time t1, PPU partition 600(0) executes the processing context of VM A. A first SMC engine 700(0) of PPU partition 600(0) sequentially executes processing context 2510(A0-1), processing context 2510(A1-1), and processing context 2510(A2-1). At the same time, a second SMC engine 700(2) of PPU partition 600(0) sequentially executes processing context 2510(A3-1), processing context 2510(A4-1), and processing context 2510(A5-1). At time t1, PPU partition 600(0) stops executing processing context associated with VM A and switches to processing context associated with VM B. Processing context 2510(A2-1) and processing context 2510(A5-1) are context switched out of the corresponding SMC engines 700(0) and 700(2). PPU partition 600(0) is reconfigured from having two SMC engines 700(0) (which includes two GPCs 230(0) and 230(1)) and 700(2) (which includes two GPCs 230(2) and 230(3)) to having one SMC engine 700(0) (which includes four GPCs 230(0), 230(1), 230(2), and 230(3). Once the reconfiguration is complete, PPU partition 600(0) begins executing processing context 2510(B0-1) of VM B.
[0224] From time t1 to time t4, PPU partition 600(0) executes the processing context of VM B. The first SMC engine 700(0) of PPU partition 600(0) sequentially executes processing context 2510(B0-1), processing context 2510(B1-1), and processing context 2510(B2-1). At time t4, PPU partition 600(0) stops executing the processing context associated with VM B and switches to the processing context associated with VM A. Processing context 2510(B2-1) is context switched out of the corresponding SMC engine 700(0). The PPU partition 600(0) is reconfigured from having one SMC engine 700(0) (which includes four GPCs 230(0), 230(1), 230(2), and 230(3)) to having two SMC engines 700(0) (which includes two GPCs 230(0) and 230(1)) and 700(2) (which includes two GPCs 230(2) and 230(3)). Once the reconfiguration is complete, the PPU partition 600(0) begins executing the processing contexts 2510(A2-2) and 2510(A5-2) of VM A. Starting at time t4, the PPU partition 600(0) again executes the processing context of VM A. The first SMC engine 700(0) of the PPU partition 600(0) executes the processing context 2510(A2-2) and the processing context 2510(A0-2) in sequence. Meanwhile, the second SMC engine 700 ( 2 ) of the PPU partition 600 ( 0 ) sequentially executes the processing context 2510 ( A5 - 2 ) and the processing context 2510 ( A3 - 2 ).
[0225] As shown in PPU partition timeline 2502(4), PPU partition 600(4) time slices between two VMs, referred to as VM C and VM D. Timeline 2502(4) illustrates time slicing between VM C and VM D, where the processing context associated with VM C is shown as context 2510(Cx-y) and the processing context associated with VM D is shown as context 2510(Dx-y). From time t0 to time t2, PPU partition 600(4) is in an idle state and does not execute any processing context. From time t2 to time t5, PPU partition 600(4) executes the processing context of VM C. The first SMC engine 700(4) of PPU partition 600(4) sequentially executes processing context 2510(C0-1) and processing context 2510(C1-1). At time t5, PPU partition 600(4) stops executing processing context 2510(C1-1) and is reconfigured to begin executing the processing context of VM D. Starting at time t5, PPU partition 600(4) executes the processing context of VM D. The first SMC engine 700(4) of PPU partition 600(4) executes processing context 2510(D1-1).
[0226] As shown in PPU partition timeline 2502(6), PPU partition 600(6) time slices between two VMs, referred to as VM E and VM F. Timeline 2502(6) shows time slicing between VM E and VM F, where the processing context associated with VM E is shown in the form of context 2510(Ex-y) and the processing context associated with VM F is shown in the form of context 2510(Fx-y). From time t0 to time t3, PPU partition 600(6) is in an idle state and does not execute any processing context. From time t3 to time t6, PPU partition 600(6) executes the processing context of VM E. The first SMC engine 700(6) of PPU partition 600(6) executes processing context 2510(E0-1) and processing context 2510(E1-1) in sequence. At time t6, PPU partition 600(6) stops executing processing context 2510(E1-1) and is reconfigured to begin executing the processing context of VM F. Starting at time t6, PPU partition 600(4) executes the processing context of VM F. The first SMC engine 700(6) of PPU partition 600(6) executes processing context 2510(F0-1).
[0227] As shown in PPU partition timeline 2502(7), PPU partition 600(7) is time sliced within a single VM, referred to as VMG. Timeline 2502(6) shows the time slice of VM G, where the processing context associated with VM G is shown in the form of context 2510(Gx-y). From time t0 to time t5, PPU partition 600(7) is in an idle state and does not execute any processing context. Starting at time t5, PPU partition 600(7) executes processing context G. The first SMC engine 700(7) of PPU partition 600(7) executes processing context 2510(G0-1).
[0228] In this manner, each of the PPU partitions 600(0), 600(4), 600(6), and 600(7) time slices between processing contexts corresponding to one or more VMs. Each of the PPU partitions 600(0), 600(4), 600(6), and 600(7) transitions from one processing context to another processing context independently of one another. For example, PPU partition 600(0) may switch from executing one processing context for a particular VM to another processing context for the same or a different VM, regardless of whether any one or more of the PPU partitions 600(4), 600(6), and 600(7) switches processing contexts. During the time period shown, each of the PPU partitions 600(0), 600(4), 600(6), and 600(7) maintains a constant number of PPU slices 610. In some embodiments, the number of PPU slices 610 for each PPU partition 600 may vary, as now described.
[0229] Fig.26 is another set of timelines 2600 showing the Figure 2 The functionality of the processing context shown in the timeline set 2600 is similar to that of the VM-level time slice associated with the PPU 200. Fig.25 2500, except as further described below. As shown, timeline set 2600 includes, but is not limited to, four PPU partition timelines 2602(0), 2602(4), 2602(6), and 2602(7). In some embodiments, PPU partition timelines 2602(0), 2602(4), 2602(6), and 2602(7) may correspond to, respectively. Figure 6 PPU partitions 600(0), 600(4), 600(6) and 600(7).
[0230] As shown in PPU partition timeline 2602(0), PPU partition 600(0) time slices between two VMs, referred to as VM A and VM B. Timeline 2602(0) illustrates time slicing between VM A and VM B, where a processing context associated with VM A is shown as context 2610(Ax-y) and a processing context associated with VM B is shown as context 2610(Bx-y). From time t0 to time t1, PPU partition 600(0) executes a processing context of VM B. A first SMC engine 700(0) of PPU partition 600(0) sequentially executes processing context 2610(B1-1) and processing context 2610(B2-1). At time t1, PPU partition 600(0) stops executing a processing context associated with VM B and switches to a processing context associated with VM A. Processing context 2610(B2-1) is context switched out of the corresponding SMC engine 700(0). The PPU partition 600(0) is reconfigured from having one SMC engine 700(0) (which includes four GPCs 230(0), 230(1), 230(2), and 230(3)) to having two SMC engines 700(0) (which includes two GPCs 230(0) and 230(1)) and 700(2) (which includes two GPCs 230(2) and 230(3)). Once the reconfiguration is complete, the PPU partition 600(0) begins executing the processing context of VM A. From time t1 to time t3, the PPU partition 600(0) executes the processing context of A. The first SMC engine 700(0) of the PPU partition 600(0) sequentially executes processing context 2610(A2-1), processing context 2610(A0-1), and processing context 2610(A1-1). At the same time, the second SMC engine 700(2) of the PPU partition 600(0) sequentially executes the processing context 2610(A5-1), the processing context 2610(A3-1), and the processing context 2610(A4-1). At time t3, the PPU partition 600(0) stops executing the processing context associated with VM A and switches to the processing context associated with VM B. The processing context 2610(A1-1) and the processing context 2610(A4-1) are context switched out of the corresponding SMC engines 700(0) and 700(2). PPU partition 600(0) is reconfigured from having two SMC engines 700(0) (which includes two GPCs 230(0) and 230(1)) and 700(2) (which includes two GPCs 230(2) and 230(3)) to having one SMC engine 700(0) (which includes four GPCs 230(0), 230(1), 230(2), and 230(3)).Once the reconfiguration is complete, PPU partition 600(0) begins executing the processing context of VM B. Starting from time t3, PPU partition 600(0) again executes the processing context of VM B. The first SMC engine 700(0) of PPU partition 600(0) sequentially executes processing context 2 610(B2-2) and processing context 2 610(B0-1).
[0231] As shown in the PPU partition timeline 2602(4), from time t0 to time t2, the first SMC engine 700(4) of the PPU partition 600(4) sequentially executes the processing context 2610(D0-1) and the processing context 2610(D1-1). Then, the first SMC engine 700(4) of the PPU partition 600(4) is idle. As shown in the PPU partition timeline 2602(6), from time t0 to time t2, the first SMC engine 700(6) of the PPU partition 600(6) sequentially executes the processing context 2610(F0-1) and the processing context 2610(F1-1). Then, the first SMC engine 700(6) of the PPU partition 600(6) is idle. As shown in the PPU partition timeline 2602(7), the PPU partition 600(7) is idle from time t0 to time t2.
[0232] At time t2, PPU partitions 600(4), 600(6), and 600(7) are merged to form a single PPU partition 600(4) having four SMC engines 700(4)-700(7). As shown in PPU partition timeline 2604(4), the merged PPU partition 600(4) executes the processing context of VM H. Starting from time t2, the first SMC engine 700(4) of PPU partition 600(4) sequentially executes processing context 2610(H0-1), processing context 2610(H1-1), processing context 2610(H2-1), and processing context 2610(H0-2).
[0233] In this manner, the PPU partitions 600 may be merged and / or split into partitions of different sizes during time slicing. The PPU partitions 600 may be merged and / or split independently of each other. Fig.25 and 26 , each VM executes on a fixed number of SMC engines 700, resulting in a given VM executing a constant number of processing contexts simultaneously. As shown, VM A executes on two SMC engines 700 simultaneously, while the remaining VMs execute on one SMC engine 700 at a time. In some embodiments, a particular VM may vary the number of SMC engines 700 on which the VM executes, as now described.
[0234] Fig. 27is a timeline 2700 showing the Figure 2 The processing contexts shown in timeline 2700 are respectively associated with the SMC-level time slices of the PPU 200. Fig.25 and 26 The timeline sets 2500 and 2600 are substantially the same, except as further described below. As shown, timeline 2700 represents a single PPU partition timeline. In some embodiments, the PPU partition timeline represented by timeline 2700 may correspond to Figure 6 Any PPU partition 600 includes at least two SMC engines 700.
[0235] As shown in timeline 2700, a time slice of a PPU partition 600 within a single VM, referred to as VM A. Timeline 2700 illustrates a time slice of VM A, wherein the processing context associated with VM A is shown in the form of context 2710 (Ax-y). Starting at time t0, VM A is executed on two SMC engines 700. A first SMC engine 700 included in the PPU partition 600 executes processing context 2710 (A0-1) and then idles until time t1. At the same time, a second SMC engine 700 included in the PPU partition 600 executes processing context 2710 (A1-1) and then idles until time t1. The duration between time t0 and time t1 is long enough to ensure that there is enough time for processing contexts 2710 (A0-1) and 2710 (A1-1) to complete execution and for the SMC engines to enter an idle state. In some embodiments, processing context 2710 (A0-1) and processing context 2710 (A1-1) can simultaneously perform the same task on two separate SMC engines 700, thereby providing spatial redundancy. In such an embodiment, processing context 2710 (A0-1) and processing context 2710 (A1-1) can perform tasks on redundant SMC engines 700 having the same configuration as each other, and then compare the accuracy and certainty of the results.
[0236] Between time t1 and time t2, the run lists for processing contexts 2710(A0-1) and 2710(A1-1) are removed from the PPU partition 600. The PPU partition 600 is reconfigured from two SMC engines 700 to one SMC engine 700, which includes all resources of the two SMC engines 700. The PPU partition 600 then executes processing contexts 2710(B2-1), 2710(B3-1), and 2710(B4-1) using the new run lists.
[0237] Between time t2 and time t3, VM B executes on one SMC engine 700. The SMC engine 700 contained in the PPU partition 600 sequentially executes the processing context 2710 (B2-1), the processing context 2710 (B3-1), and the processing context 2710 (B4-1). In some embodiments, the processing contexts 2710 (B2-1), 2710 (B3-1), and 2710 (B4-1) can perform performance-intensive tasks, which can benefit from being executed on a single SMC engine 700 that has more computing resources than the SMC engine executing the processing contexts 2710 (A0-1) and 2710 (A1-1). The SMC engine 700 then idles until time t3. In some embodiments, the SMC engine 700 performs offline scheduling tasks during this idle period.
[0238] Between time t3 and time t4, the run lists for processing contexts 2710(B2-1), 2710(B3-1), and 2710(B4-1) are removed from the PPU partition 600. The PPU partition 600 is reconfigured from one SMC engine 700 to two SMC engines 700. The two SMC engines 700 each include a portion of the resources included in one SMC engine 700. The PPU partition 600 then uses the new run lists to execute processing contexts 2710(A0-2) and 2710(A1-2).
[0239] Starting at time t4, VM A is again executed on two SMC engines 700. The first SMC engine 700 executes processing context 2710 (A0-2) and then idles until time t5. At the same time, the second SMC engine 700 executes processing context 2710 (A1-2) and then idles until time t5. The duration between time t4 and time t5 is long enough to ensure that there is enough time for processing contexts 2710 (A0-2) and 2710 (A1-2) to complete execution and put the SMC engines into an idle state. In some embodiments, processing context 2710 (A0-2) and processing context 2710 (A1-2) can execute the same task simultaneously on two separate SMC engines 700, thereby providing spatial redundancy. In such an embodiment, processing context 2710 (A0-2) and processing context 2710 (A1-2) can execute tasks on redundant SMC engines 700 having the same configuration as each other, and then compare the accuracy and certainty of the results.
[0240] Between time t5 and time t6, the run lists for process contexts 2710(A0-2) and 2710(A1-2) are removed from PPU partition 600. PPU partition 600 is reconfigured from two SMC engines 700 to one SMC engine 700, which includes all resources of both SMC engines 700. PPU partition 600 then uses the new run list to execute process context 2710(B3-2). Starting at time t6, VM B is again executed on one SMC engine 700. SMC engine 700 executes process context 2710(B3-2).
[0241] In some embodiments, the PPU partition 600 can be quickly reconfigured between executing on one SMC engine and executing on two SMC engines, a process referred to herein as "fast reconfiguration". Fast reconfiguration increases the utilization of PPU 200 resources while providing a mechanism for multiple processing contexts to execute in different modes on a single PPU partition 600. One or both of the core driver 914 and the hardware microcode within the PPU 200 include various optimizations to implement fast reconfiguration. These optimizations are now described.
[0242] During reconfiguration, certain resources in the PPU 200, such as the FECS 530 and GPC 242 environment contexts, are not reset unless the resources generate errors. As a result, loading microcode into these resources during reconfiguration can be divided into multiple stages. In particular, the microcode loading sequence for the FECS 530 and GPC 242 can be divided into a LOAD stage and an INIT stage. For all available FECS 530 and GPC 242 environment contexts within the PPU partition 600, the LOAD stage is executed in parallel, thereby reducing the time required to load microcode into these resources. The INIT stage is executed during reconfiguration, thereby performing the initialization of FECS 530 and GPC242 context switches in parallel with reconfiguring the PPU partition 600. During the initialization stage, the PPU 200 ensures that the LOAD stage of all FECS 530 and GPC 242 environment contexts has been completed. As a result, the time to load and initialize the resources of the PPU partition 600 is reduced. In addition, PPU 200 stores a cache of standardized processing context images, referred to herein as “golden processing context images,” for each possible configuration of PPU partition 600. The appropriate golden processing context image is retrieved and loaded during the LOAD and INIT phases, further reducing the time to load and initialize the resources of PPU partition 600.
[0243] As described herein, a particular VM may change the number of SMC engines 700 on which the VM executes over time. In one specific example, VM A includes various tasks associated with an autonomous vehicle. Some tasks of an autonomous vehicle are more critical than other tasks. For example, tasks associated with autonomous driving (such as detecting traffic lights and avoiding collisions) will be considered more critical than tasks associated with the vehicle's entertainment system. These more critical tasks may need to comply with certain regulations or industry standards. One such standard assigns a classification level known as the Automotive Safety Integrity Level (ASIL). In order to increase the integrity level, ASIL includes four levels, namely ASIL-A, ASIL-B, ASIL-C, and ASIL-D. Tasks such as detecting traffic lights and avoiding collisions will be classified as ASIL-D. Less important tasks can be classified as lower ASIL levels. Tasks not related to safety (such as tasks associated with the vehicle's entertainment system) can be classified as QM, which indicates that only standard quality management specifications are applicable.
[0244] In this regard, processing contexts 2710(A0-1) and 2710(A1-1) may include two instances of the same ASIL-D level task executed simultaneously on two different SMC engines 700 of the PPU partition 600. After processing contexts 2710(A0-1) and 2710(A1-1) complete execution, the results of processing contexts 2710(A0-1) and 2710(A1-1) are compared with each other. If processing context 2710(A0-1) and processing context 2710(A1-1) generate the same result, the result has been verified and the vehicle will continue to drive based on the result. On the other hand, a fault in one or more components associated with processing context 2710(A0-1) or processing context 2710(A1-1) may cause the affected processing context to generate incorrect results. Therefore, if processing context 2710 (A0-1) and processing context 2710 (A1-1) generate different results, the results are invalid and the vehicle will perform appropriate evasive actions, such as slowly moving away from the traffic flow to the closest location.
[0245] After completing the execution of processing contexts 2710 (A0-1) and 2710 (A1-1), the PPU partition 600 is reconfigured to include only one SMC engine 700. The SMC engine 700 executes QM-level processing contexts 2710 (B2-1), 2710 (B3-1), and 2710 (B4-1) in sequence. These processing contexts include less critical tasks, such as tasks associated with the vehicle's entertainment system. After completing the processing contexts 2710 (B2-1), 2710 (B3-1), and 2710 (B4-1), the PPU partition 600 is reconfigured to include two SMC engines 700. The SMC engine 700 simultaneously executes processing contexts 2710 (A0-2) and 2710 (A1-2), which are two instances of the same ASIL-D level task. After the processing contexts 2710 ( A0 - 2 ) and 2710 ( A1 - 2 ) complete execution, the PPU partition 600 is reconfigured again to include only one SMC engine 700 , and executes the QM level processing context 2710 ( B3 - 2 ).
[0246] In this manner, the PPU partition 600 is dynamically reconfigured between multiple SMC engines 700 executing ASIL-D level tasks and a single SMC engine 700 executing QM level tasks. The duration between consecutive ASIL-D processing contexts (e.g., the duration between time t0 and time t4) is referred to as a "frame," where a portion of the frame between time t0 and time t1 is allocated for executing ASIL-D tasks.
[0247] Fig.28 2800 , PPU 200 (1) executes four VMs 2810A, 2810B, 2810C, and 2810D. Each of these VMs 2810A, 2810B, 2810C, and 2810D executes on a different SMC engine 700 included in PPU 200 (1). Similarly, PPU 200 (2) executes four VMs 2810E, 2810F, 2810G, and 2810H. Each of these VMs 2810E, 2810F, 2810G, and 2810H executes on a different SMC engine 700 included in PPU 200 (1).
[0248] Over time, VMs may be migrated from one PPU 200 to another PPU 200 for various reasons, including but not limited to preparing for system maintenance, consolidating VMs onto fewer PPUs 200 to improve utilization or save power, and gaining efficiency by migrating to different data centers. In a first example, a VM may be forced to migrate to a different PPU 200 when the system on which the VM is currently executing is to be shut down for system maintenance. In a second example, the processing context of one or more VMs may be idle for an indeterminate amount of time. If all processing contexts in one or more VMs are idle, the VM may be migrated from one PPU 200 to another PPU 200 to improve utilization or reduce energy consumption of the PPU 200. In a third example, VMs associated with a particular user or user group may be migrated from a geographically distant data center to a closer data center to improve communication latency. More generally, VMs may be migrated to different PPUs 200.
[0249] More generally, a VM may be migrated from one PPU 200 to another PPU 200 at any time when the context has been removed from the hardware by context saving. The associated operating system, e.g. Fig. 9 The guest operating system 916 can preempt a context and force the context to be saved at any time, not just when the corresponding VM is idle. A context can be preempted when all work for the corresponding VM has been completed. In addition, a context can be preempted by forcing the context to stop submitting further work and draining the current work in progress, even if the VM has other work to perform. In either case, once the context's work in progress has been drained and the context has been saved, the VM can be migrated from one PPU 200 to another PPU 200. During VM migration, the VM may suspend execution for a period of several milliseconds.
[0250] In some embodiments, a VM may be migrated only to a PPU partition 600 in another PPU 200 that has the same configuration as the PPU partition 600 that is currently executing the VM. For example, a VM may be restricted to migrating only to a PPU partition 600 in another PPU 200 that has the same number of GPCs 242 as the PPU partition 600 that is currently executing the VM. As shown in diagram 2802, four VMs are in an idle state. PPU 200(1) executes two VMs 2810A and 2810C. The other two VMs 2810B and 2810D that were previously executed on PPU 200(1) are in an idle state. Similarly, PPU 200(2) executes two VMs 2810E and 2810H. The other two VMs 2810F and 2810G that were previously executed on PPU 200(1) are in an idle state. As a result, each of PPUs 200(1) and 200(2) is not fully utilized. In this case, the currently executing VMs may be migrated to better utilize the available PPU resources. In one example, the VMs executing on PPU 200(1) may consume half of the hardware resources available on PPU 200(1). Similarly, the VMs executing on PPU 200(2) may consume half of the hardware resources available on PPU 200(2). As a result, each of PPU 200(1) and PPU 200(2) will operate at approximately 50% capacity. If all VMs executing on PPU 200(2) are migrated to PPU 200(1), PPU 200(1) will operate at approximately 100% capacity. PPU 200(2) will operate at 0% of capacity. As a result, in order to reduce power consumption, the power supply voltage to PPU 200(2) can be reduced. As shown in diagram 2804, VMs 2810E and 2810H have been migrated from PPU 200(2) to PPU 200(1). As a result, PPU 200(1) executes four VMs 2810A, 2810E, 2810C, and 2810H. Therefore, PPU 200(1) is more fully utilized. After the VM migration, PPU 200(2) no longer executes any VM. As a result, PPU 200(2) can be powered off to reduce power consumption. If another VM subsequently begins execution, PPU 200(2) can be powered on to execute the other VM.
[0251] Fig.29 is a timeline set 2900 showing the Figure 2 The functions of the timeline set 2900 are respectively related to Fig.25 and 26 Timeline 2500 and 2600 and Fig. 272700, except as further described below. As shown, the timeline set 2900 includes, but is not limited to, four PPU partition timelines 2902(0), 2902(1), 2902(2), and 2902(3). In some embodiments, the PPU partition timelines 2902(0), 2902(1), 2902(2), and 2902(3) may correspond to Figure 6 2902(0), 2902(1), 2902(2), and 2902(3), each PPU partition 600 is time sliced between five VMs, referred to as VM A through VM F. As a result, each of the five VMs is migrated between the four PPU partitions 600.
[0252] exist Fig.29 During the time period shown, VM A executes processing context 2910 (A0) on the first PPU partition, as shown in PPU partition timeline 2902 (3). VM A then migrates to the second PPU partition and executes processing context 2910 (A1) on PPU partition timeline 2902 (2). Subsequently, VM A migrates to the third PPU partition and the fourth PPU partition in sequence and executes processing context 2910 (A2) and 2910 (A3) on PPU partition timelines 2902 (1) and 2902 (0), respectively. VM A then migrates back to the first PPU partition and executes processing context 2910 (A4) on PPU partition timeline 2902 (3). Finally, VM A migrates again to the second PPU partition and executes processing context 2910 (A5) on PPU partition timeline 2902 (2).
[0253] In a similar manner, VM B, which executes processing contexts 2910 (B0) to 2910 (B4), executes on the first PPU partition 600 and then migrates between the other three PPU partitions, as shown in PPU partition timelines 2902 (0), 2902 (1), 2902 (2), and 2902 (3). The remaining three VMs are similarly migrated between the four PPU partitions, with VM C executing processing contexts 2910 (C0) to 2910 (C5), VM D executing processing contexts 2910 (D0) to 2910 (D5), and VM E executing processing contexts 2910 (E0) to 2910 (E5).
[0254] In this manner, five VMs are migrated between four PPU partitions 600, where each VM accesses substantially the same number of PPU resources. Overall, the five VMs are each able to execute approximately 80% of the time, where four PPU partitions 600 divided by five VMs equals 4 / 5 or 80%. As a result, fine-grained VM migration performs load balancing between a group of VMs regardless of the number of VMs relative to the number of PPU partitions 600.
[0255] Figure 30A-Figure 30B A method for performing Figure 2 Flowchart of method steps for time slicing a VM in a PPU 200. Figure 1-Figure 15 The method steps are described herein as a system configured to perform the method steps in any order, but one of ordinary skill in the art will understand that any system configured to perform the method steps in any order is within the scope of the present disclosure.
[0256] As shown, method 3000 begins at step 3002, where PPU 200 determines that at least one VM is about to switch from a first set of one or more processing contexts to a second set of one or more processing contexts. The VM may switch one or more processing contexts for any technically feasible reason, including but not limited to the VM completing execution of all tasks, the VM entering an idle state, the VM has executed a maximum allocated time, or the VM has generated an error.
[0257] At step 3004, the PPU 200 determines whether the PPU 200 needs to perform a PPU internal partition change to accommodate the new one or more processing contexts. A PPU internal partition change occurs when the PPU partition 600 maintains the same number of PPU slices 610 during a context switch but changes the number of active SMC engines 700 within the PPU partition 600. If the PPU 200 does not need to perform a PPU internal partition change to accommodate the new one or more processing contexts, the method 3000 proceeds to step 3008. However, if the PPU 200 needs to perform a PPU internal partition change to accommodate the new one or more processing contexts, the method 3000 proceeds to step 3006, where the PPU 200 reconfigures the PPU partition 600 to maintain the same number of PPU slices 610 while changing the number of active SMC engines 700.
[0258] At step 3008, the PPU 200 determines whether the PPU 200 needs to perform an inter-PPU partition change to accommodate the new one or more processing contexts. An inter-PPU partition change occurs when a PPU partition 600 changes the number of PPU slices 610 by merging or splitting one or more PPU partitions 600 during a context switch. As a result, the PPU 200 also changes the number of intra-PPU partitions 600. Depending on the new processing contexts, the PPU 200 may or may not change the number of SMC engines 700 active in each PPU partition 600. If the PPU 200 does not need to perform an inter-PPU partition change to accommodate the new one or more processing contexts, the method 3000 proceeds to step 3012. However, if the PPU 200 needs to perform an inter-PPU partition change to accommodate the new one or more processing contexts, the method 3000 proceeds to step 3010, where the PPU 200 reconfigures the PPU partition 600 to change the number of PPU slices 610 included in the PPU partition 600. To change the number of PPU slices 610 in the PPU partition 600, the PPU 200 merges two or more PPU partitions 600 into a single PPU partition 600. Additionally or alternatively, the PPU 200 may split the PPU partition 600 into two or more PPU partitions 600.
[0259] At step 3012, the PPU 200 determines whether all VMs executing on the given PPU 200 are idle. If one or more VMs on the given PPU 200 are not idle (active), the method proceeds to step 3018. On the other hand, if all VMs executing on the given PPU 200 are idle, the method 3000 proceeds to step 3014, where the PPU 200 determines whether one or more other PPUs 200 have resources available to execute the idle VMs. If the resources are not available on the one or more other PPUs 200, the method proceeds to step 3018. On the other hand, if the resources are not available on the one or more other PPUs 200, the method proceeds to step 3016, where the PPU 200 migrates the idle VMs to one or more other PPUs 200.
[0260] At step 3018, after performing the intra-PPU partition change, the inter-PPU partition change, and / or the VM migration, the PPU 200 begins executing the new processing context. The method 3000 then terminates. In various embodiments, the PPU 200 determines the need to perform an intra-PPU change independently of determining the need to perform an inter-PPU change. Similarly, in various embodiments, the PPU 200 determines the need to perform an inter-PPU change independently of determining the need to perform an intra-PPU change.
[0261] Privileged Register Address Map
[0262] As further described herein, Figure 5 The PRI hub 512 and internal PRI bus (not shown) enable the CPU 110 and / or any unit in the PPU 200 to read and write privileged registers, also referred to as "PRI registers": which are distributed throughout the PPU 200. As such, the PRI hub 212 is configured to map PRI bus addresses between a common address space covering all PRI bus registers and an address space defined separately for each system pipe 230. When communicating over a PCIe link (typically from the CPU), the PRI registers are accessed through a PCIe address space, referred to herein as the "base address register 0" space, or more simply, the "BAR0" address space, which is typically used for devices connected to the PCIe bus. Typically, the addressable memory range of the BAR0 address space of the PPU 200 is limited to 16 megabytes (MB) due to the large number of devices that must all fit in the BAR0 address space. The 16MB address range is sufficient to access the privileged registers of a single SMC engine 700. However, to support multiple SMC engines 700, the address range may exceed 16MB. Therefore, the PRI hub 512 provides two addressing modes to support execution with multiple SMC engines 300. The first addressing mode, referred to herein as the "legacy mode," is applicable to operations involving a single SMC engine 700. The second addressing mode, referred to herein as the "SMC engine addressing mode," is applicable to operations involving multiple SMC engines 700. The addressing modes are now described.
[0263] Fig.31 is a memory map according to various embodiments, which shows how the BAR0 address space 3110 is mapped to Figure 2The BAR0 address space 3110 includes, but is not limited to, a first address space 3112, a graphics register (GFX REG) address space 3114, and a second address space 3116. The privileged register address space 3120 includes, but is not limited to, a first address space 3122, a legacy graphics register address space 3124, a second address space 3126, and an SMC graphics register address space 3128(0)-3128(7). The BAR0 address space 3110 supports two addressing modes, a legacy addressing mode and an SMC addressing mode. In general, the two modes are targeted to be: (1) a legacy mode, in which the entire PPU 200 is treated as one engine with a set of PRI registers; and (2) an SMC mode, in which each PPU partition 600 is addressed as if each PPU partition 600 is a separate engine, and as if each PPU partition 600 itself is the entire PPU 200. SMC mode allows the driver software 122 to be the same when processing the entire PPU in legacy mode and when processing only on one PPU partition 600. That is, the driver can be written once and used in both legacy mode and SMC mode scenarios.
[0264] In the legacy addressing mode, the PPU 200 executes tasks as a single cluster of hardware resources, rather than as separate PPU partitions 600 with separate SMC engines 700. In the legacy mode, a memory read or write to a memory address is directed to the first address space 3112 or the second address space 3116 of the BAR0 address space 3110, accessing a corresponding memory address in the first address space 3122 or the second address space 3126 of the privileged register address space 3120, respectively. Similarly, a memory read or write to a memory address is directed to the graphics register address space 3114 of the BAR0 address space 3110, accessing a corresponding memory address in the legacy graphics register address space 3124 of the privileged register address space 3120. The legacy graphics register address space 3124 includes address ranges for various components within the PPU 200, including, but not limited to, the compute FE 540, the graphics FE 542, the SKED 550, the CWD 560, and the PDA / PDB 562. Furthermore, legacy graphics register address space 3124 includes an address range for each GPC 242. A GPC 242 is individually addressable via a dedicated address range within legacy graphics register address space 3124, minus any GPC 242 that is removed due to floor cleaning. Additionally or alternatively, legacy graphics register address space 3124 includes an address range for broadcasting data to all GPCs 242 simultaneously. These GPC broadcast address spaces may be useful when all GPCs 242 are configured identically.
[0265] In SMC addressing mode, the PPU 200 performs tasks as a separate PPU partition 600 with a separate SMC engine 700. As in the conventional mode, memory reads or writes to memory addresses are directed to the first address space 3112 or the second address space 3116 of the BAR0 address space 3110, respectively accessing corresponding memory addresses in the first address space 3122 or the second address space 3126 of the privileged register address space 3120. In SMC mode, SMC graphics register address spaces 3128(0)-3128(7) are provided to respectively access various components in each of the SMC engines 700(0)-700(7). )-540(7), graphics FEs 542(0)-542(7), SKEDs 550(0)-550(7), CWDs 560(0)-560(7), and PDA / PDBs 562(0)-562(7). The GPCs 242 of a particular corresponding SMC engine 700 are individually addressable via a dedicated address range within the legacy graphics register address space 3124, minus any GPCs 242 that are removed due to floor sweeping. Additionally or alternatively, the legacy graphics register address space 3124 includes an address range for broadcasting data to all GPCs 242 of a particular corresponding SMC engine 700 simultaneously. The BAR0 address space 3110 provides two mechanisms for accessing the SMC graphics register address space 3128(0)-3128(7).
[0266] In the first mechanism, the graphics register address space 3114 of the BAR0 address space 3110 is mapped to one of the SMC graphics register address spaces 3128(0)-3128(7) in the privileged register address space 3120. A specific address within the BAR0 address space 3110 accesses the SMC window register. The SMC window register includes two fields. The two fields include an SMC enable field and an SMC index field. The SMC enable field contains a binary logic value, which is either FALSE or TRUE. If the SMC enable field is FALSE, the BAR0 address space 3110 accesses the privileged register address space 3120 in a conventional addressing mode, as described herein. If the SMC enable field is TRUE, the BAR0 address space 3110 accesses the privileged register address space 3120 in an SMC addressing mode based on the value of the SMC index field. The value of the SMC index field specifies which SMC engine 700 is currently mapped to the BAR0 address space 3110. For example, if the value of the SMC index field is 0, then the SMC graphics register address space 3128(0) of the privileged register address space 3120 will be mapped to the graphics register address space 3114 of the BAR0 address space 3110. Similarly, if the value of the SMC index field is 1, then the SMC graphics register address space 3128(1) of the privileged register address space 3120 will be mapped to the register address space 3114 of the graphics BAR0 address space 3110, and so on. Memory reads or writes to memory addresses are directed to the graphics register address space 3114 of the BAR0 address space 3110, accessing the corresponding memory address within the SMC graphics register address space 3128 specified by the SMC index field. Through this first mechanism, the SMC graphics register address space 3128 of the SMC engine 700 specified by the SMC index field can be accessed, while access to the remaining SMC graphics register address spaces 3128 is prohibited.
[0267] In a second mechanism, certain privileged components (e.g., hypervisor 124) can access individual addresses at any of the SMC graphics register address spaces 3128(0)-3128(7). This second mechanism accesses the SMC graphics register address space 3128(0)-3128(7) via two specific addresses in BAR0 address space 3110. One of the two addresses accesses the SMC address register. The other of the two addresses accesses the SMC data register. A specific memory address at any location in the SMC graphics register address space 3128(0)-3128(7) can be accessed in two steps. In the first step, an address is written to the SMC address register that corresponds to an address in the SMC graphics register address space 3128(0)-3128(7). In the second step, the SMC data register is read or written with a data value. Reading or writing the SMC data register in BAR0 address space 3110 results in a corresponding read or write to the privileged register address space 3120 at the memory address specified in the SMC address register. The SMC address register is then dereferenced, thereby enabling the SMC address register and the SMC data register for subsequent transactions. Reading or writing the SMC address register does not result in a read or write to the privileged register address space 3120.
[0268] Fig.32 According to various embodiments, Figure 2 Flowchart of steps of a method for addressing a privileged register address space in a PPU 200. Figures 1 to 17 The method steps are described herein as a system configured to perform the method steps in any order, but one of ordinary skill in the art will understand that any system configured to perform the method steps in any order is within the scope of the present disclosure.
[0269] As shown, the method 3200 begins at step 3202, where the PPU 200 detects a memory access to the privileged register address space 3120. More specifically, the PPU 200 detects a memory access to the BAR0 address space 3110. At step 3204, the PPU 200 determines whether the memory access is to the graphics register address space 3114. If the memory access is not to the graphics register address space 3114, the method 3200 proceeds to step 3214, where the PPU 200 generates a memory transaction to the address specified by the memory access. The method 3200 then terminates.
[0270] Returning to step 3204, if the memory access is to the graphics register address space 3114, the method 3200 proceeds to step 3206, where the PPU 200 determines whether the memory access is in legacy mode. If the memory access is in legacy mode, the method 3200 proceeds to step 3214, where the PPU 200 generates a memory transaction to the address specified by the memory access. The method 3200 then terminates. On the other hand, if the memory access is not in legacy mode, the method 3200 proceeds to step 3208, where the PPU 200 determines whether the memory access is in windowed mode.
[0271] If the memory access is in window mode, the method proceeds to step 3212, where the PPU 200 generates a memory transaction based on the value of the SMC index field specified in the SMC window register. The value of the SMC index field specifies which SMC engine 700 is currently mapped to the BAR0 address space 3110. For example, if the value of the SMC index field is 0, then the SMC graphics register address space 3128 (0) of the privileged register address space 3120 will be mapped to the graphics register address space 3114 of the BAR0 address space 3110. Similarly, if the value of the SMC index field is 1, then the SMC graphics register address space 3128 (1) of the privileged register address space 3120 will be mapped to the graphics register address space 3114 of the BAR0 address space 3110, and so on. A memory read or write to a memory address is directed to the graphics register address space 3114 of the BAR0 address space 3110, accessing the corresponding memory address within the SMC graphics register address space 3128 specified by the SMC index field. Through this first mechanism, the SMC graphics register address space 3128 of the SMC engine 700 specified by the SMC index field may be accessed, while access to the remaining SMC graphics register address space 3128 is prohibited. The method 3200 then terminates.
[0272] Returning to step 3208, if the memory access is not in window mode, the method proceeds to step 3210, where the PPU 200 generates a memory transaction based on the values of the SMC address register and the SMC data register. More specifically, the PPU 200 accesses the SMC graphics register address space 3128(0)-3128(7) through two specific addresses within the BAR0 address space 3110. One of the two addresses accesses the SMC address register. The other of the two addresses accesses the SMC data register. A specific memory address anywhere in the SMC graphics register address space 3128(0)-3128(7) can be accessed in two steps. In the first step, an address is written to the SMC address register that corresponds to an address in the SMC graphics register address space 3128(0)-3128(7). In the second step, the SMC data register is read or written with a data value. Reading or writing the SMC data register in BAR0 address space 3110 causes a corresponding read or write to the privileged register address space 3120 at the memory address specified in the SMC address register. The SMC address register is then dereferenced, thereby enabling the SMC address register and the SMC data register for subsequent transactions. Reading or writing the SMC address register does not cause a read or write to the privileged register address space 3120. Then, method 3200 terminates.
[0273] Performance monitoring using multiple SMC engines
[0274] As discussed further herein, performance monitors (PMs), e.g. Figure 2 PM 236, Figure 3 PM 360 and Figure 4 The PM 430 monitors the overall performance and / or resource consumption of the corresponding components included in the PPU 200. A performance monitor (PM) is included in a performance monitoring system that provides performance monitoring and performance analysis across multiple SMC engines 700. The performance monitoring system simultaneously or substantially simultaneously profiles multiple VMs and processing contexts executed in the VMs. The performance monitoring system isolates multiple virtual machines and multiple processing contexts executed in the VMs from each other with respect to how to generate and capture performance data to prevent leakage of performance data between VMs. The PM 232 and associated counters in the performance data monitoring system track the attributes of performance data to a specific SMC engine 700. In the case of shared resources and units, where the attributes are not traceable to a specific SMC engine 700, a device with a higher privilege entity (such as the hypervisor 124) collects performance data of the shared resources and units. When a VM migrates to other partitioned PPUs 600 and / or other PPUs 200, the performance monitoring system simultaneously profiles the profile compute engine and the graphics engine, as well as the profile VM. The performance monitoring system is now described.
[0275] Fig.33 According to various embodiments, Figure 2 200 , a block diagram of a performance monitoring system 3300 for a PPU 200 of the embodiment of the present invention. As shown in the figure, the performance monitoring system 3300 includes, but is not limited to, a performance monitor 3310, a selection multiplexer 3320, a monitoring bus 3330, and a performance multiplexer unit 3340. The performance monitor 3310 and the selection multiplexer 3320 together constitute a performance monitor module (PMM). Each GPC 242, each partition unit 262, and each system pipeline 230 includes at least one PMM. The performance multiplexer unit 3340 is included in each unit being monitored. The logic within the performance multiplexer unit 3340 is included in one or more of the FE 540, the SKED 550, the CWD 560, and / or other suitable functional units. As further described herein, all components of the various performance monitoring systems 3300 are integrated with the performance monitor aggregator ( Fig.33 Except as further described below, the functionality of the performance monitoring system 3300 is similar to that of the Figure 2 PM 236, Figure 3 The PM 360 and PM 430 are basically the same.
[0276] In operation, the performance multiplexer units 3340(0)-3340(P) enable programmable selection of groups of signals in the PPU 200 that can be monitored by the corresponding performance monitor 3340. Each performance multiplexer 3340 can select a group of signals sent to the monitor bus 3330. A subset of the signals from the monitor bus 3330 is selected for monitoring by the selection multiplexer 3320. Signals from the monitor bus 3330 that are not selected by the selection multiplexer 3320 are not monitored. Signals from within the PPU 200 are connected to the performance multiplexer units 3340(0)-3340(P) in groups so that one group is selected at a time for monitoring. The performance multiplexer unit 3340 multiplexes the signals so that the signals in a particular signal group are selected as a group. The selection inputs of the multiplexers included in the performance multiplexer unit 3340 are programmed by one or more registers included in the privileged register address space 3120. As a result, the specific signals transmitted by performance multiplexer unit 3340 to monitoring bus 3330 are programmable.
[0277] The monitor bus 3330 receives groups of signals from the performance multiplexer units 3340(0)-3340(P). Each signal sent to the monitor bus 3330 is connected as an input to each selection multiplexer 3320.
[0278] The selection multiplexer 3320 includes a set of individual multiplexers 3322(0)-3322(M) and 3324(0)-3324(N). The input side of each multiplexer 3322(0)-3322(M) and 3324(0)-3324(N) receives all signals from the monitoring bus 3330 and selects one signal to send. The selection inputs of the multiplexers 3322(0)-3322(M) and 3324(0)-3324(N) are programmed by one or more registers contained in the privileged register address space 3120. As a result, a specific set of signals sent by the selection multiplexer 3320 is programmable. The selection multiplexer 3320 sends the selected signal to the performance monitor 3310.
[0279] Due to the composition of programmable performance multiplexer unit 3340 and programmable selection multiplexer 3320, the specific PPU signals monitored by performance monitor 3310 are programmable.
[0280] The performance monitor 3310 includes a performance counter array 3312, a shadow counter array 3314, and a trigger function table 3316. The performance monitor 3310 receives a signal sent by a selection multiplexer 3320. More specifically, the shadow counter array 3314 receives a signal sent by multiplexers 3322(0)-3322(M). Similarly, the trigger function table 3316 receives a signal sent by multiplexers 3324(0)-3324(N). As further described, the counters within the shadow counter array 3314 are updated based on the signals received from the multiplexers 3322(0)-3322(M) and various trigger conditions. Typically, the shadow counter array 3314 includes a set of one or more signal counters, each of which increments whenever the signal received from the corresponding multiplexer 3322 is in a specific logic state. Based on certain signals in the form of trigger conditions, the values in the shadow counter array 3314 are transmitted to the performance counter array 3312. Performance counter array 3312 includes a set of one or more signal counters corresponding to the signal counters included in the shadow counter array.
[0281] In one mode of operation, after transferring to performance counter array 3312, the counters in shadow counter array 3314 are reset to zero so that the shadow counter values stored in shadow counter array 3314 always correspond to activity since the previous trigger.
[0282] Performance monitor 3310 may be configured according to various counting modes that define the number of counters included in performance counter array 3312 and shadow counter array 3314. The counting modes also define how and when performance counter array 3312 and shadow counter array 3314 are triggered, and how data from performance counter array 3312 is transmitted to other devices within PPU 200. These various counting modes may be divided into two main performance monitoring modes - non-streaming performance monitoring and streaming performance monitoring.
[0283] In the non-streaming performance monitoring mode, the trigger function table 3316 is programmed to combine the signals received from the multiplexers 3324 (0)-3324 (N) according to certain specified logical signal expressions. When the conditions of one or more of these logical signal expressions are met, the trigger function table 3316 sends a signal to the performance counter array 3312 in the form of a logical trigger 3350. In response to receiving the logical trigger 3350, the performance counter array 3312 samples and stores the current value in the shadow counter array 3314. The value in the performance counter array 3312 is then read through the privileged register address space 3120.
[0284] In streaming performance monitoring mode, the performance monitor aggregator (PMA) sends a signal in the form of a PMA trigger 3352 to the performance counter array 3312. In response to receiving the PMA trigger 3352, the performance counter array 3312 samples and stores the current value in the shadow counter array 3314. The performance monitor 3310 generates a performance monitor (PMM) record, which may include, but is not limited to, the value in the performance counter array 3312 when the PMA trigger 3352 is received from the PMA, a count of the total number of PMA triggers to which the performance monitor 3310 responds, the SMC engine ID, and the PMMID that uniquely identifies the PMM that generated the record in the system. These PMM records are then sent to the PMM router associated with one or more performance monitors 3310. The PMM router then sends the PMM record to the PMA. In some embodiments, the PMM ID for each performance monitor 3310 can be programmed via one or more registers contained in the privileged register address space 3120.
[0285] Typically, a specific performance monitor 3310 in a specific performance monitoring system 3300 in PPU 200 is located in the same clock frequency domain as the signal monitored by the specific performance monitor 3310. However, the specific performance monitor 3310 may be located in the same clock frequency domain or in a different clock frequency domain relative to another performance monitor in PPU 200.
[0286] Various configurations of the performance multiplexer unit 3340 are now described.
[0287] Figure 34A-Figure 34B It shows the various embodiments Fig.33 Various configurations of the performance multiplexer unit 3340.
[0288] like Fig.34A As shown, a first configuration of performance multiplexer unit 3340(0) includes, but is not limited to, signal group A 3420(0)-3420(P), signal group B 3430(0)-3430(Q), and multiplexers 3412(0) and 3412(1). In operation, multiplexer 3412(0) selects one of signal groups A 3420(0)-3420(P), wherein each of signal groups A 3420(0)-3420(P) is a subgroup of a larger signal group C. Multiplexer 3412(0) selects one of signal groups A 3420(0)-3420(P) and transmits the selected signal group to monitoring bus 3330. Similarly, multiplexer 3412(1) selects one of signal group B 3430(0)-3430(Q), wherein each of signal group B 3430(0)-3430(Q) is a subgroup of a larger signal group D. Multiplexer 3412(1) selects one of signal group B 3430(0)-3430(Q) and transmits the selected subgroup to monitor bus 3330. The select input of multiplexer 3412 included in performance multiplexer unit 3340(0) is programmed by one or more registers contained in privileged register address space 3120. As a result, the set of signals sent by performance multiplexer unit 3340(0) is programmable.
[0289] like Fig.34BAs shown, a second configuration of the performance multiplexer unit 3340(1) includes, but is not limited to, signal group C 3440(0)-3440(R) and multiplexer 3412(2). In operation, multiplexer 3412(2) selects one of signal groups C 3440(0)-3440(R), wherein each of signal groups C 3440(0)-3440(R) is a subgroup of a larger signal group E. Multiplexer 3412(2) selects one of signal groups C 3440(0)-3440(R) and transmits the selected signal group to monitoring bus 3330. In the configuration of performance multiplexer unit 3340(1), several signals are transmitted to multiple signal groups. In particular, signal C1 3450 is sent to signal group C 3440(0) and signal group C 3440(1). Similarly, signal C2 3452 is sent to signal group C 3440(1) and signal group C 3440(2). The select input of multiplexer 3412(2) included in performance multiplexer unit 3340(1) is programmed by one or more registers included in privileged register address space 3120. As a result, the set of signals transmitted by performance multiplexer unit 3340(1) is programmable. The configuration of performance multiplexer unit 3340(1) may be useful in making signals available in multiple signal groups 3440 to facilitate visibility of certain signal groups in a single pass of performance monitoring system 3300.
[0290] Fig.35 According to various embodiments, Figure 2 2 is a block diagram of a performance monitor aggregation system 3500 of a PPU 200. As shown, the performance monitor aggregation system 3500 includes, but is not limited to, GPCs 242(0)-242(M), partition units 262(0)-262(N), crossbar unit 250, control crossbar and SMC arbiter 510, PM management system 3530, and performance analysis system 3540.
[0291] In operation, GPCs 242(0)-242(M) perform various processing tasks for one or more system pipelines 230. Each GPC 242 includes multiple parallel processing cores capable of executing a large number of threads simultaneously, and has any degree of independence and / or isolation from other GPCs 242. Each GPC 242(0)-242(M) includes one or more PMs 360(0)-360(M) and GPC PMM routers 3514(0)-3514(M). The functions of PMs 360(0)-360(M) are substantially similar to those of Fig.333310. PMs 360(0)-360(M) generate PMM records that include performance data for corresponding GPCs 242(0)-242(M). PMs 360(0)-360(M) send these PMM records to and receive data from corresponding GPC PMM routers 3514(0)-3514(M). GPC PMM routers 3514(0)-3514(M) transmit the PMM records to PM management system 3530 via crossbar unit 250.
[0292] Partition units 262(0)-262(N) provide access to PPU memory ( Fig.35 Each partition unit 262 performs memory access operations using different DRAMs in parallel with each other, thereby efficiently utilizing the available memory bandwidth of the PPU memory. Each of the partition units 262(0)-262(N) includes one or more PMs 430(0)-430(N) and partition unit (PU) PMM routers 3524(0)-3524(N). The functions of PMs 430(0)-430(N) are substantially similar to those of Fig.33 The performance monitor 3310 of the embodiment of the present invention. The PM 430(0)-430(N) generates PMM records, which include the performance data of the corresponding partition units 262(0)-262(N). The PM 430(0)-430(N) sends these PMM records to the corresponding PU PMM routers 3524(0)-3524(N) and receives data from them. The PU PMM routers 3524(0)-3524(N) in turn transmit the PMM records to the PM management system 3530 by controlling the crossbar switch and the SMC arbiter 510.
[0293] PM management system 3530 controls the collection of PMM records and stores PMM records for reporting purposes. PM management system 3530 includes but is not limited to system performance monitor 3532, system (SYS) PMM router 3534, performance monitor aggregator (PMA) 3536, high speed hub (HSHUB) and transmission logic 3539.
[0294] The functions of the PM 3532 system are basically similar to Fig.33 The system PM 3532 generates PMM records that contain performance data for system-wide components that are not contained in a particular GPC 242 or partition unit 262. The system PM 3532 sends these PMM records to and receives data from the system PMM router 3534. The system PMM router 3534 in turn sends the PMM records to the PMA 3536.
[0295] The PMA 3536 generates triggers for various performance monitors including PMs 360(0)-360(M), PMs 430(0)-430(N), and system PMs 3532. The PMA 3536 generates these triggers by two techniques. In the first technique, the PMA 3536 generates triggers in response to signals sent by each system pipeline 230 when the system pipeline 230 receives commands from the host interface 220. In the second technique, the PMA 3536 generates triggers by periodically sending programmed controlled trigger pulses to the PMs. Typically, the performance monitor aggregation system 3500 includes at least one programmable trigger pulse generator corresponding to each system pipeline 230 in addition to another trigger pulse generator independent of any system pipeline 230. The PMA 3536 sends trigger signals to the GPC PMM routers 3514(0)-3514(M) and the PU PMs 3524(0)-3524(N) by controlling the crossbar switch and the SMC arbiter 510. PMA 3536 sends the trigger directly to the system PMM router 3534 through the communication link inside the PM management system 3530. The PMM router sends the PMA trigger to the corresponding PM. In some embodiments, these triggers take the following combination Fig.36 Describes the form of the trigger packet.
[0296] Fig.36 According to various embodiments, Fig.35 The format of the trigger packet associated with the performance monitor aggregation system 3500. The purpose of the trigger packet is to convey information about the trigger source to the performance monitor 3310. Each of the performance monitors 3310 uses this information to determine whether to respond to a specific trigger. In this regard, the trigger packet contains information that can be used by each performance monitor 3310 to determine whether to respond to a specific trigger packet. Each performance monitor 3310 associated with a specific SMC engine 700 is programmed with the SMC engine ID corresponding to the SMC engine 700. Such a performance monitor 3310 responds to each SMC trigger packet including the same SMC engine ID. Each performance monitor 3310 that is not associated with a specific SMC engine 700 or each performance monitor 3310 that is programmed with an invalid SMC engine ID does not respond to each SMC trigger packet. On the contrary, such a performance monitor 3310 responds to a shared trigger packet.
[0297] Figure 3600 shows the general format of a triggered packet. As shown, Figure 3600 includes a packet type 3602 indicating that the packet is PM triggered, a trigger type 3604, and a trigger payload 3606. PM trigger type 3604 is an enumerated value that identifies the type of trigger format. For example, in order to identify three different types of triggered packets, PM trigger type 3604 can be a 2-bit value. Trigger payload 3606 includes different data based on PM trigger type 3604. Three different types of triggered packets are now described, wherein the three types of triggered packets correspond to the three categories of performance monitoring data (traditional data, each SMC data, and shared data).
[0298] Diagram 3610 shows the format of a legacy trigger packet. The legacy trigger packet includes a packet type 3602 indicating that the packet is a PM trigger and a trigger type 3614 indicating that the trigger packet is a legacy trigger packet. The trigger payload 3606 of the legacy trigger packet includes an unused field 3616.
[0299] Diagram 3620 illustrates the format of each SMC trigger packet. Each SMC trigger packet includes a packet type 3602 indicating that the packet is a PM trigger and a trigger type 3624 indicating that the trigger packet is an SMC trigger packet. The trigger payload 3606 of each SMC trigger packet includes an SMC engine ID field. The SMC engine ID field 3626 identifies the specific SMC engine 700 to which the trigger applies.
[0300] Diagram 3630 shows the format of a shared trigger packet. The shared trigger packet includes a packet type 3602 indicating that the packet is a PM trigger and a trigger type 3634 indicating that the trigger packet is a shared trigger packet. The trigger payload 3606 of the shared trigger packet includes unused fields.
[0301] The type of trigger packet sent by the PMA 3536 is determined by the trigger source and one or more registers included in the privileged register address space 3120 corresponding to each trigger source, which are programmed to indicate the type of trigger packet that the PMA should send for that source. In one mode of operation, the PMA is programmed so that the trigger packets generated in response to a source associated with an SMC engine are each SMC trigger packets that set the SMC engine ID to the corresponding SMC engine, and the trigger packets generated in response to a source not associated with an SMC engine are shared trigger packets. Trigger packets can be generated at any technically feasible rate, with a maximum of one trigger packet per computing cycle.
[0302] In response to receiving a trigger packet from PMA 3536, each PM checks the trigger type and the trigger payload to determine whether the PM should respond to the trigger. Each PM responds to a traditional trigger packet unconditionally. In the case of a per SMC trigger packet, the PM responds only if the SMC engine ID contained in the trigger payload matches the SMC engine ID assigned to the PM via register privileged register address space 3120 programming. An invalid SMC engine ID programmed in this register ensures that the PM does not respond to each SMC trigger packet. In the case of a shared trigger packet, the PM responds only if the PM has been programmed to respond via a register in privileged register address space 3120. In one operating mode, all PMs that are uniquely assigned to an SMC engine as monitoring units are programmed to respond to each SMC trigger packet with a corresponding SMC engine ID payload, while all other PMs are programmed to respond to shared trigger packets, but not to each SMC trigger packet.
[0303] In the event that the PM determines that a response to a trigger is warranted, the PM samples the counter included in the corresponding PM. The responding PM then transmits a PMM record that includes the sampled counter value, the total number of triggers responded to, the SMC engine ID assigned to the PM, and the PMM ID that uniquely identifies the PM in the system. The PMM router, PMA 3536, and / or performance analysis system 3540 use the PMM ID to identify which PM sent the corresponding PMM record. The PMM router then sends the record to the PMA 3536. More specifically, the GPC PMM routers 3514 (0)-3514 (M) send the PMM record to the PMA 3536 via the crossbar switch unit 250 and the high-speed hub 3538. The PU PMM routers 3524 (0)-3524 (N) send the PMM record to the PMA 3536 via the control crossbar switch and the SMC arbiter 510. The system PMM router 3534 transmits the PMM record directly to the PM management system 3530 via a communication link. In this way, PMA 3536 receives PMM records from all relevant PMs in PPU 200.
[0304] In some embodiments, when a trigger is transmitted, the PMA 3536 additionally generates a PMA record. Generally, the purpose of a PMA record is to record the timestamp at which a specific PMA trigger is generated, and to associate a PMM record corresponding to the performance monitor 3310 that is triggered by the PMA in response to that timestamp. A PMA record includes, but is not limited to, a timestamp, an SMC engine ID associated with the source of the PMA trigger, the total number of triggers generated by the source with the same SMC engine ID, and associated metadata. When the performance monitor 3310 receives a trigger, the performance monitor 3310 also generates a PMM record with a trigger count. Subsequently, when the PMA record and the PMM record are parsed, a PMM record with a specific trigger count can be associated with a PMA record with the same trigger count. In this way, a timestamp corresponding to a PMM record is established based on the timestamp of the associated PMA record. As a result, the behavior of the PPU 200 reflected by the PMM record is accurately associated with the time range defined by two adjacent PMA triggers from the same source.
[0305] When receiving PMM records and generating PMA records, PMA 3536 stores PMM records and PMA records in the form of record buffers in data storage in PPU memory through high-speed hub 3538. High-speed hub 3538 sends PMM records and PMA records to partition units 262 (0)-262 (N). Then, the partition units store PMM records and PMA records in record buffers in PPU memory. Additionally or optionally, high-speed hub 3538 transmits PMM records and PMA records to performance analysis system 3540 via transmission logic 3539. In some embodiments, high-speed hub 3538, transmission logic 3539 and performance analysis system 3540 can communicate with each other via PCIe link. A user can view PMM records and PMA records on performance analysis system 3540 to characterize the behavior of PPU 200 as reflected in PMM records. Performance analysis system 3540 collects PMM records and PMA records with the same trigger count. The performance analysis system 3540 then uses the timestamp from the PMA record and the performance data from the PMM record with the same trigger count to determine the timestamp associated with the performance data. In some embodiments, the performance analysis system 3540 can access the performance record buffer as a virtual memory. As a result of placing the performance record buffer in different virtual address spaces, the performance monitoring data of different SMC engines 700 can be isolated from each other, as now described.
[0306] The PMA 3536 provides isolation of performance monitoring data between several SMC engines 700. In particular, the PMA 3536 classifies PMM records and PMA records into categories based on the operating mode. When the PPU 200 operates in the traditional mode, the PPU 200 executes the task as a single hardware resource cluster, rather than as a separate PPU partition 600 with a separate SMC engine 700. In the traditional mode, the PMA 3536 stores the PMM records and PMA records in a single category as a single set of performance monitoring data. When the PPU 200 operates in the SMC mode, the PPU 200 executes the task as a separate PPU partition 600 with a separate SMC engine 700. In the SMC mode, the PMA 3536 classifies and stores the PMM records and PMA records in different categories using the SMC engine ID field of the PMM records and PMA records. Records with each SMC engine ID are stored in a separate data storage in the form of a record buffer that is accessible from a different virtual address space matching the corresponding SMC engine 700. As described above, PMM records and PMA records that are not traceable to a specific SMC engine 700 contain invalid SMC engine IDs. The PMA stores these records in a separate data store in the form of an SMC record buffer in a virtual address space that is accessible by any authorized entity with sufficient privileges to access the data of all SMC engines 700. Such authorized entities include, but are not limited to, the hypervisor 124 in a virtual environment and the root user or operating system kernel in a non-virtual environment. Each SMC engine 700 can access some or all of the performance monitoring data in the non-SMC PMA record buffer by requesting the data from the authorized entity.
[0307] In some embodiments, PMA 3536 is configured such that a trigger corresponding to each SMC engine 700 is generated to coincide with a context switch event of the same SMC engine 700. In such an embodiment, PM is configured such that the counters in shadow counter array 3314 are reset to zero after each trigger so that data transmitted from PM to PMA 3536 per SMC engine ID while time slicing is enabled is attributable to a single context or VM.
[0308] Fig.37 According to various embodiments, a method for monitoring Figure 2 Flowchart of method steps for the performance of the PPU 200. Although combined Figure 1-Figure 23 The method steps are described herein as a system with a plurality of method steps, but one of ordinary skill in the art will understand that any system configured to perform the method steps in any order is within the scope of the present disclosure.
[0309] As shown, method 3700 begins at step 3702, where PMA 3536 generates and sends a trigger to PM 3310. In addition, PMA 3536 generates a corresponding PMA record, which includes a timestamp and, optionally, an SMC engine ID corresponding to the source of the trigger. PM 3310 receives the trigger to sample performance data.
[0310] In step 3704, in response to receiving the trigger, PM 3310 determines whether a response to the trigger is guaranteed. Each PM 3310 checks the trigger type and trigger payload from the trigger packet to determine whether PM 3310 should respond to the trigger. Each PM 3310 responds to the traditional trigger packet unconditionally. In the case of each SMC trigger packet, PM 3310 responds only when the SMC engine ID contained in the trigger payload matches the SMC engine ID assigned to PM 3310 by register programming in the privileged register address space 3120. The invalid amount of SMC engine ID programmed in the register ensures that PM 3310 does not respond to each SMC trigger packet. In the case of a shared trigger packet, PM 3310 responds only when PM 3310 has been programmed to respond through the register in the privileged register address space 3120. In one mode of operation, all PMs 3310 that are uniquely assigned to the SMC engine 700 as monitoring units are programmed to respond to each SMC trigger packet with a corresponding SMC engine ID payload. All other PMs 3310 are programmed to respond to the shared trigger packets, but not to the per-SMC trigger packets.
[0311] If a response is required, performance counter array 3312 samples and stores the current value in shadow counter array 3314. In non-streaming mode, other components can read the values in the performance counter array through privileged register interface hub 512.
[0312] In streaming mode, the method proceeds to step 3706, where PM 3310 sends sampled performance data and PMA 3536 receives sampled performance data from PM 3310. PM 3310 generates PMM records that include the values in performance counter array 3312 when a PMA trigger 3352 is received from PMA 3536. These PMM records are then sent to the PMM router associated with the particular performance monitor 3310. The PMM router, in turn, sends the PMM records to PMA 3536.
[0313] In step 3708, PMA 3536 classifies PMM records and PMA records into categories based on the operating mode. When PPU 200 operates in the traditional mode, PMA 3536 classifies PMM records and PMA records in a single category as a single set of performance monitoring data. When PPU 200 operates in the SMC mode, PMA 3536 classifies PMM records and PMA records into different categories using the SMC engine ID field of the PMM records and PMA records. As described above, PMM records and PMA records that cannot be traced back to a specific SMC engine 700 contain an invalid SMC engine ID. PMA 3536 classifies these PMM records and PMA records into a separate category.
[0314] At step 3710, the PMA 3536 stores the PMM records and / or PMA records in a PMA record buffer in the PPU memory. When the PPU 200 operates in the traditional mode, the PMA 3536 stores the PMM records and PMA records in a single category as a single set of performance monitoring data. When the PPU 200 operates in the SMC mode, the PMA 3536 stores the PMM records and PMA records associated with each SMC engine ID in a separate data store accessible from a different virtual address space matching the corresponding SMC engine 700. The PMA 3536 stores the PMM records and PMA records that cannot be traced back to a specific SMC engine 700 in a separate data store in the form of a non-SMC PMA record buffer in a virtual address space that is accessible only to any authorized entity with sufficient privileges to access the data of all SMC engines 700. Such authorized entities include, but are not limited to, the hypervisor 124 in a virtualized environment and the root user or operating system kernel in a non-virtualized environment. Each SMC engine 700 may access some or all of the performance monitoring data in the non-SMC PMA record buffer by requesting the data from an authorized entity.
[0315] More specifically, PMA 3536 transmits PMM records and PMA record streams to high-speed hub 3538. High-speed hub 3538 sends the PMM records and PMA transmissions to partition units 262(0)-262(N). The partition units then store the PMM records and PMA records in PMA record buffers in PPU memory. A PMA record buffer for each record is selected based on the SMC engine ID field of the record so that each record ultimately resides in PPU memory that is accessible in a virtual address space that matches the SMC engine corresponding to the PMA record buffer. PMM records and PMA records that are not traceable to a specific SMC engine 700 in a separate data store can only be accessed by authorized entities.
[0316] In step 3712, PMA 3536 sends the PMA record and / or PMM record to performance analysis system 3540 via high-speed hub 3538 and transmission logic 3539. Additionally or alternatively, performance analysis system 3540 accesses the PMA record and / or PMM record through one or more virtual addresses in the virtual address space. Typically, performance analysis system 3540 includes a software application executed on CPU 110 and / or any other technically feasible processor. Performance analysis system 3540 directly accesses virtual memory to access PMA record and / or PMM record. Virtual memory can be associated with PPU 200 and / or CPU 110. A user can view PMA record and / or PMM record on performance analysis system 3540 to characterize the behavior of PPU 200 as reflected in PMA record and / or PMM record. Then, method 3700 terminates.
[0317] Power and clock frequency management of the SMC engine
[0318] Complex systems, such as Figure 2 The PPU 200 may consume a lot of power. More specifically, certain components within the PPU 200 may have different power consumption levels from each other at different points in time. In one example, components in one PPU partition 600 may perform computationally and / or graphics intensive tasks, thereby increasing power consumption relative to other PPU partitions 600. In another example, due to leakage current and related factors, the PPU partition 600 may consume power even when it is idle. In addition, the increased power consumption may result in higher operating temperatures, which may in turn result in reduced performance. As a result, the PPU 200 includes power and clock frequency management that takes into account how power consumption within one PPU partition 600 may negatively impact the performance of other PPU partitions 600.
[0319] Fig.38 According to various embodiments, Figure 2 2 is a block diagram of a power and clock frequency management system 3800 of a PPU 200. The power and clock frequency management system 3800 includes, but is not limited to, circuit subsections 3810(0)-3810(N), a power gate controller 3820, and a clock frequency controller 3830.
[0320] Each of the circuit subsections 3810(0)-3810(N) includes any set of components contained in the PPU 200 at any level of granularity. In this regard, each of the circuit subsections 3810(0)-3810(N) may include, but is not limited to, the system pipeline 230, the PPU partition 600, the PPU slice 610, the SMC engine 700, or any technically feasible subset thereof.
[0321] In operation, the power gate controller 3820 monitors the activity state of the circuit subsections 3810(0)-3810(N). If the power gate controller 3820 determines that a particular circuit subsection (such as the circuit subsection 3810(2)) is in an idle state, the power gate controller 3820 reduces the supply voltage of the circuit subsection 3810(2) to a voltage that is less than the operating voltage but maintains the data stored in the memory. Alternatively, the power gate controller 3820 can remove power from the circuit subsection 3810(2), thereby shutting down the circuit subsection 3810(2). Subsequently, if the circuit subsection 3810(2) is needed to perform certain tasks, the power gate controller 3820 increases the supply voltage of the circuit subsection 3180(2) to a voltage suitable for operation.
[0322] Clock frequency controller 3830 monitors power consumption of circuit subsections 3810(0)-3810(N). If clock frequency controller 3830 determines that a particular circuit subsection (e.g., circuit subsection 3810(3)) consumes more power relative to other circuit subsections 3810, clock frequency controller 3830 reduces the frequency of a clock signal associated with circuit subsection 3810(3). As a result, the power consumed by circuit subsection 3810(3) is reduced. Subsequently, if clock frequency controller 3830 determines that circuit subsection 3810(3) consumes less power relative to other circuit subsections 3810, clock frequency controller 3830 increases the frequency of a clock signal associated with circuit subsection 3810(3), thereby increasing the performance of circuit subsection 3810(3).
[0323] In this manner, the power gate controller 3820 and the clock frequency controller 3830 reduce the overall power consumption of the PPU 200 and reduce the negative impact of one PPU partition 600 on another PPU partition 600 due to temperature effects.
[0324] Fig.39 is a method for managing Figure 2 Flow chart of the steps of a method for calculating the power consumption of the PPU 200. Figures 1 to 25 The method steps are described herein as a system configured to perform the method steps in any order, but one of ordinary skill in the art will understand that any system configured to perform the method steps in any order is within the scope of the present disclosure.
[0325] As shown, the method 3900 begins at step 3902, where the power and clock frequency management system 3800 of the PPU 200 monitors the activity status of the VMs executed on the various circuit sub-portions 3810 of the PPU 200. At step 3904, the power and clock frequency management system 3800 determines whether any circuit sub-portion 3810 is in an idle state. If no circuit sub-portion 3810 is in an idle state, the method proceeds to step 3908. However, if one or more circuit sub-portions 3810 are in an idle state, the method proceeds to step 3906, where the power and clock frequency management system 3800 reduces the power supply voltage of the idle circuit sub-portions 3810. In particular, the power gate controller 3820 in the power and clock frequency management system 3800 reduces the supply voltage of the circuit sub-portion 3810 (2) to a voltage that is less than the operating voltage but maintains the data stored in the memory. Alternatively, the power gate controller 3820 can remove power from the circuit sub-portion 3810 (2), thereby shutting down the circuit sub-portion 3810 (2).
[0326] In step 3908, the power and clock frequency management system 3800 monitors the power consumption of each SMC engine 700 in the PPU. In step 3910, the power and clock frequency management system 3800 determines whether one or more SMC engines 700 are consuming too much power relative to other MC engines 700. If no SMC engine 700 is consuming too much power, the method proceeds to step 3902 to continue monitoring. However, if one or more SMC engines 700 are consuming too much power, the method proceeds to step 3912, where the clock frequency controller 3830 within the power and clock frequency management system 3800 reduces the clock frequency of one or more circuit sub-sections 3810 associated with the SMC engine 700 that is consuming too much power. The method then proceeds to step 3902 to continue monitoring.
[0327] In summary, various embodiments include a parallel processing unit (PPU) that can be divided into partitions. Each partition is configured to simultaneously execute processing tasks associated with multiple processing contexts. A given partition includes one or more logical groupings or "slices" of GPU resources. Each slice provides sufficient computing, graphics, and memory resources to simulate the operation of an entire PPU. A hypervisor executing on the CPU performs various techniques to partition the PPU on behalf of an administrator user. A guest user is assigned to a partition and can then perform processing tasks within that partition that are isolated from any other guest user assigned to any other partition.
[0328] One technical advantage of the disclosed technology over the prior art is that, using the disclosed technology, the PPU can support multiple processing contexts simultaneously and be functionally isolated from each other. Therefore, multiple CPU processes can effectively utilize PPU resources via multiple different processing contexts without interfering with each other. Another technical advantage of the disclosed technology is that because the PPU can be partitioned into isolated computing environments using the disclosed technology, the PPU can support a more reliable form of multi-tenancy relative to the prior art methods that rely on processing sub-contexts to provide multi-tenant functionality. Therefore, when the disclosed technology is implemented, the PPU becomes more suitable for cloud-based deployments, where access to different partitions within the same PPU can be provided to different and potentially competing entities. These technical advantages represent one or more technical advances over the prior art methods.
[0329] 1. In some embodiments, a computer-implemented method comprises: generating a first signal to sample performance data of multiple engines included in a processor, wherein the performance data is captured by one or more performance monitors; receiving the performance data from the one or more performance monitors based on the first signal; extracting a first subset of the performance data associated with a first engine included in the multiple engines; and storing the first subset of the performance data in a first data storage accessible to the first engine.
[0330] 2. The computer-implemented method of clause 1, wherein the first data store is inaccessible to all other engines included in the plurality of engines.
[0331] 3. The computer-implemented method according to clause 1 or 2 further includes: extracting a portion of the performance data that cannot be traced to any engine included in the multiple engines; and storing the portion of the performance data that cannot be traced to any engine in a second data storage.
[0332] 4. The computer-implemented method of any of clauses 1-3, wherein the second data store is accessible to an authorized entity associated with the processor and is inaccessible to all engines included in the plurality of engines.
[0333] 5. A computer-implemented method according to any one of clauses 1-4, wherein generating the first signal to sample the performance data comprises: sending the first signal to a signal counter array included in a first performance monitor included in the one or more performance monitors; and sampling at least a portion of the performance data via the signal counter array.
[0334] 6. A computer-implemented method according to any one of clauses 1-5, wherein generating the first signal to sample the performance data comprises: combining one or more signals received by a first performance monitor included in the one or more performance monitors according to a logical signal expression; determining a condition that satisfies the logical signal expression; in response, sending the first signal to a signal counter array included in the first performance monitor; and sampling at least a portion of the performance data via the signal counter array.
[0335] 7. A computer-implemented method as recited in any one of clauses 1-6, wherein the performance data is based on a first signal group received via a first multiplexer.
[0336] 8. A computer-implemented method as recited in any one of clauses 1-7, wherein the performance data is further based on a second signal group received via a second multiplexer.
[0337] 9. The computer-implemented method of any one of clauses 1-8, further comprising: extracting a second subset of the performance data associated with a second engine included in the plurality of engines; and storing the second subset of the performance data in a second data store accessible to the second engine.
[0338] 10. In some embodiments, a non-transitory computer-readable medium stores program instructions that, when executed by a processor, cause the processor to perform the following steps: generate a first signal to sample performance data of multiple engines included in the processor; based on the first signal, receive the performance data; extract a subset of the performance data associated with a first engine included in the multiple engines; and store the subset of the performance data in a first data storage accessible to the first engine.
[0339] 11. The non-transitory computer-readable medium of clause 10, wherein generating the first signal to sample the performance data comprises: sending the first signal to a signal counter array included in a performance monitor; and sampling at least a portion of the performance data via the signal counter array.
[0340] 12. A non-transitory computer-readable medium according to clause 10 or 11, wherein generating the first signal to sample the performance data includes: combining one or more signals received by a performance monitor according to a logical signal expression; determining a condition that satisfies the logical signal expression; in response, sending the first signal to a signal counter array included in the performance monitor; and sampling at least a portion of the performance data via the signal counter array.
[0341] 13. The non-transitory computer-readable medium of any of clauses 10-12, wherein the performance data is based on a first signal group received via a first multiplexer.
[0342] 14. The non-transitory computer-readable medium of any of clauses 10-13, wherein the performance data is further based on a second signal group received via a second multiplexer.
[0343] 15. The non-transitory computer-readable medium of any of clauses 10-14, wherein the performance data is based on a first performance monitor associated with a first clock signal domain and a second performance monitor associated with a second clock signal domain.
[0344] 16. The non-transitory computer-readable medium of any of clauses 10-15, wherein the performance data is associated with a duration between the first signal and the second signal to sample the performance data of the plurality of engines.
[0345] 17. The non-transitory computer-readable medium of any of clauses 10-16, wherein the first signal coincides with a first context switch event associated with the first engine, and the second signal coincides with a second context switch event associated with the first engine.
[0346] 18. A system comprising: a memory storing a software application; and a processor, which, when executing the software application, is configured to perform the following steps: generate a first signal to sample performance data of multiple engines included in the processor; enable one or more performance monitors to capture the performance data based on the first signal; receive the performance data from the one or more performance monitors; extract a subset of the performance data associated with a first engine included in the multiple engines; and store the subset of the performance data in a first data storage accessible to the first engine.
[0347] 19. The system of clause 18, wherein the processor executes a plurality of virtual machines, and further comprising: determining that no virtual machine included in the plurality of virtual machines is utilizing a first circuit subportion included in the processor; and reducing a power supply voltage associated with the first circuit subportion.
[0348] 20. A system according to clause 18 or 19, wherein each circuit sub-portion included in the plurality of circuit sub-portions is associated with a different engine included in the plurality of engines, and further comprising: determining that a first circuit sub-portion included in the plurality of circuit sub-portions consumes more power than each other circuit sub-portion included in the plurality of circuit sub-portions; and reducing the frequency of a clock signal associated with the first circuit sub-portion.
[0349] Any and all combinations of any claim elements recited in any claim and / or any elements described in this application, in any manner, are within the intended scope of the present embodiments and protection.
[0350] The description of the various embodiments has been given for purposes of illustration, but is not intended to be exhaustive or limited to the embodiments disclosed. Many modifications and variations will be apparent to those of ordinary skill in the art without departing from the scope and spirit of the described embodiments.
[0351] Aspects of embodiments of the present invention may be implemented as systems, methods or computer program products. Therefore, aspects of the present disclosure may take full hardware embodiments, full software embodiments (including firmware, resident software, microcode, etc.) or embodiments combining software and hardware aspects, which are generally referred to herein as "modules", "systems" or "computers". In addition, any hardware and / or software technology, process, function, component, engine, module or system described in the present disclosure may be implemented as circuits or circuit sets. In addition, aspects of the present disclosure may take the form of a computer program product embodied in one or more computer-readable media embodied with computer-readable program code thereon.
[0352] Any combination of one or more computer-readable media can be utilized. The computer-readable medium can be a computer-readable signal medium or a computer-readable storage medium. Computer-readable storage media can be, for example, but not limited to, electronic, magnetic, optical, electromagnetic, infrared or semiconductor systems, devices or equipment, or any suitable combination of the foregoing. More specific examples (non-exhaustive list) of computer-readable storage media will include the following: an electrical connection with one or more wires, a portable computer floppy disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any other suitable combination of the foregoing. In the context of this article, a computer-readable storage medium can be any tangible medium that can contain or store a program for use by or in conjunction with an instruction execution system, device or equipment.
[0353] Aspects of the present disclosure are described above with reference to the flowchart and / or block diagram of the method, device (system) and computer program product according to the embodiment of the present disclosure. It will be understood that each frame of the flowchart illustration and / or block diagram and the combination of frames in the flowchart illustration and / or block diagram can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer or other programmable data processing device to produce a machine. When the instruction is executed by the processor of a computer or other programmable data processing device, the function / action specified in the flowchart and / or block diagram box can be realized. Such a processor can be, but is not limited to, a general-purpose processor, a special-purpose processor, an application-specific processor or a field programmable gate array.
[0354] The flow chart and block diagram in the accompanying drawings illustrate the architecture, function and operation that the system, method and computer program product according to various embodiments of the present invention may realize.In this regard, each box in the flow chart or block diagram can represent a module, a fragment or a part of a code, which includes one or more executable instructions for realizing one or more specified logical functions.It should also be noted that in some alternative embodiments, the function indicated in the box may not occur in the order indicated in the figure.For example, depending on the function involved, two boxes shown in succession can actually be executed substantially at the same time, or sometimes these boxes can be executed in reverse order.It should also be noted that each box of the block diagram and / or flow chart description and the combination of the boxes in the block diagram and / or flow chart description can be realized by a combination of a system based on special-purpose hardware or special-purpose hardware and computer instructions that performs a specified function or action.
[0355] While the foregoing is directed to embodiments of the present disclosure, other and further embodiments of the disclosure may be devised without departing from the basic scope thereof, and the scope of the disclosure is determined by the claims that follow.
Claims
1. A computer-implemented method comprising: generating a first signal to sample performance data of a plurality of engines included in a processor, wherein the performance data is captured by a plurality of performance monitors; receiving the performance data from the plurality of performance monitors based on the first signal, wherein the performance data includes a first identifier that identifies a performance monitor included in the plurality of performance monitors that generated the performance data; extracting a plurality of performance data subsets from the performance data based on a plurality of second identifiers included in the performance data, wherein a first performance data subset of the plurality of performance data includes the first identifier, wherein each of the plurality of performance data subsets corresponds to a different second identifier included in the plurality of second identifiers, and wherein each of the plurality of second identifiers corresponds to a different engine included in the plurality of engines; as well as Each of the plurality of performance data subsets is stored in a different data store among a plurality of data stores, wherein for each of the plurality of performance data subsets, the different data store among the plurality of data stores is determined based on the second identifier included in the performance data subset, wherein each data store included in the plurality of data stores is isolated from access by other engines included in the plurality of engines except for a corresponding engine, and wherein each of the plurality of engines accesses a corresponding data store among the plurality of data stores via a different virtual address space. 2 . The computer-implemented method of claim 1 , wherein each data store included in the plurality of data stores is inaccessible to all other engines included in the plurality of engines.
3. The computer-implemented method of claim 1 , further comprising: extracting a portion of the performance data that is not traceable to any engine included in the plurality of engines; as well as The portion of the performance data that is not traceable to any engine is stored in a first data store. 4 . The computer-implemented method of claim 3 , wherein the first data store is accessible to an authorized entity associated with the processor and is inaccessible to all engines included in the plurality of engines.
5. The computer-implemented method of claim 1 , wherein generating the first signal to sample the performance data comprises: sending the first signal to a signal counter array included in a first performance monitor included in the plurality of performance monitors; as well as At least a portion of the performance data is sampled via the signal counter array.
6. The computer-implemented method of claim 1 , wherein generating the first signal to sample the performance data comprises: combining one or more signals received by a first performance monitor included in the plurality of performance monitors according to a logical signal expression; Determining a condition that satisfies the logic signal expression; in response, sending the first signal to a signal counter array included in the first performance monitor; as well as At least a portion of the performance data is sampled via the signal counter array.
7. The computer-implemented method of claim 1, wherein the performance data is based on a first signal group received via a first multiplexer.
8. The computer-implemented method of claim 7, wherein the performance data is further based on a second signal group received via a second multiplexer.
9. A non-transitory computer readable medium storing program instructions which, when executed by a processor, cause the processor to perform the following steps: generating a first signal to sample performance data of a plurality of engines included in a processor, wherein the performance data is captured by a plurality of performance monitors; receiving the performance data from the plurality of performance monitors based on the first signal, wherein the performance data includes a first identifier that identifies a performance monitor included in the plurality of performance monitors that generated the performance data; extracting a plurality of performance data subsets from the performance data based on a plurality of second identifiers included in the performance data, wherein a first performance data subset of the plurality of performance data includes the first identifier, wherein each of the plurality of performance data subsets corresponds to a different second identifier included in the plurality of second identifiers, and wherein each of the plurality of second identifiers corresponds to a different engine included in the plurality of engines; as well as Each of the plurality of performance data subsets is stored in a different data store among a plurality of data stores, wherein for each of the plurality of performance data subsets, the different data store among the plurality of data stores is determined based on the second identifier included in the performance data subset, wherein each data store included in the plurality of data stores is isolated from access by other engines included in the plurality of engines except for a corresponding engine, and wherein each of the plurality of engines accesses a corresponding data store among the plurality of data stores via a different virtual address space.
10. The non-transitory computer-readable medium of claim 9, wherein generating the first signal to sample the performance data comprises: sending the first signal to a signal counter array included in a performance monitor; as well as At least a portion of the performance data is sampled via the signal counter array.
11. The non-transitory computer-readable medium of claim 9, wherein generating the first signal to sample the performance data comprises: combining one or more signals received by the performance monitor according to a logical signal expression; Determining a condition that satisfies the logic signal expression; in response, sending the first signal to a signal counter array included in the performance monitor; as well as At least a portion of the performance data is sampled via the signal counter array.
12. The non-transitory computer-readable medium of claim 9, wherein the performance data is based on a first signal group received via a first multiplexer.
13. The non-transitory computer-readable medium of claim 12, wherein the performance data is further based on a second signal group received via a second multiplexer.
14. The non-transitory computer-readable medium of claim 9, wherein the performance data is based on a first performance monitor associated with a first clock signal domain and a second performance monitor associated with a second clock signal domain. 15 . The non-transitory computer-readable medium of claim 9 , wherein the performance data is associated with a duration between the first signal and the second signal to sample the performance data of the plurality of engines. 16 . The non-transitory computer-readable medium of claim 15 , wherein the first signal coincides with a first context switch event associated with a first engine, and the second signal coincides with a second context switch event associated with the first engine.
17. A computer system comprising: a memory that stores software applications; as well as The processor, when executing the software application, is configured to perform the following steps: generating a first signal to sample performance data of a plurality of engines included in the processor; causing a plurality of performance monitors to capture the performance data based on the first signal; receiving the performance data from the one or more performance monitors, wherein the performance data includes a first identifier that identifies a performance monitor included in the plurality of performance monitors that generated the performance data; extracting a plurality of performance data subsets from the performance data based on a plurality of second identifiers included in the performance data, wherein a first performance data subset of the plurality of performance data includes the first identifier, wherein each of the plurality of performance data subsets corresponds to a different second identifier included in the plurality of second identifiers, and wherein each of the plurality of second identifiers corresponds to a different engine included in the plurality of engines; as well as Each of the plurality of performance data subsets is stored in a different data store among a plurality of data stores, wherein for each of the plurality of performance data subsets, the different data store among the plurality of data stores is determined based on the second identifier included in the performance data subset, wherein each data store included in the plurality of data stores is isolated from access by other engines included in the plurality of engines except for a corresponding engine, and wherein each of the plurality of engines accesses a corresponding data store among the plurality of data stores via a different virtual address space.
18. The system of claim 17, wherein the processor executes a plurality of virtual machines, and further comprising: determining that no virtual machine included in the plurality of virtual machines is utilizing a first circuit subportion included in the processor; as well as A supply voltage associated with the first circuit subsection is reduced.
19. The system of claim 17, wherein each circuit subsection included in the plurality of circuit subsections is associated with a different engine included in the plurality of engines, and further comprising: determining that a first circuit subsection included in the plurality of circuit subsections consumes more power than each other circuit subsection included in the plurality of circuit subsections; as well as The frequency of a clock signal associated with the first circuit subsection is reduced.
Citation Information
Patent Citations
Method and apparatus for power management of a processor in a virtual environment
US20130155073A1
System, method, and computer program product for collecting execution statistics for graphics processing unit workloads
US20150355996A1
Computing in parallel processing environments
US8738860B1