Graphics processor thread mid-term preemption
By introducing a hardware-supported thread preemption mechanism in the graphics processor, the waiting time overhead problem caused by software-assisted preemption in the prior art is solved, and execution efficiency and hardware utilization are improved.
Patent Information
- Application Number
- CN202411652005.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2024-01-17
- Filing Date
- 2024-11-19
- Publication Date
- 2025-07-18
AI Technical Summary
The existing graphics processors rely on software assistance in the mid-term preemption of threads, resulting in large wait time overhead and affecting execution efficiency.
By introducing a hardware-supported thread preemption mechanism in the graphics processor, the graphics processor's thread dispatch hardware is used to save and restore thread state and avoid software intervention.
It realizes efficient mid-thread preemption without affecting performance, reduces waiting time overhead and improves hardware utilization.
Smart Images

Figure CN120339032A_ABST
Abstract
Description
Technical Field The present disclosure generally relates to data processing via a graphics processor, and more particularly to methods for enabling mid-thread preemption within a general purpose graphics processor. Background Art Hardware utilization can be improved by increasing the granularity at which preemption can occur. Instruction-level preemption has been performed in a graphics processor using software-assisted context switching. However, software-assisted context switching has latency overheads that reduce the performance and efficiency of performing instruction-level preemption. Brief Description of the Drawings Embodiments described herein are illustrated by way of example and not limitation in the figures of the accompanying drawings, in which like reference numerals indicate like elements and in which: Figure 1 is a block diagram of a computer system configured to implement one or more aspects of the embodiments described herein; Figures 2A - 2D illustrates a parallel processor component; Figures 3A - 3C is a block diagram of a graphics multiprocessor and a multi-processor based GPU; Figures 4A - 4F illustrates an exemplary architecture in which multiple GPUs are communicatively coupled to multiple multi-core processors; Figure 5 illustrates a graphics processing pipeline; Figure 6 illustrates a machine learning software stack; Figure 7 illustrates a general purpose graphics processing unit; Figure 8 illustrates a multi-GPU computing system; Figures 9A - 9B illustrates layers of an exemplary deep neural network; Figure 10 illustrates an exemplary recurrent neural network; Figure 11 illustrates training and deployment of a deep neural network; Figure 12A is a block diagram illustrating distributed learning; Figure 12B is a block diagram illustrating a programmable network interface and a data processing unit; Figure 13 illustrates an exemplary inference system on a chip (SOC) adapted to perform inference using a trained model; Figure 14 is a block diagram of a processing system; Figures 15A - 15CIllustrated computing system and graphics processor; Figures 16A - 16C Block diagram illustrating additional graphics processor and compute accelerator architectures; Figure 17 Block diagram of the graphics processing engine of a graphics processor; Figures 18A - 18C Illustrated thread execution logic including an array of processing elements employed in a graphics processor core; Figure 19 Illustrated die of a multi-die processor according to an embodiment; Figure 20 Block diagram illustrating a graphics processor instruction format; Figure 21 Block diagram of an additional graphics processor architecture; Figures 22A - 22B Illustrated graphics processor command format and command sequence; Figure 23 Illustrated exemplary graphics software architecture for a data processing system; Figure 24A Block diagram illustrating an IP core development system; Figure 24B Illustrated cross-sectional side view of an integrated circuit package component; Figure 24C Illustrated package component including a hardware logic die connecting multiple units to a substrate (e.g., a base die); Figure 24D Illustrated package component including interchangeable dies; Figure 25 Block diagram illustrating an exemplary system-on-chip integrated circuit; Figures 26A - 26B Block diagram illustrating an exemplary graphics processor for use within a SoC; Figure 27 Block diagram of a data processing system according to an embodiment; Figure 28 Illustrated processing resource architecture according to an embodiment; Figure 29 Block diagram of a system including a GPGPU device according to an embodiment; Figure 30 Illustration of a system for dispatching a thread group to processing resources according to an embodiment; Figure 31 Illustration of a system for facilitating thread dispatching and execution on a graphics processor according to an embodiment; Figure 32 Shows a TSB for storing status information of multiple sub-slices according to an embodiment; Figure 33Shows an ODB for storing the state of an over-dispatched thread group (ODTG) according to an embodiment; Figure 34 Shows high-level operations of in-thread preemption based on a walker according to an embodiment; Figure 35 Illustrates a method for performing in-thread preemption of a thread group executed by a graphics processor according to an embodiment; Figure 36 Illustrates a method for dispatching and preempting over-dispatched threads; Figure 37 Is a block diagram of a computing device including a graphics processor according to an embodiment. Detailed Description Current parallel graphics data processing includes systems and methods developed to perform specific operations on graphics data, such as, for example, linear interpolation, tessellation, rasterization, texture mapping, depth testing, etc. Traditionally, graphics processors have used fixed-function computing units to process graphics data. However, recently, multiple parts of the graphics processor have been made programmable, enabling such processors to support a wider variety of operations for processing vertex data and fragment data. Typically, to further improve performance, graphics processors implement processing techniques such as pipelining, which attempt to process as much graphics data as possible in parallel through different parts of the graphics pipeline. A parallel graphics processor with a single-instruction, multiple-thread (SIMT) architecture is designed to maximize the amount of parallel processing in the graphics pipeline. In the SIMT architecture, parallel thread groups attempt to execute program instructions synchronously together as frequently as possible to improve processing efficiency. A graphics processing unit (GPU) is communicatively coupled to a host / processor core to accelerate, for example, graphics operations, machine learning operations, pattern analysis operations, and / or various general-purpose GPU (GPGPU) functions. The GPU can be communicatively coupled to the host processor / core via a bus or another interconnect (e.g., a high-speed interconnect such as PCIe or NVLink). Alternatively, the GPU can be integrated on the same package or die as the core and communicatively coupled to the core via an internal processor bus / interconnect (i.e., within the package or die). Regardless of the manner in which the GPU is connected, the processor core can allocate work to the GPU in the form of a sequence of commands / instructions contained in a work descriptor. The GPU then uses dedicated circuitry / logic to efficiently process these commands / instructions. In the following description, numerous specific details are set forth in order to provide a more thorough understanding. However, it will be apparent to one of ordinary skill in the art that embodiments described herein may be practiced without one or more of these specific details. In other instances, well-known features have not been described so as not to obscure the details of the current embodiments. GPU preemption enables a workload being executed in the GPU to replace another workload. Preemption can occur at various granularities, including command-buffer-level preemption, command-level preemption, thread-group-level preemption, thread-level preemption, and mid-thread preemption at the instruction level. To perform mid-thread preemption, the execution state of the thread is saved in a manner that allows the hardware to correctly resume at the exact instruction where the thread was preempted. Mid-thread preemption (MTP) in the GPU has previously been performed with software assistance. However, relying on software assistance introduces latency into the preemption process. Using the techniques described herein, the GPU hardware can be configured to save the thread state without requiring intervention from the graphics driver software to perform a software-assisted context switch. System Overview Figure 1FIG. is a block diagram of a computing system 100 configured to implement one or more aspects of the embodiments described herein. The computing system 100 includes a processing subsystem 101 having one or more processors 102 and a system memory 104 communicating via an interconnect path that may include a memory hub 105. The memory hub 105 may be a separate component within a chipset component or may be integrated within one or more of the processors 102. The memory hub 105 is coupled to an I / O subsystem 111 via a communication link 106. The I / O subsystem 111 includes an I / O hub 107 that enables the computing system 100 to receive input from one or more input devices 108. In addition, the I / O hub 107 enables a display controller (which may be included within one or more of the processors 102) to provide output to one or more display devices 110A. In one embodiment, one or more of the display devices 110A coupled to the I / O hub 107 may include a local, internal, or embedded display device. The processing subsystem 101 includes, for example, one or more parallel processors 112 coupled to the memory hub 105 via a bus or other communication link 113. The communication link 113 may be one of any number of standard-based communication link technologies or protocols, such as, but not limited to, PCI Express (PCIe), or may be a vendor-specific communication interface or fabric. The one or more parallel processors 112 may form a parallel or vector processing system within a computing complex that may include a large number of processing cores and / or processing clusters, such as, for example, a many integrated core (MIC) processor. For example, the one or more parallel processors 112 form a graphics processing subsystem that may output pixels to a display device of one of the one or more display devices 110A coupled via the I / O hub 107. The one or more parallel processors 112 may also include a display controller and a display interface (not shown) for enabling a direct connection to one or more display devices 110B. Within the I / O subsystem 111, the system storage unit 114 may be connected to the I / O hub 107 to provide a storage mechanism for the computing system 100. The I / O switch 116 may be used to provide an interface mechanism to enable connections between the I / O hub 107 and other components, such as network adapter 118 and / or wireless network adapter 119 that may be integrated into the platform, and various other devices that may be added via one or more plug-in devices 120. The (one or more) plug-in devices 120 may also include, for example, one or more external graphics processor devices, graphics cards, and / or computing accelerators. The network adapter 118 may be an Ethernet adapter or another wired network adapter. The wireless network adapter 119 may include one or more of the following: Wi-Fi, Bluetooth, near field communication (NFC), or other network devices including one or more wireless radio devices. The computing system 100 may include other components not explicitly shown, including USB or other port connections, optical storage drives, video capture devices, etc., which may also be connected to the I / O hub 107. The Figure 1 communication paths interconnecting the various components therein may be implemented using any suitable protocol, such as a PCI (Peripheral Component Interconnect)-based protocol (e.g., PCI Express), or any other bus or point-to-point communication interface and / or (one or more) protocols, such as NVLink high-speed interconnect, Compute Express Link TM (ComputeExpress Link TM , CXL TM)(e.g., CXL.mem), Infinity Fabric (IF), Ethernet (IEEE 802.3), remote direct memory access (RDMA), InfiniBand, Internet Wide Area RDMA Protocol (iWARP), Transmission Control Protocol (TCP), User Datagram Protocol (UDP), quick UDP Internet Connections (QUIC), RDMA over Converged Ethernet (RoCE), Intel QuickPath Interconnect (QPI), Intel Ultra Path Interconnect (UPI), Intel On-Chip System Fabric (IOSF), Omnipath, HyperTransport, Advanced Microcontroller Bus Architecture (AMBA) interconnect, OpenCAPI, Gen-Z, Cache Coherent Interconnect for Accelerators (CCIX), 3GPP Long Term Evolution (LTE) (4G), 3GPP 5G and its variants, or wired or wireless interconnect protocols known in the art. In some examples, protocols such as non-volatile memory express over Fabrics (NVMe-oF) or NVMe may be used to copy or store data to the virtualized storage node. One or more parallel processors 112 may include circuitry optimized for graphics and video processing (including, for example, video output circuitry) and form a graphics processing unit (GPU). Alternatively or additionally, as described in more detail herein, one or more parallel processors 112 may include circuitry optimized for general purpose processing while retaining the underlying computational architecture. The components of computing system 100 may be integrated on a single integrated circuit with one or more other system elements. For example, one or more parallel processors 112, memory hub 105, processor(s) 102, and I / O hub 107 may be integrated into a system-on-chip (SoC) integrated circuit. Alternatively, the components of computing system 100 may be integrated into a single package to form a system-in-package configuration. In one embodiment, at least some of the components of computing system 100 may be integrated into a multi-chip module (MCM), which may be interconnected with other multi-chip modules into a modular computing system. It will be appreciated that the computing system 100 shown herein is illustrative, and variations and modifications are possible. The connection topology may be modified as needed, including the number and arrangement of bridges, the number of processor(s) 102, and the number of parallel processor(s) 112. For example, system memory 104 may be connected directly to processor(s) 102 rather than through a bridge, while other devices communicate with system memory 104 via memory hub 105 and processor(s) 102. In other alternative topologies, parallel processor(s) 112 are connected to I / O hub 107 or directly to one of processor(s) 102 rather than to memory hub 105. In other embodiments, I / O hub 107 and memory hub 105 may be integrated into a single chip. It is also possible that two or more processor sets 102 are attached via multiple sockets, which may be coupled to two or more instances of parallel processor(s) 112. Some of the specific components shown herein are optional and may not be included in all implementations of computing system 100. For example, any number of plug-in cards or peripheral devices may be supported, or some components may be eliminated. Additionally, some architectures may use different terms for components similar to those illustrated Figure 1 herein. For example, memory hub 105 may be referred to as a north bridge in some architectures, while I / O hub 107 may be referred to as a south bridge. Figure 2AFIG. 200 shows a parallel processor. The parallel processor 200 can be a GPU, GPGPU, etc. as described herein. Various components of the parallel processor 200 can be implemented using one or more integrated circuit devices, such as programmable processors, application specific integrated circuits (ASICs), or field programmable gate arrays (FPGAs). The illustrated parallel processor 200 can be one or more of the parallel processors 112 shown in Figure 1 . Figure 1 One or more of the (one or more) parallel processors 112 shown in . The parallel processor 200 includes a parallel processing unit 202. The parallel processing unit includes an I / O unit 204 that enables communication with other devices, including other instances of the parallel processing unit 202. The I / O unit 204 can be directly connected to other devices. For example, the I / O unit 204 is connected to other devices via a hub or switch interface, such as a memory hub 105. The connection between the memory hub 105 and the I / O unit 204 forms a communication link 113. Within the parallel processing unit 202, the I / O unit 204 is connected to a host interface 206 and a memory crossbar 216, where the host interface 206 receives commands related to performing processing operations, and the memory crossbar 216 receives commands related to performing memory operations. When the host interface 206 receives a command buffer via the I / O unit 204, the host interface 206 can direct the work operations for executing those commands to a front end 208. In one embodiment, the front end 208 is coupled to a scheduler 210 that is configured to distribute commands or other work items to an array of processing clusters 212. The scheduler 210 ensures that the array of processing clusters 212 is properly configured and in an active state before tasks are distributed to the processing clusters within the array of processing clusters 212. The scheduler 210 can be implemented via firmware logic executed on a microcontroller. The scheduler 210 implemented by the microcontroller can be configured to perform complex scheduling and work distribution operations at both a coarse-grained and fine-grained level, enabling fast preemption and context switching of threads executing on the array of processing clusters 212. Preferably, the host software can authenticate the workload scheduled on the array of processing clusters 212 via one of a plurality of graphics processing doorbells. In other examples, polling for new workloads or interrupts can be used to identify or indicate the availability of work to be performed. The workload can then be automatically distributed across the array of processing clusters 212 by the scheduler 210 logic within the scheduler microcontroller. The processing cluster array 212 may include up to "N" processing clusters (e.g., cluster 214A, cluster 214B to cluster 214N). Each of the clusters 214A - 214N in the processing cluster array 212 may execute a large number of concurrent threads. The scheduler 210 may use various scheduling and / or work distribution algorithms to allocate work to the clusters 214A - 214N in the processing cluster array 212, and these scheduling and / or work distribution algorithms may vary depending on the workload generated for each type of program or computation. Scheduling may be handled dynamically by the scheduler 210, or may be assisted in part by compiler logic during the compilation of the program logic configured to be executed by the processing cluster array 212. Optionally, different clusters 214A - 214N in the processing cluster array 212 may be assigned to process different types of programs or to perform different types of computations. The processing cluster array 212 may be configured to perform various types of parallel processing operations. For example, the processing cluster array 212 is configured to perform general - purpose parallel computing operations. For example, the processing cluster array 212 may include logic for performing processing tasks that include filtering of video and / or audio data, performing modeling operations including physical operations, and performing data transformation. The processing cluster array 212 is configured to perform parallel graphics processing operations. In such embodiments where the parallel processor 200 is configured to perform graphics processing operations, the processing cluster array 212 may include additional logic to support the execution of such graphics processing operations, including but not limited to texture sampling logic for performing texture operations, and tessellation logic and other vertex processing logic. Additionally, the processing cluster array 212 may be configured to execute graphics - processing - related shader programs, such as but not limited to vertex shaders, tessellation shaders, geometry shaders, and pixel shaders. The parallel processing unit 202 may transfer data from the system memory via the I / O unit 204 for processing. During processing, the transferred data may be stored in on - chip memory (e.g., parallel processor memory 222) during processing and then written back to the system memory. In embodiments where a parallel processing unit 202 is used to perform graphics processing, the scheduler 210 may be configured to divide the processing workload into tasks of approximately equal size to better enable the distribution of graphics processing operations to the multiple clusters 214A - 214N in the processing cluster array 212. In some of these embodiments, portions of the processing cluster array 212 may be configured to perform different types of processing. For example, a first portion may be configured to perform vertex shading and topology generation, a second portion may be configured to perform tessellation and geometry shading, and a third portion may be configured to perform pixel shading or other screen space operations to produce a rendered image for display. Intermediate data produced by one or more of the clusters 214A - 214N may be stored in a buffer to allow the intermediate data to be transferred between the clusters 214A - 214N for further processing. During operation, the processing cluster array 212 may receive processing tasks to be executed via the scheduler 210, which receives commands defining the processing tasks from the front end 208. For graphics processing operations, the processing tasks may include the data to be processed and indices of state parameters and commands defining how the data is to be processed (e.g., what program is to be executed), such as surface (patch) data, primitive data, vertex data, and / or pixel data. The scheduler 210 may be configured to fetch the index corresponding to the task, or may receive the index from the front end 208. The front end 208 may be configured to ensure that the processing cluster array 212 is configured in an active state before the workload specified by the incoming command buffer (e.g., batch buffer, push buffer, etc.) is initiated. Each instance of one or more instances of the parallel processing unit 202 may be coupled to the parallel processor memory 222. The parallel processor memory 222 may be accessed via a memory crossbar 216 that may receive memory requests from the processing cluster array 212 as well as the I / O unit 204. The memory crossbar 216 may access the parallel processor memory 222 via a memory interface 218. The memory interface 218 may include a plurality of partition units (e.g., partition unit 220A, partition unit 220B, up to partition unit 220N) each of which may be coupled to a portion (e.g., a memory cell) of the parallel processor memory 222. The number of partition units 220A - 220N may be configured to be equal to the number of memory cells such that the first partition unit 220A has a corresponding first memory cell 224A, the second partition unit 220B has a corresponding second memory cell 224B, and the Nth partition unit 220N has a corresponding Nth memory cell 224N. In other embodiments, the number of partition units 220A - 220N may not be equal to the number of memory devices. The memory cells 224A - 224N may include various types of memory devices including dynamic random-access memory (DRAM) or graphics random access memory such as synchronous graphics random access memory (SGRAM) including graphics double data rate (GDDR) memory. Optionally, the memory cells 224A - 224N may also include 3D stacked memory including but not limited to high bandwidth memory (HBM). Those skilled in the art will appreciate that the specific implementation of the memory cells 224A - 224N may vary and may be selected from one of a variety of conventional designs. Rendering targets such as frame buffers or texture maps may be stored across the memory cells 224A - 224N, allowing the partition units 220A - 220N to write portions of each rendering target in parallel to efficiently utilize the available bandwidth of the parallel processor memory 222. In some embodiments, local instances of the parallel processor memory 222 may be excluded in favor of a unified memory design that utilizes system memory in conjunction with local cache memory. Optionally, any one of the clusters 214A - 214N in the processing cluster array 212 has the ability to process data to be written to any one of the memory cells 224A - 224N within the parallel processor memory 222. The memory crossbar 216 can be configured to transfer the output of each cluster 214A - 214N to any of the partition units 220A - 220N or to another cluster 214A - 214N, and this other cluster 214A - 214N can perform additional processing operations on the output. Each cluster 214A - 214N can communicate with the memory interface 218 through the memory crossbar 216 to read from or write to various external memory devices. In one embodiment among the embodiments having the memory crossbar 216, the memory crossbar 216 has a connection to the memory interface 218 to communicate with the I / O unit 204 and has a connection to a local instance of the parallel processor memory 222, enabling the processing units within different processing clusters 214A - 214N to communicate with the system memory or other memories not local to the parallel processing unit 202. Generally, the memory crossbar 216 may be capable of using virtual channels to separate the traffic flow between the clusters 214A - 214N and the partition units 220A - 220N. Although a single instance of the parallel processing unit 202 is illustrated within the parallel processor 200, any number of instances of the parallel processing unit 202 may be included. For example, multiple instances of the parallel processing unit 202 may be provided on a single plug - in card, or multiple plug - in cards may be interconnected. For example, the parallel processor 200 may be a plug - in device, such as Figure 1 the plug - in device 120, which may be a graphics card (such as a discrete graphics card including one or more GPUs, one or more memory devices, and device - to - device or network or fabric interfaces). Different instances of the parallel processing unit 202 can be configured to interoperate even if the different instances have different numbers of processing cores, different amounts of local parallel processor memory, and / or other configuration differences. Optionally, some instances of the parallel processing unit 202 may include higher - precision floating - point units relative to other instances. Systems containing one or more instances of the parallel processing unit 202 or the parallel processor 200 can be implemented in various configurations and form factors, including but not limited to, desktop computers, laptop computers, or handheld personal computers, servers, workstations, game consoles, and / or embedded systems. The orchestrator can use one or more of the following to form composite nodes for workload execution: decomposed processor resources, cache resources, memory resources, storage resources, and networking resources. In one embodiment, the parallel processing unit 202 may be partitioned into multiple instances. Those multiple instances may be configured to execute workloads associated with different clients in an isolated manner, such that a predetermined quality of service is provided for each client. For example, each cluster 214A - 214N may be partitioned and isolated from other clusters, allowing the array of processing clusters 212 to be divided into multiple computing partitions or instances. In such a configuration, workloads executed on isolated partitions are protected from errors or inaccuracies associated with different workloads executed on different partitions. The partitioning units 220A - 220N may be configured to enable dedicated and / or isolated paths to the memories of the clusters 214A - 214N associated with the respective computing partitions. This data path isolation enables the computing resources within a partition to communicate with one or more assigned memory units 224A - 224N without being disturbed by the activities of other partitions. Figure 2B is a block diagram of the partitioning unit 220. The partitioning unit 220 may be an instance of one of the partitioning units 220A - 220N of Figure 2A . As illustrated, the partitioning unit 220 includes an L2 cache 221, a frame buffer interface 225, and a ROP 226 (raster operation unit). The L2 cache 221 is a read / write cache configured to perform load and store operations received from the memory crossbar 216 and the ROP 226. Read misses and urgent write-back requests are output by the L2 cache 221 to the frame buffer interface 225 for processing. Updates may also be sent via the frame buffer interface 225 to the frame buffer for processing. In one embodiment, the frame buffer interface 225 interfaces with the memory unit 224 of the memory units 224A - 224N within the Figure 2A parallel processor memory 222. The partitioning unit 220 may additionally or alternatively interface with a memory unit of the memory units in the parallel processor memory via a memory controller (not shown). In a graphics application, the ROP 226 is a processing unit that performs raster operations such as stencil, z-test, blending, etc. The ROP 226 then outputs the processed graphics data, which is stored in the graphics memory. In some embodiments, the ROP 226 includes or is coupled to a codec (CODEC) 227, which includes compression logic for compressing depth or color data written to the memory or L2 cache 221 and decompressing depth or color data read from the memory or L2 cache 221. The compression logic can be lossless compression logic that utilizes one or more of a variety of compression algorithms. The type of compression performed by the CODEC 227 can vary based on the statistical characteristics of the data to be compressed. For example, in one embodiment, delta color compression is performed on a per-tile basis for depth and color data. In one embodiment, the CODEC 227 includes compression and decompression logic that can compress and decompress computational data associated with machine learning operations. The CODEC 227 can, for example, compress sparse matrix data for sparse machine learning operations. The CODEC 227 can also compress sparse matrix data encoded in a sparse matrix format (e.g., coordinate list encoding (COO), compressed sparse row (CSR), compressed sparse column (CSC), etc.) to generate compressed and encoded sparse matrix data. The compressed and encoded sparse matrix data can be decompressed and / or decoded before being processed by a processing element, or the processing element can be configured to consume the compressed, encoded, or compressed and encoded data for processing. The ROP 226 can be included within each processing cluster (e.g., Figure 2A clusters 214A - 214N) rather than being included within the partitioning unit 220. In such embodiments, read and write requests for pixel data rather than pixel fragment data are transmitted through the memory crossbar 216. The processed graphics data can be displayed on a display device (such as, Figure 1 one of the one or more display devices 110A - 110B), routed for further processing by one or more processors 102, or routed for further processing by Figure 2A one of the processing entities within the parallel processor 200. Figure 2C is a block diagram of a processing cluster 214 within a parallel processing unit. For example, the processing cluster is Figure 2AAn instance of one of the processing clusters 214A - 214N in the processing cluster 214. The processing cluster 214 can be configured to execute many threads in parallel, where the term "thread" refers to an instance of a particular program executed on a particular set of input data. Optionally, single-instruction, multiple-data (SIMD) instruction-issuing techniques can be used to support the concurrent execution of a large number of threads without providing multiple independent instruction units. Alternatively, single-instruction, multiple-thread (SIMT) techniques can be used to support the concurrent execution of a large number of generally synchronous threads using a common instruction unit configured to issue instructions to a set of processing engines within each processing cluster in the processing cluster. Different from SIMD execution mechanisms where all processing engines typically execute the same instruction, SIMT execution allows different threads to more easily follow divergent execution paths through a given thread program. Those skilled in the art will understand that the SIMD processing mechanism represents a functional subset of the SIMT processing mechanism. The operation of the processing cluster 214 can be controlled via a pipeline manager 232 that distributes processing tasks to the SIMT parallel processors. The pipeline manager 232 receives instructions from Figure 2A the scheduler 210 and manages the execution of those instructions via the graphics multiprocessor 234 and / or the texture unit 236. The illustrated graphics multiprocessor 234 is an exemplary instance of a SIMT parallel processor. However, various types of SIMT parallel processors of different architectures can be included within the processing cluster 214. One or more instances of the graphics multiprocessor 234 can be included within the processing cluster 214. The graphics multiprocessor 234 can process data, and a data crossbar 240 can be used to distribute the processed data to one of a plurality of possible destinations, including facilitating the exchange of data between graphics multiprocessors within the processing cluster 214. The pipeline manager 232 can facilitate the distribution of the processed data by specifying the destination for the processed data to be distributed via the data crossbar 240. Each graphics multiprocessor 234 within the processing cluster 214 can include the same set of functional execution logic (e.g., arithmetic logic units, load-store units, etc.). The functional execution logic can be configured in a pipelined manner, in which new instructions can be issued before the previous instructions are completed. The functional execution logic supports various operations, including integer and floating-point arithmetic, comparison operations, boolean operations, bit shifts, and the calculation of various algebraic functions. Different operations can be executed using the same functional unit hardware, and any combination of functional units can exist. The instructions transmitted to processing cluster 214 constitute a thread. A set of threads executed across a collection of parallel processing engines is a thread group. The thread group executes the same program on different input data. Each thread within the thread group may be assigned to a different processing engine within graphics multiprocessor 234. The thread group may include fewer threads than the number of processing engines within graphics multiprocessor 234. When the thread group includes fewer threads than the number of processing engines, one or more of the processing engines may be idle during the cycle in which the thread group is being processed. The thread group may also include more threads than the number of processing engines within graphics multiprocessor 234. When the thread group includes more threads than the number of processing engines within graphics multiprocessor 234, processing may be performed in consecutive clock cycles. Optionally, multiple thread groups may be executed concurrently on graphics multiprocessor 234. Graphics multiprocessor 234 may include an internal cache memory to perform load and store operations. Optionally, graphics multiprocessor 234 may forego the internal cache and use the cache memory within processing cluster 214 (e.g., level 1 (L1) cache 248). Each graphics multiprocessor 234 also has access to a second level (L2) cache within a partitioning unit (e.g., Figure 2A partitioning units 220A - 220N), which are shared among all processing clusters 214 and may be used to transfer data between threads. Graphics multiprocessor 234 may also access off-chip global memory, which may include one or more of local parallel processor memory and / or system memory. Any memory external to parallel processing unit 202 may be used as global memory. In embodiments in which processing cluster 214 includes multiple instances of graphics multiprocessor 234, common instructions and data may be shared and stored in L1 cache 248. Each processing cluster 214 may include an MMU 245 (memory management unit) configured to map virtual addresses to physical addresses. In other embodiments, one or more instances of MMU 245 may reside in Figure 2Awithin the memory interface 218. The MMU 245 includes a set of page table entries (PTEs) for mapping virtual addresses to physical addresses of the die, and optionally includes a cache line index. The MMU 245 may include a translation lookaside buffer (TLB) or cache that may reside within the graphics multiprocessor 234 or the L1 cache 248 of the processing cluster 214. The physical address is processed to distribute surface data access locality, thereby allowing efficient request interleaving among the partition units. The cache line index may be used to determine whether a request to a cache line is a hit or a miss. In graphics and computing applications, the processing cluster 214 may be configured such that each graphics multiprocessor 234 is coupled to a texture unit 236 for performing texture mapping operations, e.g., determining texture sample locations, reading texture data, and filtering texture data. The texture data is read from an internal texture L1 cache (not shown), or in some embodiments, from the L1 cache within the graphics multiprocessor 234, and fetched from the L2 cache, local parallel processor memory, or system memory as needed. Each graphics multiprocessor 234 outputs the processed tasks to the data crossbar 240 to provide the processed tasks to another processing cluster 214 for further processing, or stores the processed tasks in the L2 cache, local parallel processor memory, or system memory via the memory crossbar 216. The preROP 242 (pre-raster operations unit) is configured to receive data from the graphics multiprocessor 234 and direct the data to the ROP units, which may be located together with partition units (e.g., Figure 2A partition units 220A-220N as described herein). The preROP 242 unit may perform optimizations for color blending, organize pixel color data, and perform address translation. It will be appreciated that the core architecture described herein is illustrative, and variations and modifications are possible. Any number of processing units (e.g., graphics multiprocessors 234, texture units 236, preROP 242, etc.) may be included within the processing cluster 214. Further, although only one processing cluster 214 is shown, the parallel processing unit as described herein may include any number of instances of the processing cluster 214. Optionally, each processing cluster 214 may be configured to operate independently of other processing clusters 214 using separate and distinct processing units, L1 caches, L2 caches, etc. Figure 2DAn example of a graphics multiprocessor 234 is shown, where the graphics multiprocessor 234 is coupled to the pipeline manager 232 of the processing cluster 214. The graphics multiprocessor 234 has an execution pipeline that includes, but is not limited to, an instruction cache 252, an instruction unit 254, an address mapping unit 256, a register file 258, one or more general-purpose graphics processing unit (GPGPU) cores 262, and one or more load / store units 266. The GPGPU cores 262 and the load / store units 266 are coupled to a cache memory 272 and a shared memory 270 via a memory and cache interconnect 268. The graphics multiprocessor 234 may additionally include tensor and / or ray tracing cores 263, which include hardware logic for accelerating matrix and / or ray tracing operations. The instruction cache 252 may receive a stream of instructions to be executed from the pipeline manager 232. The instructions are cached in the instruction cache 252 and dispatched for execution by the instruction unit 254. The instruction unit 254 may dispatch the instructions as a thread group (e.g., a warp), where each thread in the thread group is assigned to a different execution unit within the GPGPU core 262. Instructions may access any one of a local address space, a shared address space, or a global address space by specifying an address within a unified address space. The address mapping unit 256 may be used to translate an address in the unified address space into a different memory address accessible by the load / store unit 266. The register file 258 provides a collection of registers for the functional units of the graphics multiprocessor 234. The register file 258 provides temporary storage for the operands of the data paths connected to the functional units (e.g., the GPGPU cores 262, the load / store units 266) of the graphics multiprocessor 234. The register file 258 may be partitioned among each of the functional units such that each functional unit is assigned a dedicated portion of the register file 258. For example, the register file 258 may be partitioned among different groups of units executed by the graphics multiprocessor 234. The GPGPU cores 262 may each include a floating point unit (FPU) and / or an integer arithmetic logic unit (ALU) for executing instructions of the graphics multiprocessors 234. In some implementations, the GPGPU cores 262 may include hardware logic that would otherwise reside within the tensor and / or ray tracing cores 263. The GPGPU cores 262 may be architecturally similar or architecturally different. For example and in one embodiment, a first portion of the GPGPU cores 262 includes a single-precision FPU and an integer ALU, while a second portion of the GPGPU cores includes a double-precision FPU. Optionally, the FPU may implement the IEEE 754-2008 standard for floating-point arithmetic or enable variable-precision floating-point arithmetic. The graphics multiprocessors 234 may additionally include one or more fixed-function or special-function units for performing specific functions, such as copy rectangle or pixel blend operations. One or more of the GPGPU cores may also include fixed-function or special-function logic. The GPGPU cores 262 may include SIMD logic capable of executing a single instruction on multiple sets of data. Optionally, the GPGPU cores 262 may physically execute SIMD4, SIMD8, and SIMD16 instructions and logically execute SIMD1, SIMD2, and SIMD32 instructions. The SIMD instructions for the GPGPU cores may be generated at compile time by a shader compiler or automatically generated when executing a program written and compiled for a single program multiple data (SPMD) or SIMT architecture. Multiple threads of a program configured for the SIMT execution model may be executed via a single SIMD instruction. For example and in one embodiment, eight SIMT threads that perform the same or similar operations may be executed in parallel via a single SIMD8 logic unit. The memory and cache interconnect 268 is an interconnect network that connects each of the functional units in the graphics multiprocessor 234 to the register file 258 and to the shared memory 270. For example, the memory and cache interconnect 268 is a crossbar interconnect that allows the load / store unit 266 to implement load and store operations between the shared memory 270 and the register file 258. The register file 258 can operate at the same frequency as the GPGPU core 262, so data transfer between the GPGPU core 262 and the register file 258 is very low latency. The shared memory 270 can be used to enable communication between threads executing on the functional units within the graphics multiprocessor 234. The cache memory 272 can be used as a data cache, for example, to cache texture data passed between the functional units and the texture unit 236. The shared memory 270 can also be used as a managed cached program. The shared memory 270 and the cache memory 272 can be coupled to the data crossbar 240 to enable communication with other components of the processing cluster. Threads executing on the GPGPU core 262 can also programmatically store data in the shared memory in addition to the automatically cached data stored in the cache memory 272. Figures 3A - 3C Illustrates an additional graphics multiprocessor according to an embodiment. Figures 3A - 3B Illustrates graphics multiprocessors 325, 350, which are related to the graphics multiprocessor 234 of Figure 2C and can be used in place of one of those graphics multiprocessors. Thus, any disclosure of a feature in connection with the graphics multiprocessor 234 herein also discloses the corresponding combination with the graphics multiprocessor(s) 325, 350, but is not limited thereto. Figure 3C Illustrates a graphics processing unit (GPU) 380 that includes a collection of dedicated graphics processing resources arranged as multi-core groups 365A - 365N, which correspond to the graphics multiprocessors 325, 350. The illustrated graphics multiprocessors 325, 350 and multi-core groups 365A - 365N can be streaming multiprocessors (SMs) capable of simultaneously executing a large number of execution threads. Figure 3A The graphics multiprocessor 325 of Figure 2DMultiple additional instances of execution resource units of the graphics multiprocessor 234. For example, the graphics multiprocessor 325 may include multiple instances of instruction units 332A - 332B, register heaps 334A - 334B, and (one or more) texture units 344A - 344B. The graphics multiprocessor 325 also includes multiple sets of graphics or computing execution units (e.g., GPGPU cores 336A - 336B, tensor cores 337A - 337B, ray tracing cores 338A - 338B) and multiple sets of load / store units 340A - 340B. The execution resource units have a common instruction cache 330, texture and / or data cache memory 342, and shared memory 346. Each component can communicate via the interconnect structure 327. The interconnect structure 327 may include one or more crossbars to enable communication between the components of the graphics multiprocessor 325. The interconnect structure 327 can be a separate, high-speed network structure layer on which each component of the graphics multiprocessor 325 is stacked. The components of the graphics multiprocessor 325 communicate with remote components via the interconnect structure 327. For example, cores 336A - 336B, 337A - 337B, and 338A - 338B can each communicate with the shared memory 346 via the interconnect structure 327. The interconnect structure 327 can arbitrate the communication within the graphics multiprocessor 325 to ensure fair bandwidth allocation between components. Figure 3B The graphics multiprocessor 350 includes multiple sets of execution resources 356A - 356D, where, as Figure 2D and Figure 3A illustrated, each set of execution resources includes multiple instruction units, register heaps, GPGPU cores, and load / store units. The execution resources 356A - 356D can work in cooperation with (one or more) texture units 360A - 360D for texture operations while sharing the instruction cache 354 and the shared memory 353. For example, the execution resources 356A - 356D can share the instruction cache 354, the shared memory 353, and multiple instances of texture and / or data cache memories 358A - 358B. Each component can communicate via an interconnect structure 352 similar to the interconnect structure 327 of Figure 3A . Those skilled in the art will understand that Figure 1 , Figures 2A - 2D and Figures 3A - 3BThe architecture described herein is descriptive and not restrictive in terms of the scope of the current embodiments. Thus, the techniques described herein may be implemented on any suitably configured processing unit without departing from the scope of the embodiments described herein, including but not limited to: one or more mobile application processors; one or more desktop or server central processing units (CPUs), including multi-core CPUs; one or more parallel processor units such as, Figure 2A the parallel processing unit 202 as well as one or more graphics processors or dedicated processing units. The parallel processors or GPGPUs described herein may be communicatively coupled to a host / processor core to accelerate graphics operations, machine learning operations, pattern analysis operations, and various general-purpose GPU (GPGPU) functions. The GPU may be communicatively coupled to the host processor / core via a bus or another interconnect (e.g., a high-speed interconnect such as PCIe, NVLink, or other known protocols, standardized protocols, or proprietary protocols). In other embodiments, the GPU may be integrated on the same package or die as the core and communicatively coupled to the core via an internal processor bus / interconnect (i.e., within the package or die). Regardless of how the GPU is connected, the processor core may allocate work to the GPU in the form of a sequence of commands / instructions contained in a work descriptor. The GPU then uses dedicated circuitry / logic to efficiently process these commands / instructions. Figure 3C Illustrated is a graphics processing unit (GPU) 380 that includes a collection of dedicated graphics processing resources arranged as multi-core groups 365A - 365N. While details are provided for only a single multi-core group 365A, it will be appreciated that the other multi-core groups 365B - 365N may be equipped with the same or similar collections of graphics processing resources. The details described with respect to multi-core groups 365A - 365N may also apply to any graphics multiprocessor 234, 325, 350 described herein. As illustrated, multi-core group 365A may include a collection of graphics cores 370, a collection of tensor cores 371, and a collection of ray tracing cores 372. A scheduler / dispatcher 368 schedules and dispatches graphics threads for execution on the respective cores 370, 371, 372. A collection of register files 369 stores operand values used by the cores 370, 371, 372 when executing graphics threads. These register files may include, for example, integer registers for storing integer values, floating-point registers for storing floating-point values, vector registers for storing packed data elements (integer and / or floating-point data elements), and tile registers for storing tensor / matrix values. The tile registers may be implemented as a combined collection of vector registers. One or more combined first-level (L1) cache and shared memory units 373 store graphics data locally within each multi-core group 365A, such as texture data, vertex data, pixel data, light data, bounding volume data, etc. One or more texture units 374 can also be used to perform texture operations, such as texture mapping and sampling. A second-level (L2) cache 375 shared by all multi-core groups 365A - 365N or a subset of multi-core groups 365A - 365N stores graphics data and / or instructions for multiple concurrent graphics threads. As illustrated, the L2 cache 375 can be shared across multiple multi-core groups 365A - 365N. One or more memory controllers 367 couple the GPU 380 to a memory 366, which can be system memory (e.g., DRAM) and / or dedicated graphics memory (e.g., GDDR6 memory). Input / Output (I / O) circuitry 363 couples the GPU 380 to one or more I / O devices 362, such as a digital signal processor (DSP), a network controller, or a user input device. On-chip interconnects can be used to couple the I / O devices 362 to the GPU 380 and the memory 366. One or more I / O memory management units (IOMMUs) 364 of the I / O circuitry 363 directly couple the I / O devices 362 to the system memory 366. Optionally, the IOMMU 364 manages a plurality of page table sets for mapping virtual addresses to physical addresses in the system memory 366. The I / O devices 362, the (one or more) CPUs 361, and the (one or more) GPUs 380 can then share the same virtual address space. In one implementation of the IOMMU 364, the IOMMU 364 supports virtualization. In this case, the IOMMU 364 can manage a first page table set for mapping guest / graphics virtual addresses to guest / graphics physical addresses and a second page table set for mapping guest / graphics physical addresses to system / host physical addresses (e.g., within the system memory 366). The base addresses of each of the first page table set and the second page table set can be stored in control registers and swapped out during a context switch (e.g., such that a new context is provided access to the relevant page table sets). Although not illustrated in Figure 3C each of the cores 370, 371, 372, and / or multi-core groups 365A - 365N can include translation lookaside buffers (TLBs) for caching guest virtual to guest physical translations, guest physical to host physical translations, and guest virtual to host physical translations. (One or more) CPUs 361, GPUs 380, and I / O devices 362 may be integrated on a single semiconductor chip and / or chip package. The illustrated memory 366 may be integrated on the same chip or may be coupled to the memory controller 367 via an off-chip interface. In one implementation, the memory 366 includes GDDR6 memory that shares the same virtual address space as other physical system-level memories, but the basic principles described herein are not limited to this particular implementation. Tensor cores 371 may include multiple execution units specifically designed to perform matrix operations, which are the basic computational operations for performing deep learning operations. For example, synchronous matrix multiplication operations may be used for neural network training and inference. Tensor cores 371 may perform matrix processing using various operand precisions, including single-precision floating point (e.g., 32 bits), half-precision floating point (e.g., 16 bits), integer (16 bits), byte (8 bits), and nibble (4 bits). For example, neural network implementations extract features of each rendered scene, potentially combining details from multiple frames to build a high-quality final image. In deep learning implementations, schedulable parallel matrix multiplication work may be used for execution on tensor cores 371. Training of neural networks in particular requires a large number of matrix dot product operations. To handle the inner product formulation of an N x N x N matrix multiplication, tensor cores 371 may include at least N dot product processing elements. Before matrix multiplication begins, a complete matrix is loaded into the on-chip registers, and for each of the N loops, at least one column of the second matrix is loaded. For each loop, there are N dot products to be processed. Depending on the specific implementation, matrix elements can be stored with different precisions, including 16-bit words, 8-bit bytes (e.g., INT8), and 4-bit nibbles (e.g., INT4). Different precision modes can be specified for the tensor core 371 to ensure that the most efficient precision is used for different workloads (e.g., inference workloads, which can tolerate quantization down to bytes and nibbles). The supported formats additionally include 64-bit floating point (FP64) and non-IEEE floating point formats, such as the bfloat16 format (e.g., Brain floating point), a 16-bit floating point format with one sign bit, eight exponent bits, and eight significand bits (seven of which are explicitly stored). One embodiment includes support for a reduced-precision tensor floating point (TF32) mode that performs calculations using the range of FP32 (8 bits) and the precision of FP16 (10 bits). Reduced-precision TF32 operations can be performed on FP32 inputs with higher performance relative to FP32 and increased precision relative to FP16 and produce FP32 outputs. In one embodiment, one or more 8-bit floating point formats (FP8) are supported. In one embodiment, the tensor core 371 supports a sparse operation mode for matrices in which the vast majority of values are zero. The tensor core 371 includes support for sparse input matrices encoded in a sparse matrix representation (e.g., coordinate list encoding (COO), compressed sparse row (CSR), compressed sparse column (CSC), etc.). The tensor core 371 also includes support for a compressed sparse matrix representation in cases where the sparse matrix representation can be further compressed. Compressed matrix data, encoded matrix data, and / or compressed and encoded matrix data, along with associated compression and / or encoding metadata, can be read by the tensor core 371, and non-zero values can be extracted. For example, for a given input matrix A, non-zero values can be loaded from at least a portion of the compressed and / or encoded representation of matrix A. Based on the positions of the non-zero values in matrix A (which can be determined from the index or coordinate metadata associated with the non-zero values), the corresponding values in input matrix B can be loaded. Depending on the operation to be performed (e.g., multiplication), if the corresponding value is a zero value, the loading of the value from input matrix B can be bypassed. In one embodiment, the pairing of values for certain operations (such as multiplication operations) can be pre-scanned by the scheduler logic, and only operations between non-zero inputs are scheduled. Depending on the dimensions of matrices A and B and the operation to be performed, the output matrix C can be dense or sparse. In the case where the output matrix C is sparse and depending on the configuration of the tensor core 371, the output matrix C can be output in a compressed format, a sparse encoding, or a compressed sparse encoding. The ray tracing core 372 can accelerate ray tracing operations for both real-time ray tracing implementations and non-real-time ray tracing implementations. Specifically, the ray tracing core 372 can include a ray traversal / intersection circuitry component that uses a bounding volume hierarchy (BVH) to perform ray traversal and identify intersections between rays enclosed within the BVH volume and primitives. The ray tracing core 372 can also include circuitry components for performing depth testing and culling (e.g., using a Z-buffer or a similar arrangement). In one implementation, the ray tracing core 372 performs traversal and intersection operations in cooperation with the image denoising techniques described herein, at least part of which can be performed on the tensor core 371. For example, the tensor core 371 can implement a deep learning neural network to perform denoising of frames generated by the ray tracing core 372. However, the (one or more) CPUs 361, the graphics core 370, and / or the ray tracing core 372 can also implement all or part of the denoising and / or deep learning algorithms. In addition, as described above, a distributed approach for denoising can be adopted, where the GPU 380 is in a computing device coupled to other computing devices via a network or a high-speed interconnect. According to this distributed approach, the interconnected computing devices can share neural network learning / training data to improve the speed at which the entire system learns to perform denoising for different types of image frames and / or different graphics applications. The ray tracing core 372 can handle all BVH traversals and / or ray-primitive intersections, thus sparing the graphics core 370 from being overloaded with thousands of instructions per ray. For example, each ray tracing core 372 includes a first set of specialized circuitry components for performing bounding box tests (e.g., for traversal operations) and / or a second set of specialized circuitry components for performing ray-triangle intersection tests (e.g., intersecting traversed rays). Thus, for example, the multi-core group 365A can simply initiate a ray probe, and the ray tracing core 372 independently performs ray traversal and intersection and returns hit data (e.g., hit, no hit, multiple hits, etc.) to the thread context. When the ray tracing core 372 performs traversal and intersection operations, the other cores 370, 371 are freed up to perform other graphics or computing work. Optionally, each ray tracing core 372 can include a traversal unit for performing BVH test operations and / or an intersection unit for performing ray-primitive intersection tests. The intersection unit generates "hit", "no hit", or "multiple hits" responses, which the intersection unit provides to the appropriate thread. During traversal and intersection operations, the execution resources of other cores (e.g., the graphics core 370 and the tensor core 371) are freed up to perform other forms of graphics work. In an optional embodiment described below, a hybrid rasterization / ray tracing method is used in which work is distributed between the graphics core 370 and the ray tracing core 372. The ray tracing core 372 (and / or other cores 370, 371) may include hardware support for a ray tracing instruction set, such as Microsoft's DirectX Ray Tracing (DXR), which includes the DispatchRays command; and ray generation shaders, closest hit shaders, any hit shaders, and miss shaders, which are capable of assigning a unique set of shaders and textures to each object. Another ray tracing platform that may be supported by the ray tracing core 372, the graphics core 370, and the tensor core 371 is the Vulkan API (e.g., Vulkan version 1.1.85, or later versions). However, note that the basic principles described herein are not limited to any particular ray tracing ISA. Generally, the respective cores 372, 371, 370 may support a ray tracing instruction set that includes instructions / functions for one or more of the following: ray generation, closest hit, any hit, ray-primitive intersection, per-primitive and hierarchical bounding box construction, miss, traversal, and exception. More specifically, preferred embodiments include ray tracing instructions for performing one or more of the following functions: Light Generation —— Ray generation instructions may be executed for each pixel, sample, or other user-defined work assignment. Nearest Hit —— Closest hit instructions may be executed to locate the closest intersection of a ray with a primitive within a scene. Any Hit —— Any hit instructions identify multiple intersections between a ray and primitives within a scene, potentially identifying a new closest intersection. Intersection —— Intersection instructions perform ray-primitive intersection tests and output the results. Primitive - by - Primitive Bounding Box Construction —— This instruction builds a bounding box around a given primitive or set of primitives (e.g., when building a new BVH or other acceleration data structure). Miss —— Indicates that a ray misses all geometry within a scene or a specified region of the scene. Access —— Indicates the sub-volumes that a ray will traverse. Exception —— Includes various types of exception handlers (e.g., called for various error conditions). In one embodiment, the ray tracing core 372 may be adapted to accelerate general computing operations that may use computational techniques similar to ray intersection tests. A computational framework may be provided that enables shader programs to be compiled into low-level instructions and / or primitives for performing general computing operations via the ray tracing core. Exemplary computational problems that may benefit from computational operations performed on the ray tracing core 372 include computations involving the propagation of light beams, waves, rays, or particles within a coordinate space. Interactions associated with that propagation may be computed relative to geometries or meshes within the coordinate space. For example, computations associated with the propagation of electromagnetic signals through an environment may be accelerated via the use of instructions or primitives executed via the ray tracing core. Refraction and reflection of signals occurring through objects in the environment may be computed as a direct ray tracing simulation. The ray tracing core 372 may also be used to perform computations that are not directly similar to ray tracing. For example, the ray tracing core 372 may be used to accelerate mesh projection, mesh refinement, and volume sampling computations. General coordinate space computations may also be performed, such as nearest neighbor computations. For example, a set of points near a given point may be discovered by defining a bounding box around the given point in the coordinate space. Subsequently, the BVH and ray tracing logic within the ray tracing core 372 may be used to determine the set of point intersections within the bounding box. The intersections constitute the origin and the nearest neighbors of that origin. Computations performed using the ray tracing core 372 may be performed in parallel with computations performed on the graphics core 372 and the tensor core 371. The shader compiler may be configured to compile compute shaders or other general graphics processing programs into low-level primitives that can be parallelized across the graphics core 370, the tensor core 371, and the ray tracing core 372. Techniques for GPU - to - Host Processor Interconnection Figure 4A FIG. illustrates an exemplary architecture in which a plurality of GPUs 410-413 (e.g., such as the parallel processor 200 shown in Figure 2A are communicatively coupled to a plurality of multi-core processors 405-406 via high-speed links 440A-440D (e.g., buses, point-to-point interconnects, etc.). Depending on the implementation, the high-speed links 440A-440D may support communication throughputs of 4 GB / s, 30 GB / s, 80 GB / s, or higher. A variety of interconnect protocols may be used, including but not limited to PCIe 4.0 or 5.0 and NVLink 2.0. However, the basic principles described herein are not limited to any particular communication protocol or throughput. Two or more of GPUs 410-413 may be interconnected via high-speed links 442A-442B, which may be implemented using the same or different protocols / links as those used for high-speed links 440A-440D. Similarly, two or more of multi-core processors 405-406 may be connected via high-speed link 443, which may be a symmetric multi-processor (SMP) bus operating at 20 GB / s, 30 GB / s, 120 GB / s, or lower or higher speeds. Alternatively, Figure 4A all communication between the various system components shown in Figure 4A may be implemented using the same protocol / link (e.g., via a common interconnect structure). However, as mentioned, the basic principles described herein are not limited to any particular type of interconnect technology. Each of multi-core processors 405 and 406 may be communicatively coupled to processor memories 401-402 via memory interconnects 430A-430B, respectively, and each GPU 410-413 is communicatively coupled to GPU memories 420-423 via GPU memory interconnects 450A-450D, respectively. Memory interconnects 430A-430B and 450A-450D may utilize the same or different memory access technologies. By way of example and not limitation, processor memories 401-402 and GPU memories 420-423 may be volatile memories such as dynamic random access memory (DRAM) (including stacked DRAM), graphics DDR SDRAM (GDDR) (e.g., GDDR5, GDDR6), or high bandwidth memory (HBM), and / or may be non-volatile memories such as 3D XPoint / Optane or Nano-Ram. For example, a portion of the memory may be volatile memory and another portion may be non-volatile memory (e.g., using a two-level memory (2LM) hierarchy). The memory subsystem as described herein may be compatible with several memory technologies such as double data rate versions published by JEDEC (Joint Electronic Device Engineering Council). As described below, although each of the processors 405-406 and GPUs 410-413 may be physically coupled to specific memories 401-402, 420-423 respectively, a unified memory architecture may be implemented in which the same virtual system address space (also referred to as the "effective address" space) is distributed among all of the various physical memories. For example, each of the processor memories 401-402 may contain 64GB of system memory address space, and each of the GPU memories 420-423 may contain 32GB of system memory address space (resulting in a total of 256GB of addressable memory in this example). Figure 4B Additional optional details of the interconnection between the multi-core processor 407 and the graphics acceleration module 446 are illustrated. The graphics acceleration module 446 may include one or more GPU chips integrated on a line card, which is coupled to the processor 407 via a high-speed link 440. Alternatively, the graphics acceleration module 446 may be integrated on the same package or chip as the processor 407. The illustrated processor 407 includes a plurality of cores 460A-460D, each of which has a translation lookaside buffer 461A-461D and one or more caches 462A-462D. The cores may include various other components for executing instructions and processing data, and these components are not illustrated to avoid obscuring the basic principles of the components described herein (e.g., instruction fetch unit, branch prediction unit, decoder, execution unit, reorder buffer, etc.). The caches 462A-462D may include a first-level (L1) cache and a second-level (L2) cache. In addition, one or more shared caches 456 may be included in the cache hierarchy and shared by the core set 460A-460D. For example, one embodiment of the processor 407 includes 24 cores, each of which has its own L1 cache, twelve shared L2 caches, and twelve shared L3 caches. In this embodiment, one of the L2 cache and the L3 cache is shared by two adjacent cores. The processor 407 and the graphics accelerator integrated module 446 are connected to the system memory 441, which may include the processor memories 401-402. Coherence is maintained for the data and instructions stored in the caches 462A-462D, 456, and the system memory 441 via inter-core communication through the coherence bus 464. For example, each cache may have cache coherence logic / circuit components associated therewith to communicate via the coherence bus 464 in response to detected reads or writes to a particular cache line. In one implementation, a cache snooping protocol is implemented via the coherence bus 464 to snoop on cache accesses. Cache snooping / coherence techniques are well understood by those skilled in the art and will not be described in detail herein to avoid obscuring the basic principles described herein. Agent circuitry 425 may be provided to communicatively couple the graphics acceleration module 446 to the coherence bus 464, allowing the graphics acceleration module 446 to participate as a peer of the cores in a cache coherence protocol. Specifically, interface 435 provides connectivity to the agent circuitry 425 via a high-speed link 440 (e.g., a PCIe bus, NVLink, etc.), and interface 437 connects the graphics acceleration module 446 to the high-speed link 440. In one implementation, the accelerator integrated circuit 436 provides cache management, memory access, context management, and interrupt management services on behalf of the plurality of graphics processing engines 431, 432, ……, N of the graphics acceleration module 446. Each of the graphics processing engines 431, 432, ……, N may include separate graphics processing units (GPUs). Alternatively, the graphics processing engines 431, 432, ……, N may include different types of graphics processing engines within a GPU, such as graphics execution units, media processing engines (e.g., video encoders / decoders), samplers, and block image transfer (BLIT) engines. In other words, the graphics acceleration module may be a GPU having a plurality of graphics processing engines 431-432, ……, N, or the graphics processing engines 431-432, ……, N may be separate GPUs integrated on a common package, line card, or chip. The graphics processing engines 431-432, ……, N may be configured using any of the graphics processor or computing accelerator architectures described herein. The accelerator integrated circuit 436 may include a memory management unit (MMU) 439 for performing various memory management functions, such as virtual-to-physical memory translation (also known as effective-to-real memory translation) and a memory access protocol for accessing system memory 441. The MMU 439 may also include a translation lookaside buffer (TLB) (not shown) for caching virtual / effective-to-physical / real address translations. In one implementation, the cache 438 stores commands and data for efficient access by the graphics processing engines 431, 432, ……, N. Data stored in the cache 438 and the graphics memories 433-434, ……, M may be kept coherent with the core caches 462A-462D, 456, and the system memory 441. As mentioned, this may be accomplished via the agent circuitry 425, which participates in the cache coherence mechanism on behalf of the cache 438 and the memories 433-434, ……, M (e.g., sending updates related to modifications / accesses of cache lines on the processor caches 462A-462D, 456 to the cache 438 and receiving updates from the cache 438). The register set 445 stores context data for the threads to be executed by the graphics processing engines 431 - 432, …, N, and the context management circuit 448 manages these thread contexts. For example, the context management circuit 448 can perform save and restore operations to save and restore the contexts of the respective threads during a context switch (e.g., where the first thread is saved and the second thread is restored so that the second thread can be executed by the graphics processing engine). For example, during a context switch, the context management circuit 448 can store the current register values into a specified area in the memory (e.g., identified by a context pointer). When returning to this context, it can then restore the register values. The interrupt management circuit 447 can, for example, receive interrupts from system devices and process the interrupts received from the system devices. In one implementation, the MMU 439 translates the virtual / valid addresses from the graphics processing engine 431 into actual / physical addresses in the system memory 441. Optionally, the accelerator integrated circuit 436 supports multiple (e.g., 4, 8, 16) graphics accelerator modules 446 and / or other accelerator devices. The graphics accelerator modules 446 can be dedicated to a single application executed on the processor 407 or can be shared among multiple applications. Optionally, a virtualized graphics execution environment is provided in which the resources of the graphics processing engines 431 - 432, …, N are shared among multiple applications, virtual machines (VMs), or containers. The resources can be subdivided into “slices” that are allocated to different VMs and / or applications based on the processing requirements and priorities associated with the VMs and / or applications or based on a predefined partitioning profile for the graphics accelerator modules 446. VMs and containers can be used interchangeably herein. A virtual machine (VM) can be software that runs an operating system and one or more applications. A VM can be defined by a specification, configuration file, virtual disk file, non - volatile random access memory (NVRAM) settings file, and log files, and is backed up by the physical resources of the host computing platform. A VM can include an operating system (OS) or application environment installed on software that emulates dedicated hardware. End - users have the same experience on a virtual machine as they would have on dedicated hardware. Specialized software called a hypervisor fully emulates the CPU, memory, hard disk, network, and other hardware resources of a PC client or server, enabling virtual machines to share resources. The hypervisor can emulate multiple virtual hardware platforms isolated from each other, allowing virtual machines to run on the same underlying physical host Servers, VMware ESXi, and other operating systems. A container can be a software package of applications, configurations, and dependencies, so that an application can run reliably from one computing environment to another. Containers can share the operating system installed on a server platform and run as isolated processes. A container can be a software package that contains everything required for software to run (such as system tools, libraries, and settings). Containers are not installed like traditional software programs, which allows containers to be isolated from other software and from the operating system itself. The isolated nature of containers provides several benefits. First, the software in a container will run in the same way in different environments. For example, a container including PHP and MySQL can run in exactly the same way on both a computer and machine. Second, containers provide increased security because the software will not affect the host operating system. While installed applications may change system settings and modify resources (such as the Windows registry), a container is able to modify only the settings within that container. Accordingly, the accelerator integrated circuit 436 acts as a bridge to the system for the graphics acceleration module 446 and provides address translation and system memory caching services. In one embodiment, to facilitate the bridging functionality, the accelerator integrated circuit 436 may further include shared I / O 497 (e.g., PCIe, USB, or other components) and hardware to enable system control of voltage, clock control, performance, thermal, and security. The shared I / O 497 may utilize separate physical connections or may span the high-speed link 440. Additionally, the accelerator integrated circuit 436 may provide virtualization facilities for the host processor to manage virtualization of the graphics processing engine, interrupts, and memory management. Since the hardware resources of the graphics processing engines 431 - 432, …, N are explicitly mapped to the actual address space seen by the host processor 407, any host processor can use valid address values to directly address these resources. An optional function of the accelerator integrated circuit 436 is to physically separate the graphics processing engines 431 - 432, …, N such that they appear to the system as independent units. One or more graphics memories 433 - 434, …, M may be respectively coupled to each of the graphics processing engines 431 - 432, …, N. The graphics memories 433 - 434, …, M store the instructions and data being processed by each of the graphics processing engines 431 - 432, …, N. The graphics memories 433 - 434, …, M can be volatile memories such as DRAM (including stacked DRAM), GDDR memory (e.g., GDDR5, GDDR6), or HBM, and / or can be non-volatile memories such as 3D XPoint / Optane, Samsung Z-NAND, or Nano-Ram. To reduce the data traffic on the high-speed link 440, a biasing technique can be used to ensure that the data stored in the graphics memories 433-434, …, M is the data that will be most frequently used by the graphics processing engines 431-432, …, N and preferably not used (or at least not frequently used) by the cores 460A-460D. Similarly, the biasing mechanism attempts to keep the data required by the cores (and preferably not the graphics processing engines 431-432, …, N) in the system memory 441 and the caches 462A-462D, 456 of the cores. According to Figure 4C the variant shown in Figure 4B , the accelerator integrated circuit 436 is integrated within the processor 407. The graphics processing engines 431-432, …, N communicate directly with the accelerator integrated circuit 436 via the high-speed link 440, through the interfaces 437 and 435 (which can again utilize any form of bus or interface protocol). The accelerator integrated circuit 436 can perform the same operations as those described with respect to , but potentially with a higher throughput considering the close proximity of the accelerator integrated circuit 436 to the coherence bus 464 and the caches 462A-462D, 456. The described embodiments can support different programming models, which include a dedicated process programming model (without graphics acceleration module virtualization) and a shared programming model (with virtualization). The latter can include a programming model controlled by the accelerator integrated circuit 436 and a programming model controlled by the graphics acceleration module 446. In an embodiment of the dedicated process model, the graphics processing engines 431, 432, …, N can be dedicated to a single application or process under a single operating system. A single application can funnel requests from other applications to the graphics engines 431, 432, …, N, thus providing virtualization within a VM / partition. For a shared programming model, the graphics acceleration module 446 or individual graphics processing engines 431-432, ..., N use a process handle to select process elements. The process elements may be stored in system memory 441 and may be addressable using the effective address to physical address translation techniques described herein. The process handle may be an implementation-specific value provided to the host process when registering its context with the graphics processing engines 431-432, ..., N (i.e., calling system software to add the process element to a process element linked list). The lower 16 bits of the process handle may be the offset of the process element within the process element linked list. Figure 4D FIG. illustrates an exemplary accelerator integration slice 490. As used herein, "slice" includes a designated portion of the processing resources of the accelerator integrated circuit 436. An application virtual address space 482 within system memory 441 stores process elements 483. The process elements 483 may be stored in response to a GPU call 481 from an application 480 executing on the processor 407. The process elements 483 contain the process state of the corresponding application 480. A work descriptor (WD) 484 contained within the process element 483 may be a single job requested by the application or may include a pointer to a job queue. In the latter case, the WD 484 is a pointer to a job request queue within the application's address space 482. The graphics acceleration module 446 and / or individual graphics processing engines 431-432, ..., N may be shared by all processes in the system or a subset of the processes in the system. For example, the techniques described herein may include infrastructure for establishing process state and sending the WD 484 to the graphics acceleration module 446 to initiate a job in a virtualized environment. In one implementation, the dedicated process programming model is implementation-specific. In this model, a single process owns the graphics acceleration module 446 or a separate graphics processing engine 431. Since the graphics acceleration module 446 is owned by a single process, when the graphics acceleration module 446 is assigned, the hypervisor initializes the accelerator integrated circuit 436 for the owning partition, and the operating system initializes the accelerator integrated circuit 436 for the owning process. In operation, the WD fetch unit 491 in the accelerator integrated slice 490 fetches the next WD 484, which includes an indication of work to be completed by one of the graphics processing engines in the graphics processing engine of the graphics acceleration module 446. As shown, data from the WD 484 can be stored in the register 445 and used by the MMU 439, the interrupt management circuit 447, and / or the context management circuit 448. For example, the MMU 439 may include segment / page walk circuit components for accessing the segment table / page table 486 within the OS virtual address space 485. The interrupt management circuit 447 can process the interrupt event 492 received from the graphics acceleration module 446. When performing a graphics operation, the effective address 493 generated by the graphics processing engines 431 - 432, …, N is translated into a physical address by the MMU 439. The same set of registers 445 can be replicated for each graphics processing engine 431 - 432, …, N and / or the graphics acceleration module 446 and can be initialized by the hypervisor or the operating system. Each of these replicated registers can be included in the accelerator integrated slice 490. In one embodiment, each graphics processing engine 431 - 432, …, N can be presented to the hypervisor 496 as a different graphics processor device. QoS settings can be configured for the clients of a particular graphics processing engine 431 - 432, …, N, and data isolation between the clients of each engine can be enabled. Exemplary registers that can be initialized by the hypervisor are shown in Table 1. Table 1 - Hypervisor Initialized Registers 1 Slice Control Register 2 Real Address (RA) Scheduled Process Region Pointer 3 Permission Mask Override Register 4 Interrupt Vector Table Entry Offset 5 Interrupt Vector Table Entry Limit 6 Status Register 7 Logical Partition ID 8 Real Address (RA) Hypervisor Accelerator Utilization Record Pointer 9 Storage Descriptor Register Exemplary registers that can be initialized by the operating system are shown in Table 2. Table 2 - Operating System Initialized Registers 1 Process and Thread Identification 2 Effective Address (EA) Context Save / restore Pointer 3 Virtual Address (VA) Accelerator Utilization Record Pointer 4 Virtual Address (VA) Storage Segment Table Pointer 5 Permission Mask 6 Work Descriptor Each WD 484 can be specific to a particular graphics acceleration module 446 and / or graphics processing engine 431 - 432, …, N. It contains all the information that the graphics processing engines 431 - 432, …, N need to complete their work, or it can be a pointer to the memory location where the command queue for the work to be completed by the application has been established. Figure 4E Additional optional details of the shared model are illustrated. It includes the hypervisor physical address space 498 in which the process element list 499 is stored. The hypervisor physical address space 498 is accessible via the hypervisor 496, which virtualizes the graphics acceleration module engine for the operating system 495. The shared programming model allows all processes or a subset of processes from all partitions in the system or a subset of partitions in the system to use the graphics acceleration module 446. There are two programming models in which the graphics acceleration module 446 is shared by multiple processes and partitions: time-sharing and graphics-directed sharing. In this model, the hypervisor 496 owns the graphics acceleration module 446 and makes its functionality available to all operating systems 495. To enable the graphics acceleration module 446 to support virtualization by the hypervisor 496, the graphics acceleration module 446 may comply with the following requirements: 1) The job requests of the application must be autonomous (i.e., the state does not need to be maintained between jobs), or the graphics acceleration module 446 must provide a context save and restore mechanism. 2) The graphics acceleration module 446 must ensure that the job requests of the application are completed within the specified amount of time, including any translation errors, or the graphics acceleration module 446 provides the ability to preempt the processing of jobs. 3) The graphics acceleration module 446 must be guaranteed fairness between processes when operating in the directed sharing programming model. For a directional sharing model, application 480 may be required to make an operating system 495 system call with the graphics acceleration module 446 type, work descriptor (WD), authority mask register (AMR) value, and context save / restore area pointer (CSRP). The graphics acceleration module 446 type describes the target acceleration function for the system call. The graphics acceleration module 446 type can be a system-specific value. The WD is specifically formatted for the graphics acceleration module 446, and the WD may take the following forms: a graphics acceleration module 446 command, a valid address pointer to a user-defined structure, a valid address pointer to a command queue, or any other data structure for describing the work to be done by the graphics acceleration module 446. In one embodiment, the AMR value is the AMR state to be used for the current process. The value passed to the operating system is similar to the application that sets the AMR. If the accelerator integrated circuit 436 and the graphics acceleration module 446 implementation do not support the User Authority MaskOverride Register (UAMOR), the operating system may apply the current UAMOR value to the AMR value before passing the AMR in the hypervisor call. The hypervisor 496 may optionally apply the current Authority Mask Override Register (AMOR) value before placing the AMR in the process element 483. The CSRP can be one of the registers 445 that contains the valid address of a region in the address space 482 of the application for the graphics acceleration module 446 to use to save and restore the context state. If no state is required to be saved between jobs or when a job is preempted, the pointer is optional. The context save / restore area can be pinned system memory. Upon receiving the system call, the operating system 495 may verify that the application 480 is registered and has been granted permission to use the graphics acceleration module 446. The operating system 495 then calls the hypervisor 496 with the information shown in Table 3. Table 3 - OS Call Parameters to the Hypervisor Upon receiving the hypervisor call, the hypervisor 496 verifies that the operating system 495 is registered and has been granted permission to use the graphics acceleration module 446. The hypervisor 496 then places the process element 483 in the linked list of process elements for the corresponding graphics acceleration module 446 type. The process element may include the information shown in Table 4. Table 4 - Process Element Information 1 Work Descriptor (WD) 2 Permission Mask Register (AMR) Value (Potentially Masked). 3 Effective Address (EA) Context Save / restore Region Pointer (CSRP) 4 Process ID (PID) and Optional Thread ID (TID) 5 Virtual Address (VA) Accelerator Utilization Record Pointer (AURP) 6 Virtual Address of Storage Segment Table Pointer (SSTP) 7 Logical Interrupt Service Number (LISN) 8 Interrupt Vector Table, Derived from Hypervisor Call Arguments 9 State Register (state register, SR) Value 10 Logical Partition ID (logical partition ID, LPID) 11 Real Address (RA) Hypervisor Accelerator Utilization Record Pointer 12 Storage Descriptor Register (Storage Descriptor Register, SDR) The hypervisor may initialize the registers 445 of multiple accelerator integration slices 490. As Figure 4F Illustrated, in an optional implementation, a unified memory addressable via a common virtual memory address space is employed, and the common virtual memory address space is used to access the physical processor memories 401-402 and the GPU memories 420-423. In this implementation, operations executed on the GPUs 410-413 utilize the same virtual / effective memory address space to access the processor memories 401-402 and vice versa, thus simplifying programmability. A first portion of the virtual / effective address space may be allocated to the processor memory 401, a second portion may be allocated to the second processor memory 402, a third portion may be allocated to the GPU memory 420, and so on. The entire virtual / effective memory space (sometimes referred to as the effective address space) may thus be distributed across each of the processor memories 401-402 and the GPU memories 420-423, allowing any processor or GPU to access the physical memory using a virtual address mapped to any physical memory. One or more bias / coherency management circuit components 494A-494E may be provided within one or more of the MMUs 439A-439E. These bias / coherency management circuit components 494A-494E ensure cache coherency between the caches of the host processor (e.g., 405) and the GPUs 410-413 and implement a bias technique for a physical memory indicating where certain types of data should be stored. Although Figure 4F Multiple instances of the bias / coherency management circuit components 494A-494E are illustrated, the bias / coherency circuit components may be implemented within the MMU of one or more host processors 405 and / or within the accelerator integrated circuit 436. The GPU-attached memories 420-423 can be mapped as part of the system memory and accessed using shared virtual memory (SVM) technology, but do not suffer from the typical performance drawbacks associated with full system cache coherence. The ability of the GPU-attached memories 420-423 to be accessed as system memory without heavy cache coherence overhead provides a beneficial operating environment for GPU migration. This arrangement allows the host processor 405 to set up operation objects and access computation results without the overhead of traditional I / O DMA data copying. Such traditional copying involves driver calls, interrupts, and memory mapped I / O (MMIO) accesses, all of which are inefficient compared to simple memory accesses. At the same time, the ability to access the GPU-attached memories 420-423 without cache coherence overhead can be critical to the execution time of migrated computations. For example, in the presence of a large amount of streaming write memory traffic, the cache coherence overhead can significantly reduce the effective write bandwidth seen by the GPUs 410-413. The efficiency of operation object setup, result access, and GPU computation all play a role in determining the effectiveness of GPU migration. The choice between GPU bias and host processor bias can be driven by a bias tracker data structure. For example, a bias table can be used, which can be a page-granularity structure (i.e., controlled at the granularity of memory pages), and this page-granularity structure includes 1 or 2 bits per GPU-attached memory page. The bias table can be implemented in the stolen memory ranges of one or more of the GPU-attached memories 420-423, with or without a bias cache in the GPUs 410-413 (e.g., for caching frequently / recently used entries of the bias table). Alternatively, the entire bias table can be maintained within the GPU. In one implementation, prior to an actual access to the GPU memory, the bias table entry associated with each access to the GPU-attached memories 420-423 is accessed, resulting in the following operations. First, local requests from the GPUs 410-413 for pages found in GPU bias are directly forwarded to the corresponding GPU memories 420-423. Local requests from the GPUs for pages found in host bias are forwarded to the processor 405 (e.g., via a high-speed link as discussed above). Optionally, requests from the processor 405 for pages found in host processor bias complete the request as a normal memory read. Alternatively, requests involving pages in GPU bias can be forwarded to the GPUs 410-413. If the GPU is not currently using the page, the GPU can then transition the page to host processor bias. The bias state of a page can be changed through software-based mechanisms, hardware-assisted software-based mechanisms, or for a limited set of cases, through pure hardware-based mechanisms. A mechanism for changing the bias state employs an API call (e.g., OpenCL), which in turn calls the device driver of the GPU. The device driver then sends a message (or enqueues a command descriptor) to the GPU that instructs the GPU to change the bias state and perform a cache dump purge operation in the host for some transitions. A cache dump purge operation is required for the transition from host processor 405 bias to GPU bias, but not for the reverse transition. Cache coherence can be maintained by temporarily rendering pages that are GPU-biased and not cacheable by the host processor 405. To access these pages, processor 405 can request access from GPU 410, which, depending on the implementation, may or may not grant immediate access. Therefore, to reduce communication between host processor 405 and GPU 410, it is beneficial to ensure that the GPU-biased pages are those that the GPU needs but not the host processor 405, and vice versa. Graphics Processing Pipeline Figure 5 FIG. illustrates a graphics processing pipeline 500. A graphics multiprocessor (such as, for example, the graphics multiprocessor 234 as in Figure 2D the graphics multiprocessor 325 as in Figure 3A the graphics multiprocessor 350 as in Figure 3B can implement the illustrated graphics processing pipeline 500. The graphics multiprocessor can be included within a parallel processing subsystem as described herein, such as a parallel processor 200 as in Figure 2A which can be associated with and can be used in place of one of the (one or more) parallel processors 112 as in Figure 1 Various parallel processing systems can implement the graphics processing pipeline 500 via one or more instances of a parallel processing unit (e.g., the parallel processing unit 202 as in Figure 2A For example, a shader unit (e.g., the graphics multiprocessor 234 as in Figure 2C can be configured to perform one or more of the following functions: vertex processing unit 504, tessellation control processing unit 508, tessellation evaluation processing unit 512, geometry processing unit 516, and fragment / pixel processing unit 524. The functions of data assembler 502, primitive assembler 506, 514, 518, tessellation unit 510, rasterizer 522, and raster operation unit 526 can also be performed by other processing engines and corresponding partitioning units within a processing cluster (e.g., the processing cluster 214 as in Figure 2A as in Figure 2Ais performed by the partitioning units 220A - 220N). The graphics processing pipeline 500 may also be implemented using dedicated processing units for one or more functions. It is also possible that one or more portions of the graphics processing pipeline 500 are performed by parallel processing logic within a general - purpose processor (e.g., a CPU). Optionally, one or more portions of the graphics processing pipeline 500 may access on - chip memory (e.g., the parallel processor memory 222 as in Figure 2A via a memory interface 528, and this memory interface 528 may be an instance of the memory interface 218 of Figure 2A . The graphics processor pipeline 500 may also be implemented via a multi - core group 365A as in Figure 3C . The data assembler 502 is a processing unit that can collect vertex data for surfaces and primitives. The data assembler 502 then outputs the vertex data, which includes vertex attributes, to the vertex processing unit 504. The vertex processing unit 504 is a programmable execution unit that executes a vertex shader program to illuminate and transform the vertex data as specified by the vertex shader program. The vertex processing unit 504 reads data stored in a cache, local, or system memory for use in processing the vertex data and can be programmed to transform the vertex data from an object - based coordinate representation to a world - space coordinate space or a normalized device coordinate space. A first instance of the primitive assembler 506 receives vertex attributes from the vertex processing unit 504. The primitive assembler 506 reads the stored vertex attributes as needed and constructs graphics primitives for processing by the tessellation control processing unit 508. Graphics primitives include triangles, line segments, points, patches, etc. supported by various graphics processing application programming interfaces (APIs). The tessellation control processing unit 508 treats the input vertices as control points of a geometric patch. The control points are transformed from an input representation from the patch (e.g., the basis of the patch) to a representation suitable for use in surface evaluation by the tessellation evaluation processing unit 512. The tessellation control processing unit 508 may also calculate the tessellation factor for the edges of the geometric patch. The tessellation factor is applied to individual edges and quantifies the view - dependent level of detail associated with that edge. The tessellation unit 510 is configured to receive the tessellation factors for the edges of the patch and to tessellate the patch surface into multiple geometric primitives (such as line, triangle, or quadrilateral primitives), which are sent to the tessellation evaluation processing unit 512. The tessellation evaluation processing unit 512 operates on the parameterized coordinates of the tessellated patch to generate a surface representation and vertex attributes for each vertex associated with the geometric primitive. A second instance of the primitive assembler 514 receives vertex attributes from the tessellation evaluation processing unit 512 as needed, reads the stored vertex attributes, and constructs graphics primitives for processing by the geometry processing unit 516. The geometry processing unit 516 is a programmable execution unit that executes a geometry shader program to transform the graphics primitives received from the primitive assembler 514 as specified by the geometry shader program. The geometry processing unit 516 can be programmed to subdivide a graphics primitive into one or more new graphics primitives and calculate parameters used to rasterize those new graphics primitives. The geometry processing unit 516 may be able to add or delete elements in the geometry stream. The geometry processing unit 516 outputs parameters and vertices specifying the new graphics primitives to the primitive assembler 518. The primitive assembler 518 receives the parameters and vertices from the geometry processing unit 516 and constructs graphics primitives for processing by the viewport scaling, culling, and clipping unit 520. The geometry processing unit 516 reads data stored in the parallel processor memory or system memory for use in processing geometry data. The viewport scaling, culling, and clipping unit 520 performs clipping, culling, and viewport scaling and outputs the processed graphics primitives to the rasterizer 522. The rasterizer 522 may perform depth culling and other depth-based optimizations. The rasterizer 522 also performs a scan conversion on the new graphics primitives to generate fragments and outputs those fragments and associated coverage data to the fragment / pixel processing unit 524. The fragment / pixel processing unit 524 is a programmable execution unit configured to execute a fragment shader program or a pixel shader program. The fragment / pixel processing unit 524 transforms the fragments or pixels received from the rasterizer 522 as specified by the fragment or pixel shader program. For example, the fragment / pixel processing unit 524 can be programmed to perform operations including but not limited to texture mapping, shading, blending, texture correction, and perspective correction to produce shaded fragments or pixels that are output to the raster operations unit 526. The fragment / pixel processing unit 524 can read data stored in the parallel processor memory or system memory for use in processing fragment data. The fragment or pixel shader program can be configured to shade at a sample, pixel, tile, or other granularity depending on the sampling rate configured for the processing unit. The raster operations unit 526 is a processing unit that performs raster operations and outputs pixel data to be stored in the graphics memory (e.g., as Figure 2A the parallel processor memory 222 in Figure 1In the system memory 104), the processed graphic data to be displayed on one or more display devices 110A - 110B or for further processing by one of the one or more processors 102 or parallel processors 112. These raster operations include, but are not limited to, stenciling, z - testing, blending, etc. The raster operation unit 526 can be configured to compress the z - data or color data written to the memory and decompress the z - data or color data read from the memory. Machine Learning Overview The architectures described above can be applied to perform training and inference operations using machine - learning models. Machine learning has been successful in solving many kinds of tasks. The computations that occur when training and using machine - learning algorithms (e.g., neural networks) are naturally well - suited for efficient parallel implementations. Accordingly, parallel processors such as general - purpose graphics processing units (GPGPUs) have played an important role in the practical implementation of deep neural networks. Parallel graphics processors with a single - instruction multiple - thread (SIMT) architecture are designed to maximize the amount of parallel processing in a graphics pipeline. In the SIMT architecture, groups of parallel threads attempt to execute program instructions synchronously together as frequently as possible to increase processing efficiency. The efficiency provided by parallel machine - learning algorithm implementations allows the use of high - capacity networks and enables the training of those networks on larger data sets. Machine - learning algorithms are algorithms that are capable of learning based on a data set. For example, a machine - learning algorithm can be designed to model high - level abstractions within a data set. For example, an image - recognition algorithm can be used to determine which of several classes a given input belongs to; given an input, a regression algorithm can output a numerical value; and a pattern - recognition algorithm can be used to generate transformed text or perform text - to - speech and / or speech recognition. An exemplary type of machine - learning algorithm is a neural network. There are many types of neural networks; a simple type of neural network is a feed - forward network. A feed - forward network can be implemented as an acyclic graph in which nodes are arranged in layers. Typically, a feed - forward network topology includes an input layer and an output layer separated by at least one hidden layer. The hidden layer transforms the input received by the input layer into a representation useful for generating an output in the output layer. Network nodes are fully connected via edges to nodes in adjacent layers, but there are no edges between nodes within each layer. The data received at the nodes of the input layer of a feed - forward network is propagated (i.e., “fed forward”) to the nodes of the output layer via an activation function that calculates the state of the nodes in each successive layer of the network based on coefficients (“weights”) associated with each of the edges connecting these layers. Depending on the particular model being represented by the algorithm being executed, the output from a neural - network algorithm can take various forms. Before a machine learning algorithm can be used to model a particular problem, a training data set is used to train the algorithm. Training a neural network involves: selecting a network topology; using a set of training data representing the problem being modeled by the network; and adjusting the weights until the network model performs with minimal error for all instances of the training data set. For example, during supervised learning training of a neural network, the output produced by the network in response to an input representing an instance in the training data set is compared to the "correct" labeled output for that instance, an error signal representing the difference between the output and the labeled output is computed, and the weights associated with the connections are adjusted to minimize that error as the error signal is propagated backward through the layers of the network. When the error for each output in the output generated from the instances of the training data set is minimized, the network is considered to be "trained". The accuracy of machine learning algorithms is significantly affected by the quality of the data sets used to train the algorithms. The training process can be computationally intensive and may require a large amount of time on a conventional general-purpose processor. Accordingly, parallel processing hardware is used to train many types of machine learning algorithms. This is particularly useful for optimizing the training of neural networks because the computations performed when adjusting the coefficients in a neural network are themselves naturally amenable to parallel implementation. Specifically, many machine learning algorithms and software applications have been adapted to utilize the parallel processing hardware within a general-purpose graphics processing device. Figure 6 is a generalized diagram of a machine learning software stack 600. A machine learning application 602 is any logic that can be configured to train a neural network using a training data set or to implement machine intelligence using a trained deep neural network. The machine learning application 602 may include training and inference functionality for a neural network and / or specialized software that can be used to train a neural network before deployment. The machine learning application 602 can implement any type of machine intelligence, including but not limited to: image recognition, map creation and localization, autonomous navigation, speech synthesis, medical imaging, or language translation. Example machine learning applications 602 include but are not limited to voice-based virtual assistants, image or face recognition algorithms, autonomous navigation, and software tools used to train the machine learning models used by the machine learning application 602. Hardware acceleration for machine learning applications 602 can be enabled via a machine learning framework 604. The machine learning framework 604 can provide a library of machine learning primitives. Machine learning primitives are basic operations typically performed by machine learning algorithms. Without the machine learning framework 604, developers of machine learning algorithms would be required to create and optimize the main computational logic associated with the machine learning algorithm, and then re-optimize the computational logic when developing new parallel processors. Instead, machine learning applications can be configured to use primitives provided by the machine learning framework 604 to perform necessary calculations. Exemplary primitives include tensor convolutions, activation functions, and pooling, which are computational operations performed when training convolutional neural networks (CNNs). The machine learning framework 604 can also provide primitives to implement basic linear algebra subroutines performed by many machine learning algorithms, such as matrix and vector operations. Examples of machine learning frameworks 604 include, but are not limited to, TensorFlow, TensorRT, PyTorch, MXNet, Caffee, and other advanced machine learning frameworks. The machine learning framework 604 can process input data received from the machine learning application 602 and generate appropriate inputs to the computing framework 606. The computing framework 606 can abstract the underlying instructions provided to the GPGPU driver 608 so that the machine learning framework 604 can take advantage of hardware acceleration via the GPGPU hardware 610 without requiring the machine learning framework 604 to be very familiar with the architecture of the GPGPU hardware 610. In addition, the computing framework 606 can enable hardware acceleration for the machine learning framework 604 across various types and generations of GPGPU hardware 610. Exemplary computing frameworks 606 include CUDA computing frameworks and associated machine learning libraries, such as CUDA Deep Neural Network (cuDNN) libraries. The machine learning software stack 600 may also include a communication library or framework to facilitate multi-GPU and multi-node computing. GPGPU Machine Learning Acceleration Figure 7 A general purpose graphics processing unit 700 is shown, which may be Figure 2A Parallel processor 200 or Figure 1 (one or more) parallel processors 112. The general purpose processing unit (GPGPU) 700 can be configured to provide support for hardware acceleration of primitives provided by machine learning frameworks to accelerate the processing of computing workloads of the type associated with training deep neural networks. In addition, the GPGPU 700 can be directly linked to other instances of the GPGPU to create a multi-GPU cluster, thereby improving the training speed of deep neural networks in particular. Primitives are also supported to accelerate inference operations for deployed neural networks. The GPGPU 700 includes a host interface 702 for enabling connection with a host processor. The host interface 702 can be a PCI Express interface. However, the host interface can also be a vendor-specific communication interface or communication fabric. The GPGPU 700 receives commands from the host processor and uses a global scheduler 704 to distribute execution threads associated with those commands to a set of processing clusters 706A - 706H. The processing clusters 706A - 706H share a cache memory 708. The cache memory 708 can act as a higher-level cache for the cache memories within the processing clusters 706A - 706H. The illustrated processing clusters 706A - 706H can correspond to the processing clusters 214A - 214N as in Figure 2A . The GPGPU 700 includes memories 714A - 714B coupled to the processing clusters 706A - 706H via a set of memory controllers 712A - 712B. The memories 714A - 714B can include various types of memory devices, including dynamic random access memory (DRAM) or graphics random access memory, such as synchronous graphics random access memory (SGRAM), including graphics double data rate (GDDR) memory. The memories 714A - 714B can also include 3D stacked memory, including but not limited to high bandwidth memory (HBM). Each of the processing clusters 706A - 706H can include a set of graphics multiprocessors, such as Figure 2D the graphics multiprocessor 234 of Figure 3A the graphics multiprocessor 325 of Figure 3B the graphics multiprocessor 350 of Figure 3C or can include multicore groups 365A - 365N as in Multiple instances of the GPGPU 700 can be configured to operate as a computing cluster. The communication mechanisms used by the computing cluster for synchronization and data exchange vary across embodiments. For example, multiple instances of the GPGPU 700 communicate through the host interface 702. In one embodiment, the GPGPU 700 includes an I / O hub 709 that couples the GPGPU 700 to a GPU link 710, which enables direct connections to other instances of the GPGPU. The GPU link 710 can be coupled to a dedicated GPU-to-GPU bridge that enables communication and synchronization between multiple instances of the GPGPU 700. Optionally, the GPU link 710 is coupled to a high-speed interconnect to transfer data to and receive data from other GPGPUs or parallel processors. Multiple instances of the GPGPU 700 can be located in separate data processing systems and can communicate via a network device that can be accessed via the host interface 702. In addition to or instead of the host interface 702, the GPU link 710 can be configured to enable a connection to a host processor. Although the illustrated configuration of the GPGPU 700 can be configured for training neural networks, alternative configurations of the GPGPU 700 can be configured for deployment within high-performance or low-power inference platforms. In an inference configuration, the GPGPU 700 includes fewer processing clusters among the processing clusters 706A - 706H as compared to a training configuration. Additionally, the memory technologies associated with the memories 714A - 714B can vary between the inference configuration and the training configuration. In one embodiment, the inference configuration of the GPGPU 700 can support inference-specific instructions. For example, the inference configuration can provide support for one or more 8-bit integer or floating-point dot product instructions that are typically used during inference operations for deployed neural networks. Figure 8 Illustrate a multi-GPU computing system 800. The multi-GPU computing system 800 can include a processor 802 coupled to multiple GPGPUs 806A - 806D via a host interface switch 804. The host interface switch 804 can be a PCI Express switch device that couples the processor 802 to a PCI Express bus through which the processor 802 can communicate with a collection of GPGPUs 806A - 806D. Each of the multiple GPGPUs 806A - 806D can be Figure 7An instance of the GPGPU 700. The GPGPUs 806A - 806D can be interconnected via a set of high - speed point - to - point GPU - to - GPU links 816. The high - speed GPU - to - GPU links can be connected to each of the GPGPUs 806A - 806D via dedicated GPU links, such as Figure 7 the GPU link 710 in Figure 7 . The P2P GPU links 816 enable direct communication between each of the GPGPUs in the GPGPUs 806A - 806D without the need for communication through the host interface bus to which the processor 802 is connected. With GPU - to - GPU traffic directed to the P2P GPU links, the host interface bus remains available for system memory access or for communicating, for example, with other instances of the multi - GPU computing system 800 via one or more network devices. Although in Figure 8 the GPGPUs 806A - 806D are connected to the processor 802 via the host interface switch 804, the processor 802 may alternatively include direct support for the P2P GPU links 816 and be directly connected to the GPGPUs 806A - 806D. In one embodiment, the P2P GPU links 816 enable the multi - GPU computing system 800 to operate as a single logical GPU. Machine Learning Neural Network Implementation The computing architectures described herein can be configured to perform types of parallel processing that are particularly suitable for training and deploying neural networks for machine learning. A neural network can be generalized as a network of functions with graph relationships. As is well known in the art, there are various types of neural network implementations used in machine learning. An exemplary type of neural network is the feed - forward network as previously described. The second exemplary type of neural network is the convolutional neural network (CNN). A CNN is a specialized feedforward neural network for processing data with a known, grid-like topology, such as image data. Accordingly, CNNs are commonly used in computer vision and image recognition applications, but they can also be used for other types of pattern recognition, such as speech and language processing. The nodes in the input layer of a CNN are organized into a collection of "filters" (feature detectors inspired by receptive fields found in the retina), and the output of each filter collection is propagated to nodes in successive layers of the network. The computations used for a CNN include applying the convolution mathematical operation to each filter to produce the output of that filter. Convolution is a specialized kind of mathematical operation performed by two functions to produce a third function, which is a modified version of one of the two original functions. In convolutional network terminology, the first function to the convolution can be referred to as the input, while the second function can be referred to as the convolution kernel. The output can be referred to as the feature map. For example, the input to a convolutional layer can be a multi-dimensional data array defining the various color components of an input image. The convolution kernel can be a multi-dimensional array of parameters, where the parameters are adapted through a training process for the neural network. A recurrent neural network (RNN) is a series of feedforward neural networks that includes feedback connections between the layers. An RNN enables modeling of sequential data by sharing parameter data across different parts of the neural network. The architecture used for an RNN includes loops. These loops represent the influence of the current value of a variable on its own value at future times, since at least part of the output data from the RNN is used as feedback for processing subsequent inputs in the sequence. Because of the variable nature in which language data can be composed, this feature makes RNNs particularly useful for language processing. The accompanying drawings described below present exemplary feedforward networks, CNN networks, and RNN networks, and describe the general processes for training and deploying each of those types of networks, respectively. It will be understood that these descriptions are exemplary and non-limiting with respect to any particular embodiments described herein, and generally speaking the concepts illustrated can generally be applied to deep neural networks and machine learning techniques. The exemplary neural networks described above can be used to perform deep learning. Deep learning is machine learning using deep neural networks. Different from shallow neural networks that include only a single hidden layer, the deep neural networks used in deep learning are artificial neural networks composed of multiple hidden layers. Deeper neural networks are generally more computationally intensive to train. However, the additional hidden layers of the network enable multi-step pattern recognition, which results in a reduced output error relative to shallow machine learning techniques. Deep neural networks used in deep learning typically include a front-end network for performing feature recognition, which is coupled to a back-end network that represents a mathematical model that can perform operations (e.g., object classification, speech recognition, etc.) based on the feature representation provided to the mathematical model. Deep learning enables machine learning to be performed without having to perform manual feature engineering for the model. Instead, the deep neural network can learn features based on the statistical structure or correlations within the input data. The learned features can be provided to the mathematical model, which can map the detected features to an output. The mathematical model used by the network is typically specific to the particular task to be performed, and different models will be used to perform different tasks. Once the neural network is structured, a learning model can be applied to the network to train the network to perform a specific task. The learning model describes how to adjust the weights within the model to reduce the output error of the network. Backpropagation of errors is a commonly used method for training neural networks. An input vector is presented to the network for processing. The output of the network is compared to the desired output using a loss function, and an error value is calculated for each neuron in the output layer. Subsequently, the error values are propagated backward until each neuron has an associated error value that roughly represents the contribution of that neuron to the original output. The network can then learn from those errors using an algorithm (such as the stochastic gradient descent algorithm) to update the weights of the neural network. Figures 9A - 9B Illustrates an exemplary convolutional neural network. Figure 9A Illustrates the individual layers within a CNN. As Figure 9A shown, an exemplary CNN for modeling image processing can receive an input 902 that describes the red, green, and blue (RGB) components of an input image. The input 902 can be processed by a plurality of convolutional layers (e.g., convolutional layer 904, convolutional layer 906). The output from the plurality of convolutional layers can optionally be processed by a set of fully connected layers 908. As previously described for a feedforward network, the neurons in a fully connected layer have full connections to all the activations in the previous layer. The output from the fully connected layer 908 can be used to generate an output result from the network. Matrix multiplication can be used instead of convolution to calculate the activations within the fully connected layer 908. Not all CNN implementations utilize the fully connected layer 908. For example, in some implementations, the convolutional layer 906 can generate the output of the CNN. The convolutional layers are sparsely connected, which is different from the traditional neural network configuration found in the fully connected layer 908. Traditional neural network layers are fully connected such that each output unit interacts with each input unit. However, as illustrated, the convolutional layers are sparsely connected because the output of the convolution of the receptive field (rather than the corresponding state values of each node in the receptive field) is input to the nodes of the subsequent layer. The kernel associated with the convolutional layer performs the convolution operation, and the output of this convolution operation is sent to the next layer. The dimensionality reduction performed within the convolutional layer is one aspect that enables the CNN to scale to handle large images. Figure 9B Illustrates an exemplary computational stage within the convolutional layer of a CNN. The input 912 to the convolutional layer of the CNN can be processed in three stages of the convolutional layer 914. These three stages can include a convolution stage 916, a detector stage 918, and a pooling stage 920. The convolutional layer 914 can then output data to a successive convolutional layer. The final convolutional layer of the network can generate output feature map data or provide an input to a fully connected layer, e.g., to generate classification values for the input to the CNN. In the convolution stage 916, a number of convolutions are performed in parallel to produce a set of linear activations. The convolution stage 916 can include an affine transformation, which is any transformation that can be specified as a linear transformation plus a translation. Affine transformations include rotation, translation, scaling, and combinations of these transformations. The convolution stage computes the output of a function (e.g., a neuron) connected to a specific region in the input, which can be determined as the local region associated with the neuron. The neuron computes the dot product between the weights of the neuron and the weights of the region in the local input to which the neuron is connected. The output from the convolution stage 916 defines the set of linear activations to be processed by successive stages of the convolutional layer 914. The linear activations can be processed by the detector stage 918. In the detector stage 918, each linear activation is processed by a non - linear activation function. The non - linear activation function increases the non - linear nature of the overall network without affecting the receptive field of the convolutional layer. Several types of non - linear activation functions can be used. One particular type is the rectified linear unit (ReLU), which uses an activation function defined as f(x) = max(0, x), such that the threshold for activation is zero. The pooling stage 920 uses a pooling function that replaces the output of the convolutional layer 906 with summary statistics of nearby outputs. The pooling function can be used to introduce translational invariance into the neural network such that small translations to the input do not change the pooled output. Invariance to local translations can be useful in scenarios where the presence of a feature in the input data is more important than the exact location of the feature. Various types of pooling functions can be used during the pooling stage 920, including max pooling, average pooling, and l2 norm pooling. Additionally, some CNN implementations do not include a pooling stage. Instead, such implementations are alternative and additional convolutional stages with an increased stride relative to the previous convolutional stage. The output from the convolutional layer 914 can then be processed by the next layer 922. The next layer 922 can be an additional convolutional layer or one of the fully connected layers 908. For example, Figure 9A the first convolutional layer 904 can output to the second convolutional layer 906, and the second convolutional layer can output to the first of the fully connected layers 908. Figure 10 An exemplary recurrent neural network 1000 is illustrated. In a recurrent neural network (RNN), the previous state of the network influences the output of the network's current state. RNNs can be constructed in various ways and using various functions. The use of RNNs generally revolves around using a mathematical model to predict the future based on a previous sequence of inputs. For example, an RNN can be used to perform statistical language modeling to predict the upcoming word given a previous sequence of words. The illustrated RNN 1000 can be described as having an input layer 1002 that receives an input vector, a hidden layer 1004 that implements a recurrent function, a feedback mechanism 1005 that enables'memory' of the previous state, and an output layer 1006 that outputs the result. The RNN 1000 operates based on time steps. The state of the RNN at a given time step is affected via the feedback mechanism 1005, based on the previous time step. For a given time step, the state of the hidden layer 1004 is defined by the previous state and the input at the current time step. The initial input (x1) at the first time step can be processed by the hidden layer 1004. The second input (x2) can be processed by the hidden layer 1004 using the state information determined during the processing of the initial input (x1). A given state can be calculated as s t = f(Ux t + Ws t-1 ) where U and W are parameter matrices. The function f is generally non-linear, such as the hyperbolic tangent function (Tanh) or a variant of the rectifier function f(x) = max(0, x). However, the specific mathematical function used in the hidden layer 1004 can vary depending on the specific implementation details of the RNN 1000. In addition to the basic CNN and RNN networks described, acceleration for variants of those networks can also be enabled. An example RNN variant is the long short term memory (LSTM) RNN. The LSTM RNN is capable of learning long-term dependencies, which may be necessary for processing longer language sequences. A variant of the CNN is the convolutional deep belief network, which has a structure similar to that of the CNN and is trained in a manner similar to that of the deep belief network. A deep belief network (DBN) is a generative neural network composed of multiple layers of stochastic (random) variables. The DBN can be trained layer by layer using greedy unsupervised learning. The learned weights of the DBN can then be used to provide a pre-trained neural network by determining an optimal initial set of weights for the neural network. In a further embodiment, acceleration for reinforcement learning is enabled. In reinforcement learning, an artificial agent learns by interacting with its environment. The agent is configured to optimize certain objectives to maximize cumulative reward. Figure 11 Illustrates the training and deployment of a deep neural network. Once a given network has been structured for a task, a training dataset 1102 is used to train the neural network. Various training frameworks 1104 have been developed to enable hardware acceleration of the training process. For example, Figure 6 the machine learning framework 604 can be configured as the training framework 1104. The training framework 1104 can access the untrained neural network 1106 and enable the untrained neural network to be trained using the parallel processing resources described herein to generate the trained neural network 1108. To initiate the training process, initial weights can be selected randomly or by pre-training using a deep belief network. Subsequently, a training loop is performed in a supervised or unsupervised manner. Supervised learning is a learning method in which training is performed as a mediated operation, such as when the training data set 1102 includes the input paired with the expected output of the input, or in the case where the training data set includes an input with a known output and the output of the neural network is manually graded. The network processes the input and compares the resulting output with the expected output or a set of expected outputs. Subsequently, the error is propagated back through the system. The training framework 1104 can be adjusted to adjust the weights that control the untrained neural network 1106. The training framework 1104 can provide tools to monitor how well the untrained neural network 1106 is converging to a model suitable for generating correct answers based on the known input data. As the weights of the network are adjusted to refine the output generated by the neural network, the training process occurs repeatedly. The training process can continue until the neural network reaches a statistically desired accuracy associated with the trained neural network 1108. The trained neural network 1108 can then be deployed to perform any number of machine learning operations to generate inference results 1114 based on the input of new data 1112. Unsupervised learning is a learning method in which the network attempts to train itself using unlabeled data. Thus, for unsupervised learning, the training data set 1102 will include input data that does not have any associated output data. The untrained neural network 1106 can learn groupings within the unlabeled input and can determine how individual inputs relate to the entire data set. Unsupervised training can be used to generate self-organizing maps, which are a type of trained neural network 1108 that is capable of performing operations useful in reducing the dimensionality of the data. Unsupervised training can also be used to perform anomaly detection, which allows identification of data points in the input data set that deviate from the normal pattern of the data. Variants of supervised training and unsupervised training can also be employed. Semi-supervised learning is a technique in which a mixture of labeled data and unlabeled data with the same distribution is included in the training data set 1102. Progressive learning is a variant of supervised learning in which input data is continuously used to further train the model. Progressive learning enables the trained neural network 1108 to adapt to new data 1112 without forgetting the knowledge ingrained in the network during initial training. Whether supervised or unsupervised, the training process for particularly deep neural networks can be computationally intensive for a single computing node. A distributed network of computing nodes can be used instead of a single computing node to accelerate the training process. Figure 12A is a block diagram illustrating distributed learning. Distributed learning is a training model that uses multiple distributed computing nodes to perform supervised or unsupervised training of a neural network. Each of the distributed computing nodes can include one or more host processors and one or more general-purpose processing nodes in a general-purpose processing node, such asFigure 7 The highly parallel general-purpose graphics processing unit 700 in Figure 7 . As shown, distributed learning can be performed with a combination of model parallelism 1202, data parallelism 1204, or model and data parallelism 1206. In model parallelism 1202, different computing nodes in a distributed system can perform training computations for different parts of a single network. For example, each layer of a neural network can be trained by a different processing node of the distributed system. Benefits of model parallelism include the ability to scale to particularly large models. Partitioning the computations associated with different layers of a neural network enables the training of very large neural networks in which all layer weights would not fit into the memory of a single computing node. In some instances, model parallelism can be particularly useful when performing unsupervised training of large neural networks. In data parallelism 1204, different nodes of a distributed network have a complete instance of the model, and each node receives a different portion of the data. Results from different nodes are then combined. While different ways for data parallelism are possible, all data parallelism training methods require techniques for combining results and synchronizing model parameters between each node. Exemplary ways for combining data include parameter averaging and update-based data parallelism. Parameter averaging trains each node on a subset of the training data and sets the global parameters (e.g., weights, biases) to the average of the parameters from each node. Parameter averaging uses a central parameter server that maintains parameter data. Update-based data parallelism is similar to parameter averaging, except that updates to the model are transmitted instead of parameters from the nodes to the parameter server. Additionally, update-based data parallelism can be performed in a decentralized manner where updates are compressed and transmitted between nodes. Combined model and data parallelism 1206 can be implemented, for example, in a distributed system where each computing node includes multiple GPUs. Each node can have a complete instance of the model, where individual GPUs within each node are used to train different parts of the model. Distributed training has increased overhead compared to training on a single machine. However, the parallel processors and GPGPUs described herein can each implement various techniques to reduce the overhead of distributed training, including techniques for enabling high-bandwidth GPU-to-GPU data transfer and accelerated remote data synchronization. Figure 12Bis a block diagram showing a programmable network interface 1210 and a data processing unit. The programmable network interface 1210 is a programmable network engine that can be used to accelerate network-based computing tasks within a distributed environment. The programmable network interface 1210 can be coupled to a host system via a host interface 1270. The programmable network interface 1210 can be used to accelerate network or storage operations for the CPU or GPU of the host system. The host system can be, for example, a node of a distributed learning system for performing distributed training, such as, as Figure 12A shown in. The host system can also be a data center node within a data center. In one embodiment, access to a remote storage device containing model data can be accelerated by the programmable network interface 1210. For example, the programmable network interface 1210 can be configured to present a remote storage device as a local storage device of the host system. The programmable network interface 1210 can also accelerate remote direct memory access (RDMA) operations performed between the GPU of the host system and the GPU of a remote system. In one embodiment, the programmable network interface 1210 can enable storage functionality such as, but not limited to, NVME-oF. The programmable network interface 1210 can also accelerate encryption, data integrity, compression, and other operations for a remote storage device on behalf of the host system, thereby allowing the remote storage device to approximate the latency of a storage device directly attached to the host system. The programmable network interface 1210 can also perform resource allocation and management on behalf of the host system. Storage security operations can be migrated to the programmable network interface 1210 and performed in concert with the allocation and management of remote storage resources. Network-based operations that would otherwise be performed by the processor of the host system for managing access to a remote storage device can instead be performed by the programmable network interface 1210. In one embodiment, network and / or data security operations can be migrated from the host system to the programmable network interface 1210. A data center security policy for a data center node can be handled by the programmable network interface 1210 rather than the processor of the host system. For example, the programmable network interface 1210 can detect and mitigate attempted network-based attacks (e.g., DDoS) on the host system, thereby preventing the attack from compromising the availability of the host system. The programmable network interface 1210 may include a system-on-chip (SoC 1220) that executes an operating system via multiple processor cores 1222. The processor cores 1222 may include general-purpose processor (e.g., CPU) cores. In one embodiment, the processor cores 1222 may also include one or more GPU cores. The SoC 1220 may execute instructions stored in the memory device 1240. The storage device 1250 may store local operating system data. The storage device 1250 and the memory device 1240 may also be used to cache remote data for the host system. The network ports 1260A - 1260B enable connection to a network or fabric, facilitate network access for the SoC 1220, and facilitate network access for the host system via the host interface 1270. The programmable network interface 1210 may also include an I / O interface 1275, such as a USB interface. The I / O interface 1275 may be used to couple external devices to the programmable network interface 1210 or as a debug interface. The programmable network interface 1210 further includes a management interface 1230 that enables software on the host device to manage and configure the programmable network interface 1210 and / or the SoC 1220. In one embodiment, the programmable network interface 1210 may also include one or more accelerators or GPUs 1245 to accept the migration of parallel computing tasks from the SoC 1220, the host system, or a remote system coupled via the network ports 1260A - 1260B. Exemplary Machine Learning Applications Machine learning can be applied to solve various technical problems, including but not limited to computer vision, autonomous driving and navigation, speech recognition, and language processing. Computer vision has traditionally been one of the most active research areas for machine learning applications. The scope of computer vision applications ranges from reproducing human visual capabilities (such as recognizing faces) to creating new classes of visual capabilities. For example, a computer vision application may be configured to identify sound waves from vibrations induced in objects visible in a video. Machine learning accelerated by parallel processors enables computer vision applications to train using training data sets that are significantly larger than previously feasible, and enables inference systems to be deployed using low-power parallel processors. Machine learning accelerated by parallel processors has applications in autonomous driving, including lane and road sign recognition, obstacle avoidance, navigation, and driving control. Accelerated machine learning techniques can be used to train a driving model based on a data set that defines the appropriate response to specific training inputs. The parallel processors described herein are capable of enabling the rapid training of increasingly complex neural networks for autonomous driving solutions and of enabling the deployment of low-power inference processors in mobile platforms suitable for integration into autonomous vehicles. Deep neural networks accelerated by parallel processors have enabled machine learning approaches for automatic speech recognition (ASR). ASR involves creating a function that computes the most likely language sequence given an input acoustic sequence. Accelerated machine learning using deep neural networks has enabled an alternative to the hidden Markov model (HMM) and Gaussian mixture model (GMM) previously used for ASR. Parallel processor-accelerated machine learning can also be used to accelerate natural language processing. The automated learning process can utilize statistical inference algorithms to produce models that are robust to noisy or unfamiliar inputs. Exemplary natural language processor applications include automatic machine translation between human languages. The parallel processing platforms for machine learning can be divided into a training platform and a deployment platform. The training platform is generally highly parallel and includes optimizations for accelerating multi-GPU single-node training and multi-node multi-GPU training. Exemplary parallel processors suitable for training include Figure 7 the general-purpose graphics processing unit 700 and Figure 8 the multi-GPU computing system 800. In contrast, the deployed machine learning platforms generally include lower-power parallel processors suitable for use in products such as cameras, autonomous robots, and autonomous vehicles. In addition, machine learning techniques can also be applied to accelerate or enhance graphics processing activities. For example, a machine learning model can be trained to recognize the output generated by a GPU-accelerated application and generate an enlarged version of the output. Such techniques can be applied to accelerate the generation of high-resolution images for game applications. Various other graphics pipeline activities can benefit from the use of machine learning. For example, a machine learning model can be trained to perform tessellation operations on geometric data to increase the complexity of a geometric model, thereby allowing the automatic generation of a more detailed geometry from a relatively low-detail geometry. Figure 13FIG. illustrates an exemplary inference system-on-a-chip (SOC) 1300 suitable for performing inference using a trained model. The SOC 1300 may integrate processing components, including a media processor 1302, a vision processor 1304, a GPGPU 1306, and a multi-core processor 1308. The GPGPU 1306 may be the GPGPU described herein, such as the GPGPU 700, and the multi-core processor 1308 may be the multi-core processor described herein, such as the multi-core processors 405-406. The SOC 1300 may additionally include on-chip memory 1305, which may enable a shared on-chip data pool accessible by each of the processing components. The processing components may be optimized for low-power operation to enable deployment on various machine learning platforms, including autonomous vehicles and autonomous robots. For example, one implementation of the SOC 1300 may be used as part of a main control system for an autonomous vehicle. In the case where the SOC 1300 is configured for use in an autonomous vehicle, the SOC is designed and configured to comply with relevant functional safety standards for deployment jurisdiction. During operation, the media processor 1302 and the vision processor 1304 may work together to accelerate computer vision operations. The media processor 1302 may enable low-latency decoding of multiple high-resolution (e.g., 4K, 8K) video streams. The decoded video streams may be written to a buffer in the on-chip memory 1305. The vision processor 1304 may then parse the decoded video and perform preliminary processing operations on the frames of the decoded video to prepare the frames for processing using a trained image recognition model. For example, the vision processor 1304 may accelerate convolutional operations for a CNN used to perform image recognition on high-resolution video data, while the backend model computations are performed by the GPGPU 1306. The multi-core processor 1308 may include control logic for assisting in the ordering and synchronization of data transfers and shared memory operations performed by the media processor 1302 and the vision processor 1304. The multi-core processor 1308 may also act as an application processor for executing software applications that can utilize the inference computing capabilities of the GPGPU 1306. For example, at least part of the navigation and driving logic may be implemented in software executed on the multi-core processor 1308. Such software may directly issue computational workloads to the GPGPU 1306, or the computational workloads may be issued to the multi-core processor 1308, which may migrate at least part of those operations to the GPGPU 1306. The GPGPU 1306 may include compute clusters, such as the low-power configured processing clusters 706A - 706H within the general-purpose graphics processing unit 700. The compute clusters within the GPGPU 1306 may support instructions specifically optimized to perform inference computations on trained neural networks. For example, the GPGPU 1306 may support instructions for performing low-precision computations such as 8-bit and 4-bit integer vector operations. Additional System Overview Figure 14 is a block diagram of the processing system 1400. Figure 14 Elements with the same or similar names as elements in any other figure in this document that describe the same elements as in other figures can operate or function in a similar manner as in other figures, may include the same components, and may be linked to other entities, such as those described elsewhere in this document, but are not limited to this. The system 1400 may be used in: a single-processor desktop computer system, a multi-processor workstation system, or a server system with a large number of processors 1402 or processor cores 1407. The system 1400 may be a processing platform incorporated within a system-on-chip (SoC) integrated circuit for use in mobile devices, handheld devices, or embedded devices, such as for use within an Internet-of-things (IoT) device with wired or wireless connectivity to a local area network or a wide area network. The system 1400 may be a processing system having components corresponding to those of Figure 1 . For example, in different configurations, the (one or more) processors 1402 or the (one or more) processor cores 1407 may correspond to the (one or more) processors 102 of Figure 1 . The (one or more) graphics processors 1408 may correspond to the (one or more) parallel processors 112 of Figure 1 . The external graphics processor 1418 may be one of the (one or more) plug-in devices 120 of Figure 1 . The system 1400 may include, may be coupled with, or may be integrated within: a server-based gaming platform; a game console, including a game and media console; a mobile game console, a handheld game console, or an online game console. The system 1400 may be part of a mobile phone, a smartphone, a tablet computing device, or a mobile Internet-connected device (such as a laptop computer with low internal storage capacity). The processing system 1400 may also include, be coupled with, or be integrated within: a wearable device, such as a smartwatch wearable device; smart glasses or clothing that are enhanced with augmented reality (AR) or virtual reality (VR) features to provide visual, audio, or tactile outputs to supplement real-world visual, audio, or tactile experiences or otherwise provide text, audio, graphics, video, holographic images or video, or tactile feedback; other augmented reality (AR) devices; or other virtual reality (VR) devices. The processing system 1400 may include a television or set-top box device, or may be part of a television or set-top box device. The system 1400 may include, be coupled with, or be integrated within an autonomous vehicle, such as a bus, a tractor-trailer, a car, a motorcycle or electric cycle, an airplane or a glider (or any combination thereof). The autonomous vehicle may use the system 1400 to process the environment sensed around the vehicle. One or more processors 1402 may include one or more processor cores 1407 that are configured to process instructions that, when executed, perform operations for system or user software. At least one of the one or more processor cores 1407 may be configured to process a particular instruction set 1409. The instruction set 1409 may facilitate complex instruction set computing (CISC), reduced instruction set computing (RISC), or computing via very long instruction word (VLIW). The one or more processor cores 1407 may process different instruction sets 1409, and the different instruction sets 1409 may include instructions for facilitating the emulation of other instruction sets. The processor core 1407 may also include other processing devices, such as a digital signal processor (DSP). Processor 1402 may include cache memory 1404. Depending on the architecture, processor 1402 may have a single internal cache or multiple levels of internal caches. In some embodiments, the cache memory is shared among various components of processor 1402. In some embodiments, processor 1402 also uses an external cache (e.g., a third-level (L3) cache or a last-level cache (LLC)) (not shown), and the external cache can be shared among processor cores 1407 using known cache coherence techniques. Register file 1406 may additionally be included in processor 1402 and may include different types of registers for storing different types of data (e.g., integer registers, floating-point registers, status registers, and instruction pointer registers). Some registers may be general-purpose registers, while other registers may be dedicated to the design of processor 1402. One or more processors 1402 may be coupled to one or more interface buses 1410 to transfer communication signals, such as addresses, data, or control signals, between processor 1402 and other components in system 1400. In one of these embodiments, interface bus 1410 may be a processor bus, such as a certain version of the Direct Media Interface (DMI) bus. However, the processor bus is not limited to the DMI bus and may include one or more peripheral component interconnect buses (e.g., PCI, PCI Express), a memory bus, or other types of interface buses. For example, (one or more) processors 1402 may include an integrated memory controller 1416 and a platform controller hub 1430. Memory controller 1416 facilitates communication between memory devices and other components of system 1400, while platform controller hub (PCH) 1430 provides connections to I / O devices via a local I / O bus. The memory device 1420 can be a dynamic random access memory (DRAM) device, a static random-access memory (SRAM) device, a flash memory device, a phase change memory device, or some other memory device with suitable performance to act as the process memory. The memory device 1420 can operate, for example, as the system memory for the system 1400 to store data 1422 and instructions 1421 for use when one or more processors 1402 execute an application or a process. The memory controller 1416 is also coupled to an optional external graphics processor 1418, which can communicate with one or more graphics processors 1408 in the processor 1402 to perform graphics operations and media operations. In some embodiments, the graphics operations, media operations, and / or computing operations can be assisted by an accelerator 1412, which is a coprocessor that can be configured to perform a set of specialized graphics operations, media operations, or computing operations. For example, the accelerator 1412 can be a matrix multiplication accelerator for optimizing machine learning or computing operations. The accelerator 1412 can be a ray tracing accelerator, which can be used to perform ray tracing operations in cooperation with the graphics processor 1408. In one embodiment, an external accelerator 1419 can be used instead of or in cooperation with the accelerator 1412. A display device 1411 can be provided, which can be connected to the processor(s) 1402. The display device 1411 can be one or more of the following: an internal display device such as in a mobile electronic device or a laptop computer device; or an external display device attached via a display interface (e.g., DisplayPort, etc.). The display device 1411 can be a head mounted display (HMD), such as a stereoscopic display device for use in virtual reality (VR) applications or augmented reality (AR) applications. The platform controller hub 1430 enables peripheral devices to be connected to the memory device 1420 and the processor 1402 via a high-speed I / O bus. The I / O peripherals include, but are not limited to, an audio controller 1446, a network controller 1434, a firmware interface 1428, a wireless transceiver 1426, a touch sensor 1425, and a data storage device 1424 (e.g., non-volatile memory, volatile memory, hard disk drive, flash memory, NAND, 3D NAND, 3D XPoint / Optane, etc.). The data storage device 1424 can be connected via a storage interface (e.g., SATA) or via a peripheral bus (such as a peripheral component interconnect bus (e.g., PCI, PCI Express)). The touch sensor 1425 can include a touch screen sensor, a pressure sensor, or a fingerprint sensor. The wireless transceiver 1426 can be a Wi-Fi transceiver, a Bluetooth transceiver, or a mobile network transceiver, such as a 3G, 4G, 5G, or Long-Term Evolution (LTE) transceiver. The firmware interface 1428 enables communication with system firmware and can be, for example, a unified extensible firmware interface (UEFI). The network controller 1434 can enable a network connection to a wired network. In some embodiments, a high-performance network controller (not shown) is coupled to the interface bus 1410. The audio controller 1446 can be a multi-channel high-definition audio controller. In some of these embodiments, the system 1400 includes an optional legacy I / O controller 1440 for coupling legacy (e.g., Personal System 2 (PS / 2)) devices to the system. The platform controller hub 1430 can also be connected to one or more Universal Serial Bus (USB) controllers 1442 to connect to input devices, such as a keyboard and mouse 1443 combination, a camera 1444, or other USB input devices. It will be appreciated that the system 1400 shown is exemplary and not restrictive, as other types of data processing systems configured in different ways can also be used. For example, instances of the memory controller 1416 and the platform controller hub 1430 can be integrated into a discrete external graphics processor, such as the external graphics processor 1418. The platform controller hub 1430 and / or the memory controller 1416 can be external to one or more of the processors 1402. For example, the system 1400 can include an external memory controller 1416 and a platform controller hub 1430 that can be configured as a memory controller hub and a peripheral controller hub within a system chipset that communicates with the (one or more) processors 1402. For example, a circuit board ("sled") can be used, on which components such as a CPU, memory, and other components are placed, and on which the components such as a CPU, memory, and other components are designed to achieve enhanced thermal performance. Processing components such as processors can be located on the top side of the sled, while nearby memory such as DIMMs is located on the bottom side of the sled. As a result of the enhanced airflow provided by this design, the components can operate at higher frequencies and power levels than in a typical system, thereby improving performance. Additionally, the sled is configured for blind mating of power and data communication cables in a rack, thereby enhancing their ability to be quickly removed, upgraded, reinstalled, and / or replaced. Similarly, the individual components located on the sled (such as processors, accelerators, memory, and data storage drives) are configured to be easily upgradable due to their increased spacing from each other. In an illustrative embodiment, the components additionally include hardware authentication features for verifying their authenticity. A data center can utilize a single network architecture ("fabric") that supports multiple other network architectures, including Ethernet and Omni-Path. The sled can be coupled to a switch via optical fiber, which provides higher bandwidth and lower latency than typical twisted-pair cabling (e.g., Category 5, 5e, 6, etc.). Due to the high-bandwidth, low-latency interconnect and network architecture, the data center can, in use, centralize physically dispersed pool resources such as memory, accelerators (e.g., GPUs, graphics accelerators, FPGAs, ASICs, neural network, and / or artificial intelligence accelerators, etc.), and data storage drives, and provide them to computing resources (e.g., processors) as needed, enabling the computing resources to access the centralized resources as if the centralized resources were local. A power supply or power source can provide voltage and / or current to system 1400 or any component or system described herein. In one example, the power supply includes an AC-to-DC (alternating current to direct current) adapter for plugging into a wall outlet. Such AC power can be a renewable energy (e.g., solar) power source. In one example, the power source includes a DC power source, such as an external AC-to-DC converter. The power source or power supply can also include wireless charging hardware for charging via a proximity charging field. The power source can include an internal battery, an AC supply, an action-based power supply, a solar power supply, or a fuel cell source. Figures 15A - 15C Illustrated is a computing system and a graphics processor. Figures 15A - 15CElements with the same or similar names as elements in any other figure in this document describe elements that are the same as those in other figures, can operate or function in a similar manner as those in other figures, may include the same components, and may be linked to other entities, such as those described elsewhere in this document, but are not limited thereto. Figure 15A is a block diagram of a processor 1500, which can be a variant of one of the processors in processor 1402 and can be used in place of one of those processors. Thus, the disclosure of any feature in connection with processor 1500 in this document also discloses the corresponding combination with (one or more) processors 1402, but is not limited thereto. Processor 1500 may have one or more processor cores 1502A - 1502N, an integrated memory controller 1514, and an integrated graphics processor 1508. In the case where the integrated graphics processor 1508 is excluded, a system including the processor will include a graphics processor device within the system chipset or coupled via a system bus. Processor 1500 may include additional cores, and the additional cores are at most the additional core 1502N represented by the dashed box and include the additional core 1502N represented by the dashed box. Each of the processor cores 1502A - 1502N includes one or more internal cache units 1504A - 1504N. In some embodiments, each of the processor cores 1502A - 1502N also has access to one or more shared cache units 1506. The internal cache units 1504A - 1504N and the shared cache units 1506 represent the cache memory hierarchy within processor 1500. The cache memory hierarchy may include at least one level of instruction and data cache within each processor core and one or more levels of shared mid - level cache, such as a second - level (L2), third - level (L3), fourth - level (L4), or other levels of cache, where the highest - level cache before external memory is classified as the LLC. In some embodiments, cache coherence logic maintains coherence between the cache units 1506 and 1504A - 1504N. Processor 1500 may also include a set of one or more bus controller units 1516 and a system agent core 1510. The one or more bus controller units 1516 manage a set of peripheral buses, such as one or more PCI buses or PCI Express buses. The system agent core 1510 provides management functionality for the various processor components. The system agent core 1510 may include one or more integrated memory controllers 1514 for managing access to various external memory devices (not shown). For example, one or more of the processor cores 1502A-1502N may include support for simultaneous multithreading. The system agent core 1510 includes components for coordinating and operating the cores 1502A-1502N during multithreaded processing. The system agent core 1510 may additionally include a power control unit (PCU) that includes logic and components for regulating the power states of the processor cores 1502A-1502N and the graphics processor 1508. The processor 1500 may additionally include a graphics processor 1508 for performing graphics processing operations. In some of these embodiments, the graphics processor 1508 is coupled to a set of shared cache units 1506 and the system agent core 1510, which includes one or more integrated memory controllers 1514. The system agent core 1510 may also include a display controller 1511 for driving the graphics processor output to one or more coupled displays. The display controller 1511 may also be a separate module coupled to the graphics processor via at least one interconnect, or may be integrated within the graphics processor 1508. A ring-based interconnect 1512 can be used to couple the internal components of the processor 1500. However, alternative interconnect units may be used, such as point-to-point interconnects, switched interconnects, or other technologies, including those well known in the art. In some of these embodiments having a ring-based interconnect 1512, the graphics processor 1508 is coupled to the ring-based interconnect 1512 via an I / O link 1513. The exemplary I / O link 1513 represents at least one of a plurality of various I / O interconnects, including an on-package I / O interconnect that facilitates communication between the various processor components and a high-performance memory module 1518, such as an eDRAM module or a high-bandwidth memory (HMB) module. Optionally, each of the processor cores 1502A-1502N and the graphics processor 1508 may use the high-performance memory module 1518 as a shared last-level cache. The processor cores 1502A - 1502N can be, for example, homogeneous cores that execute the same instruction set architecture. Alternatively, the processor cores 1502A - 1502N are heterogeneous in terms of the instruction set architecture (ISA), where one or more of the processor cores 1502A - 1502N execute a first instruction set and at least one of the other cores executes a subset of the first instruction set or a different instruction set. The processor cores 1502A - 1502N can be heterogeneous in terms of the microarchitecture, where one or more cores with a relatively high power consumption are coupled to one or more power cores with a lower power consumption. As another example, the processor cores 1502A - 1502N are heterogeneous in terms of computing power. Additionally, the processor 1500 can be implemented on one or more chips or be implemented as a SoC integrated circuit that also has the illustrated components among other components. Figure 15B is a block diagram of the hardware logic of the graphics processor core block 1519 according to some embodiments described herein. In some embodiments, Figure 15B Elements having the same reference numerals (or names) as elements in any other figure herein can operate or function in a manner similar to the way described elsewhere herein. In one embodiment, the graphics processor core block 1519 is an example of a partition of a graphics processor. The graphics processor core block 1519 can be included in Figure 15A the integrated graphics processor 1508 or a discrete graphics processor, parallel processor, and / or computing accelerator. The graphics processor as described herein can include multiple graphics core blocks based on the target power and performance envelope. Each graphics processor core block 1519 can include a functional block 1530 coupled to a plurality of graphics cores 1521A - 1521F, the plurality of graphics cores 1521A - 1521F including modular blocks of fixed - function logic and general - purpose programmable logic. The graphics processor core block 1519 also includes a shared / cache memory 1536 that can be accessed by all of the graphics cores 1521A - 1521F, rasterizer logic 1537, and additional fixed - function logic 1538. In some embodiments, the functional block 1530 includes a geometry / fixed-function pipeline 1531 that can be shared by all the graphics cores in the graphics processor core block 1519. In embodiments, the geometry / fixed-function pipeline 1531 includes a 3D geometry pipeline, a video front-end unit, a thread generator and a global thread dispatcher, and a unified return buffer manager that manages the unified return buffer. In one embodiment, the functional block 1530 further includes a graphics SoC interface 1532, a graphics microcontroller 1533, and a media pipeline 1534. The graphics SoC interface 1532 provides an interface between the graphics processor core block 1519 and other core blocks within the graphics processor or compute accelerator SoC. The graphics microcontroller 1533 is a programmable sub-processor that can be configured to manage various functions of the graphics processor core block 1519, including thread dispatching, scheduling, and preemption. The media pipeline 1534 includes logic for facilitating the decoding, encoding, preprocessing, and / or post-processing of multimedia data, including image and video data. The media pipeline 1534 implements media operations via requests to the compute or sampling logic within the graphics cores 1521A - 1521F. One or more pixel backends 1535 may also be included within the functional block 1530. The pixel backend 1535 includes buffer memory for storing pixel color values and is capable of performing blending operations and lossless color compression on the rendered pixel data. In one embodiment, the graphics SoC interface 1532 enables the graphics processor core block 1519 to communicate with general-purpose application processor cores (e.g., CPUs) and / or other components within the SoC and / or with a system host CPU coupled to the SoC via a peripheral interface. The graphics SoC interface 1532 also enables communication with off-chip memory hierarchy components such as shared last-level cache memory, system RAM, and / or embedded on-chip or on-package DRAM. The SoC interface 1532 is also capable of enabling communication with fixed-function devices within the SoC such as a camera imaging pipeline and enabling the use and / or implementation of global memory atomicity that can be shared between the graphics processor core block 1519 and a CPU within the SoC. The graphics SoC interface 1532 can also implement power management control for the graphics processor core block 1519 and enable an interface between the clock domain of the graphics processor core block 1519 and other clock domains within the SoC. In one embodiment, the graphics SoC interface 1532 enables receipt of command buffers from a command streamer and a global thread dispatcher that are configured to provide commands and instructions to each of one or more graphics cores within the graphics processor. The commands and instructions can be dispatched to the media pipeline 1534 when media operations are to be performed and can be dispatched to the geometry and fixed-function pipeline 1531 when graphics processing operations are to be performed. When compute operations are to be performed, compute dispatch logic can dispatch commands to the graphics cores 1521A - 1521F, bypassing the geometry pipeline and the media pipeline. The graphics microcontroller 1533 can be configured to perform various scheduling tasks and management tasks for the graphics processor core block 1519. In one embodiment, the graphics microcontroller 1533 can execute graphics workloads and / or compute workloads scheduled on the vector engines 1522A - 1522F, 1524A - 1524F and matrix engines 1523A - 1523F, 1525A - 1525F within the graphics cores 1521A - 1521F. In this scheduling model, host software executing on the CPU core of the SoC including the graphics processor core block 1519 can submit a workload to one of the multiple graphics processor doorbells, which invokes a scheduling operation for the appropriate graphics engine. The scheduling operations include: determining which workload to run next, submitting the workload to the command stream converter, preempting an existing workload running on the engine, monitoring the progress of the workload, and notifying the host software when the workload is complete. In one embodiment, the graphics microcontroller 1533 is also capable of facilitating a low - power or idle state for the graphics processor core block 1519, thereby providing the ability to save and restore registers within the graphics processor core block 1519 across low - power state transitions independent of the operating system and / or graphics driver software on the system. The graphics processor core block 1519 can have more or fewer graphics cores 1521A - 1521F than shown, up to N modular graphics cores. For each set of N graphics cores, the graphics processor core block 1519 can also include: shared / cache memory 1536, which can be configured as shared memory or cache memory; rasterizer logic 1537; and additional fixed - function logic 1538 for accelerating various graphics and compute processing operations. Within each of the graphics cores 1521A - 1521F is a collection of execution resources available to perform graphics operations, media operations, and compute operations in response to requests made by the graphics pipeline, media pipeline, or shader program. The graphics cores 1521A - 1521F include multiple vector engines 1522A - 1522F, 1524A - 1524F, matrix acceleration units 1523A - 1523F, 1525A - 1525D, cache / shared local memory (SLM), samplers 1526A - 1526F, and ray - tracing units 1527A - 1527F. The vector engines 1522A - 1522F, 1524A - 1524F are general - purpose graphics processing units capable of performing floating - point and integer / fixed - point logical operations to serve graphics operations, media operations, or compute operations (including graphics programs, media programs, or compute / GPGPU programs). The vector engines 1522A - 1522F, 1524A - 1524F are capable of operating with variable vector widths using SIMD execution mode, SIMT execution mode, or SIMT + SIMD execution mode. The matrix acceleration units 1523A - 1523F, 1525A - 1525D include matrix - matrix and matrix - vector acceleration logic that improves the performance of matrix operations, especially low - precision and mixed - precision (e.g., INT8, FP16, BF16, FP8) matrix operations for machine learning. In one embodiment, each of the matrix acceleration units 1523A - 1523F, 1525A - 1525D includes one or more systolic arrays of processing elements capable of performing concurrent matrix multiplication or dot - product operations on matrix elements. The samplers 1526A - 1526F are capable of reading media data or texture data into memory and are capable of sampling the data in different ways based on the configured sampler state and the texture / media format being read. Threads executing on the vector engines 1522A - 1522F, 1524A - 1524F or the matrix acceleration units 1523A - 1523F, 1525A - 1525D can utilize the cache / SLM 1528A - 1528F within each of the graphics cores 1521A - 1521F. The cache / SLM 1528A - 1528F can be configured as a pool of cache memory or shared memory local to the respective graphics cores 1521A - 1521F. The ray - tracing units 1527A - 1527F within the graphics cores 1521A - 1521F include ray - traversal / intersection circuitry that uses a bounding - volume hierarchy (BVH) to perform ray traversal and identify intersections between rays and primitives enclosed within the BVH volume. In one embodiment, the ray - tracing units 1527A - 1527F include circuitry for performing depth testing and culling (e.g., using a depth buffer or similar arrangement). In one implementation, the ray - tracing units 1527A - 1527F perform traversal and intersection operations in cooperation with image denoising, at least part of which can be performed using the associated matrix acceleration units 1523A - 1523F, 1525A - 1525D. Figure 15Cis a block diagram of a general - purpose graphics processing unit (GPGPU) 1570 according to an embodiment described herein. The GPGPU 1570 can be configured as a graphics processor (e.g., graphics processor 1508) and / or a computing accelerator. The GPGPU 1570 can be interconnected with a host processor (e.g., one or more CPUs 1546) and memories 1571, 1572 via one or more system and / or memory buses. Memory 1571 can be a system memory that can be shared with one or more CPUs 1546, while memory 1572 is a device memory dedicated to the GPGPU 1570. For example, components within the GPGPU 1570 and memory 1572 can be mapped to memory addresses accessible by one or more CPUs 1546. Access to memories 1571 and 1572 can be facilitated via a memory controller 1568. The memory controller 1568 can include an internal direct memory access (DMA) controller 1569, or can include logic for performing operations otherwise performed by a DMA controller. The GPGPU 1570 includes a plurality of cache memories, which include an L2 cache 1553, an L1 cache 1554, an instruction cache 1555, and a shared memory 1556, at least a portion of which can also be partitioned as a cache memory. The GPGPU 1570 also includes a plurality of computing units 1560A - 1560N. Each computing unit 1560A - 1560N includes a set of vector registers 1561, a set of scalar registers 1562, a set of vector logic units 1563, and a set of scalar logic units 1564. The computing units 1560A - 1560N may also include a local shared memory 1565 and a program counter 1566. The computing units 1560A - 1560N can be coupled to a constant cache 1567, which can be used to store constant data, which is data that will not change during the execution of a kernel program or a shader program on the GPGPU 1570. The constant cache 1567 can be a scalar data cache, and the cached data can be fetched directly into the scalar registers 1562. During operation, one or more CPUs 1546 may write commands to registers in the GPGPU 1570, or to memory in the GPGPU 1570 that has been mapped to an accessible address space. The command processor 1557 may read commands from the registers or memory and determine how to process those commands within the GPGPU 1570. Subsequently, the thread dispatcher 1558 may be used to dispatch threads to the compute units 1560A - 1560N to execute those commands. Each compute unit 1560A - 1560N may execute threads independently of the other compute units. Additionally, each compute unit 1560A - 1560N may be independently configured for conditional computation and may conditionally output the result of the computation to memory. When the submitted commands are complete, the command processor 1557 may interrupt one or more CPUs 1546. Figures 16A - 16C FIG. illustrates a block diagram of an additional graphics processor and compute accelerator architecture provided by embodiments described herein, such as according to Figures 15A - 15C FIG. Figures 16A - 16C Elements having the same or similar names as elements in any other figure in this document describe the same elements as in the other figures, can operate or function in a similar manner as in the other figures, may include the same components, and may be linked to other entities, such as those described elsewhere in this document, but are not limited thereto. Figure 16A is a block diagram of a graphics processor 1600, which may be a discrete graphics processing unit, or may be a graphics processor integrated with multiple processing cores or other semiconductor devices, such as but not limited to memory devices or network interfaces. The graphics processor 1600 may be a variant of the graphics processor 1508 and may be used in place of the graphics processor 1508. Thus, the disclosure of any feature in connection with the graphics processor 1508 herein also discloses the corresponding combination with the graphics processor 1600, but is not limited thereto. The graphics processor may communicate via a memory - mapped I / O interface to registers on the graphics processor and using commands placed in the processor memory. The graphics processor 1600 may include a memory interface 1614 for accessing memory. The memory interface 1614 may be an interface to local memory, one or more internal caches, one or more shared external caches, and / or to system memory. Optionally, the graphics processor 1600 further includes a display controller 1602 for driving display output data to a display device 1618. The display controller 1602 includes hardware for one or more overlay planes for the display and the composition of multi-layer video or user interface elements. The display device 1618 may be an internal or external display device. In one embodiment, the display device 1618 is a head-mounted display device, such as a virtual reality (VR) display device or an augmented reality (AR) display device. The graphics processor 1600 may include a video codec engine 1606 for encoding media into one or more media encoding formats, decoding media from one or more media encoding formats, or transcoding media between one or more media encoding formats, the one or more media encoding formats including but not limited to: Moving Picture Experts Group (MPEG) formats (such as MPEG-2), Advanced Video Coding (AVC) formats (such as H.264 / MPEG-4 AVC, H.265 / HEVC, Alliance for Open Media (AOMedia) VP8, VP9), and Society of Motion Picture & Television Engineers (SMPTE) 421M / VC-1, and Joint Photographic Experts Group (JPEG) formats (such as JPEG, and Motion JPEG (MJPEG) formats). The graphics processor 1600 may include a block image transfer (BLIT) engine 1603 for performing two-dimensional (2D) rasterizer operations, including, for example, bit boundary block transfers. However, alternatively, one or more components of the graphics processing engine (GPE) 1610 may be used to perform 2D graphics operations. In some embodiments, the GPE 1610 is a computing engine for performing graphics operations, the graphics operations including three-dimensional (3D) graphics operations and media operations. The GPE 1610 may include a 3D pipeline 1612 for performing 3D operations, such as rendering three-dimensional images and scenes using processing functions for 3D primitive shapes (e.g., rectangles, triangles, etc.). The 3D pipeline 1612 includes programmable and fixed-function elements that perform various tasks within the element and / or generate execution threads to the 3D / media subsystem 1615. Although the 3D pipeline 1612 can be used to perform media operations, embodiments of the GPE 1610 also include a media pipeline 1616 that is dedicated to performing media operations, such as video post-processing and image enhancement. The media pipeline 1616 may include fixed-function or programmable logic units for performing one or more specialized media operations, such as video decoding acceleration, video deinterlacing, and video encoding acceleration, in place of, or on behalf of, the video codec engine 1606. The media pipeline 1616 may additionally include a thread generation unit for generating threads for execution on the 3D / media subsystem 1615. The generated threads perform computations for media operations on one or more graphics execution units included in the 3D / media subsystem 1615. The 3D / media subsystem 1615 may include logic for executing threads generated by the 3D pipeline 1612 and the media pipeline 1616. The pipeline may send thread execution requests to the 3D / media subsystem 1615, which includes thread dispatch logic for arbitrating and dispatching various requests for available thread execution resources. The execution resources include an array of graphics execution units for processing 3D threads and media threads. The 3D / media subsystem 1615 may include one or more internal caches for thread instructions and data. Additionally, the 3D / media subsystem 1615 may also include shared memory for sharing data between threads and for storing output data, which includes registers and addressable memory. Figure 16B Illustrated is a graphics processor 1620, which is a variant of the graphics processor 1600 and can be used in place of the graphics processor 1600 and vice versa. Thus, any disclosure of a feature in connection with the graphics processor 1600 herein also discloses the corresponding combination with the graphics processor 1620, but is not limited thereto. According to the embodiments described herein, the graphics processor 1620 has a tiled architecture. The graphics processor 1620 may include a graphics processing engine cluster 1622 that has within the graphics engine tiles 1610A - 1610D Figure 16AMultiple instances of the graphics processing engine 1610. Each graphics engine slice 1610A - 1610D can be interconnected via a set of slice interconnections 1623A - 1623F. Each graphics engine slice 1610A - 1610D can also be connected to memory modules or memory devices 1626A - 1626D via memory interconnections 1625A - 1625D. The memory devices 1626A - 1626D can use any graphics memory technology. For example, the memory devices 1626A - 1626D can be Graphics Double Data Rate (GDDR) memory. The memory devices 1626A - 1626D can be High Bandwidth Memory (HBM) modules, which can be on - die with their respective graphics engine slices 1610A - 1610D. The memory devices 1626A - 1626D can be stacked memory devices that can be stacked on top of their respective graphics engine slices 1610A - 1610D. Each graphics engine slice 1610A - 1610D and the associated memory 1626A - 1626D can reside on separate dielets that are bonded to a base die or a base substrate, as further described in Figures 24B - 24D as described in further detail. The graphics processor 1620 can be configured with a non - uniform memory access (NUMA) system in which the memory devices 1626A - 1626D are coupled to the associated graphics engine slices 1610A - 1610D. A given memory device can be accessed by a graphics engine slice different from the graphics engine slice to which the memory device is directly connected. However, the access latency to the memory devices 1626A - 1626D can be minimized when accessing the local slice. In one embodiment, a cache - coherent NUMA (ccNUMA) system is enabled, which uses the slice interconnections 1623A - 1623F to enable communication between cache controllers within the graphics engine slices 1610A - 1610D to maintain a consistent memory image when more than one cache stores the same memory location. The graphics processing engine cluster 1622 can be connected to an on-chip or on-package fabric interconnect 1624. In one embodiment, the fabric interconnect 1624 includes a network processor, a network on a chip (NoC), or another switching processor that enables the fabric interconnect 1624 to act as a packet-switching fabric interconnect for exchanging data packets between components of the graphics processor 1620. The fabric interconnect 1624 can enable communication between the graphics engine slices 1610A-1610D and components such as the video codec engine 1606 and one or more copy engines 1604. The copy engines 1604 can be used to move data out of the memory devices 1626A-1626D and memory external to the graphics processor 1620 (e.g., system memory), move data into the memory devices 1626A-1626D and memory external to the graphics processor 1620 (e.g., system memory), and move data between the memory devices 1626A-1626D and memory external to the graphics processor 1620 (e.g., system memory). The fabric interconnect 1624 can also be used to interconnect the graphics engine slices 1610A-1610D. The graphics processor 1620 can optionally include a display controller 1602 for enabling connection to an external display device 1618. The graphics processor can also be configured as a graphics accelerator or a computing accelerator. In the accelerator configuration, the display controller 1602 and the display device 1618 can be omitted. The graphics processor 1620 can be connected to a host system via a host interface 1628. The host interface 1628 can enable communication between the graphics processor 1620, the system memory, and / or other system components. The host interface 1628 can be, for example, a PCI Express bus or another type of host system interface. For example, the host interface 1628 can be an NVLink or NVSwitch interface. The host interface 1628 and the fabric interconnect 1624 can cooperate to enable multiple instances of the graphics processor 1620 to act as a single logical device. The cooperation between the host interface 1628 and the fabric interconnect 1624 can also enable the individual graphics engine slices 1610A-1610D to present themselves to the host system as different logical graphics devices. Figure 16C Illustrates a computing accelerator 1630 according to an embodiment described herein. The computing accelerator 1630 can include Figure 16BThe architectural similarity of the graphics processor 1620 and is optimized for computing acceleration. The compute engine cluster 1632 may include a set of compute engine slices 1640A - 1640D, and the set of compute engine slices 1640A - 1640D includes execution logic optimized for parallel or vector-based general computing operations. The compute engine slices 1640A - 1640D may not include fixed-function graphics processing logic, but in some embodiments, one or more of the compute engine slices 1640A - 1640D may include logic for performing media acceleration. The compute engine slices 1640A - 1640D may be connected to memories 1626A - 1626D via memory interconnects 1625A - 1625D. The memories 1626A - 1626D and the memory interconnects 1625A - 1625D may be of similar technology as in the graphics processor 1620 or may be different technologies. The compute engine slices 1640A - 1640D may also be interconnected via a set of slice interconnects 1623A - 1623F and may be connected to and / or interconnected by the fabric interconnect 1624. In one embodiment, the compute accelerator 1630 includes a large L3 cache 1636 that can be configured as a device-wide cache. The compute accelerator 1630 can also be connected to the host processor and memory via the host interface 1628 in a manner similar to Figure 16B the graphics processor 1620. The compute accelerator 1630 may also include an integrated network interface 1642. In one embodiment, the integrated network interface 1642 includes a network processor and controller logic that enables the compute engine cluster 1632 to communicate through the physical layer interconnect 1644 without data crossing the memory of the host system. In one embodiment, one of the compute engine slices 1640A - 1640D is replaced by network processor logic, and data to be transmitted or received via the physical layer interconnect 1644 can be directly transmitted to or from the memories 1626A - 1626D. Multiple instances of the compute accelerator 1630 can be combined into a single logical device via the physical layer interconnect 1644. Alternatively, each compute engine slice 1640A - 1640D can be presented as a different network-accessible compute accelerator device. Graphics Processing Engine Figure 17 is a block diagram of a graphics processing engine 1710 of a graphics processor according to some embodiments. The graphics processing engine (GPE) 1710 can be Figure 16A a certain version of the GPE 1610 shown in Figure 16B and can also represent the graphics engine slices 1610A - 1610D of Figure 17Elements with the same or similar names as elements in any other figure in this document describe the same elements as in other figures, can operate or run in a similar manner as in other figures, may include the same components, and may be linked to other entities, such as those described elsewhere in this document, but are not limited thereto. For example, in Figure 17 is also illustrated Figure 16A the 3D pipeline 1612 and the media pipeline 1616. The media pipeline 1616 is optional in some embodiments of the GPE 1710 and may not be explicitly included within the GPE 1710. For example and in at least one embodiment, a separate media and / or image processor is coupled to the GPE 1710. The GPE 1710 may be coupled to or include a command stream converter 1703 that provides a command stream to the 3D pipeline 1612 and / or the media pipeline 1616. Alternatively or additionally, the command stream converter 1703 may be directly coupled to the unified return buffer 1718. The unified return buffer 1718 may be communicatively coupled to the graphics core cluster 1714. Optionally, the command stream converter 1703 is coupled to a memory, which may be system memory, or one or more of internal cache memory and shared cache memory. The command stream converter 1703 may receive commands from the memory and send these commands to the 3D pipeline 1612 and / or the media pipeline 1616. These commands are instructions fetched from a ring buffer that stores commands for the 3D pipeline 1612 and the media pipeline 1616. The ring buffer may additionally include a batch command buffer that stores batches of multiple commands. Commands for the 3D pipeline 1612 may also include references to data stored in the memory, such as but not limited to vertex data and geometry data for the 3D pipeline 1612 and / or image data and memory objects for the media pipeline 1616. The 3D pipeline 1612 and the media pipeline 1616 process commands and data by performing operations via logic within the respective pipelines or by dispatching one or more execution threads to the graphics core cluster 1714. The graphics core cluster 1714 may include one or more graphics core blocks (e.g., graphics core block 1715A, graphics core block 1715B), each block including one or more graphics cores. Each graphics core includes a set of graphics execution resources, which includes: general and graphics-specific execution logic for performing graphics operations and computational operations; and fixed-function texture processing logic and / or machine learning and artificial intelligence acceleration logic. In various embodiments, the 3D pipeline 1612 can include fixed function and programmable logic for processing one or more shader programs by processing instructions and dispatching execution threads to the graphics core cluster 1714, such as vertex shaders, geometry shaders, pixel shaders, fragment shaders, compute shaders, or other shader programs. The graphics core cluster 1714 provides a unified block of execution resources for use in processing these shader programs. The multi-functional execution logic (e.g., execution units) within the graphics core blocks 1715A - 1715B of the graphics core cluster 1714 includes support for various 3D API shader languages and can execute multiple synchronously executed threads associated with multiple shaders. The graphics core cluster 1714 can include execution logic for performing media functions such as video and / or image processing. In addition to graphics processing operations, the execution units can also include general-purpose logic programmable to perform parallel general-purpose computing operations. The general-purpose logic can perform processing operations in parallel or in conjunction with Figure 14 the (one or more) processor cores 1407 or general-purpose logic within the cores 1502A - 1502N as in Figure 15A . Output data generated by threads executing on the graphics core cluster 1714 can output the data to memory in a unified return buffer (URB) 1718. The URB 1718 can store data for multiple threads. The URB 1718 can be used to send data between different threads executing on the graphics core cluster 1714. The URB 1718 can additionally be used for synchronization between threads executing on the graphics core cluster 1714 and fixed function logic within the shared function logic 1720. Optionally, the graphics core cluster 1714 can be scalable such that the array includes a variable number of graphics cores, each with a variable number of execution units based on the target power and performance levels of the GPE 1710. The execution resources can be dynamically scalable such that the execution resources can be enabled or disabled as needed. The graphics core cluster 1714 is coupled to shared function logic 1720, which includes multiple resources shared among the graphics cores in the graphics core array. The shared functions within the shared function logic 1720 are hardware logic units that provide specialized complementary functionality to the graphics core cluster 1714. In various embodiments, the shared function logic 1720 includes, but is not limited to, sampler 1721 logic, math 1722 logic, and inter-thread communication (ITC) 1723 logic. Additionally, one or more caches 1725 can be implemented within the shared function logic 1720. Shared functionality is implemented at least in cases where demand for a given specialized function is insufficient to be included within the graphics core cluster 1714. Instead, a single instantiation of that specialized function is implemented as an independent entity in shared functionality logic 1720 and shared between execution resources within the graphics core cluster 1714. The exact set of functions shared between the graphics core clusters 1714 and included within the graphics core cluster 1714 varies from embodiment to embodiment. Specific shared functions within the shared functionality logic 1720 that are widely used by the graphics core cluster 1714 may be included within the shared functionality logic 1716 within the graphics core cluster 1714. Optionally, the shared functionality logic 1716 within the graphics core cluster 1714 may include some or all of the logic within the shared functionality logic 1720. All logic elements within the shared functionality logic 1720 may be replicated within the shared functionality logic 1716 of the graphics core cluster 1714. Alternatively, the shared functionality logic 1720 is excluded in favor of the shared functionality logic 1716 within the graphics core cluster 1714. Graphics Processing Resources Figures 18A - 18C Execution logic including an array of processing elements employed in a graphics processor is illustrated according to embodiments described herein. Figure 18A Illustrated is a graphics core cluster according to an embodiment. Figure 18B Illustrated is a vector engine of a graphics core according to an embodiment. Figure 18C Illustrated is a matrix engine of a graphics core according to an embodiment. Figures 18A - 18C Elements having the same reference numeral as elements of any other figure herein may operate or function in any manner similar to that described elsewhere herein, but are not limited thereto. For example, Figures 18A - 18C The components can be Figure 15B The graphics processor core block 1519 and / or Figure 17 In one embodiment, Figures 18A - 18C The components have Figure 15A Graphics processor 1508 or Figure 15C The GPGPU 1570 has similar functionality to the equivalent component. like Figure 18A As shown in FIG. 1 , in one embodiment, the graphics core cluster 1714 includes a graphics core block 1715, which may be Figure 17 Graphics core block 1715 may include any number of graphics cores (e.g., graphics core 1815A, graphics core 1815B, all the way up to graphics core 1815N). Multiple instances of graphics core block 1715 may be included. In one embodiment, the elements of graphics cores 1815A-1815N have the sameFigure 15B The components of the graphics cores 1521A - 1521F have similar or equivalent functionality. In such embodiments, each of the graphics cores 1815A - 1815N includes circuit components, including but not limited to: vector engines 1802A - 1802N, matrix engines 1803A - 1803N, memory load / store units 1804A - 1804N, instruction caches 1805A - 1805N, data caches / shared local memories 1806A - 1806N, ray tracing units 1808A - 1808N, samplers 1810A - 1810N. The circuit components of the graphics cores 1815A - 1815N may additionally include fixed function logic 1812A - 1812N. The number of vector engines 1802A - 1802N and matrix engines 1803A - 1803N within the designed graphics cores 1815A - 1815N may vary based on the workload, performance, and power goals for that design. Referring to the graphics core 1815A, the vector engine 1802A and the matrix engine 1803A can be configured to perform parallel computational operations on data in various integer and floating - point data formats based on instructions associated with a shader program. Each of the vector engine 1802A and the matrix engine 1803A can act as a programmable general - purpose computing unit capable of executing multiple synchronous hardware threads and processing multiple data elements in parallel for each thread. The vector engine 1802A and the matrix engine 1803A support processing variable - width vectors in various SIMD widths, including but not limited to SIMD8, SIMD16, and SIMD32. Input data elements can be stored in registers as packed data types, and the vector engine 1802A and the matrix engine 1803A can process each element based on the data size of the element. For example, when operating on a 256 - bit wide vector, the 256 bits of the vector are stored in a register, and the vector is processed as four separate 64 - bit packed data elements (Quad - Word (QW) size data elements), eight separate 32 - bit packed data elements (Double Word (DW) size data elements), sixteen separate 16 - bit packed data elements (Word (W) size data elements), or thirty - two separate 8 - bit data elements (byte (B) size data elements). However, different vector widths and register sizes are possible. In one embodiment, the vector engine 1802A and the matrix engine 1803A can also be configured to perform SIMT operations on thread groups (e.g., 8, 16, or 32 threads) or unit groups of various sizes. Continuing with graphics core 1815A, the memory load / store unit 1804A services memory access requests issued by the vector engine 1802A, matrix engine 1803A, and / or other components of the graphics core 1815A having access to memory. The memory access requests may be processed by the memory load / store unit 1804A to load or store the requested data to / from a cache or memory, or to / from a register file associated with the vector engine 1802A and / or matrix engine 1803A. The memory load / store unit 1804A may also perform prefetch operations. Additionally referring to Figure 19 , in one embodiment, the memory load / store unit 1804A is configured to provide SIMT scatter / gather prefetch or block prefetch for data stored in memory 1910, from memory that is local to other slices via the slice interconnect 1908, or from system memory. Prefetching may be performed on a particular L1 cache (e.g., data cache / shared local memory 1806A), L2 cache 1904, or L3 cache 1906. In one embodiment, prefetching of the L3 cache 1906 automatically causes the data to be stored in the L2 cache 1904. The instruction cache 1805A stores instructions to be executed by the graphics core 1815A. In one embodiment, the graphics core 1815A also includes instruction fetch and prefetch circuitry to fetch or prefetch instructions into the instruction cache 1805A. The graphics core 1815A also includes instruction decoding logic for decoding the instructions within the instruction cache 1805A. The data cache / shared local memory 1806A may be configured as a data cache managed by a cache controller implementing a cache replacement policy and / or as a shared memory that is explicitly managed. The ray tracing unit 1808A includes circuitry for accelerating ray tracing operations. The sampler 1810A provides texture sampling for 3D operations and media sampling for media operations. The fixed function logic 1812A includes fixed function circuitry that is shared among instances of the vector engine 1802A and matrix engine 1803A. The graphics cores 1815B - 1815N can operate in a manner similar to the graphics core 1815A. The functionality of the instruction caches 1805A - 1805N, data caches / shared local memories 1806A - 1806N, ray tracing units 1808A - 1808N, samplers 1810A - 1812N, and fixed function logics 1812A - 1812N corresponds to the equivalent functionality in the graphics processor architectures described herein. For example, the instruction caches 1805A - 1805N can operate in a manner similar to Figure 15CThe instruction cache 1555 operates in a similar manner. The data cache / shared local memories 1806A - 1806N, ray tracing units 1808A - 1808N, and samplers 1810A - 1812N can operate in a manner similar to Figure 15B the caches / SLMs 1528A - 1528F, ray tracing units 1527A - 1527F, and samplers 1526A - 1526F. The fixed function logic 1812A - 1812N can include Figure 15B elements of the geometry / fixed function pipeline 1531 and / or additional fixed function logic 1538 of Figure 3C the ray tracing cores 372. In one embodiment, the ray tracing units 1808A - 1808N include circuitry for performing ray tracing acceleration operations performed by As Figure 18B shown, in one embodiment, the vector engine 1802 includes an instruction fetch unit 1837, a general register file (GRF) array 1824, an architectural register file (ARF) array 1826, a thread arbiter 1822, a dispatch unit 1830, a branch unit 1832, a set of SIMD floating point units (FPU) 1834, and in one embodiment, a set of integer SIMD ALUs 1835. The GRF 1824 and ARF 1826 include a set of general register files and architectural register files associated with each hardware thread that can be active in the vector engine 1802. In one embodiment, the per-thread architectural state is maintained in the ARF 1826, while the data used during thread execution is stored in the GRF 1824. The execution state of each thread, including the instruction pointer for each thread, can be saved in thread-specific registers in the ARF 1826. Register renaming can be used to dynamically allocate registers to hardware threads. In one embodiment, the vector engine 1802 has an architecture that is a combination of Simultaneous Multi-Threading (SMT) and Fine-Grained Interleaved Multi-Threading (IMT). The architecture has a modular configuration that can be fine-tuned at design time based on the target number of synchronized threads and the number of registers per graphics core, where the graphics core resources are divided across the logic for executing multiple synchronized threads. The number of logical threads that can be executed by the vector engine 1802 is not limited to the number of hardware threads, and multiple logical threads can be assigned to each hardware thread. In one embodiment, the vector engine 1802 may issue multiple instructions in concert, and these instructions may each be different instructions. The thread arbiter 1822 may dispatch the instructions to one of the send unit 1830, the branch unit 1832, or the (one or more) SIMD FPUs 1834 for execution. Each execution thread may access 128 general-purpose registers within the GRF 1824, where each register may store 32 bytes that can be accessed as a variable-width vector of 32-bit data elements. In one embodiment, each thread has access to 4 kilobytes within the GRF 1824, but the embodiments are not limited thereto, and more or fewer register resources may be provided in other embodiments. In one embodiment, the vector engine 1802 is partitioned into seven hardware threads that can independently perform computational operations, but the number of threads per vector engine 1802 may also vary according to the embodiment. For example, in one embodiment, up to 16 hardware threads are supported. In an embodiment where seven threads can access 4 kilobytes, the GRF 1824 may store a total of 28 kilobytes. In the case where 16 threads can access 4 kilobytes, the GRF 1824 may store a total of 64 kilobytes. Flexible addressing modes may permit addressing of registers together, effectively creating wider registers or representing strided rectangular block data structures. In one embodiment, memory operations, sampler operations, and other longer-latency system communications are dispatched via "send" instructions executed by the messaging send unit 1830. In one embodiment, branch instructions are dispatched to the dedicated branch unit 1832 to facilitate SIMD scatter and eventual gather. In one embodiment, the vector engine 1802 includes one or more SIMD floating-point units ((one or more) FPUs) 1834 for performing floating-point operations. In one embodiment, the (one or more) FPUs 1834 also support integer computations. In one embodiment, the (one or more) FPUs 1834 may execute up to M 32-bit floating-point (or integer) operations, or execute up to 2M 16-bit integer or 16-bit floating-point operations. In one embodiment, at least one of the (one or more) FPUs provides extended mathematical capabilities that support high-throughput transcendental math functions and double-precision 64-bit floating point. In some embodiments, a set of 8-bit integer SIMD ALUs 1835 also exist and may be specifically optimized to perform operations associated with machine learning computations. In one embodiment, the SIMD ALUs are replaced by a set of additional SIMD FPUs 1834 that can be configured to perform integer and floating-point operations. In one embodiment, the SIMD FPUs 1834 and the SIMD ALUs 1835 can be configured to execute SIMT programs. In one embodiment, combined SIMD+SIMT operations are supported. In one embodiment, an array of multiple instances of the vector engine 1802 may be instantiated in the graphics core. For scalability, the product architect may choose the exact number of vector engines grouped per graphics core. In one embodiment, the vector engine 1802 may execute instructions across multiple execution channels. In a further embodiment, each thread executing on the vector engine 1802 executes on a different channel. As Figure 18C shown, in one embodiment, the matrix engine 1803 includes an array of processing elements configured to perform tensor operations, the tensor operations including vector / matrix operations and matrix / matrix operations such as, but not limited to, matrix multiplication and / or dot product operations. The matrix engine 1803 may be configured with M rows and N columns of processing elements (1852AA - 1852MN), the processing elements (1852AA - 1852MN) including multiplier and adder circuits organized in a pipelined manner. In one embodiment, the processing elements 1852AA - 1852MN form the physical pipeline stages of an N-wide and M-deep systolic array that can be used to perform vector / matrix operations or matrix / matrix operations in a data-parallel manner, including matrix multiplication, fused multiply-add, dot product, or other general matrix-matrix multiplication (GEMM) operations. In one embodiment, the matrix engine 1803 supports 16-bit and 8-bit floating-point operations, as well as 8-bit, 4-bit, 2-bit, and binary integer operations. The matrix engine 1803 may also be configured to accelerate specific machine learning operations. In such embodiments, the matrix engine 1803 may be configured with support for the bfloat (brain floating point) 16-bit floating-point format, or the tensor floating-point 32-bit floating-point format (TF32), which have different numbers of mantissa bits and exponent bits relative to the Institute of Electrical and Electronics Engineers (IEEE) 754 format. In one embodiment, during each cycle, each stage may add the result of the operations performed in that stage to the output of the previous stage. In other embodiments, after a set of compute cycles, the pattern of data movement between processing elements 1852AA - 1852MN may vary based on the instructions or macro-operations being executed. For example, in one embodiment, partial sum loopback is enabled and the processing elements may alternatively add the output of the current cycle to the output generated in the previous cycle. In one embodiment, the final stage of the systolic array may be configured with a loopback to the initial stage of the systolic array. In such embodiments, the number of physical pipeline stages may be decoupled from the number of logical pipeline stages supported by the matrix engine 1803. For example, in the case where the processing elements 1852AA - 1852MN are configured as a systolic array with M physical stages, the loopback from stage M to the initial pipeline stage may enable the processing elements 1852AA - 1852MN to operate as a systolic array with, for example, 2M, 3M, 4M, etc. logical pipeline stages. In one embodiment, matrix engine 1803 includes memories 1841A - 1841N, 1842A - 1842M for storing input data in the form of row and column data for an input matrix. Memories 1842A - 1842M may be configured to store the row elements (A0 - Am) of a first input matrix, and memories 1841A - 1841N may be configured to store the column elements (B0 - Bn) of a second input matrix. The row and column elements are provided as inputs to processing elements 1852AA - 1852MN for processing. In one embodiment, the row and column elements of the input matrix may be stored in the systolic register bank 1840 within matrix engine 1803 before these elements are provided to memories 1841A - 1841N, 1842A - 1842M. In one embodiment, excluding the systolic register bank 1840, and loading memories 1841A - 1841N, 1842A - 1842M from registers in the associated vector engine (e.g., Figure 18B the GRF 1824 of vector engine 1802) or other memories of the graphics core including matrix engine 1803 (e.g., Figure 18A the data cache / shared local memory 1806A for matrix engine 1803A). The results generated by processing elements 1852AA - 1852MN are then output to an output buffer and / or written to a register bank (e.g., systolic register bank 1840, GRF 1824, data cache / shared local memory 1806A - 1806N) for further processing by other functional units of the graphics processor or for output to memory. In some embodiments, matrix engine 1803 is configured to support input sparsity, where multiplication operations in sparse regions of the input data can be bypassed by skipping the multiplication operations of operands with zero values. In one embodiment, processing elements 1852AA - 1852MN are configured to skip the execution of certain operations with zero-valued inputs. In one embodiment, the sparsity within the input matrix can be detected, and operations with known zero output values can be bypassed before being submitted to processing elements 1852AA - 1852MN. Loading zero-valued operands into the processing elements can be bypassed, and processing elements 1852AA - 1852MN can be configured to perform multiplication on non-zero input elements. Matrix engine 1803 can also be configured to support output sparsity, such that operations with results predetermined to be zero are bypassed. For input sparsity and / or output sparsity, in one embodiment, metadata is provided to processing elements 1852AA - 1852MN to indicate which processing elements and / or data channels will be active during a particular processing cycle. In one embodiment, matrix engine 1803 includes hardware for enabling operations on sparse data with a compressed representation of a sparse matrix, which stores non-zero values and metadata identifying the positions of the non-zero values within the matrix. Exemplary compressed representations include, but are not limited to, compressed tensor representations such as compressed sparse row (CSR) representation, compressed sparse column (CSC) representation, compressed sparse fiber (CSF) representation. Support for the compressed representation enables operations to be performed on the input in compressed tensor format without the compressed representation being decompressed or decoded. In such embodiments, operations can be performed only on non-zero input values, and the resulting non-zero output values can be mapped into the output matrix. In some embodiments, hardware support for machine-specific lossless data compression formats is also provided, which are used when transferring data within the hardware or across system buses. Such data can be retained in the compressed format for sparse input data, and matrix engine 1803 can use the compressed metadata for the compressed data to enable operations to be performed only on non-zero values or to bypass blocks of zero data inputs for multiplication operations. In various embodiments, input data may be provided by a programmer in a compressed tensor representation, or a codec may compress the input data into a compressed tensor representation or another sparse data encoding. In addition to supporting compressed tensor representations, streaming compression of sparse input data may be performed before the sparse input data is provided to processing elements 1852AA - 1852MN. In one embodiment, compression is performed on data written to a cache memory associated with the graphics core cluster 1714, where the compression is performed using an encoding supported by the matrix engine 1803. In one embodiment, the matrix engine 1803 includes support for inputs with structured sparsity, in which a predetermined level or pattern of sparsity is imposed on the input data. The data may be compressed to a known compression ratio, where the compressed data is processed by processing elements 1852AA - 1852MN based on metadata associated with the compressed data. Figure 19 FIG. 1900 shows a die 1900 of a multi-die processor according to an embodiment. In one embodiment, die 1900 represents Figure 16B one of the graphics engine dies 1610A - 1610D or Figure 16C one of the compute engine dies 1640A - 1640D. The die 1900 of the multi-die graphics processor includes an array of graphics core clusters (e.g., graphics core clusters 1714A, graphics core clusters 1714B, up to graphics core clusters 1714N), where each graphics core cluster has an array of graphics cores 1815A - 1815N. The die 1900 also includes a global dispatcher 1902 for dispatching threads to the processing resources of the die 1900. The die 1900 may include an L3 cache 1906 and a memory 1910 or be coupled to the L3 cache 1906 and the memory 1910. In various embodiments, the L3 cache 1906 may be excluded, or the die 1900 may include additional levels of cache, such as an L4 cache. In one embodiment, such as Figure 16B and Figure 16C in, each instance of the die 1900 in the multi-die graphics processor has an associated memory 1910. In one embodiment, the multi-die processor may be configured as a multi-chip module, in which the L3 cache 1906 and / or the memory 1910 reside on a separate die different from the graphics core clusters 1714A - 1714N. In this context, a die is an integrated circuit that is at least partially encapsulated and includes different logic units that can be assembled together with other dies into a larger package. For example, the L3 cache 1906 may be included in a dedicated cache die or reside on the same die as the graphics core clusters 1714A - 1714N. In one embodiment, the L3 cache 1906 may be included in as Figure 24Cin the illustrated active base die or active interposer. Memory fabric 1903 enables communication between graphics core clusters 1714A - 1714N, L3 cache 1906, and memory 1910. L2 cache 1904 is coupled to memory fabric 1903 and can be configured to cache transactions executed via memory fabric 1903. The die interconnect 1908 enables communication with other dies on the graphics processor and can be one of the die interconnects 1623A - 1623F of Figure 16B and Figure 16C In embodiments that exclude L3 cache 1906 from die 1900, L2 cache 1904 can be configured as a combined L2 / L3 cache. Memory fabric 1903 can be configured to route data to L3 cache 1906 or to a memory controller associated with memory 1910 based on the presence or absence of L3 cache 1906 in a particular implementation. L3 cache 1906 can be configured as a per-tile cache dedicated to the processing resources of die 1900 or can be part of a GPU-wide L3 cache. Figure 20 is a block diagram illustrating a graphics processor instruction format 2000. The graphics processor execution units support an instruction set with instructions in multiple formats. The solid boxes illustrate components that are typically included in the execution unit instructions, while the dashed boxes include optional or components that are only included in a subset of the instructions. In some embodiments, the described and illustrated graphics processor instruction format 2000 is a macro-instruction because they are the instructions supplied to the execution units, as opposed to micro-operations that result from instruction decoding once the instruction has been processed. Thus, a single instruction can cause the hardware to execute multiple micro-operations. The graphics processor execution units described herein can natively support instructions in a 128-bit instruction format 2010. Based on the selected instructions, instruction options, and number of operands, a 64-bit compact instruction format 2030 can be used for some instructions. The native 128-bit instruction format 2010 provides access to all instruction options, while some options and operations are restricted in the 64-bit format 2030. The native instructions available in the 64-bit format 2030 vary by embodiment. The instruction is partially compressed using a set of index values in index field 2013. The execution unit hardware references a set of compression tables based on the index values and uses the compression table output to reconstruct the native instruction in 128-bit instruction format 2010. Instructions of other sizes and formats can be used. For each format, the instruction opcode 2012 defines the operation for the execution unit to perform. The execution unit executes each instruction in parallel across multiple data elements of each operand. For example, in response to an addition instruction, the execution unit performs a synchronous addition operation across each color channel representing a texture element or a picture element. By default, the execution unit executes each instruction across all data channels of the operand. The instruction control field 2014 can enable control of certain execution options such as channel selection (e.g., predication) and data channel order (e.g., swizzle). For instructions in the 128-bit instruction format 2010, the execution size field 2016 limits the number of data channels to be executed in parallel. The execution size field 2016 may not be available for the 64-bit compact instruction format 2030. Some execution unit instructions have up to three operands, including two source operands src0 2020, src1 2022, and one destination operand (dest 2018). Other instructions such as, for example, data manipulation instructions, dot product instructions, multiply-add instructions, or multiply-accumulate instructions may have a third source operand (e.g., SRC2 2024). The instruction opcode 2012 determines the number of source operands. The last source operand of the instruction can be an immediate (e.g., hard-coded) value passed with the instruction. The execution unit may also support multiple destination instructions, where one or more of the destinations are implicit or implied based on the instruction and / or the specified destinations. The 128-bit instruction format 2010 may include an access / addressing mode field 2026 that, for example, specifies whether to use direct register addressing mode or indirect register addressing mode. When using direct register addressing mode, the register addresses of one or more operands are directly provided by bits in the instruction. The 128-bit instruction format 2010 may also include an access / addressing mode field 2026 that specifies the addressing mode and / or access mode of the instruction. The access mode can be used to define the data access alignment of the instruction. Access modes including a 16-byte aligned access mode and a 1-byte aligned access mode may be supported, where the byte alignment of the access mode determines the access alignment of the instruction operands. For example, when in a first mode, the instruction may use byte-aligned addressing for source and destination operands, and when in a second mode, the instruction may use 16-byte aligned addressing for all source and destination operands. The addressing mode portion of the access / addressing mode field 2026 can determine whether the instruction is to use direct addressing or indirect addressing. When using direct register addressing mode, the bits in the instruction directly provide the register addresses of one or more operands. When using indirect register addressing mode, the register addresses of one or more operands can be calculated based on the address register value and the address immediate field in the instruction. The instructions can be grouped based on the opcode 2012 bit field to simplify opcode decoding 2040. For an 8-bit opcode, bits 4, 5, and 6 allow the execution unit to determine the type of opcode. The exact opcode grouping shown is only an example. The move and logic opcode group 2042 can include data move and logic instructions (e.g., move (mov), compare (cmp)). The move and logic group 2042 can share the five least significant bits (LSB), where the move (mov) instruction takes the form of 0000xxxxb and the logic instruction takes the form of 0001xxxxb. The flow control instruction group 2044 (e.g., call, jump) includes instructions in the form of 0010xxxxb (e.g., 0x20). The miscellaneous instruction group 2046 includes a mixture of instructions, including synchronization instructions (e.g., wait, send) in the form of 0011xxxxb (e.g., 0x30). The parallel math instruction group 2048 includes per-component arithmetic instructions (e.g., add, multiply (mul)) in the form of 0100xxxxb (e.g., 0x40). The parallel math instruction group 2048 performs arithmetic operations in parallel across data channels. The vector math group 2050 includes arithmetic instructions (e.g., dp4) in the form of 0101xxxxb (e.g., 0x50). The vector math group performs arithmetic on vector operands, such as dot product calculations. In one embodiment, the illustrated opcode decoding 2040 can be used to determine which part of the execution unit will be used to execute the decoded instruction. For example, some instructions can be designated as systolic instructions to be executed by a systolic array. Other instructions (such as ray tracing instructions (not shown)) can be routed to a ray tracing core or ray tracing logic within a slice or partition of the execution logic. Graphics Pipeline Figure 21 is a block diagram of a graphics processor 2100 according to another embodiment. Figure 21 Elements having the same or similar names as elements in any other figure in this document describe the same elements as in other figures, can operate or run in a similar manner as in other figures, may include the same components, and may be linked to other entities, such as those described elsewhere in this document, but are not limited thereto. The graphics processor 2100 may include different types of graphics processing pipelines, such as a geometry pipeline 2120, a media pipeline 2130, a display engine 2140, thread execution logic 2150, and a render output pipeline 2170. The graphics processor 2100 may be a graphics processor within a multi-core processing system that includes one or more general-purpose processing cores. The graphics processor may be controlled by register writes to one or more control registers (not shown) or via commands issued through the ring interconnect 2102 to the graphics processor 2100. The ring interconnect 2102 may couple the graphics processor 2100 to other processing components, such as other graphics processors or general-purpose processors. Commands from the ring interconnect 2102 are interpreted by a command stream converter 2103, which supplies instructions to the various components of the geometry pipeline 2120 or the media pipeline 2130. The command stream converter 2103 may direct the operation of a vertex fetcher 2105, which reads vertex data from memory and executes vertex processing commands provided by the command stream converter 2103. The vertex fetcher 2105 may provide the vertex data to a vertex shader 2107, which performs coordinate space transformation and lighting operations on each vertex. The vertex fetcher 2105 and the vertex shader 2107 may execute vertex processing instructions by dispatching execution threads to the graphics cores 2152A - 2152B via a thread dispatcher 2131. The graphics cores 2152A - 2152B may be an array of vector processors having instruction sets for performing graphics operations and media operations. The graphics cores 2152A - 2152B may have attached L1 caches 2151 dedicated to each array or shared between the arrays. The caches may be configured as data caches, instruction caches, or a single cache partitioned to contain data and instructions in different partitions. The geometry pipeline 2120 may include a tessellation component for performing hardware-accelerated tessellation of 3D objects. A programmable hull shader 2111 may configure the tessellation operation. A programmable domain shader 2117 may provide a backend evaluation of the tessellation output. The tessellator 2113 may operate under the direction of the hull shader 2111 and may include dedicated logic for generating a detailed set of geometric objects based on a coarse geometric model that is provided as input to the geometry pipeline 2120. Additionally, if tessellation is not used, the tessellation components (e.g., the hull shader 2111, the tessellator 2113, and the domain shader 2117) may be bypassed. The tessellation components may operate based on data received from the vertex shader 2107. The complete geometric object can be processed by the geometry shader 2119 via one or more threads dispatched to the graphics core 2152A-2152B, or can proceed directly to the clipper 2129. The geometry shader can operate on the entire geometric object, rather than on vertices or vertex patches as in previous stages of the graphics pipeline. If tessellation is disabled, the geometry shader 2119 receives input from the vertex shader 2107. The geometry shader 2119 can be programmable by a geometry shader program to perform geometric tessellation in the case where the tessellation unit is disabled. Before rasterization, the clipper 2129 processes vertex data. The clipper 2129 can be a fixed-function clipper or a programmable clipper with clipping and geometry shader functionality. The rasterizer and depth test component 2173 in the render output pipeline 2170 can dispatch a pixel shader to convert the geometric object into a per-pixel representation. The pixel shader logic can be included in the thread execution logic 2150. Optionally, the application can bypass the rasterizer and depth test component 2173 and access the un-rasterized vertex data via the out-flow unit 2123. The graphics processor 2100 has an interconnect bus, an interconnect fabric, or some other interconnect mechanism that allows data and messages to be passed between the major components of the processor. In some embodiments, the graphics cores 2152A-2152B and associated logic units (e.g., L1 cache 2151, sampler 2154, texture cache 2158, etc.) are interconnected via data ports 2156 to perform memory accesses and communicate with the render output pipeline components of the processor. The sampler 2154, caches 2151, 2158, and graphics cores 2152A-2152B each can have separate memory access paths. Optionally, the texture cache 2158 can also be configured as a sampler cache. The render output pipeline 2170 can include a rasterizer and depth test component 2173 that converts the vertex-based object into an associated pixel-based representation. The rasterizer logic can include a windowizer / masker unit for performing fixed-function triangle and line rasterization. In some embodiments, an associated render cache 2178 and depth cache 2179 are also available. The pixel operation component 2177 performs pixel-based operations on the data, but in some instances, pixel operations associated with 2D operations (e.g., bit-block image transfer with blending) are performed by the 2D engine 2141, or by the display controller 2143 using an overlay display plane during display. The shared L3 cache 2175 can be available to all graphics components, allowing data to be shared without using the main system memory. The media pipeline 2130 may include a media engine 2137 and a video front end 2134. The video front end 2134 may receive pipeline commands from a command stream converter 2103. The media pipeline 2130 may include a separate command stream converter. The video front end 2134 may process the media commands before sending them to the media engine 2137. The media engine 2137 may include thread generation functionality for generating threads to be dispatched to thread execution logic 2150 via a thread dispatcher 2131. The graphics processor 2100 may include a display engine 2140. The display engine 2140 may be external to the processor 2100 and may be coupled to the graphics processor via a ring interconnect 2102, or some other interconnect bus or fabric. The display engine 2140 may include a 2D engine 2141 and a display controller 2143. The display engine 2140 may contain dedicated logic capable of operating independently of the 3D pipeline. The display controller 2143 may be coupled to a display device (not shown), which may be a system-integrated display device such as in a laptop computer or an external display device attached via a display device connector. The geometry pipeline 2120 and the media pipeline 2130 may be configured to perform operations based on multiple graphics and media programming interfaces and are not dedicated to any one application programming interface (API). Driver software for the graphics processor may convert API calls dedicated to a particular graphics or media library into commands that can be processed by the graphics processor. Support may be provided for all of the Open Graphics Library (OpenGL), Open Computing Language (OpenCL), and / or Vulkan graphics and compute APIs from the Khronos Group. Support may also be provided for the Direct3D library from Microsoft Corporation. Combinations of these libraries may be supported. Support may also be provided for the Open Source Computer Vision Library (OpenCV). Future APIs with a compatible 3D pipeline will also be supported if a mapping from the pipeline of the future API to the pipeline of the graphics processor can be made. Graphics Pipeline Programming Figure 22A is a block diagram illustrating a graphics processor command format 2200 for programming a graphics processing pipeline, such as the pipeline described herein in connection with Figure 16A , Figure 17 , Figure 21 described pipeline. Figure 22B is a block diagram illustrating a graphics processor command sequence 2210 in accordance with an embodiment. Figure 22AThe solid - box illustrations therein are generally components included in the graphics commands, while the dashed lines include optional components or components included only in a subset of the graphics commands. Figure 22A An exemplary graphics - processor command format 2200 includes fields for identifying a client 2202 of the command, a command operation code (opcode) 2204, and a data field 2206. A sub - opcode 2205 and a command size 2208 are also included in some commands. The client 2202 can specify the client unit of the graphics device that processes the command data. The graphics - processor command parser can examine the client field of each command to adjust further processing of the command and route the command data to the appropriate client unit. The graphics - processor client units can include a memory - interface unit, a rendering unit, a 2D unit, a 3D unit, and a media unit. Each client unit can have a corresponding processing pipeline for processing commands. Once a command is received by a client unit, the client unit reads the opcode 2204 and the sub - opcode 2205 (if present) to determine the operation to be performed. The client unit uses the information in the data field 2206 to execute the command. For some commands, an explicit command size 2208 is expected to specify the size of the command. The command parser can automatically determine the size of at least some commands based on the command opcode. Commands can be aligned by multiples of double - words. Other command formats can also be used. Figure 22B The flowchart therein illustrates an exemplary graphics - processor command sequence 2210. Software or firmware of a data - processing system featuring an exemplary graphics processor can use a certain version of the shown command sequence to establish, execute, and terminate a set of graphics operations. The sample command sequence is shown and described for illustrative purposes only, and the sample command sequence is not limited to these specific commands or this command sequence. Additionally, commands can be issued in a batch in the command sequence such that the graphics processor will process the command sequence in at least a partially concurrent manner. The graphics - processor command sequence 2210 can start with a pipeline - flush clear command 2212 to cause any active graphics pipeline to complete the current outstanding commands for the pipeline. Optionally, the 3D pipeline 2222 and the media pipeline 2224 may not operate concurrently. Executing the pipeline - flush clear causes the active graphics pipeline to complete any outstanding commands. In response to the pipeline - flush clear, the command parser for the graphics processor will pause command processing until the active drawing engine has completed the outstanding operations and the associated read cache has been invalidated. Optionally, any data marked as "dirty" in the render cache can be flushed to memory. The pipeline - flush clear command 2212 can be used for pipeline synchronization or can be used before placing the graphics processor in a low - power state. The pipeline select command 2213 can be used when a command sequence requires the graphics processor to explicitly switch between pipelines. The pipeline select command 2213 may be needed only once in the execution context before issuing pipeline commands, unless the context is issuing commands for two pipelines. A pipeline dump clear command 2212 may be needed immediately before a pipeline switch via the pipeline select command 2213. The pipeline control command 2214 can be configured to operate the graphics pipeline and can be used to program the 3D pipeline 2222 and the media pipeline 2224. The pipeline control command 2214 can configure the pipeline state for the active pipeline. The pipeline control command 2214 can be used for pipeline synchronization and for clearing data from one or more cache memories within the active pipeline before processing a batch of commands. Commands related to the return buffer state 2216 can be used to configure the set of return buffers for a corresponding pipeline for writing data. Some pipeline operations require the allocation, selection, or configuration of one or more return buffers into which intermediate data is written during processing. The graphics processor can also use one or more return buffers to store output data and perform cross-thread communication. The return buffer state 2216 can include the size and number of return buffers to select for use in a pipeline operation. The remaining commands in the command sequence differ based on the active pipeline for the operation. Based on the pipeline determination 2220, the command sequence is customized for the 3D pipeline 2222 starting with the 3D pipeline state 2230 or the media pipeline 2224 starting at the media pipeline state 2240. Commands for configuring the 3D pipeline state 2230 include 3D state set commands for vertex buffer state, vertex element state, constant color state, depth buffer state, and other state variables that will be configured before processing 3D primitive commands. The values of these commands are determined at least in part based on the particular 3D API in use. The 3D pipeline state 2230 commands may also be able to selectively disable or bypass certain pipeline elements in cases where they will not be used. The 3D primitive 2232 command can be used to submit 3D primitives to be processed by the 3D pipeline. The commands and associated parameters passed to the graphics processor via the 3D primitive 2232 command are forwarded to the vertex fetch function in the graphics pipeline. The vertex fetch function uses the 3D primitive 2232 command data to generate vertex data structures. The vertex data structures are stored in one or more return buffers. The 3D primitive 2232 command can be used to perform vertex operations on 3D primitives via the vertex shader. To process the vertex shader, the 3D pipeline 2222 dispatches shader execution threads to the graphics processor execution units. The 3D pipeline 2222 can be triggered by executing a 2234 command or event. A register write can trigger command execution. Execution can be triggered by a "go" or "kick" command in a command sequence. Command execution can use a pipeline synchronization command to trigger a command sequence dump clear through the graphics pipeline. The 3D pipeline will perform geometric processing on 3D primitives. Once the operation is complete, the resulting geometric object is rasterized, and the pixel engine colors the resulting pixels. For those operations, additional commands for controlling pixel coloring and pixel backend operations may also be included. When performing media operations, the graphics processor command sequence 2210 can follow the media pipeline 2224 path. Generally, the specific use and manner of programming for the media pipeline 2224 depends on the media or computing operation to be performed. During media decoding, specific media decoding operations can be migrated to the media pipeline. The media pipeline can also be bypassed, and media decoding can be performed entirely or partially using the resources provided by one or more general-purpose processing cores. The media pipeline can also include elements for general-purpose graphics processing unit (GPGPU) operations, where the graphics processor is used to perform SIMD vector operations using compute shader programs that are not explicitly related to the rendering of graphics primitives. The media pipeline 2224 can be configured in a manner similar to the 3D pipeline 2222. A set of commands for configuring the media pipeline state 2240 is dispatched or placed in the command queue before the media object commands 2242. The commands for the media pipeline state 2240 can include data for configuring the media pipeline elements that will be used to process the media object. This includes data for configuring video decoding and video encoding logic within the media pipeline, such as the encoding or decoding format. The commands for the media pipeline state 2240 can also support the use of one or more pointers to "indirect" state elements that point to a batch of state settings. The media object commands 2242 can supply pointers to the media objects to be processed by the media pipeline. The media objects include memory buffers that contain the video data to be processed. Optionally, all media pipeline states must be valid before issuing the media object commands 2242. Once the pipeline state is configured and the media object commands 2242 are queued, the media pipeline 2224 is triggered by executing a command 2244 or an equivalent execution event (e.g., a register write). Subsequently, the output from the media pipeline 2224 can be post-processed by operations provided by the 3D pipeline 2222 or the media pipeline 2224. GPGPU operations can be configured and executed in a manner similar to media operations. Graphics Software Architecture Figure 23The figure illustrates an exemplary graphics software architecture for a data processing system 2300. Such a software architecture may include a 3D graphics application 2310, an operating system 2320, and at least one processor 2330. The processor 2330 may include a graphics processor 2332 and one or more general processor cores 2334. The processor 2330 may be a variant of the processor 1402 or any other processor described herein. The processor 2330 may be used in place of the processor 1402 or any other processor described herein. Thus, the disclosure of any feature in combination with the processor 1402 or any other processor described herein also discloses the corresponding combination with the graphics processor 2332, but not limited thereto. In addition, Figure 23 Elements having the same or similar names as elements in any other figure in this document describe the same elements as in other figures, can operate or run in a similar manner as in other figures, may include the same components, and may be linked to other entities, such as those described elsewhere in this document, but not limited thereto. The graphics application 2310 and the operating system 2320 each execute in the system memory 2350 of the data processing system. The 3D graphics application 2310 may include one or more shader programs, and the one or more shader programs include shader instructions 2312. Shader language instructions may be in a high-level shader language, such as the High-Level Shader Language (HLSL) of Direct3D, the OpenGL Shader Language (GLSL), and so on. The application may also include executable instructions 2314 in machine language suitable for execution by the general processor core 2334. The application may also include graphics objects 2316 defined by vertex data. The operating system 2320 may be from Microsoft Corporation An operating system, a proprietary UNIX-like operating system, or an open-source UNIX-like operating system using a Linux kernel variant. The operating system 2320 may support a graphics API 2322, such as, a Direct3D API, an OpenGL API, or a Vulkan API. When the Direct3D API is in use, the operating system 2320 uses a front-end shader compiler 2324 to compile any shader instructions 2312 in HLSL into a lower-level shader language. The compilation may be just-in-time (JIT) compilation or application executable shader pre-compilation. During the compilation of a 3D graphics application 2310, high-level shaders may be compiled into low-level shaders. The shader instructions 2312 can be provided in an intermediate form, such as a version of the Standard Portable Intermediate Representation (SPIR) used by the Vulkan API. A user-mode graphics driver 2326 may include a backend shader compiler 2327 to compile the shader instructions 2312 into a hardware-specific representation. When the OpenGL API is in use, the shader instructions 2312 in the high-level GLSL language are passed to the user-mode graphics driver 2326 for compilation. The user-mode graphics driver 2326 may use the operating system kernel-mode functionality 2328 to communicate with the kernel-mode graphics driver 2329. The kernel-mode graphics driver 2329 may communicate with the graphics processor 2332 to dispatch commands and instructions. IP Core Implementation One or more aspects may be implemented by representative code stored on a machine-readable medium that represents and / or defines logic within an integrated circuit, such as a processor. For example, the machine-readable medium may include instructions that represent various logics within the processor. When read by a machine, the instructions may cause the machine to fabricate logic for performing the techniques described herein. Such representations (referred to as "IP cores") are reusable units of the logic of an integrated circuit, and these reusable units may be stored on a tangible, machine-readable medium as a hardware model that describes the structure of the integrated circuit. The hardware model may be supplied to various customers or manufacturing facilities that load the hardware model on a manufacturing machine for fabricating the integrated circuit. The integrated circuit may be fabricated such that the circuit performs the operations described in association with any of the embodiments described herein. Figure 24AFIG. is a block diagram of an IP core development system 2400 that can be used to fabricate integrated circuits to perform operations. The IP core development system 2400 can be used to generate modular, reusable designs that can be incorporated into larger designs or used to construct an entire integrated circuit (e.g., a system-on-a-chip integrated circuit). A design facility 2430 can generate a software simulation 2410 of an IP core design in a high-level programming language (e.g., C / C++). The software simulation 2410 can be used to design, test, and verify the behavior of the IP core using a simulation model 2412. The simulation model 2412 can include functional simulation, behavioral simulation, and / or timing simulation. Subsequently, a register transfer level (RTL) design 2415 can be created or synthesized from the simulation model 2412. The RTL design 2415 is an abstraction of the behavior of an integrated circuit that models the flow of digital signals between hardware registers (including the associated logic that is executed using the modeled digital signals). In addition to the RTL design 2415, lower-level designs at the logic level or transistor level can also be created, designed, or synthesized. Thus, the specific details of the initial design and simulation can vary. The RTL design 2415 or an equivalent can be further synthesized by the design facility into a hardware model 2420, which can be in a hardware description language (HDL) or some other representation of physical design data. The HDL can be further simulated or tested to verify the IP core design. A non-volatile memory 2440 (e.g., a hard disk, flash memory, or any non-volatile storage medium) can be used to store the IP core design for delivery to a third-party manufacturing facility 2465. Alternatively, the IP core design can be transmitted via a wired connection 2450 or a wireless connection 2460 (e.g., via the Internet). The manufacturing facility 2465 can then fabricate an integrated circuit that is at least partially based on the IP core design. The fabricated integrated circuit can be configured to perform operations in accordance with at least one embodiment described herein. Figure 24BCross-sectional side view of an illustrated integrated circuit package component 2470. The integrated circuit package component 2470 illustrates an implementation of one or more processor or accelerator devices as described herein. The package component 2470 includes a plurality of hardware logic units 2472, 2474 connected to a substrate 2480. The logic 2472, 2474 may be implemented at least in part in configurable logic or fixed functional logic hardware and may include one or more portions of any of the (one or more) processor cores, (one or more) graphics processors, or other accelerator devices described herein. Each logic unit 2472, 2474 may be implemented within a semiconductor die and is coupled to the substrate 2480 via an interconnect structure 2473. The interconnect structure 2473 may be configured to route electrical signals between the logic 2472, 2474 and the substrate 2480 and may include interconnects such as, but not limited to, bumps or pillars. The interconnect structure 2473 may be configured to route electrical signals such as, for example, input / output (I / O) signals and / or power or ground signals associated with the operation of the logic 2472, 2474. Optionally, the substrate 2480 may be an epoxy-based laminated substrate. The substrate 2480 may also include other suitable types of substrates. The package component 2470 may be connected to other electrical devices via package interconnects 2483. The package interconnects 2483 may be coupled to the surface of the substrate 2480 to route electrical signals to other electrical devices such as a motherboard, other chip sets, or a multi-chip module. The logic units 2472, 2474 may be electrically coupled to a bridge 2482 configured to route electrical signals between the logic 2472, logic 2474. The bridge 2482 may be a dense interconnect structure that provides routing for electrical signals. The bridge 2482 may include a bridge substrate made of glass or a suitable semiconductor material. Circuitry features may be formed on the bridge substrate to provide a chip-to-chip connection between the logic 2472, logic 2474. Although two logic units 2472, 2474 and a bridge 2482 are illustrated, embodiments described herein may include more or fewer logic units on one or more dies. The one or more dies may be connected by zero or more bridges as the bridge 2482 may be excluded when the logic is included on a single die. Alternatively, multiple dies or logic units may be connected by one or more bridges. Additionally, multiple logic units, dies, and bridges may be connected together in other possible configurations including three-dimensional configurations. Figure 24CIllustrated is a packaged component 2490 that includes hardware logic die that connect to a substrate 2480 (e.g., a base die) for multiple units. Graphics processing units, parallel processors, and / or compute accelerators as described herein may be composed of various silicon die fabricated separately. In this context, a die is an integrated circuit that is at least partially encapsulated and includes different logic units that may be assembled with other die into a larger package. Die with various sets of different IP core logic may be assembled into a single device. Additionally, die may be integrated into a base die or base die using active interposer technology. The concepts described herein enable interconnection and communication between different forms of IP within a GPU. IP cores may be fabricated using different process technologies and configured during fabrication, which avoids the complexity of converging multiple IPs onto the same fabrication process (especially on large SoCs with several styles of IP). Allowing the use of multiple process technologies improves time to market and provides a cost-effective way to create multiple product SKUs. Additionally, decomposed IP is more easily modified to be independently power gated, and components not in use for a given workload may be turned off, thereby reducing overall power consumption. In embodiments, the packaged component 2490 may include fewer or more components and die interconnected by a structure 2485 or one or more bridges 2487. The die within the packaged component 2490 may have a 2.5D arrangement using a Chip-on-Wafer-on-Substrate stack, where multiple die are stacked side-by-side on a silicon interposer that includes through-silicon vias (TSVs) to couple the die to the substrate 2480, which includes electrical connections to the package interconnect 2483. In one embodiment, the silicon interposer is an active interposer 2489 that includes embedded logic in addition to the TSVs. In such embodiments, the die within the packaged component 2490 are arranged on top of the active interposer 2489 using 3D face-to-face die stacking. The active interposer 2489 may include hardware logic for I / O 2491, cache memory 2492, and other hardware logic 2493 in addition to the interconnect structure 2485 and silicon bridges 2487. The structure 2485 enables communication between the various logic die 2472, 2474 and the logic 2491, 2493 within the active interposer 2489. The structure 2485 may be a NoC interconnect or another form of packet-switched fabric that exchanges data packets between components of the packaged component. For complex components, the structure 2485 may be a dedicated die that enables communication between the various hardware logics of the packaged component 2490. The bridge structure 2487 within the active interposer 2489 can be used to facilitate point-to-point interconnections, for example, between logic or I / O dies 2474 and memory dies 2475. In some implementations, the bridge structure 2487 can also be embedded within the substrate 2480. Hardware logic dies can include dedicated hardware logic dies 2472, logic or I / O dies 2474, and / or memory dies 2475. The hardware logic die 2472 and the logic or I / O die 2474 can be implemented at least partially in configurable logic or fixed-functional logic hardware and can include any one or more portions of the (one or more) processor cores, (one or more) graphics processors, parallel processors, or other accelerator devices described herein. The memory die 2475 can be DRAM (e.g., GDDR, HBM) memory or cache (SRAM) memory. The cache memory 2492 within the active interposer 2489 (or substrate 2480) can act as a global cache for the packaged components 2490, act as part of a distributed global cache, or act as a dedicated cache for the structure 2485. Each die can be fabricated as a separate semiconductor die and can be coupled to a base die that is embedded within or coupled to the substrate 2480. The coupling to the substrate 2480 can be performed via the interconnecting fabric 2473. The interconnecting fabric 2473 can be configured to route electrical signals between the various dies and the logic within the substrate 2480. The interconnecting fabric 2473 can include interconnects such as, but not limited to, bumps or posts. In some embodiments, the interconnecting fabric 2473 can be configured to route electrical signals such as, for example, input / output (I / O) signals and / or power or ground signals associated with the operation of the logic, I / O, and memory dies. In one embodiment, additional interconnecting fabric couples the active interposer 2489 to the substrate 2480. The substrate 2480 can be an epoxy-based laminated substrate; however, it is not limited thereto, and the substrate 2480 can also include other suitable types of substrates. The packaged component 2490 can be connected to other electrical devices via the package interconnect 2483. The package interconnect 2483 can be coupled to the surface of the substrate 2480 to route electrical signals to other electrical devices such as a motherboard, other chip sets, or a multi-chip module. The logic or I / O die 2474 and the memory die 2475 can be electrically coupled via a bridge 2487, which is configured to route electrical signals between the logic or I / O die 2474 and the memory die 2475. The bridge 2487 can be a dense interconnect fabric that provides routing for electrical signals. The bridge 2487 can include a bridge substrate made of glass or a suitable semiconductor material. Circuitry features can be formed on the bridge substrate to provide die-to-die connections between the logic or I / O die 2474 and the memory die 2475. The bridge 2487 can also be referred to as a silicon bridge or an interconnect bridge. For example, the bridge 2487 is an Embedded Multi-die Interconnect Bridge (EMIB). Alternatively, the bridge 2487 can simply be a direct connection from one die to another die. Figure 24D FIG. illustrates a package assembly 2494 that includes interchangeable dies 2495 according to an embodiment. The interchangeable dies 2495 can be assembled into standardized slots on one or more base dies 2496, 2498. The base dies 2496, 2498 can be coupled via a bridge interconnect 2497, which can be similar to other bridge interconnects described herein and can be, for example, an EMIB. Memory dies can also be connected to logic or I / O dies via bridge interconnects. The I / O and logic dies can communicate via an interconnect structure. Each of the base dies can support one or more slots in a standardized format for either logic or I / O or memory / cache. SRAM and power delivery circuitry can be fabricated into one or more of the base dies 2496, 2498, which can be fabricated using a different process technology relative to the interchangeable die 2495, with the interchangeable die 2495 stacked on top of the base die. For example, the base dies 2496, 2498 can be fabricated using a larger process technology while the interchangeable die can be fabricated using a smaller process technology. One or more of the interchangeable dies 2495 can be memory (e.g., DRAM) dies. Different memory densities can be selected for the package assembly 2494 based on the power and / or performance of the product using the package assembly 2494. Additionally, logic dies with different numbers of types of functional units can be selected based on the power and / or performance of the product at the time of assembly. Additionally, dies containing different types of IP logic cores can be inserted into the interchangeable die slots, enabling a hybrid processor design that can mix and match IP blocks of different technologies. Exemplary System - on - Chip Integrated Circuit Figures 25 - 26BFIG. illustrates an exemplary integrated circuit and associated graphics processor that can be fabricated using one or more IP cores. In addition to what is illustrated, other logic and circuitry can be included, including additional graphics processor / cores, peripheral interface controllers, or general purpose processor cores. Figures 25 - 26B Elements described with the same or similar names as elements in any other figure herein describe the same elements as in the other figures, can operate or function in a similar manner as in the other figures, can include the same components, and can be linked to other entities, such as those described elsewhere herein, but are not limited thereto. Figure 25 is a block diagram of an exemplary system-on-chip integrated circuit 2500 that can be fabricated using one or more IP cores. The exemplary integrated circuit 2500 includes one or more application processors 2505 (e.g., CPUs), at least one graphics processor 2510, which can be a variant of graphics processors 1408, 1508, 2510, or can be a variant of any graphics processor described herein and can be used in place of any of the described graphics processors. Thus, the disclosure of any feature in connection with a graphics processor herein also discloses the corresponding combination with graphics processor 2510, but is not limited thereto. The integrated circuit 2500 can additionally include an image processor 2515 and / or a video processor 2520, either of which can be a modular IP core from the same design facility or multiple different design facilities. The integrated circuit 2500 can include peripheral or bus logic, including a USB controller 2525, a UART controller 2530, an SPI / SDIO controller 2535, and an 2 S / I 2 C controller 2540. Additionally, the integrated circuit can include a display device 2545 that is coupled to one or more of a high-definition multimedia interface (HDMI) controller 2550 and a mobile industry processor interface (MIPI) display interface 2555. Storage can be provided by a flash memory subsystem 2560 (including flash memory and a flash memory controller). A memory interface can be provided via a memory controller 2565 to access SDRAM or SRAM memory devices. Some integrated circuits additionally include an embedded security engine 2570. Figures 26A - 26Bis a block diagram of an exemplary graphics processor for use within a SoC in accordance with an embodiment described herein. The illustrated graphics processor may be a variant of graphics processor 1408, 1508, 2510, or any other graphics processor described herein. The graphics processor may be used in place of graphics processor 1408, 1508, 2510, or any other graphics processor described herein. Thus, the disclosure of any feature in connection with graphics processor 1408, 1508, 2510, or any other graphics processor described herein also discloses the corresponding combination with a Figures 26A - 26B graphics processor, but is not limited thereto. Figure 26A illustrates an exemplary graphics processor 2610 of a system-on-chip integrated circuit that may be fabricated using one or more IP cores in accordance with an embodiment. Figure 26B illustrates an additional exemplary graphics processor 2640 of a system-on-chip integrated circuit that may be fabricated using one or more IP cores in accordance with an embodiment. Figure 26A The graphics processor 2610 is an example of a low-power graphics processor core. Figure 26B The graphics processor 2640 is an example of a higher-performance graphics processor core. For example, as mentioned at the beginning of this paragraph, each of the graphics processors 2610 and 2640 may be a Figure 25 variant of the graphics processor 2510. As Figure 26A shown, the graphics processor 2610 includes a vertex processor 2605 and one or more fragment processors 2615A - 2615N (e.g., 2615A, 2615B, 2615C, 2615D, up to 2615N - 1 and 2615N). The graphics processor 2610 may execute different shader programs via separate logic such that the vertex processor 2605 is optimized to perform operations for a vertex shader program, while the one or more fragment processors 2615A - 2615N perform fragment (e.g., pixel) shading operations for a fragment or pixel shader program. The vertex processor 2605 executes the vertex processing stage of the 3D graphics pipeline and generates primitives and vertex data. The (one or more) fragment processors 2615A - 2615N use the primitive data and vertex data generated by the vertex processor 2605 to produce a frame buffer that is displayed on a display device. The (one or more) fragment processors 2615A - 2615N may be optimized to execute fragment shader programs as provided in the OpenGL API, which may be used to perform operations similar to pixel shader programs as provided in the Direct 3D API. The graphics processor 2610 additionally includes one or more memory management units (MMUs) 2620A - 2620B, (one or more) caches 2625A - 2625B, and (one or more) circuit interconnections 2630A - 2630B. The one or more MMUs 2620A - 2620B provide virtual - to - physical address mapping for the graphics processor 2610 (including for the vertex processor 2605 and / or (one or more) fragment processors 2615A - 2615N), and this virtual - to - physical address mapping can reference vertex data or image / texture data stored in memory in addition to the vertex data or image / texture data stored in one or more caches 2625A - 2625B. The one or more MMUs 2620A - 2620B can be synchronized with other MMUs within the system so that each processor 2505 - 2520 can participate in a shared or unified virtual memory system, and the other MMUs within the system include one or more MMUs associated with Figure 25 one or more application processors 2505, image processors 2515, and / or video processors 2520. The components of the graphics processor 2610 can correspond to the components of other graphics processors described herein. The one or more MMUs 2620A - 2620B can correspond to Figure 2C the MMU 245. The vertex processor 2605 and fragment processors 2615A - 2615N can correspond to the graphics multiprocessor 234. According to an embodiment, the one or more circuit interconnections 2630A - 2630B enable the graphics processor 2610 to interface with other IP cores within the SoC via the internal bus of the SoC or via a direct connection. The one or more circuit interconnections 2630A - 2630B can correspond to Figure 2C the data cross - switch 240. Further correspondences can be found between the similar components of the graphics processor 2610 and the various graphics processor architectures described herein. As Figure 26B shown, the graphics processor 2640 includes Figure 26AOne or more MMUs 2620A - 2620B, caches 2625A - 2625B, and circuit interconnects 2630A - 2630B of the graphics processor 2610. The graphics processor 2640 includes one or more shader cores 2655A - 2655N (e.g., 2655A, 2655B, 2655C, 2655D, 2655E, 2655F, up to 2655N - 1 and 2655N), which provide a unified shader core architecture where a single core or type of core can execute all types of programmable shader code, including shader program code for implementing vertex shaders, fragment shaders, and / or compute shaders. The exact number of shader cores present can vary depending on the embodiment and implementation. Additionally, the graphics processor 2640 includes an inter - core task manager 2645, which acts as a thread dispatcher for dispatching execution threads to one or more shader cores 2655A - 2655N and a tiling unit 2658 for accelerating tiling operations for tile - based rendering, in which rendering operations for a scene are subdivided in image space, e.g., to take advantage of local spatial coherence within the scene or to optimize the use of internal caches. The shader cores 2655A - 2655N can, for example, correspond to Figure 2D the graphics multiprocessor 234 in Figure 3A and Figure 3B the graphics multiprocessors 325, 350 respectively, or correspond to Figure 3C the multi - core group 365A in Data Processing System for GPGPU Program Execution Figure 27 is a block diagram of a data processing system 2700 according to an embodiment. The data processing system 2700 is a heterogeneous processing system having a processor 2702, unified memory 2710, and a GPGPU 2720 including machine learning acceleration logic. The processor 2702 and the GPGPU 2720 can be any processor and GPGPU / parallel processor as described herein. For example, additionally referring to Figure 1 , the application processor 2702 can be a variant of the processors in one or more of the illustrated processors 102 and / or share an architecture with the processors in one or more of the illustrated processors 102. The GPGPU 2720 can be a variant of the parallel processors in one or more of the illustrated parallel processors 112 and / or share an architecture with the parallel processors in one or more of the illustrated parallel processors 112. Additionally referring to Figure 14, the processor 2702 can be a variant of one of the illustrated processors 1402 and / or share an architecture with one of the illustrated processors 1402, and the GPGPU 2720 can be a variant of one of the illustrated graphics processors 1408 and / or share a system architecture with one of the illustrated graphics processors 1408. The application processor 2702 can execute the instructions of the compiler 2715 stored in the system memory 2712. In one embodiment, the compiler 2715 executes on the application processor 2702 to compile the source code 2714A into the compiled code 2714B. The compiled code 2714B can include instructions executable by the application processor 2702 and / or instructions executable by the GPGPU 2720. Compilation of the instructions to be executed by the GPGPU can be facilitated by a shader or compute program compiler (such as the shader compiler 2327 and / or the shader compiler 2324 in Figure 23 ). In one embodiment, compilation of the instructions to be executed by the GPGPU can alternatively or additionally be performed at least in part via the execution logic within the GPGPU 2720. For example, the compiler 2715 can migrate certain program code analysis, translation, or compilation operations to the GPGPU 2720. During compilation, the compiler 2715 can perform operations to insert metadata, which includes hints regarding the level of data parallelism present in the compiled code 2714B and / or hints regarding the data locality associated with the threads to be dispatched based on the compiled code 2714B. The compiler 2715 can include the information necessary to perform such operations, or the operations can be performed with the assistance of the runtime library 2716. The runtime library 2716 can also assist the compiler 2715 in compiling the source code 2714A and can also include instructions that link with the compiled code 2714B at runtime to facilitate the execution of the compiled instructions on the GPGPU 2720. The compiler 2715 can also facilitate register allocation of variables via a register allocator (RA) and generate load and store instructions to move the data of the variables between memory and the registers assigned to the variables. The unified memory 2710 represents a unified address space that can be accessed by the processor 2702 and the GPGPU 2720. The unified memory may include system memory 2712 and GPGPU memory 2718. The GPGPU memory 2718 is memory within the address space of the GPGPU 2720 and may include some or all of the system memory 2712. In one embodiment, the compiled code 2714B stored in the system memory 2712 may be mapped into the GPGPU memory 2718 for access by the GPGPU 2720. The GPGPU memory 2718 also includes the GPGPU local memory 2728 of the GPGPU 2720. The GPGPU local memory 2728 may include, for example, HBM or GDDR memory. The GPGPU 2720 includes a plurality of compute blocks 2724A - 2724N, which may include one or more of the various processing resources described herein. The processing resources may be or include various different computing resources such as, for example, execution units, compute units, streaming multiprocessors, graphics multiprocessors, or multi - core groups. In one embodiment, the GPGPU 2720 additionally includes a tensor accelerator 2723 (e.g., a matrix accelerator), which may include one or more special - function compute units designed to accelerate a subset of matrix operations (e.g., dot products, etc.). Some functions to be executed by the compute blocks 2724A - 2724N may be directly scheduled to or migrated to the tensor accelerator 2723. In various embodiments, the tensor accelerator 2723 includes processing - element logic configured to efficiently execute matrix - computation operations such as multiply - and - add operations and dot - product operations used by 3D graphics or compute shader programs. In one embodiment, the tensor accelerator 2723 is capable of being configured to accelerate operations used by machine - learning frameworks. In one embodiment, the tensor accelerator 2723 is an application - specific integrated circuit specifically configured to perform a particular set of parallel matrix - multiplication and / or addition operations. In one embodiment, the tensor accelerator 2723 is a field - programmable gate array (FPGA) providing fixed - function logic that can be updated between workloads. In one embodiment, the set of compute operations that can be performed by the tensor accelerator 2723 may be limited relative to the operations that can be performed by the compute blocks 2724A - 2724N. However, the tensor accelerator 2723 is capable of performing parallel tensor operations with a significantly higher throughput relative to the compute blocks 2724A - 2724N. The tensor accelerator 2723 may also be referred to as a tensor accelerator or tensor core. In one embodiment, the logic components within the tensor accelerator 2723 may be distributed across the processing resources of the plurality of compute blocks 2724A - 2724N rather than being centrally located within a single circuit. The GPGPU 2720 may further include a set of resources that can be shared by the compute blocks 2724A - 2724N and the tensor accelerator 2723. The set of resources includes, but is not limited to, a set of registers 2725, a power and performance module 2726, and a cache 2727. In one embodiment, the registers 2725 include directly accessible registers and indirectly accessible registers, where the indirectly accessible registers are optimized for use by the tensor accelerator 2723. The power and performance module 2726 may be configured to adjust the power delivery and clock frequency of the compute blocks 2724A - 2724N to power gate idle components within the compute blocks 2724A - 2724N. In various embodiments, the cache 2727 may include an instruction cache and / or a lower - level data cache. The GPGPU 2720 may additionally include an L3 data cache 2730, which may be used to cache data accessed from the unified memory 2710 by the compute elements within the tensor accelerator 2723 and / or the compute blocks 2724A - 2724N. In one embodiment, the L3 data cache 2730 includes a shared local memory 2732, which can be shared by the compute elements within the compute blocks 2724A - 2724N and the tensor accelerator 2723. In one embodiment, the GPGPU 2720 includes instruction handling logic, such as a fetch and decode unit 2721 and a scheduler controller 2722. The fetch and decode unit 2721 includes a fetch unit and a decode unit to fetch and decode instructions for execution by one or more of the tensor accelerator 2723 or the compute blocks 2724A - 2724N. The instructions may be scheduled via the scheduler controller 2722 to appropriate functional units within the compute blocks 2724A - 2724N or the tensor accelerator. In one embodiment, the scheduler controller 2722 is an ASIC configurable to perform advanced scheduling operations. In one embodiment, the scheduler controller 2722 is a microcontroller or per - instruction low - energy processor configured to execute scheduling instructions loaded from a firmware module. Figure 28 Illustrates a processing resource architecture 2800 according to an embodiment. The processing resource architecture 2800 includes functional units that can be found in a graphics core (such as Figure 18A any of the graphics cores 1815A - 1815N of Figure 18B the vector engine 1802 and Figure 18CComponents of the matrix engine 1803. In one embodiment, a graphics core (e.g., graphics core 1815A) may include multiple instances of processing resources having a processing resource architecture 2800. A local thread dispatcher (TDL 2801) may dispatch instructions to a set 2802 of ready instruction queues associated with the hardware threads (T0-Tn) of each processing resource. Threads may be assigned to processing resources in a round-robin manner. If the instructions of a thread are not ready for execution, the dispatching of that thread may be skipped. If a processing resource does not have sufficient resources to accept an incoming thread, the TDL 2801 may skip that processing resource. Each hardware thread (T0-Tn) is a SIMD thread that can execute a single instruction on data from multiple channels. Instruction data can be read from registers in the register file 2810 and written to registers in the register file 2810. Intermediate data for calculations can be stored in an array of accumulators 2811. In various embodiments, the register file 2810 may be configured with the characteristics of other GPU register files described herein (e.g., register file 258, register files 334A-334B, register file 369, register 445, etc.). The register file 2810 may include vector registers (e.g., vector register 1561) and, in one embodiment, a limited number of scalar registers (scalar register 1562). In one embodiment, the register file 2810 may include a general register file array, such as GRF 1824. In one embodiment, the number of hardware threads is configurable, where a configurable number of registers within the register file 2810 are assigned to each thread. In one embodiment, SIMT execution is enabled by assigning multiple SIMT threads to a single SIMD hardware thread. Data elements associated with multiple SIMT threads may be mapped to multiple sub-registers within a single SIMD register file. A thread control unit 2832 controls instruction execution via the hardware threads. The thread control unit 2832 includes a thread arbiter 2804 that arbitrates which ready instructions are assigned to instruction queues 2806 associated with floating point (FPU), integer (INT), or matrix pipelines (Matrix). For the source operand read phase of instruction execution for instructions within the instruction queue 2806, the source read arbiter 2808 arbitrates access to the register file and the accumulators. Operands read from the register file 2810 can be buffered within the reuse buffer 2812, which is coupled to the matrix pipeline and the functional units of the FPU and INT. The reuse buffer 2812 can act as an operand cache that facilitates reuse of operand data by those functional units. In one embodiment, the reuse buffer 2812 is multi-level deep, allowing caching of source operands from multiple instructions. The processing resource architecture 2800 includes multiple different types of functional units, including a matrix unit 2814, a single-precision floating-point (FP32) unit 2816, a high-throughput double-precision floating-point (FP64) 2817, a low-throughput FP64 unit 2819, and an integer unit 2820. The high-throughput FP64 2817 unit can be configured to execute instructions intended for high-performance computing (HPC) and scientific computing workloads and may be found primarily in compute accelerators and graphics processors designed for servers and / or workstations. The low-throughput FP64 unit 2819 may be found primarily in consumer-grade graphics processors and / or compute accelerators. At least one embodiment can include both the high-throughput FP64 2817 unit and the low-throughput FP64 unit 2819. The processing resource architecture may also include an ARF 2818, which contains configuration registers, architectural registers, thread-specific data, and other non-general-purpose registers. The ARF 2818 can be similar to Figure 18B the ARF 1826. Different instructions are assigned to different functional units based on the instruction type. Matrix operations (such as dot product or matrix multiplication operations) are assigned to the matrix unit 2814, which can be an instance of the matrix engine 1803 of Figure 18C or another matrix unit described herein, such as one or more of the tensor cores 371 in Figure 3C . Single-precision floating-point instructions can be executed by the FP32 unit 2816 or the high-throughput FP64 unit 2817. Double-precision floating-point instructions can be executed by the high-throughput FP64 unit 2817 or the low-throughput FP64 unit 2819. Integer operations are executed by the integer unit 2820. In one embodiment, the FP32 unit 2816 and the high-throughput FP64 unit 2817 are grouped into a first ALU / FPU (e.g., FPU0), while the low-throughput FP64 unit 2819 and the integer unit 2820 are grouped into a second ALU / FPU (e.g., FPU1). Outputs from these functional units can be stored in one or more accumulators 2811 or written back to the general-purpose registers in the register file 2810 via the destination write arbiter 2822. In one embodiment, the processing resource architecture 2800 includes a Message Execution Unit (MEU 2815). In response to a request made by a functional unit within the processing resource architecture 2800, the MEU 2815 generates a message and transmits the message to a shared functional unit associated with the processing resource having the architecture 2800. Such shared functional units include, for example, a ray tracing unit, a sampler, and other fixed function logic, such as the fixed function logic 1812A, sampler 1810A, and ray tracing unit 1808A associated with the Figure 18A graphics core 1815A. The MEU 2815 may also generate a message and transmit the message to a load / store unit, such as the Figure 28 load / store circuitry 2821 and / or the memory load / store unit 1804A of the graphics core 1815A. Messages to the load / store unit are used, for example, to perform load and store operations between the register file 2810 and memory via the data cache / shared local memory 1806A. In one embodiment, the processing resource architecture 2800 includes a source crossbar 2830 coupled to the integer unit 2820. The source crossbar is capable of effecting a shuffle and reordering of incoming data according to a desired data pattern. In one embodiment, the source crossbar 2830 is coupled only to the Src0 input. In other embodiments, the source crossbar 2830 may be coupled to other source inputs (e.g., Src1, Src2). Only the Src0 input is required for the transform instructions described herein. However, other transform instructions that utilize multiple source inputs are conceivable. Source data elements may be read from registers within the register file and stored in a reuse buffer associated with the integer unit 2820. Then, the packing pattern of the source data elements may be reconfigured via the source crossbar 2830 before the data elements are provided as input to the integer unit 2820. The manner in which data elements are stored into the registers of the register file is referred to as the regioning mode of the registers. Register regioning involves partitioning the SIMD registers into smaller regions, each region capable of holding a subset of the data elements. In one embodiment, the register regions may be 1D or 2D (e.g., 1×16, 2×8, 4×4, etc.). Instructions executed by the processing resource may specify the registers and sub-registers along with regioning indications that specify the vertical and horizontal spans of the region and the number of data elements per row in the region. The vertical span indicates the number of data elements to skip to reach the next row of a multi-row region. The width specifies the number of data elements per row in the region. The horizontal span indicates which elements in the row are selected for input. For example, a horizontal span of 1 indicates that each element is selected. A horizontal span of 2 indicates that every other element is selected. GPU Thread Dispatch Hardware Figure 29 It is a block diagram of a system 2900 including a GPGPU device 2902 according to an embodiment. The GPGPU device 2902 includes a graphics engine 2908 and a plurality of compute engines 2910. The graphics engine 2908 can process a command buffer or command list of instructions received from a graphics driver. The graphics engine 2908 performs graphics operations in response to these commands. A render command stream converter (RCS 2916) associated with the graphics engine 2908 can stream render commands to perform shader operations. A thread dispatcher 2920 dispatches threads of those shader operations to processing resources within compute blocks 2720A - 2720N. The dispatched threads are executed via hardware threads within the processing resources. Similarly, a set 2910 of compute engines and the graphics engine 2908 operating asynchronously with each other can dispatch commands for compute shaders and / or GPGPU programs via a plurality of compute command stream converters (CCS n 2918). The thread dispatcher 2920 dispatches threads to processing resources within compute blocks 2720A - 2720N to execute compute shaders and / or GPGPU programs. The dispatched threads are executed via hardware threads within the processing resources. In one embodiment, the thread dispatcher 2920 includes a compute walker circuit component 2922 for facilitating the execution of compute programs on graphics processor hardware. In one embodiment, the processing device can be segmented into multiple portions including one or more compute blocks 2720A - 2720N and / or one or more portions of a single compute block. Each portion can be assigned to process commands from a specific source. In the case where a single source (e.g., CCS1 in CCSn 2918) provides commands, the entire processing device (all segments) can be assigned to process commands from a single source. When a second source provides commands (e.g., CCS1 and CCS2), some segments can be assigned to execute commands from the first source while other segments can be assigned to execute commands from the second source. Thus, commands from multiple applications can be executed simultaneously by the processing unit. In the embodiments described herein, the number of active hardware threads within the processing resources of compute blocks 2720A - 2720N is configurable. The number of registers allocated to a given hardware thread is also configurable. In one embodiment, this configuration is performed via a per - thread variable register (VRT) configuration 2906 within a non - pipeline state defined within GPU device 2902. The non - pipeline state 2904 defines state attributes that are globally applied to hardware resources, such as the processing resources within the graphics core applied to compute blocks 2720A - 2720N. Within the non - pipeline state 2904 applied to the processing resources, the VRT configuration 2906 is defined on a per - shader - stage (e.g., shader type) basis, along with enable bits that facilitate backward compatibility. For example, in one embodiment, the VRT configuration 2906 can include separate configurations for vertex shaders, tessellation shaders, geometry shaders, fragment / pixel shaders, mesh shaders, compute shaders, etc. The VRT configuration 2906 is delivered to compute blocks 2720A - 2720N by a thread dispatcher 2920 associated with the dispatch of the threads for execution.
[0381] Figure 30 FIG. is an illustration of a system 3000 for dispatching a thread group to processing resources 3020 according to an embodiment. In one embodiment, the thread group dispatch is performed using a hierarchy of hardware units that handle thread group dispatch. A command - stream transformer 2918 (e.g., Figure 29 CCSn 2918) provides thread - group information to a global compute front - end (CFEG 3010), which is shown as CCS0 - CCS3. The thread - group information includes information about one or more compute kernels to be executed by the graphics processor. The CFEG3010 then dispatches the thread group to the connected CFEs 3015 (shown as CFE0 - CFE7). Each CFE further dispatches the received thread group to the processing resources 3020. In one embodiment, the processing resources 3020 are divided into a plurality of slices 3017 (shown as slice0 - slice15), where each of the CFEs 3015 is coupled to at least two of the plurality of slices 3017. Each of the plurality of slices 3017 is a separately assignable unit of the processing resources 3020. In one embodiment, each slice can be associated with an accelerator - integrated slice 490 as in Figure 4D such that the slice can be presented as a virtualizable sub - device. In one embodiment, each slice can be dedicated to a virtual machine or a virtualized graphics execution environment, such as a container. Each of the plurality of slices 3017 can include one or more sub - slices, where each sub - slice includes one or more sets of individual processing resources and / or execution units. For example, the individual processing resources and / or execution units can have Figure 28 the processing - resource architecture 2800. The processing resources and / or execution units can also have, for example,Figure 2D architecture associated with the graphics multiprocessor 234, the computing unit 1560A, or a variant thereof.
[0382] The specific hardware composition of each slice may vary across embodiments. Refer to Figure 18A , in one embodiment, a slice may include one or more graphics cores, such as one or more of the graphics cores 1815A - 1815N. A slice may also include a graphics core block 1715, which includes all the associated graphics cores 1815A - 1815N. In one embodiment, a slice may include a graphics core cluster 1714. In such an embodiment, a sub - slice may include multiple sets of individual processing resources and / or execution units, each set including multiple vector engines 1802A - 1802N and matrix engines 1803A - 1803N. Depending on the granularity of the slice, an individual graphics core 1815A may be a single slice, a single sub - slice, and / or include multiple sub - slices. Alternatively, refer to Figure 2C , a slice may include a processing cluster 214, where each sub - slice includes one or more graphics multiprocessors 234 or one or more portions of the graphics multiprocessor 234.
[0383] In one thread - group dispatching strategy, the CFEG 3010 dispatches thread groups to the CFE in a round - robin manner, where each CFE is coupled to two consecutive slices. For example, CFE0 dispatches thread groups to slice 0 and slice 1. In this example, within a system of 16 slices 3017, there are 8 CFEs (CFE0 to CFE7). Each cluster of eight (8) slices is coupled to four (4) consecutive CFEs. If a CFE cannot accept a thread group and each dispatching cycle starts from the first CFE (CFE0), the thread - group dispatching strategy may skip that CFE. Thus, the strategy is likely to dispatch consecutive thread groups to different clusters of CFEs, such that the strategy cannot fully utilize sequential memory addresses accessed by consecutive / adjacent thread groups. However, such a dispatching strategy may lead to a slowdown in the workload, where the workload suffers a significantly lower L2 hit rate (relative to the baseline) and / or relatively high number of stall cycles at the interface. This may be due to a low data reuse rate within a particular L2 cache partition and / or potential congestion at the interface due to repetition of memory requests from multiple L2 cache partitions.
[0384] In some embodiments, to improve memory access efficiency, the following dispatch policy may be implemented: the dispatch policy dictates that a consecutive thread group be dispatched to a specific CFE cluster before switching to another cluster. In one embodiment, a thread group may be overdispatched to a cluster such that more threads or thread groups that can be concurrently executed by the cluster may be dispatched to that cluster. The overdispatched threads may be dispatched to a cluster and will execute when the resources of that cluster become available, e.g., when the executing threads or thread groups are preempted during a latency event.
[0385] Figure 31 is an illustration of a system 3100 that facilitates thread dispatch and execution on a graphics processor according to an embodiment. System 3100 includes a graphics driver 3110, a GPGPU thread generator 3120, a thread dispatch network 3130, and a GPU core 3140. In one embodiment, GPU core 3140 includes a plurality of sub-slices, where sub-slice 3142 includes a plurality of processing resources 3146. Graphics driver 3110 includes a user-mode graphics driver 2326 and a kernel-mode graphics driver 2329, as Figure 23 shown. In one embodiment, the operations performed by graphics driver 3110 are mainly performed by user-mode graphics driver 2326. Threads may be generated via GPGPU thread generator 3120 and dispatched via thread dispatch network 3130 to processing resources, including GPU core 3140. In one embodiment, thread dispatch network 3130 includes Figure 29 a thread dispatcher 2920 of Figure 30 and components of system 3000 (e.g., CFEG 3010, CFE 3015). In one embodiment, sub-slice 3142 includes a bindless thread dispatch circuitry (BTD) thread generator 3144. The BTD thread generator 3144 in sub-slice 3142 may facilitate bindless thread dispatch of bindless shaders and may generate sub-threads that execute concurrently with other threads. GPU Thread Mid - Preemption
[0386] Techniques are described herein for enabling mid - thread (instruction - level) preemption in a graphics processor without requiring software intervention by a graphics driver associated with the graphics processor. The save / restore operations for mid - thread preemption can be performed at the granularity of sub - slices 3142 using the SIP routine 3148, which is a GPU - executable kernel program executed by the processing resource 3146. The SIP routine is configured to copy the active thread state from the processing resource 3146 to a memory in hardware configured to store thread states. In one embodiment, the active thread state includes, for example, general - purpose registers, architecture - specific registers, shared local memory, and barriers for the threads active on the sub - slice at the time a mid - thread preemption command is received. The active thread state is stored at sub - slice granularity, where each sub - slice has a dedicated area in memory for storing its active thread state during preemption. In one embodiment, the SIP routine performs operations via the compute engine and copy engine of each preempted sub - slice to save the thread state during preemption and then restore the active thread state during resume. In various embodiments, the SIP routine for mid - thread preemption can be provided, for example, by the firmware or hardware of the GPU and associated with the dispatched thread group, automatically generated for the thread group during shader compilation, or hand - coded by a programmer to be executed with user - generated GPGPU code.
[0387] In one embodiment, the memory in hardware includes a Thread State Buffer (TSB 3152) and an Over - dispatch Buffer (ODB 3154). The TSB 3152 stores data for running threads. The ODB 3154 stores data for threads that have been dispatched but are not running at the time a preemption command is received. Each sub - slice includes a TSB 3152 and an ODB 3154. In one embodiment, the graphics driver 3110 allocates memory for each TSB 3152 and ODB 3154. The graphics driver 3110 can create the buffers globally or for a specific workload during hardware initialization or when mid - thread preemption is enabled. If mid - thread preemption is disabled, the buffers may be de - allocated. In one embodiment, the TSB 3152 and ODB 3154 can be created by hardware within the graphics processor, such as, for example, the graphics micro - controller 1533 in FIG. 15 ...
Claims
1. A graphics processor, comprising: A memory interface; A plurality of graphics cores coupled via a data interconnect, each of the plurality of graphics cores including a plurality of processing resources; And Circuitry for managing execution of a workload by the plurality of graphics cores, the circuitry being configured to: Dispatch a first plurality of threads on behalf of a first thread group to be executed by a first set of the plurality of processing resources; Receive a command for mid-thread preemption of execution of the first thread group; Trigger execution of a kernel program on the first set of processing resources to save an execution state of the first set of processing resources to a memory of the graphics processor; And After saving the execution state of the first set of processing resources, replace the first thread group with a second thread group on the first set of processing resources.
2. The graphics processor according to claim 1, wherein, The plurality of graphics cores are arranged in a plurality of slices, each slice having a plurality of sub-slices including a plurality of processing resources, each sub-slice being independently preemptable, and the first set of processing resources is associated with a first sub-slice.
3. The graphics processor according to claim 2, wherein, The circuitry is configured to replace the first thread group on the first sub-slice while executing a third thread group via a second sub-slice.
4. The graphics processor according to claim 3, wherein, The circuitry is to trigger execution of the kernel program to save the execution state of the first set of processing resources via an interrupt routine triggered for execution on each processing resource of the first sub-slice.
5. The graphics processor according to claim 4, wherein, The interrupt routine is to cause execution of the kernel program, which saves the execution state of the first set of processing resources to a thread state buffer within the memory of the graphics processor.
6. The graphics processor according to claim 5, wherein, Each sub-slice is to be associated with a separate thread state buffer.
7. The graphics processor according to claim 5, wherein, The execution state of the first set of processing resources includes an active thread state of the first sub-slice at the time of receiving the command for mid-thread preemption, the active thread state including general-purpose registers, architecture-specific registers, shared local memory, and barrier state.
8. The graphics processor according to claim 1, wherein the circuitry is configured to: Receive a command for triggering a resume event of the first thread group; In response to the command for triggering the resume event, dispatch a third plurality of threads to a second set of processing resources of the graphics processor, the third plurality of threads restoring the execution state saved for the first set of processing resources to the second set of processing resources; and Restore execution of the first thread group at the second set of processing resources via the third plurality of threads.
9. The graphics processor according to claim 8, wherein, The second set of processing resources is different from the first set of processing resources.
10. The graphics processor according to claim 9, wherein, The first set of processing resources is configured via kernel program code to construct a resume batch buffer that causes restoration of the execution state, the resume batch buffer being executed in response to the command for triggering the resume event.
11. A method for performing mid-thread preemption in a graphics processor configured to execute a compute program, the method comprising: A first thread group representative of a first plurality of processing resources to be executed by the first plurality of processing resources dispatches the first plurality of threads to the first plurality of processing resources of the graphics processor; Receives a command for mid-thread preemption of the first thread group; Executes a kernel program on the first plurality of processing resources to save an execution state of the first plurality of processing resources to a memory of the graphics processor; And After saving the execution state of the first plurality of processing resources, replaces the first thread group on the first plurality of processing resources with a second thread group.
12. The method according to claim 11, wherein, The graphics processor includes a plurality of graphics cores, the plurality of graphics cores being arranged in a plurality of slices, each slice having a plurality of sub-slices including a plurality of processing resources, each sub-slice being independently preemptable, and the first plurality of processing resources being associated with a first sub-slice.
13. The method according to claim 12, further comprising replacing the first thread group on the first sub-slice while executing a third thread group via a second sub-slice.
14. The method according to claim 11, further comprising: Receiving a command to resume the first thread group; Dispatching a third plurality of threads to a second plurality of processing resources of the graphics processor; And Restoring, via the third plurality of threads, the execution state saved for the first plurality of processing resources to the second plurality of processing resources.
15. The method according to claim 14, wherein, Executing the kernel program on the first plurality of processing resources to save the execution state of the first plurality of processing resources includes interrupting program code associated with the first thread group and executing the kernel program at each of the first plurality of processing resources.
16. The method according to claim 15, wherein, Also establishing a restore command buffer via the kernel program associated with saving the execution state of the first plurality of processing resources, wherein executing the restore command buffer causes the execution state of the first thread group to be restored to the second plurality of processing resources.
17. The method according to claim 16, further comprising resuming execution of the first thread group at the second plurality of processing resources, wherein, The second plurality of processing resources is different from the first plurality of processing resources and has a different number of processing resources from the first plurality of processing resources.
18. A graphics processing system, including components for executing the method according to any one of claims 11-17.
19. An apparatus, including: A memory interface; A plurality of graphics cores coupled via a data interconnect, each of the plurality of graphics cores including a plurality of processing resources; And Circuitry for managing the plurality of graphics cores to execute a workload, the circuitry being configured to: In addition to dispatching a first plurality of threads to a first set of processing resources of the plurality of processing resources, over-dispatch a second plurality of threads to the first set of processing resources on behalf of a first thread group; Begin executing the first plurality of threads via the first set of processing resources; Receive a command for mid-thread preemption of the first thread group; Save an execution state of the first set of processing resources to a first location in a memory of the graphics processor; Save a thread state of unstarted threads of the second plurality of threads to a second location in the memory of the graphics processor; And After saving the execution state of the first set of processing resources and the thread states of the unstarted threads among the second plurality of threads, replace the first thread group with a second thread group on the first set of processing resources.
20. The apparatus according to claim 19, wherein the circuit component is configured to: Receive a command triggering a recovery event that causes recovery of a previously preempted thread group; Restore the execution state of the first set of processing resources from a first location in the memory of the graphics processor to a second set of processing resources; Replay the unstarted threads of the second plurality of threads from a second location in the memory of the graphics processor to oversubscription of the second set of processing resources; And Restore the execution of the first thread group at the second set of processing resources.
21. The apparatus according to claim 20, wherein, The plurality of graphics cores are arranged in a plurality of slices, each slice having a plurality of sub-slices including a plurality of processing resources, each sub-slice being independently preemptable, and the first set of processing resources is associated with a first sub-slice.
22. The apparatus according to claim 21, wherein, The circuit component is configured to replace the first thread group on the first sub-slice during execution of a third thread group via a second sub-slice.
23. The device according to claim 22, wherein, To save the execution state of the first set of processing resources to a first location in the memory of the graphics processor, the circuit component is configured to identify active threads on the first set of processing resources and save the thread states of the active threads to a thread state buffer at the first location in the memory.
24. The apparatus according to claim 23, wherein, The circuit component is configured to index the thread states of the active threads in the thread state buffer according to hardware thread identifiers and processing resource identifiers respectively associated with each active thread.
25. The apparatus according to claim 24, wherein, The thread state buffer is associated with the first sub-slice and each sub-slice has an associated thread state buffer, and to save the thread states of the unstarted threads among the second plurality of threads, the circuit component is to save the thread state data of the unstarted threads together with the thread identifiers of each unstarted thread to an oversubscription buffer at a second location in the memory.
Citation Information
Cited By
Operator execution method, electronic equipment, storage medium and program product
CN121255478A
Operator execution method, electronic device, storage medium, and program product
CN121255478B