Unbound thread dispatch intermediate thread preemption on graphics processor

By introducing a hardware-supported intermediate thread preemption mechanism into the graphics processor, the performance bottleneck caused by software-assisted context switching is solved, and the execution efficiency and performance of the graphics processor are improved.

CN120339038APending Publication Date: 2025-07-18INTEL CORP
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202411949067.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2024-01-17
Filing Date
2024-12-27
Publication Date
2025-07-18

AI Technical Summary

Technical Problem

In the prior art, intermediate thread preemption of the graphics processor depends on software-assisted context switching, resulting in large wait time overhead and reducing the performance and efficiency of executing instruction-level preemption.

Method used

By introducing an intermediate thread preemption mechanism supported by hardware in the graphics processor, the graphics processor's thread dispatch hardware is used to save and restore the state of the preemption thread and avoid software intervention.

Benefits of technology

It realizes intermediate thread preemption without software intervention, improves the execution efficiency and performance of the graphics processor, and reduces the waiting time overhead.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120339038A_ABST
    Figure CN120339038A_ABST
Patent Text Reader

Abstract

The invention relates to unbound thread dispatch intermediate thread preemption on a graphics processor. Techniques are provided for enabling intermediate thread (instruction level) preemption in a graphics processor without software intervention by a graphics driver associated with the graphics processor. Hardware-based intermediate thread preemption is facilitated by using a thread of a graphics processor to dispatch hardware to trigger execution of a kernel program by the graphics processor that holds a thread state of preempted processing resources and facilitates subsequent recovery of the thread state to the same or a different set of processing resources.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure generally relates to data processing via a graphics processing unit, and more particularly, to methods for enabling mid-thread preemption within a general-purpose graphics processing unit. Background Art

[0002] Hardware utilization can be enhanced by increasing the granularity at which preemption can occur. Software-assisted context switching has been used in graphics processing units to perform instruction-level preemption. However, software-assisted context switching has a latency overhead, which degrades the performance and efficiency of performing instruction-level preemption. Brief Description of the Drawings

[0003] Embodiments described herein are illustrated by way of example and not limitation in the accompanying drawings, in which like reference numerals indicate like elements and in which:

[0004] Figure 1 is a block diagram of a computer system configured to implement one or more aspects of the embodiments described herein;

[0005] Figures 2A - 2D illustrates a parallel processor component;

[0006] Figures 3A - 3C is a block diagram of a graphics multiprocessor and a multi-processor based GPU;

[0007] Figures 4A - 4F illustrates an exemplary architecture in which multiple GPUs are communicatively coupled to multiple multi-core processors;

[0008] Figure 5 illustrates a graphics processing pipeline;

[0009] Figure 6 illustrates a machine learning software stack;

[0010] Figure 7 illustrates a general-purpose graphics processing unit;

[0011] Figure 8 illustrates a multi-GPU computing system;

[0012] Figures 9A - 9B illustrates the layers of an exemplary deep neural network;

[0013] Figure 10 illustrates an exemplary recurrent neural network;

[0014] Figure 11 illustrates the training and deployment of a deep neural network;

[0015] Figure 12A is a block diagram illustrating distributed learning;

[0016] Figure 12B is a block diagram illustrating a programmable network interface and a data processing unit;

[0017] Figure 13 illustrates an exemplary inference system on a chip (SOC) adapted to perform inference using a trained model;

[0018] Figure 14 is a block diagram of a processing system;

[0019] Figures 15A - 15C illustrates a computing system and a graphics processor;

[0020] Figures 16A - 16C illustrates a block diagram of an additional graphics processor and a computing accelerator architecture;

[0021] Figure 17 is a block diagram of a graphics processing engine of a graphics processor;

[0022] Figures 18A - 18C illustrates thread execution logic including an array of processing elements employed in a graphics processor core;

[0023] Figure 19 illustrates a die of a multi-die processor according to an embodiment;

[0024] Figure 20 is a block diagram illustrating a graphics processor instruction format;

[0025] Figure 21 is a block diagram of an additional graphics processor architecture;

[0026] Figures 22A - 22B illustrates a graphics processor command format and a command sequence;

[0027] Figure 23 illustrates an exemplary graphics software architecture for a data processing system;

[0028] Figure 24A is a block diagram illustrating an IP core development system;

[0029] Figure 24B illustrates a cross-sectional side view of an integrated circuit package component;

[0030] Figure 24C illustrates a package component including a hardware logic die connecting multiple units to a substrate (e.g., a base die);

[0031] Figure 24D illustrates a package component including interchangeable dies;

[0032] Figure 25is a block diagram illustrating an exemplary system-on-chip integrated circuit;

[0033] Figures 26A - 26B is a block diagram illustrating an exemplary graphics processor for use within a SoC;

[0034] Figure 27 is a block diagram of a data processing system according to an embodiment;

[0035] Figure 28 illustrates a processing resource architecture according to an embodiment;

[0036] Figure 29 is a block diagram of a system including a GPGPU device according to an embodiment;

[0037] Figure 30 is an illustration of a system for dispatching a thread group to processing resources according to an embodiment;

[0038] Figure 31 is an illustration of a system for facilitating thread dispatching and execution on a graphics processor according to an embodiment;

[0039] Figure 32 shows a TSB for storing state information of multiple sub-slices according to an embodiment;

[0040] Figure 33 shows an ODB for storing the state of an over-dispatched thread-group (ODTG) according to an embodiment;

[0041] Figure 34 shows high-level operations of walker-based intermediate thread preemption according to an embodiment;

[0042] Figure 35 illustrates thread dispatching and execution hardware associated with a sub-slice according to an embodiment;

[0043] Figure 36 illustrates concurrent execution of BTD and non-BTD threads in a sub-slice;

[0044] Figure 37 is a block diagram of a computing device including a graphics processor according to an embodiment;

[0045] Figure 38 illustrates an MTP save operation for pending over-dispatched BTD sub-threads according to an embodiment;

[0046] Figure 39 illustrates MTP save of over-dispatched threads from CFEG according to an embodiment;

[0047] Figure 40Illustration of MTP recovery operations for pre-empted active threads;

[0048] Figure 41 Illustration of MTP recovery for pending BTD child threads;

[0049] Figure 42 Illustration of MTP recovery for oversubscribed CFEG threads according to an embodiment;

[0050] Figure 43 Illustration of a system for pre-empting and preserving thread scratchpad data across intermediate threads according to an embodiment;

[0051] Figure 44 Illustration of a system for reusing scratchpad allocations after thread migration during MTP recovery according to an embodiment;

[0052] Figure 45 Illustration of a graphics processing system with support for multiple submission queues per context according to an embodiment;

[0053] Figure 46 Illustration of an exemplary computing hierarchy according to an embodiment;

[0054] Figure 47 Illustration of a system including data structures adapted to implement multi-queue and multi-context intermediate thread pre-emption according to an embodiment.

[0055] Figure 48 Illustration of a system where multi-context pre-emption is performed on sub-slices using separate buffers;

[0056] Figure 49 Illustration of a method for performing intermediate thread pre-emption on application- and device-generated threads according to an embodiment;

[0057] Figure 50 Illustration of a method for enabling thread scratchpad reuse for pre-empted threads according to an embodiment;

[0058] Figure 51 Illustration of a method for enabling intermediate thread pre-emption on a multi-context and / or multi-queue processing resource cluster according to an embodiment; and

[0059] Figure 52 Is a block diagram of a computing device including a graphics processor according to an embodiment. Detailed Description

[0060] Current parallel graphics data processing includes systems and methods developed to perform specific operations on graphics data, such as, for example, linear interpolation, tessellation, rasterization, texture mapping, depth testing, and the like. Traditionally, graphics processors have used fixed-function computing units to process graphics data. However, more recently, multiple parts of the graphics processor have been made programmable, enabling such processors to support a wider variety of operations to process vertex data and fragment data. To further improve performance, graphics processors typically implement processing techniques such as pipelining, which attempt to process as much graphics data as possible in parallel across different parts of the graphics pipeline. A parallel graphics processor with a single instruction, multiple thread (SIMT) architecture is designed to maximize the amount of parallel processing in the graphics pipeline. In the SIMT architecture, groups of parallel threads attempt to execute program instructions synchronously together as frequently as possible to improve processing efficiency.

[0061] A graphics processing unit (GPU) is communicatively coupled to a host / processor core to accelerate, for example, graphics operations, machine learning operations, pattern analysis operations, and / or various general-purpose GPU (GPGPU) functions. The GPU can be communicatively coupled to the host processor / core via a bus or another interconnect (e.g., a high-speed interconnect such as PCIe or NVLink). Alternatively, the GPU can be integrated on the same package or chip as the core and communicatively coupled to the core via an internal processor bus / interconnect (i.e., within the package or chip). Regardless of the manner in which the GPU is connected, the processor core can allocate work to the GPU in the form of a sequence of commands / instructions contained in a work descriptor. The GPU then uses dedicated circuitry / logic to efficiently process these commands / instructions.

[0062] In the following description, numerous specific details are set forth in order to provide a more thorough understanding. However, it will be apparent to one of ordinary skill in the art that embodiments described herein may be practiced without one or more of these specific details. In other instances, well-known features have not been described so as not to obscure the details of the current embodiments.

[0063] GPU preemption enables a workload being executed on a GPU to be replaced by another workload. Preemption can occur at various granularities, including command buffer-level preemption, command-level preemption, thread group-level preemption, thread-level preemption, and mid-thread preemption at the instruction level. To perform mid-thread preemption, the execution state of a thread is saved in a manner that allows the hardware to correctly resume at the precise instruction where the thread was preempted. Mid-thread preemption (MTP) in a GPU has previously been performed with software assistance. However, relying on software assistance introduces latency into the preemption process. Using the techniques described herein, the GPU hardware can be configured to save thread state without the intervention of the graphics driver software to perform software-assisted context switching. System Overview

[0064] Figure 1 is a block diagram of a computing system 100 configured to implement one or more aspects of the embodiments described herein. The computing system 100 includes a processing subsystem 101 that has one or more processors 102 and a system memory 104 communicating via an interconnect path that may include a memory hub 105. The memory hub 105 can be a separate component within a chipset component or can be integrated within one or more of the processors 102. The memory hub 105 is coupled to an I / O subsystem 111 via a communication link 106. The I / O subsystem 111 includes an I / O hub 107 that enables the computing system 100 to receive input from one or more input devices 108. Additionally, the I / O hub 107 enables a display controller (which may be included within one or more of the processors 102) to provide output to one or more display devices 110A. In one embodiment, one or more of the display devices 110A coupled to the I / O hub 107 can include a local, internal, or embedded display device.

[0065] The processing subsystem 101 includes, for example, one or more parallel processors 112 coupled to the memory hub 105 via a bus or other communication link 113. The communication link 113 can be one of any number of standard-based communication link technologies or protocols, such as, but not limited to, PCI Express (PCIe), or can be a vendor-specific communication interface or fabric. The one or more parallel processors 112 can form a parallel or vector processing system in a computing complex that can include a large number of processing cores and / or processing clusters, such as, for example, a many integrated core (MIC) processor. For example, the one or more parallel processors 112 form a graphics processing subsystem that can output pixels to a display device among one or more display devices 110A coupled via the I / O hub 107. The one or more parallel processors 112 can also include a display controller and a display interface (not shown) for enabling a direct connection to one or more display devices 110B.

[0066] Within the I / O subsystem 111, the system storage unit 114 can be connected to the I / O hub 107 to provide a storage mechanism for the computing system 100. The I / O switch 116 can be used to provide an interface mechanism to enable connections between the I / O hub 107 and other components, such as a network adapter 118 and / or a wireless network adapter 119 that can be integrated into the platform, and various other devices that can be added via one or more plug-in devices 120. The (one or more) plug-in devices 120 can also include, for example, one or more external graphics processor devices, graphics cards, and / or computing accelerators. The network adapter 118 can be an Ethernet adapter or another wired network adapter. The wireless network adapter 119 can include one or more of the following: Wi-Fi, Bluetooth, near field communication (NFC), or other network devices that include one or more wireless radio devices.

[0067] The computing system 100 can include other components not explicitly shown, including USB or other port connections, optical storage drives, video capture devices, etc., which can also be connected to the I / O hub 107. The communication paths interconnecting the various components in Figure 1 can be implemented using any suitable protocol, such as a PCI (Peripheral Component Interconnect)-based protocol (e.g., PCI Express), or any other bus or point-to-point communication interface and / or (one or more) protocols, such as NVLink high-speed interconnect, Compute Express Link TM (ComputeExpress LinkTM , CXL TM (e.g., CXL.mem), Infinity Fabric (IF), Ethernet (IEEE802.3), remote direct memory access (RDMA), InfiniBand, Internet Wide Area RDMA Protocol (iWARP), Transmission Control Protocol (TCP), User Datagram Protocol (UDP), quick UDP Internet Connections (QUIC), RDMA over Converged Ethernet (RoCE), Intel QuickPath Interconnect (QPI), Intel Ultra Path Interconnect (UPI), Intel On-Chip System Fabric (IOSF), Omnipath, HyperTransport, Advanced Microcontroller Bus Architecture (AMBA) interconnect, OpenCAPI, Gen-Z, Cache Coherent Interconnect for Accelerators (CCIX), 3GPP Long Term Evolution (LTE) (4G), 3GPP 5G and its variants, or a wired or wireless interconnect protocol known in the art. In some examples, protocols such as non-volatile memory express over Fabrics (NVMe-oF) or NVMe may be used to copy or store data to the virtualized storage node.

[0068] One or more parallel processors 112 may include circuitry optimized for graphics and video processing (including, for example, video output circuitry) and constitute a graphics processing unit (GPU). Alternatively or additionally, as described in more detail herein, one or more parallel processors 112 may include circuitry optimized for general-purpose processing while retaining the underlying computational architecture. The components of computing system 100 may be integrated with one or more other system elements on a single integrated circuit. For example, one or more parallel processors 112, memory hub 105, processor(s) 102, and I / O hub 107 may be integrated into a system-on-chip (SoC) integrated circuit. Alternatively, the components of computing system 100 may be integrated into a single package to form a system-in-package configuration. In one embodiment, at least some of the components of computing system 100 may be integrated into a multi-chip module (MCM), which may be interconnected with other multi-chip modules into a modular computing system.

[0069] It will be appreciated that the computing system 100 shown herein is illustrative, and variations and modifications are possible. The connection topology may be modified as needed, including the number and arrangement of bridges, the number of processor(s) 102, and the number of parallel processor(s) 112. For example, system memory 104 may be connected directly to processor(s) 102 rather than through a bridge, and other devices communicate with system memory 104 via memory hub 105 and processor(s) 102. In other alternative topologies, the parallel processor(s) are connected to I / O hub 107 or directly to one of the processor(s) 102 rather than to memory hub 105. In other embodiments, I / O hub 107 and memory hub 105 may be integrated into a single chip. It is also possible for two or more sets of processors 102 to be attached via multiple sockets, which may be coupled to two or more instances of the parallel processor(s) 112.

[0070] Some of the specific components shown herein are optional and may not be included in all implementations of computing system 100. For example, any number of plug-in cards or peripheral devices may be supported, or some components may be eliminated. Additionally, some architectures may use different terms for components similar to those illustrated Figure 1 herein. For example, memory hub 105 may be referred to as a north bridge in some architectures, while I / O hub 107 may be referred to as a south bridge.

[0071] Figure 2AFIG. Parallel processor 200 is illustrated. The parallel processor 200 can be a GPU, GPGPU, etc. as described herein. Various components of the parallel processor 200 can be implemented using one or more integrated circuit devices, such as programmable processors, application specific integrated circuits (ASICs), or field programmable gate arrays (FPGAs). The illustrated parallel processor 200 can be one or more of the parallel processors 112 shown in Figure 1 One or more of those shown in.

[0072] The parallel processor 200 includes a parallel processing unit 202. The parallel processing unit includes an I / O unit 204 that enables communication with other devices, including other instances of the parallel processing unit 202. The I / O unit 204 can be directly connected to other devices. For example, the I / O unit 204 is connected to other devices via a hub or switch interface, such as a memory hub 105. The connection between the memory hub 105 and the I / O unit 204 forms a communication link 113. Within the parallel processing unit 202, the I / O unit 204 is connected to a host interface 206 and a memory crossbar 216, where the host interface 206 receives commands related to performing processing operations, and the memory crossbar 216 receives commands related to performing memory operations.

[0073] When the host interface 206 receives a command buffer via the I / O unit 204, the host interface 206 can direct the work operations for executing those commands to a front end 208. In one embodiment, the front end 208 is coupled to a scheduler 210 that is configured to distribute commands or other work items to an array of processing clusters 212. The scheduler 210 ensures that the array of processing clusters 212 is properly configured and in an active state before tasks are distributed to the processing clusters within the array of processing clusters 212. The scheduler 210 can be implemented via firmware logic executed on a microcontroller. The scheduler 210 implemented by the microcontroller can be configured to perform complex scheduling and work distribution operations at both a coarse-grained and fine-grained level, enabling fast preemption and context switching of threads executing on the array of processing clusters 212. Preferably, the host software can authenticate the workload scheduled on the array of processing clusters 212 via one of a plurality of graphics processing doorbells. In other examples, polling for new workloads or interrupts can be used to identify or indicate the availability of work to be performed. The workload can then be automatically distributed across the array of processing clusters 212 by the scheduler 210 logic within the scheduler microcontroller.

[0074] The processing cluster array 212 may include up to "N" processing clusters (e.g., cluster 214A, cluster 214B to cluster 214N). Each of the clusters 214A - 214N in the processing cluster array 212 can execute a large number of concurrent threads. The scheduler 210 can use various scheduling and / or work distribution algorithms to allocate work to the clusters 214A - 214N in the processing cluster array 212, and these scheduling and / or work distribution algorithms can vary depending on the workload generated for each type of program or computation. Scheduling can be handled dynamically by the scheduler 210, or can be assisted in part by compiler logic during the compilation of the program logic configured to be executed by the processing cluster array 212. Optionally, different ones of the clusters 214A - 214N in the processing cluster array 212 can be assigned to process different types of programs or to perform different types of computations.

[0075] The processing cluster array 212 can be configured to perform various types of parallel processing operations. For example, the processing cluster array 212 is configured to perform general - purpose parallel computing operations. For example, the processing cluster array 212 may include logic for performing processing tasks, which include filtering of video and / or audio data, performing modeling operations including physical operations, and performing data transformations.

[0076] The processing cluster array 212 is configured to perform parallel graphics processing operations. In such embodiments where the parallel processor 200 is configured to perform graphics processing operations, the processing cluster array 212 may include additional logic to support the execution of such graphics processing operations, including but not limited to texture sampling logic for performing texture operations, and tessellation logic and other vertex processing logic. In addition, the processing cluster array 212 can be configured to execute graphics - processing - related shader programs, such as but not limited to vertex shaders, tessellation shaders, geometry shaders, and pixel shaders. The parallel processing unit 202 can transfer data from the system memory via the I / O unit 204 for processing. During processing, the transferred data can be stored in on - chip memory (e.g., parallel processor memory 222) during processing and then written back to the system memory.

[0077] In embodiments in which a parallel processing unit 202 is used to perform graphics processing, a scheduler 210 may be configured to divide a processing workload into tasks of approximately equal size to better enable distribution of graphics processing operations to a plurality of clusters 214A - 214N in a processing cluster array 212. In some of these embodiments, portions of the processing cluster array 212 may be configured to perform different types of processing. For example, a first portion may be configured to perform vertex shading and topology generation, a second portion may be configured to perform tessellation and geometry shading, and a third portion may be configured to perform pixel shading or other screen space operations to produce a rendered image for display. Intermediate data produced by one or more of the clusters 214A - 214N may be stored in a buffer to allow the intermediate data to be transferred between the clusters 214A - 214N for further processing.

[0078] During operation, the processing cluster array 212 may receive processing tasks to be executed via the scheduler 210, which receives commands defining the processing tasks from a front end 208. For graphics processing operations, the processing tasks may include data to be processed and indices of status parameters and commands defining how the data is to be processed (e.g., what program is to be executed), such as surface (patch) data, primitive data, vertex data, and / or pixel data. The scheduler 210 may be configured to fetch an index corresponding to the task or may receive the index from the front end 208. The front end 208 may be configured to ensure that the processing cluster array 212 is configured in an effective state before a workload specified by an incoming command buffer (e.g., a batch buffer, a push buffer, etc.) is initiated.

[0079] Each instance of one or more instances of the parallel processing unit 202 may be coupled to the parallel processor memory 222. The parallel processor memory 222 may be accessed via the memory crossbar 216, which may receive memory requests from the processing cluster array 212 as well as the I / O unit 204. The memory crossbar 216 may access the parallel processor memory 222 via the memory interface 218. The memory interface 218 may include a plurality of partition units (e.g., partition unit 220A, partition unit 220B, up to partition unit 220N) each of which may be coupled to a portion (e.g., a memory cell) of the parallel processor memory 222. The number of partition units 220A - 220N may be configured to be equal to the number of memory cells such that the first partition unit 220A has a corresponding first memory cell 224A, the second partition unit 220B has a corresponding second memory cell 224B, and the Nth partition unit 220N has a corresponding Nth memory cell 224N. In other embodiments, the number of partition units 220A - 220N may not be equal to the number of memory devices.

[0080] The memory cells 224A - 224N may include various types of memory devices, including dynamic random - access memory (DRAM) or graphics random - access memory, such as synchronous graphics random - access memory (SGRAM), including graphics double - data rate (GDDR) memory. Optionally, the memory cells 224A - 224N may also include 3D stacked memory, including but not limited to high - bandwidth memory (HBM). Those skilled in the art will appreciate that the specific implementation of the memory cells 224A - 224N may vary and may be selected from one of a variety of conventional designs. Rendering targets such as frame buffers or texture maps may be stored across the memory cells 224A - 224N, allowing the partition units 220A - 220N to write portions of each rendering target in parallel to efficiently utilize the available bandwidth of the parallel processor memory 222. In some embodiments, local instances of the parallel processor memory 222 may be excluded in favor of a unified memory design that utilizes system memory in conjunction with local cache memory.

[0081] Optionally, any one of the clusters 214A - 214N in the processing cluster array 212 has the ability to process data to be written to any one of the memory units 224A - 224N within the parallel processor memory 222. The memory crossbar 216 can be configured to transfer the output of each cluster 214A - 214N to any of the partition units 220A - 220N or to another cluster 214A - 214N, which can perform additional processing operations on the output. Each cluster 214A - 214N can communicate with the memory interface 218 through the memory crossbar 216 to read from or write to various external memory devices. In one embodiment among the embodiments having a memory crossbar 216, the memory crossbar 216 has a connection to the memory interface 218 to communicate with the I / O unit 204 and has a connection to a local instance of the parallel processor memory 222, enabling the processing units within different processing clusters 214A - 214N to communicate with the system memory or other memories not local to the parallel processing unit 202. Generally, the memory crossbar 216 can, for example, be capable of using virtual channels to separate the traffic flow between the clusters 214A - 214N and the partition units 220A - 220N.

[0082] Although a single instance of the parallel processing unit 202 is illustrated within the parallel processor 200, any number of instances of the parallel processing unit 202 can be included. For example, multiple instances of the parallel processing unit 202 can be provided on a single plug-in card, or multiple plug-in cards can be interconnected. For example, the parallel processor 200 can be a plug-in device, such as, Figure 1 the plug-in device 120, which can be a graphics card (such as a discrete graphics card including one or more GPUs, one or more memory devices, and device - to - device or network or fabric interfaces). Different instances of the parallel processing unit 202 can be configured to interoperate even if the different instances have different numbers of processing cores, different amounts of local parallel processor memory, and / or other configuration differences. Optionally, some instances of the parallel processing unit 202 can include higher - precision floating - point units relative to other instances. A system including one or more instances of the parallel processing unit 202 or the parallel processor 200 can be implemented in various configurations and form factors, including but not limited to, desktop computers, laptop computers, or handheld personal computers, servers, workstations, game consoles, and / or embedded systems. The orchestrator can use one or more of the following to form a composite node for workload execution: decomposed processor resources, cache resources, memory resources, storage resources, and networking resources.

[0083] In one embodiment, the parallel processing unit 202 may be partitioned into multiple instances. Those multiple instances may be configured to execute workloads associated with different clients in an isolated manner, thereby providing a predetermined quality of service for each client. For example, each cluster 214A - 214N may be partitioned and isolated from other clusters, allowing the processing cluster array 212 to be divided into multiple computing partitions or instances. In such a configuration, workloads executed on isolated partitions are protected from errors or inaccuracies associated with different workloads executed on different partitions. The partitioning units 220A - 220N may be configured to enable dedicated and / or isolated paths to the memories of the clusters 214A - 214N associated with the respective computing partitions. This data path isolation enables the computing resources within a partition to communicate with one or more assigned memory units 224A - 224N without being disturbed by the activities of other partitions.

[0084] Figure 2B is a block diagram of the partitioning unit 220. The partitioning unit 220 may be Figure 2A an instance of one of the partitioning units 220A - 220N. As illustrated, the partitioning unit 220 includes an L2 cache 221, a frame buffer interface 225, and a ROP 226 (raster operation unit). The L2 cache 221 is a read / write cache configured to perform load and store operations received from the memory crossbar 216 and the ROP 226. Read misses and urgent write-back requests are output by the L2 cache 221 to the frame buffer interface 225 for processing. Updates may also be sent via the frame buffer interface 225 to the frame buffer for processing. In one embodiment, the frame buffer interface 225 provides an interface to the memory unit 224 within the parallel processor memory 222 of Figure 2A the memory units 224A - 224N. The partitioning unit 220 may additionally or alternatively provide an interface to one of the memory units in the parallel processor memory via a memory controller (not shown).

[0085] In a graphics application, the ROP 226 is a processing unit that performs raster operations such as stencil, z-test, blending, etc. The ROP 226 then outputs the processed graphics data, which is stored in the graphics memory. In some embodiments, the ROP 226 includes a codec (CODEC) 227 or is coupled to the CODEC 227, and the CODEC 227 includes compression logic for compressing depth or color data written to the memory or L2 cache 221 and decompressing depth or color data read from the memory or L2 cache 221. The compression logic can be lossless compression logic that utilizes one or more of a variety of compression algorithms. The type of compression performed by the CODEC 227 can vary based on the statistical characteristics of the data to be compressed. For example, in one embodiment, delta color compression is performed on the depth and color data on a per-tile basis. In one embodiment, the CODEC 227 includes compression and decompression logic that can compress and decompress computational data associated with machine learning operations. The CODEC 227 can, for example, compress sparse matrix data for sparse machine learning operations. The CODEC 227 can also compress sparse matrix data encoded in a sparse matrix format (e.g., coordinate list encoding (COO), compressed sparse row (CSR), compressed sparse column (CSC), etc.) to generate compressed and encoded sparse matrix data. The compressed and encoded sparse matrix data can be decompressed and / or decoded before being processed by a processing element, or the processing element can be configured to consume the compressed, encoded, or compressed and encoded data for processing.

[0086] The ROP 226 can be included within each processing cluster (e.g., Figure 2A clusters 214A - 214N) rather than being included within the partitioning unit 220. In such embodiments, read and write requests for pixel data rather than pixel fragment data are transmitted through the memory crossbar 216. The processed graphics data can be displayed on a display device (such as Figure 1 one of the one or more display devices 110A - 110B), routed for further processing by the (one or more) processors 102, or routed for further processing by Figure 2A one of the processing entities within the parallel processor 200.

[0087] Figure 2Cis a block diagram of processing cluster 214 within a parallel processing unit. For example, the processing cluster is Figure 2A an instance of one of processing clusters 214A - 214N of Figure 2A . Processing cluster 214 can be configured to execute many threads in parallel, where the term "thread" refers to an instance of a particular program executed on a particular set of input data. Optionally, single-instruction, multiple-data (SIMD) instruction issue techniques can be used to support parallel execution of a large number of threads without providing multiple independent instruction units. Alternatively, single-instruction, multiple-thread (SIMT) techniques can be used to support parallel execution of a large number of generally synchronized threads using a common instruction unit configured to issue instructions to a set of processing engines within each processing cluster in the processing cluster. Different from SIMD execution mechanisms where all processing engines typically execute the same instruction, SIMT execution allows different threads to more easily follow divergent execution paths through a given thread program. Those skilled in the art will understand that SIMD processing mechanisms represent a functional subset of SIMT processing mechanisms.

[0088] The operation of processing cluster 214 can be controlled by distributing processing tasks to pipeline manager 232 of the SIMT parallel processor. Pipeline manager 232 receives instructions from Figure 2A scheduler 210 and manages the execution of those instructions via graphics multiprocessor 234 and / or texture unit 236. The illustrated graphics multiprocessor 234 is an exemplary instance of a SIMT parallel processor. However, various types of SIMT parallel processors of different architectures can be included within processing cluster 214. One or more instances of graphics multiprocessor 234 can be included within processing cluster 214. Graphics multiprocessor 234 can process data, and data crossbar 240 can be used to distribute the processed data to one of multiple possible destinations, including facilitating data exchange between graphics multiprocessors within processing cluster 214. Pipeline manager 232 can facilitate the distribution of the processed data by specifying the destination for the processed data to be distributed via data crossbar 240.

[0089] Each graphics multiprocessor 234 within processing cluster 214 may include the same set of functional execution logic (e.g., arithmetic logic unit, load-store unit, etc.). The functional execution logic can be configured in a pipelined manner, in which new instructions can be issued before previous instructions are completed. The functional execution logic supports a variety of operations, including integer and floating-point arithmetic, comparison operations, boolean operations, bit shifts, and the calculation of various algebraic functions. Different operations can be executed using the same functional unit hardware, and any combination of functional units can exist.

[0090] Instructions transmitted to processing cluster 214 constitute a thread. A set of threads executed across a collection of parallel processing engines is a thread group. The thread group executes the same program on different input data. Each thread within the thread group can be assigned to a different processing engine within graphics multiprocessor 234. The thread group may include fewer threads than the number of processing engines within graphics multiprocessor 234. When the thread group includes fewer threads than the number of processing engines, one or more of the processing engines may be idle during the cycles in which the thread group is being processed. The thread group may also include more threads than the number of processing engines within graphics multiprocessor 234. When the thread group includes more threads than the number of processing engines within graphics multiprocessor 234, processing can be performed in consecutive clock cycles. Optionally, multiple thread groups can be executed concurrently on graphics multiprocessor 234.

[0091] Graphics multiprocessor 234 may include an internal cache memory to perform load and store operations. Optionally, graphics multiprocessor 234 can forgo the internal cache and use the cache memory within processing cluster 214 (e.g., level 1 (L1) cache 248). Each graphics multiprocessor 234 also has access to a second-level (L2) cache within the partition units (e.g., Figure 2A partition units 220A - 220N), which are shared among all processing clusters 214 and can be used to transfer data between threads. Graphics multiprocessor 234 can also access off-chip global memory, which may include one or more of local parallel processor memory and / or system memory. Any memory external to parallel processing unit 202 can be used as global memory. In embodiments where processing cluster 214 includes multiple instances of graphics multiprocessor 234, common instructions and data can be shared and stored in L1 cache 248.

[0092] Each processing cluster 214 may include an MMU 245 (memory management unit) configured to map virtual addresses to physical addresses. In other embodiments, one or more instances of MMU 245 may reside inFigure 2A within the memory interface 218. The MMU 245 includes a set of page table entries (PTEs) for mapping virtual addresses to physical addresses of the chip and optionally includes cache line indices. The MMU 245 may include a translation lookaside buffer (TLB) or cache that may reside within the graphics multiprocessor 234 or the L1 cache 248 of the processing cluster 214. The physical address is processed to distribute surface data access locality, allowing for efficient request interleaving between partition units. The cache line index may be used to determine whether a request to a cache line is a hit or a miss.

[0093] In graphics and computing applications, the processing cluster 214 may be configured such that each graphics multiprocessor 234 is coupled to a texture unit 236 for performing texture mapping operations, e.g., determining texture sample positions, reading texture data, and filtering texture data. The texture data is read from an internal texture L1 cache (not shown) or, in some embodiments, from the L1 cache within the graphics multiprocessor 234 and fetched from the L2 cache, local parallel processor memory, or system memory as needed. Each graphics multiprocessor 234 outputs the processed tasks to the data crossbar 240 to provide the processed tasks to another processing cluster 214 for further processing or stores the processed tasks in the L2 cache, local parallel processor memory, or system memory via the memory crossbar 216. The preROP 242 (pre-raster operations unit) is configured to receive data from the graphics multiprocessor 234 and direct the data to ROP units that may be located with partition units (e.g., Figure 2A partition units 220A - 220N as described herein). The preROP 242 unit may perform optimizations for color blending, organize pixel color data, and perform address translation.

[0094] It will be appreciated that the core architecture described herein is illustrative and variations and modifications are possible. Any number of processing units (e.g., graphics multiprocessors 234, texture units 236, preROP 242, etc.) may be included within the processing cluster 214. Further, although only one processing cluster 214 is shown, the parallel processing unit as described herein may include any number of instances of the processing cluster 214. Optionally, each processing cluster 214 may be configured to operate independently of other processing clusters 214 using separate and distinct processing units, L1 caches, L2 caches, etc.

[0095] Figure 2DAn example of a graphics multiprocessor 234 is shown, where the graphics multiprocessor 234 is coupled to the pipeline manager 232 of the processing cluster 214. The graphics multiprocessor 234 has an execution pipeline that includes, but is not limited to, an instruction cache 252, an instruction unit 254, an address mapping unit 256, a register file 258, one or more general-purpose graphics processing unit (GPGPU) cores 262, and one or more load / store units 266. The GPGPU cores 262 and the load / store units 266 are coupled to a cache memory 272 and a shared memory 270 via a memory and cache interconnect 268. The graphics multiprocessor 234 may additionally include a tensor / and or ray tracing core 263, which includes hardware logic for accelerating matrix and / or ray tracing operations.

[0096] The instruction cache 252 may receive a stream of instructions to be executed from the pipeline manager 232. The instructions are cached in the instruction cache 252 and dispatched for execution by the instruction unit 254. The instruction unit 254 may dispatch the instructions as a thread group (e.g., a warp, a wavefront, etc.), where each thread in the thread group is assigned to a different execution unit within the GPGPU core 262. The instructions may access any one of a local address space, a shared address space, or a global address space by specifying an address within a unified address space. The address mapping unit 256 may be used to translate an address in the unified address space into a different memory address that can be accessed by the load / store unit 266.

[0097] The register file 258 provides a collection of registers for the functional units of the graphics multiprocessor 234. The register file 258 provides temporary storage for the operands of the data paths connected to the functional units (e.g., the GPGPU cores 262, the load / store units 266) of the graphics multiprocessor 234. The register file 258 may be partitioned among each of the functional units such that each functional unit is assigned a dedicated portion of the register file 258. For example, the register file 258 may be partitioned among different groups of units executed by the graphics multiprocessor 234.

[0098] The GPGPU cores 262 may each include a floating point unit (FPU) and / or an integer arithmetic logic unit (ALU) for executing instructions of the graphics multiprocessors 234. In some implementations, the GPGPU cores 262 may include hardware logic that would otherwise reside within the tensor and / or ray tracing cores 263. The GPGPU cores 262 may be architecturally similar or may be architecturally different. For example and in one embodiment, a first portion of the GPGPU core 262 includes a single-precision FPU and an integer ALU, while a second portion of the GPGPU core includes a double-precision FPU. Optionally, the FPU may implement the IEEE 754-2008 standard for floating point arithmetic or enable variable precision floating point arithmetic. The graphics multiprocessors 234 may additionally include one or more fixed function or special function units for performing specific functions such as copy rectangle or pixel blend operations. One or more of the GPGPU cores may also include fixed function or special function logic.

[0099] The GPGPU cores 262 may include SIMD logic capable of executing a single instruction on multiple sets of data. Optionally, the GPGPU cores 262 may physically execute SIMD4, SIMD8, and SIMD16 instructions and logically execute SIMD1, SIMD2, and SIMD32 instructions. SIMD instructions for the GPGPU cores may be generated at compile time by a shader compiler or automatically generated when executing a program written and compiled for a single program multiple data (SPMD) or SIMT architecture. Multiple threads of a program configured for the SIMT execution model may be executed via a single SIMD instruction. For example and in one embodiment, eight SIMT threads that perform the same or similar operations may be executed in parallel via a single SIMD8 logic unit.

[0100] The memory and cache interconnect 268 is an interconnect network that connects each of the functional units in the graphics multiprocessor 234 to the register file 258 and to the shared memory 270. For example, the memory and cache interconnect 268 is a crossbar interconnect that allows the load / store unit 266 to implement load and store operations between the shared memory 270 and the register file 258. The register file 258 can operate at the same frequency as the GPGPU core 262, so data transfer between the GPGPU core 262 and the register file 258 is very low latency. The shared memory 270 can be used to enable communication between threads executing on the functional units within the graphics multiprocessor 234. The cache memory 272 can be used as a data cache, for example, to cache texture data passed between the functional units and the texture unit 236. The shared memory 270 can also be used as a managed cached program. The shared memory 270 and the cache memory 272 can be coupled to the data crossbar 240 to enable communication with other components of the processing cluster. Threads executing on the GPGPU core 262 can also programmatically store data in the shared memory in addition to the automatically cached data stored in the cache memory 272.

[0101] Figures 3A - 3C Illustrates an additional graphics multiprocessor according to an embodiment. Figures 3A - 3B Illustrates graphics multiprocessors 325, 350, which are related to the Figure 2C graphics multiprocessor 234 and can be used in place of one of those graphics multiprocessors. Thus, any disclosure of a feature in connection with the graphics multiprocessor 234 herein also discloses the corresponding combination with the graphics multiprocessors 325, 350, but is not limited thereto. Figure 3C Illustrates a graphics processing unit (GPU) 380 that includes a collection of dedicated graphics processing resources arranged as multi-core groups 365A - 365N, which correspond to the graphics multiprocessors 325, 350. The illustrated graphics multiprocessors 325, 350 and the multi-core groups 365A - 365N can be streaming multiprocessors (SMs) capable of executing a large number of execution threads simultaneously.

[0102] Figure 3A The graphics multiprocessor 325 includes with respect to Figure 2DMultiple additional instances of execution resource units of the graphics multiprocessor 234. For example, the graphics multiprocessor 325 may include multiple instances of instruction units 332A - 332B, register heaps 334A - 334B, and (one or more) texture units 344A - 344B. The graphics multiprocessor 325 also includes multiple sets of graphics or compute execution units (e.g., GPGPU cores 336A - 336B, tensor cores 337A - 337B, ray tracing cores 338A - 338B) and multiple sets of load / store units 340A - 340B. The execution resource units have a common instruction cache 330, texture and / or data cache memory 342, and shared memory 346.

[0103] Each component can communicate via the interconnect structure 327. The interconnect structure 327 may include one or more crossbars to enable communication between the components of the graphics multiprocessor 325. The interconnect structure 327 is a separate, high - speed network structure layer on which each component of the graphics multiprocessor 325 is stacked. The components of the graphics multiprocessor 325 communicate with remote components via the interconnect structure 327. For example, cores 336A - 336B, 337A - 337B, and 338A - 338B can each communicate with the shared memory 346 via the interconnect structure 327. The interconnect structure 327 can arbitrate the communication within the graphics multiprocessor 325 to ensure fair bandwidth allocation between components.

[0104] Figure 3B The graphics multiprocessor 350 includes multiple sets of execution resources 356A - 356D, where, as Figure 2D and Figure 3A illustrated, each set of execution resources includes multiple instruction units, register heaps, GPGPU cores, and load / store units. The execution resources 356A - 356D can work in cooperation with (one or more) texture units 360A - 360D for texture operations while sharing the instruction cache 354 and the shared memory 353. For example, the execution resources 356A - 356D can share the instruction cache 354, the shared memory 353, and multiple instances of texture and / or data cache memories 358A - 358B. Each component can communicate via an interconnect structure 352 similar to the interconnect structure 327 of Figure 3A .

[0105] Those skilled in the art will understand that Figure 1 , Figures 2A - 2D and Figures 3A - 3BThe architecture described herein is descriptive and not restrictive in terms of the scope of the current embodiments. Thus, the techniques described herein may be implemented on any appropriately configured processing unit without departing from the scope of the embodiments described herein, including but not limited to: one or more mobile application processors; one or more desktop or server central processing units (CPUs), including multi-core CPUs; one or more parallel processor units such as, Figure 2A the parallel processing unit 202 as well as one or more graphics processors or dedicated processing units.

[0106] The parallel processors or GPGPUs described herein may be communicatively coupled to a host / processor core to accelerate graphics operations, machine learning operations, pattern analysis operations, and various general-purpose GPU (GPGPU) functions. The GPU may be communicatively coupled to the host processor / core via a bus or another interconnect (e.g., a high-speed interconnect such as PCIe, NVLink, or other known protocols, standardized protocols, or proprietary protocols). In other embodiments, the GPU may be integrated on the same package or die as the core and communicatively coupled to the core via an internal processor bus / interconnect (i.e., within the package or die). Regardless of the manner in which the GPU is connected, the processor core may allocate work to the GPU in the form of a sequence of commands / instructions contained in a work descriptor. The GPU then uses dedicated circuitry / logic to efficiently process these commands / instructions.

[0107] Figure 3C Illustrated is a graphics processing unit (GPU) 380 that includes a collection of dedicated graphics processing resources arranged as multi-core groups 365A - 365N. While details are provided for only a single multi-core group 365A, it will be appreciated that the other multi-core groups 365B - 365N may be equipped with the same or similar collections of graphics processing resources. The details described with respect to multi-core groups 365A - 365 may also apply to any graphics multiprocessor 234, 325, 350 described herein.

[0108] As shown, the multi-core group 365A may include a set 370 of graphics cores, a set 371 of tensor cores, and a set 372 of ray tracing cores. A scheduler / dispatcher 368 schedules and dispatches graphics threads for execution on the respective cores 370, 371, 372. A set 369 of register files stores operand values used by the cores 370, 371, 372 when executing graphics threads. These register files may include, for example, integer registers for storing integer values, floating-point registers for storing floating-point values, vector registers for storing packed data elements (integer and / or floating-point data elements), and tile registers for storing tensor / matrix values. The tile registers may be implemented as a combined set of vector registers.

[0109] One or more combined level 1 (L1) caches and shared memory units 373 store graphics data locally within each multi-core group 365A, such as texture data, vertex data, pixel data, ray data, bounding volume data, etc. One or more texture units 374 may also be used to perform texture operations, such as texture mapping and sampling. A level 2 (L2) cache 375 shared by all multi-core groups 365A - 365N or a subset of multi-core groups 365A - 365N stores graphics data and / or instructions for multiple concurrent graphics threads. As shown, the L2 cache 375 may be shared across multiple multi-core groups 365A - 365N. One or more memory controllers 367 couple the GPU 380 to a memory 366, which may be system memory (e.g., DRAM) and / or dedicated graphics memory (e.g., GDDR6 memory).

[0110] Input / Output (I / O) circuitry 363 couples the GPU 380 to one or more I / O devices 362, such as a digital signal processor (DSP), a network controller, or a user input device. On-chip interconnects may be used to couple the I / O devices 362 to the GPU 380 and the memory 366. One or more I / O memory management units (IOMMUs) 364 of the I / O circuitry 363 directly couple the I / O devices 362 to the system memory 366. Optionally, the IOMMU 364 manages a set of multiple page tables for mapping virtual addresses to physical addresses in the system memory 366. The I / O devices 362, the (one or more) CPUs 361, and the (one or more) GPUs 380 may then share the same virtual address space.

[0111] In one implementation of the IOMMU 364, the IOMMU 364 supports virtualization. In this case, the IOMMU 364 can manage a first set of page tables for mapping guest / graphics virtual addresses to guest / graphics physical addresses and a second set of page tables for mapping guest / graphics physical addresses to system / host physical addresses (e.g., within the system memory 366). The base address of each of the first set of page tables and the second set of page tables can be stored in a control register and swapped out during a context switch (e.g., such that a new context is provided access to the relevant page table set). Although not illustrated in Figure 3C , each of the cores 370, 371, 372, and / or the multi-core groups 365A - 365N can include translation lookaside buffers (TLBs) for caching guest virtual-to-guest physical translations, guest physical-to-host physical translations, and guest virtual-to-host physical translations.

[0112] (One or more) CPUs 361, GPUs 380, and I / O devices 362 can be integrated on a single semiconductor chip and / or chip package. The illustrated memory 366 can be integrated on the same chip or can be coupled to the memory controller 367 via an off-chip interface. In one implementation, the memory 366 includes GDDR6 memory that shares the same virtual address space as other physical system-level memories, but the basic principles described herein are not limited to this particular implementation.

[0113] The tensor core 371 can include a plurality of execution units specifically designed to perform matrix operations, which are fundamental computational operations for performing deep learning operations. For example, synchronous matrix multiplication operations can be used for neural network training and inference. The tensor core 371 can perform matrix processing using various operand precisions, including single-precision floating point (e.g., 32 bits), half-precision floating point (e.g., 16 bits), integer (16 bits), byte (8 bits), and nibble (4 bits). For example, a neural network implementation extracts features of each rendered scene, potentially combining details from multiple frames to construct a high-quality final image.

[0114] In a deep learning implementation, parallel matrix multiplication work can be scheduled for execution on the tensor core 371. Training of neural networks in particular requires a large number of matrix dot product operations. To handle the inner product formulation of an N x N x N matrix multiplication, the tensor core 371 can include at least N dot product processing elements. Before the matrix multiplication begins, a complete matrix is loaded into the on-chip registers, and for each of the N loops, at least one column of the second matrix is loaded. For each loop, there are N dot products to be processed.

[0115] Depending on the particular implementation, matrix elements can be stored with different precisions, including 16-bit words, 8-bit bytes (e.g., INT8), and 4-bit nibbles (e.g., INT4). Different precision modes can be specified for the tensor core 371 to ensure that the most efficient precision is used for different workloads (e.g., inference workloads, which can tolerate quantization down to bytes and nibbles). The supported formats additionally include 64-bit floating point (FP64) and non-IEEE floating point formats, such as the bfloat16 format (e.g., Brain floating point), a 16-bit floating point format with one sign bit, eight exponent bits, and eight significand bits (seven of which are explicitly stored). One embodiment includes support for a reduced precision tensor floating point (TF32) mode that performs calculations using the range of FP32 (8 bits) and the precision of FP16 (10 bits). Reduced precision TF32 operations can be performed on FP32 inputs with higher performance relative to FP32 and increased precision relative to FP16 and produce FP32 outputs. In one embodiment, one or more 8-bit floating point formats (FP32) are supported.

[0116] In one embodiment, the tensor core 371 supports a sparse operation mode for matrices in which the vast majority of values are zero. The tensor core 371 includes support for sparse input matrices encoded in a sparse matrix representation (e.g., coordinate list encoding (COO), compressed sparse row (CSR), compressed sparse column (CSC), etc.). The tensor core 371 also includes support for a compressed sparse matrix representation in cases where the sparse matrix representation can be further compressed. Compressed matrix data, encoded matrix data, and / or compressed and encoded matrix data, as well as associated compression and / or encoding metadata, can be read by the tensor core 371, and non-zero values can be extracted. For example, for a given input matrix A, non-zero values can be loaded from at least a portion of the compressed and / or encoded representation of matrix A. Based on the positions of the non-zero values in matrix A (which can be determined from the indices or coordinate metadata associated with the non-zero values), the corresponding values in the input matrix B can be loaded. Depending on the operation to be performed (e.g., multiplication), if the corresponding value is a zero value, the loading of the value from the input matrix B can be bypassed. In one embodiment, the pairing of values for certain operations (such as multiplication operations) can be pre-scanned by the scheduler logic, and only operations between non-zero inputs are scheduled. Depending on the dimensions of matrices A and B and the operation to be performed, the output matrix C can be dense or sparse. In the case where the output matrix C is sparse and depending on the configuration of the tensor core 371, the output matrix C can be output in a compressed format, sparse encoding, or compressed sparse encoding.

[0117] The ray tracing core 372 can accelerate ray tracing operations for both real-time ray tracing implementations and non-real-time ray tracing implementations. Specifically, the ray tracing core 372 can include a ray traversal / intersection circuit that uses a bounding volume hierarchy (BVH) to perform ray traversal and identify intersections between rays enclosed within the BVH volume and primitives. The ray tracing core 372 can also include circuitry for performing depth testing and culling (e.g., using a Z-buffer or a similar arrangement). In one implementation, the ray tracing core 372 performs traversal and intersection operations in cooperation with the image denoising techniques described herein, at least part of which can be performed on the tensor core 371. For example, the tensor core 371 can implement a deep learning neural network to perform denoising of frames generated by the ray tracing core 372. However, the (one or more) CPUs 361, the graphics core 370, and / or the ray tracing core 372 can also implement all or part of the denoising and / or deep learning algorithms.

[0118] In addition, as described above, a distributed approach to denoising can be employed, where the GPU 380 is in a computing device coupled to other computing devices via a network or a high-speed interconnect. According to this distributed approach, the interconnected computing devices can share neural network learning / training data to improve the speed at which the overall system learns to perform denoising for different types of image frames and / or different graphics applications.

[0119] The ray tracing core 372 can handle all BVH traversals and / or ray-primitive intersections, freeing the graphics core 370 from being overloaded with thousands of instructions per ray. For example, each ray tracing core 372 includes a first set of specialized circuitry for performing bounding box tests (e.g., for traversal operations) and / or a second set of specialized circuitry for performing ray-triangle intersection tests (e.g., intersecting the traversed rays). Thus, for example, the multi-core group 365A can simply initiate a ray probe, and the ray tracing core 372 independently performs ray traversal and intersection and returns hit data (e.g., hit, no hit, multiple hits, etc.) to the thread context. When the ray tracing core 370 performs traversal and intersection operations, the other cores 371, 372 are freed up to perform other graphics or computing work.

[0120] Optionally, each ray tracing core 372 may include a traversal unit for performing BVH test operations and / or an intersection unit for performing ray-primitive intersection tests. The intersection unit generates "hit", "miss", or "multiple hit" responses, which it provides to the appropriate threads. During traversal and intersection operations, the execution resources of other cores (e.g., the graphics core 370 and the tensor core 371) are freed to perform other forms of graphics work.

[0121] In an optional embodiment described below, a hybrid rasterization / ray tracing method is used in which work is distributed between the graphics core 370 and the ray tracing core 372.

[0122] The ray tracing core 372 (and / or other cores 370, 371) may include hardware support for a ray tracing instruction set, such as Microsoft's DirectX Ray Tracing (DXR), which includes the DispatchRays command; and ray generation shaders, closest hit shaders, any hit shaders, and miss shaders, which enable the assignment of a unique set of shaders and textures to each object. Another ray tracing platform that may be supported by the ray tracing core 372, the graphics core 370, and the tensor core 371 is the Vulkan API (e.g., Vulkan version 1.1.85, or later versions). However, note that the basic principles described herein are not limited to any particular ray tracing ISA.

[0123] Generally, the individual cores 372, 371, 370 may support a ray tracing instruction set that includes instructions / functions for one or more of the following: ray generation, closest hit, any hit, ray-primitive intersection, per-primitive and hierarchical bounding box construction, miss, traverse, and exception. More specifically, the preferred embodiment includes ray tracing instructions for performing one or more of the following functions:

[0124] Light Generation - Ray generation instructions may be executed for each pixel, sample, or other user-defined work assignment.

[0125] Nearest Hit - Closest hit instructions may be executed to locate the closest intersection of a ray with a primitive within a scene.

[0126] Any Hit - Any hit instructions identify multiple intersections between a ray and primitives within a scene, potentially identifying new closest intersections.

[0127] Intersection - Intersection instructions perform ray-primitive intersection tests and output the results.

[0128] Primitive - by - Primitive Bounding Box Construction —— This instruction builds a bounding box around a given primitive or group of primitives (e.g., when building a new BVH or other acceleration data structure).

[0129] Miss —— Indicates that the ray misses all geometries within the scene or a specified region of the scene.

[0130] Visit —— Indicates the child volume that the ray will traverse.

[0131] Exception —— Includes various types of exception handlers (e.g., called for various error conditions).

[0132] In one embodiment, the ray tracing core 372 may be adapted to accelerate general-purpose computing operations, which may use computational techniques similar to ray intersection tests for acceleration. A computational framework may be provided that enables shader programs to be compiled into low-level instructions and / or primitives for performing general-purpose computing operations via the ray tracing core. Exemplary computational problems that may benefit from computational operations performed on the ray tracing core 372 include computations involving the propagation of light beams, waves, rays, or particles within a coordinate space. Interactions associated with that propagation may be computed relative to geometries or meshes within the coordinate space. For example, computations associated with the propagation of electromagnetic signals through an environment may be accelerated via the use of instructions or primitives executed via the ray tracing core. Refraction and reflection of signals by objects in the environment may be computed as a direct ray tracing simulation.

[0133] The ray tracing core 372 may also be used to perform computations not directly similar to ray tracing. For example, the ray tracing core 372 may be used to accelerate mesh projection, mesh refinement, and volume sampling computations. General coordinate space computations may also be performed, such as nearest neighbor computations. For example, a set of points near a given point may be discovered by defining a bounding box around the point in the coordinate space. Subsequently, the BVH and ray probing logic within the ray tracing core 372 may be used to determine the set of point intersections within the bounding box. The intersections constitute the origin and the nearest neighbors of that origin. Computations performed using the ray tracing core 372 may be executed in parallel with computations performed on the graphics core 372 and the tensor core 371. The shader compiler may be configured to compile compute shaders or other general-purpose graphics programs into low-level primitives that can be parallelized across the graphics core 370, the tensor core 371, and the ray tracing core 372. Techniques for GPU - to - Host Processor Interconnection

[0134] Figure 4A Illustrated in which multiple GPUs 410-413 (e.g., such as Figure 2AThe exemplary architecture of the parallel processor 200 shown is communicatively coupled to a plurality of multi-core processors 405-406 via high-speed links 440A-440D (e.g., buses, point-to-point interconnects, etc.). Depending on the implementation, the high-speed links 440A-440D may support a communication throughput of 4GB / s, 30GB / s, 80GB / s, or higher. Various interconnect protocols may be used, including but not limited to PCIe 4.0 or 5.0 and NVLink 2.0. However, the basic principles described herein are not limited to any particular communication protocol or throughput.

[0135] Two or more of the GPUs 410-413 may be interconnected via high-speed links 442A-442B, which may be implemented using the same or different protocols / links as those used for the high-speed links 440A-440D. Similarly, two or more of the multi-core processors 405-406 may be connected via a high-speed link 443, which may be a symmetric multi-processor (SMP) bus operating at 20GB / s, 30GB / s, 120GB / s, or lower or higher speeds. Alternatively, Figure 4A All communication between the various system components shown may be implemented using the same protocol / link (e.g., via a common interconnect structure). However, as mentioned, the basic principles described herein are not limited to any particular type of interconnect technology.

[0136] Each of the multi-core processors 405 and 406 may be communicatively coupled to the processor memories 401-402 via the memory interconnects 430A-430B, respectively, and each of the GPUs 410-413 is communicatively coupled to the GPU memories 420-423 via the GPU memory interconnects 450A-450D, respectively. The memory interconnects 430A-430B and 450A-450D may utilize the same or different memory access technologies. By way of example and not limitation, the processor memories 401-402 and the GPU memories 420-423 may be volatile memories such as dynamic random access memory (DRAM) (including stacked DRAM), graphics double data rate SDRAM (GDDR) (e.g., GDDR5, GDDR6), or high bandwidth memory (HBM), and / or may be non-volatile memories such as 3D Xpoint / Optane or Nano-Ram. For example, a portion of the memory may be volatile memory and another portion may be non-volatile memory (e.g., using a two-level memory (2LM) hierarchy). The memory subsystems described herein may be compatible with several memory technologies such as double data rate versions published by JEDEC (Joint Electronic Device Engineering Council).

[0137] As described below, although the processors 405-406 and the GPUs 410-413 may be physically coupled to specific memories 401-402, 420-423, respectively, a unified memory architecture may be implemented in which the same virtual system address space (also referred to as the "effective address" space) is distributed among all the various physical memories. For example, each of the processor memories 401-402 may include a 64 GB system memory address space, and each of the GPU memories 420-423 may include a 32 GB system memory address space (resulting in a total of 256 GB of addressable memory in this example).

[0138] Figure 4B Additional optional details of the interconnection between the multi-core processor 407 and the graphics acceleration module 446 are illustrated. The graphics acceleration module 446 may include one or more GPU chips integrated on a line card that is coupled to the processor 407 via a high-speed link 440. Alternatively, the graphics acceleration module 446 may be integrated on the same package or chip as the processor 407.

[0139] The illustrated processor 407 includes multiple cores 460A - 460D, each of which has a translation lookaside buffer 461A - 461D and one or more caches 462A - 462D. The cores may include various other components for executing instructions and processing data, which are not illustrated to avoid obscuring the basic principles of the components described herein (e.g., instruction fetch unit, branch prediction unit, decoder, execution unit, reorder buffer, etc.). The caches 462A - 462D may include a first - level (L1) cache and a second - level (L2) cache. Additionally, one or more shared caches 456 may be included in the cache hierarchy and shared by a set of cores 460A - 460D. For example, one embodiment of the processor 407 includes 24 cores, each of which has its own L1 cache, twelve shared L2 caches, and twelve shared L3 caches. In this embodiment, one of the L2 caches and L3 caches is shared by two adjacent cores. The processor 407 and the graphics accelerator integration module 446 are connected to the system memory 441, which may include processor memories 401 - 402.

[0140] Coherence is maintained for data and instructions stored in the caches 462A - 462D, 456, and the system memory 441 via inter - core communication through the coherence bus 464. For example, each cache may have cache coherence logic / circuit associated therewith to communicate via the coherence bus 464 in response to a detected read or write to a particular cache line. In one implementation, a cache snooping protocol is implemented via the coherence bus 464 to snoop on cache accesses. Cache snooping / coherence techniques are well understood by those skilled in the art and will not be described in detail herein to avoid obscuring the basic principles described herein.

[0141] Agent circuitry 425 may be provided to communicatively couple the graphics acceleration module 446 to the coherence bus 464, thereby allowing the graphics acceleration module 446 to participate in the cache coherence protocol as a peer of the cores. Specifically, interface 435 provides connectivity to the agent circuitry 425 via a high - speed link 440 (e.g., PCIe bus, NVLink, etc.), and interface 437 connects the graphics acceleration module 446 to the high - speed link 440.

[0142] In one implementation, the accelerator integrated circuit 436 provides cache management, memory access, context management, and interrupt management services for multiple graphics processing engines 431, 432...N of the graphics acceleration module 446. Each of the graphics processing engines 431, 432...N may include separate graphics processing units (GPUs). Alternatively, the graphics processing engines 431, 432...N may include different types of graphics processing engines within a GPU, such as graphics execution units, media processing engines (e.g., video encoder / decoder), samplers, and block image transfer (BLIT) engines. In other words, the graphics acceleration module may be a GPU having multiple graphics processing engines 431 - 432...N, or the graphics processing engines 431 - 432...N may be separate GPUs integrated on a common package, line card, or chip. The graphics processing engines 431 - 432...N may be configured using any of the graphics processor or computing accelerator architectures described herein.

[0143] The accelerator integrated circuit 436 may include a memory management unit (MMU) 439 that performs various memory management functions such as virtual - to - physical memory translation (also known as effective - to - actual memory translation) and a memory access protocol for accessing system memory 441. The MMU 439 may also include a translation lookaside buffer (TLB) (not shown) for caching virtual / effective - to - physical / actual address translations. In one implementation, the cache 438 stores commands and data for efficient access by the graphics processing engines 431, 432...N. The data stored in the cache 438 and the graphics memories 433 - 434...M may be kept consistent with the core caches 462A - 462D, 456, and the system memory 441. As mentioned, this may be done via the proxy circuit 425, which participates in the cache coherence mechanism on behalf of the cache 438 and the memories 433 - 434...M (e.g., sending updates related to the modification / access of cache lines on the processor caches 462A - 462D, 456 to the cache 438 and receiving updates from the cache 438).

[0144] The set of registers 445 stores context data for threads to be executed by the graphics processing engines 431 - 432...N, and the context management circuitry 448 manages these thread contexts. For example, the context management circuitry 448 may perform save and restore operations to save and restore the context of each thread during a context switch (e.g., where the first thread is saved and the second thread is restored such that the second thread can be executed by the graphics processing engine). For example, during a context switch, the context management circuitry 448 may store the current register values into a specified area in memory (e.g., identified by a context pointer). When returning to that context, it may then restore the register values. The interrupt management circuitry 447 may, for example, receive interrupts from system devices and process the interrupts received from the system devices.

[0145] In one implementation, the MMU 439 translates virtual / valid addresses from the graphics processing engine 431 into actual / physical addresses in the system memory 441. Optionally, the accelerator integrated circuit 436 supports multiple (e.g., 4, 8, 16) graphics accelerator modules 446 and / or other accelerator devices. The graphics accelerator modules 446 may be dedicated to a single application executing on the processor 407 or may be shared among multiple applications. Optionally, a virtualized graphics execution environment is provided in which the resources of the graphics processing engines 431 - 432...N are shared among multiple applications, virtual machines (VMs), or containers. The resources may be subdivided into "slices" that are allocated to different VMs and / or applications based on the processing requirements and priorities associated with the VMs and / or applications or based on a predefined partitioning profile for the graphics accelerator modules 446. VMs and containers may be used interchangeably herein.

[0146] A virtual machine (VM) may be software that runs an operating system and one or more applications. The VM may be defined by a specification, configuration file, virtual disk file, non-volatile random access memory (NVRAM) settings file, and log files, and is backed up by the physical resources of the host computing platform. The VM may include an operating system (OS) or application environment installed on software that emulates dedicated hardware. End users have the same experience on the virtual machine as they would have on dedicated hardware. Specialized software called a hypervisor fully emulates the CPU, memory, hard disk, network, and other hardware resources of a PC client or server, enabling the virtual machines to share resources. The hypervisor may emulate multiple virtual hardware platforms that are isolated from each other, allowing virtual machines to run on the same underlying physical host. Servers, VMware ESXi, and other operating systems.

[0147] A container can be a software package of an application, its configuration, and dependencies, so that the application runs reliably from one computing environment to another. Containers can share the operating system installed on the server platform and run as isolated processes. A container can be a software package that contains everything needed for the software to run (such as system tools, libraries, and settings). Containers are not installed like traditional software programs, which allows them to be isolated from other software and from the operating system itself. The isolated nature of containers provides several benefits. First, the software in a container will run the same way in different environments. For example, a container including PHP and MySQL can run in exactly the same way on both computers and machines. Second, containers provide increased security because the software will not affect the host operating system. While installed applications can change system settings and modify resources (such as the Windows registry), a container can only modify the settings within that container.

[0148] Accordingly, the accelerator integrated circuit 436 acts as a bridge to the system for the graphics acceleration module 446 and provides address translation and system memory caching services. In one embodiment, to facilitate the bridging function, the accelerator integrated circuit 436 may also include shared I / O 497 (e.g., PCIe, USB, or other components) and hardware to enable system control of voltage, clock control, performance, thermal, and security. The shared I / O 497 may utilize separate physical connections or may span the high-speed link 440. Additionally, the accelerator integrated circuit 436 may provide virtualization facilities for the host processor to manage virtualization of the graphics processing engine, interrupts, and memory management.

[0149] Since the hardware resources of the graphics processing engines 431 - 432...N are explicitly mapped to the actual address space seen by the host processor 407, any host processor can use valid address values to directly address these resources. An optional function of the accelerator integrated circuit 436 is to physically separate the graphics processing engines 431 - 432...N such that they appear to the system as independent units.

[0150] One or more graphics memories 433 - 434...M may be coupled respectively to each of the graphics processing engines 431 - 432...N. The graphics memories 433 - 434...M store instructions and data processed by each of the graphics processing engines 431 - 432...N. The graphics memories 433 - 434...M may be volatile memories such as DRAM (including stacked DRAM), GDDR memory (e.g., GDDR5, GDDR6), or HBM, and / or may be non - volatile memories such as 3D Xpoint / Optane, Samsung Z - NAND, or Nano - Ram.

[0151] To reduce data traffic on the high - speed link 440, a biasing technique may be used to ensure that the data stored in the graphics memories 433 - 434...M is the data that will be most frequently used by the graphics processing engines 431 - 432...N and preferably not used (at least not frequently) by the cores 460A - 460D. Similarly, the biasing mechanism attempts to keep the data needed by the cores (and preferably not the graphics processing engines 431 - 432...N) in the system memory 441 and the caches 462A - 462D, 456 of the cores.

[0152] According to Figure 4C In the variant shown in Figure 4B the accelerator integrated circuit 436 is integrated within the processor 407. The graphics processing engines 431 - 432...N communicate directly with the accelerator integrated circuit 436 via the high - speed link 440, through the interface 437 and the interface 435 (which again may utilize any form of bus or interface protocol). The accelerator integrated circuit 436 may perform the same operations as those described with respect to

[0153] The described embodiments may support different programming models, which include a dedicated process programming model (without graphics acceleration module virtualization) and a shared programming model (with virtualization). The latter may include a programming model controlled by the accelerator integrated circuit 436 and a programming model controlled by the graphics acceleration module 446.

[0154] In an embodiment of the dedicated process model, the graphics processing engines 431, 432,..., N may be dedicated to a single application or process under a single operating system. A single application may allow other application requests to leak to the graphics engines 431, 432,..., N, thus providing virtualization within the VM / partition.

[0155] In the dedicated process programming model, the graphics processing engines 431, 432... N can be shared by multiple VM / application partitions. The sharing model requires the hypervisor to virtualize the graphics processing engines 431 - 432... N to allow access by each operating system. For a single - partition system without a hypervisor, the graphics processing engines 431 - 432... N are owned by the operating system. In both cases, the operating system can virtualize the graphics processing engines 431 - 432... N to provide access rights to each process or application.

[0156] For the shared programming model, the graphics acceleration module 446 or individual graphics processing engines 431 - 432... N use a process handle to select process elements. The process elements can be stored in the system memory 441 and can be addressable using the effective - address - to - physical - address translation techniques described herein. The process handle can be an implementation - specific value provided to the host process when registering its context with the graphics processing engines 431 - 432... N (i.e., calling system software to add the process element to the process - element linked list). The lower 16 bits of the process handle can be the offset of the process element within the process - element linked list.

[0157] Figure 4D FIG. illustrates an exemplary accelerator integration slice 490. As used herein, "slice" includes a designated portion of the processing resources of the accelerator integrated circuit 436. The application effective - address space 482 within the system memory 441 stores process elements 483. The process elements 483 can be stored in response to a GPU call 481 from an application 480 executing on the processor 407. The process elements 483 contain the process state of the corresponding application 480. The work descriptor (WD) 484 contained in the process element 483 can be a single job requested by the application or can contain a pointer to a job queue. In the latter case, the WD 484 is a pointer to the job - request queue in the application's address space 482.

[0158] The graphics acceleration module 446 and / or individual graphics processing engines 431 - 432... N can be shared by all processes in the system or a subset of the processes in the system. For example, the techniques described herein can include infrastructure for establishing process state and sending the WD 484 to the graphics acceleration module 446 to start a job in a virtualized environment.

[0159] In one implementation, the dedicated process programming model is implementation-specific. In this model, a single process owns the graphics acceleration module 446 or a separate graphics processing engine 431. Since the graphics acceleration module 446 is owned by a single process, when the graphics acceleration module 446 is assigned, the hypervisor initializes the accelerator integrated circuit 436 for the owning partition, and the operating system initializes the accelerator integrated circuit 436 for the owning process.

[0160] In operation, the WD fetch unit 491 in the accelerator integrated slice 490 fetches the next WD 484, which includes an indication of work to be completed by one of the graphics processing engines in the graphics acceleration module 446. As shown, data from the WD 484 can be stored in the register 445 and used by the MMU 439, the interrupt management circuit 447, and / or the context management circuit 448. For example, the MMU 439 can include a segment / page walk circuit for accessing the segment table / page table 486 within the OS virtual address space 485. The interrupt management circuit 447 can process the interrupt event 492 received from the graphics acceleration module 446. When performing a graphics operation, the effective address 493 generated by the graphics processing engines 431 - 432...N is translated into a physical address by the MMU 439.

[0161] The same set of registers 445 can be replicated for each graphics processing engine 431 - 432...N and / or the graphics acceleration module 446 and can be initialized by the hypervisor or the operating system. Each of these replicated registers can be included in the accelerator integrated slice 490. In one embodiment, each graphics processing engine 431 - 432...N can be presented to the hypervisor 496 as a different graphics processor device. QoS settings can be configured for the clients of a particular graphics processing engine 431 - 432...N, and data isolation between the clients of each engine can be enabled. Exemplary registers that can be initialized by the hypervisor are shown in Table 1. Table 1 - Registers Initialized by the Hypervisor 1 Slice Control Register 2 Process Region Pointer Scheduled by Real Address (RA) 3 Permission Mask Override Register 4 Interrupt Vector Table Entry Offset 5 Interrupt Vector Table Entry Limit 6 Status Register 7 Logical Partition ID 8 Real Address (RA) Hypervisor Accelerator Utilization Record Pointer 9 Storage Descriptor Register

[0162] Exemplary registers that can be initialized by the operating system are shown in Table 2. Table 2 - Registers Initialized by the Operating System 1 Process and Thread Identification 2 Effective Address (EA) Context Save / Resume Pointer 3 Virtual Address (VA) Accelerator Utilization Record Pointer 4 Virtual Address (VA) Storage Segment Table Pointer 5 Permission Mask 6 Work Descriptor

[0163] Each WD 484 can be specific to a particular graphics acceleration module 446 and / or graphics processing engines 431 - 432...N. It contains all the information that the graphics processing engines 431 - 432...N need to do their job, or it can be a pointer to a memory location where a command queue for the work to be done by the application has been established.

[0164] Figure 4E Additional optional details of the shared model are illustrated. It includes the hypervisor physical address space 498 in which the process element list 499 is stored. The hypervisor physical address space 498 is accessible via the hypervisor 496, which virtualizes the graphics acceleration module engine for the operating system 495.

[0165] The shared programming model allows all processes or a subset of processes from all partitions in the system or a subset of partitions in the system to use the graphics acceleration module 446. There are two programming models in which the graphics acceleration module 446 is shared by multiple processes and partitions: time - division sharing and graphics - directed sharing.

[0166] In this model, the system hypervisor 496 owns the graphics acceleration module 446 and makes its functionality available to all operating systems 495. To enable the graphics acceleration module 446 to support virtualization by the system hypervisor 496, the graphics acceleration module 446 may comply with the following requirements: 1) The job requests of the application must be autonomous (i.e., the state does not need to be maintained between jobs), or the graphics acceleration module 446 must provide a context save and restore mechanism. 2) The graphics acceleration module 446 must ensure that the job requests of the application are completed within the specified amount of time, including any translation errors, or the graphics acceleration module 446 provides the ability to preempt the processing of jobs. 3) The graphics acceleration module 446 must be guaranteed fairness between processes when operating in the directed - sharing programming model.

[0167] For the directional sharing model, application 480 may be required to make an operating system 495 system call with the graphics acceleration module 446 type, work descriptor (WD), authority mask register (AMR) value, and context save / restore area pointer (CSRP). The graphics acceleration module 446 type describes the target acceleration function for the system call. The graphics acceleration module 446 type can be a system-specific value. The WD is specifically formatted for the graphics acceleration module 446, and the WD can take the following forms: a graphics acceleration module 446 command, a valid address pointer to a user-defined structure, a valid address pointer to a command queue, or any other data structure used to describe the work to be done by the graphics acceleration module 446. In one embodiment, the AMR value is the AMR state to be used for the current process. The value is passed to the operating system similar to the application setting the AMR. If the accelerator integrated circuit 436 and the graphics acceleration module 446 implementation do not support the User Authority Mask Override Register (UAMOR), the operating system may apply the current UAMOR value to the AMR value before passing the AMR in the hypervisor call. The hypervisor 496 may optionally apply the current Authority Mask Override Register (AMOR) value before placing the AMR in the process element 483. The CSRP can be one of the registers 445 that contains a valid address of a region in the address space 482 of the application for the graphics acceleration module 446 to use to save and restore the context state. If no state is required to be saved between jobs or when a job is preempted, the pointer is optional. The context save / restore area can be pinned system memory.

[0168] Upon receiving the system call, the operating system 495 may verify that the application 480 is registered and has been granted permission to use the graphics acceleration module 446. The operating system 492 then calls the hypervisor 496 with the information shown in Table 3. Table 3 - OS Call Parameters to the Hypervisor

[0169] Upon receiving a hypervisor call, the hypervisor 496 verifies that the operating system 495 is registered and has been granted permission to use the graphics acceleration module 446. The hypervisor 496 then places the process element 483 in a linked list of process elements for the corresponding type of graphics acceleration module 446. The process element may include the information shown in Table 4. Table 4 - Process Element Information 1 Work Descriptor (WD) 2 Permission Mask Register (AMR) Value (Potentially Masked). 3 Effective Address (EA) Context Save / Resume Region Pointer (CSRP) 4 Process ID (PID) and Optional Thread ID (TID) 5 Virtual Address (VA) Accelerator Utilization Record Pointer (AURP) 6 Virtual Address of Storage Segment Table Pointer (SSTP) 7 Logical Interrupt Service Number (LISN) 8 Interrupt Vector Table, Derived from Hypervisor Call Parameters 9 Status Register (state register, SR) Value 10 Logical Partition ID (logical partition ID, LPID) 11 Real Address (RA) Hypervisor Accelerator Utilization Record Pointer 12 Storage Descriptor Register (Storage Descriptor Register, SDR)

[0170] The hypervisor may initialize the registers 445 of the plurality of accelerator integration slices 490.

[0171] As Figure 4F illustrated, in an optional implementation, a unified memory addressable via a common virtual memory address space is employed, and this common virtual memory address space is used to access the physical processor memories 401 - 402 and the GPU memories 420 - 423. In this implementation, operations executed on the GPUs 410 - 413 utilize the same virtual / effective memory address space to access the processor memories 401 - 402 and vice versa, thereby simplifying programmability. A first portion of the virtual / effective address space may be allocated to the processor memory 401, a second portion may be allocated to the second processor memory 402, a third portion may be allocated to the GPU memory 420, and so on. The entire virtual / effective memory space (sometimes referred to as the effective address space) may thus be distributed across each of the processor memories 401 - 402 and the GPU memories 420 - 423, allowing any processor or GPU to access a physical memory using a virtual address mapped to that physical memory.

[0172] One or more bias / coherency management circuits 494A - 494E may be provided within one or more of the MMUs 439A - 439E. These bias / coherency management circuits 494A - 494E ensure cache coherency between the cache of the host processor (e.g., 405) and the caches of the GPUs 410 - 413, and implement a bias technique for the physical memory indicating where certain types of data should be stored. Although multiple instances of the bias / coherency management circuits 494A - 494E are illustrated Figure 4F in, the bias / coherency circuits may be implemented within the MMU of one or more host processors 405 and / or within the accelerator integrated circuit 436.

[0173] The GPU-attached memories 420-423 can be mapped as part of the system memory and accessed using shared virtual memory (SVM) technology, but do not suffer from the typical performance drawbacks associated with full system cache coherence. The ability of the GPU-attached memories 420-423 to be accessed as system memory without heavy cache coherence overhead provides a beneficial operating environment for GPU migration. This arrangement allows the host processor 405 to set up the operation objects and access the computation results without the overhead of traditional I / O DMA data copying. Such traditional copying involves driver calls, interrupts, and memory mapped I / O (MMIO) accesses, all of which are inefficient compared to simple memory accesses. At the same time, the ability to access the GPU-attached memories 420-423 without cache coherence overhead can be critical to the execution time of the migrated computations. For example, in the case of a large amount of streaming write memory traffic, the cache coherence overhead can significantly reduce the effective write bandwidth seen by the GPUs 410-413. The efficiency of operation object setup, result access, and GPU computation all play a role in determining the effectiveness of GPU migration.

[0174] The choice between GPU bias and host processor bias can be driven by a bias tracker data structure. For example, a bias table can be used, which can be a page-granularity structure (i.e., controlled at the granularity of memory pages), and this page-granularity structure includes one or two bits per GPU-attached memory page. The bias table can be implemented in the stolen memory ranges of one or more of the GPU-attached memories 420-423, with or without a bias cache in the GPUs 410-413 (e.g., for caching frequently / recently used entries of the bias table). Alternatively, the entire bias table can be maintained within the GPU.

[0175] In one implementation, before an actual access to the GPU memory, bias table entries associated with each access to the memories 420 - 423 attached to the GPUs are accessed, resulting in the following operations. First, local requests from the GPUs 410 - 413 for pages that are found in the GPU bias are directly forwarded to the corresponding GPU memories 420 - 423. Local requests from the GPUs for pages that are found in the host bias are forwarded to the processor 405 (e.g., via a high - speed link as discussed above). Optionally, requests from the processor 405 for pages that are found in the host processor bias complete the request as a normal memory read. Alternatively, requests for pages that involve the GPU bias can be forwarded to the GPUs 410 - 413. If the GPU is not currently using the page, the GPU can then transition the page to the host processor bias.

[0176] The bias state of a page can be changed by a software - based mechanism, a hardware - assisted software - based mechanism, or for a limited set of cases by a pure - hardware - based mechanism.

[0177] A mechanism for changing the bias state employs an API call (e.g., OpenCL), which in turn calls the device driver of the GPU. The device driver then sends a message (or enqueues a command descriptor) to the GPU that instructs the GPU to change the bias state and perform a cache dump flush operation in the host for some transitions. A cache dump flush operation is required for the transition from the host processor 405 bias to the GPU bias, but not for the reverse transition.

[0178] Cache coherence can be maintained by temporarily rendering pages in the GPU bias that cannot be cached by the host processor 405. To access these pages, the processor 405 can request access from the GPU 410, which, depending on the implementation, may or may not grant immediate access. Thus, to reduce communication between the host processor 405 and the GPU 410, it is beneficial to ensure that the pages in the GPU bias are those that the GPU needs but not the host processor 405, and vice versa. Graphics Processing Pipeline

[0179] Figure 5 Illustrated is a graphics processing pipeline 500. Graphics multiprocessors (such as Figure 2D the graphics multiprocessor 234 in Figure 3A the graphics multiprocessor 325 of Figure 3B the graphics multiprocessor 350 of) can implement the illustrated graphics processing pipeline 500. The graphics multiprocessors can be included within a parallel processing subsystem as described herein, such as Figure 2A the parallel processor 200 of, which can be associated withFigure 1 is associated with the (one or more) parallel processors 112 and can be used in place of one of those parallel processors. Various parallel processor systems can implement the graphics processing pipeline 500 via one or more instances of a parallel processing unit (e.g., Figure 2A the parallel processing unit 202) as described herein. For example, shader units (e.g., Figure 2C the graphics multiprocessor 234) can be configured to perform the functions of one or more of the following: vertex processing unit 504, tessellation control processing unit 508, tessellation evaluation processing unit 512, geometry processing unit 516, and fragment / pixel processing unit 524. The functions of data assembler 502, primitive assembler 506, 514, 518, tessellation unit 510, rasterizer 522, and raster operation unit 526 can also be performed by other processing engines and corresponding partitioning units (e.g., Figure 2A the processing cluster 214) within a processing cluster (e.g., Figure 2A the partitioning units 220A - 220N). The graphics processing pipeline 500 can also be implemented using dedicated processing units for one or more functions. It is also possible for one or more parts of the graphics processing pipeline 500 to be executed by parallel processing logic within a general - purpose processor (e.g., CPU). Optionally, one or more parts of the graphics processing pipeline 500 can access on - chip memory (e.g., the parallel processor memory 222 as in Figure 2A ) via a memory interface 528, which can be an instance of Figure 2A the memory interface 218. The graphics processor pipeline 500 can also be implemented via a multi - core group 365A as in Figure 3C .

[0180] The data assembler 502 is a processing unit that can collect vertex data for surfaces and primitives. The data assembler 502 then outputs the vertex data, which includes vertex attributes, to the vertex processing unit 504. The vertex processing unit 504 is a programmable execution unit that executes a vertex shader program to illuminate and transform the vertex data as specified by the vertex shader program. The vertex processing unit 504 reads data stored in a cache, local, or system memory for use in processing the vertex data and can be programmed to transform the vertex data from an object - based coordinate representation to a world - space coordinate space or a normalized device coordinate space.

[0181] A first instance of primitive assembler 506 receives vertex attributes from vertex processing unit 504. The primitive assembler 506 reads stored vertex attributes as needed and constructs graphics primitives for processing by tessellation control processing unit 508. Graphics primitives include triangles, line segments, points, patches, etc. as supported by various graphics processing application programming interfaces (APIs).

[0182] Tessellation control processing unit 508 treats input vertices as control points of a geometric patch. The control points are transformed from an input representation (e.g., the basis of the patch) from the patch to a representation suitable for use in surface evaluation by tessellation evaluation processing unit 512. Tessellation control processing unit 508 may also compute tessellation factors for the edges of the geometric patch. The tessellation factors are applied to individual edges and quantify the view-dependent level of detail associated with that edge. Tessellation unit 510 is configured to receive the tessellation factors for the edges of the patch and to tessellate the patch surface into multiple geometric primitives (such as line, triangle, or quadrilateral primitives), which are transmitted to tessellation evaluation processing unit 512. Tessellation evaluation processing unit 512 operates on the parameterized coordinates of the tessellated patch to generate a surface representation and vertex attributes for each vertex associated with the geometric primitive.

[0183] A second instance of primitive assembler 514 receives vertex attributes from tessellation evaluation processing unit 512 as needed, reads stored vertex attributes, and constructs graphics primitives for processing by geometry processing unit 516. Geometry processing unit 516 is a programmable execution unit that executes a geometry shader program to transform the graphics primitives received from primitive assembler 514 as specified by the geometry shader program. Geometry processing unit 516 can be programmed to subdivide a graphics primitive into one or more new graphics primitives and compute parameters used to rasterize the new graphics primitives.

[0184] Geometry processing unit 516 may be able to add or delete elements in a geometry stream. Geometry processing unit 516 outputs parameters and vertices specifying new graphics primitives to primitive assembler 518. Primitive assembler 518 receives the parameters and vertices from geometry processing unit 516 and constructs graphics primitives for processing by viewport scale, cull, and clip unit 520. Geometry processing unit 516 reads data stored in parallel processor memory or system memory for use in processing geometric data. Viewport scale, cull, and clip unit 520 performs clipping, culling, and viewport scaling and outputs the processed graphics primitives to rasterizer 522.

[0185] The rasterizer 522 can perform depth culling and other depth-based optimizations. The rasterizer 522 also performs scan conversion on new graphics primitives to generate fragments and outputs those fragments and associated coverage data to the fragment / pixel processing unit 524. The fragment / pixel processing unit 524 is a programmable execution unit configured to execute fragment shader programs or pixel shader programs. The fragment / pixel processing unit 524 transforms the fragments or pixels received from the rasterizer 522 as specified by the fragment or pixel shader program. For example, the fragment / pixel processing unit 524 can be programmed to perform operations including but not limited to texture mapping, shading, blending, texture correction, and perspective correction to produce shaded fragments or pixels that are output to the raster operations unit 526. The fragment / pixel processing unit 524 can read data stored in the parallel processor memory or system memory for use in processing fragment data. The fragment or pixel shader program can be configured to shade at a sample, pixel, tile, or other granularity depending on the sampling rate configured for the processing unit.

[0186] The raster operations unit 526 is a processing unit that performs raster operations and outputs pixel data as processed graphics data to be stored in a graphics memory (e.g., the parallel processor memory 222 as in Figure 2A and / or the system memory 104 as in Figure 1 ), to be displayed on one or more display devices 110A - 110B or for further processing by one of the one or more processors 102 or the parallel processor 112. These raster operations include but are not limited to stencil printing, z-testing, blending, etc. The raster operations unit 526 can be configured to compress z data or color data written to the memory and decompress z data or color data read from the memory. Machine Learning Overview

[0187] The architecture described above can be applied to perform training and inference operations using machine learning models. Machine learning has been successful in solving many kinds of tasks. The computations that occur when training and using machine learning algorithms (e.g., neural networks) are naturally well-suited for efficient parallel implementations. Accordingly, parallel processors such as general-purpose graphics processing units (GPGPUs) have played an important role in the practical implementation of deep neural networks. Parallel graphics processors with a single instruction multiple thread (SIMT) architecture are designed to maximize the amount of parallel processing in the graphics pipeline. In the SIMT architecture, groups of parallel threads attempt to execute program instructions synchronously together as frequently as possible to increase processing efficiency. The efficiency provided by parallel machine learning algorithm implementations allows the use of high-capacity networks and enables the training of those networks on larger datasets.

[0188] A machine learning algorithm is an algorithm that can learn based on a data set. For example, a machine learning algorithm can be designed to model high-level abstractions within a data set. For example, an image recognition algorithm can be used to determine which of several classes a given input belongs to; given an input, a regression algorithm can output a numerical value; and a pattern recognition algorithm can be used to generate transformed text or perform text-to-speech and / or speech recognition.

[0189] An exemplary type of machine learning algorithm is a neural network. There are many types of neural networks; a simple type of neural network is a feedforward network. A feedforward network can be implemented as an acyclic graph in which nodes are arranged in layers. Typically, a feedforward network topology includes an input layer and an output layer separated by at least one hidden layer. The hidden layer transforms the input received by the input layer into a representation useful for generating an output in the output layer. Network nodes are fully connected via edges to nodes in adjacent layers, but there are no edges between nodes within each layer. Data received at the nodes of the input layer of a feedforward network is propagated (i.e., "fed forward") to the nodes of the output layer via an activation function that calculates the state of the nodes in each successive layer of the network based on coefficients ("weights") associated with each of the edges connecting these layers. Depending on the particular model being represented by the algorithm being executed, the output from a neural network algorithm can take various forms.

[0190] Before a machine learning algorithm can be used to model a particular problem, the algorithm is trained using a training data set. Training a neural network involves: selecting a network topology; using a set of training data that represents the problem being modeled by the network; and adjusting the weights until the network model performs with minimal error for all instances of the training data set. For example, during a supervised learning training process for a neural network, the output produced by the network in response to an input representing an instance in the training data set is compared to the "correct" labeled output for that instance, an error signal representing the difference between the output and the labeled output is calculated, and as the error signal is propagated backward through the layers of the network, the weights associated with the connections are adjusted to minimize that error. When the error for each output generated from an instance of the training data set is minimized, the network is considered to be "trained".

[0191] The accuracy of machine learning algorithms is significantly affected by the quality of the datasets used to train the algorithms. The training process can be computationally intensive and may take a large amount of time on conventional general-purpose processors. Accordingly, parallel processing hardware is used to train many types of machine learning algorithms. This is particularly useful for optimizing the training of neural networks, as the computations performed when adjusting the coefficients in a neural network are naturally suited to parallel implementation. Specifically, many machine learning algorithms and software applications have been adapted to utilize the parallel processing hardware within general-purpose graphics processing devices.

[0192] Figure 6 is a generalized diagram of a machine learning software stack 600. A machine learning application 602 is any logic that can be configured to train a neural network using a training dataset or to implement machine intelligence using a trained deep neural network. The machine learning application 602 can include training and inference functions for neural networks and / or specialized software that can be used to train neural networks prior to deployment. The machine learning application 602 can implement any type of machine intelligence, including but not limited to: image recognition, map creation and localization, autonomous navigation, speech synthesis, medical imaging, or language translation. Example machine learning applications 602 include but are not limited to voice-based virtual assistants, image or facial recognition algorithms, autonomous navigation, and software tools used to train the machine learning models used by the machine learning application 602.

[0193] Hardware acceleration for the machine learning application 602 can be enabled via a machine learning framework 604. The machine learning framework 604 can provide a library of machine learning primitives. Machine learning primitives are the basic operations typically performed by machine learning algorithms. Without the machine learning framework 604, developers of machine learning algorithms would need to create and optimize the main computational logic associated with the machine learning algorithm and then re-optimize the computational logic when developing new parallel processors. Instead, the machine learning application can be configured to perform the necessary computations using the primitives provided by the machine learning framework 604. Exemplary primitives include tensor convolution, activation functions, and pooling, which are computational operations performed when training a convolutional neural network (CNN). The machine learning framework 604 can also provide primitives to implement the basic linear algebra subroutines performed by many machine learning algorithms, such as matrix and vector operations. Examples of the machine learning framework 604 include but are not limited to TensorFlow, TensorRT, PyTorch, MXNet, Caffee, and other advanced machine learning frameworks.

[0194] The machine learning framework 604 can process the input data received from the machine learning application 602 and generate an appropriate input to the computing framework 606. The computing framework 606 can abstract the underlying instructions provided to the GPGPU driver 608 so that the machine learning framework 604 can utilize hardware acceleration via the GPGPU hardware 610 without the machine learning framework 604 being very familiar with the architecture of the GPGPU hardware 610. In addition, the computing framework 606 can enable hardware acceleration for the machine learning framework 604 across various types and generations of GPGPU hardware 610. Exemplary computing frameworks 606 include the CUDA computing framework and associated machine learning libraries, such as the CUDA Deep Neural Network (cuDNN) library. The machine learning software stack 600 can also include a communication library or framework to facilitate multi-GPU and multi-node computing. GPGPU Machine Learning Acceleration

[0195] Figure 7 FIG. illustrates a general purpose graphics processing unit 700, which can be Figure 2A the parallel processor 200 or Figure 1 one or more of the parallel processors 112. The general purpose processing unit (GPGPU) 700 can be configured to provide support for hardware acceleration of primitives provided by a machine learning framework to accelerate the processing of computational workloads of the type associated with training deep neural networks. In addition, the GPGPU 700 can be directly linked to other instances of GPGPUs to create a multi-GPU cluster, thereby improving the training speed, especially for deep neural networks. Primitives are also supported to accelerate inference operations for deployed neural networks.

[0196] The GPGPU 700 includes a host interface 702 for enabling connection to a host processor. The host interface 702 can be a PCI Express interface. However, the host interface can also be a vendor-specific communication interface or communication fabric. The GPGPU 700 receives commands from the host processor and uses a global scheduler 704 to distribute the execution threads associated with those commands to a set of processing clusters 706A - 706H. The processing clusters 706A - 706H share a cache memory 708. The cache memory 708 can act as a higher-level cache for the cache memories within the processing clusters 706A - 706H. The illustrated processing clusters 706A - 706H can correspond to the processing clusters 214A - 214N as in Figure 2A

[0197] The GPGPU 700 includes memories 714A-714B coupled to processing clusters 706A-706H via a set of memory controllers 712A-712B. The memories 714A-714B may include various types of memory devices, including dynamic random access memory (DRAM) or graphics random access memory, such as synchronous graphics random access memory (SGRAM), including graphics double data rate (GDDR) memory. The memories 714A-714B may also include 3D stacked memory, including but not limited to high bandwidth memory (HBM).

[0198] Each of the processing clusters 706A-706H may include a set of graphics multiprocessors, such as Figure 2D graphics multiprocessor 234 of Figure 3A graphics multiprocessor 325 of Figure 3B graphics multiprocessor 350 of or may include multi-core groups 365A-365N as in Figure 3C The graphics multiprocessors of the compute clusters include various types of integer and floating-point logic units capable of performing compute operations at a range of precisions including those suitable for machine learning computations. For example, at least a subset of the floating-point units in each compute cluster of the processing clusters 706A-706H may be configured to perform 16-bit or 32-bit floating-point operations, while a different subset of the floating-point units may be configured to perform 64-bit floating-point operations.

[0199] Multiple instances of the GPGPU 700 may be configured to operate as compute clusters. The communication mechanisms used by the compute clusters for synchronization and data exchange vary across embodiments. For example, multiple instances of the GPGPU 700 communicate via a host interface 702. In one embodiment, the GPGPU 700 includes an I / O hub 709 that couples the GPGPU 710 to a GPU link 710 that enables a direct connection to other instances of the GPGPU. The GPU link 710 may be coupled to a dedicated GPU-to-GPU bridge that enables communication and synchronization between multiple instances of the GPGPU 700. Optionally, the GPU link 710 is coupled to a high-speed interconnect to transfer data to and receive data from other GPGPUs or parallel processors. Multiple instances of the GPGPU 700 may be located in separate data processing systems and may communicate via a network device that may be accessed via the host interface 702. In addition to or instead of the host interface 702, the GPU link 710 may be configured to enable a connection to a host processor.

[0200] Although the illustrated configuration of the GPGPU 700 can be configured for training neural networks, alternative configurations of the GPGPU 700 can be configured for deployment within high-performance or low-power inference platforms. In an inference configuration, the GPGPU 700 includes fewer processing clusters among the processing clusters 706A - 706H as compared to a training configuration. Additionally, the memory technology associated with the memories 714A - 714B can differ between the inference configuration and the training configuration. In one embodiment, the inference configuration of the GPGPU 700 can support inference-specific instructions. For example, the inference configuration can provide support for one or more 8-bit integer or floating-point dot product instructions that are typically used during inference operations for deployed neural networks.

[0201] Figure 8 FIG. illustrates a multi-GPU computing system 800. The multi-GPU computing system 800 can include a processor 802 coupled to a plurality of GPGPUs 806A - 806D via a host interface switch 804. The host interface switch 804 can be a PCI Express switch device that couples the processor 802 to a PCI Express bus through which the processor 802 can communicate with the set of GPGPUs 806A - 806D. Each of the plurality of GPGPUs 806A - 806D can be an Figure 7 instance of the GPGPU 700. The GPGPUs 806A - 806D can be interconnected via a set of high-speed point-to-point GPU-to-GPU links 816. The high-speed GPU-to-GPU links can be connected to each of the GPGPUs 806A - 806D via dedicated GPU links such as Figure 7 the GPU link 710 in. The P2P GPU links 816 enable direct communication between each of the GPGPUs 806A - 806D without the need for communication through the host interface bus to which the processor 802 is connected. With GPU-to-GPU traffic directed to the P2P GPU links, the host interface bus remains available for system memory access or for communicating with other instances of the multi-GPU computing system 800, for example, via one or more network devices. Although in Figure 8 the GPGPUs 806A - 806D are connected to the processor 802 via the host interface switch 804, the processor 802 can alternatively include direct support for the P2P GPU links 816 and be directly connected to the GPGPUs 806A - 806D. In one embodiment, the P2P GPU links 816 enable the multi-GPU computing system 800 to operate as a single logical GPU. Machine Learning Neural Network Implementation

[0202] The computing architectures described herein can be configured to perform a class of parallel processing particularly suited for training and deploying neural networks for machine learning. A neural network can be generalized as a network of functions with graphical relationships. As is well known in the art, there are various types of neural network implementations used in machine learning. One exemplary type of neural network is the feedforward network as previously described.

[0203] A second exemplary type of neural network is the convolutional neural network (CNN). A CNN is a specialized feedforward neural network for processing data with a known, grid-like topology, such as image data. Accordingly, CNNs are commonly used in computer vision and image recognition applications, but they can also be used for other types of pattern recognition, such as speech and language processing. The nodes in the input layer of a CNN are organized as a collection of "filters" (feature detectors inspired by receptive fields found in the retina), and the output of each filter collection is propagated to nodes in successive layers of the network. The computations for a CNN include applying convolutional mathematical operations to each filter to produce the output of that filter. Convolution is a specialized kind of mathematical operation performed by two functions to produce a third function, which is a modified version of one of the two original functions. In convolutional network terminology, the first function to the convolution can be referred to as the input, while the second function can be referred to as the convolution kernel. The output can be referred to as the feature map. For example, the input to a convolutional layer can be a multi-dimensional data array defining the various color components of an input image. The convolution kernel can be a multi-dimensional array of parameters, where the parameters are adapted through the training process for the neural network.

[0204] A recurrent neural network (RNN) is a series of feedforward neural networks that includes feedback connections between the layers. An RNN enables modeling of sequential data by sharing parameter data across different parts of the neural network. The architecture for an RNN includes loops. These loops represent the influence of the current value of a variable on its own value at future times, since at least part of the output data from the RNN is used as feedback for processing subsequent inputs in the sequence. This feature makes the RNN particularly useful for language processing due to the variable nature in which language data can be composed.

[0205] The figures described below present exemplary feedforward networks, CNN networks, and RNN networks, and describe the general processes for training and deploying each of those types of networks, respectively. It will be understood that these descriptions are exemplary and non-limiting with respect to any particular embodiments described herein, and generally the concepts illustrated can be generally applied to deep neural networks and machine learning techniques.

[0206] The exemplary neural networks described above can be used to perform deep learning. Deep learning is machine learning using deep neural networks. Different from shallow neural networks that only include a single hidden layer, the deep neural networks used in deep learning are artificial neural networks composed of multiple hidden layers. Deeper neural networks are generally more computationally intensive to train. However, the additional hidden layers of the network enable multi-step pattern recognition, which results in a reduced output error relative to shallow machine learning techniques.

[0207] The deep neural networks used in deep learning typically include a front-end network for performing feature recognition, which is coupled to a back-end network that represents a mathematical model that can perform operations (such as object classification, speech recognition, etc.) based on the feature representation provided to the mathematical model. Deep learning enables machine learning to be performed without the need to perform manual feature engineering for the model. Instead, the deep neural network can learn features based on the statistical structure or correlation within the input data. The learned features can be provided to a mathematical model, which can map the detected features to an output. The mathematical model used by the network is typically dedicated to the specific task to be performed, and different models will be used to perform different tasks.

[0208] Once the neural network is structured, a learning model can be applied to the network to train the network to perform a specific task. The learning model describes how to adjust the weights within the model to reduce the output error of the network. Backpropagation of errors is a commonly used method for training neural networks. An input vector is presented to the network for processing. The output of the network is compared with the desired output using a loss function, and an error value is calculated for each neuron in the output layer. Subsequently, the error values are propagated backward until each neuron has an associated error value that roughly represents the contribution of that neuron to the original output. The network can then learn from those errors using an algorithm (such as the stochastic gradient descent algorithm) to update the weights of the neural network.

[0209] Figures 9A - 9B Illustrates an exemplary convolutional neural network. Figure 9A Illustrates the various layers within the CNN. As Figure 9AAs shown, an exemplary CNN for modeling image processing can receive an input 902 that describes the red, green, and blue (RGB) components of an input image. The input 902 can be processed by a plurality of convolutional layers (e.g., convolutional layer 904, convolutional layer 906). The output from the plurality of convolutional layers can optionally be processed by a set of fully connected layers 908. As previously described for feedforward networks, neurons in a fully connected layer have full connections to all activations in the previous layer. The output from the fully connected layer 908 can be used to generate an output result from the network. Matrix multiplication rather than convolution can be used to compute the activations within the fully connected layer 908. Not all CNN implementations utilize the fully connected layer 908. For example, in some implementations, the convolutional layer 906 can generate the output of the CNN.

[0210] Convolutional layers are sparsely connected, which is different from the traditional neural network configuration found in the fully connected layer 908. Traditional neural network layers are fully connected such that each output unit interacts with each input unit. However, as illustrated, convolutional layers are sparsely connected because the output of the convolution of the receptive field (rather than the corresponding state values of each node in the receptive field) is input to the nodes of the subsequent layer. The kernel associated with the convolutional layer performs a convolution operation, and the output of this convolution operation is sent to the next layer. The dimensionality reduction performed within the convolutional layer is one aspect that enables CNNs to scale to handle large images.

[0211] Figure 9B Illustrates an exemplary computational stage within the convolutional layer of a CNN. The input 912 to the convolutional layer of the CNN can be processed in three stages of the convolutional layer 914. These three stages can include a convolution stage 916, a detector stage 918, and a pooling stage 920. The convolutional layer 914 can then output the data to a successive convolutional layer. The final convolutional layer of the network can generate output feature map data or provide an input to a fully connected layer, e.g., to generate classification values for the input to the CNN.

[0212] A number of convolutions are performed in parallel in the convolution stage 916 to produce a set of linear activations. The convolution stage 916 can include an affine transformation, which is any transformation that can be specified as a linear transformation plus a translation. Affine transformations include rotation, translation, scaling, and combinations of these transformations. The convolution stage computes the output of a function (e.g., a neuron) connected to a specific region in the input, which can be determined as the local region associated with the neuron. The neuron computes the dot product between the weights of the neuron and the weights of the region in the local input to which the neuron is connected. The output from the convolution stage 916 defines the set of linear activations processed by successive stages of the convolutional layer 914.

[0213] The linear activations can be processed by the detector stage 918. In the detector stage 918, each linear activation is processed by a non-linear activation function. The non-linear activation function increases the non-linearity of the overall network without affecting the receptive field of the convolutional layer. Several types of non-linear activation functions can be used. One particular type is the rectified linear unit (ReLU), which uses an activation function defined as f(x) = max(0, x), such that the threshold for activation is zero.

[0214] The pooling stage 920 uses a pooling function that replaces the output of the convolutional layer 906 with summary statistics of nearby outputs. The pooling function can be used to introduce translational invariance into the neural network, such that small translations to the input do not change the pooled output. Local translational invariance can be useful in scenarios where the presence of features in the input data is more important than the exact location of the features. Various types of pooling functions can be used during the pooling stage 920, including max pooling, average pooling, and l2-norm pooling. Additionally, some CNN implementations do not include a pooling stage. Instead, such implementations are alternative and additional convolutional stages with an increased stride relative to the previous convolutional stage.

[0215] The output from the convolutional layer 914 can then be processed by the next layer 922. The next layer 922 can be an additional convolutional layer or one of the layers in the fully-connected layer 908. For example, Figure 9A the first convolutional layer 904 can output to the second convolutional layer 906, and the second convolutional layer can output to the first layer in the fully-connected layer 908.

[0216] Figure 10Illustrated is an exemplary recurrent neural network 1000. In a recurrent neural network (RNN), the previous state of the network affects the output of the current state of the network. RNNs can be built in various ways and using various functions. The use of RNNs generally revolves around using a mathematical model to predict the future based on a previous sequence of inputs. For example, an RNN can be used to perform statistical language modeling to predict the upcoming word given a previous sequence of words. The illustrated RNN 1000 can be described as having an input layer 1002 that receives an input vector, a hidden layer 1004 for implementing a recurrent function, a feedback mechanism 1005 for enabling the'memory' of the previous state, and an output layer 1006 for outputting a result. The RNN 1000 operates based on time steps. The state of the RNN at a given time step is affected via the feedback mechanism 1005 based on the previous time step. For a given time step, the state of the hidden layer 1004 is defined by the previous state and the input at the current time step. The initial input (x1) at the first time step can be processed by the hidden layer 1004. The second input (x2) can be processed by the hidden layer 1004 using the state information determined during the processing of the initial input (x1). A given state can be calculated as s t = f(Ux t + Ws t-1 ), where U and W are parameter matrices. The function f is generally non-linear, such as the hyperbolic tangent function (Tanh) or a variant of the rectifier function f(x) = max(0, x). However, the specific mathematical function used in the hidden layer 1004 can vary depending on the specific implementation details of the RNN 1000.

[0217] In addition to the basic CNN and RNN networks described, acceleration can also be enabled for variants of those networks. An example RNN variant is the long short term memory (LSTM) RNN. The LSTM RNN is capable of learning long-term dependencies, which may be necessary for processing longer language sequences. A variant of the CNN is the convolutional deep belief network, which has a structure similar to the CNN and is trained in a manner similar to the deep belief network. A deep belief network (DBN) is a generative neural network consisting of multiple layers of stochastic (random) variables. The DBN can be trained layer by layer using greedy unsupervised learning. The learned weights of the DBN can then be used to provide a pre-trained neural network by determining an optimal initial set of weights for the neural network. In a further embodiment, acceleration for reinforcement learning is enabled. In reinforcement learning, an artificial agent learns by interacting with its environment. The agent is configured to optimize certain objectives to maximize the cumulative reward.

[0218] Figure 11 Illustrated is the training and deployment of a deep neural network. Once a given network has been structured for a task, a training data set 1102 is used to train the neural network. A variety of training frameworks 1104 have been developed to enable hardware acceleration of the training process. For example, Figure 6 the machine learning framework 604 can be configured as the training framework 1104. The training framework 1104 can access the untrained neural network 1106 and enable the untrained neural network to be trained using the parallel processing resources described herein to generate a trained neural network 1108.

[0219] To initiate the training process, weights can be randomly selected or initial weights can be selected by using a deep belief network for pre-training. Subsequently, a training loop is performed in a supervised or unsupervised manner.

[0220] Supervised learning is a learning method in which training is performed as a mediated operation, such as when the training data set 1102 includes the input paired with the desired output of the input, or when the training data set includes inputs with known outputs and the outputs of the neural network are manually graded. The network processes the input and compares the resulting output with the expected output or set of desired outputs. Subsequently, the error is propagated back through the system. The training framework 1104 can be adjusted to adjust the weights controlling the untrained neural network 1106. The training framework 1104 can provide tools to monitor how well the untrained neural network 1106 is converging to a model suitable for generating correct answers based on the known input data. As the weights of the network are adjusted to refine the output generated by the neural network, the training process occurs repeatedly. The training process can continue until the neural network reaches a statistically desired accuracy associated with the trained neural network 1108. The trained neural network 1108 can then be deployed to perform any number of machine learning operations to generate inference results 1114 based on the input of new data 1112.

[0221] Unsupervised learning is a learning method in which the network attempts to train itself using unlabeled data. Thus, for unsupervised learning, the training data set 1102 will include input data that does not have any associated output data. The untrained neural network 1106 can learn groupings within the unlabeled input and can determine how individual inputs relate to the overall data set. Unsupervised training can be used to generate a self-organizing map, which is a type of trained neural network 1108 that is capable of performing operations useful in reducing the dimensionality of data. Unsupervised training can also be used to perform anomaly detection, which allows identification of data points in the input data set that deviate from the normal pattern of the data.

[0222] Variants of supervised and unsupervised training can also be employed. Semi-supervised learning is a technique in which a mixture of labeled and unlabeled data having the same distribution is included in the training data set 1102. Progressive learning is a variant of supervised learning in which input data is continuously used to further train the model. Progressive learning enables the trained neural network 1108 to adapt to new data 1112 without forgetting the knowledge embedded in the network during initial training.

[0223] Whether supervised or unsupervised, the training process for a particularly deep neural network can be computationally too intensive for a single computing node. A distributed network of computing nodes can be used instead of a single computing node to accelerate the training process.

[0224] Figure 12A is a block diagram illustrating distributed learning. Distributed learning is a training model that uses multiple distributed computing nodes to perform supervised or unsupervised training of a neural network. Each of the distributed computing nodes can include one or more host processors or one or more general-purpose processing nodes in a general-purpose processing node such as Figure 7 the highly parallel general-purpose graphics processing unit 700 in. As illustrated, distributed learning can be performed with model parallelism 1202, data parallelism 1204, or a combination of model and data parallelism 1206.

[0225] In model parallelism 1202, different computing nodes in a distributed system can perform training computations for different parts of a single network. For example, each layer of a neural network can be trained by a different processing node of the distributed system. Benefits of model parallelism include the ability to scale to particularly large models. Partitioning the computations associated with different layers of a neural network enables the training of very large neural networks in which all layer weights would not fit into the memory of a single node. In some instances, model parallelism can be particularly useful when performing unsupervised training of large neural networks.

[0226] In data parallelism 1204, different nodes of a distributed network have a complete instance of the model, and each node receives a different portion of the data. Results from different nodes are then combined. While different ways of implementing data parallelism are possible, all data parallelism training approaches require techniques for combining the results and synchronizing the model parameters across each node. Exemplary ways of combining data include parameter averaging and update-based data parallelism. Parameter averaging trains each node on a subset of the training data and sets the global parameters (e.g., weights, biases) to the average of the parameters from each node. Parameter averaging uses a central parameter server that maintains the parameter data. Update-based data parallelism is similar to parameter averaging, except that updates to the model are transmitted instead of the parameters from the nodes to the parameter server. Additionally, update-based data parallelism can be performed in a decentralized manner where the updates are compressed and transmitted between nodes.

[0227] Combined model and data parallelism 1206 can be implemented, for example, in a distributed system where each computing node includes multiple GPUs. Each node can have a complete instance of the model, where individual GPUs within each node are used to train different parts of the model.

[0228] Distributed training has increased overhead relative to training on a single machine. However, the parallel processors and GPGPUs described herein can each implement various techniques to reduce the overhead of distributed training, including techniques for enabling high-bandwidth GPU-to-GPU data transfer and accelerated remote data synchronization.

[0229] Figure 12B is a block diagram illustrating a programmable network interface 1210 and a data processing unit. The programmable network interface 1210 is a programmable network engine that can be used to accelerate network-based computing tasks within a distributed environment. The programmable network interface 1210 can be coupled to a host system via a host interface 1270. The programmable network interface 1210 can be used to accelerate network or storage operations for the CPU or GPU of the host system. The host system can be, for example, a node of a distributed learning system for performing distributed training, such as, as Figure 12A shown. The host system can also be a data center node within a data center.

[0230] In one embodiment, access to a remote storage device containing model data can be accelerated by the programmable network interface 1210. For example, the programmable network interface 1210 can be configured to present the remote storage device as a local storage device of the host system. The programmable network interface 1210 can also accelerate remote direct memory access (RDMA) operations performed between the GPU of the host system and the GPU of the remote system. In one embodiment, the programmable network interface 1210 can enable storage functions such as, but not limited to, NVME-oF. The programmable network interface 1210 can also accelerate encryption, data integrity, compression, and other operations for the remote storage device on behalf of the host system, thereby allowing the remote storage device to approximate the latency of a storage device directly attached to the host system.

[0231] The programmable network interface 1210 can also perform resource allocation and management on behalf of the host system. Storage security operations can be migrated to the programmable network interface 1210 and executed in coordination with the allocation and management of remote storage resources. Network-based operations for managing access to the remote storage device that would otherwise be performed by the processor of the host system can instead be performed by the programmable network interface 1210.

[0232] In one embodiment, network and / or data security operations can be migrated from the host system to the programmable network interface 1210. The data center security policy for the data center node can be disposed of by the programmable network interface 1210 rather than the processor of the host system. For example, the programmable network interface 1210 can detect and mitigate attempted network-based attacks (e.g., DDoS) on the host system, thereby preventing the attack from compromising the availability of the host system.

[0233] The programmable network interface 1210 may include a system-on-chip (SoC 1220) that executes an operating system via a plurality of processor cores 1222. The processor cores 1222 may include general-purpose processor (e.g., CPU) cores. In one embodiment, the processor cores 1222 may also include one or more GPU cores. The SoC 1220 may execute instructions stored in the memory device 1240. The storage device 1250 may store local operating system data. The storage device 1250 and the memory device 1240 may also be used to cache remote data for the host system. The network ports 1260A - 1260B enable connection to a network or fabric, facilitate network access for the SoC 1220, and facilitate network access for the host system via the host interface 1270. The programmable network interface 1210 may also include an I / O interface 1275, such as a USB interface. The I / O interface 1275 may be used to couple external devices to the programmable network interface 1210 or to couple as a debug interface. The programmable network interface 1210 further includes a management interface 1230 that enables software on the host device to manage and configure the programmable network interface 1210 and / or the SoC 1220. In one embodiment, the programmable network interface 1210 may also include one or more accelerators or GPUs 1245 to accept the migration of parallel computing tasks from the SoC 1220, the host system, or a remote system coupled via the network ports 1260A - 1260B. Exemplary Machine Learning Applications

[0234] Machine learning can be applied to solve various technical problems, including but not limited to computer vision, autonomous driving and navigation, speech recognition, and language processing. Computer vision has traditionally been one of the most active research areas for machine learning applications. The scope of computer vision applications ranges from reproducing human visual capabilities (such as recognizing faces) to creating new classes of visual capabilities. For example, a computer vision application can be configured to identify sound waves from vibrations induced in objects visible in a video. Machine learning accelerated by parallel processors enables computer vision applications to train using training data sets that are significantly larger than previously feasible, and enables inference systems to be deployed using low-power parallel processors.

[0235] Machine learning accelerated by parallel processors has applications in autonomous driving, including lane and road sign recognition, obstacle avoidance, navigation, and driving control. The accelerated machine learning techniques can be used to train a driving model based on a data set that defines the appropriate response to specific training inputs. The parallel processors described herein are capable of enabling the rapid training of increasingly complex neural networks for autonomous driving solutions and of enabling the deployment of low-power inference processors in mobile platforms suitable for integration into autonomous vehicles.

[0236] Deep neural networks accelerated by parallel processors have enabled machine learning approaches for automatic speech recognition (ASR). ASR involves creating a function that computes the most likely language sequence given an input acoustic sequence. Accelerated machine learning using deep neural networks has enabled an alternative to the hidden Markov model (HMM) and Gaussian mixture model (GMM) previously used for ASR.

[0237] Parallel processor-accelerated machine learning can also be used to accelerate natural language processing. The automated learning process can utilize statistical inference algorithms to produce models that are robust to noisy or unfamiliar inputs. Exemplary natural language processor applications include automatic machine translation between human languages.

[0238] The parallel processing platforms for machine learning can be divided into a training platform and a deployment platform. The training platform is generally highly parallel and includes optimizations for accelerating multi-GPU single-node training and multi-node multi-GPU training. Exemplary parallel processors suitable for training include Figure 7 the general-purpose graphics processing unit 700 and Figure 8 the multi-GPU computing system 800. In contrast, deployed machine learning platforms generally include lower-power parallel processors suitable for use in products such as cameras, autonomous robots, and autonomous vehicles.

[0239] In addition, machine learning techniques can also be applied to accelerate or enhance graphics processing activities. For example, a machine learning model can be trained to recognize the output generated by a GPU-accelerated application and generate an enlarged version of the output. Such techniques can be applied to accelerate the generation of high-resolution images for game applications. Various other graphics pipeline activities can benefit from the use of machine learning. For example, a machine learning model can be trained to perform tessellation operations on geometric data to increase the complexity of a geometric model, thereby allowing the automatic generation of a more detailed geometry from a geometry with relatively low detail.

[0240] Figure 13FIG. illustrates an exemplary inference system-on-a-chip (SOC) 1300 suitable for performing inference using a trained model. The SOC 1300 may integrate processing components, including a media processor 1302, a vision processor 1304, a GPGPU 1306, and a multi-core processor 1308. The GPGPU 1306 may be the GPGPU described herein, such as the GPGPU 700, and the multi-core processor 1308 may be the multi-core processor described herein, such as the multi-core processors 405-406. The SOC 1300 may additionally include on-chip memory 1305, which may enable a shared on-chip data pool accessible by each of the processing components. The processing components may be optimized for low-power operation to enable deployment on various machine learning platforms, including autonomous vehicles and autonomous robots. For example, one implementation of the SOC 1300 may be used as part of a main control system for an autonomous vehicle. In the case where the SOC 1300 is configured to be used in an autonomous vehicle, the SOC is designed and configured to comply with relevant functional safety standards for deployment jurisdiction.

[0241] During operation, the media processor 1302 and the vision processor 1304 may work together to accelerate computer vision operations. The media processor 1302 may enable low-latency decoding of multiple high-resolution (e.g., 4K, 8K) video streams. The decoded video streams may be written to a buffer in the on-chip memory 1305. The vision processor 1304 may then parse the decoded video and perform preliminary processing operations on the frames of the decoded video to prepare the frames for processing using a trained image recognition model. For example, the vision processor 1304 may accelerate the convolutional operations for a CNN that performs image recognition on high-resolution video data, while the backend model computations are performed by the GPGPU 1306.

[0242] The multi-core processor 1308 may include control logic for assisting in the ordering and synchronization of data transfers and shared memory operations performed by the media processor 1302 and the vision processor 1304. The multi-core processor 1308 may also act as an application processor for executing software applications that can utilize the inference computing capabilities of the GPGPU 1306. For example, at least part of the navigation and driving logic may be implemented in software executed on the multi-core processor 1308. Such software may directly issue computing workloads to the GPGPU 1306, or the computing workloads may be issued to the multi-core processor 1308, which may migrate at least part of those operations to the GPGPU 1306.

[0243] The GPGPU 1306 may include computing clusters, such as the processing clusters 706A - 706H in a low - power configuration within the general - purpose graphics processing unit 700. The computing clusters within the GPGPU 1306 may support instructions that are specifically optimized to perform inference computations on trained neural networks. For example, the GPGPU 1306 may support instructions for performing low - precision computations such as 8 - bit and 4 - bit integer vector operations. Additional System Overview

[0244] Figure 14 is a block diagram of the processing system 1400. Figure 14 Elements with the same or similar names as elements in any other figure herein describe elements that are the same as those in other figures, can operate or function in a similar manner as those in other figures, may include the same components, and may be linked to other entities, such as those described elsewhere herein, but are not limited thereto. The system 1400 can be used in: a single - processor desktop computer system, a multi - processor workstation system, or a server system with a large number of processors 1402 or processor cores 1407. The system 1400 can be a processing platform incorporated within a system - on - a - chip (SoC) integrated circuit for use in mobile devices, handheld devices, or embedded devices, such as for use within an Internet - of - things (IoT) device with wired or wireless connectivity to a local area network or a wide area network.

[0245] The system 1400 can be a processing system having components corresponding to Figure 1 those components. For example, in different configurations, the (one or more) processors 1402 or (one or more) processor cores 1407 may correspond to Figure 1 the (one or more) processors 102. The (one or more) graphics processors 1408 may correspond to Figure 1 the (one or more) parallel processors 112. The external graphics processor 1418 can be Figure 1 one of the (one or more) plug - in devices 120.

[0246] System 1400 may include, may be coupled with, or may be integrated within: a server-based gaming platform; a game console, including a game and media console; a mobile game console, a handheld game console, or an online game console. System 1400 may be part of a mobile phone, a smartphone, a tablet computing device, or a mobile Internet-connected device (such as a laptop computer with low internal storage capacity). Processing system 1400 may also include, be coupled with, or be integrated within: a wearable device, such as a smartwatch wearable device; smart glasses or clothing that are enhanced with augmented reality (AR) or virtual reality (VR) features to provide visual, audio, or tactile output to supplement real-world visual, audio, or tactile experiences or otherwise provide text, audio, graphics, video, holographic images or video, or tactile feedback; other augmented reality (AR) devices; or other virtual reality (VR) devices. Processing system 1400 may include a television or a set-top box device, or may be part of a television or a set-top box device. System 1400 may include an autonomous vehicle, be coupled with an autonomous vehicle, or be integrated within an autonomous vehicle, such as a bus, a tractor-trailer, a car, a motorcycle or electric cycle, an airplane, or a glider (or any combination thereof). The autonomous vehicle may use system 1400 to process the environment sensed around the vehicle.

[0247] One or more processors 1402 may include one or more processor cores 1407 that are configured to process instructions that, when executed, perform operations for system and user software. At least one of the one or more processor cores 1407 may be configured to process a particular instruction set 1409. The instruction set 1409 may facilitate complex instruction set computing (CISC), reduced instruction set computing (RISC), or computing via very long instruction word (VLIW). The one or more processor cores 1407 may process different instruction sets 1409, and the different instruction sets 1409 may include instructions for facilitating the emulation of other instruction sets. The processor cores 1407 may also include other processing devices, such as a digital signal processor (DSP).

[0248] The processor 1402 may include a cache memory 1404. Depending on the architecture, the processor 1402 may have a single internal cache or multiple levels of internal caches. In some embodiments, the cache memory is shared among various components of the processor 1402. In some embodiments, the processor 1402 also uses an external cache (e.g., a third-level (L3) cache or a last-level cache (LLC)) (not shown), and the external cache can be shared among the processor cores 1407 using known cache coherence techniques. A register file 1406 may additionally be included in the processor 1402 and may include different types of registers for storing different types of data (e.g., integer registers, floating-point registers, status registers, and instruction pointer registers). Some registers may be general-purpose registers, while other registers may be dedicated to the design of the processor 1402.

[0249] One or more processors 1402 may be coupled to one or more interface buses 1410 to transfer communication signals, such as address, data, or control signals, between the processor 1402 and other components in the system 1400. In one of these embodiments, the interface bus 1410 may be a processor bus, such as a certain version of the Direct Media Interface (DMI) bus. However, the processor bus is not limited to the DMI bus and may include one or more peripheral component interconnect buses (e.g., PCI, PCI Express), a memory bus, or other types of interface buses. For example, the (one or more) processors 1402 may include an integrated memory controller 1416 and a platform controller hub 1430. The memory controller 1416 facilitates communication between the memory device and other components of the system 1400, while the platform controller hub (PCH) 1430 provides connections to I / O devices via a local I / O bus.

[0250] The memory device 1420 can be a dynamic random access memory (DRAM) device, a static random-access memory (SRAM) device, a flash memory device, a phase change memory device, or some other memory device with suitable performance to act as a process memory. The memory device 1420 can operate, for example, as a system memory for the system 1400 to store data 1422 and instructions 1421 for use when one or more processors 1402 execute an application or process. The memory controller 1416 is also coupled to an optional external graphics processor 1418, which can communicate with one or more graphics processors 1408 in the processor 1402 to perform graphics operations and media operations. In some embodiments, the graphics operations, media operations, and / or computing operations can be assisted by an accelerator 1412, which is a coprocessor that can be configured to perform a set of specialized graphics operations, media operations, or computing operations. For example, the accelerator 1412 can be a matrix multiplication accelerator for optimizing machine learning or computing operations. The accelerator 1412 can be a ray tracing accelerator, which can be used to perform ray tracing operations in cooperation with the graphics processor 1408. In one embodiment, an external accelerator 1419 can be used instead of the accelerator 1412, or can be used in cooperation with the accelerator 1412.

[0251] A display device 1411 can be provided, which can be connected to the processor(s) 1402. The display device 1411 can be one or more of the following: an internal display device, such as in a mobile electronic device or a laptop computer device; or an external display device attached via a display interface (e.g., DisplayPort, etc.). The display device 1411 can be a head mounted display (HMD), such as a stereoscopic display device for use in virtual reality (VR) applications or augmented reality (AR) applications.

[0252] The platform controller hub 1430 can enable peripheral devices to be connected to the memory device 1420 and the processor 1402 via a high-speed I / O bus. The I / O peripheral devices include, but are not limited to, an audio controller 1446, a network controller 1434, a firmware interface 1428, a wireless transceiver 1426, a touch sensor 1425, and a data storage device 1424 (e.g., non-volatile memory, volatile memory, hard disk drive, flash memory, NAND, 3D NAND, 3D Xpoint / Optane, etc.). The data storage device 1424 can be connected via a storage interface (e.g., SATA) or via a peripheral bus (such as a peripheral component interconnect bus (e.g., PCI, PCI Express)). The touch sensor 1425 can include a touch screen sensor, a pressure sensor, or a fingerprint sensor. The wireless transceiver 1426 can be a Wi-Fi transceiver, a Bluetooth transceiver, or a mobile network transceiver, such as a 3G, 4G, 5G, or Long-Term Evolution (LTE) transceiver. The firmware interface 1428 enables communication with the system firmware and can be, for example, a unified extensible firmware interface (UEFI). The network controller 1434 can enable a network connection to a wired network. In some embodiments, a high-performance network controller (not shown) is coupled to the interface bus 1410. The audio controller 1446 can be a multi-channel high-definition audio controller. In some of these embodiments, the system 1400 includes an optional legacy I / O controller 1440 for coupling legacy (e.g., Personal System 2 (PS / 2)) devices to the system. The platform controller hub 1430 can also be connected to one or more Universal Serial Bus (USB) controllers 1442 to connect to input devices, such as a keyboard and mouse 1443 combination, a camera 1444, or other USB input devices.

[0253] It will be appreciated that the system 1400 shown is exemplary and not restrictive, as other types of data processing systems configured in different ways may also be used. For example, instances of the memory controller 1416 and the platform controller hub 1430 may be integrated into a discrete external graphics processor, such as the external graphics processor 1418. The platform controller hub 1430 and / or the memory controller 1416 may be external to one or more processors 1402. For example, the system 1400 may include an external memory controller 1416 and a platform controller hub 1430, which may be configured as a memory controller hub and a peripheral controller hub within a system chipset that communicates with the (one or more) processors 1402.

[0254] For example, a circuit board ("sled") may be used, on which components such as CPUs, memories, and other components are placed, and on which the components (such as CPUs, memories, and other components) are designed to achieve improved thermal performance. Processing components such as processors may be located on the top side of the sled, while nearby memories such as DIMMs are located on the bottom side of the sled. As a result of the enhanced airflow provided by this design, the components can operate at higher frequencies and power levels than in a typical system, thereby improving performance. Additionally, the sled is configured for blind mating of power and data communication cables in a rack, thereby enhancing their ability to be quickly removed, upgraded, reinstalled, and / or replaced. Similarly, the individual components located on the sled (such as processors, accelerators, memories, and data storage drives) are configured to be easily upgraded due to their increased spacing from each other. In an illustrative embodiment, the components additionally include hardware authentication features for proving their authenticity.

[0255] A data center may utilize a single network architecture ("fabric") that supports multiple other network architectures, including Ethernet and Omnipath. The sled may be coupled to a switch via optical fiber, which provides higher bandwidth and lower latency than typical twisted pair cabling (e.g., Category 5, Category 5e, Category 6, etc.). Due to the high-bandwidth, low-latency interconnect and network architecture, the data center in use may centralize physically dispersed resources such as memories, accelerators (e.g., GPUs, graphics accelerators, FPGAs, ASICs, neural network, and / or artificial intelligence accelerators, etc.), and data storage drives, and provide them to computing resources (e.g., processors) as needed, enabling the computing resources to access the centralized resources as if the centralized resources were local.

[0256] A power supply or power source can supply voltage and / or current to system 1400 or any component or system described herein. In one example, the power supply includes an AC to DC (alternating current to direct current) adapter for insertion into a wall outlet. Such AC power can be a renewable energy (e.g., solar) power source. In one example, the power source includes a DC power source, such as an external AC to DC converter. The power source or power supply can also include wireless charging hardware for charging via a proximity charging field. The power source can include an internal battery, an AC supply, an action-based power supply, a solar power supply, or a fuel cell source.

[0257] Figures 15A - 15C Illustrative computing systems and graphics processors. Figures 15A - 15C Elements having the same or similar names as elements in any other figure herein describe the same elements as in other figures, can operate or function in a similar manner as in other figures, may include the same components, and may be linked to other entities, such as those described elsewhere herein, but are not limited thereto.

[0258] Figure 15Ais a block diagram of a processor 1500, which can be a variant of one of the processors in processor 1402 and can be used in place of one of those processors. Thus, the disclosure of any feature in connection with processor 1500 herein also discloses the corresponding combination with (one or more) processors 1402, but is not limited thereto. Processor 1500 may have one or more processor cores 1502A - 1502N, an integrated memory controller 1514, and an integrated graphics processor 1508. In the case where the integrated graphics processor 1508 is excluded, a system including the processor will include a graphics processor device within the system chipset or coupled via a system bus. Processor 1500 may include additional cores, and the additional cores are at most the additional core 1502N represented by the dashed box and include the additional core 1502N represented by the dashed box. Each of the processor cores 1502A - 1502N includes one or more internal cache units 1504A - 1504N. In some embodiments, each of the processor cores 1502A - 1502N also has access to one or more shared cache units 1506. The internal cache units 1504A - 1504N and the shared cache units 1506 represent the cache memory hierarchy within processor 1500. The cache memory hierarchy may include at least one level of instruction and data cache within each processor core and one or more levels of shared mid - level cache, such as a second - level (L2), third - level (L3), fourth - level (L4), or other levels of cache, where the highest - level cache before external memory is classified as LLC. In some embodiments, cache coherence logic maintains coherence between the cache units 1506 and 1504A - 1504N.

[0259] Processor 1500 may also include a set 1516 of one or more bus controller units and a system agent core 1510. The one or more bus controller units 1516 manage a set of peripheral buses, such as one or more PCI buses or PCI Express buses. The system agent core 1510 provides management functions for the various processor components. The system agent core 1510 may include one or more integrated memory controllers 1514 for managing access to various external memory devices (not shown).

[0260] For example, one or more of the processor cores 1502A - 1502N may include support for simultaneous multithreading operations. The system agent core 1510 includes components for coordinating and operating the cores 1502A - 1502N during multithreaded processing. The system agent core 1510 may additionally include a power control unit (PCU) that includes logic and components for regulating the power states of the processor cores 1502A - 1502N and the graphics processor 1508.

[0261] The processor 1500 may additionally include a graphics processor 1508 for performing graphics processing operations. In some of these embodiments, the graphics processor 1508 is coupled to a set of shared cache units 1506 and the system agent core 1510, which includes one or more integrated memory controllers 1514. The system agent core 1510 may also include a display controller 1511 for driving the graphics processor output to one or more coupled displays. The display controller 1511 may also be a separate module coupled to the graphics processor via at least one interconnect, or may be integrated within the graphics processor 1508.

[0262] A ring - based interconnect 1512 may be used to couple the internal components of the processor 1500. However, alternative interconnect units may be used, such as point - to - point interconnects, switched interconnects, or other techniques, including those well - known in the art. In some of these embodiments with a ring - based interconnect 1512, the graphics processor 1508 is coupled to the ring - based interconnect 1512 via an I / O link 1513.

[0263] The exemplary I / O link 1513 represents at least one of a variety of I / O interconnects, including an on - package I / O interconnect that facilitates communication between the various processor components and a high - performance memory module 1518, such as an eDRAM module or a high - bandwidth memory (HMB) module. Optionally, each of the processor cores 1502A - 1502N and the graphics processor 1508 may use the high - performance memory module 1518 as a shared last - level cache.

[0264] The processor cores 1502A - 1502N can be, for example, homogeneous cores that execute the same instruction set architecture. Alternatively, the processor cores 1502A - 1502N are heterogeneous in terms of instruction set architecture (ISA), where one or more of the processor cores 1502A - 1502N execute a first instruction set, while at least one of the other cores executes a subset of the first instruction set or a different instruction set. The processor cores 1502A - 1502N can be heterogeneous in terms of microarchitecture, where one or more cores with relatively high power consumption are coupled with one or more power cores with lower power consumption. As another example, the processor cores 1502A - 1502N are heterogeneous in terms of computing power. Additionally, the processor 1500 can be implemented on one or more chips or as a System - on - Chip (SoC) integrated circuit that also has the illustrated components in addition to other components.

[0265] Figure 15B is a block diagram of the hardware logic of the graphics processor core block 1519 according to some embodiments described herein. In some embodiments, Figure 15B Elements having the same reference numerals (or names) as elements in any other figure herein can operate or function in a manner similar to the way described elsewhere herein. In one embodiment, the graphics processor core block 1519 is an example of a partition of a graphics processor. The graphics processor core block 1519 can be included in Figure 15A the integrated graphics processor 1508 or a discrete graphics processor, a parallel processor, and / or a computing accelerator. The graphics processor as described herein can include multiple graphics core blocks based on a target power and performance envelope. Each graphics processor core block 1519 can include a functional block 1530 coupled to a plurality of graphics cores 1521A - 1521F, the plurality of graphics cores 1521A - 1521F including modular blocks of fixed - function logic and general - purpose programmable logic. The graphics processor core block 1519 also includes a shared / cache memory 1536 that can be accessed by all of the graphics cores 1521A - 1521F, rasterizer logic 1537, and additional fixed - function logic 1538.

[0266] In some embodiments, the functional block 1530 includes a geometry / fixed function pipeline 1531 that can be shared by all the graphics cores in the graphics processor core block 1519. In embodiments, the geometry / fixed function pipeline 1531 includes a 3D geometry pipeline, a video front-end unit, a thread generator and a global thread dispatcher, and a unified return buffer manager that manages the unified return buffer. In one embodiment, the functional block 1530 also includes a graphics SoC interface 1532, a graphics microcontroller 1533, and a media pipeline 1534. The graphics SoC interface 1532 provides an interface between the graphics processor core block 1519 and other core blocks within the graphics processor or the compute accelerator SoC. The graphics microcontroller 1533 is a programmable sub-processor that can be configured to manage various functions of the graphics processor core block 1519, including thread dispatching, scheduling, and preemption. The media pipeline 1534 includes logic for facilitating the decoding, encoding, preprocessing, and / or postprocessing of multimedia data, including image and video data. The media pipeline 1534 implements media operations via requests to the compute or sampling logic within the graphics cores 1521A - 1521F. One or more pixel backends 1535 may also be included within the functional block 1530. The pixel backend 1535 includes buffer memories for storing pixel color values and is capable of performing blending operations and lossless color compression on the rendered pixel data.

[0267] In one embodiment, the graphics SoC interface 1532 enables the graphics processor core block 1519 to communicate with general-purpose application processor cores (e.g., CPUs) and / or other components within the system host CPU that is within the SoC or coupled to the SoC via a peripheral interface. The graphics SoC interface 1532 also enables communication with off-chip memory hierarchy components such as shared last-level cache memory, system RAM, and / or embedded on-chip or package-on-chip DRAM. The SoC interface 1532 is also capable of enabling communication with fixed-function devices within the SoC such as camera imaging pipelines, and enabling and / or implementing global memory atomicity that can be shared between the graphics processor core block 1519 and the CPU within the SoC. The graphics SoC interface 1532 can also implement power management control for the graphics processor core block 1519 and enable an interface between the clock domain of the graphics processor core block 1519 and other clock domains within the SoC. In one embodiment, the graphics SoC interface 1532 enables receipt of command buffers from a command stream converter and a global thread dispatcher that are configured to provide commands and instructions to each of one or more graphics cores within the graphics processor. The commands and instructions can be dispatched to the media pipeline 1534 when media operations are to be performed and can be dispatched to the geometry and fixed-function pipeline 1531 when graphics processing operations are to be performed. When compute operations are to be performed, compute dispatch logic can dispatch commands to the graphics cores 1521A-1521F, thereby bypassing the geometry pipeline and the media pipeline.

[0268] The graphics microcontroller 1533 can be configured to perform various scheduling tasks and management tasks for the graphics processor core block 1519. In one embodiment, the graphics microcontroller 1533 can execute graphics workloads and / or compute workloads scheduled on the vector engines 1522A - 1522F, 1524A - 1524F and matrix engines 1523A - 1523F, 1525A - 1525F within the graphics cores 1521A - 1521F. In this scheduling model, host software executing on the CPU core of the SoC including the graphics processor core block 1519 can submit a workload to one of a plurality of graphics processor doorbells, which invokes a scheduling operation for the appropriate graphics engine. The scheduling operations include: determining which workload to run next, submitting the workload to the command stream converter, preempting an existing workload running on the engine, monitoring the progress of the workload, and notifying the host software when the workload is complete. In one embodiment, the graphics microcontroller 1533 is also capable of facilitating a low - power or idle state for the graphics processor core block 1519, thereby providing the ability to save and restore registers within the graphics processor core block 1519 across low - power state transitions independent of the operating system and / or graphics driver software on the system.

[0269] The graphics processor core block 1519 can have more or fewer graphics cores 1521A - 1521F than shown, up to N modular graphics cores. For each set of N graphics cores, the graphics processor core block 1519 can also include: a shared / cache memory 1536, which can be configured as either shared memory or cache memory; rasterizer logic 1537; and additional fixed - function logic 1538 for accelerating various graphics and compute processing operations.

[0270] Within each of the graphics cores 1521A - 1521F is included a set of execution resources that can be used to perform graphics operations, media operations, and compute operations in response to requests made by the graphics pipeline, media pipeline, or shader program. The graphics cores 1521A - 1521F include multiple vector engines 1522A - 1522F, 1524A - 1524F, matrix acceleration units 1523A - 1523F, 1525A - 1525D, cache / shared local memory (SLM), samplers 1526A - 1526F, and ray tracing units 1527A - 1527F.

[0271] The vector engines 1522A - 1522F, 1524A - 1524F are general - purpose graphics processing units capable of performing floating - point and integer / fixed - point logical operations to serve graphics operations, media operations, or compute operations (including graphics programs, media programs, or compute / GPGPU programs). The vector engines 1522A - 1522F, 1524A - 1524F are capable of operating with variable vector widths using SIMD execution mode, SIMT execution mode, or SIMT+SIMD execution mode. The matrix acceleration units 1523A - 1523F, 1525A - 1525D include matrix - matrix and matrix - vector acceleration logic that improves the performance of matrix operations, especially low - precision and mixed - precision (e.g., INT8, FP16, BF16, FP8) matrix operations for machine learning. In one embodiment, each of the matrix acceleration units 1523A - 1523F, 1525A - 1525D includes one or more systolic arrays of processing elements capable of performing concurrent matrix multiplication or dot - product operations on matrix elements.

[0272] The samplers 1526A - 1526F are capable of reading media data or texture data into memory and sampling the data in different ways based on the configured sampler state and the texture / media format being read. Threads executing on the vector engines 1522A - 1522F, 1524A - 1524F or the matrix acceleration units 1523A - 1523F, 1525A - 1525D can utilize the cache / SLM 1528A - 1528F within each of the graphics cores 1521A - 1521F. The cache / SLM 1528A - 1528F can be configured as a pool of cache memory or shared memory local to each of the corresponding graphics cores 1521A - 1521F. The ray - tracing units 1527A - 1527F within the graphics cores 1521A - 1521F include ray - traversal / intersection circuitry for performing ray traversal using a bounding - volume hierarchy (BVH) and identifying intersections between rays enclosed within the BVH volume and primitives. In one embodiment, the ray - tracing units 1527A - 1527F include circuitry for performing depth testing and culling (e.g., using a depth buffer or similar arrangement). In one implementation, the ray - tracing units 1527A - 1527F perform traversal and intersection operations in cooperation with image denoising, at least part of which can be performed using the associated matrix acceleration units 1523A - 1523F, 1525A - 1525D.

[0273] Figure 15Cis a block diagram of a general-purpose graphics processing unit (GPGPU) 1570 according to an embodiment described herein. The GPGPU 1570 can be configured as a graphics processor (e.g., graphics processor 1508) and / or a compute accelerator. The GPGPU 1570 can be interconnected with a host processor (e.g., one or more CPUs 1546) and memories 1571, 1572 via one or more system and / or memory buses. Memory 1571 can be a system memory that can be shared with one or more CPUs 1546, and memory 1572 is a device memory dedicated to the GPGPU 1570. For example, components within the GPGPU 1570 and memory 1572 can be mapped to memory addresses that can be accessed by one or more CPUs 1546. Access to memories 1571 and 1572 can be facilitated via a memory controller 1568. The memory controller 1568 can include an internal direct memory access (DMA) controller 1569, or can include logic for performing operations otherwise performed by the DMA controller.

[0274] The GPGPU 1570 includes a plurality of cache memories, which include an L2 cache 1553, an L1 cache 1554, an instruction cache 1555, and a shared memory 1556, at least a portion of which can also be partitioned as a cache memory. The GPGPU 1570 also includes a plurality of compute units 1560A - 1560N. Each compute unit 1560A - 1560N includes a set of vector registers 1561, a set of scalar registers 1562, a set of vector logic units 1563, and a set of scalar logic units 1564. The compute units 1560A - 1560N can also include a local shared memory 1565 and a program counter 1566. The compute units 1560A - 1560N can be coupled to a constant cache 1567, which can be used to store constant data, which is data that does not change during the execution of a kernel program or a shader program on the GPGPU 1570. The constant cache 1567 can be a scalar data cache, and the cached data can be fetched directly into the scalar registers 1562.

[0275] During operation, one or more CPUs 1546 may write commands to registers in GPGPU 1570, or to memory in GPGPU 1570 that has been mapped into an accessible address space. Command processor 1557 may read commands from registers or memory and determine how to process those commands within GPGPU 1570. Thread dispatcher 1558 may then be used to dispatch threads to computing units 1560A-1560N to execute those commands. Each computing unit 1560A-1560N may execute threads independently of other computing units. In addition, each computing unit 1560A-1560N may be independently configured for conditional computation and may conditionally output the results of the computation to memory. Command processor 1557 may interrupt one or more CPUs 1546 when the submitted commands are completed.

[0276] Figures 16A - 16C The diagram is described in this article, for example according to Figures 15A - 15C The embodiments provide block diagrams of additional graphics processor and computing accelerator architectures. Figures 16A - 16C Elements having the same or similar names as elements of any other figures herein describe the same elements as in the other figures, can operate or function in a similar manner as in the other figures, may include the same components, and may be linked to other entities such as, but not limited to, those described elsewhere herein.

[0277] Figure 16A 1 is a block diagram of a graphics processor 1600, which may be a discrete graphics processing unit, or may be a graphics processor integrated with multiple processing cores or other semiconductor devices, such as, but not limited to, memory devices or network interfaces. Graphics processor 1600 may be a variant of graphics processor 1508 and may be used in place of graphics processor 1508. Therefore, the disclosure of any feature herein in conjunction with graphics processor 1508 also discloses the corresponding combination with graphics processor 1600, but is not limited thereto. The graphics processor may communicate via a memory-mapped I / O interface to registers on the graphics processor and using commands placed into processor memory. Graphics processor 1600 may include a memory interface 1614 for accessing memory. Memory interface 1614 may be an interface to local memory, one or more internal caches, one or more shared external caches, and / or to system memory.

[0278] Optionally, the graphics processor 1600 further includes a display controller 1602 for driving display output data to a display device 1618. The display controller 1602 includes hardware for one or more overlay planes for the display and the composition of multi-layered video or user interface elements. The display device 1618 may be an internal or external display device. In one embodiment, the display device 1618 is a head-mounted display device, such as a virtual reality (VR) display device or an augmented reality (AR) display device. The graphics processor 1600 may include a video codec engine 1606 for encoding media into one or more media encoding formats, decoding media from one or more media encoding formats, or transcoding media between one or more media encoding formats, the one or more media encoding formats including but not limited to: Moving Picture Experts Group (MPEG) formats (such as MPEG-2), Advanced Video Coding (AVC) formats (such as H.264 / MPEG-4 AVC, H.265 / HEVC, Alliance for Open Media (AOMedia) VP8, VP9), and Society of Motion Picture & Television Engineers (SMPTE) 421M / VC-1, and Joint Photographic Experts Group (JPEG) formats (such as JPEG, and Motion JPEG (MJPEG) formats).

[0279] The graphics processor 1600 may include a block image transfer (BLIT) engine 1603 for performing two-dimensional (2D) rasterizer operations, including, for example, bit boundary block transfers. However, alternatively, one or more components of the graphics processing engine (GPE) 1610 may be used to perform 2D graphics operations. In some embodiments, the GPE 1610 is a computing engine for performing graphics operations, the graphics operations including three-dimensional (3D) graphics operations and media operations.

[0280] The GPE 1610 may include a 3D pipeline 1612 for performing 3D operations, such as rendering three-dimensional images and scenes using processing functions for 3D primitive shapes (e.g., rectangles, triangles, etc.). The 3D pipeline 1612 includes programmable and fixed-function elements that perform various tasks within the element and / or generate execution threads to the 3D / media subsystem 1615. Although the 3D pipeline 1612 can be used to perform media operations, embodiments of the GPE 1610 also include a media pipeline 1616 that is dedicated to performing media operations, such as video post-processing and image enhancement.

[0281] The media pipeline 1616 may include fixed-function or programmable logic units for performing one or more specialized media operations, such as video decode acceleration, video deinterlacing, and video encode acceleration, in place of, or on behalf of, the video codec engine 1606. The media pipeline 1616 may additionally include a thread generation unit for generating threads for execution on the 3D / media subsystem 1615. The generated threads execute computations for media operations on one or more graphics execution units included in the 3D / media subsystem 1615.

[0282] The 3D / media subsystem 1615 may include logic for executing the threads generated by the 3D pipeline 1612 and the media pipeline 1616. The pipelines may send thread execution requests to the 3D / media subsystem 1615, which includes thread dispatch logic for arbitrating and dispatching various requests for available thread execution resources. The execution resources include an array of graphics execution units for processing 3D threads and media threads. The 3D / media subsystem 1615 may include one or more internal caches for thread instructions and data. Additionally, the 3D / media subsystem 1615 may also include shared memory for sharing data between threads and for storing output data, which includes registers and addressable memory.

[0283] Figure 16B Illustrated is a graphics processor 1620, which is a variant of the graphics processor 1600 and can be used in place of the graphics processor 1600 and vice versa. Thus, any disclosure of a feature in connection with the graphics processor 1600 herein also discloses the corresponding combination with the graphics processor 1620, but is not limited thereto. According to the embodiments described herein, the graphics processor 1620 has a tiled architecture. The graphics processor 1620 may include a graphics processing engine cluster 1622 that has within the graphics engine tiles 1610A - 1610D Figure 16AMultiple instances of the graphics processing engine 1610. Each graphics engine slice 1610A-1610D can be interconnected via a set of slice interconnects 1623A-1623F. Each graphics engine slice 1610A-1610D can also be connected to memory modules or memory devices 1626A-1626D via memory interconnects 1625A-1625D. The memory devices 1626A-1626D can use any graphics memory technology. For example, the memory devices 1626A-1626D can be graphics double data rate (GDDR) memories. The memory devices 1626A-1626D can be high bandwidth memory (HBM) modules that can be on-die with their respective graphics engine slices 1610A-1610D. The memory devices 1626A-1626D can be stacked memory devices that can be stacked on top of their respective graphics engine slices 1610A-1610D. Each graphics engine slice 1610A-1610D and the associated memory 1626A-1626D can reside on separate dielets that are bonded to a base die or a base substrate, as further described in Figures 24B - 24D as described in further detail.

[0284] The graphics processor 1620 can be configured with a non-uniform memory access (NUMA) system in which the memory devices 1626A-1626D are coupled to the associated graphics engine slices 1610A-1610D. A given memory device can be accessed by a graphics engine slice different from the one to which the memory device is directly connected. However, when accessing the local slice, the access latency to the memory devices 1626A-1626D can be minimized. In one embodiment, a cache coherent NUMA (ccNUMA) system is enabled that uses the slice interconnects 1623A-1623F to enable communication between cache controllers within the graphics engine slices 1610A-1610D to maintain a consistent memory image when more than one cache stores the same memory location.

[0285] The graphics processing engine cluster 1622 can be connected to the on-chip or on-package fabric interconnect 1624. In one embodiment, the fabric interconnect 1624 includes a network processor, a network on a chip (NoC), or another switching processor that enables the fabric interconnect 1624 to act as a packet-switched fabric interconnect for exchanging data packets between components of the graphics processor 1620. The fabric interconnect 1624 can enable communication between the graphics engine slices 1610A-1610D and components such as the video codec 1606 and one or more copy engines 1604. The copy engines 1604 can be used to move data out of the memory devices 1626A-1626D and memory external to the graphics processor 1620 (e.g., system memory), move data into the memory devices 1626A-1626D and memory external to the graphics processor 1620 (e.g., system memory), and move data between the memory devices 1626A-1626D and memory external to the graphics processor 1620 (e.g., system memory). The fabric interconnect 1624 can also be used to interconnect the graphics engine slices 1610A-1610D. The graphics processor 1620 can optionally include a display controller 1602 for enabling connection to an external display device 1618. The graphics processor can also be configured as a graphics accelerator or a compute accelerator. In the accelerator configuration, the display controller 1602 and the display device 1618 can be omitted.

[0286] The graphics processor 1620 can be connected to the host system via the host interface 1628. The host interface 1628 can enable communication between the graphics processor 1620, the system memory, and / or other system components. The host interface 1628 can be, for example, a PCI Express bus or another type of host system interface. For example, the host interface 1628 can be an NVLink or NVSwitch interface. The host interface 1628 and the fabric interconnect 1624 can cooperate to enable multiple instances of the graphics processor 1620 to act as a single logical device. The cooperation between the host interface 1628 and the fabric interconnect 1624 can also enable the individual graphics engine slices 1610A-1610D to present themselves to the host system as different logical graphics devices.

[0287] Figure 16C FIG. illustrates a compute accelerator 1630 according to an embodiment described herein. The compute accelerator 1630 can include Figure 16BThe architectural similarity of the graphics processor 1620 and is optimized for computing acceleration. The compute engine cluster 1632 may include a set of compute engine slices 1640A - 1640D, and the set of compute engine slices 1640A - 1640D includes execution logic optimized for parallel or vector-based general computing operations. The compute engine slices 1640A - 1640D may not include fixed-function graphics processing logic, but in some embodiments, one or more of the compute engine slices 1640A - 1640D may include logic for performing media acceleration. The compute engine slices 1640A - 1640D may be connected to memories 1626A - 1626D via memory interconnects 1625A - 1625D. The memories 1626A - 1626D and the memory interconnects 1625A - 1625D may be of a similar technology as in the graphics processor 1620 or may be different technologies. The compute engine slices 1640A - 1640D may also be interconnected via a set of slice interconnects 1623A - 1623F and may be connected to and / or interconnected by the fabric interconnect 1624. In one embodiment, the compute accelerator 1630 includes a large L3 cache 1636 that can be configured as a device-wide cache. The compute accelerator 1630 can also be connected to the host processor and memory via the host interface 1628 in a manner similar to Figure 16B the graphics processor 1620.

[0288] The compute accelerator 1630 may also include an integrated network interface 1642. In one embodiment, the integrated network interface 1642 includes a network processor and controller logic that enables the compute engine cluster 1632 to communicate through the physical layer interconnect 1644 without data crossing the memory of the host system. In one embodiment, one of the compute engine slices 1640A - 1640D is replaced by network processor logic, and data to be transmitted or received via the physical layer interconnect 1644 can be directly transmitted to or from the memories 1626A - 1626D. Multiple instances of the compute accelerator 1630 can be combined into a single logical device via the physical layer interconnect 1644. Alternatively, each compute engine slice 1640A - 1640D can be presented as a different network-accessible compute accelerator device. Graphics Processing Engine

[0289] Figure 17 is a block diagram of a graphics processing engine 1710 of a graphics processor according to some embodiments. The Graphics Processing Engine (GPE) 1710 can be Figure 16A a certain version of the GPE 1610 shown in Figure 16B and can also represent Figure 17Elements having the same or similar names as elements in any other figure herein describe the same elements as in other figures, can operate or run in a similar manner as in other figures, may include the same components, and may be linked to other entities such as those described elsewhere herein, but are not limited thereto. For example, in Figure 17 is also illustrated Figure 16A the 3D pipeline 1612 and the media pipeline 1616. The media pipeline 1616 is optional in some embodiments of the GPE 1710 and may not be explicitly included within the GPE 1710. For example and in at least one embodiment, a separate media and / or image processor is coupled to the GPE 1710.

[0290] The GPE 1710 may be coupled to or include a command stream converter 1703 that provides a command stream to the 3D pipeline 1612 and / or the media pipeline 1616. Alternatively or additionally, the command stream converter 1703 may be directly coupled to the unified return buffer 1718. The unified return buffer 1718 may be communicatively coupled to the graphics core cluster 1714. Optionally, the command stream converter 1703 is coupled to a memory, which may be a system memory, or one or more of an internal cache memory and a shared cache memory. The command stream converter 1703 may receive commands from the memory and send these commands to the 3D pipeline 1612 and / or the media pipeline 1616. These commands are indications taken from a ring buffer that stores commands for the 3D pipeline 1612 and the media pipeline 1616. The ring buffer may additionally include a batch command buffer that stores a batch of multiple commands. Commands for the 3D pipeline 1612 may also include references to data stored in the memory, such as but not limited to vertex data and geometry data for the 3D pipeline 1612 and / or image data and memory objects for the media pipeline 1616. The 3D pipeline 1612 and the media pipeline 1616 process commands and data by performing operations via logic within the respective pipelines or by dispatching one or more execution threads to the graphics core cluster 1714. The graphics core cluster 1714 may include one or more graphics core blocks (e.g., graphics core block 1715A, graphics core block 1715B), each block including one or more graphics cores. Each graphics core includes a set of graphics execution resources, which includes: general and graphics-specific execution logic for performing graphics operations and computational operations; and fixed function texture processing logic and / or machine learning and artificial intelligence acceleration logic.

[0291] In various embodiments, the 3D pipeline 1612 can include fixed function and programmable logic for processing one or more shader programs by processing instructions and dispatching execution threads to the graphics core cluster 1714, such as vertex shaders, geometry shaders, pixel shaders, fragment shaders, compute shaders, or other shader programs. The graphics core cluster 1714 provides a unified execution resource block for use in processing these shader programs. The multi-functional execution logic (e.g., execution units) within the graphics core blocks 1715A - 1715B of the graphics core cluster 1714 includes support for various 3D API shader languages and can execute multiple synchronized execution threads associated with multiple shaders.

[0292] The graphics core cluster 1714 can include execution logic for performing media functions such as video and / or image processing. In addition to graphics processing operations, the execution units can also include general-purpose logic programmable to perform parallel general-purpose computing operations. The general-purpose logic can perform processing operations in parallel or in conjunction Figure 14 with the (one or more) processor cores 1407 or general-purpose logic within the cores 1502A - 1502N as in Figure 15A .

[0293] Output data generated by threads executing on the graphics core cluster 1714 can output the data to memory in a unified return buffer (URB) 1718. The URB 1718 can store data for multiple threads. The URB 1718 can be used to send data between different threads executing on the graphics core cluster 1714. The URB 1718 can additionally be used for synchronization between threads executing on the graphics core cluster 1714 and fixed function logic within the shared function logic 1720.

[0294] Optionally, the graphics core cluster 1714 can be scalable such that the array includes a variable number of graphics cores, each with a variable number of execution units based on the target power and performance levels of the GPE 1710. The execution resources can be dynamically scalable such that the execution resources can be enabled or disabled as needed.

[0295] The graphics core cluster 1714 is coupled to shared function logic 1720, which includes multiple resources shared among the graphics cores in the graphics core array. The shared functions within the shared function logic 1720 are hardware logic units that provide specialized complementary functions to the graphics core cluster 1714. In various embodiments, the shared function logic 1720 includes, but is not limited to, sampler 1721 logic, math 1722 logic, and inter-thread communication (ITC) 1723 logic. Additionally, one or more caches 1725 can be implemented within the shared function logic 1720.

[0296] The shared functions are implemented at least in cases where the demand for a given specialized function is not sufficient to be included within the graphics core cluster 1714. Instead, a single instantiation of that specialized function is implemented as a stand-alone entity within the shared function logic 1720 and is shared among the execution resources within the graphics core cluster 1714. The exact set of functions shared among the graphics core clusters 1714 and included within the graphics core clusters 1714 varies by embodiment. Specific shared functions widely used by the graphics core cluster 1714 within the shared function logic 1720 can be included within the shared function logic 1716 within the graphics core cluster 1714. Optionally, the shared function logic 1716 within the graphics core cluster 1714 can include some or all of the logic within the shared function logic 1720. All of the logic elements within the shared function logic 1720 can be replicated within the shared function logic 1716 of the graphics core cluster 1714. Alternatively, the shared function logic 1720 is excluded in favor of the shared function logic 1716 within the graphics core cluster 1714. Graphics Processing Resources

[0297] Figures 18A - 18C Illustrates execution logic including an array of processing elements employed in a graphics processor according to embodiments described herein. Figure 18A Illustrates a graphics core cluster according to an embodiment. Figure 18B Illustrates a vector engine of a graphics core according to an embodiment. Figure 18C Illustrates a matrix engine of a graphics core according to an embodiment. Figures 18A - 18C Elements having the same reference numerals as elements in any other figure herein can operate or function in any manner similar to the ways described elsewhere herein, but are not limited thereto. For example, Figures 18A - 18C the elements can be considered in the context of Figure 15B the graphics processor core block 1519 and / or Figure 17 the graphics core blocks 1715A - 1715B of Figures 18A - 18C In one embodiment, the elements have the same as Figure 15Athe graphics processing unit 1508 or Figure 15C equivalent components of the GPGPU 1570 have similar functions.

[0298] such as Figure 18A As shown, in one embodiment, the graphics core cluster 1714 includes a graphics core block 1715, and the graphics core block 1715 can be Figure 17 the graphics core block 1715A or the graphics core block 1715B. The graphics core block 1715 can include any number of graphics cores (e.g., graphics core 1815A, graphics core 1815B, up to graphics core 1815N), and multiple instances of the graphics core block 1715 can be included. In one embodiment, the elements of the graphics cores 1815A - 1815N have functions similar or equivalent to those of Figure 15B the elements of the graphics cores 1521A - 1521F. In such embodiments, each of the graphics cores 1815A - 1815N includes circuitry including, but not limited to: vector engines 1802A - 1802N, matrix engines 1803A - 1803N, memory load / store units 1804A - 1804N, instruction caches 1805A - 1805N, data caches / shared local memories 1806A - 1806N, ray tracing units 1808A - 1808N, samplers 1810A - 1810N. The circuitry of the graphics cores 1815A - 1815N can additionally include fixed function logic 1812A - 1812N. The number of vector engines 1802A - 1802N and matrix engines 1803A - 1803N within the designed graphics cores 1815A - 1815N can vary based on the workload, performance, and power targets for that design.

[0299] With reference to graphics core 1815A, vector engine 1802A and matrix engine 1803A can be configured to perform parallel computing operations on data in various integer and floating-point data formats based on instruction pairs associated with shader programs. Each vector engine 1802A and matrix engine 1803A can act as a programmable general-purpose computing unit capable of executing multiple synchronous hardware threads to process multiple data elements in parallel for each thread. Vector engine 1802A and matrix engine 1803A support processing variable-width vectors in various SIMD widths, including but not limited to SIMD8, SIMD16, and SIMD32. Input data elements can be stored in registers as packed data types, and vector engine 1802A and matrix engine 1803A can process each element based on the data size of the element. For example, when operating on a 256-bit wide vector, the 256 bits of the vector are stored in a register, and the vector is processed as four separate 64-bit packed data elements (Quad-Word (QW) size data elements), eight separate 32-bit packed data elements (Double Word (DW) size data elements), sixteen separate 16-bit packed data elements (Word (W) size data elements), or thirty-two separate 8-bit data elements (byte (B) size data elements). However, different vector widths and register sizes are possible. In one embodiment, vector engine 1802A and matrix engine 1803A can also be configured to perform SIMT operations on groups of units and groups of threads of various sizes (e.g., 8, 16, or 32 threads).

[0300] Continuing with graphics core 1815A, memory load / store unit 1804A services memory access requests issued by vector engine 1802A, matrix engine 1803A, and / or other components of graphics core 1815A having access to memory. Memory access requests can be processed by memory load / store unit 1804A to load or store the requested data to / from a cache or memory, or to / from a register file associated with vector engine 1802A and / or matrix engine 1803A. Memory load / store unit 1804A can also perform prefetch operations. Additionally refer to Figure 19, in one embodiment, the memory load / store unit 1804A is configured to provide SIMT scatter / gather prefetch or block prefetch for data stored in the memory 1910, from memory that is local to other tiles via the tile interconnect 1908, or from the system memory. Prefetching can be performed on a specific L1 cache (e.g., the data cache / shared local memory 1806A), the L2 cache 1904, or the L3 cache 1906. In one embodiment, prefetching the L3 cache 1906 automatically causes the data to be stored in the L2 cache 1904.

[0301] The instruction cache 1805A stores instructions to be executed by the graphics core 1815A. In one embodiment, the graphics core 1815A also includes instruction fetch and prefetch circuitry for fetching or prefetching instructions into the instruction cache 1805A. The graphics core 1815A also includes instruction decoding logic for decoding the instructions within the instruction cache 1805A. The data cache / shared local memory 1806A can be configured as a data cache managed by a cache controller implementing a cache replacement policy and / or configured as a shared memory that is explicitly managed. The ray tracing unit 1808A includes circuitry for accelerating ray tracing operations. The sampler 1810A provides texture sampling for 3D operations and media sampling for media operations. The fixed function logic 1812A includes fixed function circuitry that is shared among instances of the vector engine 1802A and the matrix engine 1803A. The graphics cores 1815B - 1815N can operate in a manner similar to the graphics core 1815A.

[0302] The functions of the instruction caches 1805A - 1805N, the data caches / shared local memories 1806A - 1806N, the ray tracing units 1808A - 1808N, the samplers 1810A - 1812N, and the fixed function logics 1812A - 1812N correspond to the equivalent functions in the graphics processor architectures described herein. For example, the instruction caches 1805A - 1805N can operate in a manner similar to Figure 15C the instruction cache 1555. The data caches / shared local memories 1806A - 1806N, the ray tracing units 1808A - 1808N, and the samplers 1810A - 1812N can operate in a manner similar to Figure 15B the caches / SLMs 1528A - 1528F, the ray tracing units 1527A - 1527F, and the samplers 1526A - 1526F. The fixed function logics 1812A - 1812N can include Figure 15B elements of the geometry / fixed function pipeline 1531 and / or additional fixed function logic 1538. In one embodiment, the ray tracing units 1808A - 1808N include circuitry for performing theFigure 3C A circuit for ray tracing acceleration operations performed by the ray tracing core 372.

[0303] As Figure 18B shown, in one embodiment, the vector engine 1802 includes an instruction fetch unit 1837, a general register file (GRF) array 1824, an architectural register file (ARF) array 1826, a thread arbiter 1822, a send unit 1830, a branch unit 1832, a set of SIMD floating point units (FPUs) 1834, and in one embodiment, a set of integer SIMD ALUs 1835. The GRF 1824 and the ARF 1826 include a set of general register files and architectural register files associated with each hardware thread that can be active in the vector engine 1802. In one embodiment, the per-thread architectural state is maintained in the ARF 1826, while data used during thread execution is stored in the GRF 1824. The execution state of each thread, including the instruction pointer for each thread, can be saved in thread-specific registers in the ARF 1826. Register renaming can be used to dynamically allocate registers to hardware threads.

[0304] In one embodiment, the vector engine 1802 has an architecture that is a combination of Simultaneous Multi-Threading (SMT) and Fine-Grained Interleaved Multi-Threading (IMT). This architecture has a modular configuration that can be fine-tuned at design time based on the target number of synchronized threads and the number of registers per graphics core, where the graphics core resources are divided across the logic for executing multiple synchronized threads. The number of logical threads that can be executed by the vector engine 1802 is not limited to the number of hardware threads, and multiple logical threads can be assigned to each hardware thread.

[0305] In one embodiment, the vector engine 1802 may issue multiple instructions in concert, and these instructions may each be different instructions. The thread arbiter 1822 may dispatch the instructions to one of the send unit 1830, the branch unit 1832, or the (one or more) SIMD FPUs 1834 for execution. Each execution thread may access 128 general-purpose registers within the GRF 1824, where each register may store 32 bytes that can be accessed as a variable-width vector with a data element of 32 bytes. In one embodiment, each thread has access to 4 kilobytes within the GRF 1824, but the embodiments are not limited thereto, and more or fewer register resources may be provided in other embodiments. In one embodiment, the vector engine 1802 is partitioned into seven hardware threads that can independently perform computational operations, but the number of threads for each vector engine 1802 may also vary according to the embodiment. For example, in one embodiment, up to 16 hardware threads are supported. In an embodiment where seven threads can access 4 kilobytes, the GRF 1824 may store a total of 28 kilobytes. In the case where 16 threads can access 4 kilobytes, the GRF 1824 may store a total of 64 kilobytes. Flexible addressing modes may permit addressing of the registers together, thereby effectively creating wider registers or representing strided rectangular block data structures.

[0306] In one embodiment, memory operations, sampler operations, and other longer-latency system communications are dispatched via "send" instructions executed by the message passing send unit 1830. In one embodiment, branch instructions are dispatched to the dedicated branch unit 1832 to facilitate SIMD scatter and eventual gather.

[0307] In one embodiment, vector engine 1802 includes one or more SIMD floating point units ((one or more) FPUs) 1834 for performing floating point operations. In one embodiment, the (one or more) FPUs 1834 also support integer computations. In one embodiment, the (one or more) FPUs 1834 may execute up to M 32-bit floating point (or integer) operations, or execute up to 2M 16-bit integer or 16-bit floating point operations. In one embodiment, at least one of the (one or more) FPUs provides extended math capabilities that support high throughput transcendental math functions and double precision 64-bit floating point. In some embodiments, a set 1835 of 8-bit integer SIMD ALUs also exists and may be specifically optimized to perform operations associated with machine learning computations. In one embodiment, the SIMD ALUs are replaced by a set 1834 of additional SIMD ALUs configurable to perform integer and floating point operations. In one embodiment, the SIMD FPU 1834 and SIMD ALU 1835 are configurable to execute SIMT programs. In one embodiment, combined SIMD+SIMT operations are supported.

[0308] In one embodiment, an array of multiple instances of vector engine 1802 may be instantiated within a graphics core. For scalability, the product architect may choose the exact number of vector engines grouped per graphics core. In one embodiment, vector engine 1802 may execute instructions across multiple execution lanes. In a further embodiment, each thread executing on vector engine 1802 is executed on a different lane.

[0309] As Figure 18CAs shown, in one embodiment, matrix engine 1803 includes an array of processing elements configured to perform tensor operations, which include vector / matrix operations and matrix / matrix operations, such as but not limited to matrix multiplication and / or dot product operations. Matrix engine 1803 can be configured with M rows and N columns of processing elements (1852AA - 1852MN), and the processing elements (PE 1852AA - PE 1852MN) include multiplier and adder circuits organized in a pipelined manner. In one embodiment, processing elements 1852AA - 1852MN form the physical pipeline stages of an N-wide and M-deep systolic array, which can be used to perform vector / matrix operations or matrix / matrix operations in a data-parallel manner, including matrix multiplication, fused multiply-add, dot product, or other general matrix-matrix multiplication (GEMM) operations. In one embodiment, matrix engine 1803 supports 16-bit and 8-bit floating-point operations, as well as 8-bit, 4-bit, 2-bit, and binary integer operations. Matrix engine 1803 can also be configured to accelerate specific machine learning operations. In such embodiments, matrix engine 1803 can be configured with support for the bfloat (brain floating-point) 16-bit floating-point format, which has a different number of mantissa bits and exponent bits relative to the Institute of Electrical and Electronics Engineers (IEEE) 754 format, or the tensor floating-point 32-bit floating-point format (TF32).

[0310] In one embodiment, during each cycle, each stage can add the result of the operation performed in that stage to the output of the previous stage. In other embodiments, after a set of computational cycles, the pattern of data movement between processing elements 1852AA - 1852MN can vary based on the instruction or macro operation being executed. For example, in one embodiment, partial sum loopback is enabled, and the processing element can alternatively add the output of the current cycle to the output generated in the previous cycle. In one embodiment, the final stage of the systolic array can be configured with a loopback to the initial stage of the systolic array. In such embodiments, the number of physical pipeline stages can be decoupled from the number of logical pipeline stages supported by matrix engine 1803. For example, in the case where processing elements 1852AA - 1852MN are configured as an M-physical-stage systolic array, the loopback from stage M to the initial pipeline stage can enable processing elements 1852AA - 1852MN to operate as a systolic array with, for example, 2M, 3M, 4M, etc. logical pipeline stages.

[0311] In one embodiment, matrix engine 1803 includes memories 1841A - 1841N, 1842A - 1842M for storing input data in the form of row and column data for an input matrix. Memories 1842A - 1842M are configurable to store row elements (A0 - Am) of a first input matrix, and memories 1841A - 1841N are configurable to store column elements (B0 - Bn) of a second input matrix. The row elements and column elements are provided as inputs to processing elements 1852AA - 1852MN for processing. In one embodiment, the element rows and column elements of the input matrix may be stored in systolic register bank 1840 within matrix engine 1803 before these elements are provided to memories 1841A - 1841N, 1842A - 1842M. In one embodiment, excluding the systolic register bank 1840, memories 1841A - 1841N, 1842A - 1842M are loaded from registers in an associated vector engine (e.g., Figure 18B GRF1824 of vector engine 1802) or other memories of a graphics core including matrix engine 1803 (e.g., Figure 18A data cache / shared local memory 1806A for matrix engine 1803A). Results generated by processing elements 1852AA - 1852MN are then output to an output buffer and / or written to a register bank (e.g., systolic register bank 1840, GRF 1824, data cache / shared local memory 1806A - 1806N) for further processing by other functional units of the graphics processor or for output to memory.

[0312] In some embodiments, matrix engine 1803 is configured with support for input sparsity, where multiplication operations for sparse regions of the input data can be bypassed by skipping multiplication operations for operands having zero values. In one embodiment, processing elements 1852AA - 1852MN are configured to skip execution of certain operations having zero-valued inputs. In one embodiment, the sparsity within the input matrix can be detected, and operations having known zero output values can be bypassed before being submitted to processing elements 1852AA - 1852MN. Loading zero-valued operands into the processing elements can be bypassed, and processing elements 1852AA - 1852MN can be configured to perform multiplication on non-zero input elements. Matrix engine 1803 can also be configured with support for output sparsity such that operations having results predetermined to be zero can be bypassed. For input sparsity and / or output sparsity, in one embodiment, metadata is provided to processing elements 1852AA - 1852MN to indicate which processing elements and / or data channels will be active during a given processing cycle.

[0313] In one embodiment, matrix engine 1803 includes hardware for enabling operations on sparse data having a compressed representation of a sparse matrix, where the sparse matrix stores non-zero values and metadata identifying the locations of the non-zero values within the matrix. Exemplary compressed representations include, but are not limited to, compressed tensor representations such as compressed sparse row (CSR) representation, compressed sparse column (CSC) representation, compressed sparse fiber (CSF) representation. Support for compressed representations enables operations to be performed on inputs in compressed tensor format without the compressed representation being decompressed or decoded. In such embodiments, operations can be performed only on non-zero input values, and the resulting non-zero output values can be mapped into the output matrix. In some embodiments, hardware support for machine-specific lossless data compression formats is also provided, which are used when transferring data within the hardware or across a system bus. Such data can be retained in a compressed format for sparse input data, and matrix engine 1803 can use compressed metadata for the compressed data to enable operations to be performed only on non-zero values or to bypass blocks of zero data inputs for multiplication operations.

[0314] In various embodiments, input data can be provided by a programmer in a compressed tensor representation, or a codec can compress the input data into a compressed tensor representation or another sparse data encoding. Additionally, to support compressed tensor representations, streaming compression of sparse input data can be performed before the input data is provided to processing elements 1852AA - 1852MN. In one embodiment, compression is performed on data written to a cache memory associated with graphics core cluster 1714, where the compression is performed using an encoding supported by matrix engine 1803. In one embodiment, matrix engine 1803 includes support for inputs having a structured sparsity, in which a predetermined level or pattern of sparsity is imposed on the input data. The data can be compressed to a known compression ratio, where the compressed data is processed by compression elements 1852AA - 1852MN according to the metadata associated with the compressed data.

[0315] Figure 19 Illustrates slice 1900 of a multi-slice processor according to an embodiment. In one embodiment, slice 1900 represents Figure 16B graphics engine slices 1610A - 1610D or Figure 16COne of the compute engine slices 1640A - 1640D. The slice 1900 of the multi - slice graphics processor includes an array of graphics core clusters (e.g., graphics core cluster 1714A, graphics core cluster 1714B, up to graphics core cluster 1714N), where each graphics core cluster has an array of graphics cores 1815A - 1815N. The slice 1900 also includes a global dispatcher 1902 for dispatching threads to the processing resources of the slice 1900.

[0316] The slice 1900 may include an L3 cache 1906 and a memory 1910 or be coupled to the L3 cache 1906 and the memory 1910. In various embodiments, the L3 cache 1906 may be excluded, or the slice 1900 may include additional levels of cache, such as an L4 cache. In one embodiment, such as Figure 16B and Figure 16C in, each instance of the slice 1900 of the multi - slice graphics processor has an associated memory 1910. In one embodiment, the multi - slice processor may be configured as a multi - chip module, in which the L3 cache 1906 and / or the memory 1910 reside on a separate die different from the graphics core clusters 1714A - 1714N. In this context, a die is an integrated circuit that is at least partially encapsulated and includes different logic units that can be assembled with other dice into a larger package. For example, the L3 cache 1906 may be included in a dedicated cache die or reside on the same die as the graphics core clusters 1714A - 1714N. In one embodiment, the L3 cache 1906 may be included in an active base die or an active interposer as Figure 24C illustrated.

[0317] The memory fabric 1903 enables communication between the graphics core clusters 1714A - 1714N, the L3 cache 1906, and the memory 1910. The L2 cache 1904 is coupled to the memory fabric 1903 and can be configured to cache transactions executed via the memory fabric 1903. The slice interconnect 1908 enables communication with other slices on the graphics processor and can be Figure 16B and Figure 16COne of the tile interconnects 1623A - 1623F. In embodiments where the L3 cache 1906 is excluded from the tile 1900, the L2 cache 1904 may be configured as a combined L2 / L3 cache. The memory structure 1903 may be configured to route data to the L3 cache 1906 or to the memory controller associated with the memory 1910 based on the presence or absence of the L3 cache 1906 in a particular implementation. The L3 cache 1906 may be configured as a per-tile cache that is dedicated to the processing resources of the tile 1900 or may be part of a GPU-wide L3 cache.

[0318] Figure 20 is a block diagram illustrating the graphics processor instruction format 2000. The graphics processor execution units support an instruction set with instructions in a variety of formats. The solid boxes illustrate components that are typically included in the execution unit instructions, while the dashed boxes include optional or components that are only included in a subset of the instructions. In some embodiments, the graphics processor instruction format 2000 described and illustrated is a macro-instruction because they are instructions supplied to the execution unit, as opposed to micro-operations that result from instruction decoding once the instruction is processed. Thus, a single instruction can cause the hardware to perform multiple micro-operations.

[0319] The graphics processor execution units described herein may natively support instructions in the 128-bit instruction format 2010. Based on the selected instructions, instruction options, and number of operands, the 64-bit compact instruction format 2030 may be used for some instructions. The native 128-bit instruction format 2010 provides access to all instruction options, while some options and operations are restricted in the 64-bit format 2030. The native instructions available in the 64-bit format 2030 vary by embodiment. The instruction is partially compressed using a set of index values in the index field 2013. The execution unit hardware references a set of compression tables based on the index values and uses the compression table output to reconstruct the native instruction in the 128-bit instruction format 2010. Instructions of other sizes and formats may be used.

[0320] For each format, instruction opcode 2012 defines the operation to be performed by the execution unit. The execution unit executes each instruction in parallel across multiple data elements of each operand. For example, in response to an add instruction, the execution unit performs a synchronous add operation across each color channel representing a texture element or a picture element. By default, the execution unit executes each instruction across all data channels of the operand. Instruction control field 2014 can enable control of certain execution options such as channel selection (e.g., predication) and data channel order (e.g., swizzle). For instructions in the 128-bit instruction format 2010, execution size field 2016 limits the number of data channels to be executed in parallel. Execution size field 2016 may not be available for the 64-bit compact instruction format 2030.

[0321] Some execution unit instructions have up to three operands, including two source operands src0 2020, src1 2022, and one destination operand (dest 2018). Other instructions (such as, for example, data manipulation instructions, dot product instructions, multiply-add instructions, or multiply-accumulate instructions) may have a third source operand (e.g., SRC2 2024). Instruction opcode 2012 determines the number of source operands. The last source operand of an instruction can be an immediate (e.g., hard-coded) value passed with the instruction. The execution unit can also support multiple destination instructions, where one or more of the destinations are implied or implicit based on the instruction and / or the specified destination.

[0322] The 128-bit instruction format 2010 can include an access / addressing mode field 2026 that, for example, specifies the use of direct register addressing mode or indirect register addressing mode. When using the direct register addressing mode, the register addresses of one or more operands are directly provided by bits in the instruction.

[0323] The 128-bit instruction format 2010 can also include an access / addressing mode field 2026 that specifies the addressing mode and / or access mode of the instruction. The access mode can be used to define the data access alignment of the instruction. Access modes including a 16-byte aligned access mode and a 1-byte aligned access mode can be supported, where the byte alignment of the access mode determines the access alignment of the instruction operands. For example, when in a first mode, the instruction can use byte-aligned addressing for source and destination operands, and when in a second mode, the instruction can use 16-byte aligned addressing for all source and destination operands.

[0324] The addressing mode portion of access / addressing mode field 2026 can determine whether the instruction is to use direct addressing or indirect addressing. When using direct register addressing mode, the bits in the instruction directly provide the register addresses of one or more operands. When using indirect register addressing mode, the register addresses of one or more operands can be calculated based on the address register value and the address immediate number field in the instruction.

[0325] Instructions can be grouped based on the opcode 2012 bit field to simplify opcode decoding 2040. For an 8-bit opcode, bits 4, 5, and 6 allow the execution unit to determine the type of opcode. The exact opcode grouping shown is only an example. The move and logic opcode group 2042 can include data move and logic instructions (e.g., move (mov), compare (cmp)). The move and logic group 2042 can share the five least significant bits (LSB), where the move (mov) instruction takes the form of 0000xxxxb and the logic instruction takes the form of 0001xxxxb. The flow control instruction group 2044 (e.g., call, jump) includes instructions in the form of 0010xxxxb (e.g., 0x20). The miscellaneous instruction group 2046 includes a mix of instructions, including synchronization instructions (e.g., wait, send) in the form of 0011xxxxb (e.g., 0x30). The parallel math instruction group 2048 includes per-component arithmetic instructions (e.g., add, multiply (mul)) in the form of 0100xxxxb (e.g., 0x40). The parallel math instruction group 2048 executes arithmetic operations in parallel across data channels. The vector math group 2050 includes arithmetic instructions (e.g., dp4) in the form of 0101xxxxb (e.g., 0x50). The vector math group performs arithmetic on vector operands, such as dot product calculations. In one embodiment, the illustrated opcode decoding 2040 can be used to determine which part of the execution unit will be used to execute the decoded instruction. For example, some instructions can be designated as systolic instructions to be executed by the systolic array. Other instructions (such as ray tracing instructions (not shown)) can be routed to the ray tracing core or ray tracing logic within a slice or partition of the execution logic. Graphics Pipeline

[0326] Figure 21 is a block diagram of a graphics processor 2100 according to another embodiment. Figure 21Elements having the same or similar names as elements in any other figure in this document describe the same elements as in other figures, can operate or function in a similar manner as in other figures, may include the same components, and may be linked to other entities such as those described elsewhere in this document, but are not limited thereto.

[0327] The graphics processor 2100 may include different types of graphics processing pipelines, such as, a geometry pipeline 2120, a media pipeline 2130, a display engine 2140, thread execution logic 2150, and a render output pipeline 2170. The graphics processor 2100 may be a graphics processor within a multi-core processing system that includes one or more general-purpose processing cores. The graphics processor may be controlled by register writes to one or more control registers (not shown) or via commands issued through the ring interconnect 2102 to the graphics processor 2100. The ring interconnect 2102 may couple the graphics processor 2100 to other processing components (such as, other graphics processors or general-purpose processors). Commands from the ring interconnect 2102 are interpreted by a command stream converter 2103, which supplies instructions to the various components of the geometry pipeline 2120 or the media pipeline 2130.

[0328] The command stream converter 2103 may direct the operation of a vertex fetcher 2105, which reads vertex data from memory and executes vertex processing commands provided by the command stream converter 2103. The vertex fetcher 2105 may provide the vertex data to a vertex shader 2107, which performs coordinate space transformation and lighting operations on each vertex. The vertex fetcher 2105 and the vertex shader 2107 may execute vertex processing instructions by dispatching execution threads to the graphics cores 2152A - 2152B via a thread dispatcher 2131.

[0329] The graphics cores 2152A - 2152B may be an array of vector processors having instruction sets for performing graphics operations and media operations. The graphics cores 2152A - 2152B may have attached L1 caches 2151 dedicated to each array or shared between the arrays. The caches may be configured as data caches, instruction caches, or a single cache partitioned to contain data and instructions in different partitions.

[0330] The geometry pipeline 2120 may include a tessellation component for performing hardware-accelerated tessellation of 3D objects. The programmable hull shader 2111 may configure the tessellation operation. The programmable domain shader 2117 may provide a backend evaluation of the tessellation output. The tessellator 2113 may operate under the direction of the hull shader 2111 and may include dedicated logic for generating a detailed set of geometric objects based on a coarse geometric model, which is provided as input to the geometry pipeline 2120. Additionally, if tessellation is not used, the tessellation components (e.g., the hull shader 2111, the tessellator 2113, and the domain shader 2117) may be bypassed. The tessellation components may operate based on data received from the vertex shader 2107.

[0331] The complete geometric object may be processed by the geometry shader 2119 via one or more threads dispatched to the graphics cores 2152A - 2152B, or may proceed directly to the clipper 2129. The geometry shader may operate on the entire geometric object, rather than on vertices or vertex patches as in previous stages of the graphics pipeline. If tessellation is disabled, the geometry shader 2119 receives input from the vertex shader 2107. The geometry shader 2119 may be programmable by a geometry shader program to perform geometric tessellation in the case where the tessellation unit is disabled.

[0332] Before rasterization, the clipper 2129 processes the vertex data. The clipper 2129 may be a fixed-function clipper or a programmable clipper with clipping and geometry shader functionality. The rasterizer and depth test component 2173 in the render output pipeline 2170 may dispatch the pixel shader to convert the geometric object to a per-pixel representation. The pixel shader logic may be included in the thread execution logic 2150. Optionally, the application may bypass the rasterizer and depth test component 2173 and access the un-rasterized vertex data via the egress unit 2123.

[0333] The graphics processor 2100 has an interconnect bus, an interconnect fabric, or some other interconnect mechanism that allows data and messages to be passed between the major components of the processor. In some embodiments, the graphics cores 2152A - 2152B and associated logic units (e.g., L1 cache 2151, sampler 2154, texture cache 2158, etc.) are interconnected via a data port 2156 to perform memory accesses and communicate with the render output pipeline components of the processor. The sampler 2154, caches 2151, 2158, and the graphics cores 2152A - 2152B each may have separate memory access paths. Optionally, the texture cache 2158 may also be configured as a sampler cache.

[0334] The rendering output pipeline 2170 may include a rasterizer and depth test component 2173 that converts vertex-based objects into associated pixel-based representations. The rasterizer logic may include a windower / masker unit for performing fixed-function triangle and line rasterization. In some embodiments, associated render cache 2178 and depth cache 2179 are also available. The pixel operation component 2177 performs pixel-based operations on the data, but in some instances, pixel operations associated with 2D operations (e.g., bit block image transfer with blending) are performed by the 2D engine 2141 or, when displayed, by the display controller 2143 using an overlay display plane instead. The shared L3 cache 2175 may be available to all graphics components, allowing data to be shared without using the main system memory.

[0335] The media pipeline 2130 may include a media engine 2137 and a video front end 2134. The video front end 2134 may receive pipeline commands from the command stream converter 2103. The media pipeline 2130 may include a separate command stream converter. The video front end 2134 may process the media commands before sending them to the media engine 2137. The media engine 2137 may include a thread generation function for generating threads for dispatch to thread execution logic 2150 via a thread dispatcher 2131.

[0336] The graphics processor 2100 may include a display engine 2140. The display engine 2140 may be external to the processor 2100 and may be coupled to the graphics processor via a ring interconnect 2102, or some other interconnect bus or fabric. The display engine 2140 may include a 2D engine 2141 and a display controller 2143. The display engine 2140 may contain dedicated logic capable of operating independently of the 3D pipeline. The display controller 2143 may be coupled to a display device (not shown), which may be a system-integrated display device such as in a laptop computer or an external display device attached via a display device connector.

[0337] The geometry pipeline 2120 and the media pipeline 2130 can be configured to perform operations based on multiple graphics and media programming interfaces and are not dedicated to any one application programming interface (API). Driver software for the graphics processor can translate API calls dedicated to a particular graphics or media library into commands that can be processed by the graphics processor. Support can be provided for all of the Open Graphics Library (OpenGL), Open Computing Language (OpenCL), and / or Vulkan graphics and compute APIs from the Khronos Group. Support can also be provided for the Direct3D library from Microsoft Corporation. Combinations of these libraries can be supported. Support can also be provided for the Open Source Computer Vision Library (OpenCV). Future APIs with compatible 3D pipelines will also be supported if a mapping can be made from the pipelines of the future APIs to the pipelines of the graphics processor. Graphics Pipeline Programming

[0338] Figure 22A is a block diagram illustrating a graphics processor command format 2200 for programming a graphics processing pipeline, such as the pipeline described herein in connection with Figure 16A , Figure 17 , Figure 21 described pipelines. Figure 22B is a block diagram illustrating a graphics processor command sequence 2210 in accordance with an embodiment. Figure 22A The solid boxes in generally illustrate components that are included in a graphics command, while the dashed boxes include components that are optional or are only included in a subset of the graphics commands. Figure 22A The exemplary graphics processor command format 2200 of includes fields for a client 2202 that identifies the command, a command operation code (opcode) 2204, and a data field 2206. A sub-opcode 2205 and a command size 2208 are also included in some commands.

[0339] The client 2202 can specify the client units of the graphics device that process command data. The graphics processor command parser can check the client fields of each command to adjust the further processing of the command and route the command data to the appropriate client unit. The graphics processor client units can include a memory interface unit, a rendering unit, a 2D unit, a 3D unit, and a media unit. Each client unit can have a corresponding processing pipeline for processing commands. Once a command is received by a client unit, the client unit reads the opcode 2204 and the sub-opcode 2205 (if present) to determine the operation to be performed. The client unit uses the information in the data field 2206 to execute the command. For some commands, an explicit command size 2208 is expected to specify the size of the command. The command parser can automatically determine the size of at least some of the commands in the command based on the command opcode. The commands can be aligned via multiples of a double word. Other command formats can also be used.

[0340] Figure 22B The flowchart in FIG. shows an exemplary graphics processor command sequence 2210. The software or firmware of a data processing system featuring an exemplary graphics processor can use a certain version of the shown command sequence to establish, execute, and terminate a set of graphics operations. The sample command sequence is shown and described for illustrative purposes only, and the sample command sequence is not limited to these specific commands or this command sequence. Additionally, the commands can be issued in a batch in the command sequence such that the graphics processor will process the command sequence in at least a partially concurrent manner.

[0341] The graphics processor command sequence 2210 can start with a pipeline dump clear command 2212 to cause any active graphics pipeline to complete the current outstanding commands for the pipeline. Optionally, the 3D pipeline 2222 and the media pipeline 2224 may not operate concurrently. Performing the pipeline dump clear causes the active graphics pipeline to complete any outstanding commands. In response to the pipeline dump clear, the command parser for the graphics processor will pause command processing until the active drawing engine completes the outstanding operations and the associated read cache is invalidated. Optionally, any data marked "dirty" in the render cache can be dumped and cleared to memory. The pipeline dump clear command 2212 can be used for pipeline synchronization or can be used before putting the graphics processor in a low power state.

[0342] When the command sequence requires the graphics processor to explicitly switch between pipelines, a pipeline select command 2213 can be used. The pipeline select command 2213 may only be needed once in the execution context before issuing pipeline commands, unless the context is to issue commands for both pipelines. A pipeline dump clear command 2212 may be required immediately before the pipeline switch via the pipeline select command 2213.

[0343] The pipeline control command 2214 can be configured to operate the graphics pipeline and can be used to program the 3D pipeline 2222 and the media pipeline 2224. The pipeline control command 2214 can configure the pipeline state for the active pipeline. The pipeline control command 2214 can be used for pipeline synchronization and for clearing data from one or more cache memories within the active pipeline before processing a batch of commands.

[0344] Commands related to the return buffer state 2216 can be used to configure the set of return buffers for the corresponding pipeline for writing data. Some pipeline operations require the allocation, selection, or configuration of one or more return buffers into which intermediate data is written during processing. The graphics processor can also use one or more return buffers to store output data and perform cross-thread communication. The return buffer state 2216 can include the size and number of return buffers to select for the set of pipeline operations.

[0345] The remaining commands in the command sequence vary based on the active pipeline for the operation. Based on the pipeline determination 2220, the command sequence is customized for the 3D pipeline 2222 starting with the 3D pipeline state 2230 or the media pipeline 2224 starting at the media pipeline state 2240.

[0346] Commands for configuring the 3D pipeline state 2230 include 3D state setting commands for vertex buffer state, vertex element state, constant color state, depth buffer state, and other state variables that will be configured before processing 3D primitive commands. The values of these commands are determined at least in part based on the specific 3D API in use. The 3D pipeline state 2230 commands can also be able to selectively disable or bypass certain pipeline elements if they will not be used.

[0347] The 3D primitive 2232 commands can be used to submit 3D primitives to be processed by the 3D pipeline. The commands and associated parameters passed to the graphics processor via the 3D primitive 2232 commands are forwarded to the vertex fetch function in the graphics pipeline. The vertex fetch function uses the 3D primitive 2232 command data to generate vertex data structures. The vertex data structures are stored in one or more return buffers. The 3D primitive 2232 commands can be used to perform vertex operations on the 3D primitives via the vertex shader. To process the vertex shader, the 3D pipeline 2222 dispatches shader execution threads to the graphics processor execution units.

[0348] The 3D pipeline 2222 can be triggered by executing 2234 commands or events. A register can be written to trigger command execution. Execution can be triggered via "go" or "kick" commands in a command sequence. Command execution can use pipeline synchronization commands to trigger a command sequence dump clearance through the graphics pipeline. The 3D pipeline will perform geometric processing on 3D primitives. Once the operation is complete, the resulting geometric object is rasterized, and the pixel engine colors the resulting pixels. For those operations, additional commands for controlling pixel coloring and pixel backend operations can also be included.

[0349] When performing media operations, the graphics processor command sequence 2210 can follow the media pipeline 2224 path. Generally, the specific uses and ways of programming the media pipeline 2224 depend on the media or computing operations to be performed. During media decoding, specific media decoding operations can be migrated to the media pipeline. The media pipeline can also be bypassed, and media decoding can be performed entirely or partially using the resources provided by one or more general-purpose processing cores. The media pipeline can also include elements for general-purpose graphics processing unit (GPGPU) operations, where the graphics processor is used to perform SIMD vector operations using compute shader programs that are not explicitly related to the rendering of graphics primitives.

[0350] The media pipeline 2224 can be configured in a manner similar to the 3D pipeline 2222. The set of commands for configuring the media pipeline state 2240 is dispatched or placed in the command queue before the media object commands 2242. The commands for the media pipeline state 2240 can include data for configuring the media pipeline elements that will be used to process the media object. This includes data for configuring video decoding and video encoding logic within the media pipeline, such as encoding or decoding formats. The commands for the media pipeline state 2240 can also support the use of one or more pointers to "indirect" state elements that point to a bulk of state settings.

[0351] The media object commands 2242 can supply pointers to the media objects to be processed by the media pipeline. The media objects include memory buffers that contain the video data to be processed. Optionally, all media pipeline states must be valid before issuing the media object commands 2242. Once the pipeline state is configured and the media object commands 2242 are queued, the media pipeline 2224 is triggered via an execute command 2244 or an equivalent execution event (e.g., a register write). Subsequently, the output from the media pipeline 2224 can be post-processed by operations provided by the 3D pipeline 2222 or the media pipeline 2224. GPGPU operations can be configured and executed in a manner similar to media operations. Graphics Software Architecture

[0352] Figure 23 FIG. illustrates an exemplary graphics software architecture for a data processing system 2300. Such a software architecture may include a 3D graphics application 2310, an operating system 2320, and at least one processor 2330. The processor 2330 may include a graphics processor 2332 and one or more general-purpose processor cores 2334. The processor 2330 may be a variant of the processor 1402 or any other processor described herein. The processor 2330 may be used in place of the processor 1402 or any other processor described herein. Thus, the disclosure of any feature in connection with the processor 1402 or any other processor described herein also discloses the corresponding combination with the graphics processor 2330, but is not limited thereto. In addition, Figure 23 Elements having the same or similar names as elements in any other figure in this document describe the same elements as in the other figures, can operate or run in a similar manner as in the other figures, may include the same components, and may be linked to other entities, such as those described elsewhere in this document, but are not limited thereto. The graphics application 2310 and the operating system 2320 each execute in the system memory 2350 of the data processing system.

[0353] The 3D graphics application 2310 may include one or more shader programs, the one or more shader programs including shader instructions 2312. The shader language instructions may be in a high-level shader language, such as the High-Level Shader Language (HLSL) of Direct3D, the OpenGL Shader Language (GLSL), and so on. The application may also include executable instructions 2314 in machine language suitable for execution by the general-purpose processor core 2334. The application may also include graphics objects 2316 defined by vertex data.

[0354] The operating system 2320 may be from Microsoft Corporation An operating system, a proprietary UNIX-like operating system, or an open-source UNIX-like operating system using a Linux kernel variant. The operating system 2320 may support a graphics API 2322, such as, a Direct3D API, an OpenGL API, or a Vulkan API. When the Direct3D API is in use, the operating system 2320 uses a front-end shader compiler 2324 to compile any shader instructions 2312 in HLSL into a lower-level shader language. The compilation can be just-in-time (JIT) compilation or application executable shader pre-compilation. During the compilation of a 3D graphics application 2310, high-level shaders can be compiled into low-level shaders. The shader instructions 2312 can be provided in an intermediate form, such as a version of the Standard Portable Intermediate Representation (SPIR) used by the Vulkan API.

[0355] The user-mode graphics driver 2326 may include a backend shader compiler 2327 to compile the shader instructions 2312 into a hardware-specific representation. When the OpenGL API is in use, the shader instructions 2312 in the high-level GLSL language are passed to the user-mode graphics driver 2326 for compilation. The user-mode graphics driver 2326 may use the operating system kernel-mode functionality 2328 to communicate with the kernel-mode graphics driver 2329. The kernel-mode graphics driver 2329 may communicate with the graphics processor 2332 to dispatch commands and instructions. IP Core Implementation Method

[0356] One or more aspects may be implemented by representative code stored on a machine-readable medium that represents and / or defines logic within an integrated circuit, such as a processor. For example, the machine-readable medium may include instructions that represent various logics within the processor. When read by a machine, the instructions may cause the machine to fabricate logic for performing the techniques described herein. Such representations (referred to as "IP cores") are reusable units of the logic of an integrated circuit, and these reusable units can be stored as a hardware model that describes the structure of the integrated circuit on a tangible, machine-readable medium. The hardware model can be supplied to various customers or manufacturing facilities that load the hardware model on manufacturing machines for fabricating integrated circuits. Integrated circuits can be fabricated such that the circuits perform the operations described in association with any of the embodiments described herein.

[0357] Figure 24AFIG. is a block diagram of an IP core development system 2400 that can be used to fabricate integrated circuits to perform operations. The IP core development system 2400 can be used to generate modular, reusable designs that can be incorporated into larger designs or used to build an entire integrated circuit (e.g., a SOC integrated circuit). A design facility 2430 can generate a software simulation 2410 of an IP core design in a high-level programming language (e.g., C / C++). The software simulation 2410 can be used to design, test, and verify the behavior of the IP core using a simulation model 2412. The simulation model 2412 can include functional simulation, behavioral simulation, and / or timing simulation. Subsequently, a register transfer level (RTL) design 2415 can be created or synthesized from the simulation model 2412. The RTL design 2415 is an abstraction of the behavior of an integrated circuit that models the flow of digital signals between hardware registers (including the associated logic that is executed using the modeled digital signals). In addition to the RTL design 2415, lower-level designs at the logic level or transistor level can also be created, designed, or synthesized. Thus, the specific details of the initial design and simulation can vary.

[0358] The RTL design 2415 or an equivalent can be further synthesized by the design facility into a hardware model 2420, which can be in a hardware description language (HDL) or some other representation of physical design data. The HDL can be further simulated or tested to verify the IP core design. A non-volatile memory 2440 (e.g., a hard disk, flash memory, or any non-volatile storage medium) can be used to store the IP core design for delivery to a third-party manufacturing facility 2465. Alternatively, the IP core design can be transmitted via a wired connection 2450 or a wireless connection 2460 (e.g., via the Internet). The manufacturing facility 2465 can then fabricate an integrated circuit that is at least partially based on the IP core design. The fabricated integrated circuit can be configured to perform operations in accordance with at least one embodiment described herein.

[0359] Figure 24BCross-sectional side view of an illustrated integrated circuit package component 2470. The integrated circuit package component 2470 illustrates an implementation of one or more processor or accelerator devices as described herein. The package component 2470 includes a plurality of hardware logic units 2472, 2474 connected to a substrate 2480. The logic 2472, 2474 may be implemented at least in part in configurable logic or fixed-function logic hardware and may include one or more portions of any of the (one or more) processor cores, (one or more) graphics processors, or other accelerator devices described herein. Each logic unit 2472, 2474 may be implemented within a semiconductor die and is coupled to the substrate 2480 via an interconnect structure 2473. The interconnect structure 2473 may be configured to route electrical signals between the logic 2472, 2474 and the substrate 2480 and may include interconnects such as, but not limited to, bumps or pillars. The interconnect structure 2473 may be configured to route electrical signals such as, for example, input / output (I / O) signals and / or power or ground signals associated with the operation of the logic 2472, 2474. Optionally, the substrate 2480 may be an epoxy-based laminated substrate. The substrate 2480 may also include other suitable types of substrates. The package component 2470 may be connected to other electrical devices via a package interconnect 2483. The package interconnect 2483 may be coupled to the surface of the substrate 2480 to route electrical signals to other electrical devices such as a motherboard, other chipset, or multi-chip module.

[0360] The logic units 2472, 2474 may be electrically coupled to a bridge 2482 that is configured to route electrical signals between the logic 2472 and the logic 2474. The bridge 2482 may be a dense interconnect structure that provides routing for electrical signals. The bridge 2482 may include a bridge substrate made of glass or a suitable semiconductor material. Circuitry features may be formed on the bridge substrate to provide a chip-to-chip connection between the logic 2472 and the logic 2474.

[0361] Although two logic units 2472, 2474 and a bridge 2482 are illustrated, embodiments described herein may include more or fewer logic units on one or more dies. The one or more dies may be connected by zero or more bridges since the bridge 2482 may be excluded when the logic is included on a single die. Alternatively, multiple dies or logic units may be connected by one or more bridges. Additionally, multiple logic units, dies, and bridges may be connected together in other possible configurations including three-dimensional configurations.

[0362] Figure 24CIllustrated is a packaged component 2490 that includes hardware logic die that connect to multiple cells of a substrate 2480 (e.g., a base die). Graphics processing units, parallel processors, and / or compute accelerators as described herein can be composed of various silicon die fabricated separately. In this context, a die is an integrated circuit that is at least partially encapsulated and that includes different logic units that can be assembled with other die into a larger package. Die with various sets of different IP core logic can be assembled into a single device. Additionally, die can be integrated into a base die or base die using active interposer technology. The concepts described herein enable interconnection and communication between different forms of IP within a GPU. IP cores can be fabricated using different process technologies and configured during fabrication, which avoids the complexity of converging multiple IPs into the same manufacturing process, especially for large SoCs with several styles of IP. Allowing the use of multiple process technologies improves time to market and provides a cost-effective way to create multiple product SKUs. Additionally, the decomposed IP is more easily modified to be independently power gated, and components not in use for a given workload can be turned off, thus reducing overall power consumption.

[0363] In embodiments, the packaged component 2490 can include fewer or more components and die interconnected by a structure 2485 or one or more bridges 2487. The die within the packaged component 2490 can have a 2.5D arrangement using a Chip-on-Wafer-on-Substrate stack, where multiple die are stacked side-by-side on a silicon interposer that includes through-silicon vias (TSV) to couple the die to the substrate 2480, which includes electrical connections to the package interconnect 2483.

[0364] In one embodiment, the silicon interposer is the active interposer 2489, which includes embedded logic in addition to the TSVs. In such embodiments, the die within the package assembly 2490 are arranged on top of the active interposer 2489 using 3D face-to-face die stacking. The active interposer 2489 may further include hardware logic for I / O 2491, cache memory 2492, and other hardware logic 2493 in addition to the interconnect structure 2485 and the silicon bridge 2487. The structure 2485 enables communication between the various logic dies 2472, 2474 and the logic 2491, 2493 within the active interposer 2489. The structure 2485 can be a NoC interconnect or another form of packet-switched fabric that exchanges data packets between the components of the package assembly. For complex components, the structure 2485 can be a dedicated die that enables communication between the various hardware logics of the package assembly 2490.

[0365] The bridge structure 2487 within the active interposer 2489 can be used to facilitate point-to-point interconnects, for example, between a logic or I / O die 2474 and a memory die 2475. In some implementations, the bridge structure 2487 can also be embedded within the substrate 2480.

[0366] The hardware logic dies can include dedicated hardware logic dies 2472, logic or I / O dies 2474, and / or memory dies 2475. The hardware logic die 2472 and the logic or I / O die 2474 can be implemented at least in part in configurable logic or fixed-function logic hardware and can include one or more portions of any of the (one or more) processor cores, (one or more) graphics processors, parallel processors, or other accelerator devices described herein. The memory die 2475 can be DRAM (e.g., GDDR, HBM) memory or cache (SRAM) memory. The cache memory 2492 within the active interposer 2489 (or the substrate 2480) can act as a global cache for the package assembly 2490, act as part of a distributed global cache, or act as a dedicated cache for the structure 2485.

[0367] Each die can be fabricated as a separate semiconductor die and can be coupled to a base die that is embedded within or coupled to a substrate 2480. The coupling to the substrate 2480 can be performed via an interconnect structure 2473. The interconnect structure 2473 can be configured to route electrical signals between various dies and logic within the substrate 2480. The interconnect structure 2473 can include interconnects such as, but not limited to, bumps or pillars. In some embodiments, the interconnect structure 2473 can be configured to route electrical signals such as, for example, input / output (I / O) signals and / or power or ground signals associated with the operation of logic, I / O, and memory dies. In one embodiment, an additional interconnect structure couples an active interposer 2489 to the substrate 2480.

[0368] The substrate 2480 can be an epoxy-based laminated substrate, however, it is not limited thereto, and the substrate 2480 can also include other suitable types of substrates. The packaged component 2490 can be connected to other electrical devices via a package interconnect 2483. The package interconnect 2483 can be coupled to the surface of the substrate 2480 to route electrical signals to other electrical devices such as, for example, a motherboard, other chip sets, or a multi-chip module.

[0369] The logic or I / O die 2474 and the memory die 2475 can be electrically coupled via a bridge 2487 that is configured to route electrical signals between the logic or I / O die 2474 and the memory die 2475. The bridge 2487 can be a dense interconnect structure that provides routing for electrical signals. The bridge 2487 can include a bridge substrate formed of glass or a suitable semiconductor material. Circuitry features can be formed on the bridge substrate to provide die-to-die connections between the logic or I / O die 2474 and the memory die 2475. The bridge 2487 can also be referred to as a silicon bridge or an interconnect bridge. For example, the bridge 2487 is an Embedded Multi-die Interconnect Bridge (EMIB). Alternatively, the bridge 2487 can simply be a direct connection from one die to another die.

[0370] Figure 24DFIG. illustrates a packaged component 2494 including an interchangeable die 2495 according to an embodiment. The interchangeable die 2495 can be assembled into standardized sockets on one or more base dies 2496, 2498. The base dies 2496, 2498 can be coupled via a bridge interconnect 2497, which can be similar to other bridge interconnects described herein and can be, for example, EMIB. Memory dies can also be connected to logic or I / O dies via a bridge interconnect. The I / O and logic dies can communicate via an interconnect structure. Each of the base dies can support one or more sockets in a standardized format for one of logic or I / O or memory / cache.

[0371] SRAM and power delivery circuitry can be fabricated into one or more of the base dies 2496, 2498, which can be fabricated using a different process technology relative to the interchangeable die 2495, with the interchangeable die 2495 stacked on top of the base die. For example, the base dies 2496, 2498 can be fabricated using a larger process technology while the interchangeable die can be fabricated using a smaller process technology. One or more of the interchangeable dies 2495 can be memory (e.g., DRAM) dies. Different memory densities can be selected for the packaged component 2494 based on the power and / or performance for the product using the packaged component 2494. Additionally, logic dies with different numbers of types of functional units can be selected based on the power and / or performance for the product at assembly. Additionally, dies containing different types of IP logic cores can be inserted into the interchangeable die sockets, enabling a hybrid processor design that can mix and match IP blocks of different technologies. Exemplary System - on - Chip Integrated Circuit

[0372] Figures 25 - 26B FIG. illustrates an exemplary integrated circuit and associated graphics processor that can be fabricated using one or more IP cores. In addition to what is illustrated, other logic and circuitry can be included, including additional graphics processors / cores, peripheral interface controllers, or general processor cores. Figures 25 - 26B Elements with the same or similar names as elements in any other figure herein describe the same elements as in other figures, can operate or function in a similar manner as in other figures, can include the same components, and can be linked to other entities, such as those described elsewhere herein, but not limited thereto.

[0373] Figure 25FIG. is a block diagram of an exemplary system-on-chip integrated circuit 2500 that can be fabricated using one or more IP cores. The exemplary integrated circuit 2500 includes one or more application processors 2505 (e.g., CPUs), at least one graphics processor 2510, and the at least one graphics processor 2510 can be a variant of the graphics processors 1408, 1508, 2510, or can be any graphics processor described herein and used in place of any of the described graphics processors. Thus, the disclosure of any feature in connection with a graphics processor herein also discloses the corresponding combination with the graphics processor 2510, but is not limited thereto. The integrated circuit 2500 can additionally include an image processor 2515 and / or a video processor 2520, and either the image processor 2515 or the video processor 2520 can be a modular IP core from the same design facility or multiple different design facilities. The integrated circuit 2500 can include peripheral or bus logic, including a USB controller 2525, a UART controller 2530, an SPI / SDIO controller 2535, and an 2 S / I 2 C controller 2540. Additionally, the integrated circuit can include a display device 2545 that is coupled to one or more of a high-definition multimedia interface (HDMI) controller 2550 and a mobile industry processor interface (MIPI) display interface 2555. Storage can be provided by a flash memory subsystem 2560 (including flash memory and a flash memory controller). A memory interface can be provided via a memory controller 2565 to obtain access to SDRAM or SRAM memory devices. Some integrated circuits additionally include an embedded security engine 2570.

[0374] Figures 26A - 26B FIG. is a block diagram of an exemplary graphics processor for use within a SoC in accordance with an embodiment described herein. The illustrated graphics processor can be a variant of the graphics processors 1408, 1508, 2510, or any other graphics processor described herein. The graphics processor can be used in place of the graphics processors 1408, 1508, 2510, or any other graphics processor described herein. Thus, the disclosure of any feature in connection with the graphics processors 1408, 1508, 2510, or any other graphics processor described herein also discloses the corresponding combination with Figures 26A - 26B the graphics processor, but is not limited thereto. Figure 26A FIG. illustrates an exemplary graphics processor 2610 of a system-on-chip integrated circuit that can be fabricated using one or more IP cores in accordance with an embodiment.Figure 26B FIG. additional exemplary graphics processor 2640 of a system-on-chip integrated circuit that can be fabricated using one or more IP cores, according to an embodiment. Figure 26A The graphics processor 2610 is an example of a low-power graphics processor core. Figure 26B The graphics processor 2640 is an example of a higher-performance graphics processor core. For example, as mentioned at the beginning of this paragraph, each of the graphics processors 2610 and 2640 can be Figure 25 a variant of the graphics processor 2510.

[0375] As Figure 26A shown, the graphics processor 2610 includes a vertex processor 2605 and one or more fragment processors 2615A-2615N (e.g., 2615A, 2615B, 2615C, 2615D, up to 2615N-1 and 2615N). The graphics processor 2610 can execute different shader programs via separate logic, such that the vertex processor 2605 is optimized to perform operations for vertex shader programs, while the one or more fragment processors 2615A-2615N perform fragment (e.g., pixel) shading operations for fragment or pixel shader programs. The vertex processor 2605 executes the vertex processing stage of the 3D graphics pipeline and generates primitives and vertex data. The (one or more) fragment processors 2615A-2615N use the primitive data and vertex data generated by the vertex processor 2605 to produce a frame buffer that is displayed on a display device. The (one or more) fragment processors 2615A-2615N can be optimized to execute fragment shader programs as provided in the OpenGL API, which can be used to perform operations similar to those of pixel shader programs as provided in the Direct 3D API.

[0376] The graphics processor 2610 additionally includes one or more memory management units (MMUs) 2620A - 2620B, (one or more) caches 2625A - 2625B, and (one or more) circuit interconnects 2630A - 2630B. The one or more MMUs 2620A - 2620B provide virtual - to - physical address mapping for the graphics processor 2610 (including for the vertex processor 2605 and / or (one or more) fragment processors 2615A - 2615N), and this virtual - to - physical address mapping can reference vertex data or image / texture data stored in memory in addition to vertex data or image / texture data stored in one or more caches 2625A - 2625B. The one or more MMUs 2620A - 2620B can be synchronized with other MMUs within the system such that each processor 2505 - 2520 can participate in a shared or unified virtual memory system, and the other MMUs within the system include one or more MMUs associated with Figure 25 one or more application processors 2505, image processors 2515, and / or video processors 2520. The components of the graphics processor 2610 can correspond to the components of other graphics processors described herein. The one or more MMUs 2620A - 2620B can correspond to the MMU 245 of Figure 2C . The vertex processor 2605 and fragment processors 2615A - 2515N can correspond to the graphics multiprocessor 234. According to an embodiment, the one or more circuit interconnects 2630A - 2630B enable the graphics processor 2610 to interface with other IP cores within the SoC via the internal bus of the SoC or via a direct connection. The one or more circuit interconnects 2630A - 2630B can correspond to the data cross - switch 240 of Figure 2C . Further correspondences can be found between the similar components of the graphics processor 2610 and the various graphics processor architectures described herein.

[0377] As Figure 26B shown, the graphics processor 2640 includes Figure 26AOne or more MMUs 2620A - 2620B, caches 2625A - 2625B, and circuit interconnections 2630A - 2630B of the graphics processor 2610. The graphics processor 2640 includes one or more shader cores 2655A - 2655N (e.g., 2655A, 2655B, 26555C, 2655D, 2655E, 2655F, up to 2655N - 1 and 2655N), which provide a unified shader core architecture where a single core or type of core can execute all types of programmable shader code, including shader program code for implementing vertex shaders, fragment shaders, and / or compute shaders. The exact number of shader cores present can vary depending on the embodiment and implementation. Additionally, the graphics processor 2640 includes an inter - core task manager 2645, which acts as a thread dispatcher for dispatching execution threads to one or more shader cores 2655A - 2655N and a tiling unit 2658 for accelerating tiling operations for tile - based rendering, in which rendering operations for a scene are subdivided in image space, e.g., to take advantage of local spatial coherence within the scene or optimize the use of internal caches. The shader cores 2655A - 2655N can, for example, correspond to Figure 2D the graphics multiprocessor 234 in Figure 3A and Figure 3B the graphics multiprocessors 325, 350 respectively, or correspond to Figure 3C the multi - core group 365A in Data Processing System for GPGPU Program Execution

[0378] Figure 27 is a block diagram of a data processing system 2700 according to an embodiment. The data processing system 2700 is a heterogeneous processing system having a processor 2702, unified memory 2710, and a GPGPU 2720 including machine - learning acceleration logic. The processor 2702 and the GPGPU 2720 can be any of the processors and GPGPUs / parallel processors as described herein. For example, with further reference to Figure 1 , the application processor 2702 can be a variant of the processors in one or more of the illustrated processors 102 and / or share an architecture with the processors in one or more of the illustrated processors 102. The GPGPU 2720 can be a variant of the parallel processors in one or more of the illustrated parallel processors 112 and / or share an architecture with the parallel processors in one or more of the illustrated parallel processors 112. With further reference to Figure 14, the processor 2702 can be a variant of one of the illustrated processors 1402 and / or share an architecture with one of the illustrated processors 1402, and the GPGPU 2720 can be a variant of one of the illustrated graphics processors 1408 and / or share an architecture with one of the illustrated graphics processors 1408.

[0379] The application processor 2702 can execute instructions for the compiler 2715 stored in the system memory 2712. In one embodiment, the compiler 2715 executes on the application processor 2702 to compile the source code 2714A into the compiled code 2714B. The compiled code 2714B can include instructions executable by the application processor 2702 and / or instructions executable by the GPGPU 2720. The compilation of the instructions to be executed by the GPGPU can be facilitated by a shader or compute program compiler (such as Figure 23 the shader compiler 2327 and / or the shader compiler 2324 in). In one embodiment, the compilation of the instructions to be executed by the GPGPU can alternatively or additionally be performed at least in part via execution logic within the GPGPU 2720. For example, the compiler 2715 can migrate certain program code analysis, translation, or compilation operations to the GPGPU 2720.

[0380] During compilation, the compiler 2715 can perform operations for inserting metadata that includes hints about the level of data parallelism present in the compiled code 2714B and / or hints about the data locality associated with the threads to be dispatched based on the compiled code 2714B. The compiler 2715 can include the information necessary for performing such operations, or these operations can be performed with the assistance of the runtime library 2716. The runtime library 2716 can also assist the compiler 2715 in compiling the source code 2714A and can also include instructions that are linked with the compiled code 2714B at runtime to facilitate the execution of the compiled instructions on the GPGPU 2720. The compiler 2715 can also facilitate register allocation of variables via a register allocator (RA) and generate load and store instructions for moving data for the variable between memory and the registers assigned to the variable.

[0381] The unified memory 2710 represents a unified address space that can be accessed by the processor 2702 and the GPGPU 2720. The unified memory may include a system memory 2712 and a GPGPU memory 2718. The GPGPU memory 2718 is memory within the address space of the GPGPU 2720 and may include some or all of the system memory 2712. In one embodiment, the compiled code 2714B stored in the system memory 2712 may be mapped into the GPGPU memory 2718 for access by the GPGPU 2720. The GPGPU memory 2718 also includes the GPGPU local memory 2728 of the GPGPU 2720. The GPGPU local memory 2728 may include, for example, HBM or GDDR memory.

[0382] The GPGPU 2720 includes a plurality of compute blocks 2724A - 2724N, which may include one or more of the various processing resources described herein. The processing resources may be or may include various different computing resources such as, for example, execution units, compute units, stream multiprocessors, graphics multiprocessors, or multi - core groups. In one embodiment, the GPGPU 2720 additionally includes a tensor accelerator 2723 (e.g., a matrix accelerator), which may include one or more dedicated compute units designed to accelerate a subset of matrix operations (e.g., dot products, etc.). Some functions to be executed by the compute blocks 2724A - 2724N may be directly scheduled or migrated to the tensor accelerator 2723. In embodiments, the tensor accelerator 2723 includes processing element logic configured to efficiently execute matrix computation operations such as multiplication and addition operations and dot product operations used by 3D graphics or compute shader programs. In one embodiment, the tensor accelerator 2723 may be configured to accelerate operations used by a machine learning framework. In one embodiment, the tensor accelerator 2723 is an application - specific integrated circuit explicitly configured to perform a particular set of parallel matrix multiplication and / or addition operations. In one embodiment, the tensor accelerator 2723 is a field programmable gate array (FPGA) that provides fixed - function logic that can be updated between workloads. In one embodiment, the set of compute operations executable by the tensor accelerator 2723 may be limited relative to the operations executable by the compute blocks 2724A - 2724N. However, the tensor accelerator 2723 can execute parallel tensor operations with significantly higher throughput relative to the compute blocks 2724A - 2724N. The tensor accelerator 2723 may also be referred to as a tensor accelerator or a tensor core. In one embodiment, the logic components within the tensor accelerator 2723 may be distributed across the processing resources of the plurality of compute blocks 2724A - 2724N rather than being centrally located within a single circuit.

[0383] The GPGPU 2720 may further include a set of resources that can be shared by the compute blocks 2724A - 2724N and the tensor accelerator 2723. The set of resources includes, but is not limited to, a set of registers 2725, a power and performance module 2726, and a cache 2727. In one embodiment, the registers 2725 include directly accessible and indirectly accessible registers, where the indirectly accessible registers are optimized for use by the tensor accelerator 2723. The power and performance module 2726 may be configured to adjust the power delivery and clock frequency for the compute blocks 2724A - 2724N to power gate idle components within the compute blocks 2724A - 2724N. In various embodiments, the cache 2727 may include an instruction cache and / or a lower - level data cache.

[0384] The GPGPU 2720 may additionally include an L3 data cache 2730, which may be used to cache data accessed from the unified memory 2710 by the tensor accelerator 2723 and / or compute elements within the compute blocks 2724A - 2724N. In one embodiment, the L3 data cache 2730 includes a shared local memory 2732, which can be shared by the compute elements within the compute blocks 2724A - 2724N and the tensor accelerator 2723.

[0385] In one embodiment, the GPGPU 2720 includes instruction handling logic, such as a fetch and decode unit 2721 and a scheduler controller 2722. The fetch and decode unit 2721 includes a fetch unit and a decode unit for fetching instructions and decoding the instructions for execution by one or more of the compute blocks 2724A - 2724N or by the tensor accelerator 2723. The instructions may be scheduled via the scheduler controller 2722 to appropriate functional units within the compute blocks 2724A - 2724N or the tensor accelerator. In one embodiment, the scheduler controller 2722 is an ASIC configured to perform high - level scheduling operations. In one embodiment, the scheduler controller 2722 is a microcontroller or a low - per - instruction - energy - consuming processor configured to execute scheduling instructions loaded from a firmware module.

[0386] Figure 28 FIG. illustrates a processing resource architecture 2800 according to an embodiment. The processing resource architecture 2800 includes functional units that can be found in a graphics core (such as Figure 18A any of the graphics cores 1815A - 1815N of the graphics core), including Figure 18B the vector engine 1802 and Figure 18CComponents of the matrix engine 1803. In one embodiment, a graphics core (e.g., graphics core 1815A) may include multiple instances of processing resources having a processing resource architecture 2800. A local thread dispatcher (TDL 2801) may dispatch instructions to a set 2802 of ready instruction queues associated with hardware threads (T0-Tn) for each processing resource. Threads may be distributed to processing resources in a round-robin manner. If the instructions for a thread are not ready for execution, the dispatch of that thread may be skipped. If a processing resource does not have sufficient resources to accept an incoming thread, TDL 2801 may skip that processing resource.

[0387] Each hardware thread (T0-Tn) is a SIMD thread that can execute a single instruction on multiple channels of data. Data for instructions may be read from registers in the register file 2810 and written to registers in the register file 2810. Intermediate data for computations may be stored in an array of accumulators 2811. In embodiments, the register file 2810 may be configured to have the characteristics of other GPU register files described herein (e.g., register file 258, register files 334A-334B, register file 369, register 445, etc.). The register file 2810 may include vector registers (e.g., vector register 1561) and, in one embodiment, a limited number of scalar registers (scalar register 1562). In one embodiment, the register file 2810 may include a general register file array, such as GRF 1824. In one embodiment, the number of hardware threads is configurable, where a configurable number of registers within the register file 2810 are allocated to each thread. In one embodiment, SIMT execution is enabled by assigning multiple SIMT threads to a single SIMD hardware thread. Data elements associated with multiple SIMT threads may be mapped to multiple sub-registers within a single SIMD register file. The thread control unit 2832 executes via hardware thread control instructions. The thread control unit 2832 includes a thread arbiter 2804 that arbitrates which of the ready instructions are assigned to instruction queues 2806 associated with floating point (FPU), integer (INT), or matrix pipelines (Matrix).

[0388] The source read arbiter 2808 arbitrates access to the register file and accumulators for the source operand read phase of instruction execution for instructions within the instruction queue 2806. Operands read from the register file 2810 may be buffered within a reuse buffer 2812 that is coupled to the functional units of the FPU, INT, and matrix pipelines. The reuse buffer 2812 can act as an operand cache that facilitates reuse of operand data by those functional units. In one embodiment, the reuse buffer 2812 is multi-level deep, allowing caching of source operands from multiple instructions. The processing resource architecture 2800 includes a variety of different types of functional units, including a matrix unit 2814, a single-precision floating-point (FP32) unit 2816, a high-throughput double-precision floating-point (FP64) 2817, a low-throughput FP64 unit 2819, and an integer unit 2820. The high-throughput FP64 2817 unit can be configured to execute instructions intended for use by high-performance compute (HPC) and scientific computing workloads and can be found primarily in graphics processors and compute accelerators designed for servers and / or workstations. The low-throughput FP64 unit 2819 can be found primarily in consumer-grade graphics processors and / or compute accelerators. At least one embodiment can include both the high-throughput FP64 2817 unit and the low-throughput FP64 unit 2819. The processing resource architecture may also include an ARF 2818 that contains configuration registers, architectural registers, thread-specific data, and other non-general-purpose registers. The ARF 2818 can be similar to Figure 18B the ARF 1826.

[0389] Based on the instruction type, different instructions are assigned to different functional units. Matrix operations such as dot product or matrix multiplication operations are assigned to the matrix unit 2814, which can be Figure 18C an instance of the matrix engine 1803, or another matrix unit described herein (such as Figure 3Cone or more Tensor Cores 371). Single-precision floating-point instructions may be executed by the FP32 units 2816 or the high-throughput FP64 units 2817. Double-precision floating-point instructions may be executed by the high-throughput FP64 units 2817 or the low-throughput FP64 units 2819. Integer operations are executed by the integer units 2820. In one embodiment, the FP32 units 2816 and the high-throughput FP64 units 2817 are grouped into a first ALU / FPU (e.g., FPU0), while the low-throughput FP64 units 2819 and the integer units 2820 are grouped into a second ALU / FPU (e.g., FPU1). Outputs from these functional units may be written back to the general-purpose registers in the register file 2810 or stored in one or more accumulators 2811 via the destination write arbiter 2822.

[0390] In one embodiment, the processing resource architecture 2800 includes a Message Execution Unit (MEU 2815). In response to requests from functional units within the processing resource architecture 2800, the MEU 2815 generates messages and transmits the messages to shared functional units associated with the processing resource having the architecture 2800. Such shared functional units include, for example, a ray tracing unit, a sampler, and other fixed-function logic, such as Figure 18A the ray tracing unit 1808A, the sampler 1810A, and the fixed-function logic 1812A associated with the graphics core 1815A of. The MEU 2815 may also generate messages and transmit the messages to a load / store unit, such as the memory load / store unit 1804A of the graphics core 1815A and / or Figure 28 the load / store circuitry 2821 of. Messages to the load / store unit are used to perform load and store operations between the register file 2810 and memory (e.g., via the data cache / shared local memory 1806A).

[0391] In one embodiment, the processing resource architecture 2800 includes a source crossbar 2830 coupled to the integer units 2820. The source crossbar implements shuffling and reordering of incoming data according to the desired data pattern. In one embodiment, the source crossbar 2830 is coupled only to the Src0 (source 0) input. In other embodiments, the source crossbar 2830 may be coupled to other source inputs (e.g., Src1 (source 1), Src2 (source 2)). For use with the transform instructions described herein, only the Src0 input is required. However, other transform instructions that utilize multiple source inputs are contemplated.

[0392] Source data elements can be read from registers within a register file and stored in a reuse buffer associated with the integer unit 2820. Subsequently, before the data elements are provided as input to the integer unit 2820, the packed mode of the source data elements can be reconstructed via the source crossbar 2830. The manner in which data elements are stored into the registers of the register file is referred to as the regioning mode of the registers. Register regioning involves partitioning SIMD registers into smaller regions, each region being capable of holding a subset of the data elements. In one embodiment, the register regions can be 1D or 2D (e.g., 1x16, 2x8, 4x4, etc.). Instructions executed by the processing resources can specify registers and sub-registers as well as regioning pointers that specify the vertical and horizontal spans of the regions and the number of data elements per row in the regions. The vertical span indicates the number of data elements to skip to reach the next row in a multi-row region. The width specifies the number of data elements per row in the region. The horizontal span indicates which elements in the selected row are to be used as input. For example, a horizontal span of one indicates that each element is selected. A horizontal span of two indicates that every other element is selected. GPU Thread Dispatch Hardware

[0393] Figure 29 FIG. 6 is a block diagram of a system 2900 including a GPGPU device 2902 according to an embodiment. The GPGPU device 2902 includes a graphics engine 2908 and a plurality of compute engines 2910. The graphics engine 2908 can process a command list or command buffer of instructions received from a graphics driver. The graphics engine 2908 performs graphics operations in response to these commands. A render command stream converter (RCS 2916) associated with the graphics engine 2908 can stream render commands for performing shader operations. A thread dispatcher 2920 dispatches threads for these shader operations to processing resources within compute blocks 2720A - 2720N. The dispatched threads are executed via hardware threads within the processing resources. Similarly, a collection of compute engines 2910, which operate asynchronously with each other, and the graphics engine 2908 can dispatch commands for compute shaders and / or GPGPU programs via a plurality of compute command stream converters (CCSn 2918). The thread dispatcher 2920 dispatches threads to processing resources within compute blocks 2720A - 2720N to execute compute shaders and / or GPGPU programs. The dispatched threads are executed via hardware threads within the processing resources.

[0394] In one embodiment, the thread dispatcher 2920 includes a compute walker circuit system 2922 to facilitate the execution of compute programs on the graphics processor hardware. In one embodiment, the processing device may be segmented into multiple parts, the multiple parts including one or more compute blocks 2720A - 2720N and / or one or more parts including a single compute block. Each part may be assigned to process commands from a specific source. In the case where a single source (e.g., CCS1 in CCSn 2918) provides commands, the entire processing device (all segments) may be assigned to process commands from that single source. When there is a second source to provide commands (e.g., CCS1 and CCS2), some segments may be assigned to execute commands from the first source, and other segments may be assigned to execute commands from the second source. Thus, commands from multiple applications may be executed simultaneously by the processing unit.

[0395] In the embodiments described herein, the number of active hardware threads within the processing resources of the compute blocks 2720A - 2720N is configurable. The number of registers assigned to a given hardware thread is also configurable. In one embodiment, this configuration is performed via a variable register per thread (VRT) configuration 2906 within a non - pipelined state defined within the GPU device 2902. The non - pipelined state 2904 defines state attributes that are globally applied to hardware resources (such as the processing resources within the graphics core that are globally applied to the compute blocks 2720A - 2720N). Within the non - pipelined state 2904 applied to the processing resources, the VRT configuration 2906 is defined per shader stage (e.g., shader type) and has enable bits to facilitate backward compatibility. For example, in one embodiment, the VRT configuration 2906 may include separate configurations for vertex shaders, tessellation shaders, geometry shaders, fragment / pixel shaders, mesh shaders, compute shaders, etc. The VRT configuration 2906 is delivered by the thread dispatcher 2920 associated with the dispatching of threads to the compute blocks 2720A - 2720N for execution.

[0396] Figure 30 is an illustration of a system 3000 for dispatching a thread group to processing resources 3020 according to an embodiment. In one embodiment, the thread group dispatching is performed using a hierarchy of hardware units that handle the dispatching of the thread group. The command stream converter 2918 (e.g., Figure 29The CCSn 2918) provides thread group information to the global compute front end (CFEG 3010), and these command stream converters are shown as CCS0 - CCS3. The thread group information includes information for one or more compute kernels to be executed by the graphics processor. The CFEG 3010 then dispatches the thread group to the connected CFE 3015 (shown as CFE0–CFE7). Each CFE further dispatches the received thread group to the processing resources 3020. In one embodiment, the processing resources 3020 are divided into multiple slices 3017 (shown as slice0 - slice15), where each CFE in the CFE 3015 is coupled to at least two slices among the multiple slices 3017. Each slice among the multiple slices 3017 is an individually assignable unit of the processing resources 3020. In one embodiment, each slice can be associated with an accelerator integrated slice 490 such as Figure 4D to enable the slice to be presented as a virtualizable sub-device. In one embodiment, each slice can be dedicated to a virtual machine or a virtualized graphics execution environment, such as a container. Each slice among the multiple slices 3017 can include one or more sub-slices, where each sub-slice includes one or more sets of individual processing resources and / or execution units. For example, the individual processing resources and / or execution units can have Figure 28 the processing resource architecture 2800. The processing resources and / or execution units can also have an architecture associated with, for example, Figure 2D the graphics multiprocessor 234, the compute unit 1560A, or variants thereof.

[0397] The specific hardware composition of each slice can vary across embodiments. Referring to Figure 18A , in one embodiment, a slice can include one or more graphics cores, such as one or more of the graphics cores 1815A - 1815N. A slice can also include a graphics core block 1715, and the graphics core block 1715 includes all the associated graphics cores 1815A - 1815N. In one embodiment, a slice can include a graphics core cluster 1714. In such embodiments, a sub-slice can include a group of individual processing resources and / or execution units, and each group includes multiple vector engines 1802A - 1802N and matrix engines 1803A...

Claims

1. A graphics processor, comprising: A memory interface; A plurality of graphics cores, the plurality of graphics cores being coupled via a data interconnect, each graphics core of the plurality of graphics cores including a plurality of processing resources; And Circuitry for managing the execution of a workload by the plurality of graphics cores, the circuitry being configured to: Receive a command for intermediate thread preemption of threads for execution by a first plurality of processing resources of a graphics core; Save the execution state of a first set of threads of the first plurality of processing resources, the first set of threads including threads in an active state; Save the thread state of a second set of threads of the first plurality of processing resources, the second set of threads including threads in a non-active state; And After the execution state of the first set of threads and the thread state of the second set of threads are saved, replace the first set of threads and the second set of threads on the first plurality of processing resources.

2. The graphics processor according to claim 1, further comprising: A local thread generator for generating the first set of threads for the first plurality of processing resources; And A device thread generator for generating child threads for execution via the first plurality of processing resources in response to a request from a parent thread executed via the first plurality of processing resources.

3. The graphics processor according to claim 2, wherein, To save the execution state of the first set of threads, the circuitry is configured to: Save the execution state of a first subset of the first set of threads, the first subset being generated via the local thread generator; And Save the execution state of a second subset of the first set of threads, the second subset being generated via the device thread generator.

4. The graphics processor according to claim 2 or 3, wherein, To save the thread state of the second set of threads, the circuitry is configured to: Save the thread state of a first subset of the second set of threads, the first subset being generated via the local thread generator; And Save the thread state of a second subset of the second set of threads, the second subset being generated via the device thread generator.

5. The graphics processor according to claim 2 or 3, wherein, The circuitry is configured to receive a command for triggering a recovery event for a second plurality of processing resources, the recovery event for restoring the first set of threads and the second set of threads to the second plurality of processing resources, wherein the second plurality of processing resources is different from the first plurality of processing resources.

6. The graphics processor according to claim 5, wherein, The graphics processor is configured to, in response to the recovery event: Generate a first recovery thread set for the second plurality of processing resources via the local thread generator; Generate a second recovery thread set for the second plurality of processing resources via a device thread generator associated with the second plurality of processing resources; And Generate a third recovery thread set for the second plurality of processing resources via the local thread generator.

7. The graphics processor according to claim 6, the circuitry being configured to: Restore the execution state of the first set of threads via the first recovery thread set; Restore the thread states of a first subset of the second thread set via the second set of restore threads; and Restore the thread states of a second subset of the second thread set via the third set of restore threads.

8. The graphics processor of claim 7, wherein the circuitry is configured to: Save and restore the execution state of the first thread set via a first location in memory; Save and restore the thread states of a first subset of the second thread set via a second location in memory; and Save and restore the thread states of a second subset of the second thread set via a third location in memory.

9. The graphics processor according to claim 8, wherein, The memory is an on-die memory of the graphics processor.

10. The graphics processor of claim 8 or 9, wherein the circuitry is configured to: Save the execution state and thread states via the execution resources of the first plurality of processing resources; and Restore the execution state and thread states via the execution resources of the second plurality of processing resources.

11. A method, comprising: Receiving a command for intermediate thread preemption of threads of a first plurality of processing resources for executing a graphics core; Saving the execution state of a first thread set of the first plurality of processing resources, the first thread set including threads in an active state; Saving the thread states of a second thread set of the first plurality of processing resources, the second thread set including threads in an inactive state; and And Replacing the first thread set and the second thread set on the first plurality of processing resources after the execution state of the first thread set and the thread states of the second thread set are saved.

12. The method of claim 11, further comprising: Generating the first thread set for the first plurality of processing resources via a local thread generator; and And Generating child threads for execution via the first plurality of processing resources via a device thread generator, the child threads being generated in response to a request from a parent thread executed via the first plurality of processing resources.

13. The method according to claim 12, wherein, Saving the execution state of the first thread set includes: Saving the execution state of a first subset of the first thread set, the first subset being generated via the local thread generator; and Saving the execution state of a second subset of the first thread set, the second subset being generated via the device thread generator.

14. The method according to claim 12, wherein, Saving the thread states of the second thread set includes: Saving the thread states of a first subset of the second thread set, the first subset being generated via the local thread generator; and Saving the thread states of a second subset of the second thread set, the second subset being generated via the device thread generator.

15. The method according to claim 12, further comprising: Receiving a command for triggering a restore event for a second plurality of processing resources, the restore event for restoring the first thread set and the second thread set to the second plurality of processing resources, wherein the second plurality of processing resources is different from the first plurality of processing resources.

16. The method according to claim 15, further comprising: In response to the restore event: Generating a first set of restore threads for the second plurality of processing resources via the local thread generator; Generating a second set of recovery threads for the second plurality of processing resources via a device thread generator associated with the second plurality of processing resources; and Generating a third set of recovery threads for the second plurality of processing resources via the local thread generator.

17. The method according to claim 16, further comprising: Restoring the execution state of the first set of threads via the first set of recovery threads; Restoring the thread states of a first subset of the second set of threads via the second set of recovery threads; and Restoring the thread states of a second subset of the second set of threads via the third set of recovery threads.

18. The method according to claim 17, further comprising: Saving and restoring the execution state of the first set of threads via a first location in a memory; Saving and restoring the thread states of the first subset of the second set of threads via a second location in the memory; and Saving and restoring the thread states of the second subset of the second set of threads via a third location in the memory.

19. The method according to claim 18, wherein, The memory is an on-die memory of a graphics processor.

20. The method according to claim 18, further comprising: Saving the execution state and the thread state via an execution resource of the first plurality of processing resources; and Restoring the execution state and the thread state via an execution resource of the second plurality of processing resources.

21. A graphics processing system, comprising: A memory device; and A plurality of processing resources, the plurality of processing resources being arranged as a plurality of slices, each slice including a plurality of sub-slices, the sub-slices including portions of the plurality of processing resources of the slice in the plurality of slices, the slice in the plurality of slices being configurable for assignment to a virtualized graphics environment, the sub-slices including circuitry configured to perform operations including: Dispatching a plurality of parent threads generated for execution via the sub-slice to the processing resources of the sub-slice; Generating a first sub-thread via a device thread generator of the sub-slice in response to a request from a first parent thread among the plurality of parent threads; Receiving a command for performing an intermediate thread preemption of a first thread group, the first thread group including the first parent thread and the first sub-thread; Saving the thread states of the first parent thread and the first sub-thread to the storage device; and Replacing the first thread group with a second thread group to be executed in place of the first thread group.

22. The graphics processing system according to claim 21, the operation further comprising: Assigning a stack identifier to each parent thread generated within the sub-slice; and Associating the first sub-thread with the stack identifier assigned to the first parent thread.

23. The graphics processing system according to claim 21 or 22, wherein the operation further comprises: Overdispatching the first sub-thread to a processing resource such that the first sub-thread is automatically executed when a physical thread of the processing resource is available.

24. The graphics processing system according to claim 23, the operation further comprising: Receiving the command for performing the intermediate thread preemption of the first thread group after the first sub-thread starts execution; and The physical thread of the processing resource saves the thread state of the first sub-thread, and the thread state includes the execution state of the physical thread.

25. The graphics processing system according to claim 23, wherein the operation further comprises: receiving the command for pre-empting the intermediate thread of the first thread group before the first sub-thread starts to execute; generating a first pseudo-thread via the device thread generator of the sub-slice; and saving the thread state of the first sub-thread via the first pseudo-thread.

Citation Information

Cited By

  • Instruction synchronization method and artificial intelligence chip

    CN121070625A