Data Locality Enhancement of Graphics Processing Unit

By integrating a GPU with a SIMT architecture and dedicated circuitry, the GPU communicatively coupled to a host processor enhances data locality and parallel processing efficiency, addressing challenges in handling complex graphics and machine learning operations.

JP7695034B2Active Publication Date: 2025-06-18INTEL CORP
View PDF 7 Cites 0 Cited by

Patent Information

Application Number
JP2020189363
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2020-11-11
Filing Date
2020-11-13
Publication Date
2025-06-18
Estimated Expiration
2040-11-13

AI Technical Summary

Technical Problem

Current graphics processing units (GPUs) face challenges in maximizing data locality and parallel processing efficiency, particularly in handling complex graphics and machine learning operations.

Method used

The integration of a GPU communicatively coupled to a host processor, utilizing a SIMT architecture and dedicated circuitry for efficient processing, enhances data locality and parallel processing by assigning work through series of commands/instructions in a work descriptor.

Benefits of technology

This configuration significantly improves processing efficiency, enabling GPUs to handle complex graphics and machine learning operations more effectively by maximizing parallel processing and data locality.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007695034000005
    Figure 0007695034000005
  • Figure 0007695034000006
    Figure 0007695034000006
  • Figure 0007695034000007
    Figure 0007695034000007
Patent Text Reader

Abstract

To provide data locality enhancement for graphics processing units.SOLUTION: Embodiments described herein provide an apparatus comprising a plurality of processing resources including a first processing resource and a second processing resource, a memory communicatively coupled to the first processing resource and the second processing resource, and a processor. The processor is configured to receive data dependence for one or more tasks including one or more producer tasks executed on the first processing resource and one or more consumer tasks executed on the second processing resource, and move data output from one or more producer tasks executed on the first processing resource to a cache memory communicatively coupled to the second processing resource. Other embodiments may be described and claimed.SELECTED DRAWING: Figure 27
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application claims priority to U.S. Provisional Patent Application No. 62 / 935,716, filed November 15, 2019, by Christopher J. Hughes et al., entitled “DATA LOCALITY ENHANCEMENT FOR GRAPHICS PROCESSING UNITS,” the entire disclosure of which is hereby incorporated by reference.

[0002] This disclosure generally relates to data processing, and more specifically, to data processing by a general-purpose graphics processing unit.

Background Art

[0003] Current parallel graphics data processing includes systems and methods developed to perform certain operations on graphics data, such as linear interpolation, tessellation, rasterization, texture mapping, depth testing, etc. Traditionally, graphics processors have used fixed-function calculation units to process graphics data. However, more recently, some graphics processors have been made programmable, enabling such processors to support a wider range of operations for processing vertex and fragment data.

[0004] To further improve performance, graphics processors typically implement processing techniques such as pipelining to process as much graphics data as possible in parallel through different parts of the graphics pipeline. A parallel graphics processor with a single instruction, multiple thread (SIMT) architecture is designed to maximize the amount of parallel processing in a graphics pipeline. In a SIMT architecture, to increase processing efficiency, multiple groups of parallel threads attempt to execute program instructions synchronously with each other as frequently as possible. An overview of software and hardware for SIMT architectures can be found in Shane Cook's CUDA Programming Chapter 3, pages 37-51 (2013).

Brief Description of the Drawings

[0005] To enable a more detailed understanding of the above-described features of the present embodiment, a more specific description of the embodiment briefly summarized above may be made by referring to the embodiments shown in part in the accompanying drawings. However, it should be noted that the accompanying drawings merely illustrate typical embodiments and should not be regarded as limiting the scope of the embodiments.

Figure 1

Figure 2A

Figure 2B

Figure 2C

Figure 2D

Figure 3A

Figure 3B

Figure 3C

Figure 4A

Figure 4B

Figure 4C

Figure 4D

Figure 4E

Figure 4F

Figure 5

Figure 6

Figure 7

Figure 8

Figure 9A

Figure 9B

Figure 10

Figure 11

Figure 12A

Figure 12B

Figure 13

Figure 14

Figure 15A

Figure 15B

Figure 15C

Figure 16A

Figure 16B

Figure 16C

Figure 17

Figure 18A

Figure 18B

Figure 19

Figure 20

Figure 21

Figure 22A

Figure 22B

Figure 23

Figure 24A

Figure 24B

Figure 24C

Figure 24D

Figure 25

Figure 26A

Figure 26B

Figure 27

Figure 28

Figure 29

DETAILED DESCRIPTION OF THE INVENTION

[0006] A graphics processing unit (GPU) is communicatively coupled to a host / processor core, for example, to accelerate graphics operations, machine learning operations, pattern analysis operations, and / or various general-purpose GPU (GPGPU) functions. The GPU can be communicatively coupled to the host processor / core over a bus or other interconnect (e.g., a high-speed interconnect such as PCIe or NVLink). Alternatively, the GPU can be integrated on the same package or chip as the core and communicatively coupled to the core over an internal processor bus / interconnect (i.e., within the package or chip). Regardless of how the GPU is connected, the processor core can assign work to the GPU in the form of a series of commands / instructions included in a work descriptor. And the GPU uses dedicated circuitry / logic for efficiently processing those commands / instructions.

[0007] In the following description, numerous specific details are set forth in order to provide a more thorough understanding. However, it will be understood by those skilled in the art that the embodiments described herein may be practiced without one or more of those specific details. Also, well-known mechanisms are not described in order to avoid obscuring the details of the present embodiments.

[0008] System Overview FIG. 1 is a block diagram illustrating a computing system 100 configured to implement one or more aspects of the embodiments described herein. The computing system 100 includes a processing subsystem 101 having one or more processors 102 and a system memory 104 that communicate via an interconnect path that may include a memory hub 105. The memory hub 105 may be a separate component within a chipset component or may be integrated within one or more of the processors 102. The memory hub 105 is coupled to an I / O subsystem 111 via a communication link 106. The I / O subsystem 111 includes an I / O hub 107 that enables the computing system 100 to receive inputs from one or more input devices 108. Further, the I / O hub 107 may enable a display controller, which may be included within one or more of the processors 102, to provide outputs to one or more display devices 110A. In one embodiment, the one or more display devices 110A coupled to the I / O hub 107 may include local, internal, or embedded display devices.

[0009] The processing subsystem 101 includes, for example, one or more parallel processors 112 coupled to a memory hub 105 via a bus or other communication link 113. The communication link 113 may be one of several standards-based communication link technologies or protocols, such as, but not limited to, PCI Express, or alternatively, a vendor-specific communication interface or communication fabric. The one or more parallel processors 112 may form a computationally intensive parallel or vector processing system that can include a number of processing cores and / or processing clusters, such as, for example, a many integrated core (MIC) processor. For example, the one or more parallel processors 112 may form a graphics processing subsystem that can output pixels to one of one or more display devices 110A coupled via an I / O hub 107. The one or more parallel processors 112 may also include a display controller and display interface (not shown) that enables a direct connection to one or more display devices 110B.

[0010] Within the I / O subsystem 111, the system storage unit 114 can be connected to the I / O hub 107 to provide a storage mechanism for the computing system 100. An interface mechanism can be used to enable connections between the I / O hub 107 and other components, such as a network adapter 118 and / or a wireless network adapter 119 that can be integrated into the platform, and various other devices that can be added via one or more add-in devices 120. The (one or more) add-in devices 120 can also include, for example, one or more external graphics processor devices and / or computing accelerators. The network adapter 118 can be an Ethernet adapter or other wired network adapter. The wireless network adapter 119 can include one or more of Wi-Fi, Bluetooth, near field communication (NFC), or other network devices including one or more wireless radios.

[0011] Computing system 100 can include other components not explicitly shown, including USB or other port connections, optical storage drives, video capture devices, and the like, which can also be connected to I / O hub 107. The communication paths interconnecting the various components of FIG. 1 can be implemented using a suitable protocol such as, for example, a PCI (Peripheral Component Interconnect)-based protocol (e.g., PCI-Express), or other buses or point-to-point communication interfaces and / or protocols such as, for example, NVLink high-speed interconnect, Compute Express Link (registered trademark) (CXL (registered trademark)) (e.g., CXL.mem), Infinity Fabric (IF), Ethernet (IEEE 802.3), Remote Direct Memory Access (RDMA), InfiniBand, Internet Wide Area RDMA Protocol (iWARP), Transmission Control Protocol (TCP), User Datagram Protocol (UDP), quick UDP Internet Connections (QUIC), RDMA over Converged Ethernet (RoCE), Intel QuickPath Interconnect (QPI), Intel Ultra Path Interconnect (UPI), Intel On-Chip System Fabric (IOSF), Omnipath, HyperTransport, Advanced Microcontroller Bus Architecture (AMBA), OpenCAPI, Gen-Z, Cache Coherent Interconnect for Accelerators (CCIX), 3GPP Long Term Evolution (LTE) (4G), 3GPP 5G, and variations thereof, or technically known wired or wireless interconnect protocols.In some examples, data can be replicated or stored in a virtualized storage node using a protocol such as, for example, Non-Volatile Memory Express (NVMe) over Fabrics (NVMe-oF) or NVMe.

[0012] One or more parallel processors 112 can incorporate circuits optimized for graphics and video processing, including, for example, a video output circuit, and can constitute a Graphics Processing Unit (GPU). Alternatively, or in addition, one or more parallel processors 112 can incorporate circuits optimized for general-purpose processing while preserving the underlying computing architecture, as will be described in more detail herein. Components of computing system 100 may be integrated with one or more other system elements on a single integrated circuit. For example, one or more parallel processors 112, memory hub 105, processor(s) 102, and I / O hub 107 can be integrated into a System-on-Chip (SoC) integrated circuit. Alternatively, components of computing system 100 may be integrated into a single package to form a System-in-Package (SIP) configuration. In one embodiment, at least some of the components of computing system 100 may be integrated into a Multi-Chip Module (MCM) and interconnected with other multi-chip modules to form a modular computing system.

[0013] It should be understood that the computing system 100 shown here is exemplary and can be modified and changed. The connection topology, including the number and configuration of bridges, the number of processors 102, and the number of parallel processors 112, can be changed as desired. For example, the system memory 104 can be directly connected to the (one or more) processors 102 without going through a bridge, while other devices communicate with the system memory 104 via the memory hub 105 and the (one or more) processors 102. In other alternative topologies, the (one or more) parallel processors 112 are connected to the I / O hub 107 instead of the memory hub 105, or directly connected to one of the one or more processors 102. In other embodiments, the I / O hub 107 and the memory hub 105 may be integrated on a single chip. It is also possible to attach two or more sets of processor sets 102 via a plurality of sockets that can be coupled to two or more of the parallel processors 112.

[0014] Some of the specific components shown here are optional and may not be included in all implementations of the computing system 100. For example, any number of add-in cards or peripheral devices may be supported, or some components may be removed. Also, some architectures may use different terms for components similar to those shown in FIG. 1. For example, the memory hub 105 may be called a north bridge in some architectures, and the I / O hub 107 may be called a south bridge.

[0015] FIG. 2A illustrates a parallel processor 200. The parallel processor 200 can be a GPU, GPGPU, or the like as described herein. The various components of the parallel processor 200 can be implemented using one or more integrated circuit devices such as, for example, programmable processors, application specific integrated circuits (ASICs), or field programmable gate arrays (FPGAs). The illustrated parallel processor 200 can be the (one or more) parallel processors 112 shown in FIG. 1 or one of them.

[0016] The parallel processor 200 includes a parallel processing unit 202. The parallel processing unit includes an I / O unit 204 that enables communication with other devices, including other instances of the parallel processing unit 202. The I / O unit 204 may be directly connected to other devices. For example, the I / O unit 204 is connected to other devices via a hub or switch interface such as, for example, a memory hub 105. The connection between the memory hub 105 and the I / O unit 204 forms a communication link 113. Within the parallel processing unit 202, the I / O unit 204 is connected to a host interface 206 and a memory crossbar 216. The host interface 206 receives commands for performing processing operations, and the memory crossbar 216 receives commands for performing memory operations.

[0017] When the host interface 206 receives command buffers via the I / O unit 204, the host interface 206 can instruct the front end 208 to perform the work operations for executing those commands. In one embodiment, the front end 208 is coupled to a scheduler 210 configured to distribute commands or other work items to the processing cluster array 212. The scheduler 210 ensures that the processing cluster array 212 is properly configured and in an active state before tasks are distributed to the processing clusters of the processing cluster array 212. The scheduler 210 can be implemented via firmware logic executed on a microcontroller. The scheduler 210 implemented on a microcontroller can be configured to perform complex scheduling and work distribution operations at coarse and fine granularities, enabling rapid preemption and context switching of threads executing on the processing cluster array 212. Preferably, the host software can demonstrate a workload for scheduling for the processing cluster array 212 via one of a plurality of graphics processing portals. In other examples, new workloads or polling of interrupts can be used to identify or indicate the availability of work to be performed. And the workload can be automatically distributed across the processing cluster array 212 by the scheduler 210 logic within the scheduler microcontroller.

[0018] The processing cluster array 212 can include up to "N" processing clusters (e.g., cluster 214A, cluster 214B, up to cluster 214N). Each cluster 214A - 214N of the processing cluster array 212 can execute a number of simultaneous threads. The scheduler 210 can assign work to the clusters 214A - 214N of the processing cluster array 212 using various scheduling and / or work distribution algorithms that can vary depending on the workload that occurs for each type of program or calculation. The scheduling can be dynamically processed by the scheduler 210 or can be partially assisted by the compiler logic during the compilation of the program logic configured for execution by the processing cluster array 212. Optionally, different clusters 214A - 214N of the processing cluster array 212 can be assigned to process different types of programs or to perform different types of calculations.

[0019] The processing cluster array 212 can be configured to perform various types of parallel processing operations. For example, the processing cluster array 212 can be configured to perform general-purpose parallel computing operations. For example, the processing cluster array 212 can include logic to perform tasks including filtering of video and / or audio data, performing modeling operations including physical operations, and performing data conversions.

[0020] The processing cluster array 212 is configured to perform parallel graphics processing operations. In such embodiments where the parallel processor 200 is configured to perform graphics processing operations, the processing cluster array 212 can include additional logic for supporting the execution of such graphics processing operations, including but not limited to texture sampling logic for performing texture processing, as well as tessellation logic and other vertex processing logic. Further, the processing cluster array 212 can be configured to execute shader programs related to graphics processing, such as but not limited to vertex shaders, tessellation shaders, geometry shaders, and pixel shaders. The parallel processing unit 202 can transfer data from the system memory through the I / O unit 204 for processing. In processing, the transferred data can be stored in on-chip memory (e.g., parallel processor memory 222) only during processing and then written back to the system memory.

[0021] In embodiments where graphics processing is performed using the parallel processing unit 202, the scheduler 210 can be configured to divide the processing workload into tasks of approximately equal size to better enable the distribution of graphics processing operations to the multiple clusters 214A - 214N of the processing cluster array 212. In some of these embodiments, portions of the processing cluster array 212 can be configured to perform different types of processing. For example, for creating an image rendered for display, the first portion can be configured to perform vertex shading and topology generation, the second portion can be configured to perform tessellation and geometry shading, and the third portion can be configured to perform pixel shading or other screen space operations. Intermediate data generated by one or more of the clusters 214A - 214N can be stored in a buffer to enable the transmission of the intermediate data between the clusters 214A - 214N for further processing.

[0022] In operation, the processing cluster array 212 may receive processing tasks to be executed via a scheduler 210 that receives commands defining the processing tasks from the front end 208. In a graphics processing operation, the processing task may include an index of data to be processed, such as surface (patch) data, primitive data, vertex data, and / or pixel data, and commands and state parameters defining how the data is to be processed (e.g., which program is to be executed). The scheduler 210 may be configured to fetch an index corresponding to the task or may receive the index from the front end 208. The front end 208 may be configured to ensure that the processing cluster array 212 is configured in an active state before the workload specified by an incoming command buffer (e.g., batch buffer, push buffer, etc.) is started.

[0023] Each of one or more instances of the parallel processing unit 202 may be coupled to the parallel processor memory 222. The parallel processor memory 222 can be accessed via a memory crossbar 216 that can receive memory requests from the processing cluster array 212 and the I / O unit 204. The memory crossbar 216 can access the parallel processor memory 222 via a memory interface 218. The memory interface 218 can include a plurality of partition units (e.g., partition unit 220A, partition unit 220B, or partition unit 220N), and each partition unit can be coupled to a portion (e.g., a memory unit) of the parallel processor memory 222. The number of partition units 220A - 220N can be configured to be equal to the number of memory units such that the first partition unit 220A has a corresponding first memory unit 224A, the second partition unit 220B has a corresponding second memory unit 224B, and the Nth partition unit 220N has a corresponding Nth memory unit 224N. In other embodiments, the number of partition units 220A - 220N may not be equal to the number of memory devices.

[0024] Memory units 224A - 224N can include various types of memory devices, including dynamic random access memory (DRAM), or graphics random access memory such as synchronous graphics random access memory (SGRAM) including, for example, graphics double data rate (GDDR) memory. Optionally, memory units 224A - 224N can also include 3D stacked memory, including but not limited to high bandwidth memory (HBM). As will be understood by those skilled in the art, the specific implementation of memory units 224A - 224N can vary and can be selected from a variety of conventional designs. Rendering targets such as frame buffers or texture maps can be stored across memory units 224A - 224N, and partition units 220A - 220N writing portions of each rendering target in parallel enables efficient use of the available bandwidth of parallel processor memory 222. Advantageously for a unified memory design that utilizes system memory with local cache memory, in some embodiments, the local instance of parallel processor memory 222 may be excluded.

[0025] Optionally, any of the clusters 214A - 214N of the processing cluster array 212 has the ability to process data that will be written to any of the memory units 224A - 224N within the parallel processor memory 222. The memory crossbar 216 can be configured to transfer the output of each cluster 214A - 214N to any of the partition units 220A - 220N or to another cluster 214A - 214N (where further processing operations can be performed on the output). Each cluster 214A - 214N can communicate with the memory interface 218 via the memory crossbar 216 to read from or write to various external memory devices. In one embodiment with a memory crossbar 216, the memory crossbar 216 has a connection to the memory interface 218 for communicating with the I / O unit 204 and a connection to a local instance of the parallel processor memory 222 that enables processing units within different processing clusters 214A - 214N to communicate with the system memory or other memory that is not local to the parallel processing unit 202. Generally, the memory crossbar 216 can potentially separate traffic streams between the clusters 214A - 214N and the partition units 220A - 220N, for example, using multiple virtual channels.

[0026] Only one instance of the parallel processing unit 202 is shown within the parallel processor 200, but any number of instances of the parallel processing unit 202 may be included. For example, multiple instances of the parallel processing unit 202 may be provided on a single add-in card, or alternatively, multiple add-in cards may be interconnected. For example, the parallel processor 200 may be a graphics card such as a discrete graphics card including, for example, one or more GPUs, one or more memory devices, and an inter-device interface or a network interface or a fabric interface, and may be an add-in device such as the add-in device 120 of FIG. 1. Different instances of the parallel processing unit 202 can be configured to interoperate even if those different instances have different numbers of processing cores, different amounts of local parallel processor memory, and / or other configurational differences. Optionally, some instances of the parallel processing unit 202 can include a higher-precision floating-point unit than other instances. A system incorporating one or more instances of the parallel processing unit 202 or the parallel processor 200 can be implemented in a variety of configurations and form factors including, but not limited to, desktop, laptop, or handheld personal computers, servers, workstations, game consoles, and / or embedded systems. An orchestrator can form a composite node for workload execution using one or more of non-agglomerated processor resources, cache resources, memory resources, storage resources, and networking resources.

[0027] FIG. 2B is a block diagram of partition unit 220. Partition unit 220 may be an instance of one of partition units 220A-220N of FIG. 2A. As shown, partition unit 220 includes an L2 cache 221, a frame buffer interface 225, and a ROP 226 (raster operation unit). L2 cache 221 is a read / write cache configured to perform load and store operations received from memory crossbar 216 and ROP 226. Read misses and urgent write-back requests are output by L2 cache 221 to frame buffer interface 225 for processing. Updates may also be sent to the frame buffer via frame buffer interface 225 for processing. In one embodiment, frame buffer interface 225 interfaces with one of the memory units within the parallel processor memory, such as, for example, memory units 224A-224N of FIG. 2A (e.g., within parallel processor memory 222). Partition unit 220 may also, additionally or alternatively, interface with one of the memory units within the parallel processor memory via a memory controller (not shown).

[0028] In graphics applications, ROP226 is a processing unit that performs raster operations such as stencil, z-test, blending, and the like. ROP226 then outputs the processed graphics data stored in the graphics memory. In some embodiments, ROP226 includes, or is coupled with, CODEC227 that includes compression logic for compressing depth or color data written to memory or L2 cache 221 and decompressing depth or color data read from memory or L2 cache 221. The compression logic can be reversible compression logic that uses one or more of a plurality of compression algorithms. The type of compression performed by ROP226, CODEC227 can vary based on the statistical characteristics of the data to be compressed. For example, in one embodiment, delta color compression is performed on depth and color data on a tile-by-tile basis. In one embodiment, CODEC227 includes compression and decompression logic that can compress and decompress computational data related to machine learning operations. CODEC227 can, for example, compress sparse matrix (sparse matrix) data for sparse machine learning operations. CODEC227 can also compress sparse matrix data encoded in a sparse matrix format (e.g., coordinate list encoding (COO), compressed sparse row (CSR), compressed sparse column (CSC), etc.) to generate compressed and encoded sparse matrix data. The compressed and encoded sparse matrix data can be decompressed and / or decoded before being processed by a processing element, or the processing element can be configured to consume the compressed, encoded, or compressed and encoded data for processing.

[0029] ROP226 may instead be included within each processing cluster (e.g., clusters 214A - 214N of FIG. 2A) rather than within partition unit 220. In such embodiments, read and write requests regarding pixel data are transmitted on memory crossbar 216 instead of pixel fragment data. The processed graphics data is displayed on a display device, such as one of the one or more display devices 110 of FIG. 1, routed for further processing by (one or more) processors 102, or routed for further processing by one of the processing entities within parallel processor 200 of FIG. 2A.

[0030] FIG. 2C is a block diagram of processing cluster 214 within the parallel processing unit. For example, this processing cluster is an instance of one of processing clusters 214A - 214N of FIG. 2A. Processing cluster 214 can be configured to execute multiple threads in parallel, where the term "thread" refers to an instance of a particular program executed on a particular input data set. Optionally, single instruction multiple data (SIMD) instruction - issuing techniques can be used to support parallel execution of multiple threads without providing multiple independent instruction units. Alternatively, single instruction multiple thread (SIMT) techniques can be used, using a common instruction unit configured to issue instructions to a set of processing engines within each of the processing clusters, to support parallel execution of a generally synchronized set of multiple threads. Unlike SIMD execution regimes where typically all processing engines execute the same instruction, SIMT execution allows different threads to more easily follow different execution paths within a given thread program. As will be understood by those skilled in the art, the SIMD processing regime represents a functional subset of the SIMT processing regime.

[0031] The operation of the processing cluster 214 can be controlled via a pipeline manager 232 that distributes processing tasks to SIMT parallel processors. The pipeline manager 232 receives instructions from the scheduler 210 of FIG. 2A and manages the execution of those instructions by the graphics multiprocessor 234 and / or the texture unit 236. The illustrated graphics multiprocessor 234 is an exemplary instance of a SIMT parallel processor. However, various types of SIMT parallel processors of different architectures can be included within the processing cluster 214. One or more instances of the graphics multiprocessor 234 can be included within the processing cluster 214. The graphics multiprocessor 234 can process data and, using the data crossbar 240, distribute the processed data to one of a plurality of possible destinations including other shader units. The pipeline manager 232 can assist in the distribution of the processed data by specifying the destination to which the processed data is to be distributed via the data crossbar 240.

[0032] Each graphics multiprocessor 234 within the processing cluster 214 can include an equivalent set of functional execution logic (e.g., arithmetic logic units, load-store units, etc.). The functional execution logic can be configured in a pipelined manner and can issue new instructions before the preceding instructions are completed. The functional execution logic supports a variety of operations including integer and floating-point arithmetic, comparison operations, boolean operations, bit shifts, and the calculation of various algebraic functions. Different operations may be executed using the same functional unit hardware, and any combination of functional units may exist.

[0033] The instructions transmitted to the processing cluster 214 constitute threads. A set of threads that are executed across a set of parallel processing engines is a thread group. The thread group executes the same program for different input data. Each thread within the thread group can be assigned to a different processing engine within the graphics multiprocessor 234. The thread group may include fewer threads than the number of processing engines within the graphics multiprocessor 234. If the thread group includes fewer threads than the number of processing engines, one or more of the processing engines may be in an idle state during the cycles in which the thread group is being processed. The thread group may also include more threads than the number of processing engines within the graphics multiprocessor 234. If the thread group includes more threads than the number of processing engines within the graphics multiprocessor 234, processing can be executed over a plurality of consecutive clock cycles. Optionally, multiple thread groups can be executed simultaneously on the graphics multiprocessor 234.

[0034] The graphics multiprocessor 234 may include an internal cache memory for performing load and store operations. Optionally, the graphics multiprocessor 234 may defer the internal cache and use the cache memory (e.g., level 1 (L1) cache 248) within the processing cluster 214. Each graphics multiprocessor 234 also has access to a level 2 (L2) cache within a partitioning unit (e.g., partitioning units 220A - 220N of FIG. 2A) that can be shared among all processing clusters 214 and used to transfer data between threads. The graphics multiprocessor 234 can also access off-chip global memory, which may include one or more local parallel processor memories and / or system memory. Any memory external to the parallel processing unit 202 may be used as global memory. Embodiments in which the processing cluster 214 includes multiple instances of the graphics multiprocessor 234 can share common instructions and data, which may be stored in the L1 cache 248.

[0035] Each processing cluster 214 may include an MMU 245 (memory management unit) configured to map virtual addresses to physical addresses. In other embodiments, one or more instances of the MMU 245 may be present within the memory interface 218 of FIG. 2A. The MMU 245 includes a set of page table entries (PTEs) used to map virtual addresses to the physical addresses of tiles and optionally includes a cache line index. The MMU 245 may include a translation lookaside buffer (TLB) or cache that may be present within the graphics multiprocessor 234 or the L1 cache or the processing cluster 214. The physical addresses are processed to disperse the surface data access locations to enable efficient request interleaving among the partitioning units. The cache line index may be used to determine whether a request for a cache line is a hit or a miss.

[0036] In graphics and computing applications, the processing cluster 214 can be configured such that each graphics multiprocessor 234 is coupled to a texture unit 236 that performs texture mapping operations, such as determining texture sample positions, reading texture data, and filtering texture data. The texture data is read from an internal texture L1 cache (not shown) or, in some embodiments, from the L1 cache within the graphics multiprocessor 234 and, if necessary, fetched from the L2 cache, local parallel processor memory, or system memory. Each graphics multiprocessor 234 outputs the processed task to the data crossbar 240 for providing the processed task to another processing cluster 214 for further processing or for storing the processed task in the L2 cache, local parallel processor memory, or system memory via the memory crossbar 216. The pre-ROP 242 (Pre-Raster Operations unit) is configured to receive data from the graphics multiprocessor 234 and direct the data to a ROP unit that can be arranged with the partitioning units described herein (e.g., partitioning units 220A - 220N of FIG. 2A). The pre-ROP 242 unit can perform optimizations for color mixing, compositing pixel color data, and performing address translation.

[0037] It should be understood that the core architectures described herein are exemplary and can be modified and changed. For example, any number of processing units, such as the graphics multiprocessor 234, texture unit 236, pre-ROP 242, etc., may be included within the processing cluster 214. Also, although only one processing cluster 214 is shown, the parallel processing units described herein may include any number of instances of the processing cluster 214. Optionally, each processing cluster 214 can be configured to operate independently of other processing clusters 214 using separate and different processing units, L1 caches, L2 caches, etc.

[0038] FIG. 2D shows an example of the graphics multiprocessor 234, which is coupled to the pipeline manager 232 of the processing cluster 214. The graphics multiprocessor 234 has an execution pipeline including, but not limited to, an instruction cache 252, an instruction unit 254, an address mapping unit 256, a register file 258, one or more general-purpose graphics processing unit (GPGPU) cores 262, and one or more load / store units 266. The GPGPU cores 262 and the load / store units 266 are coupled to the cache memory 272 and the shared memory 270 via a memory cache interconnect 268. The graphics multiprocessor 234 may further include a tensor and / or ray tracing core 263 including hardware logic for accelerating matrix operations and / or ray tracing operations.

[0039] The instruction cache 252 may receive a stream of instructions to be executed from the pipeline manager 232. The instructions are cached in the instruction cache 252 and dispatched for execution by the instruction unit 254. The instruction unit 254 can dispatch instructions as a thread group, with each thread of the thread group assigned to a different execution unit within the GPGPU core 262. Instructions can access any of the local, shared, or global address spaces by specifying an address within the unified address space. The address mapping unit 256 can be used to convert an address in the unified address space to another memory address accessible by the load / store unit 266.

[0040] The register file 258 provides a set of registers for the functional units of the graphics multiprocessor 234. The register file 258 provides temporary storage of operands connected to the data paths of the functional units of the graphics multiprocessor 234 (e.g., the GPGPU core 262, the load / store unit 266). The register file 258 can be divided among each of the functional units such that each functional unit is assigned a dedicated portion of the register file 258. For example, the register file 258 can be divided among different warps executed by the graphics multiprocessor 234.

[0041] Each GPGPU core 262 can include a floating-point unit (FPU) and / or an integer arithmetic logic unit (ALU) used to execute instructions of the graphics multiprocessor 234. In some implementations, the GPGPU core 262 can include hardware logic that would otherwise be present within the tensor and / or ray tracing core 263. These GPGPU cores 262 may be the same in architecture or different in architecture. For example, in one embodiment, the first part of the GPGPU core 262 includes a single-precision FPU and an integer ALU, and the second part of the GPGPU core includes a double-precision FPU. Optionally, the FPU can implement the IEEE 754-2008 standard for floating-point operations or can enable variable-precision floating-point operations. The graphics multiprocessor 234 can further include one or more fixed-function or special-function units for performing specific functions such as copy rectangle or pixel blending operations. One or more of the GPGPU cores can also include fixed or special-function logic.

[0042] The GPGPU core 262 can include SIMD logic capable of executing a single instruction on multiple data sets. Optionally, the GPGPU core 262 can physically execute SIMD4, SIMD8, and SIMD16 instructions and logically execute SIMD1, SIMD2, and SIMD32 instructions. SIMD instructions for the GPGPU core can be generated at compile time by a shader compiler or can be automatically generated when executing a program written and compiled for a single-program multiple-data (SPMD) architecture or a SIMT architecture. Multiple threads of a program configured for the SIMT execution model can be executed via a single SIMD instruction. For example, in one embodiment, eight SIMT threads performing the same or similar operations can be executed in parallel by a single SIMD8 logic unit.

[0043] The memory cache interconnect 268 is an interconnect network that connects each of the functional units of the graphics multiprocessor 234 to the register file 258 and the shared memory 270. For example, the memory cache interconnect 268 is a crossbar interconnect that enables the load / store unit 266 to implement load and store operations between the shared memory 270 and the register file 258. The register file 258 can operate at the same frequency as the GPGPU core 262, and thus, data transfer between the GPGPU core 262 and the register file 258 is very low latency. The shared memory 270 can be used to enable communication between threads executed on the functional units within the graphics multiprocessor 234. The cache memory 272 can be used as a data cache, for example, to cache texture data communicated between the functional units and the texture unit 236. The shared memory 270 can also be used as a program-controlled cache. The shared memory 270 and the cache memory 272 can be coupled to the data crossbar 240 to enable communication with other components of the processing cluster. Threads executed on the GPGPU core 262 can programmatically store data in the shared memory in addition to automatically cached data stored in the cache memory 272.

[0044] Figures 3A - 3C illustrate a further graphics multiprocessor according to an embodiment. Figures 3A - 3B relate to the graphics multiprocessor 234 of Figure 2C and show graphics multiprocessors 325, 350 that can be used in place of one of them. Thus, any disclosure of features in combination with the graphics multiprocessor 234 here also discloses corresponding combinations with the graphics multiprocessors 325, 350, but is not so limited. Figure 3C shows a graphics processing unit (GPU) 380 that includes a dedicated set of graphics processing resources organized into multi-core groups 365A - 365N corresponding to the graphics multiprocessors 325, 350. The illustrated graphics multiprocessors 325, 350 and multi-core groups 365A - 365N can be streaming multiprocessors capable of simultaneous execution of a large number of execution threads.

[0045] The graphics multiprocessor 325 of Figure 3A includes instances of a plurality of additional execution resource units with respect to the graphics multiprocessor 234 of Figure 2D. For example, the graphics multiprocessor 325 may include multiple instances of instruction units 332A - 332B, register files 334A - 334B, and texture units 344A - 344B. The graphics multiprocessor 325 also includes multiple sets of graphics or compute execution units (e.g., GPGPU cores 336A - 336B, tensor cores 337A - 337B, ray tracing cores 338A - 338B) and multiple sets of load / store units 340A - 340B. These execution resource units have a common instruction cache 330, texture and / or data cache memory 342, and shared memory 346.

[0046] These various components can communicate via an interconnect fabric 327. The interconnect fabric 327 can include one or more crossbar switches that enable communication between the various components of the graphics multiprocessor 325. The interconnect fabric 327 can be a separate high-speed network fabric layer in which the components of the graphics multiprocessor 325 are stacked. The components of the graphics multiprocessor 325 communicate with remote components via the interconnect fabric 327. For example, cores 336A-336B, 337A-337B, and 338A-338B can each communicate with the shared memory 346 via the interconnect fabric 327. The interconnect fabric 327 can arbitrate communications within the graphics multiprocessor 325 to ensure fair bandwidth allocation between components.

[0047] The graphics multiprocessor 350 of FIG. 3B includes multiple sets of execution resources 356A-356D, with each set of execution resources including multiple instruction units, register files, GPGPU cores, and load / store units as shown in FIGS. 2D and 3A. The execution resources 356A-356D can operate in cooperation with (one or more) texture units 360A-360D for texture operations while sharing an instruction cache 354 and a shared memory 353. For example, the execution resources 356A-356D can share the instruction cache 354 and the shared memory 353 with multiple instances of texture and / or data cache memories 358A-358B. These various components can communicate via an interconnect fabric 352 similar to the interconnect fabric 327 of FIG. 3A.

[0048] As will be understood by those skilled in the art, the architectures described in FIGS. 1, 2A-2D, and 3A-3B are illustrative and not limiting with respect to the scope of the present embodiment. Accordingly, the techniques described herein may include, without limitation and without departing from the scope of the embodiments described herein, the following, namely, one or more mobile application processors, one or more desktop or server central processing units (CPUs) including a multi-core CPU, one or more parallel processing units such as the parallel processing unit 202 of FIG. 2A, and one or more graphics processors or special-purpose processing units, and may be implemented on any suitably configured processing unit.

[0049] The parallel processors or GPGPUs described herein may be communicatively coupled to a host / processor core to accelerate graphics operations, machine learning operations, pattern analysis operations, and various general-purpose GPU (GPGPU) functions. The GPU may be communicatively coupled to the host processor / core via a bus or other interconnect (e.g., a high-speed interconnect such as PCIe or NVLink), NVLink, or other known protocol, standardized protocol, or proprietary protocol. In other embodiments, the GPU may be integrated on the same package or die as the core and communicatively coupled to the core on an internal processor bus / interconnect (i.e., within the package or die). Regardless of how the GPU is connected, the processor core may assign work to the GPU in the form of a series of commands / instructions included in a work descriptor. And the GPU uses dedicated circuitry / logic for efficiently processing those commands / instructions.

[0050] FIG. 3C illustrates a graphics processing unit (GPU) 380 that includes a dedicated set of graphics processing resources organized into multi-core groups 365A-365N. Although only the details of a single multi-core group 365A are presented, it is understood that the other multi-core groups 365B-365N may comprise the same or similar sets of graphics processing resources. The details described with respect to multi-core groups 365A-365N may apply to any of the graphics multiprocessors 234, 325, 350 described herein.

[0051] As shown, multi-core group 365A may include a set of graphics cores 370, a set of tensor cores 371, and a set of ray tracing cores 372. A scheduler / dispatcher 368 schedules and dispatches graphics threads to be executed on the various cores 370, 371, 372. A set of register files 369 stores operand values used by cores 370, 371, 372 when executing graphics threads. These may include, for example, integer registers for storing integer values, floating-point registers for storing floating-point values, vector registers for storing packed data elements (integer and / or floating-point data elements), and tile registers for storing tensor / matrix values. The tile registers may be implemented as a combination of sets of vector registers.

[0052] One or more combined level 1 (L1) caches and shared memory unit 373 locally store graphics data such as texture data, vertex data, pixel data, ray data, boundary volume data, etc. within each multicore group 365A. One or more texture units 374 may also be used to perform texture mapping operations such as texture mapping and sampling. Level 2 (L2) cache 375, shared by all or a subset of the multicore groups 365A - 365N, stores graphics data and / or instructions regarding a plurality of simultaneously parallel graphics threads. As shown in the figure, the L2 cache 375 may be shared across a plurality of multicore groups 365A - 365N. One or more memory controllers 367 couple the GPU 380 to a memory 366 that may be a system memory (e.g., DRAM) and / or a dedicated graphics memory (e.g., GDDR6 memory).

[0053] Input / output (I / O) circuit 363 couples the GPU 380 to one or more I / O devices 362 such as, for example, a digital signal processor (DSP), a network controller, or a user input device. An on-chip interconnect may be used to couple the I / O device 362 to the GPU 380 and the memory 366. One or more I / O memory management units (IOMMUs) 364 of the I / O circuit 363 directly couple the I / O device 362 to the system memory 366. Optionally, the IOMMU 364 manages multiple sets of page tables for mapping virtual addresses to physical addresses within the system memory 366. And the I / O device 362, the (one or more) CPUs 361, and the (one or more) GPUs 380 may share the same virtual address space.

[0054] In one implementation of the IOMMU 364, the IOMMU 364 supports virtualization. In this case, it may manage a first page table set for mapping guest / graphics virtual addresses to guest / graphics physical addresses and a second page table set for mapping guest / graphics physical addresses to system / host physical addresses (e.g., within the system memory 366). The base address of each of the first and second page table sets is stored in a control register and can be exchanged during a context switch (e.g., to provide access to the page table set corresponding to the new context). Although not shown in FIG. 3C, each of the cores 370, 371, 372, and / or the multicore groups 365A - 365N may include a translation lookaside buffer (TLB) for caching guest virtual-guest physical conversion, guest physical-host physical conversion, and guest virtual-host physical conversion.

[0055] The CPU 361, GPU 380, and I / O device 362 may be integrated on a single semiconductor chip and / or chip package. The illustrated memory 366 may be integrated on the same chip or may be coupled to the memory controller 367 via an off-chip interface. In one embodiment, the memory 366 has GDDR6 memory that shares the same virtual address space as other physical system-level memories, but the basic principles described herein are not limited to this particular implementation.

[0056] Tensor Core 371 may include a plurality of execution units specifically designed to perform matrix operations, which are basic computational operations used to execute deep learning operations. For example, simultaneous matrix multiplication operations may be used for neural network training and inference. Tensor Core 371 may perform matrix processing using a variety of operand precisions, including single-precision floating point (e.g., 32 bits), half-precision floating point (e.g., 16 bits), integer word (16 bits), byte (8 bits), and nibble (4 bits). For example, a neural network implementation may extract features of each scene to be rendered by combining details from potentially multiple frames to build a high-quality final image.

[0057] In a deep deep learning implementation, parallel matrix multiplication operations may be scheduled for execution on Tensor Core 371. Training of a neural network, in particular, requires a fairly large number of matrix dot product operations. To process the inner product formula for an N×N×N matrix multiplication, Tensor Core 371 may include at least N dot product processing elements. Before matrix multiplication begins, an entire matrix is loaded into a tile register, and at least one column of the second matrix is loaded in each of N cycles. In each cycle, there are N dot products to be processed.

[0058] The matrix elements can be stored with different precisions, including 16-bit words, 8-bit bytes (e.g., INT8), and 4-bit nibbles (e.g., INT4), depending on the specific implementation. Different precision modes may be specified for these tensor cores 371 to ensure that the most efficient precision is used for various workloads (e.g., inference workloads that can tolerate quantization to bytes and nibbles). The supported formats further include the 64-bit floating-point (FP64) format, non-IEEE floating-point formats such as the bfloat16 format (e.g., brain floating-point), a 16-bit floating-point format with 1 sign bit, 8 exponent bits, and 8 most significant bits of which 7 are explicitly stored. One embodiment includes support for a reduced-precision tensor float format (TF32) that has a range of FP32 (8 bits) with a precision of FP16 (10 bits). By performing reduced-precision TF32 operations on FP32 inputs, FP32 outputs can be generated with higher performance than FP32 and higher precision than FP16.

[0059] In one embodiment, tensor core 371 supports a sparse operation mode for matrices where most of the values are zero. Tensor core 371 includes support for sparse input matrices encoded in a sparse matrix representation (e.g., coordinate list encoding (COO), compressed sparse row (CSR), compressed sparse column (CSC), etc.). Tensor core 371 also includes support for a compressed sparse matrix representation when the sparse matrix representation can be further compressed. Compressed, encoded, and / or compressed and encoded matrix data can be prepared by tensor core 371 together with the associated compression and / or encoding metadata, and non-zero values can be extracted. For example, for a given input matrix A, non-zero values can be loaded from at least a partially compressed and / or encoded representation of matrix A. Based on the positions in matrix A of the non-zero values, which can be determined from the index or coordinate metadata for the non-zero values, corresponding values in input matrix B can be loaded. Depending on the operation being performed (e.g., multiplication), the loading of values from input matrix B can be bypassed if the corresponding values are zero values. In one embodiment, the pairing of values for a particular operation, such as a multiplication operation, can be pre-scanned by scheduler logic, and only operations between non-zero inputs are scheduled. Depending on the dimensions of matrices A and B and the operation being performed, output matrix C may be dense or sparse. If output matrix C is sparse, depending on the configuration of tensor core 371, output matrix C can be output in a compressed format, a sparse encoding, or a compressed sparse encoding.

[0060] The ray tracing core 372 can accelerate ray tracing operations in both real-time and non-real-time ray tracing implementations. In particular, the ray tracing core 372 can include a ray traversal / intersection circuit that performs ray traversal using a bounding volume hierarchy (BVH) and identifies intersections between primitives enclosed within the BVH volume and rays. The ray tracing core 372 can also include circuitry for performing depth testing and culling (e.g., using a Z-buffer or similar configuration). In one embodiment, the ray tracing core 372 executes traversal and intersection operations in cooperation with the image noise removal techniques described herein, at least a portion of which can be executed on tensor core 371. For example, tensor core 371 can implement a deep learning neural network for performing noise removal on frames generated by ray tracing core 372. However, the (one or more) CPUs 361, graphics core 370, and / or ray tracing core 372 can also implement all or part of the noise removal and / or deep learning algorithms.

[0061] Furthermore, as described above, a distributed approach to noise removal may be adopted in which a GPU 380 is present within a computing device coupled to other computing devices on a network or high-speed interconnect. In this distributed approach, interconnected computing devices can share neural network learning / training data to improve the speed at which the system as a whole learns to perform noise removal for various types of image frames and / or different graphics applications.

[0062] The ray tracing core 372 can save the graphics core 370 from being overloaded with thousands of instructions per ray by handling all BVH traversals and / or ray-primitive intersections. For example, each ray tracing core 372 includes a first set of special circuits for performing bounding box tests (e.g., traversal operations), and / or a second set of special circuits for performing ray-triangle intersection tests (e.g., traversed intersection rays). Thus, for example, the multi-core group 365A simply fires a ray probe, and the ray tracing core 372 independently performs ray traversal and intersection, and returns hit data (e.g., hit, miss, multiple hits, etc.) to the thread context. While the ray tracing core 372 performs traversal and intersection operations, the other cores 370, 371 are freed up to perform other graphics or computing tasks.

[0063] Optionally, each ray tracing core 372 may include a traversal unit for performing BVH test operations and / or an intersection unit for performing ray-primitive intersection tests. The intersection unit generates a "hit", "miss", or "multiple hits" response and provides it to the appropriate thread. During traversal and intersection operations, the execution resources of other cores (e.g., graphics core 370 and tensor core 371) are freed up to perform other forms of graphics work.

[0064] In one embodiment of the option described below, a hybrid rasterization / ray tracing approach is used in which work is distributed between the graphics core 370 and the ray tracing core 372.

[0065] The ray tracing core 372 (and / or other cores 370, 371) can include a ray tracing instruction set such as Microsoft's DirectX Ray Tracing (DXR) including, for example, the DispatchRays command, and can include hardware support for ray generation, closest hit, any hit, and miss shaders, and can enable the assignment of a set of unique shaders and textures to each object. Another ray tracing platform that can be supported by the ray tracing core 372, the graphics core 370, and the tensor core 371 is Vulkan 1.1.85. However, the basic principles described herein are not limited to a particular ray tracing ISA.

[0066] Generally, the various cores 372, 371, 370 can support a ray tracing instruction set including instructions / functions related to one or more of ray generation, closest hit, any hit, ray-primitive intersection, primitive unit and hierarchical bounding box construction, miss, visit, and exception. More specifically, a preferred embodiment includes ray tracing instructions for performing one or more of the following functions.

[0067] Ray Generation - Ray generation instructions can be executed for each pixel, sample, or other user-defined work assignment.

[0068] Closest Hit - Closest hit instructions can be executed to locate the closest intersection of a ray with a primitive in the scene.

[0069] Any Hit - Any hit instructions identify multiple intersections between a ray and primitives in the scene to potentially identify a new closest intersection.

[0070] Intersection - The Intersection instruction performs a ray-primitive intersection test and outputs the result.

[0071] Per-primitive Bounding box Construction - This instruction constructs a bounding box around a given primitive or group of primitives (e.g., when constructing a new BVH or other acceleration data structure).

[0072] Miss - Indicates that the ray does not hit all the geometry in the scene or a specific region of the scene.

[0073] Visit - Indicates the child volume that the ray is going to cross.

[0074] Exceptions - Includes various types of exception handlers (e.g., called on various error conditions).

[0075] In one embodiment, the ray tracing core 372 can be adapted to accelerate general-purpose computing operations that can be accelerated using computational techniques similar to ray intersection tests. A computational framework can be provided that enables compilation of shader programs into low-level instructions and / or primitives for performing general-purpose computing operations via the ray tracing core. Exemplary computational problems that can benefit from computational operations executed on the ray tracing core 372 include computations involving the propagation of beams, waves, rays, or particles in a coordinate space. Interactions associated with that propagation can be computed with respect to geometries or meshes in the coordinate space. For example, computations related to electromagnetic signal propagation in an environment can be accelerated through the use of instructions or primitives executed by the ray tracing core. Diffraction and reflection of signals by objects in the environment can be computed as a direct ray tracing analogy.

[0076] The ray tracing core 372 can also be used to perform calculations that are not directly similar to ray tracing. For example, using the ray tracing core 245, mesh projection, mesh refinement, and volume sampling calculations can be accelerated. General coordinate space calculations such as nearest neighbor calculations can also be performed. For example, a set of points near a given point can be discovered by defining a bounding box in the coordinate space around that point. Then, using the BVH and ray probe logic within the ray tracing core 372, the set of point intersections within the bounding box can be determined. Those intersections constitute the origin and the one nearest to the origin. The calculations performed using the ray tracing core 372 can be executed in parallel with the calculations performed by the graphics core 370 and the tensor core 371. The shader compiler can be configured to compile a compute shader or other general-purpose graphics processing program into low-level primitives that can be parallelized across the graphics core 370, the tensor core 371, and the ray tracing core 372.

[0077] Techniques related to GPU-to-host processor interconnect FIG. 4A shows an exemplary architecture in which a plurality of GPUs 410-413, such as the parallel processor 200 shown in FIG. 2A, are communicatively coupled to a plurality of multi-core processors 405-406 over high-speed links 440A-440D (e.g., a bus, a point-to-point interconnect, etc.). The high-speed links 440A-440D can support a communication throughput of 4 GB / s, 30 GB / s, 80 GB / s or more, depending on the implementation. Various interconnect protocols can be used, including but not limited to PCIe 4.0 or 5.0 and NVLink 2.0. However, the basic principles described herein are not limited to a particular communication protocol or throughput.

[0078] Two or more of GPUs 410-413 can be interconnected on high-speed links 442A-442B that can be implemented using the same or different protocols / links as those used for high-speed links 440A-440D. Similarly, two or more of multi-core processors 405-406 can be connected on high-speed link 443 that can be a symmetric multi-processor (SMP) bus operating at a speed of 20 GB / s, 30 GB / s, 120 GB / s, or less or more. Alternatively, all communication between the various system components shown in FIG. 4A may be achieved using the same protocol / link (e.g., on a common interconnect fabric). However, as described above, the basic principles described herein are not limited to a particular type of interconnect technology.

[0079] Each multi-core processor 405-406 can be communicatively coupled to processor memories 401-402 via memory interconnects 430A-430B, respectively, and each GPU 410-413 can be communicatively coupled to GPU memories 420-423 via GPU memory interconnects 450A-450D, respectively. Memory interconnects 430A-430B and 450A-450D can utilize the same or different memory access technologies. By way of example and not limitation, processor memories 401-402 and GPU memories 420-423 may be volatile memories such as, for example, dynamic random access memory (DRAM) (including stacked DRAM), graphics double data rate SDRAM (GDDR) (e.g., GDDR5, GDDR6), or high bandwidth memory (HBM), and / or may be non-volatile memories such as, for example, 3D XPoint / Optan or Nano-Ram. For example, some of these memories may be volatile memories and some may be non-volatile memories (e.g., using a two-level memory (2LM) hierarchy). The memory subsystem described herein can be compatible with many memory technologies such as, for example, double data rate versions released by JEDEC (Joint Electron Device Engineering Council).

[0080] As will be described below, although the various processors 405-406 and GPUs 410-413 can each be physically coupled to specific memories 401-402, 420-423, a unified memory architecture may be implemented in which the same virtual system address space (also referred to as the “effective address” space) is distributed among all of the various physical memories. For example, each of the processor memories 401-402 can have a 64 GB system memory address space, and each of the GPU memories 420-423 can have a 32 GB system memory address space (in this example, a total of 256 GB of addressable memory is obtained).

[0081] FIG. 4B illustrates further details of an option regarding the interconnection between the multi-core processor 407 and the graphics acceleration module 446. The graphics acceleration module 446 can include one or more GPU chips integrated on a line card coupled to the processor 407 via a high-speed link 440. Alternatively, the graphics acceleration module 446 may be integrated on the same package or chip as the processor 407.

[0082] The illustrated processor 407 includes a plurality of cores 460A - 460D, each having a translation lookaside buffer 461A - 461D and one or more caches 462A - 462D. The cores are not shown to avoid obscuring the basic principles of the components described herein (e.g., instruction fetch unit, branch prediction unit, decoder, execution unit, reorder buffer, etc.), but may include various other components for executing instructions and processing data. Caches 462A - 462D may have level 1 (L1) and level 2 (L2) caches. Further, one or more shared caches 456 may be included in the cache hierarchy and shared by a set of cores 460A - 460D. For example, one embodiment of processor 407 includes 24 cores each having its own L1 cache, 12 shared L2 caches, and 12 shared L3 caches. In this embodiment, one of the L2 and L3 caches is shared by two adjacent cores. Processor 407 and the graphics accelerator integration module 446 are connected to a system memory 441 which may include processor memories 401 - 402.

[0083] Coherence is maintained with respect to data and instructions stored in the various caches 462A - 462D, 456, and system memory 441 via core - to - core communication on coherence bus 464. For example, each cache may have associated cache coherence logic / circuitry to communicate on coherence bus 464 in response to a detected read or write to a particular cache line. In one embodiment, a cache snooping protocol is implemented on coherence bus 464 to stealthily examine (snoop) cache accesses. Cache snooping / coherence techniques are well understood by those skilled in the art and will not be described in detail here so as not to obscure the basic principles described herein.

[0084] A proxy circuit 425 is provided that communicatively couples a graphics acceleration module 446 to a coherence bus 464, enabling the graphics acceleration module 446 to participate in a cache coherence protocol as a peer of these cores. In particular, an interface 435 provides a connection to the proxy circuit 425 over a high-speed link 440 (e.g., a PCIe bus, NVLink, etc.), and an interface 437 connects the graphics acceleration module 446 to the high-speed link 440.

[0085] In one embodiment, an accelerator integration circuit 436 provides cache management, memory access, context management, and interrupt management services in place of a plurality of graphics processing engines 431, 432, N of the graphics acceleration module 446. Each of the graphics processing engines 431, 432, N may have a separate graphics processing unit. Alternatively, the graphics processing engines 431, 432, N may have different types of graphics processing engines within the GPU, such as, for example, a graphics execution unit, a media processing engine (e.g., a video encoder / decoder), a sampler, and a transfer (blit) engine. In other words, the graphics acceleration module may be a GPU having a plurality of graphics processing engines 431, 432, N, or the graphics processing engines 431, 432, N may be individual GPUs integrated on a common package, line card, or chip.

[0086] The accelerator integration circuit 436 may include a memory management unit (MMU) 439 that executes various memory management functions such as virtual-to-physical memory conversion (also referred to as valid-to-real memory conversion) and a memory access protocol for accessing the system memory 441. The MMU 439 may also include a translation lookaside buffer (TLB) (not shown) that caches virtual / valid-to-physical / real address conversions. In one embodiment, the cache 438 stores commands and data for efficient access by the graphics processing engines 431, 432, N. The data stored in the cache 438 and the graphics memories 433-434 may be kept coherent with the core caches 462A-462D, 456 and the system memory 441. As described above, this may be achieved via a proxy circuit 425 that participates in the cache coherence mechanism for the cache 438 and the memories 433, 434, M (e.g., sending updates related to changes / accesses of cache lines on the processor caches 462A-462D, 456 to the cache 438 and receiving updates from the cache 438).

[0087] A set of registers 445 stores context data of threads executed by the graphics processing engines 431, 432, N, and a context management circuit 448 manages the thread contexts. For example, the context management circuit 448 may perform save and restore operations for saving and restoring the contexts of various threads in a context switch (e.g., the first thread is saved and the second thread is restored so that the second thread can be executed by the graphics processing engine). For example, on a context switch, the context management circuit 448 may store the current register values in a specified area in memory (e.g., specified by a context pointer). And it may restore the register values when returning to the context. The interrupt management circuit 447 may receive and process interrupts received from system devices, for example.

[0088] In one implementation, the virtual / valid address from the graphics processing engine 431 is translated by the MMU 439 into a real / physical address within the system memory 441. Optionally, the accelerator integration circuit 436 supports a plurality of (e.g., 4, 8, 16) graphics accelerator modules 446 and / or other accelerator devices. The graphics accelerator module 446 may be dedicated to a single application running on the processor 407 or may be shared among multiple applications. Optionally, a virtualized graphics execution environment is provided in which the resources of the graphics processing engines 431, 432, N are shared among multiple applications, virtual machines (VMs), or containers. The resources may be subdivided into "slices" that are allocated to different VMs and / or applications based on the processing requirements and priorities associated with the VMs and / or applications. VMs and containers may be used interchangeably herein.

[0089] A virtual machine (VM) can be software that runs an operating system and one or more applications. A VM can be defined by a specification, a configuration file, a virtual disk file, a non-volatile random access memory (NVRAM) configuration file, and a log file, and is supported by the physical resources of a host computing platform. A VM can include an operating system (OS) or application environment installed in software that emulates dedicated hardware. An end user can have the same experience on a virtual machine as they would on dedicated hardware. Special software called a hypervisor fully emulates the CPU, memory, hard disk, network, and other hardware resources of a PC client or server, enabling virtual machines to share resources. A hypervisor can emulate multiple virtual hardware platforms isolated from each other, enabling Linux®, Windows® Server, VMware ESXi, and other operating systems to run on the same physical host underlying the virtual machines.

[0090] Containers can be software packages for applications, configurations, and dependencies, so that those applications can operate reliably on one computing environment against another computing environment. Containers can share the operating system installed on the server platform and run as isolated processes. A container can be a software package that houses everything necessary for software to run, such as system tools, libraries, and settings. Containers are not installed like traditional software programs and allow containers to be isolated from other software and the operating system itself. The isolation of containers provides several benefits. First, the software within a container will operate the same in different environments. For example, a container containing PHP and MySQL can operate in the same way on both a Linux® computer and a Windows® machine. Second, containers provide additional security because the software does not affect the host operating system. Installed applications can change system settings and modify resources such as the Windows registry, but a container can only change the settings within the container.

[0091] Accordingly, the accelerator integration circuit 436 acts as a bridge to the system for the graphics acceleration module 446, providing address translation and system memory cache services. In one embodiment, to facilitate the bridging function, the accelerator integration circuit 436 may also include shared I / O 497 (e.g., PCIe, USB, or others), and hardware for enabling system control of voltage, clocking, performance, heat, and security. The shared I / O 497 may utilize separate physical connections or cross the high-speed link 440. Additionally, the accelerator integration circuit 436 may provide virtualization functions for the host processor to manage virtualization of the graphics processing engine, interrupts, and memory management.

[0092] Since the hardware resources of the graphics processing engines 431, 432, N are explicitly mapped to the physical address space seen by the host processor 407, any host processor can directly address those resources using valid address values. One optional function of the accelerator integration circuit 436 is to physically separate the graphics processing engines 431, 432, N so that they appear to the system as independent units.

[0093] One or more graphics memories 433, 434, M may each be coupled to each of the graphics processing engines 431, 432, N. The graphics memories 433, 434, M store the instructions and data processed by each of the graphics processing engines 431, 432, N. The graphics memories 433, 434, M may be volatile memories such as, for example, DRAM (including stacked DRAM), GDDR memory (e.g., GDDR5, GDDR6), or HBM, and / or non-volatile memories such as, for example, 3D XPoint / Optane, Samsung Z-NAND, or Nano-Ram.

[0094] To reduce data traffic on the high-speed link 440, biasing techniques can be used to ensure that the data stored in the graphics memories 433, 434, M is data that will be most frequently used by the graphics processing engines 431, 432, N and preferably not (or at least not as frequently) used by the cores 460A - 460D. Similarly, this biasing mechanism attempts to keep data that is required by the cores (and preferably not required by the graphics processing engines 431, 432, N) in the core caches 462A - 462D, 456 and the system memory 441.

[0095] According to a variation shown in FIG. 4C, the accelerator integration circuit 436 is integrated within the processor 407. The graphics processing engines 431, 432, N communicate directly with the accelerator integration circuit 436 over the high-speed link 440 via the interface 437 and the interface 435 (which can also use any form of bus or interface protocol). The accelerator integration circuit 436 can have a higher throughput given its proximity to the coherence bus 464 and the caches 462A - 462D, 456, and can perform the same operations as described with respect to FIG. 4B.

[0096] The described embodiments can support multiple different programming models, including a dedicated process programming model (without graphics acceleration module virtualization) and a shared programming model (with virtualization). The latter can include a programming model controlled by the accelerator integration circuit 436 and a programming model controlled by the graphics acceleration module 446.

[0097] In an embodiment of the dedicated process model, the graphics processing engines 431, 432, …, N can be dedicated to a single application or process under a single operating system. That single application can stream other application requests into the graphics processing engines 431, 432, …, N and provide virtualization within the VM / partition.

[0098] In the dedicated process programming model, the graphics processing engines 431, 432, N can be shared by multiple VM / application partitions. The shared model requires a system hypervisor that virtualizes the graphics processing engines 431, 432, N to enable access by each operating system. In the case of a single partition system without a hypervisor, the graphics processing engines 431, 432, N are owned by the operating system. In either case, the operating system can virtualize the graphics processing engines 431, 432, N to provide access to each process or application.

[0099] In the case of the shared programming model, the graphics acceleration module 446 or the individual graphics processing engines 431, 432, N select process elements using a process handle. The process elements are stored in the system memory 441 and can be addressed using the effective address-to-physical address translation techniques described herein. The process handle can be an implementation-specific value provided to the host process when registering its context with the graphics processing engines 431, 432, N (i.e., calling system software to add the process element to the process element link list). The lower 16 bits of the process handle can be the offset of the process element within the process element link list.

[0100] FIG. 4D shows an exemplary accelerator integration slice 490. As used herein, a "slice" has a particular portion of the processing resources of the accelerator integration circuit 436. An application valid address space 482 within system memory 441 stores a process element 483. The process element 483 may be stored in response to a GPU call 481 from an application 480 running on the processor 407. The process element 483 includes the process state of the corresponding application 480. A work descriptor (WD) 484 included in the process element 483 may be a single job requested by the application, or may include a pointer to a queue of jobs. In the latter case, the WD 484 is a pointer to the job request queue in the application's address space 482.

[0101] The graphics acceleration module 446 and / or the individual graphics processing engines 431, 432, N can be shared by all or a subset of the processes in the system. For example, the techniques described herein may include an infrastructure for setting up a process state and sending the WD 484 to the graphics acceleration module 446 to initiate a job in a virtualized environment.

[0102] In one implementation, the dedicated process programming model is implementation-specific. In this model, a single process owns the graphics acceleration module 446 or an individual graphics processing engine 431. Since the graphics acceleration module 446 is owned by a single process, the hypervisor initializes the accelerator integration circuit 436 for the owning partition, and the operating system initializes the accelerator integration circuit 436 for the process that owns the graphics acceleration module 446 when it is assigned.

[0103] In operation, the WD fetch unit 491 within the accelerator integration slice 490 fetches the next WD 484 that contains an indication indicating work to be performed by one of the graphics processing engines of the graphics acceleration module 446. The data from the WD 484 is stored in the register 445 and used by the MMU 439, the interrupt management circuit 447, and / or the context management circuit 448 as shown. For example, the MMU 439 may include a segment / page walk circuit for accessing the segment / page table 486 within the OS virtual address space 485. The interrupt management circuit 447 may process an interrupt event (INT) 492 received from the graphics acceleration module 446. When executing a graphics operation, the effective address 493 generated by the graphics processing engines 431, 432, N is translated into a physical address by the MMU 439.

[0104] The same set of registers 445 is replicated for each of the graphics processing engines 431, 432, N and / or the graphics acceleration module 446 and can be initialized by the hypervisor or the operating system. Each of these replicated registers can be included in the accelerator integration slice 490. In one embodiment, each of the graphics processing engines 431, 432, N can be presented to the hypervisor 496 as a separate graphics processor device. QoS settings can be set for the clients of a particular graphics processing engine 431, 432, N, enabling data isolation between those clients of each engine. Exemplary registers that can be initialized by the hypervisor are shown in Table 1.

Table 1

[0105] Exemplary registers that can be initialized by the operating system are shown in Table 2.

Table 2

[0106] Each WD484 may be specific to a particular graphics acceleration module 446 and / or graphics processing engines 431, 432, N. It includes all the information necessary for the graphics processing engines 431, 432, N to perform their work, or it may be a pointer to a memory location that sets a command queue for the work to be done by the application.

[0107] Figure 4E illustrates further details of the shared model option. It includes a hypervisor physical address space 498 in which a process element list 499 is stored. The hypervisor physical address space 498 is accessible via a hypervisor 496 that virtualizes the graphics acceleration module engine for the operating system 495.

[0108] The shared programming model enables all or a subset of processes from all or a subset of partitions within the system to use the graphics acceleration module 446. There are two programming models in which the graphics acceleration module 446 is shared by multiple processes and partitions: time-slice sharing and graphics-oriented sharing.

[0109] In this model, the system hypervisor 496 owns the graphics acceleration module 446 and makes its functions available to all operating systems 495. In order for the graphics acceleration module 446 to support virtualization by the system hypervisor 496, the graphics acceleration module 446 may comply with the following requirements: 1) The job requests of the application must be autonomous (i.e., the state does not need to be maintained between jobs), or the graphics acceleration module 446 must provide a context save and restore mechanism; 2) The job requests of the application are guaranteed by the graphics acceleration module 446 to complete within a specified amount of time, including conversion errors, or the graphics acceleration module 446 provides the ability to prefetch job processing; 3) The graphics acceleration module 446 must ensure fairness between processes when operating in the specified shared programming model.

[0110] In the shared model, application 480 may be required to make an operating system 495 system call using a graphics acceleration module 446 type, a work descriptor (WD), an authority mask register (AMR) value, and a context save / restore area pointer (CSRP). The graphics acceleration module 446 type describes the target acceleration function for the system call. The graphics acceleration module 446 type may be a system-specific value. The WD is specifically formatted for the graphics acceleration module 446 and may be in the form of a graphics acceleration module 446 command, a valid address pointer to a user-defined structure, a valid address pointer to a command queue, or other data structure that describes the work to be done by the graphics acceleration module 446. In one embodiment, the AMR value is the AMR state used by the current process. The value passed to the operating system is the same as the application that sets the AMR. If the implementation of the accelerator integration circuit 436 and the graphics acceleration module 446 does not support the user authority mask override register (UAMOR), the operating system may apply the current UAMOR value to the AMR value before passing the AMR in a hypervisor call. The hypervisor 496 is optional and may apply the current authority mask override register (AMOR) value before placing the AMR in the process element 483. The CSRP may be one of the registers 445 that contains the valid address of an area within the application's address space 482 for the graphics acceleration module 446 to save and restore the context state. This pointer is optional if there is no need to save the state between jobs or when the job is preempted. The context save / restore area may be pinned system memory.

[0111] Upon receiving a system call, the operating system 495 may verify that the application 480 is registered and has been granted permission to use the graphics acceleration module 446. The operating system 495 then calls the hypervisor 496 using the information shown in Table 3.

Table 3

[0112] Upon receiving a hypervisor call, the hypervisor 496 verifies that the operating system 495 is registered and has been granted permission to use the graphics acceleration module 446. The hypervisor 496 then inserts the process element 483 into the process element link list for the corresponding graphics acceleration module 446 type. The process element may include the information shown in Table 4.

Table 4

[0113] The hypervisor may initialize the plurality of accelerator integration slice 490 registers 445.

[0114] As illustrated in FIG. 4F, in one implementation of the option, unified memory that can be addressed via a common virtual memory address space used to access physical processor memories 401-402 and GPU memories 420-423 is used. In this implementation, operations executed on GPUs 410-413 utilize the same virtual / effective memory address space to access processor memories 401-402, and vice versa, thereby simplifying programmability. A first portion of the virtual / effective address space can be allocated to processor memory 401, a second portion to a second processor memory 402, a third portion to GPU memory 420, and so on. The entirety of the virtual / effective memory space (sometimes referred to as the effective address space) can thereby be distributed across each of processor memories 401-402 and GPU memories 420-423, enabling any processor or GPU to access any physical memory using the virtual address mapped to that memory.

[0115] One or more of MMUs 439A-439E may be provided with bias / coherence management circuits 494A-494E to implement a biasing technique that indicates the physical memory in which a particular type of data should be stored to ensure cache coherence between the caches of host processors (e.g., 405) and GPUs 410-413. Multiple instances of bias / coherence management circuits 494A-494E are shown in FIG. 4F, but the bias / coherence circuit may be implemented within the MMU of one or more host processors 405 and / or within accelerator integration circuit 436.

[0116] The GPU-attached memories 420-423 can be mapped as part of the system memory, can be accessed using shared virtual memory (SVM) technology without suffering the typical performance drawbacks associated with full system cache coherence, and can be accessed without incurring the onerous cache coherence overheads associated with accessing system memory. The ability of the GPU-attached memories 420-423 to be accessed as system memory without onerous cache coherence overheads provides a beneficial operating environment for GPU offload. This configuration allows the host processor 405 software to set up operands and access computation results without the overhead of traditional I / O DMA data copies. Such traditional copies involve driver calls, interrupts, and memory-mapped I / O (MMIO) accesses, all of which are less efficient than simple memory accesses. At the same time, the ability to access the GPU-attached memories 420-423 without cache coherence overheads can be very important for the execution time of offload computations. In the case of substantial streaming write memory traffic, for example, cache coherence overheads can significantly reduce the effective write bandwidth seen by the GPUs 410-413. The efficiency of operand setup, result access, and GPU computation all play a role in determining the effectiveness of GPU offload.

[0117] The choice between GPU bias and host processor bias can be driven by a bias tracker data structure. For example, a bias table can be used that has a page granularity structure (i.e., controlled at the granularity of memory pages) that includes one or two bits per page of the GPU-attached memory. The bias table can be implemented within the stolen memory ranges of one or more of the GPU-attached memories 420-423, with or without using a bias cache within the GPUs 410-413 (e.g., caching the frequently / most recently used entries of the bias table). Alternatively, the entire bias table may be maintained within the GPU.

[0118] In one implementation, bias table entries associated with each access to the GPU-attached memories 420-423 are accessed prior to the actual access to the GPU memory, causing the following operations. First, local requests from the GPUs 410-413 to find their pages at the GPU bias are transferred directly to the corresponding GPU memories 420-423. Local requests from the GPUs to find their pages at the host bias are transferred to the processor 405 (e.g., over the high-speed link described above). Optionally, requests from the processor 405 to find the requested page at the host processor bias complete the request as a normal memory read. Alternatively, requests directed to a GPU-biased page may be transferred to the GPUs 410-413. And the GPU may migrate the page to the host processor bias if it is not currently in use.

[0119] The bias state of a page can be changed by any of a software-based mechanism, a software-based mechanism assisted by hardware, or, in limited cases, a purely hardware-based mechanism.

[0120] One mechanism for changing the bias state uses an API call (e.g., OpenCL), which in turn calls the GPU's device driver, which in turn sends a message (or adds a command descriptor to the queue) to the GPU instructing it to change the bias state and, in some migrations, perform a cache flush operation within the host. The cache flush operation is necessary for the migration from the host processor 405 bias to the GPU bias but not for the reverse migration.

[0121] Cache coherence can be maintained by temporarily making GPU-biased pages non-cacheable by host processor 405. To access these pages, processor 405 may request access from GPU 410, which may or may not immediately grant access depending on the implementation. Thus, to reduce communication between host processor 405 and GPU 410, it is beneficial to ensure that GPU-biased pages are pages that are needed by the GPU but not by host processor 405, and vice versa.

[0122] Graphics processing pipeline FIG. 5 illustrates a graphics processing pipeline 500. Graphics multiprocessors such as the graphics multiprocessor 234 as in FIG. 2D, the graphics multiprocessor 325 of FIG. 3A, the graphics multiprocessor 350 of FIG. 3B, etc. may implement the illustrated graphics processing pipeline 500. The graphics multiprocessor may be related to, for example, the (one or more) parallel processors 112 of FIG. 1 and may be included within the parallel processing subsystem described herein, such as the parallel processor 200 of FIG. 2A which may be used in place of one of them. Various parallel processing systems may implement the graphics processing pipeline 500 via one or more instances of the parallel processing units (e.g., the parallel processing unit 202 of FIG. 2A) described herein. For example, a shader unit (e.g., the graphics multiprocessor 234 of FIG. 2C) may be configured to execute one or more functions of the vertex processing unit 504, the tessellation control processing unit 508, the tessellation evaluation processing unit 512, the geometry processing unit 516, and the fragment / pixel processing unit 524. The functions of the data assembler 502, the primitive assemblers 506, 514, 518, the tessellation unit 510, the rasterizer 522, and the raster operation unit 526 may also be performed by other processing engines within a processing cluster (e.g., the processing cluster 214 of FIG. 2A) and the corresponding partitioning units (e.g., the partitioning units 220A-220N of FIG. 2A). The graphics processing pipeline 500 may also be implemented using dedicated processing units for one or more functions. One or more portions of the graphics processing pipeline 500 may also be executed by parallel processing logic within a general-purpose processor (e.g., a CPU). Optionally, one or more portions of the graphics processing pipeline 500 may access on-chip memory (e.g., the parallel processor memory 222 as in FIG. 2A) via a memory interface 528 which may be an instance of the memory interface 218 of FIG. 2A.The graphics processor pipeline 500 may also be implemented via a multi-core group 365A as in FIG. 3C.

[0123] The data assembler 502 is a processing unit that can collect vertex data and primitives of a surface. And the data assembler 502 outputs vertex data including vertex attributes to the vertex processing unit 504. The vertex processing unit 504 is a programmable execution unit that executes a vertex shader program to illuminate and transform vertex data as specified by the vertex shader program. The vertex processing unit 504 reads data stored in a cache, local memory, or system memory for use when processing vertex data, and can also be programmed to convert vertex data from an object-based coordinate representation to a world space coordinate space or a normalized device coordinate space.

[0124] A first instance of the primitive assembler 506 receives vertex attributes from the vertex processing unit 504. The primitive assembler 506 reads vertex attributes stored as necessary and constructs graphics primitives for processing by the tessellation control processing unit 508. Graphics primitives include triangles, line segments, points, patches, etc., as supported by various graphics processing application programming interfaces (APIs).

[0125] The tessellation control processing unit 508 treats the input vertices as control points of a geometric patch. The control points are converted from the input representation from the patch (e.g., the base of the patch) to a representation suitable for use in surface evaluation by the tessellation evaluation processing unit 512. The tessellation control processing unit 508 can also calculate tessellation factors for the edges of the geometric patch. The tessellation factor is applied to a single edge and quantifies the view-dependent level of detail associated with that edge. The tessellation unit 510 receives the tessellation factors for the edges of the patch and is configured to tessellate the patch into a plurality of geometric primitives, such as lines, triangles, or quadrilateral primitives, which are sent to the tessellation evaluation processing unit 512. The tessellation evaluation processing unit 512 operates on the parameterized coordinates of the subdivided patch and generates surface representations and vertex attributes for each vertex associated with the geometric primitive.

[0126] A second instance of the primitive assembler 514 receives vertex attributes from the tessellation evaluation processing unit 512, reads out vertex attributes stored as necessary, and constructs graphics primitives for processing by the geometry processing unit 516. The geometry processing unit 516 is a programmable execution unit that executes a geometry shader program to transform the graphics primitive received from the primitive assembler 514 as specified by the geometry shader program. The geometry processing unit 516 can be programmed to subdivide the graphics primitive into one or more new graphics primitives and calculate the parameters used to rasterize those new graphics primitives.

[0127] The geometry processing unit 516 may be able to add or remove elements within the geometry stream. The geometry processing unit 516 outputs parameters and vertices that define a new graphics primitive to the primitive assembler 518. The primitive assembler 518 receives the parameters and vertices from the geometry processing unit 516 and constructs a graphics primitive for processing by the viewport scale, cull, and clip unit 520. The geometry processing unit 516 reads data stored in the parallel processor memory or the system memory for use in processing geometry data. The viewport scale, cull, and clip unit 520 performs clipping, culling, and viewport scaling and outputs the processed graphics primitive to the rasterizer 522.

[0128] The rasterizer 522 can perform depth culling and other depth-based optimizations. The rasterizer 522 also performs scan conversion on new graphics primitives to generate fragments and outputs those fragments and associated coverage data to the fragment / pixel processing unit 524. The fragment / pixel processing unit 524 is a programmable execution unit configured to execute a fragment shader program or a pixel shader program. The fragment / pixel processing unit 524 transforms the fragments or pixels received from the rasterizer 522 as specified by the fragment or pixel shader program. For example, the fragment / pixel processing unit 524 can be programmed to perform operations including, but not limited to, texture mapping, shading, blending, texture correction, and depth correction to generate shaded fragments or pixels that are output to the raster operation unit 526. The fragment / pixel processing unit 524 can read data stored in either parallel processor memory or system memory for use when processing fragment data. The fragment or pixel shader program can be configured to perform shading at the sample, pixel, tile, or other granularity according to the sampling rate set for the processing unit.

[0129] The raster operation unit 526 performs raster operations including, but not limited to, stenciling, z-testing, blending, and the like, and outputs pixel data as processed graphics data to be stored in graphics memory (e.g., parallel processor memory 222 as shown in FIG. 2A and / or system memory 104 as shown in FIG. 1), or to be displayed on one or more display devices 110, or for further processing by one of one or more processors 102 or (one or more) parallel processors 112. The raster operation unit 526 may be configured to compress z or color data written to memory and decompress z or color data read from memory.

[0130] Overview of Machine Learning The above architecture can be applied to perform training and inference operations using a machine learning model. Machine learning has been successful in solving many types of tasks. The computations that occur when training and using machine learning algorithms (e.g., neural networks) are inherently suitable for efficient parallel implementation. Thus, parallel processors such as general-purpose graphics processing units (GPGPUs) have played an important role in the practical implementation of deep neural networks. Parallel graphics processors having a single instruction multiple thread (SIMT) architecture are designed to maximize the amount of parallel processing in a graphics pipeline. In the SIMT architecture, groups of parallel threads attempt to execute program instructions in synchronization with each other as frequently as possible to increase processing efficiency. The efficiency provided by parallel machine learning algorithm implementations enables the use of large networks and enables those networks to be trained on larger datasets.

[0131] A machine learning algorithm is an algorithm that can learn based on a set of data. For example, a machine learning algorithm can be designed to model high-level abstraction concepts within a dataset. For example, using an image recognition algorithm, it can be determined to which of several categories a given input belongs; a regression algorithm can output a numerical value given an input; using a pattern recognition algorithm, a translated text can be generated, or text-to-speech conversion and / or speech recognition can be performed.

[0132] An exemplary type of machine learning algorithm is a neural network. There are many types of neural networks, and a simple type of neural network is a feedforward network. A feedforward network can be implemented as an acyclic graph in which nodes are arranged in layers. Typically, a feedforward network topology includes an input layer and an output layer separated by at least one hidden layer. The hidden layer transforms the input received by the input layer into a representation useful for generating the output in the output layer. Network nodes are fully connected to the nodes of adjacent layers via edges, but there are no edges between nodes within each layer. The data received at the nodes of the input layer of the feedforward network is propagated (i.e., "fed forward") to the nodes of the output layer via an activation function that calculates the state of the nodes of each successive layer within the network based on the coefficients ("weights") associated with each of the edges connecting the layers. Depending on the particular model represented by the algorithm being executed, the output from the neural network algorithm can take various forms.

[0133] Before a machine learning algorithm can be used to model a particular problem, the algorithm is trained using a training data set. Training a neural network involves selecting a network topology, using a set of training data that represents the problem being modeled by the network, and adjusting the weights until the network model performs with minimal error for all instances of the training data set. For example, in a supervised learning training process for a neural network, the output generated by the network in response to an input representing an instance within the training data set is compared to the output labeled as "correct" for that instance, an error signal representing the difference between the output and the labeled output is calculated, and the weights associated with the connections are adjusted to minimize the error as the error signal is backpropagated through the layers of the network. When the error for each output generated from an instance of the training data set is minimized, the network is considered "trained".

[0134] The accuracy of a machine learning algorithm can be significantly affected by the quality of the data set used to train the algorithm. The training process is computationally intensive and can require a large amount of time on conventional general-purpose processors. Therefore, parallel processing hardware is used to train many types of machine learning algorithms. This is particularly useful for optimizing the training of neural networks because the calculations performed when adjusting the coefficients within a neural network are inherently suitable for parallel implementation. Specifically, many machine learning algorithms and software applications have been adapted to use the parallel processing hardware within general-purpose graphics processing devices.

[0135] FIG. 6 is a generalized diagram of a machine learning software stack 600. The machine learning application 602 can be any logic configured to train a neural network using a training data set or to implement machine intelligence using a trained deep neural network. The machine learning application 602 can include the training and inference functions of the neural network and / or special software that can be used to train the neural network before deployment. The machine learning application 602 can implement any type of machine intelligence including, but not limited to, image recognition, mapping and localization, autonomous navigation, speech synthesis, medical imaging, or language translation. Examples of the machine learning application 602 include, but are not limited to, voice-based virtual assistants, image or face recognition algorithms, autonomous navigation, and software tools used to train the machine learning models used by the machine learning application 602.

[0136] Hardware acceleration for the machine learning application 602 can be enabled via the machine learning framework 604. The machine learning framework 604 can provide a library of machine learning primitives. Machine learning primitives are the basic operations typically executed by machine learning algorithms. Without the machine learning framework 604, developers of machine learning algorithms would be required to create and optimize the main computational logic associated with the machine learning algorithm and then re-optimize the computational logic when new parallel processors are developed. Instead, this machine learning application can be configured to perform the necessary computations using the primitives provided by the machine learning framework 604. Typical primitives include tensor convolution, activation functions, and pooling, which are computational operations performed when training a convolutional neural network (CNN). The machine learning framework 604 can also provide primitives that implement basic linear algebra subprograms executed by many machine learning algorithms, such as matrix and vector operations. Examples of the machine learning framework 604 include, but are not limited to, TensorFlow, TensorRT, PyTorch, MXNet, Caffee, and other high-level machine learning frameworks.

[0137] The machine learning framework 604 can process the input data received from the machine learning application 602 to generate appropriate input to the computing framework 606. The computing framework 606 can abstract the basic instructions provided to the GPGPU driver 608 to enable the machine learning framework 604 to utilize hardware acceleration via the GPGPU hardware 610 without the machine learning framework 604 needing detailed knowledge of the architecture of the GPGPU hardware 610. Further, the computing framework 606 can enable hardware acceleration of the machine learning framework 604 across various types and generations of GPGPU hardware 610. An exemplary computing framework 606 includes a CUDA computing framework and related machine learning libraries such as, for example, the CUDA Deep Neural Network (cuDNN) library. The machine learning software stack 600 can also include a communication library or framework to support multi-GPU and multi-node computing.

[0138] GPGPU Machine Learning Acceleration FIG. 7 illustrates a general purpose graphics processing unit 700 that may be the parallel processor 200 of FIG. 2A or the (one or more) parallel processors 112 of FIG. 1. The general purpose graphics processing unit (GPGPU) 700 may be configured to provide support for hardware acceleration of primitives provided by a machine learning framework to accelerate the processing of the types of computational workloads associated with training a deep neural network. Further, the GPGPU 700 can be directly linked to other instances of GPGPU to create a multi-GPU cluster, particularly to improve the training speed of deep neural networks. Primitives for accelerating the inference operation of a deployed neural network are also supported.

[0139] The GPGPU 700 includes a host interface 702 that enables connection to a host processor. The host interface 702 can be a PCI Express interface. However, the host interface can also be a vendor - specific communication interface or communication fabric. The GPGPU 700 receives commands from the host processor and uses a global scheduler 704 to distribute the execution threads associated with those commands to a set of processing clusters 706A - 706H. The processing clusters 706A - 706H share a cache memory 708. The cache memory 708 can function as a high - level cache for the cache memories within the processing clusters 706A - 706H. The illustrated processing clusters 706A - 706H can correspond to the processing clusters 214A - 214N as in Figure 2A.

[0140] The GPGPU 700 includes memories 714A - 714B coupled to the processing clusters 706A - 706H via a set of memory controllers 712A - 712B. The memories 714A - 714B can include various types of memory devices, including dynamic random access memory (DRAM), or graphics random access memory such as synchronous graphics random access memory (SGRAM) including, for example, graphics double data rate (GDDR) memory. The memories 714A - 714B can also include, but are not limited to, 3D stacked memory including high - bandwidth memory (HBM).

[0141] Each of the processing clusters 706A-706H may include a set of graphics multiprocessors such as, for example, the graphics multiprocessor 234 of FIG. 2D, the graphics multiprocessor 325 of FIG. 3A, the graphics multiprocessor 350 of FIG. 3B, or may include multicore groups 365A-365N as in FIG. 3C. The graphics multiprocessors of the compute cluster include multiple types of integer and floating-point logic units capable of performing compute operations with a range of precision including the precision suitable for machine learning computations. For example, at least a subset of the floating-point units within each of the processing clusters 706A-706H can be configured to perform 16-bit or 32-bit floating-point operations, and different subsets of those floating-point units can be configured to perform 64-bit floating-point operations.

[0142] Multiple instances of the GPGPU 700 can be configured to operate as a compute cluster. The communication mechanisms used by the compute cluster for synchronization and data exchange vary depending on the embodiment. For example, those multiple instances of the GPGPU 700 communicate via the host interface 702. In one embodiment, the GPGPU 700 includes an I / O hub 709 that couples the GPGPU 700 to a GPU link 710 that enables a direct connection to other instances of the GPGPU. The GPU link 710 can be coupled to a dedicated GPU-to-GPU bridge that enables communication and synchronization between multiple instances of the GPGPU 700. Optionally, the GPU link 710 is coupled to a high-speed interconnect for sending and receiving data to and from other GPGPUs or parallel processors. Multiple instances of the GPGPU 700 can be placed in separate data processing systems and communicate via a network device accessible via the host interface 702. The GPU link 710 may be configured to enable a connection to the host processor in addition to, or instead of, the host interface 702.

[0143] The illustrated configuration of the GPGPU 700 can be configured to train neural networks, but alternative configurations of the GPGPU 700 can be configured for deployment within a high-performance or low-power inference platform. In the inference configuration, the GPGPU 700 includes fewer processing clusters 706A-706H than in the training configuration. Also, the memory technology associated with memories 714A-714B can differ between the inference configuration and the training configuration. In one embodiment, the GPGPU 700 in the inference configuration can support instructions specific to inference. For example, the inference configuration can provide support for one or more 8-bit integer dot product instructions that are commonly used in the inference operation of a deployed neural network.

[0144] FIG. 8 illustrates a multi-GPU computing system 800. The multi-GPU computing system 800 can include a processor 802 coupled to a plurality of GPGPUs 806A-806D via a host interface switch 804. The host interface switch 804 can be a PCI Express switch device that couples the processor 802 to a PCI Express bus through which the processor 802 can communicate with a set of GPGPUs 806A-806D. Each of the plurality of GPGPUs 806A-806D can be an instance of the GPGPU 700 of FIG. 7. The GPGPUs 806A-806D can be interconnected via a GPU-to-GPU link 816, which is a set of high-speed point-to-point links. The high-speed GPU-to-GPU link can be connected to each of the GPGPUs 806A-806D via a dedicated GPU link such as the GPU link 710 in FIG. 7, for example. The P2P GPU link 816 enables direct communication between each of the GPGPUs 806A-806D without requiring communication on the host interface bus to which the processor 802 is connected. By directing GPU-to-GPU traffic to the P2P GPU link, the host interface bus remains available for system memory access or for communication with other instances of the multi-GPU computing system 800, for example, via one or more network devices. In FIG. 8, the GPGPUs 806A-806D are shown connected to the processor 802 via the host interface switch 804, but alternatively, the processor 802 can include direct support for the P2P GPU link 816 and be directly connected to the GPGPUs 806A-806D. In one embodiment, the P2P GPU link 816 enables the multi-GPU computing system 800 to operate as a single logical GPU.

[0145] Machine learning neural network implementation The computing architecture described herein can be configured to perform a type of parallel processing particularly suitable for training and deploying neural networks for machine learning. A neural network can be generalized as a network of functions having a graph relationship. As is well known in the art, there are various types of neural network implementations used in machine learning. One exemplary type of neural network is a feedforward network, as described above.

[0146] A second exemplary type of neural network is a convolutional neural network (CNN). A CNN is a special type of feedforward neural network for processing data having a known grid-like topology, such as, for example, image data. Thus, CNNs are commonly used in computer vision applications and image recognition applications, although they can also be used for other types of pattern recognition, such as, for example, audio and language processing. The nodes in the CNN input layer are organized into a set of “filters” (feature detectors inspired by the receptive fields found in the retina), and the output of each set of filters is propagated to nodes in subsequent layers of the network. The computation of a CNN involves applying a convolution (convolution) mathematical operation to each filter to generate the output of that filter. Convolution is a special type of mathematical operation performed by two functions to produce a third function that is a modified version of one of these two original functions. In the terminology of convolutional networks, the first function for convolution is called the input, and the second function can be called the convolution kernel. The output can be called a feature map. For example, the input to a convolutional layer can be a multi-dimensional array of data defining the various color components of an input image. The convolution kernel can be a multi-dimensional array of parameters, and those parameters are adapted by the training process of the neural network.

[0147] A recurrent neural network (RNN) is a family of feedforward neural networks that includes feedback connections between layers. By sharing parameter data across different parts of the neural network, RNNs enable the modeling of sequential data. The architecture of an RNN contains cycles. A cycle represents the influence of the current value of a variable on its own value at a future point in time when at least a portion of the output data from the RNN is used as feedback for processing subsequent inputs within a sequence. This mechanism makes RNNs particularly useful for language processing due to the variability by which language data can be structured.

[0148] The figures described below present exemplary feedforward CNN and RNN networks and describe general processes for training and deploying each of these types of networks. It is to be understood that these descriptions are not limiting with respect to the specific embodiments described herein but are exemplary, and the concepts illustrated generally may be applied to deep neural networks and machine learning techniques in general.

[0149] The above-described exemplary neural networks can be used to perform deep learning. Deep learning is machine learning using deep neural networks. The deep neural networks used in deep learning are artificial neural networks composed of multiple hidden layers, in contrast to shallow neural networks that include only a single hidden layer. Deeper neural networks are generally more computationally intensive to train. However, the additional hidden layers of the network enable multi-step pattern recognition that results in a reduced output error compared to shallow machine learning techniques.

[0150] Deep neural networks used in deep learning typically include a front-end network that performs feature recognition coupled to a back-end network, where the back-end network represents a mathematical model that can perform operations (e.g., object classification, speech recognition, etc.) based on the feature representations provided to the model. Deep learning enables machine learning to be performed without the need to perform handcrafted feature engineering on the model. Instead, deep neural networks can learn features based on the statistical structure or correlations in the input data. The learned features can be provided to a mathematical model that can map the detected features to the output. The mathematical models used by the network are generally specialized for the specific task being performed, and different models will be used to perform different tasks.

[0151] Once a neural network is constructed, a learning model can be applied to the network to train the network to perform a specific task. The learning model describes how to adjust the weights within the model to reduce the output error of the network. Backpropagation of error is a common method used to train neural networks. An input vector is presented to the network for processing. The output of the network is compared to the desired output using a loss function, and an error value is calculated for each neuron in the output layer. The error values are then propagated backwards until each neuron has an associated error value that roughly represents its contribution to the original output. The network can then learn from these errors using an algorithm such as the stochastic gradient descent algorithm and update the weights of the neural network.

[0152] Figures 9A-9B show an exemplary convolutional neural network. Figure 9A illustrates various layers within the CNN. As shown in Figure 9A, an exemplary CNN used to model image processing can receive an input 902 that describes the red, green, and blue (RGB) components of an input image. The input 902 can be processed by a plurality of convolutional layers (e.g., convolutional layer 904, convolutional layer 906). The output from these multiple convolutional layers can optionally be processed by a set of fully connected layers 908. Neurons within the fully connected layers have a complete connection to all activations in the previous layer, as described above with respect to feedforward networks. Using the output from the fully connected layer 908, an output result from the network can be generated. The activations within the fully connected layer 908 can be calculated using matrix multiplication instead of convolution. Not all CNN implementations use the fully connected layer 908. For example, in some embodiments, the convolutional layer 906 can generate the output of the CNN.

[0153] Convolutional layers are sparsely connected, which is different from the traditional neural network configuration seen in the fully connected layer 908. Traditional neural network layers are fully connected such that all output units interact with all input units. However, convolutional layers are sparsely connected as the output of the convolution of the field is input to the nodes of the subsequent layer (instead of the respective state values of each of the nodes within the field) as illustrated. The kernel associated with the convolutional layer performs the convolution operation and its output is sent to the next layer. The dimensionality reduction performed within the convolutional layer is one aspect that allows the CNN to scale to process large images.

[0154] Figure 9B shows an exemplary computational stage within the convolutional layer of a CNN. The input 912 to the CNN's convolutional layer can be processed in three stages of the convolutional layer 914. Those three stages can include a convolution stage 916, a detector stage 918, and a pooling stage 920. And the convolutional layer 914 can output data to a subsequent convolutional layer. The last convolutional layer of the network can generate output feature map data or provide an input to a fully connected layer, and for example, can generate classification values for the input to the CNN.

[0155] The convolution stage 916 performs several convolutions in parallel to generate a set of linear activations. The convolution stage 916 can include an affine transformation, which is a transformation that can be defined as a linear transformation plus a translation. Affine transformations include rotation, translation, scaling, and combinations of these transformations. The convolution stage calculates the output of a function (such as a neuron, etc.) connected to a specific region within the input, and the specific region can be determined as the local region associated with that neuron. The neuron calculates the dot product between the weights of the neuron and the local input region to which the neuron is connected. The output from the convolution stage 916 defines a set of linear activations that are processed by subsequent stages of the convolutional layer 914.

[0156] Those linear activations can be processed by the detector stage 918. At the detector stage 918, each linear activation is processed by a non-linear activation function. The non-linear activation function increases the non-linearity of the whole network without affecting the receptive field of the convolutional layer. Several types of non-linear activation functions can be used. One particular type is the rectified linear unit (ReLU), which uses an activation function defined as f(x) = max(0, x) such that the activation is thresholded at zero.

[0157] The pooling stage 920 uses a pooling function that replaces the output of the convolutional layer 906 with a summary statistic of the neighborhood outputs. Using the pooling function, translational invariance can be introduced into the neural network so that small translations of the input do not change the pooled output. Invariance to local translations can be useful in scenarios where the presence of a feature in the input data is more important than the exact position of the feature. At the pooling stage 920, various types of pooling functions can be used, including max pooling, average pooling, and l2-norm pooling. Also, some CNN implementations do not include a pooling stage. Instead, such implementations are replaced with an additional convolutional stage with a larger stride than the preceding convolutional stage.

[0158] And the output from the convolutional layer 914 can be processed by the next layer 922. The next layer 922 can be either a further convolutional layer or one of the fully connected layers 908. For example, the first convolutional layer 904 in FIG. 9A can output to the second convolutional layer 906, and the second convolutional layer can output to the first of the fully connected layers 908.

[0159] Figure 10 shows an exemplary recurrent neural network 1000. In a recurrent neural network (RNN), the previous state of the network affects the output of the current state of the network. RNNs can be constructed in a variety of ways using a variety of functions. The use of RNNs generally centers around using a mathematical model to predict the future based on a previous sequence of inputs. For example, statistical language modeling can be performed using an RNN to predict a coming word given a previous sequence of words. The illustrated RNN 1000 can be described as having an input layer 1002 that receives an input vector, a hidden layer 1004 that implements a recurrence function, a feedback mechanism 1005 that enables "memory" of the previous state, and an output layer 1006 that outputs a result. The RNN 1000 operates based on time steps. The state of the RNN at a given time step is affected based on the previous time step via the feedback mechanism 1005. At a given time step, the state of the hidden layer 1004 is determined by the previous state and the input at the current time step. The initial input (x1) at the first time step can be processed by the hidden layer 1004. The second input (x2) can be processed by the hidden layer 1004 using the state information determined during the processing of the initial input (x1). A given state is s t = f(Ux t + Ws t-1 ) and can be calculated, where U and W are parameter matrices. The function f is generally non-linear, such as, for example, the hyperbolic tangent function (Tanh) or a variant of the rectifier function f(x) = max(0, x). However, the specific mathematical function used in the hidden layer 1004 can vary depending on the specific implementation details of the RNN 1000.

[0160] In addition to the basic CNN and RNN networks described, acceleration regarding variations of these networks can be enabled. One variation of the RNN is the long short term memory (LSTM) RNN. The LSTM RNN can learn the long-term dependencies that may be required to process longer language sequences. One variation for the CNN is the convolutional deep belief network that has a structure similar to the CNN and is trained in the same manner as the deep belief network. The deep belief network (DBN) is a generative neural network composed of multiple layers of probabilistic (random) variables. The DBN can be trained layer by layer using greedy unsupervised learning. And, by determining an optimal set of initial weights for the neural network using the learned weights of the DBN, a pre-trained neural network can be provided. In a further embodiment, acceleration for reinforcement learning is enabled. In reinforcement learning, an artificial agent learns by interacting with its environment. The agent is configured to optimize certain goals in order to maximize cumulative rewards.

[0161] Figure 11 illustrates the training and deployment of a deep neural network. Once a given network is structured for a task, the neural network is trained using a training data set 1102. To enable hardware acceleration of the training process, various training frameworks 1104 have been developed. For example, the machine learning framework 604 of FIG. 6 can be configured as the training framework 1104. The training framework 1104 can take in an untrained neural network 1106 and enable the untrained neural net to be trained using the parallel processing resources described herein to generate a trained neural net 1108.

[0162] To start the training process, the initial weights can be selected randomly or by pre-training using a deep belief network. Then, the training cycle is executed either with or without a teacher.

[0163] Supervised learning is a learning method in which training is performed as an intermediary operation, for example, when the training dataset 1102 includes inputs paired with their desired outputs, or when the training dataset includes inputs with known outputs and the outputs of the neural network are manually graded. The network processes the inputs and compares the resulting outputs with a set of expected or desired outputs. Then, the error is propagated backward through the system. The training framework 1104 can adjust the weights that control the untrained neural network 1106. The training framework 1104 can provide a tool to monitor how well the untrained neural network 1106 is converging towards a model suitable for generating the correct answers based on known input data. The training process is repeated as the weights of the network are adjusted to refine the outputs generated by the neural network. The training process can continue until the neural network reaches a statistically desirable accuracy associated with the trained neural network 1108. Then, the trained neural network 1108 can be deployed to perform any number of machine learning operations to generate inference results 1114 based on the input of new data 1112.

[0164] Unsupervised learning is a learning method in which a network attempts to train itself using unlabeled data. Therefore, in unsupervised learning, the training dataset 1102 will include input data without associated output data. The untrained neural network 1106 can learn to group within the unlabeled inputs and determine how individual inputs relate to the entire dataset. Using unsupervised training, a self-organized map can be generated, which is a type of trained neural network 1108 that can perform operations useful for reducing the dimensionality of the data. Unsupervised training can also be used to perform anomaly detection, which enables the identification of data points in an input dataset that deviate from the normal patterns of the data.

[0165] Variations on supervised and unsupervised training can also be used. Semi-supervised learning is a technique in which the training dataset 1102 includes a mixture of labeled and unlabeled data from the same distribution. Incremental learning is a variant of supervised learning in which input data is continuously used to further train the model. Incremental learning enables the trained neural network 1108 to adapt to new data 1112 without forgetting the knowledge that has permeated the network during initial training.

[0166] Regardless of whether it is supervised or unsupervised, the training process, especially for deep neural networks, can be overly computationally intensive for a single computing node. Instead of using a single computing node, a distributed network of computing nodes can be used to accelerate the training process.

[0167] FIG. 12A is a block diagram illustrating distributed learning. Distributed learning is a training model that performs supervised or unsupervised training of a neural network using a plurality of distributed computing nodes. Each of the distributed computing nodes can include one or more host processors and one or more of a plurality of general-purpose processing nodes, such as the highly parallel general-purpose graphics processing unit 700 as in FIG. 7. As shown, distributed learning can be performed using model parallel processing 1202, data parallel processing 1204, or parallel processing 1206 of a combination of model and data.

[0168] In model parallel processing 1202, different computing nodes within the distributed system can perform training calculations for different parts of a single network. For example, each layer of a neural network can be trained by different processing nodes of the distributed system. The advantages of model parallel processing include being able to scale particularly well to large-scale models. Dividing the calculations associated with different layers of a neural network enables the training of very large neural networks where the weights of all layers do not fit into the memory of a single computing node. In some examples, model parallel processing can be particularly useful when performing unsupervised training of large neural networks.

[0169] In data parallel processing 1204, different nodes of a distributed network have a complete instance of the model, and each node receives a different portion of the data. Then, the results from different nodes are combined. Although different approaches to data parallel processing are possible, all data parallel training approaches require techniques to synchronize model parameters among nodes and combine the results. Exemplary approaches for combining data include parameter averaging and update-based data parallel processing. Parameter averaging trains each node on a subset of the training data and sets the global parameters (e.g., weights, biases) to the average of the parameters from each node. Parameter averaging uses a central parameter server to maintain the parameter data. Update-based data parallel processing is similar to parameter averaging, except that instead of transferring parameters from the nodes to the parameter server, updates to the model are transferred. Also, update-based data parallel processing can be executed in a decentralized manner, and the updates are compressed and transferred among the nodes.

[0170] Combined model-data parallel processing 1206 can be implemented, for example, in a distributed system where each computing node includes multiple GPUs. Each node can have a complete instance of the model, and separate GPUs within each node are used to train different parts of the model.

[0171] Distributed training has increased the overhead compared to training on a single machine. However, each of the parallel processors and GPGPUs described herein can implement various techniques to reduce the overhead of distributed training, including techniques that enable wide-bandwidth GPU-to-GPU data transfer and accelerated remote data synchronization.

[0172] FIG. 12B is a block diagram illustrating a programmable network interface 1210 and a data processing unit. The programmable network interface 1210 is a programmable network engine that can be used to accelerate network-based computing tasks in a distributed environment. The programmable network interface 1210 can be coupled to a host system via a host interface 1270. Using the programmable network interface 1210, network operations or storage operations related to the CPU or GPU of the host system can be accelerated. The host system can be, for example, a node of a distributed learning system used to perform distributed training as shown in FIG. 12A. The host system can also be a data center node within a data center.

[0173] In one embodiment, access to remote storage containing model data can be accelerated by the programmable network interface 1210. For example, the programmable network interface 1210 can be configured to present a remote storage device to the host system as a local storage device. The programmable network interface 1210 can also accelerate remote direct memory access (RDMA) operations executed between the GPU of the host system and the GPU of the remote system. In one embodiment, the programmable network interface 1210 can enable storage functions such as, but not limited to, NVME-oF. The programmable network interface 1210 can also accelerate cryptographic operations, data integrity operations, compression operations, and other operations for remote storage on behalf of the host system, enabling the remote storage to approach the latency of a storage device directly connected to the host system.

[0174] The programmable network interface 1210 can also perform resource allocation and management on behalf of the host system. Storage security operations can be offloaded to the programmable network interface 1210 and executed in coordination with the allocation and management of remote storage resources. Network-based operations for managing access to remote storage, which would otherwise be executed by the processor of the host system, can instead be executed by the programmable network interface 1210.

[0175] In one embodiment, network and / or data security operations can be offloaded from the host system to the programmable network interface 1210. The data center security policy of the data center node can be handled by the programmable network interface 1210 instead of the processor of the host system. For example, the programmable network interface 1210 can detect and suppress network-based attacks (e.g., DDoS) attempted against the host system to prevent the attacks from degrading the availability of the host system.

[0176] The programmable network interface 1210 can include a system-on-chip (SoC 1220) that executes an operating system by a plurality of processor cores 1222. The processor cores 1222 can include general-purpose processor (e.g., CPU) cores. In one embodiment, the processor cores 1222 can also include one or more GPU cores. The SoC 1220 can execute instructions stored in the memory device 1240. The storage device 1250 can store local operating system data. The storage device 1250 and the memory device 1240 can also be used to cache remote data regarding the host system. The network ports 1260A - 1260B enable connection to a network or fabric and assist in network access for the SoC 1220 via the host interface 1270 regarding the host system. The programmable network interface 1210 can also include an I / O interface 1275 such as, for example, a USB interface. The I / O interface 1275 can be used to couple an external device to the programmable network interface 1210 or can be used as a debug interface. The programmable network interface 1210 also includes a management interface 1230 that enables software on the host device to manage and configure the programmable network interface 1210 and / or the SoC 1220. In one embodiment, the programmable network interface 1210 can also include one or more accelerators or GPUs 1245 for accepting offloading of parallel computing tasks from the SoC 1220, the host system, or a remote system coupled via the network ports 1260A - 1260B.

[0177] Exemplary machine learning applications Machine learning can be applied to solve a variety of technical problems including, but not limited to, computer vision, autonomous driving and navigation, speech recognition, and language processing. Computer vision has traditionally been one of the most active research areas regarding machine learning applications. The uses of computer vision range from the reproduction of human visual capabilities such as recognizing faces to the creation of new categories of visual capabilities. For example, computer vision applications can be configured to recognize sound waves from vibrations induced by visible objects in a video. Machine learning with parallel processor acceleration enables computer vision applications to be trained using much larger training data sets than previously achievable and also enables inference systems to be deployed using low-power parallel processors.

[0178] Machine learning with parallel processor acceleration has applications in autonomous driving including lane and road sign recognition, obstacle avoidance, navigation, and driving control. Using accelerated machine learning techniques, driving models can be trained based on data sets that define appropriate responses to specific training inputs. The parallel processors described herein enable the rapid training of increasingly complex neural networks used in autonomous driving solutions and also enable the deployment of low-power inference processors in mobile platforms suitable for integration into autonomous vehicles.

[0179] Deep neural networks with parallel processor acceleration enable a machine learning approach to automatic speech recognition (ASR). ASR involves generating a function that, given an input acoustic sequence, computes the most likely language sequence. Accelerated machine learning using deep neural networks enables the replacement of hidden Markov models (HMMs) and Gaussian mixture models (GMMs), which have been used in ASR.

[0180] Machine learning with parallel processor acceleration can also be used to accelerate natural language processing. An automated learning procedure can utilize statistical inference algorithms to produce a model that is robust to incorrect or unfamiliar inputs. Exemplary natural language processor applications include machine translation automatically between human languages.

[0181] The parallel processing platforms used for machine learning can be divided into a training platform and a deployment platform. The training platform is generally highly parallel and includes optimizations for accelerating single-node multi-GPU training and multi-node multi-GPU training. Exemplary parallel processors suitable for training include the general-purpose graphics processing unit 700 of FIG. 7 and the multi-GPU computing system 800 of FIG. 8. In contrast, the deployed machine learning platform generally includes low-power parallel processors suitable for use in products such as cameras, autonomous robots, and autonomous vehicles.

[0182] Also, machine learning techniques can be applied to accelerate or enhance graphics processing activities. For example, a machine learning model can be trained to recognize the output generated by a GPU-accelerated application and generate an enlarged version of that output. Such techniques can be applied to accelerate the generation of high-resolution images for gaming applications. Various other graphics pipeline activities can benefit from the use of machine learning. For example, by training a machine learning model, a tessellation operation can be performed on geometry data to increase the complexity of the geometry model, enabling the automatic generation of detailed geometry from relatively low-detail geometry.

[0183] FIG. 13 shows an exemplary inference system - on - chip (SOC) 1300 suitable for performing inference using a trained model. The SOC 1300 can integrate processing components including a media processor 1302, a vision processor 1304, a GPGPU 1306, and a multi - core processor 1308. The GPGPU 1306 can be a GPGPU as described herein, such as, for example, GPGPU 700, and the multi - core processor 1308 can be a multi - core processor as described herein, such as, for example, multi - core processors 405 - 406. The SOC 1300 can further include on - chip memory 1305 that can implement a shared on - chip data pool accessible by each of these processing components. These processing components can be optimized for low - power operation to enable deployment to a variety of machine - learning platforms including autonomous vehicles and autonomous robots. For example, one implementation of the SOC 1300 can be used as part of the main control system of an autonomous vehicle. When the SOC 1300 is configured for use in an autonomous vehicle, the SOC is designed and configured to comply with the relevant functional safety standards of the deployment jurisdiction.

[0184] In operation, the media processor 1302 and the vision processor 1304 can operate in cooperation to accelerate computer vision operations. The media processor 1302 can enable low - latency decoding of multiple high - resolution (e.g., 4K, 8K) video streams. The decoded video stream can be written to a buffer within the on - chip memory 1305. Then, the vision processor 1304 can perform preliminary processing operations on the frames of the decoded video in preparation for analyzing the decoded video and processing the frames using a trained image recognition model. For example, the vision processor 1304 can accelerate convolution operations related to a CNN used to perform image recognition on high - resolution video data, while the back - end model calculations are performed by the GPGPU 1306.

[0185] The multi-core processor 1308 can include control logic that assists in sequencing and synchronizing data transfers and shared memory operations executed by the media processor 1302 and the vision processor 1304. The multi-core processor 1308 can also function as an application processor that executes software applications that can utilize the inference computing capabilities of the GPGPU 1306. For example, at least a portion of the navigation driving logic can be implemented in software executed on the multi-core processor 1308. Such software can issue compute workloads directly to the GPGPU 1306, or the compute workloads can be issued to the multi-core processor 1308, and the multi-core processor 1308 can offload at least a portion of those operations to the GPGPU 1306.

[0186] The GPGPU 1306 can include compute clusters, such as a low-power configuration of processing clusters 706A - 706H within the general-purpose graphics processing unit 700, for example. The compute clusters within the GPGPU 1306 can support instructions that are particularly optimized for performing inference computations on trained neural networks. For example, the GPGPU 1306 can support instructions for performing low-precision computations, such as 8-bit and 4-bit integer vector operations.

[0187] Overview of the further system FIG. 14 is a block diagram of a processing system 1400. Elements of FIG. 14 that have the same or similar names as elements in any of the other figures herein describe the same elements in those other figures and can operate or function in a similar manner, can have the same components, and can be coupled to other entities as described elsewhere herein, but are not so limited. System 1400 can be used in a single-processor desktop system, a multi-processor workstation system, or a server system having a number of processors 1402 or processor cores 1407. System 1400 can be a processing platform incorporated within a system-on-chip (SoC) integrated circuit used in a mobile, handheld, or embedded device, such as within an Internet of Things (IoT) device with a wired or wireless connection to a local area network or a wide area network.

[0188] System 1400 can be a processing system having components corresponding to the components of FIG. 1. For example, in various configurations, one or more processors 1402 or one or more processor cores 1407 can correspond to one or more processors 102 of FIG. 1. One or more graphics processors 1408 can correspond to one or more parallel processors 112 of FIG. 1. An external graphics processor 1418 can be one of the one or more add-in devices 120 of FIG. 1.

[0189] System 1400 may include, be coupled with, or integrated within a server-based game platform, a game console (including a game and media console), a mobile game console, a handheld game console, or an online game console. System 1400 may also be part of a mobile Internet-connected device such as a cellular phone, smartphone, tablet computing device, or laptop with low internal storage capacity. Processing system 1400 may also include, be coupled with, or integrated within a wearable device such as, for example, a smartwatch; smart eyewear or clothing enhanced with augmented reality (AR) or virtual reality (VR) functionality that provides visual, audio, or tactile output to complement a real-world visual, audio, or tactile experience, or otherwise provides text, audio, graphics, video, holographic images or video, or tactile feedback; other augmented reality (AR) devices; or other virtual reality (VR) devices. Processing system 1400 may also include or be part of a television device or set-top box device. System 1400 may include, be coupled with, or integrated within an autonomous vehicle such as, for example, a bus, tractor trailer, automobile, motor or electric cycle, airplane or glider (or any combination thereof). The autonomous vehicle may use System 1400 to process the environment sensed around the vehicle.

[0190] One or more processors 1402 may include one or more processor cores 1407 that, when executed, process instructions to perform operations related to system software or user software. At least one of the one or more processor cores 1407 may be configured to process a particular instruction set 1409. The instruction set 1409 may support computing via a plurality of instruction set computing (CISC), reduced instruction set computing (RISC), or very long instruction word (VLIW). The one or more processor cores 1407 may process different instruction sets 1409 that may include instructions to support the emulation of other instruction sets. The processor core 1407 may also include other processing devices such as, for example, a digital signal processor (DSP).

[0191] The processor 1402 may include a cache memory 1404. Depending on the architecture, the processor 1402 may have a single internal cache or multiple levels of internal caches. In some embodiments, the cache memory is shared among various components of the processor 1402. In some embodiments, the processor 1402 may also use an external cache (e.g., a level 3 (L3) cache or a last level cache (LLC)) (not shown) that may be shared among the processor cores 1407 using known cache coherence techniques. The processor 1402 may further include a register file 1406, which may include different types of registers (e.g., integer registers, floating point registers, status registers, and instruction pointer registers) for storing different types of data. Some registers may be general purpose registers, and other registers may be specific to the design of the processor 1402.

[0192] To transmit communication signals, such as address, data, or control signals, between the processor 1402 and other components within the system 1400, one or more processors 1402 may be coupled to one or more interface buses 1410. The interface bus 1410 can be a processor bus, such as a version of the Direct Media Interface (DMI) bus, in one of these embodiments. However, the processor bus is not limited to the DMI bus and may include one or more peripheral component interconnect buses (e.g., PCI, PCI Express), memory buses, or other types of interface buses. For example, the (one or more) processors 1402 may include an integrated memory controller 1416 and a platform controller hub 1430. The memory controller 1416 facilitates communication between the memory device and other components of the system 1400, and the platform controller hub (PCH) 1430 provides connections to I / O devices via a local I / O bus.

[0193] Memory device 1420 can be a dynamic random access memory (DRAM) device, a static random access memory (SRAM) device, a flash memory device, a phase change memory device, or other memory device with performance suitable for functioning as process memory. Memory device 1420 can operate as the system memory of system 1400, for example, to store data 1422 and instructions 1421 for use when one or more processors 1402 execute an application or process. Memory controller 1416 is also coupled to an optional external graphics processor 1418 that can communicate with one or more graphics processors 1408 within processor 1402 to perform graphics and media operations. In some embodiments, operations for graphics, media, or computing may be assisted by an accelerator 1412, which can be a coprocessor configured to perform a specialized set of operations for graphics, media, or computing. For example, accelerator 1412 can be a matrix multiplication accelerator used to optimize machine learning or computing operations. Accelerator 1412 can also be a ray tracing accelerator that can be used in cooperation with graphics processor 1408 to perform ray tracing operations. In one embodiment, an external accelerator 1419 may be used instead of or in cooperation with accelerator 1412.

[0194] A display device 1411 is provided and can be connected to (one or more) processors 1402. Display device 1411 can be one or more of an internal display, such as in a mobile electronics device or a laptop device, or an external display device attached via a display interface (e.g., DisplayPort, etc.). Display device 1411 can also be a head-mounted display (HMD), such as a stereoscopic display device used in virtual reality (VR) applications or augmented reality (AR) applications.

[0195] The platform controller hub 1430 may enable peripheral devices to be connected to the memory device 1420 and the processor 1402 via a high-speed I / O bus. The I / O peripheral devices may include, but are not limited to, an audio controller 1446, a network controller 1434, a firmware interface 1428, a wireless transceiver 1426, a touch sensor 1425, and a data storage device 1424 (e.g., non-volatile memory, volatile memory, hard disk drive, flash memory, NAND, 3D NAND, 3D XPoint / Optane, etc.). The data storage device 1424 may be connected via a storage interface (e.g., SATA) or via a peripheral bus such as a peripheral component interconnect bus (e.g., PCI, PCI Express). The touch sensor 1425 may include a touch screen sensor, a pressure sensor, or a fingerprint sensor. The wireless transceiver 1426 may be a Wi-Fi (registered trademark) transceiver, a Bluetooth (registered trademark) transceiver, or a mobile network transceiver such as a 3G, 4G, 5G, or Long Term Evolution (LTE) transceiver. The firmware interface 1428 may enable communication with system firmware and may be, for example, a unified extensible firmware interface (UEFI). The network controller 1434 may enable a network connection to a wired network. In some embodiments, a high-performance network controller (not shown) is coupled to the interface bus 1410. The audio controller 1446 may be a multi-channel high-resolution audio controller. In some of these embodiments, the system 1400 includes an optional legacy I / O controller 1440 for coupling legacy (e.g., Personal System 2 (PS / 2)) devices to the system.The platform controller hub 1430 can also be connected to one or more universal serial bus (USB) controllers 1442 to connect input devices such as, for example, a combination of a keyboard and a mouse 1443, a camera 1444, or other USB input devices.

[0196] It is to be understood that the illustrated system 1400 is exemplary and not limiting, as other types of data processing systems configured differently may also be used. For example, instances of the memory controller 1416 and the platform controller hub 1430 may be integrated into a discrete external graphics processor such as, for example, the external graphics processor 1418. The platform controller hub 1430 and / or the memory controller 1416 may be external to one or more of the processors 1402. For example, the system 1400 may include an external memory controller 1416 and a platform controller hub 1430, and they may be configured as a memory controller hub and a peripheral device controller hub within a system chipset that communicates with the processor(s) 1402.

[0197] For example, a circuit board (“sled”) on which components such as a CPU, memory, and other components are arranged can be used and designed for enhanced thermal performance. For example, a processing component such as a processor can be placed on the top surface side of the sled, and near-memory such as a DIMM can be placed on the bottom surface side of the sled. As a result of the enhanced air flow provided by this design, these components can operate at higher frequencies and power levels than in a typical system, thereby enhancing performance. Further, the sled is configured to be connectable without being regarded as power cables and data communication cables within a rack, thereby enhancing the ability to quickly remove, upgrade, reinstall, and / or replace them. Similarly, individual components placed on the sled, such as a processor, accelerator, memory, and data storage drive, are configured to be easily upgradable by increasing the spacing between them. In this exemplary embodiment, these components further include a hardware authentication mechanism for proving their authenticity.

[0198] The data center can utilize a single network architecture (“fabric”) that supports multiple other network architectures including Ethernet (registered trademark) and Omni-Path. Threads can be coupled to switches via optical fibers that provide higher bandwidth and lower latency than typical twisted pair cables (e.g., Category 5, Category 5e, Category 6, etc.). The high-bandwidth and low-latency interconnects and network architecture enable the data center to pool resources such as memory, accelerators (e.g., GPUs, graphics accelerators, FPGAs, ASICs, neural networks, and / or artificial intelligence accelerators, etc.), and physically unaggregated data storage drives during use and provide them to computational resources (e.g., processors) as needed, allowing the computational resources to access the pooled resources as if they were local.

[0199] A power supply or power source may provide voltage and / or current to system 1400 or any of the components or systems described herein. In one example, the power supply includes an AC-DC (alternating current - direct current) adapter that plugs into a wall outlet. Such AC power can be a renewable energy (e.g., solar power) source. In one example, the power supply includes a DC power source such as an external AC-DC converter. The power supply or power source may also include wireless charging hardware that charges upon proximity to a charging field. The power supply can include an internal battery, an AC power source, a motion-based power source, a solar power source, or a fuel cell power source.

[0200] Figures 15A-15C illustrate a computing system and a graphics processor. Elements of Figures 15A-15C that have the same or similar names as elements in any of the other figures herein describe the same elements in those other figures and can operate or function in a similar manner, can have the same components, and can be coupled to other entities as described elsewhere herein, but are not so limited.

[0201] FIG. 15A is a block diagram of a processor 1500 that can be a variation of one of the processors 1402 and can be used in place of one of them. Thus, any disclosure of features in combination with the processor 1500 here also discloses the corresponding combination with the (one or more) processors 1402, but is not so limited. The processor 1500 can include one or more processor cores 1502A - 1502N, an integrated memory controller 1514, and an integrated graphics processor 1508. If the integrated graphics processor 1508 is excluded, a system including this processor will include a graphics processor device within the system chipset or coupled via the system bus. The processor 1500 can include additional cores up to an additional core 1502N represented by the dashed box. Each of the processor cores 1502A - 1502N includes one or more internal cache units 1504A - 1504N. In some embodiments, each processor core 1502A - 1502N also has access to one or more shared cache units 1506. The internal cache units 1504A - 1504N and the shared cache units 1506 represent the cache memory hierarchy within the processor 1500. The cache memory hierarchy can include at least one level of instruction and data cache within each processor core and one or more levels of shared intermediate level cache, such as level 2 (L2), level 3 (L3), level 4 (L4) or other levels of cache, and the highest level of cache before external memory is classified as the LLC. In some embodiments, cache coherence logic manages the coherence among the various cache units 1506 and 1504A - 1504N.

[0202] Processor 1500 may also include a set of one or more bus controller units 1516 and a system agent core 1510. The one or more bus controller units 1516 manage a set of peripheral buses, such as one or more PCI or PCI Express buses, for example. The system agent core 1510 provides management functions related to various processor components. The system agent core 1510 may include one or more integrated memory controllers 1514 that manage access to various external memory devices (not shown).

[0203] For example, one or more of the processor cores 1502A - 1502N may include support for simultaneous multithreading. The system agent core 1510 includes components for coordinating the operation of cores 1502A - 1502N during multithreaded processing. The system agent core 1510 may further include a power control unit that includes logic and components for adjusting the power states of processor cores 1502A - 1502N and graphics processor 1508.

[0204] Processor 1500 may further include a graphics processor 1508 that performs graphics processing operations. In some of these embodiments, the graphics processor 1508 couples to a set of shared cache units 1506 and a system agent core 1510 that includes one or more integrated memory controllers 1514. The system agent core 1510 may also include a display controller 1511 for driving the graphics processor output to one or more attached displays. The display controller 1511 may also be a separate module coupled to the graphics processor via at least one interconnect, or may be integrated within the graphics processor 1508.

[0205] To couple the internal components of the processor 1500, a ring-based interconnect unit 1512 may be used. However, alternative interconnect units may be used, such as, for example, point-to-point interconnects, switched interconnects, or other techniques including technically well-known techniques. In some of these embodiments having a ring-based interconnect 1512, the graphics processor 1508 couples to the ring-based interconnect 1512 via an I / O link 1513.

[0206] Exemplary I / O link 1513 represents at least one of a number of variations of I / O interconnects, including an on-package I / O interconnect that facilitates communication between various processor components and a high-performance embedded memory module 1518, such as, for example, an eDRAM module. Optionally, each of the processor cores 1502A - 1502N and the graphics processor 1508 can use the embedded memory module 1518 as a shared last-level cache.

[0207] The processor cores 1502A - 1502N can be, for example, homogeneous cores that execute the same instruction set architecture. Alternatively, the processor cores 1502A - 1502N are heterogeneous in terms of instruction set architecture, with one or more of the processor cores 1502A - 1502N executing a first instruction set and at least one of the other cores executing a subset of the first instruction set or a different instruction set. The processor cores 1502A - 1502N may be heterogeneous from a microarchitecture perspective, with one or more cores having a relatively high power consumption coupled with one or more power cores having a lower power consumption. As another example, the processor cores 1502A - 1502N are heterogeneous from a computing power perspective. Also, the processor 1500 can be implemented on one or more chips or, alternatively, as a SoC integrated circuit having the illustrated components in addition to other components.

[0208] FIG. 15B is a block diagram of the hardware logic of a graphics processor core 1519 according to some embodiments described herein. The graphics processor core 1519, which may also be referred to as a core slice, can be one or more graphics cores within a modular graphics processor. The graphics processor core 1519 illustrates one graphics core slice, and the graphics processors described herein can include multiple graphics core slices based on target power and performance limits. Each graphics processor core 1519 can include a fixed function block 1530 coupled to a plurality of sub-cores 1521A - 1521F, also referred to as sub-slices, which includes modular blocks of general purpose and fixed function logic.

[0209] The fixed function block 1530 can include a geometry / fixed function pipeline 1531 that can be shared by all sub-cores within the graphics processor core 1519, for example, in a low performance and / or low power graphics processor implementation. The geometry / fixed function pipeline 1531 can include a 3D fixed function pipeline (e.g., a 3D pipeline 1612 as in FIG. 16A described later), a video front end unit, a thread spooler and thread dispatcher, and a unified return buffer manager that manages a unified return buffer (e.g., the unified return buffer 1718 in FIG. 17 described later).

[0210] The fixed function block 1530 may also include a graphics SoC interface 1532, a graphics microcontroller 1533, and a media pipeline 1534. The graphics SoC interface 1532 provides an interface between the graphics processor core 1519 and other processor cores within the system-on-chip integrated circuit. The graphics microcontroller 1533 is a programmable sub-processor configurable to manage various functions of the graphics processor core 1519, including thread dispatch, scheduling, and preemption. The media pipeline 1534 (e.g., the media pipeline 1616 of FIGS. 16A and 17) includes logic to assist in decoding, encoding, pre-processing, and / or post-processing multimedia data, including image and video data. The media pipeline 1534 performs media operations via requests to the computational logic or sampling logic within sub-cores 1521-1521F.

[0211] The SoC interface 1532 enables the graphics processor core 1519 to communicate with a general-purpose application processor core (e.g., a CPU) and / or other components within the SoC, including memory hierarchy elements such as a shared last-level cache memory, system RAM, and / or embedded on-chip or on-package DRAM. The SoC interface 1532 can also enable communication with fixed-function devices within the SoC, such as a camera imaging pipeline, allow and / or implement the use of global memory atomics that can be shared between the graphics processor core 1519 and the CPU within the SoC. The SoC interface 1532 can also implement power management control for the graphics processor core 1519 and enable an interface between the clock domain of the graphics core 1519 and other clock domains within the SoC. Optionally, the SoC interface 1532 enables receipt of command buffers from a command streamer and a global thread dispatcher configured to provide commands and instructions to each of one or more graphics cores within the graphics processor. The commands and instructions can be sent to the media pipeline 1534 when media operations are executed and can be sent to a geometry & fixed-function pipeline (e.g., geometry & fixed-function pipeline 1531, geometry & fixed-function pipeline 1537) when graphics processing operations are executed.

[0212] The graphics microcontroller 1533 can be configured to perform various scheduling and management tasks regarding the graphics processor core 1519. In one configuration, the graphics microcontroller 1533 can perform, for example, scheduling of graphics workloads and / or compute workloads for various graphics parallel engines within the execution unit (EU) arrays 1522A - 1522F, 1524A - 1524F in sub - cores 1521A - 1521F. In this workload scheduling, host software running on the CPU cores of the SoC, including the graphics processor core 1519, can present a workload to one of a plurality of graphics processor doorbells, which then initiates scheduling operations on an appropriate graphics engine. The scheduling operations include determining which workload to execute next, presenting the workload to the command streamer, preempting existing workloads running on the engine, monitoring the progress of the workload, and notifying the host software when the workload is complete. Optionally, the graphics microcontroller 1533 can also assist in the low - power or idle state of the graphics processor core 1519 and can provide the graphics processor core 1519 with the ability to save and restore registers within the graphics processor core 1519 across low - power state transitions, independent of the operating system and / or graphics driver software on the system.

[0213] The graphics processor core 1519 may have a number greater than or less than the illustrated sub-cores 1521A - 1521F up to N modular sub-cores. For each set of N sub-cores, the graphics processor core 1519 may also include shared function logic 1535, shared and / or cache memory 1536, geometry / fixed function pipeline 1537, and additional fixed function logic 1538 that accelerates various graphics and compute processing operations. The shared function logic 1535 may include logic units (e.g., samplers, math, and / or inter-thread communication logic) related to the shared function logic 1720 of FIG. 17 that can be shared by each of the N sub-cores within the graphics processor core 1519. The shared and / or cache memory 1536 can serve as a last-level cache for the set of N sub-cores 1521A - 1521F within the graphics processor core 1519 and can also function as shared memory accessible by multiple sub-cores. The geometry / fixed function pipeline 1537 can be included instead of the geometry / fixed function pipeline 1531 within the fixed function block 1530 and can include the same or similar logic units.

[0214] The graphics processor core 1519 may include additional fixed function logic 1538 that can include various fixed function acceleration logic used by the graphics processor core 1519. Optionally, the additional fixed function logic 1538 may include an additional geometry pipeline used for position-only shading. In position-only shading, there are two geometry pipelines, namely, the full geometry pipeline within the geometry / fixed function pipelines 1538, 1531, and a cull pipeline that can be included within the additional fixed function logic 1538. For example, the cull pipeline can be a trimmed version of the full geometry pipeline. The full pipeline and the cull pipeline can execute different instances of the same application where each instance has a different context. Position-only shading can hide the long runs of discarded triangles and, in some cases, allow shading to complete even faster. For example, the cull pipeline logic within the additional fixed function logic 1538 can execute the position shader in parallel with the main application and generally produce critical results faster than the full pipeline. This is because the cull pipeline fetches and shades only the vertex position attributes without performing rasterization and pixel rendering to the frame buffer. The cull pipeline can use the generated critical results to compute visibility information for all triangles, regardless of whether those triangles are culled. The full pipeline (which can be referred to as the repipe pipeline in this example) uses that visibility information to skip culled triangles and shade only the visible triangles, which are ultimately passed to the rasterization phase.

[0215] Optionally, the additional fixed function logic 1538 may also include machine learning acceleration logic, such as fixed function matrix multiplication logic, for implementation including optimizations for machine learning training or inference.

[0216] Within each of the graphics sub-cores 1521A - 1521F, there is a set of execution resources that can be used to execute graphics operations, media operations, and compute operations in response to requests by a graphics pipeline, a media pipeline, or a shader program. The graphics sub-cores 1521A - 1521F include multiple EU arrays 1522A - 1522F, 1524A - 1524F, thread dispatch and inter-thread communication (TD / IC) logic 1523A - 1523F, 3D (e.g., texture) samplers 1525A - 1525F, media samplers 1526A - 1526F, shader processors 1527A - 1527F, and shared local memory (SLM) 1528A - 1528F. Each of the EU arrays 1522A - 1522F, 1524A - 1524F includes a plurality of execution units, which are general-purpose graphics processing units capable of performing floating-point and integer / fixed-point logical operations in the service of graphics operations, media operations, or compute operations, and include graphics, media, or shader programs. The TD / IC logic 1523A - 1523F performs local thread dispatch and thread control operations for the execution units within the sub-core and assists in communication between threads executing on the execution units of the sub-core. The 3D samplers 1525A - 1525F can read texture or other 3D graphics-related data into memory. The 3D samplers can read texture data differently based on the set sample state and the texture format associated with a given texture. The media samplers 1506A - 1506F can perform similar read operations based on the type and format associated with media data. For example, each of the graphics sub-cores 1521A - 1521F may alternately include unified 3D and media samplers. Threads executing on the execution units within each of the sub-cores 1521A - 1521F can use the shared local memory 1528A - 1528F within each sub-core to enable threads executing within a thread group to perform execution using a shared pool of on-chip memory.

[0217] FIG. 15C is a block diagram of a general-purpose graphics processing unit (GPGPU) 1570 that can be configured as a graphics processor and / or a compute accelerator, such as a graphics processor 1508, according to the embodiments described herein. The GPGPU 1570 can be interconnected with a host processor (e.g., one or more CPUs 1546) and memories 1571, 1572 via one or more system buses and / or memory buses. Memory 1571 can be a system memory shareable with one or more CPUs 1546, while memory 1572 is a device memory dedicated to the GPGPU 1570. For example, components within the GPGPU 1570 and the device memory 1572 can be mapped to memory addresses accessible by one or more CPUs 1546. Access to memories 1571 and 1572 can be assisted by a memory controller 1568. The memory controller 1568 may include an internal direct memory access (DMA) controller 1569, or alternatively, can include logic to perform operations that would otherwise be performed by a DMA controller.

[0218] The GPGPU 1570 includes a plurality of cache memories including an L2 cache 1553, an L1 cache 1554, an instruction cache 1555, and a shared memory 1556 (at least a portion of which can also be partitioned as cache memory). The GPGPU 1570 also includes a plurality of computing units 1560A - 1560N. Each computing unit 1560A - 1560N includes a set of a vector register 1561, a scalar register 1562, a vector logic unit 1563, and a scalar logic unit 1564. The computing units 1560A - 1560N can also include a local shared memory 1565 and a program counter 1566. The computing units 1560A - 1560N can be coupled with a constant cache 1567 which can be used to store constant data which is data that does not change during the execution of a kernel or shader program executed on the GPGPU 1570. The constant cache 1567 can be a scalar data cache and the cached data can be directly fetched into the scalar register 1562.

[0219] In operation, one or more CPUs 1546 can write commands to registers or memory within the GPGPU 1570 that are mapped to an accessible address space. A command processor 1557 can read commands from the registers or memory and determine how those commands are to be processed within the GPGPU 1570. Then, using a thread dispatcher 1558, threads can be dispatched to the computing units 1560A - 1560N to execute those commands. Each computing unit 1560A - 1560N can execute threads independently of other computing units. Further, each computing unit 1560A - 1560N can be independently configured for conditional computing and can output the results of the computation to memory conditionally. The command processor 1557 can interrupt one or more CPUs 1546 when the presented commands are completed.

[0220] Figures 16A-16C illustrate block diagrams of a further graphics processor and a compute accelerator architecture provided by the embodiments described herein, for example, according to Figures 15A-15C. Elements in Figures 16A-16C having the same or similar names as elements in any of the other figures herein describe the same elements in those other figures and can operate or function in a similar manner, can have the same components, and can be coupled to other entities as described elsewhere herein, but are not so limited.

[0221] Figure 16A is a block diagram of a graphics processor 1600, which may be a discrete graphics processing unit or, alternatively, a graphics processor integrated with multiple processing cores or other semiconductor devices such as, but not limited to, a memory device or a network interface. The graphics processor 1600 can be a variation of the graphics processor 1508 and can be used in place of the graphics processor 1508. Accordingly, any disclosure of features in combination with the graphics processor 1508 herein also discloses corresponding combinations with the graphics processor 1600, but is not so limited. The graphics processor can communicate with commands placed in the processor memory via an I / O interface that is memory-mapped to registers on the graphics processor. The graphics processor 1600 can include a memory interface 1614 for accessing memory. The memory interface 1614 can be an interface to local memory, one or more internal caches, one or more shared external caches, and / or system memory.

[0222] Optionally, the graphics processor 1600 also includes a display controller 1602 that drives display output data to the display device 1618. The display controller 1602 includes hardware for one or more overlay planes for the display and composition of multiple layers of video or user interface elements. The display device 1618 may be an internal display device or an external display device. In one embodiment, the display device 1618 is a head-mounted display device such as, for example, a virtual reality (VR) display device or an augmented reality (AR) display device. The graphics processor 1600 may include a video codec engine 1606 that encodes, decodes, or transcodes media to, from, or between one or more media encoding formats including, but not limited to, Moving Picture Experts Group (MPEG) formats such as MPEG-2, Advanced Video Coding (AVC) formats such as H.264 / MPEG-4 AVC, H.265 / HEVC, Alliance for Open Media (AOMedia) VP8, VP9, and Society of Motion Picture & Television Engineers (SMPTE) 421M / VC-1, Joint Photographic Experts Group (JPEG) formats such as JPEG, and Motion JPEG (MJPEG) formats.

[0223] The graphics processor 1600 may include a block image transfer (BLIT) engine 1603 that performs two-dimensional (2D) rasterizer operations including, for example, bit boundary block transfers. However, alternatively, 2D graphics operations may be performed using one or more components of the graphics processing engine (GPE) 1610. In some embodiments, the GPE 1610 is a computing engine for performing graphics operations including three-dimensional (3D) graphics operations and media operations.

[0224] The GPE 1610 may include a 3D pipeline 1612 for performing 3D operations, such as rendering three-dimensional images and scenes using, for example, processing functions that act on 3D primitive shapes (e.g., rectangles, triangles, etc.). The 3D pipeline 1612 includes programmable and fixed-function elements that perform various tasks within the element and / or generate execution threads for the 3D / media subsystem 1615. Although media operations may be performed using the 3D pipeline 1612, one embodiment of the GPE 1610 also includes a media pipeline 1616 that is specifically used for performing media operations such as video post-processing and image enhancement.

[0225] The media pipeline 1616 may include fixed-function or programmable logic units that perform one or more special media operations, such as video decoding acceleration, video deinterlacing, and video encoding acceleration, instead of or in addition to the video codec engine 1606. The media pipeline 1616 may further include a thread generation unit that generates threads for execution on the 3D / media subsystem 1615. The generated threads execute calculations for media operations on one or more graphics execution units included in the 3D / media subsystem 1615.

[0226] The 3D / media subsystem 1615 may include logic for executing threads generated by the 3D pipeline 1612 and the media pipeline 1616. These pipelines may send thread execution requests to the 3D / media subsystem 1615, which includes thread dispatch logic that arbitrates various requests and dispatches them to available thread execution resources. The execution resources include an array of graphics execution units that process 3D threads and media threads. The 3D / media subsystem 1615 may include one or more internal caches for thread instructions and data. Additionally, the 3D / media subsystem 1615 may also include shared memory that includes registers and addressable memory for sharing data between threads and storing output data.

[0227] FIG. 16B illustrates a graphics processor 1620 that is a variation of the graphics processor 1600 and can be used in place of the graphics processor 1600 (and vice versa). Thus, any disclosure of features in combination with the graphics processor 1600 herein also discloses corresponding combinations with the graphics processor 1620, but is not so limited. The graphics processor 1620, according to the embodiments described herein, has a tiled architecture. The graphics processor 1620 may include a graphics processing engine cluster 1622 having a plurality of instances of the graphics processing engine 1610 of FIG. 16A within the graphics engine tiles 1610A - 1610D. Each of the graphics engine tiles 1610A - 1610D can be interconnected via a set of tile interconnects 1623A - 1623F. Each of the graphics engine tiles 1610A - 1610D can also be connected to a memory module or memory device 1626A - 1626D via memory interconnects 1625A - 1625D. The memory devices 1626A - 1626D can use any graphics memory technology. For example, the memory devices 1626A - 1626D can be graphics double data rate (GDDR) memory. The memory devices 1626A - 1626D can be high bandwidth memory (HBM) modules that are on - die with their respective graphics engine tiles 1610A - 1610D. The memory devices 1626A - 1626D can be stacked memory devices that can be stacked on top of their respective graphics engine tiles 1610A - 1610D. Each of the graphics engine tiles 1610A - 1610D and the associated memories 1626A - 1626D can be present on separate chiplets that are bonded to a base die or base substrate, as will be described in more detail in FIGS. 24B - 24D.

[0228] The graphics processor 1620 may be configured in a non-uniform memory access (NUMA) system where memory devices 1626A - 1626D are coupled with associated graphics engine tiles 1610A - 1610D. A given memory device may be accessed by a graphics engine tile other than the tile to which it is directly connected. However, the access latency to memory devices 1626A - 1626D can be shortest when accessing the local tile. In one embodiment, a cache coherent NUMA (ccNUMA) system is implemented using interconnects 1623A - 1623F to enable communication between cache controllers within graphics engine tiles 1610A - 1610D to maintain a consistent memory image when two or more caches store the same memory location.

[0229] The graphics processing engine cluster 1622 can be connected to an on-chip or on-package fabric interconnect 1624. In one embodiment, the fabric interconnect 1624 includes a network processor, a network-on-chip (NoC), or another switching processor such that the fabric interconnect 1624 can function as a packet switching fabric interconnect that exchanges data packets among components of the graphics processor 1620. The fabric interconnect 1624 can enable communication between the graphics engine tiles 1610A-1610D and components such as, for example, video codec engines 1606 and one or more copy engines 1604. The copy engine 1604 can be used to move data from, to, or between the memory devices 1626A-1626D and memory external to the graphics processor 1620 (e.g., system memory). The fabric interconnect 1624 can also be used to interconnect the graphics engine tiles 1610A-1610D with each other. The graphics processor 1620 can optionally include a display controller 1602 that enables connection to an external display device 1618. The graphics processor may also be configured as a graphics accelerator or a compute accelerator. In the accelerator configuration, the display controller 1602 and the display device 1618 can be omitted.

[0230] The graphics processor 1620 can be connected to the host system via the host interface 1628. The host interface 1628 can enable communication between the graphics processor 1620, the system memory, and / or other system components. The host interface 1628 can be, for example, a PCI Express bus or another type of host system interface. For example, the host interface 1628 can be an NVLink or NVSwitch interface. The host interface 1628 and the fabric interconnect 1624 can cooperate to enable multiple instances of the graphics processor 1620 to operate as a single logical device. The cooperation between the host interface 1628 and the fabric interconnect 1624 can also enable the individual graphics engine tiles 1610A-1610D to be presented to the host system as separate logical graphics devices.

[0231] FIG. 16C illustrates a compute accelerator 1630 in accordance with the embodiments described herein. The compute accelerator 1630 can include architectural similarities with the graphics processor 1620 of FIG. 16B and is optimized for compute acceleration. A compute engine cluster 1632 can include a set of compute engine tiles 1640A-1640D that include execution logic optimized for parallel or vector-based general-purpose compute operations. The compute engine tiles 1640A-1640D may not include fixed-function graphics processing logic, although in some embodiments, one or more of the compute engine tiles 1640A-1640D may include logic for performing media acceleration. The compute engine tiles 1640A-1640D can be connected to memories 1626A-1626D via memory interconnects 1625A-1625D. The memories 1626A-1626D and the memory interconnects 1625A-1625D may be of the same technology as the graphics processor 1620 or different. The graphics compute engine tiles 1640A-1640D can also be interconnected via a set of tile interconnects 1623A-1623F and connected to and / or interconnected by a fabric interconnect 1624. In one embodiment, the compute accelerator 1630 can include a large L3 cache 1636 that can be configured as a device-wide cache. The compute accelerator 1630 can also be connected to a host processor and memory via a host interface 1628 in the same manner as the graphics processor 1620 of FIG. 16B.

[0232] The compute accelerator 1630 may also include an integrated network interface 1642. In one embodiment, the network interface 1642 includes network processor and controller logic that enables the compute engine cluster 1632 to communicate on the physical layer interconnect 1644 without the need for data to traverse the memory of the host system. In one embodiment, one of the compute engine tiles 1640A - 1640D is replaced with network processor logic, and data transmitted or received via the physical layer interconnect 1644 can be transmitted directly to or from the memories 1626A - 1626D. Multiple instances of the compute accelerator 1630 may be coupled via the physical layer interconnect 1644 to a single logical device. Alternatively, the various compute engine tiles 1640A - 1640D may be presented as separate network - accessible compute accelerator devices.

[0233] Graphics Processing Engine FIG. 17 is a block diagram of a graphics processing engine 1710 of a graphics processor according to some embodiments. The graphics processing engine (GPE) 1710 can be a version of the GPE 1610 shown in FIG. 16A and can also represent the graphics engine tiles 1610A - 1610D of FIG. 16B. Elements of FIG. 17 with the same or similar names as elements in any other figure here describe the same elements in that other figure and can operate or function in a similar manner, can have the same components, and can be coupled to other entities as described elsewhere in this, but are not so limited. For example, the 3D pipeline 1612 and the media pipeline 1616 of FIG. 16A are also shown in FIG. 17. The media pipeline 1616 is optional in some embodiments of the GPE 1710 and may not be explicitly included within the GPE 1710. For example, in at least one embodiment, a separate media and / or image processor is coupled to the GPE 1710.

[0234] GPE1710 can be coupled to, or can include, a command streamer 1703 that provides a command stream to the 3D pipeline 1612 and / or the media pipeline 1616. Alternatively, or in addition, the command streamer 1703 may be directly coupled to the unified return buffer 1718. The unified return buffer 1718 can be communicatively coupled to the graphics core array 1714. Optionally, the command streamer 1703 is coupled to a memory that can be system memory, or one or more of an internal cache memory and a shared cache memory. The command streamer 1703 can receive commands from the memory and send those commands to the 3D pipeline 1612 and / or the media pipeline 1616. These commands are central directives (instructions) fetched from a ring buffer that stores commands for the 3D pipeline 1612 and the media pipeline 1616. The ring buffer can further include a batch command buffer that stores batches of multiple commands. Commands for the 3D pipeline 1612 can also include references to data stored in memory, such as, but not limited to, vertex and geometry data for the 3D pipeline 1612, and / or image data and memory objects for the media pipeline 1616. The 3D pipeline 1612 and the media pipeline 1616 process commands and data by executing operations via logic within each pipeline, or by dispatching one or more execution threads to the graphics core array 1714. The graphics core array 1714 can include one or more blocks of graphics cores (e.g., (one or more) graphics cores 1715A, (one or more) graphics cores 1715B), with each block including one or more graphics cores. Each graphics core includes a set of graphics execution resources that include general-purpose and graphics-specific execution logic for performing graphics operations and computational operations, as well as fixed-function texture processing and / or machine learning and artificial intelligence acceleration logic.

[0235] In various embodiments, the 3D pipeline 1612 can include fixed function and programmable logic that processes instructions and dispatches execution threads to the graphics core array 1714 to process one or more shader programs, such as vertex shaders, geometry shaders, pixel shaders, fragment shaders, compute shaders, or other shader programs. The graphics core array 1714 provides a unified block of execution resources used when processing these shader programs. The general-purpose execution logic (e.g., execution units) within the one or more graphics cores 1714A - 1714B of the graphics core array 1714 includes support for various 3D API shader languages and can execute multiple concurrent execution threads associated with multiple shaders.

[0236] The graphics core array 1714 can include execution logic that executes media functions such as video and / or image processing. The execution units can include general-purpose logic programmable to execute parallel general-purpose computing operations in addition to graphics processing operations. The general-purpose logic can execute processing operations in parallel or with the general-purpose logic in the one or more processor cores 1407 of FIG. 14 or cores 1502A - 1502N as in FIG. 15A.

[0237] The output data generated by executing threads on the graphics core array 1714 can be output data to memory within the unified return buffer (URB) 1718. The URB 1718 can store data for multiple threads. The URB 1718 can be used to send data between different threads executed on the graphics core array 1714. The URB 1718 can further be used for synchronization between threads on the graphics core array 1714 and the fixed function logic within the shared function logic 1720.

[0238] Optionally, the graphics core array 1714 may be scalable such that the array includes variable graphics cores, each having a variable number of execution units based on the target power and performance levels of the GPE 1710. The execution resources may be dynamically scalable such that the execution resources can be enabled or disabled as needed.

[0239] The graphics core array 1714 is coupled with shared function logic 1720 that includes a plurality of resources shared among the graphics cores within the graphics core array. The shared functions within the shared function logic 1720 are hardware logic units that provide special supplemental functions to the graphics core array 1714. In various embodiments, the shared function logic 1720 includes, but is not limited to, sampler logic 1721, math logic 1722, and inter-thread communication (ITC) logic 1723. Additionally, one or more caches 1725 may be implemented within the shared function logic 1720.

[0240] The shared functions are implemented at least when the requirements for a given special function are insufficient to be included within the graphics core array 1714. Instead, a single instantiation of that special function is implemented as a stand-alone entity within the shared function logic 1720 and shared among the execution resources within the graphics core array 1714. The exact set of functions that are shared among and included within the graphics core arrays 1714 varies by embodiment. Certain shared functions within the shared function logic 1720 that are widely used by the graphics core arrays 1714 may be included within the shared function logic 1716 within the graphics core arrays 1714. Optionally, the shared function logic 1716 within the graphics core arrays 1714 can include some or all of the logic within the shared function logic 1720. All logic elements within the shared function logic 1720 may be replicated within the shared function logic 1716 of the graphics core arrays 1714. Alternatively, the shared function logic 1716 within the graphics core arrays 1714 is selected and the shared function logic 1720 is excluded.

[0241] Execution unit Figures 18A-18B illustrate a thread execution logic 1800 that includes an array of processing elements used in a graphics processor core according to embodiments described herein. Elements in Figures 18A-18B that have the same or similar names as elements in any other figure herein describe the same elements in that other figure and can operate or function in a similar manner, can have the same components, and can be coupled to other entities as described elsewhere herein, but are not so limited. Figures 18A-18B show an overview of the thread execution logic 1800 that can represent the hardware logic shown in each sub-core 1521A-1521F of Figure 15B. Figure 18A represents the execution units within a general-purpose graphics processor, and Figure 18B represents the execution units that can be used within a compute accelerator.

[0242] As shown in FIG. 18A, the thread execution logic 1800 may include a shader processor 1802, a thread dispatcher 1804, an instruction cache 1806, a scalable execution unit array including a plurality of graphics execution units 1808A - 1808N, a sampler 1810, a shared local memory 1811, a data cache 1812, and a data port 1814. Optionally, the scalable execution unit array may be dynamically scalable by enabling or disabling one or more execution units (e.g., any of graphics execution units 1808A, 1808B, 1808C, 1808D, …, 1808N - 1, and 1808N) based on the computational requirements of the workload. The components included may be interconnected via an interconnect fabric that links to each of the components. The thread execution logic 1800 may include one or more connections to memory, such as system memory or cache memory, via, for example, the instruction cache 1806, the data port 1814, the sampler 1810, and one or more of the graphics execution units 1808A - 1808N. Each execution unit (e.g., 1808A) may be a stand - alone programmable general - purpose computing unit capable of executing multiple simultaneous hardware threads while processing multiple data elements in parallel for each thread. In various embodiments, the array of execution units 1808A - 1808N is scalable to include any number of individual execution units.

[0243] In some embodiments, the graphics execution units 1808A-1808N may be used primarily to execute shader programs. The shader processor 1802 may process various shader programs and dispatch execution threads associated with the shader programs via the thread dispatcher 1804. The thread dispatcher may arbitrate thread start requests from the graphics and media pipeline and instantiate the requested threads on one or more execution units within the graphics execution units 1808A-1808N. For example, the geometry pipeline may dispatch vertex shaders, tessellation shaders, and geometry shaders to the thread execution logic for processing. Optionally, the thread dispatcher 1804 may also process runtime thread generation requests from the executing shader programs.

[0244] In some embodiments, the graphics execution units 1808A-1808N may support an instruction set that includes native support for a number of standard 3D graphics shader instructions so that shader programs from a graphics library (e.g., Direct 3D and OpenGL) are executed with minimal translation. The execution units support vertex and geometry processing (e.g., vertex programs, geometry programs, vertex shaders), pixel processing (e.g., pixel shaders, fragment shaders), and general-purpose processing (e.g., compute and media shaders). Each of the graphics execution units 1808A-1808N can multi-issue single instruction multiple data (SIMD) execution, and the multi-threaded operation enables an efficient execution environment when facing larger latency memory accesses. Each hardware thread within each execution unit has a dedicated high-bandwidth register file and an associated independent thread state. Execution is multi-issued clock-by-clock to a pipeline capable of integer arithmetic, single-precision and double-precision floating-point arithmetic, SIMD branching capabilities, logical operations, transcendental operations, and other miscellaneous operations. While waiting for data from one of the memory or shared functions, the dependency logic within the execution units 1808A-1808N puts the waiting threads to sleep until the requested data is returned. While the waiting threads are sleeping, the hardware resources can be devoted to processing other threads. For example, during the latency associated with vertex shader operations, the execution unit can execute operations related to pixel shaders, fragment shaders, or other types of shader programs, including different vertex shaders such as vertex shader 2107 shown in FIG. 21. Various embodiments can be applied to use execution by single instruction multiple threads (SIMT) instead of, or in addition to, the use of SIMD. References to SIMD cores or operations also apply to SIMT, or to SIMD in combination with SIMT.

[0245] Each execution unit within the graphics execution units 1808A - 1808N operates on an array of data elements. The number of data elements is the "execution size", i.e., the number of channels of the instruction. An execution channel is the logical unit of execution related to data element access, masking, and flow control within an instruction. The number of channels may be independent of the number of physical arithmetic logic units (ALUs), floating - point units (FPUs), or other logic units (e.g., tensor cores, ray tracing cores, etc.) in a particular graphics processor. Also, the graphics execution units 1808A - 1808N may support integer and floating - point data types.

[0246] The execution unit instruction set includes SIMD instructions. Various data elements can be stored in registers as packed data types, and the execution unit processes various elements based on the data size of the elements. For example, when operating on a 256 - bit - wide vector, a 256 - bit vector is stored in a register, and the execution unit operates on the vector as 4 separate, 64 - bit packed data elements (data elements of quad - word (QW) size), 8 separate 32 - bit packed data elements (data elements of double - word (DW) size), 16 separate 16 - bit packed data elements (data elements of word (W) size), or 32 separate 8 - bit data elements (data elements of byte (B) size). However, different vector widths and register sizes are also possible.

[0247] Optionally, one or more execution units can be combined into fused graphics execution units 1809A - 1809N that have thread control logic (1807A - 1807N) common to the fused EUs. Multiple EUs can be fused into an EU group. Each EU within a fused EU group can be configured to execute separate SIMD hardware threads. The number of EUs within a fused EU group can vary according to embodiments. Also, various SIMD widths per EU can be executed, including but not limited to SIMD8, SIMD16, and SIMD32. Each fused graphics execution unit 1809A - 1809N includes at least two execution units. For example, fused execution unit 1809A includes a first EU 1808A, a second EU 1808B, and thread control logic 1807A common to the first EU 1808A and the second EU 1808B. Thread control logic 1807A controls the threads executed on fused graphics execution unit 1809A and enables each EU within fused execution units 1809A - 1809N to execute using a common instruction pointer register.

[0248] One or more internal instruction caches (e.g., 1806) are included in thread execution logic 1800 to cache thread instructions for the execution units. One or more data caches (e.g., 1812) can be included in thread execution logic 1800 to cache thread data during thread execution. Also, threads executed on execution logic 1800 can also store explicitly managed data in shared local memory 1811. Sampler 1810 can be included to provide texture sampling for 3D operations and media sampling for media operations. Sampler 1810 can include special texture or media sampling functions that process texture or media data during the sampling process before providing the sampled data to the execution units.

[0249] In execution, the graphics and media pipeline sends thread start requests to the thread execution logic 1800 via thread generation and dispatch logic. When a group of geometric objects is processed and rasterized into pixel data, pixel processor logic (e.g., pixel shader logic, fragment shader logic, etc.) within the shader processor 1802 is called to further compute output information and write the results to an output surface (e.g., color buffer, depth buffer, stencil buffer, etc.). The pixel shader or fragment shader can compute the values of various vertex attributes that are interpolated across the rasterized objects. Next, the pixel processor logic within the shader processor 1802 can execute a pixel or fragment shader program supplied by an application programming interface (API). To execute the shader program, the shader processor 1802 dispatches threads to execution units (e.g., 1808A) via the thread dispatcher 1804. The shader processor 1802 can access texture data of a texture map stored in memory using texture sampling logic within the sampler 1810. Arithmetic operations on the texture data and input geometry data compute pixel color data for each geometric fragment or discard one or more pixels from further processing.

[0250] Also, the data port 1814 may provide a memory access mechanism for the thread execution logic 1800 to output processed data to memory for further processing on the graphics processor output pipeline. The data port 1814 can include or be coupled to one or more cache memories (e.g., data cache 1812) that cache data for memory access via the data port 1814.

[0251] Optionally, the execution logic 1800 can also include a ray tracer 1805 that can provide a ray tracing acceleration function. The ray tracer 1805 can support a ray tracing instruction set that includes instructions / functions for ray generation. The ray tracing instruction set may be the same as or different from the ray tracing instruction set supported by the ray tracing core 372 of FIG. 3C.

[0252] FIG. 18B shows exemplary internal details of the execution unit 1808. The graphics execution unit 1808 can include an instruction fetch unit 1837, a general-purpose register file array (GRF) 1824, an architecture register file array (ARF) 1826, a thread arbiter 1822, a dispatch unit 1830, a branch unit 1832, a set of SIMD floating-point units (FPUs) 1834, and optionally, a set of dedicated integer SIMD ALUs 1835. The GRF 1824 and ARF 1826 include a set of general-purpose register files and architecture register files associated with each simultaneous hardware thread that can be active in the graphics execution unit 1808. The architecture state for each thread can be maintained in the ARF 1826, and the data used during thread execution is stored in the GRF 1824. The execution state of each thread can be held in thread-specific registers within the ARF 1826, including the instruction pointer for each thread.

[0253] The graphics execution unit 1808 may have an architecture that is a combination of Simultaneous Multi-Threading (SMT) and fine-grained Interleaved Multi-Threading (IMT). This architecture can have a module configuration that is fine-tunable at design time based on the target number of simultaneous threads and the number of registers per execution unit, and the execution unit resources are divided across the logic used to execute multiple simultaneous threads. The number of logical threads that can be executed by the graphics execution unit 1808 is not limited to the number of hardware threads, and multiple logical threads can be assigned to each hardware thread.

[0254] Optionally, the graphics execution unit 1808 can issue multiple instructions, each of which can be a different instruction, simultaneously. The thread arbiter 1822 of the graphics execution unit 1808 can dispatch those instructions to one of the transmission unit 1830, the branch unit 1832, or the (one or more) SIMD FPU 1834 for execution. Each execution thread can access 128 general-purpose registers in the GRF 1824, where each register can store 32 bytes that can be accessed as a SIMD8-element vector of 32-bit data elements. Each execution unit thread can have access to 4K bytes in the GRF 1824, but embodiments are not so limited, and in other embodiments, more or fewer register resources may be provided. The graphics execution unit 1808 can be divided into 7 hardware threads that can execute computing operations independently, but the number of threads per execution unit can also vary according to embodiments. For example, up to 16 hardware threads can be supported. In an exemplary embodiment where 7 threads can access 4K bytes, the GRF 1824 can store a total of 28K bytes. In another exemplary embodiment where 16 threads can access 4K bytes, the GRF 1824 can store a total of 64K bytes. However, the number of threads per execution unit is not limited to these examples and can be more or less than the given number. The flexible addressing mode can enable addressing multiple registers together to effectively construct wider registers or to represent a strided rectangular block data structure.

[0255] In addition, or alternatively, memory operations, sampler operations, and other longer-latency system communications may be dispatched via "send" instructions executed by the message passing transmission unit 1830. Branch instructions can be dispatched to a dedicated branch unit 1832 that supports SIMD branching and final convergence.

[0256] The graphics execution unit 1808 may include one or more SIMD floating point units (FPUs) 1834 that execute floating point operations. The (one or more) FPUs 1834 may also support integer computations. In some examples, the (one or more) FPUs 1834 can SIMD execute up to M 32-bit floating point (or integer) operations, or up to 2M 16-bit integer or 16-bit floating point operations. Optionally, at least one of the (one or more) FPUs provides an extended math function that supports high throughput transcendental functions and double precision 184-bit floating point. A set of 8-bit integer SIMD ALUs 1835 may also be present and may be optimized to perform operations particularly related to machine learning computations.

[0257] Optionally, an array of multiple instances of the graphics execution unit 1808 may be instantiated in a graphics subcore group (e.g., a sub-slice). For scalability, the product designer can select the exact number of execution units per subcore group. The execution unit 1808 may execute instructions across multiple execution channels. Also, each thread executed on the graphics execution unit 1808 may execute on a different channel.

[0258] FIG. 19 shows an exemplary further execution unit 1900. Elements of FIG. 19 having the same or similar names as elements in any other of these figures describe the same elements in that other figure and can operate or function in a similar manner, can have the same components, and can be coupled to other entities as described elsewhere herein, but are not so limited. Execution unit 1900 can be, for example, a compute-optimized execution unit used in compute engine tiles 1640A-1640D as in FIG. 16C, but is not so limited. Execution unit 1900 may also be used in graphics engine tiles 1610A-1610D as in FIG. 16B. Execution unit 1900 can include a thread control unit 1901, a thread state unit 1902, an instruction fetch / prefetch unit 1903, and an instruction decode unit 1904. Execution unit 1900 can further include a register file 1906 that stores registers that can be allocated to hardware threads within the execution unit. Execution unit 1900 can further include a send unit 1907 and a branch unit 1908. Send unit 1907 and branch unit 1908 can operate in a similar manner to send unit 1830 and branch unit 1832 of graphics execution unit 1808 of FIG. 18B.

[0259] Execution unit 1900 can also include a compute unit 1910 that includes a plurality of different types of functional units. Compute unit 1910 can also include an ALU 1911, a systolic array 1912, and a math unit 1913. ALU 1911 includes an array of arithmetic logic units. ALU 1911 can be configured to perform 64-bit, 32-bit, and 16-bit integer and floating point operations across multiple processing lanes and data channels and also for multiple hardware threads and / or software threads. ALU 1911 can perform integer operations and floating point operations simultaneously (e.g., within the same clock cycle).

[0260] The systolic array 1912 includes a network of data processing units of width W and depth D that can be used to perform vector operations or other data parallel operations systolicly. The systolic array 1912 can be configured to perform various matrix operations, including inner product, outer product, and general matrix-matrix multiplication (GEMM) operations. The systolic array 1912 can support 16-bit floating point operations, as well as 8-bit, 4-bit, 2-bit, and binary integer operations. The systolic array 1912 can be configured to accelerate machine learning operations. The systolic array 1912 may be configured to have support for bfloat16, (brain floating point) 16-bit floating point format, or tensor floating 32-bit floating point format (TF32) with a different number of mantissa and exponent bits than the IEEE 754 format. The FP64 format may also be supported.

[0261] In one embodiment, systolic array 1912 includes hardware that accelerates sparse matrix operations. Multiplication operations on sparse regions of input data can be bypassed without sacrificing throughput. Block sparsity within the input matrix can be detected and operations with known output values can be bypassed. In one embodiment, systolic array 1912 includes hardware that enables operations on sparse data with a compressed representation. The compressed representation of a sparse matrix stores non-zero values and metadata that defines the positions of those non-zero values within the matrix. Exemplary compressed representations include, but are not limited to, compressed tensor representations such as compressed sparse row (CSR), compressed sparse column (CSC), compressed sparse fiber (CSF) representations, etc. Support for compressed representations enables operations to be performed without the need to decompress or decode the compressed representation for inputs in the compressed tensor format. In such embodiments, operations need only be performed on non-zero input values, and the resulting non-zero output values can be mapped to the output matrix. In some embodiments, hardware support is also provided for machine-specific reversible data compression formats used when transmitting data within the hardware or across the system bus. Such data can be held in compressed format for sparse input data, and systolic array 1912 can use the compressed metadata of the compressed data to enable operations to be performed only on non-zero values or to bypass blocks of zero data inputs for multiplication operations.

[0262] Math unit 1913 can be configured to perform a specific subset of mathematical operations in an efficient and lower power way than ALU unit 1911. Math unit 1913 can include math logic found in shared functionality logic of a graphics processing engine provided by other described embodiments, such as math logic 1722 of shared functionality logic 1720 of FIG. 17. Math unit 1913 can be configured to perform 32-bit and 64-bit floating point operations.

[0263] The thread control unit 1901 includes logic for controlling the execution of threads within the execution unit. The thread control unit 1901 can include thread arbitration logic for starting, stopping, and pre-empting the execution of threads within the execution unit 1900. The thread state unit 1902 can be used to store thread states regarding the threads assigned to execute on the execution unit 1900. Storing thread states within the execution unit 1900 enables rapid pre-emption of threads when those threads are blocked or become idle. The instruction fetch / prefetch unit 1903 can fetch instructions from the instruction cache of a higher level of execution logic (e.g., instruction cache 1806 as in FIG. 18A). The instruction fetch / prefetch unit 1903 can also issue prefetch requests for instructions to be loaded into the instruction cache based on the analysis of the currently executing threads. The instruction decoding unit 1904 can be used to decode the instructions to be executed by the computing unit. The instruction decoding unit 1904 may be used as an auxiliary decoder to decode complex instructions into component micro-operations.

[0264] The execution unit 1900 further includes a register file 1906 that can be used by the hardware threads executing on the execution unit 1900. The registers within the register file 1906 can be divided among the logic used to execute multiple simultaneous threads within the computing unit 1910 of the execution unit 1900. The number of logical threads that can be executed by the graphics execution unit 1900 is not limited to the number of hardware threads, and multiple logical threads can be assigned to each hardware thread. The size of the register file 1906 can vary across embodiments based on the number of supported hardware threads. Register renaming may be used to dynamically assign registers to hardware threads.

[0265] Figure 20 is a block diagram illustrating the graphics processor instruction format 2000. The graphics processor execution unit supports an instruction set with instructions of multiple formats. The solid boxes indicate components generally included in the execution unit instructions, and the dashed lines include components that are optional or included only in a subset of the instructions. In some embodiments, the graphics processor instruction format 2000 illustrated and described herein is a macro instruction in that they are instructions supplied to the execution unit rather than micro-operations resulting from instruction decoding of processed instructions. Thus, a single instruction can cause the hardware to execute multiple micro-operations.

[0266] The graphics processor execution unit described herein may natively support instructions in the 128-bit instruction format 2010. The 64-bit compressed instruction format 2030 is available for some instructions based on the selected instructions, instruction options, and number of operands. The native 128-bit instruction format 2010 provides access to all instruction options, while in the 64-bit format 2030, some options and operations are restricted. The native instructions available in the 64-bit format 2030 vary by embodiment. The instructions are partially compressed using a set of index values in the index field 2013. The execution unit hardware references a set of compression tables based on the index values and uses the compression table output to reconstruct the native instructions in the 128-bit instruction format 2010. Other sizes and formats of instructions may be used.

[0267] For each format, the instruction opcode 2012 specifies the operation that the execution unit performs. The execution unit executes each instruction in parallel across multiple data elements of each operand. For example, in response to an add instruction, the execution unit performs a simultaneous addition operation across each color channel representing a texture element or a picture element. By default, the execution unit executes each instruction across all data channels of the operand. The instruction control field 2014 may enable control of specific execution options such as, for example, channel selection (e.g., prediction) and data channel order (e.g., swizzle). In the instructions of the 128-bit instruction format 2010, the execution size field 2016 limits the number of data channels that are executed in parallel. The execution size field 2016 may not be available in the 64-bit compact instruction format 2030.

[0268] Some execution unit instructions have up to three operands including two source operands src0 2020, src1 2022 and one destination 2018. The execution unit may support dual destination instructions, in which case one of the destinations is implied. Data manipulation instructions can have a third source operand (e.g., SRC2 2024), and the instruction opcode 2012 determines the number of source operands. The last source operand of the instruction can be an immediate value (e.g., a hard-coded value) passed with the instruction.

[0269] The 128-bit instruction format 2010 may include an access / address mode field 2026 that specifies, for example, whether direct register addressing mode or indirect register addressing mode is used. When direct register addressing mode is used, the register addresses of one or more operands are provided directly by bits within the instruction.

[0270] The 128-bit instruction format 2010 may also include an access / address mode field 2026 that specifies the address mode and / or access mode for the instruction. The access mode may be used to define the data access alignment for the instruction. Access modes may be supported that include an access mode with 16-byte alignment and an access mode with 1-byte alignment, and the byte alignment of the access mode determines the access alignment of the instruction operand. For example, when in a first mode, the instruction may use byte-aligned addressing for source and destination operands, and when in a second mode, the instruction may use 16-byte-aligned addressing for all source and destination operands.

[0271] The address mode portion of the access / address mode field 2026 may determine whether the instruction should use direct addressing or indirect addressing. When the direct register addressing mode is used, bits within the instruction directly provide the register addresses of one or more operands. When the indirect register addressing mode is used, the register addresses of one or more operands may be calculated based on the address immediate field and the address register value within the instruction.

[0272] Instructions can be grouped based on the opcode 2012 bit field to simplify opcode decoding 2040. In an 8-bit opcode, bits 4, 5, and 6 enable the execution unit to determine the type of opcode. The illustrated detailed opcode grouping is merely an example. The move and logic opcode group 2042 can include data movement and logic instructions (e.g., move (mov), compare (cmp)). The move and logic group 2042 can share the lower 5 bits, with the move (mov) instruction being in the form 0000xxxxb and the logic instruction being in the form 0001xxxxb. The flow control instruction group 2044 (e.g., call, jmp) includes instructions in the form 0010xxxxb (e.g., 0x20). The mixed instruction group 2046 includes a mix of instructions including synchronization instructions (e.g., wait, send) in the form 0011xxxxb (e.g., 0x30). The parallel math processing instruction group 2048 includes per-component arithmetic instructions (e.g., add, mul) in the form 0100xxxxb (e.g., 0x40). The parallel math processing instruction group 2048 executes arithmetic operations in parallel across multiple data channels. The vector math processing group 2050 includes arithmetic instructions (e.g., dp4) in the form 0101xxxxb (e.g., 0x50). The vector math processing group executes arithmetic operations such as dot product calculations on vector operands. The illustrated opcode decoding 2040 can be used, in one embodiment, to determine which part of the execution unit to use to execute the decoded instructions. For example, some instructions can be designated as systolic instructions to be executed by the systolic array. Other instructions, such as ray tracing instructions (not shown), can be routed to a ray tracing core or ray tracing logic within a slice or partition of the execution logic.

[0273] Graphics pipeline FIG. 21 is a block diagram of a graphics processor 2100 according to another embodiment. Elements of FIG. 21 with the same or similar names as elements in any of the other figures described in this document describe the same elements as in those other figures, can operate or function in a similar manner, can have the same components, and can be coupled to other entities as described elsewhere in this document, but are not so limited.

[0274] The graphics processor 2100 may include different types of graphics processing pipelines, such as, for example, a geometry pipeline 2120, a media pipeline 2130, a display engine 2140, thread execution logic 2150, and a rendering output pipeline 2170. The graphics processor 2100 may be a graphics processor within a multi-core processing system that includes one or more general-purpose processing cores. The graphics processor can be controlled by register writes to one or more control registers (not shown), or can be controlled via commands issued to the graphics processor 2100 via a ring interconnect 2102. The ring interconnect 2102 can couple the graphics processor 2100 to other processing components, such as, for example, other graphics processors or general-purpose processors. Commands from the ring interconnect 2102 are interpreted by a command streamer 2103 that supplies instructions to individual components of the geometry pipeline 2120 or the media pipeline 2130.

[0275] The command streamer 2103 can instruct the operation of the vertex fetcher 2105 that reads vertex data from memory and executes vertex processing commands provided by the command streamer 2103. The vertex fetcher 2105 provides vertex data to the vertex shader 2107, and the vertex shader 2107 executes coordinate space transformation and lighting operations for each vertex. The vertex fetcher 2105 and the vertex shader 2107 can execute vertex processing instructions by dispatching execution threads to the execution units 2152A - 2152B via the thread dispatcher 2131.

[0276] The execution units 2152A - 2152B can be an array of vector processors having an instruction set for executing graphics and media operations. The execution units 2152A - 2152B can have attached L1 caches 2151 that are either specific to each array or shared among the arrays. This cache can be configured as a data cache, an instruction cache, or a single cache divided to accommodate data and instructions in different partitions.

[0277] The geometry pipeline 2120 can include a tessellation component that performs hardware - accelerated tessellation of 3D objects. The programmable hull shader 2111 can set the tessellation operation. The programmable domain shader 2117 can provide back - end evaluation of the tessellation output. The tessellator 2113 can operate under the instruction of the hull shader 2111 and can include special - purpose logic that generates a set of detailed geometric objects based on a coarse geometric model provided as input to the geometry pipeline 2120. Also, when tessellation is not used, the tessellation components (e.g., the hull shader 2111, the tessellator 2113, and the domain shader 2117) can be bypassed. The tessellation components can operate based on the data received from the vertex shader 2107.

[0278] The completed geometric object can be processed by the geometry shader 2119 via one or more threads dispatched to the execution units 2152A - 2152B, or can proceed directly to the clipper 2129. The geometry shader operates on the entire geometric object rather than on vertices or vertex patches as in the previous stages of the graphics pipeline. When tessellation is disabled, the geometry shader 2119 receives input from the vertex shader 2107. The geometry shader 2119 may be programmable by a geometry shader program that performs geometry tessellation when the tessellation unit is disabled.

[0279] Before rasterization, the clipper 2129 processes the vertex data. The clipper 2129 may be a fixed - function clipper or a programmable clipper having clipping and geometry shader functionality. To convert the geometric object into a per - pixel representation, the rasterizer and depth test component 2173 within the rendering output pipeline 2170 may dispatch a pixel shader. The pixel shader logic may be included in the thread execution logic 2150. Optionally, the application can bypass the rasterizer and depth test component 2173 and access the non - rasterized vertex data via the stream - out unit 2123.

[0280] The graphics processor 2100 has an interconnect bus, an interconnect fabric, or some other interconnect mechanism that enables passing data and messages among the major components of the processor. In some embodiments, the execution units 2152A - 2152B and related logic units (e.g., L1 cache 2151, sampler 2154, texture cache 2158, etc.) are interconnected via data ports 2156 to execute memory accesses and communicate with the rendering output pipeline components of the processor. The sampler 2154, caches 2151, 2158, and execution units 2152A - 2152B may each have separate memory access paths. Optionally, the texture cache 2158 may also be configured as a sampler cache.

[0281] The rendering output pipeline 2170 may include a rasterizer / depth test component 2173 that converts vertex - based objects to their associated pixel - based representations. The rasterizer logic may include a window / mask unit that performs fixed - function triangle and line rasterization. In some embodiments, an accompanying rendering cache 2178 and depth cache 2179 are also available. The pixel operation component 2177 performs pixel - based operations on the data, although in some examples, pixel operations related to 2D operations (e.g., bit - block image transfer with blending) are performed by the 2D engine 2141 or, at display time, replaced by the display controller 2143 using an overlay display plane. The shared L3 cache 2175 is available to all graphics components, enabling sharing of data without using the main system memory.

[0282] The media pipeline 2130 may include a media engine 2137 and a video front end 2134. The video front end 2134 may receive pipeline commands from the command streamer 2103. The media pipeline 2130 may include a separate command streamer. The video front end 2134 may process media commands before sending the commands to the media engine 2137. The media engine 2137 may include a thread generation function that generates threads dispatched to the thread execution logic 2150 via the thread dispatcher 2131.

[0283] The graphics processor 2100 may include a display engine 2140. This display engine 2140 is external to the processor 2100 and may be coupled to the graphics processor via the ring interconnect 2102 or some other interconnect bus or fabric. The display engine 2140 may include a 2D engine 2141 and a display controller 2143. The display engine 2140 may include special-purpose logic operable independently of the 3D pipeline. The display controller 2143 can be coupled to a display device (not shown), which may be an integrated system display device such as in a laptop computer or an external display device attached via a display device connector.

[0284] The geometry pipeline 2120 and the media pipeline 2130 can be configured to operate based on multiple graphics and media programming interfaces and are not specific to any one application programming interface (API). Driver software for the graphics processor can convert API calls specific to a particular graphics or media library into commands that can be processed by the graphics processor. Support can be provided for the Open Graphics Library (OpenGL), Open Computing Language (OpenCL), and / or Vulkan graphics and compute APIs (all from the Khronos Group). Support can also be provided for the Direct3D library from Microsoft. Combinations of these libraries may be supported. Support can also be provided for the Open Source Computer Vision Library (OpenCV). Future APIs with compatible 3D pipelines will also be supported if mapping from the pipeline of these future APIs to the pipeline of this graphics processor can be done.

[0285] Graphics Pipeline Programming FIG. 22A is a block diagram illustrating a graphics processor command format 2200 used to program a graphics processing pipeline, such as the pipeline described herein in connection with FIGS. 16A, 17, and 21, for example. FIG. 22B is a block diagram illustrating a graphics processor command sequence 2210 according to one embodiment. The solid box in FIG. 22A indicates components generally included in a graphics command, and the dashed lines include components that are optional or included only in a subset of the graphics commands. The exemplary graphics processor command format 2200 of FIG. 22A includes a data field for identifying a client 2202, a command operation code (opcode) 2204, and data 2206 related to the command. Some commands also include a sub-opcode 2205 and a command size 2208.

[0286] Client 2202 may specify a client unit of a graphics device that processes command data. A graphics processor command parser may examine the client field of each command to adjust further processing of the command and route the command data to an appropriate client unit. The graphics processor client units may include a memory interface unit, a rendering unit, a 2D unit, a 3D unit, and a media unit. Each client unit may have a corresponding processing pipeline for processing commands. When a command is received by a client unit, the client unit reads the opcode 2204 and, if present, the sub-opcode 2205 and determines the operation to be performed. The client unit executes the command using the information in the data field 2206. For some commands, an explicit command size 2208 is expected to specify the size of the command. The command parser may automatically determine at least a portion of the size of the command based on the command opcode. Commands may be aligned in multiples of double words. Other command formats may also be used.

[0287] The flowchart of FIG. 22B shows an exemplary graphics processor command sequence 2210. The software or firmware of a data processing system featuring an exemplary graphics processor may use one version of the shown command sequence to set up, execute, and terminate a set of graphics operations. A certain sample command sequence is illustrated and described for illustrative purposes only and is not limited to these specific commands or this command sequence. Also, these commands may be issued as a batch of commands within the command sequence such that the graphics processor processes the sequence of commands at least partially simultaneously.

[0288] The graphics processor command sequence 2210 begins with a pipeline flush command 2212 that can complete the commands currently pending in the active graphics pipeline. Optionally, the 3D pipeline 2222 and the media pipeline 2224 may be assumed not to operate simultaneously. The pipeline flush is executed to complete the commands pending in the active graphics pipeline. In response to the pipeline flush, the command parser of the graphics processor pauses command processing until the active rendering engine completes the operations in progress and the associated read cache is invalidated. Optionally, the data in the rendering cache marked as 'dirty' can be flushed to memory. The pipeline flush command 2212 can be used for pipeline synchronization or before putting the graphics processor into a low-power state.

[0289] When a command sequence requires that a graphics processor perform an explicit switch between pipelines, the pipeline selection command 2213 may be used. The pipeline selection command 2213 may be required only once within an execution context before issuing pipeline commands, provided that the context issues commands for both pipelines. A pipeline flush command 2212 may be required immediately before a pipeline switch by the pipeline selection command 2213.

[0290] The pipeline control command 2214 may configure the graphics pipeline for operation and may be used to program the 3D pipeline 2222 and the media pipeline 2224. The pipeline control command 2214 may set the pipeline state for the active pipeline. The pipeline control command 2214 may be used for pipeline synchronization and to clear data from one or more cache memories within the active pipeline before processing a batch of commands.

[0291] Commands related to the return buffer state 2216 may be used to configure a set of return buffers for each pipeline to write data. Some pipeline operations require the allocation, selection, or setting of one or more return buffers into which the operation writes intermediate data during processing. The graphics processor may also use one or more return buffers to store output data and to perform cross-thread communication. The return buffer state 2216 may include selecting the size and number of return buffers to use for a set of pipeline operations.

[0292] The remaining commands within the command sequence vary based on the active pipeline of the operation. Based on the pipeline determination 2220, the command sequence is adjusted to match either the 3D pipeline 2222 that begins with the 3D pipeline state 2230, or the media pipeline 2224 that begins with the media pipeline state 2240.

[0293] The commands that make up the 3D pipeline state 2230 include 3D state setting commands regarding the vertex buffer state, vertex element state, constant color state, depth buffer state, and other state variables that should be set before the 3D primitive commands are processed. The values of these commands are determined at least in part based on the specific 3D API being used. The 3D pipeline state 2230 commands may also be able to selectively disable or bypass those specific pipeline elements if they are not used.

[0294] 3D primitive 2232 commands can be used to present 3D primitives processed by the 3D pipeline. The commands and associated parameters passed to the graphics processor via the 3D primitive 2232 commands are transferred to the vertex fetch function within the graphics pipeline. The vertex fetch function uses the 3D primitive 2232 command data to generate a vertex data structure. The vertex data structure is stored in one or more return buffers. With the 3D primitive 2232 commands, vertex operations on the 3D primitives can be performed by the vertex shader. To process the vertex shader, the 3D pipeline 2222 dispatches shader execution threads to the graphics processor execution units.

[0295] The 3D pipeline 2222 can be triggered by an execution 2234 command or event. A register can write to trigger command execution. The execution may be triggered by a 'go' or 'kick' command within a command sequence. Command execution may be triggered using a pipeline synchronization command that flushes the command sequence from the entire graphics pipeline. The 3D pipeline performs geometry processing on 3D primitives. When the operation is complete, the resulting geometric object is rasterized, and the resulting pixels are shaded by the pixel engine. For these operations, additional commands may also be included to control pixel shading and pixel backend operations.

[0296] When performing media operations, the graphics processor command sequence 2210 can traverse the path of the media pipeline 2224. Generally, the specific use and method of programming for the media pipeline 2224 depend on the media or the computational operations being performed. Specific media decoding operations can be offloaded to the media decoding pipeline during media decoding. The media pipeline can also be bypassed, and media decoding can be performed using resources provided by one or more general-purpose processing cores, either in whole or in part. The media pipeline may also include elements for general-purpose graphics processor unit (GPGPU) operations, using the graphics processor to perform SIMD vector operations using a compute shader program that is not explicitly related to the rendering of graphics primitives.

[0297] The media pipeline 2224 can be configured in the same manner as the 3D pipeline 2222. A set of commands for configuring the media pipeline state 2240 is dispatched or placed in the command queue before the media object commands 2242. The commands for the media pipeline state 2240 may include data for configuring the media pipeline elements that are to be used to process the media object. This may include data for configuring video decoding and video encoding logic within the media pipeline, such as an encoding format or a decoding format. The commands for the media pipeline state 2240 may also support the use of one or more pointers to "indirect" state elements that include a batch of state settings.

[0298] The media object commands 2242 may supply a pointer to the media object for processing by the media pipeline. The media object includes a memory buffer that contains the video data to be processed. Optionally, all media pipeline states must be valid before issuing the media object commands 2242. Once the pipeline state is configured and the media object commands 2242 are queued, the media pipeline 2224 is triggered by an execute command 2244 or an equivalent execution event (e.g., a register write). Then, the output from the media pipeline 2224 can be post-processed by operations provided by the 3D pipeline 2222 or the media pipeline 2224. GPGPU operations can be configured and executed in the same manner as media operations.

[0299] Graphics Software Architecture Figure 23 shows an exemplary graphics software architecture for a data processing system 2300. Such a software architecture may include a 3D graphics application 2310, an operating system 2320, and at least one processor 2330. The processor 2330 may include a graphics processor 2332 and one or more general-purpose processor cores 2334. The processor 2330 may be a variation of the processor 1402 or any other processor described herein. The processor 2330 may be used in place of the processor 1402 or any other processor described herein. Thus, the disclosure of any feature in combination with the processor 1402 or any other processor described herein also discloses the corresponding combination with the graphics processor 2330, but is not so limited. Also, elements of Figure 23 that have the same or similar names as elements in any other figure herein describe the same elements in that other figure and can operate or function in a similar manner, can have the same components, and can be coupled to other entities as described elsewhere herein, but are not so limited. The graphics application 2310 and the operating system 2320 are each executed in the system memory 2350 of the data processing system.

[0300] The 3D graphics application 2310 may include one or more shader programs that include shader instructions 2312. The shader language instructions may be those of a high-level shader language such as, for example, Direct3D's High-level Shader Language (HLSL), OpenGL Shader Language (GLSL). This application may also include machine language executable instructions 2314 suitable for execution by the general-purpose processor core 2334. This application may also include a graphics object 2316 defined by vertex data.

[0301] The operating system 2320 can be a Microsoft® Windows® operating system from Microsoft, a proprietary UNIX®-like operating system, or an open-source UNIX®-like operating system using a variant of the Linux® kernel. The operating system 2320 can support a graphics API 2322 such as, for example, the Direct3D API, the OpenGL API, or the Vulkan API. When the Direct3D API is used, the operating system 2320 uses a front-end shader compiler 2324 to compile shader instructions 2312 in HLSL into a lower-level shader language. The compilation can be just-in-time (JIT) compilation or the application can perform shader pre-compilation. The high-level shader can be compiled into a low-level shader during compilation of the 3D graphics application 2310. The shader instructions 2312 can be provided in an intermediate form such as a version of the Standard Portable Intermediate Representation (SPIR) used, for example, by the Vulkan API.

[0302] The user-mode graphics driver 2326 can include a back-end shader compiler 2327 that converts the shader instructions 2312 into a hardware-specific representation. When the OpenGL API is used, the shader instructions 2312 in the GLSL high-level language are passed to the user-mode graphics driver 2326 for compilation. The user-mode graphics driver 2326 can communicate with the kernel-mode graphics driver 2329 using the operating system kernel-mode functionality 2328. The kernel-mode graphics driver 2329 can communicate with the graphics processor 2332 to dispatch commands and instructions.

[0303] IP Core Implementation One or more aspects may be implemented by an expression code stored in a machine-readable medium that represents and / or defines logic within an integrated circuit, such as a processor. For example, the machine-readable medium may contain instructions that represent various logic within the processor. When read by a machine, those instructions may cause the machine to fabricate logic to perform the techniques described herein. Such an expression, known as an “IP core,” is a reusable logic unit for an integrated circuit that may be stored in a tangible machine-readable medium as a hardware model that describes the structure of the integrated circuit. The hardware model may be supplied to various customers or manufacturing facilities, and the customer or manufacturing facility loads the hardware model into a manufacturing machine that fabricates the integrated circuit. The integrated circuit may be fabricated so that the circuit performs the operations described in connection with any of the embodiments described herein.

[0304] FIG. 24A is a block diagram illustrating an IP core development system 2400 that can be used to manufacture an integrated circuit that executes operations according to one embodiment. The IP core development system 2400 can be used to generate a modular reusable design that can be incorporated into a larger design or used to build an entire integrated circuit (e.g., an SOC integrated circuit). The design facility 2430 can generate a software simulation 2410 of the IP core design in a high-level programming language (e.g., C / C++). The software simulation 2410 can be used to design, test, and verify the behavior of the IP core using a simulation model 2412. The simulation model 2412 can include functional simulation, behavioral simulation, and / or timing simulation. Next, a register transfer level (RTL) design 2415 can be created or synthesized from the simulation model 2412. The RTL design 2415 is an abstraction of the behavior of an integrated circuit that models the flow of digital signals between hardware registers and includes the associated logic that is executed using the modeled digital signals. In addition to the RTL design 2415, lower-level designs at the logic level or transistor level can also be created, designed, or synthesized. Thus, the specific details of the initial design and simulation can vary.

[0305] The RTL design 2415 or its equivalent can be further synthesized by a design facility into a hardware model 2420, which can be in the form of a hardware description language (HDL) or some other representation of physical design data. The HDL can be further simulated or tested to verify the IP core design. The IP core design can be stored using a non-volatile memory 2440 (e.g., a hard disk, flash memory, or any non-volatile storage medium) for delivery to a third-party manufacturing facility 2465. Alternatively, the IP core design can be transmitted over a wired connection 2450 or a wireless connection 2460 (e.g., via the Internet). And the manufacturing facility 2465 can manufacture an integrated circuit based at least in part on the IP core design. The manufactured integrated circuit can be configured to perform operations in accordance with at least one embodiment described herein.

[0306] FIG. 24B illustrates a side cross-sectional view of an integrated circuit package assembly 2470. The integrated circuit package assembly 2470 exemplifies the implementation of one or more of the processors or accelerator devices described herein. The package assembly 2470 includes a plurality of units of hardware logic 2472, 2474 connected to a substrate 2480. The logic 2472, 2474 can be implemented, at least in part, in configurable logic or fixed-function logic hardware and can include one or more portions of any of the (one or more) processor cores, (one or more) graphics processors, or other accelerator devices described herein. The logic 2472, 2474 of each unit can be implemented within a semiconductor die and coupled to the substrate 2480 via an interconnect structure 2473. The interconnect structure 2473 can be configured to route electrical signals between the logic 2472, 2474 and the substrate 2480 and can include interconnects such as, but not limited to, bumps or pillars. The interconnect structure 2473 can be configured to route electrical signals related to the operation of the logic 2472, 2474, such as input / output (I / O) signals and / or power or ground signals. Optionally, the substrate 2480 can be an epoxy-based laminated substrate. The substrate 2480 can also include other suitable types of substrates. The package assembly 2470 can be connected to other electrical devices via package interconnects 2483. The package interconnects 2483 are coupled to the surface of the substrate 2480 and can route electrical signals to other electrical devices such as, for example, a motherboard, other chip sets, or multi-chip modules.

[0307] The units of logic 2472, 2474 can be electrically coupled to a bridge 2482 configured to route electrical signals between the logic 2472, 2474. The bridge 2482 can be a high-density interconnect structure that provides a route for the electrical signals. The bridge 2482 can include a bridge substrate composed of glass or a suitable semiconductor material. An electrical routing mechanism can be formed on the bridge substrate to provide an inter-chip connection between the logic 2472, 2474.

[0308] Although two units of logic 2472, 2474, and the bridge 2482 are illustrated, the embodiments described herein can include more or fewer logic units on one or more dies. If the logic is included on a single die, the bridge 2482 can be eliminated, so the one or more dies can be connected by zero or more bridges. Alternatively, multiple dies or units of logic can be connected by one or more bridges. Also, other possible configurations, including three-dimensional configurations, can connect multiple logic units, dies, and bridges together.

[0309] FIG. 24C illustrates a package assembly 2490 that includes a plurality of units of hardware logic chiplets connected to a substrate 2480 (e.g., a base die). The graphics processing units, parallel processors, and / or compute accelerators described herein can be composed of a variety of separately manufactured silicon chiplets. In this context, a chiplet is at least a partially packaged integrated circuit that includes discrete logic units that can be assembled with other chiplets into a larger package. A variety of sets of chiplets with different IP core logics can be assembled into a single device. Additionally, the chiplets can be integrated into a base die or base chiplet using active interposer technology. The concepts described herein enable interconnection and communication between different forms of IP within a GPU. Multiple IP cores can be manufactured using different process technologies and assembled during manufacturing, which avoids the complexity of converging a large number of IPs into the same manufacturing process, especially for large SoCs with several flavors of IP. Enabling the use of multiple process technologies improves the time to market and provides a cost-effective approach to generating a large number of product SKUs. Additionally, unaggregated IPs can be easily power gated independently to power off components not used on a given workload and reduce overall power consumption.

[0310] In various embodiments, the package assembly 2490 can include fewer or greater numbers of components and chiplets interconnected by a fabric 2485 or one or more bridges 2487. The chiplets within the package assembly 2490 may have a 2.5D configuration. This uses a chip-on-wafer-on-substrate stacking where multiple dies are stacked adjacent to each other on a silicon interposer, the silicon interposer includes through-silicon vias (TSVs) that couple the chiplets to the substrate 2480, and the substrate 2480 includes electrical connections to the package interconnects 2483.

[0311] In one embodiment, the silicon interposer is an active interposer 2489 that includes embedded logic in addition to TSVs. In such an embodiment, the dielets within the package assembly 2490 are placed on top of the active interposer 2489 using a 3D face-to-face die stack. The active interposer 2489 can include hardware logic for I / O 2491, cache memory 2492, and other hardware logic 2493 in addition to the interconnect fabric 2485 and silicon bridges 2487. The fabric 2485 enables communication between the various logic dielets 2472, 2474 and the logic 2491, 2493 within the active interposer 2489. The fabric 2485 can be a NoC interconnect or another form of packet-switching fabric that exchanges data packets between the components of the package assembly. In the case of a complex assembly, the fabric 2485 can be a dedicated dielet that enables communication between the various hardware logics of the package assembly 2490.

[0312] The bridge structure 2487 within the active interposer 2489 can be used, for example, to facilitate point-to-point interconnects between a logic or I / O dielet 2474 and a memory dielet 2475. In some implementations, the bridge structure 2487 can be embedded within the substrate 2480.

[0313] The hardware logic chiplet can include a special-purpose hardware logic chiplet 2472, a logic or I / O chiplet 2474, and / or a memory chiplet 2475. The hardware logic chiplet 2472 and the logic or I / O chiplet 2474 can be implemented, at least in part, in configurable logic or fixed-function logic hardware, and can include one or more portions of any of the (one or more) processor cores, (one or more) graphics processors, parallel processors, or other accelerator devices described herein. The memory chiplet 2475 can be DRAM (e.g., GDDR, HBM) memory or cache (SRAM) memory. The cache memory 2492 within the active interposer 2489 (or substrate 2480) can function as a global cache of the package assembly 2490, a part of a distributed global cache, or a dedicated cache of the fabric 2485.

[0314] Each chiplet can be manufactured as a separate semiconductor die and coupled to a base die embedded within or coupled to the substrate 2480. The coupling to the substrate 2480 can be performed via an interconnect structure 2473. The interconnect structure 2473 can be configured to route electrical signals between the chiplet and various logic within the substrate 2480. The interconnect structure 2473 can include interconnects such as, but not limited to, bumps or pillars. In some embodiments, the interconnect structure 2473 can be configured to route electrical signals such as input / output (I / O) signals and / or power or ground signals related to the operation of the logic chiplet, I / O chiplet, and memory chiplet. In one embodiment, an additional interconnect structure couples the active interposer 2489 to the substrate 2480.

[0315] The substrate 2480 can be an epoxy-based laminated substrate, but is not limited thereto. The substrate 2480 may also include other suitable types of substrates. The package assembly 2490 can be connected to other electrical devices via the package interconnect 2483. The package interconnect 2483 is coupled to the surface of the substrate 2480 and can route electrical signals to other electrical devices such as, for example, a motherboard, other chip sets, or a multi-chip module.

[0316] The logic or I / O chiplet 2474 and the memory chiplet 2475 can be electrically coupled via a bridge 2487 configured to route electrical signals between the logic or I / O chiplet 2474 and the memory chiplet 2475. The bridge 2487 can be a high-density interconnect structure that provides a route for electrical signals. The bridge 2487 can include a bridge substrate made of glass or a suitable semiconductor material. An electrical routing mechanism can be formed on the bridge substrate to provide an inter-chip connection between the logic or I / O chiplet 2474 and the memory chiplet 2475. The bridge 2487 can also be referred to as a silicon bridge or an interconnect bridge. For example, the bridge 2487 is an Embedded Multi-die Interconnect Bridge (EMIB). Alternatively, the bridge 2487 can simply be a direct connection from one chiplet to another chiplet.

[0317] FIG. 24D illustrates a package assembly 2494 that includes a replaceable chiplet 2495 according to one embodiment. The replaceable chiplet 2495 can be assembled into standardized slots on one or more base chiplets 2496, 2498. The base chiplets 2496, 2498 can be coupled via a bridge interconnect 2497, which can be similar to other bridge interconnects described herein, e.g., can be an EMIB. Through the bridge interconnect, a memory chiplet can also be connected to a logic or I / O chiplet. The I / O and logic chiplets can communicate via an interconnect fabric. Each of the base chiplets can support one or more slots that are in a standardized form for one of logic or I / O or memory / cache.

[0318] One or more of the base chiplets 2496, 2498, which can be manufactured using a different process technology than the replaceable chiplet 2495 stacked on top of the base chiplet, can have SRAM and power delivery circuits manufactured therein. For example, the base chiplets 2496, 2498 can be manufactured using a larger die process technology, and the replaceable chiplet can be manufactured using a smaller die process technology. One or more of the replaceable chiplets 2495 can be memory (e.g., DRAM) chiplets. Different memory densities can be selected for the package assembly 2494 based on the power and / or performance targeted by the product using the package assembly 2494. Further, based on the power and / or performance targeted by the product, different numbers of different types of functional units can be selected for the logic chiplet at the time of assembly. Further, chiplets including different types of IP logic cores can be inserted into the replaceable chiplet slots, enabling a hybrid processor design that can successfully combine IP blocks of different technologies.

[0319] Exemplary system-on-chip integrated circuit Figures 25-26B illustrate an exemplary integrated circuit and associated graphics processor that can be manufactured using one or more IP cores. In addition to those shown, other logic and circuitry can be included, including additional graphics processors / cores, peripheral interface controllers, or general-purpose processor cores. Elements in Figures 25-26B having the same or similar names as elements in any of the other figures herein describe the same elements in those other figures and can operate or function in a similar manner thereto, can have the same components, and can be coupled to other entities as described elsewhere herein, but are not so limited.

[0320] Figure 25 is a block diagram illustrating an exemplary system-on-chip integrated circuit 2500 that can be manufactured using one or more IP cores. The exemplary integrated circuit 2500 includes one or more application processors 2505 (e.g., CPUs) and at least one graphics processor 2510, which can be a variation of graphics processors 1408, 1508, 2510, or any of the graphics processors described herein and can be used in place of any of the graphics processors described. Accordingly, any disclosure of features in combination with the graphics processors described herein also discloses corresponding combinations with graphics processor 2510, but is not so limited. The integrated circuit 2500 can further include an image processor 2515 and / or a video processor 2520, both of which can be modular IP cores from the same or multiple different design facilities. The integrated circuit 2500 includes a USB controller 2525, a UART controller 2530, an SPI / SDIO controller 2535, and I 2 S / I 2It may include peripheral logic or bus logic including a C controller 2540. Further, the integrated circuit may include a display device 2545 coupled to one or more of a high-definition multimedia interface (HDMI (registered trademark)) controller 2550 and a mobile industry processor interface (MIPI) display interface 2555. Storage may be provided by a flash memory subsystem 2560 including a flash memory and a flash memory controller. A memory interface may be provided by a memory controller 2565 for access to SDRAM or SRAM memory devices. Some integrated circuits further include an embedded security engine 2570.

[0321] Figures 26A-26B are block diagrams showing exemplary graphics processors used within an SoC according to the embodiments described herein. These graphics processors may be variations of graphics processors 1408, 1508, 2510, or any other graphics processor described herein. These graphics processors may be used in place of graphics processors 1408, 1508, 2510, or any other graphics processor described herein. Thus, the disclosure of any feature in combination with graphics processors 1408, 1508, 2510, or the graphics processors described herein also discloses the corresponding combination with the graphics processors of FIGS. 26A-26B, but is not so limited. FIG. 26A shows an exemplary graphics processor 2610 of a system-on-chip integrated circuit that may be manufactured using one or more IP cores according to one embodiment. FIG. 26B shows an exemplary further graphics processor 2640 of a system-on-chip integrated circuit that may be manufactured using one or more IP cores according to one embodiment. The graphics processor 2610 of FIG. 26A is an example of a low-power graphics processor core. The graphics processor 2640 of FIG. 26B is an example of a higher-performance graphics processor core. For example, each of the graphics processors 2610 and 2640 may be a variation of the graphics processor 2510 of FIG. 25, as described at the beginning of this paragraph.

[0322] As shown in FIG. 26A, the graphics processor 2610 includes a vertex processor 2605 and one or more fragment processors 2615A - 2615N (e.g., 2615A, 2615B, 2615C, 2615D, up to 2615N - 1, and 2615N). The graphics processor 2610 is optimized such that the vertex processor 2605 executes operations related to vertex shader programs, while these one or more fragment processors 2615A - 2615N can execute multiple different shader programs via separate logic so as to execute fragment (e.g., pixel) shading operations related to fragment or pixel shader programs. The vertex processor 2605 executes the vertex processing stage of the 3D graphics pipeline to generate primitives and vertex data. The (one or more) fragment processors 2615A - 2615N use the primitives and vertex data generated by the vertex processor 2605 to generate a frame buffer to be displayed on a display device. The (one or more) fragment processors 2615A - 2615N may be optimized to execute fragment shader programs provided by the OpenGL API. The OpenGL API can be used to perform similar operations as pixel shader programs provided by the Direct 3D API.

[0323] The graphics processor 2610 further includes one or more memory management units (MMUs) 2620A-2620B, one or more caches 2625A-2625B, and one or more circuit interconnects 2630A-2630B. The one or more MMUs 2620A-2620B provide virtual-to-physical address mapping for the graphics processor 2610, including the vertex processor 2605 and / or one or more fragment processors 2615A-2615N. This mapping can reference vertex or image / texture data stored in memory in addition to vertex or image / texture data stored in one or more caches 2625A-2625B. The one or more MMUs 2620A-2620B can be synchronized with other MMUs in the system, including one or more MMUs associated with one or more of the application processors 2505, image processors 2515, and / or video processors 2520 of FIG. 25, so that each processor 2505-2520 can participate in a shared or unified virtual memory system. The components of the graphics processor 2610 can correspond to the components of other graphics processors described herein. The one or more MMUs 2620A-2620B can correspond to the MMU 245 of FIG. 2C. The vertex processor 2605 and the fragment processors 2615A-2615N can correspond to the graphics multiprocessor 234. The one or more circuit interconnects 2630A-2630B enable the graphics processor 2610 to interface with other IP cores within the SoC, according to an embodiment, via the internal bus of the SoC or via a direct connection. The one or more circuit interconnects 2630A-2630B can correspond to the data crossbar 240 of FIG. 2C. Further correspondences can be found between the various graphics processor architectures described herein and the similar components of the graphics processor 2610.

[0324] As shown in FIG. 26B, the graphics processor 2640 includes one or more MMUs 2620A - 2620B, (one or more) caches 2625A - 2625B, and (one or more) circuit interconnects 2630A - 2630B of the graphics processor 2610 of FIG. 26A. The graphics processor 2640 includes one or more shader cores 2655A - 2655N (2655A, 2655B, 2655C, 2655D, 2655E, 2655F, up to 2655N - 1, and 2655N), which provide a unified shader core architecture where a single core or type of core can execute all types of programmable shader code, including shader program code that implements vertex shaders, fragment shaders, and / or compute shaders. The exact number of shader cores present can vary between embodiments and implementations. Further, the graphics processor 2640 includes an inter - core task manager 2645 that functions as a thread dispatcher to dispatch execution threads to the one or more shader cores 2655A - 2655N, and a tiling unit 2658 that accelerates tiling operations related to tile - based rendering where the rendering operation of a scene is subdivided within the image space, for example, to utilize local spatial coherence within the scene or to optimize the use of internal caches. The shader cores 2655A - 2655N can correspond to, for example, the graphics multiprocessor 234 as in FIG. 2D, or each of the graphics multiprocessors 325, 350 of FIGS. 3A and 3B, or the multi - core group 365A of FIG. 3C.

[0325] Publication of the task graph to the memory hierarchy A task graph is a set of vertices and directed edges. Each vertex corresponds to some computation, and each edge represents a dependency, which is most often a data dependency (e.g., the output of one computation is the input of another computation).

[0326] A neural network can be represented as a task graph, which is done in some popular deep learning frameworks such as, for example, TensorFlow. Each vertex can represent a computation at a number of possible computational granularities. For example, a vertex may represent a coarse-grained one such as the forward propagation of an entire layer in a neural network, or a finer-grained one such as the data type conversion of a single data point. The coarser-grained vertices are simpler but may expose lower parallelism to the hardware.

[0327] A task graph can have many possible representations, and in some cases, software can provide the entire graph at once or piecemeal. In any case, a graph (or subgraph) has a set of vertices and a set of edges. Each vertex is a computation (e.g., a function or kernel) and the parameters / inputs for that computation. This can be represented, for example, by a tuple such as (function pointer "x", integer value 0, data pointer "y"). This is a fairly common configuration in task scheduling systems such as Intel's TBB. Each edge is a source vertex and a destination vertex and can thus be represented by a pair of indices, pointers (source, destination).

[0328] Software can make library calls to convey a task graph (or subgraph) to the hardware. Depending on what the hardware does with the graph, the device driver or runtime can convey the entire graph to the hardware (e.g., copy the information to a specific predetermined location or provide the address in the graphics memory where the graph is stored), or split the graph into subgraphs and convey only one piece at a time. For example, for each vertex of the graph, the GPU can enqueue the kernel to execute and pass a set of destination vertex ids for that task that can be an extension to an existing kernel descriptor.

[0329] The task graph exposes parallelism / dependencies and communication. In the simplest case where software tells the destination id of each task, the hardware keeps track of which part of the GPU (e.g., which execution unit or SM) is assigned to execute those (one or more) destinations, and can proactively copy the output data from our tasks to be near the destinations. This can be the L1 cache for a particular SM, or a shared cache near the destination.

[0330] In some embodiments, the hardware can do even more with the information. For example, the hardware can assign the destination task to the same hardware unit as the source so that data copies are eliminated (or at least reduced in the case of multiple destinations).

[0331] FIG. 27 is a schematic diagram of a matrix multiplication operation according to some embodiments. Referring to FIG. 27, in some examples, the software interface can be enhanced to enable an application to represent data dependencies across various tasks. The system hardware has knowledge of where tasks are scheduled and moves the output from "producer" (e.g., 2710, 2715) tasks to the cache of "consumer" (e.g., 2720) tasks. In some examples, data from producer tasks 2710, 2715 can be input into cache 2725 communicatively coupled to consumer task 2720.

[0332] Context-aware predictor for controlling prefetch and cacheability FIG. 28 is a schematic diagram of a context-aware predictor 2800 according to the embodiments described herein. Referring to FIG. 28, some applications exhibit pseudo-random patterns in the execution of memory accesses, which can make it difficult to implement heuristic techniques for effectively controlling the hardware within the memory system. For example, making useful decisions about whether to cache a given line or whether to move a page it is on requires the ability to accurately predict whether that cache line will be accessed again in the near future. Most heuristic techniques rely on a sequence of historical addresses (e.g., cache line addresses or page addresses) to predict the future. Those heuristic techniques typically implement a very simple set of rules for making decisions. For example, the N most recently accessed pages can be tracked under the assumption that they are likely to be re-accessed in the near future.

[0333] If accurate predictions cannot be made, the efficiency within the memory system can degrade. For example, a copy of data that will never be touched again can be unnecessarily stored in precious high-speed memory near the computer engine. Further, the space used for that copy needs to be created by pushing out something else that might have been useful.

[0334] Some algorithms are suitable for very "regular" access patterns (e.g., streams of consecutive addresses) where prediction is easy with simple heuristic techniques, while some produce "irregular" access patterns by nature.

[0335] One class of deep learning algorithms that lead to irregular accesses are known as embeddings 2812. These are a technique for mapping very sparse data into dense vectors that will then typically be fed into a neural network 2820 such as a multilayer perceptron (MLP) for example.

[0336] Embedding is generally used to map individual words to short vectors of floating-point values in language processing workloads and in recommendation systems, for example, such that related words (e.g., "mother" and "father") have a small distance between them. In the forward pass of the embedding, a given input value can be looked up in one or more tables and the corresponding table entry retrieved. This is very similar to a set of hash table lookups. Similar to hash table lookups, for a given input stream, the controller will (intentionally) retrieve a stream of seemingly unrelated table entries and the access pattern is pseudo-random by design.

[0337] Existing discovery techniques have not been successful in making good decisions about irregular access patterns, but neural networks are an alternative approach that can be adopted. Neural networks are often successful in identifying patterns even in input streams that appear random.

[0338] Also, one class of neural networks that may handle particularly difficult cases is context-aware networks. Context-aware networks take context information about the input as input or infer it for themselves at another part of the network. And that context information serves as additional input that guides the output to the network. For example, if one has a speech recognition network capable of handling multiple languages, first being able to identify (or being told) what language the speech is in can do a much better job of recognizing specific segments of the speech.

[0339] In some examples, a neural network (which may be a context-aware neural network, but does not necessarily have to be) can be added to a system that is executing a workload that includes embeddings. The embedding output is fed into this prediction network, and recommendations for hardware are made.

[0340] Embeddings are often used in systems that have context information. For example, recommendation systems have long taken in user information and some context (e.g., "that user is visiting web page X") and made recommendations based on that information to maximize the likelihood that the user will click on them, such as what ads to show that user. Thus, some or all of this context information can be passed to a context-aware network as one of its inputs.

[0341] The new neural network may be implemented purely in hardware or may be a combination of hardware and software. For example, a programmable neural network accelerator can be introduced into a processor, and a device driver or other system software can include a program for running a pre-trained network on that accelerator. The user application can include calls to an API to indicate to the device driver or system software that it is performing an embedding lookup so that some or all of the input / output can be passed to the accelerator. Alternatively, the hardware can attempt to detect that the application is performing an embedding lookup and automatically pass the detected input / output completely to the accelerator.

[0342] In some embodiments, predictions from a neural network can be used to control aspects of the memory subsystem. Two specific examples are page copying and cache management policies.

[0343] When an accelerator such as a GPU touches a page in system memory (rather than GPU memory), the hardware and / or system software must determine whether that page should be copied to GPU memory. This can lead to faster future access if the page is likely to be touched again in the near future. If the page is touched enough times in the near future, this can also save bandwidth between system memory and the GPU, since otherwise all of those accesses would generate traffic between system memory and the GPU. However, if the page is not touched enough times in the near future, copying the page to GPU memory can consume more bandwidth than would be consumed by a small (e.g., 64B) access (since the entire page needs to be read to copy it). Also, the optimal decision depends on which data in GPU memory will be replaced to make space for the new page.

[0344] Similarly, when a processor (or accelerator) has a cache hierarchy, it may have a choice of whether to insert data that is not at a particular hierarchical level into the cache. The trade-off is similar to the page replacement problem. However, in a cache, the controller may sometimes influence future choices (i.e., replacement policies) of which data to evict. Thus, in addition to determining whether data should be inserted into a given cache, when the controller decides that it should be inserted, the controller may also have to decide how to set the metadata associated with that line with respect to the replacement policy. For example, it may decide to insert a cache line but (incorrectly) mark it as an "unused recently" line, with the result that it may then be the next one evicted from that set in the cache.

[0345] Hardware-Based Data Prefetching Current GPU prefetching techniques are based on software (i.e., load hoisting). There are cases where the software cannot calculate the address before the time when it hoists the WOD (i.e., prefetches) to hide latency. In such cases, a hardware prefetcher is required to start the prefetch. This unit can be in any of the EU (execution unit), LSC (load store cache), or the next hierarchical level.

[0346] FIG. 29 is a flowchart illustrating operations implemented by a hardware-based prefetcher 2900 according to the embodiments described herein. Referring to FIG. 29, in some examples, the prefetcher 2900 implements an instruction 2910 for monitoring load / store operations, an instruction 2915 for learning a prefetch stride, an instruction 2920 for establishing reliability, and an instruction 2925 for starting a prefetch operation.

[0347] Exemplary examples of the technologies disclosed herein are presented below. Embodiments of the technology may include any one or more and any combination of the examples described below.

[0348] Example 1 includes an apparatus having a plurality of processing resources including a first processing resource and a second processing resource, a memory communicatively coupled to the first processing resource and the second processing resource, and a processor, wherein the processor receives data dependencies regarding one or more tasks having one or more producer tasks executed on the first processing resource and one or more consumer tasks executed on the second processing resource, and moves data output from the one or more producer tasks executed on the first processing resource to a cache memory communicatively coupled to the second processing resource.

[0349] Example 2 includes the matters according to Example 1, wherein the one or more tasks are represented as a task graph as tasks connected by edges.

[0350] Example 3 includes matters according to any one of Examples 1-2, and the processor maps the one or more tasks to the plurality of processing resources.

[0351] Example 4 includes matters according to any one of Examples 1-3, and the processor adds a kernel to a queue for execution by one of the plurality of processing resources.

[0352] Example 5 includes matters according to any one of Examples 1-4, and the processor passes one or more destination identifiers regarding the one or more tasks to the plurality of processing resources.

[0353] Example 6 includes matters according to any one of Examples 1-5, and the cache memory has an L1 cache.

[0354] Example 7 includes matters according to any one of Examples 1-6, and the L1 cache is shared among a plurality of processing resources.

[0355] Example 8 includes a computer-implemented method, the method including receiving, by a processor, data dependencies regarding one or more tasks having one or more producer tasks executed on a first processing resource and one or more consumer tasks executed on a second processing resource, and moving data output from the one or more producer tasks executed on the first processing resource to a cache memory communicatively coupled to the second processing resource.

[0356] Example 9 includes matters according to Example 8, and the one or more tasks are represented as a task graph as tasks connected by edges.

[0357] Example 10 includes matters according to any one of Examples 8-9, and further includes mapping the one or more tasks to the plurality of processing resources.

[0358] Example 11 includes the matters according to any one of Examples 8-10, and further includes adding a kernel to a queue for execution by one of the plurality of processing resources.

[0359] Example 12 includes the matters according to any one of Examples 8-11, and further includes passing one or more destination identifiers related to the one or more tasks to the plurality of processing resources.

[0360] Example 13 includes the matters according to any one of Examples 8-12, and the cache memory has an L1 cache.

[0361] Example 14 includes the matters according to any one of Examples 8-13, and the L1 cache is shared among a plurality of processing resources.

[0362] Example 15 includes a non-transitory computer-readable medium having one or more instructions, which, when executed on at least one processor, cause the at least one processor to receive a data dependency regarding one or more tasks having one or more producer tasks executed on a first processing resource and one or more consumer tasks executed on a second processing resource, and move data output from the one or more producer tasks executed on the first processing resource to a cache memory communicatively coupled to the second processing resource.

[0363] Example 16 includes the matters according to Example 15, and the one or more tasks are represented as a task graph as tasks connected by edges.

[0364] Example 17 includes the matters according to any one of Examples 15-16, and stores instructions that, when executed by one or more processors, cause the one or more processors to map the one or more tasks to the plurality of processing resources.

[0365] Example 18 includes matters according to any of Examples 15 - 17, and stores instructions that, when executed by one or more processors, cause the one or more processors to add a kernel to a queue for execution by one of the plurality of processing resources.

[0366] Example 19 includes matters according to any of Examples 15 - 18, and has storing instructions that, when executed by one or more processors, cause the one or more processors to pass one or more destination identifiers regarding the one or more tasks to the plurality of processing resources.

[0367] Example 20 includes matters according to any of Examples 15 - 19, and the cache memory has an L1 cache.

[0368] Example 21 includes matters according to any of Examples 15 - 20, and the L1 cache is shared among a plurality of processing resources.

[0369] Example 22 is an apparatus having a plurality of processing resources including a first processing resource and a second processing resource, a memory communicatively coupled to the first processing resource and the second processing resource, and a processor, wherein the processor receives one or more embeddings into a context - aware neural network and generates a hardware prediction using the one or more embeddings.

[0370] Example 23 includes matters according to Example 22, and the embedding provides context information to the context - aware neural network.

[0371] Example 24 includes matters according to any one of Examples 21 - 22, and the processor bypasses a cache operation in response to the one or more embeddings.

[0372] Example 25 is an apparatus having a plurality of processing resources including a first processing resource and a second processing resource, a memory communicatively coupled to the first processing resource and the second processing resource, and a processor, where the processor monitors load / store operations, learns one or more prefetch strides, establishes a confidence level, and initiates a prefetch operation.

[0373] The detailed description above includes references to the accompanying drawings that form a part of the detailed description. The drawings illustrate, by way of example, specific embodiments that may be implemented. These embodiments are also referred to herein as "examples." Such examples may include elements in addition to those illustrated or described. However, examples including the elements illustrated or described are also contemplated. Also contemplated are examples using combinations or replacements of any of the elements (or one or more aspects thereof) illustrated or described with respect to a particular example (or one or more aspects thereof) or with respect to other examples (or one or more aspects thereof) illustrated or described herein.

[0374] Publications, patents, and patent documents referenced within this document are hereby incorporated as if individually incorporated. In the event of inconsistent usage between this document and the documents so incorporated, the usage in the incorporated documents shall supplement that of this document, and for irresolvable inconsistencies, the usage in this document shall govern.

[0375] In this document, the terms "a" or "an" are used to include one or more, independent of any other instance or usage of "at least one" or "one or more" as is common in patent documents. Also, "a set of" includes one or more elements. In this document, the term "or" is used to mean non-exclusive unless otherwise specified, so that "A or B" includes "not B but A", "not A but B", and "A and B". In the appended claims, the terms "including" and "in which" are used as plain English equivalents of the terms "comprising" and "wherein", respectively. Also, in the following claims, the terms "including" and "comprising" are open-ended, i.e., a system, apparatus, article, or process that includes elements in addition to those listed after those terms in the claim is still considered to fall within the scope of that claim. Also, in the following claims, the terms "first", "second", "third", etc. are used merely as labels and are not intended to imply any numerical order among them.

[0376] The term "logic instruction" as referred to herein relates to an expression that can be understood by one or more machines that perform one or more logical operations. For example, a logic instruction can have instructions interpretable by a processor compiler to perform one or more operations on one or more data objects. However, this is merely an example of machine-readable instructions and the example is not limiting in this regard.

[0377] The term "computer-readable medium" as referred to herein relates to a medium capable of maintaining a representation perceivable by one or more machines. For example, a computer-readable medium may have one or more storage devices for storing computer-readable instructions or data. Such storage devices may have a storage medium such as, for example, an optical, magnetic, or semiconductor storage medium. However, this is merely an example of a computer-readable medium, and the examples are not limited in this regard.

[0378] The term "logic" as referred to herein relates to a structure for performing one or more logical operations. For example, logic may have a circuit that provides one or more output signals based on one or more input signals. Such a circuit may have a finite state machine that receives digital inputs and provides digital outputs, or a circuit that provides one or more analog output signals in response to one or more analog input signals. Such a circuit may be provided in an application specific integrated circuit (ASIC) or a field programmable gate array (FPGA). Also, logic may have machine-readable instructions stored in a memory, combined with a processing circuit that executes the machine-readable instructions. However, these are merely examples of structures that may provide logic, and the examples are not limiting in this regard.

[0379] Some of the methods described herein may be embodied as logic instructions on a computer-readable medium. When executed on a processor, the logic instructions cause the processor to be programmed as a special purpose machine implementing the described method. When the processor is configured by the logic instructions to execute the methods described herein, it constitutes a structure for executing the described methods. Alternatively, the methods described herein may be reduced to logic on, for example, a field programmable gate array (FPGA), an application specific integrated circuit (ASIC), or the like.

[0380] In the specification and claims, the terms "coupled" and "connected" may be used along with their derivatives. In certain examples, "connected" may be used to indicate that two or more elements are in direct physical or electrical contact with each other. "Coupled" may sometimes mean that two or more elements are in direct physical or electrical contact with each other. However, "coupled" may also mean that two or more elements are not in direct contact with each other but still may cooperate or interact with each other.

[0381] References in the specification to "an example" or "some examples" mean that a particular mechanism, structure, or characteristic described in connection with that example is included in at least one implementation. The appearances of the phrase "in one example" in various places in the specification may or may not all refer to the same example.

[0382] The above description is intended to be illustrative rather than limiting. For example, the above examples (or one or more aspects thereof) may be used in combination with each other. Other embodiments may be used by, for example, those skilled in the art upon consideration of the above description. The abstract is provided to enable the reader to quickly ascertain the essence of the technical disclosure. It is submitted with the understanding that it will not be used to interpret or limit the scope or meaning of the claims. Also, in the above detailed description, various features may be grouped together in order to streamline the disclosure. However, the claims may not describe all of the features disclosed herein. Embodiments may feature a subset of those features because embodiments may include fewer features than those disclosed in a particular example. Also, embodiments may include fewer features than those disclosed in a particular example. Accordingly, the following claims are incorporated into the detailed description, with each claim standing on its own as a separate embodiment. The scope of the embodiments disclosed herein should be determined with reference to the appended claims, along with the full scope of equivalents to which those claims are entitled.

[0383] Examples have been described using terms specific to structural mechanisms and / or methodological acts. It should be understood, however, that the subject matter of the claims may not be limited to the specific features or acts described. Rather, those specific mechanisms and acts are disclosed as examples for carrying out the subject matter of the claims.

[0384] The foregoing description and drawings should be regarded as illustrative rather than restrictive. Those skilled in the art will understand that various changes and modifications can be made to the embodiments described herein without departing from the broader spirit and scope of the invention as set forth in the appended claims.

[0385] An embodiment may be provided as a computer program product that includes one or more machine-readable media storing machine-executable instructions that, when executed by one or more machines, such as a computer, a computer network, or other electronic devices, may cause the one or more machines to perform operations in accordance with the embodiments described herein. The machine-readable media may include, but are not limited to, floppy (registered trademark) disks, optical disks, CD-ROMs (compact disk read-only memories), and magneto-optical disks, ROMs, RAMs, EPROMs (erasable programmable read-only memories), EEPROMs (electrically erasable programmable read-only memories), magnetic or optical cards, flash memories, or other types of media / machine-readable media suitable for storing machine-executable instructions.

[0386] Alternatively, an embodiment may be downloadable as a computer program product, where the program is transferred from a remote computer (e.g., a server) to a requesting computer (e.g., a client) by one or more data signals embodied and / or modulated in a carrier wave or other propagation medium via a communication link (e.g., a modem and / or network connection).

[0387] As those skilled in the art will understand from the above description, the broad technology of the embodiments can be implemented in various forms. Therefore, although the embodiments have been described in relation to their specific examples, other changes will become apparent to those skilled in the art upon consideration of the drawings, the specification, and the following claims, and thus the true scope of the embodiments should not be so limited.

[0388] As those skilled in the art will understand from the above description, the broad technology of the embodiments can be implemented in various forms. Therefore, although the embodiments have been described in relation to their specific examples, other changes will become apparent to those skilled in the art by studying the drawings, the specification, and the following claims, and thus the true scope of the embodiments should not be so limited.

Claims

1. A plurality of processing resources including a first processing resource and a second processing resource, a memory communicatively coupled to the first processing resource and the second processing resource, a processor, having, wherein the processor receives a task graph representing one or more tasks of a neural network, the one or more tasks being connected by edges representing data dependencies, the one or more tasks having one or more producer tasks to be executed on the first processing resource and one or more consumer tasks to be executed on the second processing resource, and moves data output from one or more producer tasks executed on the first processing resource to a cache memory communicatively coupled to the second processing resource, an apparatus.

2. wherein the processor maps the one or more tasks to the plurality of processing resources, The apparatus according to claim 1.

3. wherein the processor adds a kernel to a queue for execution by one of the plurality of processing resources, The apparatus according to claim 2.

4. wherein the processor passes one or more destination identifiers for the one or more tasks to the plurality of processing resources, The apparatus according to claim 3.

5. The apparatus according to claim 1, wherein the cache memory has an L1 cache.

6. The apparatus according to claim 5, wherein the L1 cache is shared among a plurality of processing resources.

7. A computer-implemented method, A processor receives a task graph representing one or more tasks of a neural network, the one or more tasks being connected by edges representing data dependencies, the one or more tasks including one or more producer tasks executed on a first processing resource and one or more consumer tasks executed on a second processing resource, and move data output from the one or more producer tasks executed on the first processing resource to a cache memory communicatively coupled to the second processing resource, A method comprising the steps of: **Claim 8** mapping the one or more tasks to a plurality of processing resources including the first processing resource and the second processing resource, The method according to claim 7, further comprising: **Claim 9** adding a kernel to a queue for execution by one of the plurality of processing resources, The method according to claim 8, further comprising: **Claim 10** passing one or more destination identifiers for the one or more tasks to the plurality of processing resources, The method according to claim 9, further comprising: **Claim 11** The method according to claim 7, wherein the cache memory has an L1 cache. **Claim 12** The method according to claim 11, wherein the L1 cache is shared among the plurality of processing resources. **Claim 13** A non-transitory machine-readable medium storing instructions that, when executed by one or more processors, cause the one or more processors to Receive a task graph representing one or more tasks of a neural network, wherein the one or more tasks are connected by edges representing data dependencies, and the one or more tasks include one or more producer tasks executed on a first processing resource and one or more consumer tasks executed on a second processing resource, and Move data output from the one or more producer tasks executed on the first processing resource to a cache memory communicatively coupled to the second processing resource. A non-transitory machine-readable medium. **Claim 14** When executed by one or more processors, cause the one or more processors to Map the one or more tasks to a plurality of processing resources including the first processing resource and the second processing resource. The non-transitory machine-readable medium according to claim 13, storing instructions. **Claim 15** When executed by one or more processors, cause the one or more processors to Add a kernel to a queue for execution by one of the plurality of processing resources. The non-transitory machine-readable medium according to claim 14, storing instructions. **Claim 16** When executed by one or more processors, cause the one or more processors to Pass one or more destination identifiers related to the one or more tasks to the plurality of processing resources. The non-transitory machine-readable medium according to claim 15, storing instructions. **Claim 17** The non-transitory machine-readable medium according to claim 13, wherein the cache memory has an L1 cache. **Claim 18** The non-transitory machine-readable medium according to claim 17, wherein the L1 cache is shared among a plurality of processing resources.

Citation Information

Patent Citations

  • Method of generating code which is executable by processor, storage area management method, and code generation program

    JP2011128803A

  • Information processing unit and information processing method

    JP2013134670A

  • Memory sharing via unified memory architecture

    JP2017208124A

  • Method and apparatus for filtered coarse pixel shading

    JP2019061713A

  • Multiprocessor, cache synchronization control method and program therefor

    US20110047335A1