Cross-die multicasting from high-bandwidth memory in a graphics processing environment

The GPU with programmable units and pipelining techniques addresses the limitations of fixed function units in graphics processors, enhancing parallel processing efficiency and performance in SIMT architectures through efficient workload distribution and dedicated graphics operations.

DE102025104571A1Pending Publication Date: 2025-09-18INTEL CORP
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
DE102025104571
Authority / Receiving Office
DE · DE
Patent Type
Applications
Current Assignee / Owner
Priority Date
2025-01-06
Filing Date
2025-02-07
Publication Date
2025-09-18

AI Technical Summary

Technical Problem

Current graphics processors face challenges in efficiently processing graphics data due to the limitations of fixed function computational units and the need for improved parallel processing techniques, particularly in systems with single instruction, multiple thread (SIMT) architectures.

Method used

Implementing a graphics processing unit (GPU) with programmable computational units and pipelining techniques to maximize parallel processing across the graphics pipeline, utilizing a scheduler to distribute workloads efficiently among processing clusters, and incorporating dedicated circuitry for graphics operations such as texture sampling and rasterization.

Benefits of technology

Enhances processing efficiency by allowing for a greater variety of operations and improved performance in graphics data processing, particularly in SIMT architectures, by optimizing workload distribution and utilizing dedicated graphics processing resources.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000000_0000_ABST
    Figure 00000000_0000_ABST
Patent Text Reader

Abstract

Disclosed is an apparatus for enabling cross-die multicasting from high-bandwidth memory in a graphics processing environment. The apparatus comprises a first processing die including: an array of processing cores, each including processing resources, shared local memory (SLM), and a local direct memory access (DMA) component; and a cache memory unit communicatively coupled to the array of processing cores, the cache memory unit partitioned with a shared memory cache communicatively coupled to a remote shared memory cache of remote processing dies and including a shared memory DMA component configured to: copy data from a high-bandwidth memory (HBM) of the device to the shared memory cache;Determining that multicast is enabled for the shared memory cache; and if multicast is enabled for the shared memory cache, multicasting the data to the remote shared memory cache of the remote processing dies communicatively coupled to the first processing die;
Need to check novelty before this filing date? Find Prior Art

Description

CROSS-REFERENCE TO RELATED APPLICATIONS

[0001] This application is a non-provisional application of U.S. Provisional Patent Application No. 63 / 564,278, filed March 12, 2024, which is incorporated herein by reference. BACKGROUND OF REVELATION

[0002] Current parallel graphics processing involves systems and methods designed to perform specific operations on graphics data, such as linear interpolation, tessellation, rasterization, texture mapping, depth testing, and so on. Traditionally, graphics processors use fixed-function compute units to process graphics data. Recently, however, parts of graphics processors have been made programmable, allowing such processors to support a wider variety of operations for processing vertex and fragment data.

[0003] To further improve performance, GPUs typically implement processing techniques such as pipelining, which attempt to process as much graphics data as possible in parallel across different parts of the graphics pipeline. Parallel GPUs using SIMT (Single Instruction, Multiple Thread) architectures are designed to maximize the amount of parallel processing in the graphics pipeline. In a SIMT architecture, groups of parallel threads attempt to execute program instructions synchronously together as often as possible to increase processing efficiency. A general overview of software and hardware for SIMT architectures can be found in Shane Cook, "CUDA Programming," Chapter 3, pages 37–51 (2013). BRIEF DESCRIPTION OF THE DRAWINGS

[0004] The embodiments described herein are illustrated by way of example and not limitation in the figures of the accompanying drawings, in which like reference numerals indicate similar elements, and in which: Fig. 1 is a block diagram illustrating a computer system configured to implement one or more aspects of the embodiments described herein; Fig. 2A - 2E illustrate parallel processor components, including graphics multiprocessors; Fig. 3 shows a graphics processing unit including dedicated sets of graphics processing resources arranged in multi-core groups; Fig. 4A-4E illustrate an example architecture in which a plurality of GPUs are communicatively coupled to a plurality of multi-core processors; Fig. 5 illustrates a graphics processing pipeline; Fig.Figure 6 illustrates a machine learning software stack; Fig. 7 illustrates a general-purpose graphics processing unit; Fig. Figure 8 illustrates a multi-GPU computing system; Fig. Figures 9A-9B illustrate layers of example deep neural networks. Fig. 10A-10B illustrate example language models; Fig. 11 illustrates training and deployment of a deep neural network; Fig. 12A is a block diagram illustrating distributed learning; Fig. 12B is a block diagram illustrating a programmable network interface and a data processing unit; Fig. 13 illustrates an exemplary inference system-on-chip (SOC) suitable for performing inference using a trained model; Fig.14 is a block diagram of a processing system; Fig. Figures 15A-15C illustrate computing systems and graphics processors. Fig. 16 is a block diagram of a graphics processor, which may be a discrete or integrated graphics processing unit; Fig. 17A-17B illustrate block diagrams of additional graphics processor and compute accelerator architectures; Fig. 18A-18C illustrate thread execution logic including an array of processing elements deployed in a graphics processor core; Fig. 19 illustrates a tile of a multi-tile processor according to one embodiment. Fig. Figure 20 is a block diagram illustrating the instruction formats of a graphics processor; Fig. 21 is a block diagram of an additional graphics processor architecture; Fig.22A-22B illustrate a graphics processor command format and a graphics processor command sequence; Fig. 23 illustrates an exemplary graphics software architecture for a data processing system; Fig. 24 is a block diagram illustrating an IP core development system; Fig. 25A illustrates a cross-sectional side view of an integrated circuit package assembly including multiple units of hardware logic chiplets connected to a substrate (e.g., base die); Fig. 25B illustrates a package assembly containing replaceable chiplets; Fig. 26 is a block diagram illustrating a system-on-chip integrated circuit; Fig.27 is a block diagram illustrating an exemplary multi-die GPU compute system for providing high-bandwidth cross-die multicasting from memory in a graphics processing environment, according to implementations herein; Fig. 28 is a block diagram illustrating an exemplary multi-die GPU computing system for providing high-bandwidth cross-die multicasting from memory according to implementations herein; Fig. 29 is a flowchart illustrating one embodiment of a method for providing hardware support for cross-die multicasting from high-bandwidth memory in a graphics processing environment; and Fig.30 is a flowchart illustrating one embodiment of a method for providing programmable support for cross-die multicasting from high-bandwidth memory in a graphics processing environment. DETAILED DESCRIPTION

[0005] Current parallel graphics processing involves systems and methods designed to perform specific operations on graphics data, such as linear interpolation, tessellation, rasterization, texture mapping, depth testing, and so on. Traditionally, graphics processors use fixed-function compute units to process graphics data. Recently, however, parts of graphics processors have been made programmable, allowing such processors to support a wider variety of operations for processing vertex and fragment data.

[0006] To further improve performance, graphics processors typically implement processing techniques such as pipelining, which attempt to process as much graphics data as possible in parallel across the various parts of the graphics pipeline. Parallel graphics processors with SIMT (Single Instruction, Multiple Thread) architectures are designed to maximize the amount of parallel processing in the graphics pipeline. In a SIMT architecture, groups of parallel threads attempt to execute program instructions synchronously together as often as possible to increase processing efficiency.

[0007] A graphics processing unit (GPU) is communicatively coupled to host processor cores to accelerate, for example, graphics operations, machine learning operations, pattern analysis operations, and / or various general-purpose GPU (GPGPU) functions. The GPU may be communicatively coupled to the host processor cores via a bus or other interconnect (e.g., a high-speed interconnect such as PCIe or NVLink). Alternatively, the GPU may be integrated into the same package or chip as the cores and communicatively coupled to the cores via an internal processor bus / interconnect (i.e., within the package or chip). Regardless of how the GPU is connected, the processor cores may allocate work to the GPU in the form of sequences of commands / instructions contained in a work descriptor.The GPU then uses dedicated circuitry / logic to efficiently process these commands / instructions.

[0008] Although some techniques described here are discussed primarily in the context of a GPU, the techniques may also be implemented in other types of processors, including, but not limited to, general-purpose processors and accelerator devices such as artificial intelligence accelerators, vision processors, and neural processing units.

[0009] In the following description, numerous specific details are set forth in order to provide a more thorough understanding. However, it will be apparent to one skilled in the art that the presently described embodiments may be practiced without one or more of these specific details. In other instances, well-known features have been omitted in order to clearly illustrate the details of the present embodiments. System overview

[0010] Fig.1 is a block diagram illustrating a computing system 100 configured to implement one or more aspects of the embodiments described herein. Computing system 100 includes a processing subsystem 101 having one or more processors 102 and a system memory 104 communicating via an interconnect path that may include a memory hub 105. Memory hub 105 may be a separate component within a chipset component or may be integrated with the one or more processors 102. Memory hub 105 is coupled to an I / O subsystem 111 via a communication link 106. I / O subsystem 111 includes an I / O hub 107, which may enable computing system 100 to receive input from one or more input devices 108.The I / O hub 107 may additionally enable a display controller, which may be included in the one or more processors 102, to provide outputs to one or more display devices 110A. In one embodiment, the one or more display devices 110A coupled to the I / O hub 107 may include a local, internal, or embedded display device.

[0011] For example, the processing subsystem 101 includes one or more parallel processors 112 coupled to the memory hub 105 via a communication link 113, such as a bus or fabric. The communication link 113 may be one of any number of standards-based communication link technologies or protocols, such as, but not limited to, PCI Express, or it may be a vendor-specific communication interface or fabric. The one or more parallel processors 112 may form a compute-focused parallel or vector processing system that may include a large number of processing cores and / or processing clusters, such as an integrated multi-core processor (MIC processor).For example, the one or more parallel processors 112 form a graphics processing subsystem that can output pixels to one or more display devices 110A coupled via the I / O hub 107. The one or more parallel processors 112 may also include a display controller (not shown) and a display interface (not shown) to enable direct connection to one or more display devices 110B.

[0012] Within the I / O subsystem 111, a system storage unit 114 may connect to the I / O hub 107 to provide a storage mechanism for the computing system 100. An I / O switch 116 may be used to provide an interface mechanism to enable connections between the I / O hub 107 and other components, such as a network adapter 118 and / or a wireless network adapter 119, that may be integrated into the platform, and various other devices that may be added via one or more add-in devices 120. The add-in device(s) 120 may also include, for example, one or more external graphics processing devices, graphics cards, and / or compute accelerators. The network adapter 118 may be an Ethernet adapter or other wired network adapter.The wireless network adapter 119 may include one or more of a Wi-Fi, Bluetooth, near field communication (NFC), or other network device that includes one or more wireless radios.

[0013] The computing system 100 may include other components not explicitly shown, including USB or other port connections, optical storage drives, video capture devices, and the like, which may also be connected to the I / O hub 107. The communication paths connecting the various components in Fig.1 can be implemented using any suitable protocols, such as PCI (Peripheral Component Interconnect) based protocols (e.g. PCI Express) or any other bus or point-to-point communication interfaces and / or protocols, such as the NVLink high-speed interconnect, Compute Express Link™ (CXL™) (e.g. CXL.mem), Infinity Fabric (IF), Ethernet (IEEE 802.3), Remote Direct Memory Access (RDMA), InfiniBand, Internet Wide Area RDMA Protocol (iWARP), Transmission Control Protocol (TCP), User Datagram Protocol (UDP), quick UDP Internet Connections (QUIC), RDMA over Converged Ethernet (RoCE), Intel QuickPath Interconnect (QPI), Intel Ultra Path Interconnect (UPI), Intel On-Chip System Fabric (IOSF), OmniPath, HyperTransport, Advanced Microcontroller Bus Architecture (AMBA) interconnect, OpenCAPI, Gen-Z, Cache Coherent Interconnect for Accelarators (CCIX), 3GPP Long Term Evolution (LTE) (4G), 3GPP 5G and variations thereof or wired or wireless interconnect protocols known in the art. In some examples, data may be transferred using a protocol such as NVMe (Non-Volatile Memory Express). NVMe-oF (Non-Volatile Memory Express over Fabrics) or NVMe, can be copied or stored in virtualized data storage nodes.

[0014] The one or more parallel processors 112 may include circuitry optimized for graphics and video processing, including, for example, video output circuitry, and forming a graphics processing unit (GPU). Alternatively or additionally, the one or more parallel processors 112 may include circuitry optimized for general-purpose processing while maintaining the underlying computing architecture described in more detail herein. Components of the computing system 100 may be integrated with one or more other system elements on a single integrated circuit. For example, the one or more parallel processors 112, the memory hub 105, the one or more processors 102, and the I / O hub 107 may be integrated into a system-on-chip (SoC) integrated circuit.Alternatively, the components of computing system 100 may be integrated into a single package to form a system-in-package (SIP) configuration. In one embodiment, at least a portion of the components of computing system 100 may be integrated into a multi-chip module (MCM), which may be interconnected with other multi-chip modules to form a modular computing system. In some configurations, computing system 100 includes, in addition to the one or more processors 102 and the one or more parallel processors 112, one or more accelerator devices 130 coupled to memory hub 105. The one or more accelerator devices 130 are configured to perform domain-specific acceleration of workloads to handle tasks that are computationally intensive or utilize high throughput.The one or more accelerator devices 130 may reduce the load imposed on the one or more processors 102 and / or the one or more parallel processors 112 of the computing system 100. The one or more accelerator devices 130 may include, but are not limited to, intelligent network interface cards, compute units, cryptographic accelerators, storage accelerators, artificial intelligence (AI) accelerators, neural processing units (NPUs), storage accelerators, and / or video transcoding accelerators.

[0015] It should be understood that the computing system 100 shown herein is illustrative and that variations and modifications are possible. The interconnect topology, including the number and arrangement of bridges, the number of processor(s) 102, and the number of parallel processor(s) 112, may be modified as desired. For example, the system memory 104 may be connected directly to the one or more processors 102 rather than via a bridge, while other devices communicate with the system memory 104 via the memory hub 105 and the one or more processors 102. In other alternative topologies, the one or more parallel processors 112 are connected to the I / O hub 107 or directly to one of the one or more processors 102 rather than to the memory hub 105. In other embodiments, the I / O hub 107 and the memory hub 105 may be integrated into a single chip.It is also possible for two or more sets of processors 102 to be mounted across multiple sockets, which may be coupled to two or more instances of the one or more parallel processors 112.

[0016] Some of the specific components shown herein are optional and may not be included in all implementations of computing system 100. For example, any number of add-in cards or peripherals may be supported, or some components may be eliminated. Furthermore, some architectures may use different terminology for components similar to those shown in Fig. 1. For example, in some architectures, the memory hub 105 may be referred to as a northbridge, while the I / O hub 107 may be referred to as a southbridge.

[0017] Fig.2A illustrates a parallel processor 200. The parallel processor 200 may be a GPU, GPGPU, or the like, as described herein. The various components of the parallel processor 200 may be implemented using one or more integrated circuit devices, such as programmable processors, application-specific integrated circuits (ASICs), or field-programmable gate arrays (FPGAs). The parallel processor 200 may be one or more of the Fig. 1 shown parallel processors 112.

[0018] The parallel processor 200 includes a parallel processing unit 202. The parallel processing unit includes an I / O unit 204 that enables communication with other devices, including other instances of the parallel processing unit 202. The I / O unit 204 may be directly connected to other devices. For example, the I / O unit 204 is connected to other devices using a hub or switch interface, such as the storage hub 105. The connections between the storage hub 105 and the I / O unit 204 form a communication link 113.Within the parallel processing unit 202, the I / O unit 204 is connected to a host interface 206 and a memory crossbar 216, where the host interface 206 receives commands directed to perform processing operations, and the memory crossbar 216 receives commands directed to perform memory operations. In one embodiment, the I / O unit 204 is configured to enable secure I / O operations via TEE (Trusted Execution Environment) I / O support. The TEE EA enables trusted I / O virtualization, where a trust relationship can be established directly between a secure virtual environment, such as a trusted virtual machine, and the parallel processor 200 or secure partitions of the parallel processor.

[0019] When host interface 206 receives a command buffer via I / O unit 204, host interface 206 may direct work operations to a front end 208 to perform those commands. In one embodiment, front end 208 is coupled to a scheduler 210 configured to dispatch commands or other work items to a processing cluster array 212. Scheduler 210 ensures that processing cluster array 212 is properly configured and in a valid state before dispatching tasks to the processing clusters of processing cluster array 212. Scheduler 210 may be implemented by firmware logic executing in a microcontroller.The microcontroller-implemented scheduler 210 is configurable to perform complex scheduling and work distribution operations at both coarse and fine granularity, enabling rapid deferral and context switching of threads executing in the processing cluster array 212. Preferably, the host software can indicate the workloads to be scheduled in the processing cluster array 212 through one of several graphics processing doorbells. In other examples, polling for new workloads or interrupts can be used to identify or indicate the availability of work to be performed. The workloads can then be automatically distributed across the processing cluster array 212 by the logic of the scheduler 210 within the scheduler's microcontroller.

[0020] Processing cluster array 212 may include up to "N" processing clusters (e.g., cluster 214A, cluster 214B, through cluster 214N). Each cluster 214A through 214N of processing cluster array 212 may execute a large number of concurrent threads. Scheduler 210 may allocate work to clusters 214A through 214N of processing cluster array 212 using various scheduling and / or work distribution algorithms, which may vary depending on the workload incurred for each type of program or computation. Scheduling may be handled dynamically by scheduler 210 or may be partially assisted by compiler logic during compilation of program logic configured for execution by processing cluster array 212.Optionally, different clusters 214A through 214N of the processing cluster array 212 may be assigned to process different types of programs or to perform different types of calculations.

[0021] The processing cluster array 212 may be configured to perform various types of parallel processing operations. For example, the processing cluster array 212 is configured to perform general-purpose parallel computing operations. For example, the processing cluster array 212 may include logic for performing processing tasks, including filtering video and / or audio data, performing modeling operations that include physics operations, and performing data transformations.

[0022] The processing cluster array 212 is configured to perform parallel graphics processing operations. In embodiments where the parallel processor 200 is configured to perform graphics processing operations, the processing cluster array 212 may include additional logic to support the execution of these graphics processing operations, including, but not limited to, texture sampling logic for performing texture operations, as well as tessellation logic and other vertex processing logic. The processing cluster array 212 may also be configured to execute shader programs related to graphics processing, such as, but not limited to, vertex shaders, tessellation shaders, geometry shaders, and pixel shaders. The parallel processing unit 202 may transfer data from system memory for processing via the I / O unit 204.The transferred data can be stored in an on-chip memory (e.g., the parallel processor memory 222) during processing and then written back to the system memory.

[0023] In embodiments where parallel processing unit 202 is used to perform graphics processing, scheduler 210 may be configured to divide the processing workload into approximately equal-sized tasks to better facilitate distribution of graphics processing operations across multiple clusters 214A-214N of processing cluster array 212. In some of these embodiments, portions of processing cluster array 212 may be configured to perform other types of processing. For example, a first portion may be configured to perform vertex shading and topology generation, a second portion may be configured to perform tessellation and geometry shading, and a third portion may be configured to perform pixel shading or other screen-space operations to produce a rendered image for a display.Intermediate data generated by one or more of the clusters 214A through 214N may be stored in buffers to allow the intermediate data to be transferred between the clusters 214A through 214N for further processing.

[0024] During operation, the processing cluster array 212 may receive processing tasks to be executed via the scheduler 210, which receives commands defining processing tasks from the front end 208. For graphics processing operations, processing tasks may include indices of data to be processed, e.g., surface (patch) data, primitive data, vertex data, and / or pixel data, as well as state parameters and commands that define how the data should be processed (e.g., which program should be executed). The scheduler 210 may be configured to retrieve the indices corresponding to the tasks or may receive the indices from the front end 208. The front end 208 may be configured to ensure that the processing cluster array 212 has been configured to a valid state before initiating the workload specified by buffers for incoming commands (e.g., stack buffers, push buffers, etc.).

[0025] Each of the one or more instances of parallel processing unit 202 may be coupled to parallel processor memory 222. Parallel processor memory 222 may be accessed via memory crossbar 216, which may receive memory requests from processing cluster array 212 as well as I / O unit 204. Memory crossbar 216 may access parallel processor memory 222 via a memory interface 218. Memory interface 218 may include a plurality of partition units (e.g., partition unit 220A, partition unit 220B, through partition unit 220N), each of which may be coupled to a portion (e.g., a memory unit) of parallel processor memory 222.The number of partition units 220A to 220N may be configured to be equal to the number of storage units such that each partition unit 220A0 to 220N has a corresponding storage unit 224A to 224N. In other embodiments, the number of partition units 220A to 220N need not be equal to the number of storage devices.

[0026] Memory units 224A-224N may include various types of memory devices, including dynamic random access memory (DRAM) or graphics random access memory, such as synchronous graphics random access memory (SGRAM), including double data rate graphics memory (GDDR). Optionally, memory units 224A-224N may also include 3D stacked memories, including, but not limited to, high bandwidth memories (HBM). Those skilled in the art will appreciate that the specific implementation of memory units 224A-224N may vary and may be selected from any of several conventional designs.Render targets, such as frame buffers or texture maps, may be stored across memory units 224A through 224N, allowing partition units 220A through 220N to write portions of each render target in parallel to efficiently utilize the available bandwidth of parallel processor memory 222. In some embodiments, a local instance of parallel processor memory 222 may be eliminated in favor of a unified memory design that utilizes system memory in conjunction with a local cache.

[0027] Optionally, any of the clusters 214A-214N of the processing cluster array 212 has the capability to process data written to any of the memory units 224A-224N within the parallel processor memory 222. The memory crossbar 216 may be configured to transfer the output of each cluster 214A-214N to any partition unit 220A-220N or to another cluster 214A-214N, which may perform additional processing operations on the output. Each cluster 214A-214N may communicate with the memory interface 218 via the memory crossbar 216 to read from or write to various external memory devices.In one embodiment with the memory crossbar 216, the memory crossbar 216 includes a connection to the memory interface 218 for communicating with the I / O unit 204, as well as a connection to a local instance of the parallel processor memory 222 to enable the processing units within the various processing clusters 214A to 214N to communicate with system memory or other memory not locally contained within the parallel processing unit 202. In principle, the memory crossbar 216 may, for example, be capable of using virtual channels to separate traffic flows between the clusters 214A to 214N and the partition units 220A to 220N.

[0028] Although a single instance of the parallel processing unit 202 is illustrated within the parallel processor 200, any number of instances of the parallel processing unit 202 may be included. For example, multiple instances of the parallel processing unit 202 may be provided on a single add-in card, or multiple add-in cards may be interconnected. For example, the parallel processor 200 may be an add-in device, such as the add-in device 120 of Fig.1, which may be a graphics card, such as a discrete graphics card, including one or more GPUs, one or more memory devices, and device-to-device or network or fabric interfaces. The various instances of the parallel processing unit 202 may be configured to operate with each other, even if the different instances have different numbers of processing cores, different amounts of local parallel processor memory, and / or other configuration differences. Optionally, some instances of the parallel processing unit 202 may include higher-precision floating-point units relative to other instances.Systems including one or more instances of the parallel processing unit 202 or parallel processor 200 may be implemented in a variety of configurations and form factors, including, but not limited to, desktop, laptop, or handheld personal computers, servers, workstations, gaming consoles, and / or embedded systems. An orchestrator may form federated nodes for executing workloads using one or more of the following: disaggregated processor resources, cache resources, memory resources, data storage resources, and network resources.

[0029] In one embodiment, the parallel processing unit 202 may be partitioned into multiple instances. These multiple instances may be configured to execute workloads associated with different clients in an isolated manner, enabling a predetermined quality of service to be provided to each client. For example, each cluster 214A through 214N may be split and isolated from other clusters, enabling the processing cluster array 212 to be divided into multiple compute partitions or instances. In such a configuration, workloads executing in an isolated partition are protected from disruptions or errors associated with another workload executing in a different partition.Partition units 220A through 220N may be configured to enable a dedicated and / or isolated path to storage for clusters 214A through 214N associated with the respective compute partitions. This datapath isolation allows the compute resources within a partition to communicate with one or more assigned storage units 224A through 224N without being subject to interference from the activities of other partitions. In one embodiment, datapath isolation may be enhanced by encrypting the data in the memory of the various partitions with encryption keys associated with a corresponding partition, thus securing data on a per-partition basis while at rest and in transition.

[0030] Fig.2B is a block diagram of a partition unit 220. The partition unit 220 may be an instance of one of the partition units 220A to 220N of Fig. 2A. As illustrated, the partition unit 220 includes an L2 cache 221, a frame buffer interface 225, and a raster operation unit (ROP) 226. The L2 cache 221 is a read / write cache configured to perform load and store operations received from the memory crossbar 216 and the ROP 226. Read misses and write-back requests are issued from the L2 cache 221 to the frame buffer interface 225 for processing. Updates may also be sent to the frame buffer via the frame buffer interface 225 for processing. In one embodiment, the frame buffer interface 225 is connected to a memory unit 224 of the memory units 224A-224N within the parallel processor memory 222. Fig.2A. The partition unit 220 may additionally or alternatively be connected to one of the memory units in the parallel processor memory via a memory controller (not shown).

[0031] In graphics applications, the ROP 226 is a processing unit that performs raster operations such as stencil, z-test, blending, and the like. The ROP 226 then outputs processed graphics data, which is stored in graphics memory. In some embodiments, the ROP 226 includes or is coupled to a CODEC 227, which includes compression logic to compress depth or color data written to memory or L2 cache 221 and to decompress depth or color data read from memory or L2 cache 221. The compression logic may be lossless compression logic using one or more of several compression algorithms. The type of compression performed by the CODEC 227 may vary based on the statistical properties of the data to be compressed.For example, in one embodiment, delta color compression is performed on depth and color data on a per-tile basis. In one embodiment, CODEC 227 includes compression and decompression logic that can compress and decompress computational data associated with machine learning operations. For example, CODEC 227 can compress sparse matrix data for sparse machine learning operations. CODEC 227 can also compress sparse matrix data encoded in a sparse matrix format (e.g., coordinate list encoding (COO), compressed sparse row (CSR), compressed sparse column (CSC), etc.) to produce compressed and encoded sparse matrix data.The compressed and encoded sparse matrix data may be decompressed and / or decoded before being processed by processing elements, or the processing elements may be configured to consume compressed, encoded, or compressed and encoded data for processing. In one embodiment, CODEC 227 may be configured as a general-purpose data compression engine for use in GPU database acceleration and high-volume data analytics.

[0032] The ROP 226 can be installed in any processing cluster (e.g., clusters 214A-214N of Fig.2A) rather than contained in the partition unit 220. In such an embodiment, read and write requests for pixel data are transmitted via the memory crossbar 216 instead of pixel fragment data. The processed graphics data may be displayed on a display device, such as one of the one or more display devices 110A, 110B of Fig. 1, for further processing by the one or more processors 102 or for further processing by one of the processing units within the parallel processor 200 of the Fig. 2A.

[0033] Fig. 2C is a block diagram of a processing cluster 214 in a parallel processing unit. For example, the processing cluster 214 is an instance of one of the processing clusters 214A through 214N of Fig.2A. The processing cluster 214 may be configured to execute many threads in parallel, where the term "thread" refers to an instance of a specific program executing on a specific set of input data. Optionally, single-instruction multiple-data (SIMD) instruction issuing techniques may be used to support parallel execution of a large number of threads without providing multiple independent instruction units. Alternatively, single-instruction multi-threaded (SIMT) techniques are used to support parallel execution of a large number of generally synchronized threads, using a common instruction unit configured to issue instructions to a set of processing engines within each of the processing clusters.Unlike a SIMD execution regime, where all processing engines typically execute identical instructions, SIMT execution allows different threads to more easily follow divergent execution paths through a given threaded program. Those skilled in the art will understand that a SIMD processing regime is a functional subset of a SIMT processing regime.

[0034] The operation of the processing cluster 214 can be controlled by a pipeline manager 232, which distributes processing tasks to SIMT parallel processors. The pipeline manager 232 receives instructions from the scheduler 210. Fig.2A and manages the execution of these instructions via a graphics multiprocessor 234 and / or a texture unit 236. The graphics multiprocessor 234 is an example instance of a SIMT parallel processor. However, various types of SIMT parallel processors with different architectures may be included within the processing cluster 214. One or more instances of the graphics multiprocessor 234 may be included within a processing cluster 214. The illustrated graphics multiprocessor 234 may also be referred to as a streaming multiprocessor (SM) and is capable of executing a large number of execution threads concurrently.

[0035] The graphics multiprocessor 234 may process data, and a data crossbar 240 may be used to distribute the processed data to one of several possible destinations, including instances of the graphics multiprocessor 234 within the processing cluster 214. The pipeline manager 232 may facilitate the distribution of processed data by specifying destinations for processed data to be distributed via the data crossbar 240. Each graphics multiprocessor 234 within the processing cluster 214 may include an identical set of functional execution logic (e.g., arithmetic logic units, load-store units, etc.). The functional execution logic may be configured in a pipeline-like manner, where new instructions may be issued before previous instructions have completed.The functional execution logic supports a wide variety of operations, including integer and floating-point arithmetic, comparison operations, Boolean operations, bit shifting, and calculations of various algebraic functions. The same hardware consisting of functional units could be leveraged to perform other operations, and any combination of functional units can be present.

[0036] The instructions transferred to the processing cluster 214 form a thread. A set of threads executing across the set of parallel processing engines is a thread group. A thread group executes the same program on different input data. Each thread within a thread group can be assigned to a different processing engine within a graphics multiprocessor 234. A thread group can include fewer threads than the number of processing engines within the graphics multiprocessor 234. If a thread group includes fewer threads than the number of processing engines, one or more of the processing engines may be idle during the cycles in which that thread group is processing. A thread group can also include more threads than the number of processing engines within the graphics multiprocessor 234.If the thread group includes more threads than the number of processing engines within the graphics multiprocessor 234, the processing may be performed distributed over consecutive clock cycles. Optionally, multiple thread groups may execute concurrently in the graphics multiprocessor 234.

[0037] The graphics multiprocessor 234 may include an internal cache to perform load and store operations. Alternatively, the graphics multiprocessor 234 may forgo an internal cache and utilize a cache (e.g., a Level 1 (L1) cache 248) within the processing cluster 214. Each graphics multiprocessor 234 also has access to Level 2 (L2) caches within the partition units (e.g., partition units 220A through 220N of the Fig.2A) that are shared by all processing clusters 214 and that can be used to transfer data between threads. The graphics multiprocessor 234 can also access a global off-chip memory, which can include one or more of the local parallel processor memory and / or the system memory. Any memory external to the parallel processing unit 202 can be used as global memory. Embodiments in which the processing cluster 214 includes multiple instances of the graphics multiprocessor 234 can share common instructions and data, which can be stored in the L1 cache 248.

[0038] Each processing cluster 214 may include a memory management unit (MMU) 245 configured to map virtual addresses to physical addresses. In other embodiments, one or more instances of the MMU 245 may be included within the memory interface 218. Fig.2A. The MMU 245 includes a set of page table entries (PTEs) used to map a virtual address to a physical address of a tile and optionally to a cache line index. The MMU 245 may include translation lookaside buffers (TLBs) or caches that may be located within the graphics multiprocessor 234 or the L1 cache 248 of the processing cluster 214. The physical address is processed to distribute access locality of surface data to enable efficient interleaving of requests between partition units. The cache line index may be used to determine whether a request for a cache line is a hit or a miss.

[0039] In graphics and computing applications, a processing cluster 214 may be configured such that each graphics multiprocessor 234 is coupled to a texture unit 236 for performing texture mapping operations, such as determining texture sampling positions, reading texture data, and filtering the texture data. Texture data is read from an internal texture L1 cache (not shown) or, in some embodiments, from the L1 cache within the graphics multiprocessor 234 and is retrieved from an L2 cache, local parallel processor memory, or system memory. Each graphics multiprocessor 234 issues processed tasks to the data crossbar 240 to provide the processed tasks to another processing cluster 214 for further processing or to store the processed task in an L2 cache, local parallel processor memory, or system memory via the memory crossbar 216.A pre-ROP 242 (pre-raster operation unit) is configured to receive data from the graphics multiprocessor 234 and to forward data to ROP units located in the partition units described herein (e.g., partition units 220A through 220N in . Fig. 2A). The preROP unit 242 can perform optimizations for color mixing, organize pixel color data, and perform address translations.

[0040] It should be understood that the core architecture described herein is illustrative and that variations and modifications are possible. Any number of processing units, e.g., the graphics multiprocessor 234, the texture units 236, the preROPs 242, etc., may be included within a processing cluster 214. A parallel processing cluster as described herein may include any number of instances of the processing cluster 214. Optionally, each processing cluster 214 may be configured to operate independently of other instances of the processing cluster 214 using, for example, separate and individual processing units, L1 caches, L2 caches, etc.

[0041] Fig.2D shows an example of the graphics multiprocessor 234, where the graphics multiprocessor 234 is coupled to the pipeline manager 232 of the processing cluster 214. The graphics multiprocessor 234 has an execution pipeline including, but not limited to, an instruction cache 252, an instruction unit 254, an address mapping unit 256, a register file 258, one or more general-purpose graphics processing unit (GPGPU) cores 262, and one or more load / store units 266. The GPGPU cores 262 and the load / store units 266 are coupled to the cache memory 272 and the shared memory 270 via a memory and cache interconnect 268. The graphics multiprocessor 234 may additionally include ray tracing cores 263, which include hardware logic for accelerating ray tracing operations, and tensor cores 264, which include hardware logic for accelerating tensor operations (e.g., matrix operations).Instruction cache 252 may receive a stream of instructions to be executed from pipeline manager 232. The instructions are cached in instruction cache 252 and dispatched for execution by instruction unit 254. Instruction unit 254 may dispatch instructions as thread groups (e.g., chains), with each thread of the thread group being assigned to a different execution unit within GPGPU core 262. An instruction may access any local, shared, or global address space by specifying an address within a unified address space. Address mapping unit 256 may be used to translate addresses in the unified address space into a unique memory address accessible by load / store units 266.

[0042] Register file 258 provides a set of registers for the functional units of graphics multiprocessor 234. Register file 258 provides temporary storage for operands associated with the data paths of the functional units (e.g., GPGPU cores 262, load / store units 266) of graphics multiprocessor 234. Register file 258 may be partitioned between each of the functional units, so that each functional unit is assigned a dedicated section of register file 258. For example, register file 258 may be partitioned between the various concatenations executed by graphics multiprocessor 234.

[0043] The GPGPU cores 262 may each include floating-point units (FPUs) and / or integer arithmetic logic units (ALUs) used to execute instructions of the graphics multiprocessor 234. In some implementations, the GPGPU cores 262 may include hardware logic that may otherwise be located in the tensor cores 264 and / or ray tracing cores 263. The GPGPU cores 262 may have a similar architecture or differ in their architecture. For example, and in one embodiment, a first portion of the GPGPU cores 262 includes a single-precision FPU and an integer ALU, while a second portion of the GPGPU cores includes a double-precision FPU. The FPUs can either implement the IEEE standard 754-2008 for floating-point arithmetic or enable variable-precision floating-point arithmetic.In one embodiment, the single-precision FPUs or a separate set of FPUs are configurable to perform operations on 16-bit floating-point operands, such as operands in a half-precision or bfloat16 format (e.g., Brain floating-point), which is a 16-bit floating-point format with one sign bit, eight exponent bits, and eight significand bits, seven of which are explicitly stored. The FPUs within one or more of the GPGPU cores 262 may also support one or more 8-bit floating-point formats. Supported 8-bit floating-point formats include the E4M3 format, which has a 4-bit exponent and a 3-bit mantissa, and the E5M2 format, which has a 5-bit exponent and a 2-bit mantissa. The graphics multiprocessor 234 may also include one or more fixed function or special function units to perform specific functions such as copy rectangle or pixel blending operations.One or more of the GPGPU cores may also include fixed or special function logic.

[0044] The GPGPU cores 262 may include SIMD logic capable of executing a single instruction on multiple data sets. Optionally, the GPGPU cores 262 may physically execute SIMD4, SIMD8, and SIMD16 instructions and logically execute SIMD1, SIMD2, and SIMD32 instructions. The SIMD instructions for the GPGPU cores may be generated at compile time by a shader compiler or automatically generated when executing programs written and compiled for single program multiple data (SPMD) or SIMT architectures. Multiple threads of a program configured for the SIMT execution model may execute via a single SIMD instruction. For example, and in one embodiment, eight SIMT threads performing the same or similar operations may execute in parallel as a single SIMD8 instruction.In one embodiment, a chain of 32 SIMT threads can be executed as a single SIMD32 instruction. Chain divergence can be handled by multiple SIMD instructions.

[0045] The memory and cache interconnect 268 is an interconnect network that connects each of the functional units of the graphics multiprocessor 234 to the register file 258 and to the shared memory 270. For example, the memory and cache interconnect 268 is a crossbar interconnect that allows the load / store unit 266 to implement load and store operations between the shared memory 270 and the register file 258. The register file 258 can operate at the same frequency as the GPGPU cores 262, so that data transfer between the GPGPU cores 262 and the register bank 258 has very low latency. The shared memory 270 can be used to enable communication between threads executing in the functional units within the graphics multiprocessor 234. The shared memory 270 can also be used as a program-managed cache.For example, cache 272 may be used as an automatically managed data cache to cache texture data communicated between the functional units and texture unit 236. Shared memory 270 and cache 272 may be coupled to data crossbar 240 to enable communication with other components of the processing cluster, thereby enabling cooperative execution of cluster workgroups across multiple graphics multiprocessors within a processing cluster. Threads executing in GPGPU cores 262 may programmatically store data in the shared memory in addition to the automatically cached data stored in cache 272.In one embodiment, the shared memory 270 and the cache memory 272 may be combined into a single configurable memory unit that may be selectively configured as the cache memory 272 or the shared memory 270.

[0046] Fig. Figure 2E illustrates a graphics multiprocessor 235 having an alternative configuration with respect to the graphics multiprocessor 234 of Fig. 2D. The disclosure of any features in combination with the graphics multiprocessor 235 described herein also discloses a corresponding combination with the graphics multiprocessor 234 of Fig. 2D, but is not limited to it. The graphics multiprocessor 235 of the Fig. 2E includes several additional instances of execution resource units 286A to 286D with respect to the graphics multiprocessor 234 of the Fig.2D. For example, the graphics multiprocessor 235 may include a plurality of instruction units 254A through 254D, the register files 258A through 258D, and the texture units 280A through 280D. The graphics multiprocessor 235 may also include a plurality of sets of graphics or compute execution units (e.g., the GPGPU cores 262A through 262D, the ray tracing cores 263A through 263D, and the tensor cores 264A through 264D) and a plurality of sets of load / store units 266A through 266D. The execution resources 286A to 286D may cooperate with one or more texture units 280A to 280D for texture operations while sharing an instruction cache 252 and a shared memory 270 and cache memories 272A, 272B. In one embodiment, the execution resources 286A to 286D additionally include multifunction units (MUFU) and / or special function units (SFU) (e.g.,the MFU 267A to 267D), which are used to perform specialized mathematical operations, such as transcendental operations including exponential, logarithmic and trigonometric functions.

[0047] The various components may communicate via an interconnect fabric 290. The interconnect fabric 290 may include one or more crossbar switches to enable communication between the various components of the graphics multiprocessor 235. The GPGPU cores 262A to 262D, the ray tracing cores 263A, 263B, and the tensor cores 264A to 264D may each communicate with each other using the shared memory 270 via the interconnect fabric 290. The interconnect fabric 290 may mediate communication within the graphics multiprocessor 235 to ensure fair bandwidth allocation between the components. In one embodiment, the interconnect fabric 290 is a separate high-speed network fabric layer upon which each component of the graphics multiprocessor 235 is stacked. The components of the graphics multiprocessor 235 can also communicate with remote components via the interconnect fabric 290.

[0048] In one embodiment, the graphics multiprocessor 235 includes a tensor transfer engine 292, which is a copy engine configurable to accelerate the movement of tensor data into and out of the graphics multiprocessor 235. The tensor transfer engine 292 can accelerate tensor memory operations by asynchronously performing address generation and data movement operations for N-dimensional blocks of tensor data, offloading operations that would otherwise be performed manually by program code executed by the graphics multiprocessor 235. The tensor transfer engine 292 is configurable to copy data, for example, between the shared memory 270 and / or cache memory 272A, 272B of the graphics multiprocessor 235 and memory external to the graphics multiprocessor 235, such as the global graphics processor memory (e.g., the parallel processor memory 222).In one embodiment, data transfers performed by the tensor transfer engine 292 may be configured to selectively bypass various levels of intermediate data storage between the source and destination memory. For example, a transfer between global memory and shared memory 270 may skip register files 258A through 258D. In one embodiment, threads during asynchronous tensor transfers may be synchronized via a non-blocking barrier synchronization mechanism.

[0049] In various embodiments, the graphics multiprocessor 235 may be tailored for specific use cases through the inclusion or exclusion of certain components, allowing for various implementations of the graphics multiprocessor 235 tailored to target performance, capability, and range characteristics. For example, compute-oriented variants of the graphics multiprocessor 235 that will not perform graphics operations may exclude the ray tracing cores 263A through 263D. Fully graphics-oriented variants may exclude the tensor transfer engine 292, while graphics-oriented variants additionally configured to accelerate neural network inference may include at least one version of the tensor transfer engine 292.

[0050] Experts will understand that the Fig. 1 and Fig.2A-2E are descriptive and do not limit the scope of the present embodiments. Thus, the techniques described herein may be implemented in any properly configured processing device, including, but not limited to, one or more mobile application processors, one or more desktop or server central processing units (CPUs) including multi-core CPUs, one or more parallel processing units, such as the parallel processing unit 202 of Fig. 2A, as well as one or more graphics processors or dedicated processing units, without departing from the scope of the embodiments described herein.

[0051] The parallel processor or GPGPU, as described herein, may be communicatively coupled to host processor / cores to accelerate graphics operations, machine learning operations, pattern analysis operations, and various general-purpose GPU (GPGPU) functions. The GPU may be communicatively coupled to the host processor / cores via a bus or other interconnect (e.g., a high-speed interconnect such as PCIe or NVLink, or other known protocols, standardized protocols, or proprietary protocols). In other embodiments, the GPU may be integrated into the same package or chip as the cores and communicatively coupled to the cores via an internal processor bus / interconnect (i.e., within the package or chip).Regardless of how the GPU is connected, the processor cores can assign work to the GPU in the form of sequences of commands / instructions contained in a work descriptor. The GPU then uses dedicated circuitry / logic to efficiently process these commands / instructions.

[0052] Fig. Figure 3 illustrates a graphics processing unit (GPU 380) that includes dedicated sets of graphics processing resources arranged in multi-core groups 365A-365N. The multi-core groups 365A-365N correspond to the graphics multiprocessor 234 of Fig. 2D or the graphics multiprocessor 235 Fig.2E. Although the details of a single example of multi-core groups 365A through 365N (e.g., multi-core group 365A) are provided, it should be understood that the other multi-core groups 365B through 365N may be equipped with the same or similar sets of graphics processing resources. The details described with respect to multi-core groups 365A through 365N may also be applied to graphics multiprocessor 234 or graphics multiprocessor 235, as described herein.

[0053] As illustrated, a multi-core group 365A may include graphics cores 370, tensor cores 371, and ray tracing cores 372. The graphics cores 370 are analogous to the GPGPU cores 262A through 262D and are configurable to execute instructions to perform graphics and / or general-purpose computing operations. A scheduler / dispatcher 368 schedules and dispatches the graphics threads for execution on the various cores in the multi-core group 365A. Register files 369 are included that store operand values ​​used by the cores when performing graphics or general-purpose computing operations for executing threads. These register files 369 may, for example, contain registers configurable to store integer values ​​or floating point values, and they may contain vector registers for storing packed integer and / or floating point data elements and tile registers for storing tensor / matrix values.The tile registers can be implemented as multidimensional registers containing combined sets of vector registers.

[0054] One or more combined Level 1 (L1) caches and shared memory units 373 store graphics data, such as texture data, vertex data, pixel data, ray data, bounding volume data, etc., locally within each multi-core group 365A. One or more texture units 374 may also be used to perform texturing operations, such as texture mapping and sampling. A Level 2 (L2) cache 375, shared by all or a subset of the multi-core groups 365A through 365N, stores graphics data and / or instructions for multiple concurrent graphics threads. As illustrated, the L2 cache 375 may be shared across a plurality of multi-core groups 365A through 365N. One or more memory controllers 367 couple the GPU 380 to a memory 366, which may be system memory (e.g., DRAM) and / or dedicated graphics memory (e.g., GDDR6 memory).

[0055] Input / output circuitry (I / O circuitry 363) couples the GPU 380 to one or more I / O devices 362, such as digital signal processors (DSPs), network controllers, or user input devices. An on-chip interconnect may be used to couple the I / O devices 362 to the GPU 380 and the memory 366. One or more I / O memory management units (IOMMLTs 364) of the I / O circuitry 363 couple the I / O devices 362 directly to the memory 366. The IOMMU 364 optionally manages multiple sets of page tables to map virtual addresses to physical addresses in the system memory 366. The I / O devices 362, the one or more CPUs 361, and the GPU 380 may then share the same virtual address space.

[0056] In one implementation of the IOMMU 364, the IOMMU 364 supports virtualization. In this case, it can manage a first set of page tables to map virtual guest / graphics addresses to physical guest / graphics addresses and a second set of page tables to map the physical guest / graphics addresses to physical system / host addresses (e.g., within memory 366). The base addresses of both the first and second sets of page tables can be stored in control registers and swapped out upon a context switch (e.g., so that the new context is provided with access to the relevant set of page tables). Although this is Fig.3, each of the cores within multi-core groups 365A through 365N may include address translation buffers (TLBs) to cache guest virtual to physical address translations, guest physical to physical address translations, and guest virtual to physical address translations.

[0057] The CPU(s) 361, the GPU 380, and the I / O devices 362 may be integrated on a single semiconductor chip and / or chip package. The memory 366 may be integrated into the same chip or coupled to the one or more memory controllers 367 via an off-chip interface. In one implementation, the memory 366 comprises GDDR6 memory that shares the same virtual address space with other physical system-level memories, although the underlying principles described herein are not limited to this particular implementation.

[0058] The tensor cores 371 may include a variety of execution units specifically designed to perform matrix operations, which are the computational operations used to perform deep learning operations. For example, concurrent matrix multiplication operations may be used for training and inference of neural networks. The tensor cores 371 may perform matrix processing using a variety of operand precisions, including single-precision floating point (e.g., 32 bits), half-precision floating point (e.g., 16 bits), integer words (16 bits), bytes (8 bits), and nibbles (4 bits). For example, a neural network implementation extracts features from each rendered scene, possibly combining details from multiple frames (single images) to create a high-quality final image.

[0059] In deep learning implementations, parallel matrix multiplication work can be scheduled for execution on the tensor cores 371. In particular, training neural networks utilizes a significant number of matrix dot product operations. To process an inner product formulation of an N x N x N matrix multiplication, the tensor cores 371 can include at least N dot product processing elements. Before matrix multiplication begins, an entire matrix is ​​loaded into tile registers, and at least one column of a second matrix is ​​loaded into each cycle for N cycles. In each cycle, N dot products are processed.

[0060] Matrix elements can be stored with different precisions depending on the specific implementation, including 16-bit words, 8-bit bytes (e.g., INT8), and 4-bit nibbles (e.g., INT4). Different precision modes can be specified for the tensor cores 371 to ensure that the most efficient precision is used for different workloads (e.g., for inference workloads that can tolerate quantization into bytes and nibbles). Supported formats additionally include 64-bit floating-point (FP64) and non-IEEE floating-point formats, such as the bfloat16 format. One embodiment includes support for a reduced-precision tensor floating-point mode (TF32) that performs computations using the range of FP32 (8 bits) and the precision of FP16 (10 bits).Reduced-precision TF32 operations can be performed on FP32 inputs and produce FP32 outputs with higher performance relative to FP32 and greater precision relative to FP16. In one embodiment, one or more of 8-bit floating-point (FP8), 6-bit floating-point (FP6), and 4-bit floating-point (FP4) formats are supported, including floating-point formats represented as microscale (MX) formats.

[0061] In one embodiment, tensor cores 371 support a sparse mode of operation for matrices in which the vast majority of values ​​are zero. Tensor cores 371 include support for sparse input matrices encoded in a sparse matrix representation (e.g., coordinate list encoding (COO), compressed sparse row (CSR), compressed sparse column (CSC), etc.). Tensor cores 371 also include support for compressed sparse matrix representations in case the sparse matrix representation needs to be further compressed. Compressed, encoded, and / or compressed and encoded matrix data, along with associated compression and / or encoding metadata, may be read by tensor cores 371, and the non-zero values ​​may be extracted.For example, for a given input matrix A, a non-zero value may be loaded from the compressed and / or encoded representation of at least a portion of matrix A. Based on the position in matrix A for the non-zero value, which may be determined from index or coordinate metadata associated with the non-zero value, a corresponding value may be loaded into input matrix B. Depending on the operation to be performed (e.g., multiply), loading the value from input matrix B may be skipped if the corresponding value is a zero value. In one embodiment, the value pairs for certain operations, such as multiply operations, may be pre-scanned by scheduling logic, and operations between non-zero inputs are scheduled.Depending on the dimensions of matrix A and matrix B and the operation to be performed, the output matrix C may be dense or sparse. If the output matrix C is sparse, and depending on the configuration of the tensor kernels 371, the output matrix C may be output in a compressed format, a sparse encoding, or a compressed sparse encoding.

[0062] Ray tracing cores 372 may accelerate ray tracing operations for both real-time and non-real-time ray tracing implementations. In particular, ray tracing cores 372 may include ray tracing or ray intersection circuitry for performing ray tracing using bounding body hierarchies (BVHs) and for identifying intersections between rays and primitives enclosed within the BVH volumes. Ray tracing cores 372 may also include circuitry for performing depth checks and depth suppression (e.g., using a Z-buffer or similar arrangement). In one implementation, the ray tracing cores 372 perform traversal and intersection operations along with the image denoising techniques described herein, at least a portion of which may be performed on the tensor cores 371.For example, the tensor cores 371 implement a deep learning neural network to denoise images generated by the ray tracing cores 372. However, the one or more CPUs 361, graphics cores 370, and / or ray tracing cores 372 may also implement all or part of the denoising and / or deep learning algorithms.

[0063] Additionally, as described above, a distributed denoising approach may be employed, where the GPU 380 is located in a computing device coupled to other computing devices via a network or high-speed interconnect. In this distributed approach, the interconnected computing devices may share the neural network learning / training data to increase the speed at which the overall system learns to perform denoising for different types of image frames and / or different graphics applications.

[0064] Ray tracing cores 372 can handle all BVH traversals and / or ray primitive intersections, preventing graphics cores 370 from being overloaded with thousands of instructions per ray. For example, each ray tracing core 372 includes a first set of specialized circuitry for performing bounding box tests (e.g., for traversal operations) and / or a second set of specialized circuitry for performing ray triangle intersection tests (e.g., for intersecting rays that have been traversed). Thus, for example, multi-core group 365A can simply launch a ray probe, and ray tracing cores 372 independently perform ray traversal and ray intersection and return hit data (e.g., a hit, no hit, multiple hits, etc.) to the thread context.The graphics cores 370 and the tensor cores 371 are then free for other graphics or computational tasks, while the ray tracing cores 372 perform the traversal and intersection operations. Optionally, each ray tracing core 372 can include a traversal unit for performing BVH checks and / or an intersection unit that performs ray primitive intersection tests. The intersection unit generates a "hit," "no hit," or "multi-hit" response, which it provides to the appropriate thread. During the traversal and intersection operations, the execution resources of the other cores (e.g., the graphics cores 370 and tensor cores 371) are free to perform other forms of graphics work. In an optional embodiment described below, a hybrid rasterization / ray tracing approach is used in which rendering operations are distributed between the graphics cores 370 and the ray tracing cores 372.

[0065] Ray tracing cores 372 may include hardware support for a ray tracing instruction set, such as Microsoft's DirectX Ray Tracing (DXR), which includes a DispatchRays instruction, as well as ray generation, closest-hit, any-hit, and miss shaders, allowing the allocation of sets of shaders and textures for each object. Another ray tracing platform that may be supported by ray tracing cores 372, graphics cores 370, and tensor cores 371 is the Vulkan API (e.g., Vulkan version 1.1.85 and later). However, it should be noted that the underlying principles described herein are not limited to any particular ray tracing ISA.In general, ray tracing cores 372, tensor cores 371, and graphics cores 370 may support a ray tracing instruction set that includes instructions / functions for one or more of ray generation, nearest hit, any hit, ray primitive intersection, per-primitive hierarchical bounding box construction, no hit, visit, and exceptions. More specifically, one embodiment includes ray tracing instructions for performing one or more of the following functions:

[0066] Ray tracing - Ray generation instructions can be executed for each pixel, sample, or any other custom work task.

[0067] Closest hit - A closest hit statement can be executed to locate the closest intersection of a ray with primitives within a scene.

[0068] Any hit - An any-hit command identifies multiple intersections between a ray and primitives within a scene to potentially identify a new nearest intersection point.

[0069] Intersect - An Intersect statement performs a ray primitive intersection test and returns a result.

[0070] Per-Primitive Bounding Box Construction - This instruction creates a bounding box around a specific primitive or group of primitives (e.g., when constructing a new BVH or other acceleration data structure).

[0071] Miss - Indicates that a ray misses all geometry within a scene or a specific region of a scene.

[0072] Visit - specifies the follow-up volumes a beam will traverse.

[0073] Exceptions - contain different types of exception handlers (e.g. called for different error conditions).

[0074] In one embodiment, ray tracing cores 372 may be configured to accelerate general-purpose computational operations using computational techniques analogous to ray-intersection testing. A computational framework may be provided that enables shader programs to be compiled into low-level instructions and / or primitives that perform general-purpose computational operations via the ray tracing cores. Example computational problems that may benefit from computational operations performed on ray tracing cores 372 include computations involving ray, wave, beam, or particle propagation within a coordinate space. Interactions associated with this propagation may be computed relative to a geometry or mesh within the coordinate space.For example, calculations related to the propagation of electromagnetic signals in an environment can be accelerated by using instructions or primitives executed via the ray tracing kernels. Diffraction and reflection of the signals from objects in the environment can be calculated as direct ray tracing analogies.

[0075] Ray tracing cores 372 can also be used to perform computations that are not directly analogous to ray tracing. For example, mesh projection, mesh refinement, and volume sampling computations can be accelerated using ray tracing cores 372. Generic coordinate space computations, such as nearest neighbor computations, can also be performed. For example, the set of points near a given point can be discovered by defining a bounding box in coordinate space around the point. BVH and ray probe logic within ray tracing cores 372 can then be used to determine the set of point intersections within the bounding box. The intersections form the origin point and the nearest neighbors of that origin point. Computations performed using ray tracing cores 372 can be performed in parallel with computations on graphics cores 370 and tensor cores 371.A shader compiler can be configured to compile a compute shader or other general-purpose graphics processing program into low-level primitives that can be parallelized across the graphics cores 370, the tensor cores 371, and the ray tracing cores 372. Techniques for connecting a GPU to a host processor

[0076] Fig. Figure 4A illustrates an example architecture in which multiple GPUs 410-413, e.g., such as the one shown in Fig.2A, are communicatively coupled to multiple multi-core processors 405-406 via high-speed links 440A-440D (e.g., buses, point-to-point interconnects, etc.). The high-speed links 440A-440D may support communication throughput of 4 GB / s, 30 GB / s, 80 GB / s, or higher, depending on the implementation. Various interconnect protocols may be used, including, but not limited to, PCIe 4.0, PCIe 5.0, PCIe 6.0, and various NVLink and NVlink C2C (chip-to-chip interconnect) protocols (e.g., NVLink v5). However, the underlying principles described herein are not limited to any particular communication protocol or data throughput.

[0077] Two or more of the GPUs 410-413 may be interconnected via high-speed links 442A-442B, which may be implemented using the same or different protocols / links than those used for the high-speed links 440A-440D. Similarly, two or more of the multi-core processors 405, 406 may be interconnected via high-speed link 443, which may be symmetric multiprocessor (SMP) buses operating at 20 GB / s, 30 GB / s, 120 GB / s, or lower or higher speeds. Alternatively, all communication between the various Fig. 4A using the same protocols / connections (e.g., via a common interconnect fabric). However, as mentioned, the underlying principles described here are not limited to any specific type of interconnect technology.

[0078] Each of multi-core processor 405 and multi-core processor 406 may be communicatively coupled to a processor memory 401-402 via memory interconnects 430A-430B, respectively, and each GPU 410-413 is communicatively coupled to a GPU memory 420-423 via GPU memory interconnects 450A-450D, respectively. Memory interconnects 430A, 430B, and 450A-450D may utilize the same or different memory access technologies. By way of example and not limitation, the processor memories 401, 402 and the GPU memories 420 to 423 may be volatile memories, such as dynamic random access memories (DRAMs) (including stacked DRAMs), graphics DDR SDRAM (GDDR) (e.g., GDDR5, GDDR6), or high bandwidth memory (HBM), and / or they may be non-volatile memories, such as 3D XPoint / Optane or Nano-RAM.For example, part of the memory may be volatile memory and another part may be non-volatile memory (e.g., using a two-level memory hierarchy (2LM hierarchy)). A memory subsystem as described here may be compatible with a variety of memory technologies, such as double-data-rate versions published by JEDEC (Joint Electronic Device Engineering Council).

[0079] As described below, although the various processors 405-406 and GPUs 410-413 may each be physically coupled to a specific processor memory 401-402 and GPU memory 420-423, a unified memory architecture may be implemented in which the same virtual system address space (also referred to as the "effective address space") is distributed among all of the various physical memories. For example, processor memories 401, 402 may each comprise 64 GB of system memory address space, and GPU memories 420-423 may each comprise 32 GB of system memory address space (resulting in a total of 256 GB of addressable memory in this example).

[0080] Fig.4B illustrates additional optional details for a connection between a processor 407 and a graphics accelerator 446. The graphics accelerator 446 may include one or more GPU chips integrated on a line card coupled to the processor 407 via the high-speed interconnect 440. Alternatively, the graphics processor 446 may be integrated in the same package or chip as the processor 407. The processor 407 includes a plurality of cores 460A through 460D, each having a translation buffer (TLB) 461A through 461D and one or more caches 462A through 462D. The cores may include numerous other components for executing instructions and processing data, which are not illustrated to avoid obscuring the underlying principles of the components described herein (e.g.,Instruction fetch units, branch prediction units, decoders, execution units, reorder buffers, etc.) may be obfuscated. Caches 462A through 462D may include Level 1 (L1) caches and Level 2 (L2) caches. Additionally, one or more shared caches 456 may be included in the caching hierarchy and shared by sets of cores 460A through 460D. For example, one embodiment of processor 407 includes 24 cores, each with its own L1 cache, twelve shared L2 caches, and twelve shared L3 caches. In this embodiment, one of the L2 and L3 caches is shared between any two adjacent cores. Processor 407 is coupled to system memory 441, which may include processor memories 401, 402.

[0081] Coherency is maintained for data and instructions stored in the various caches 462A-462D, the one or more shared caches 456, and system memory 441 via inter-core communication over a coherency bus 464. For example, each cache may have dedicated cache coherency logic / circuitry associated with it to communicate with it over the coherency bus 464 in response to detected reads or writes to particular cache lines. In one implementation, a cache sniffing protocol is implemented over the coherency bus 464 to sniff cache accesses. Cache sniffing / coherency techniques are well known to those skilled in the art and are not described in detail here to avoid obscuring the underlying principles described herein.A proxy circuit 425 may be provided that communicatively couples the graphics accelerator 446 to the coherence bus 464, thereby enabling the graphics accelerator 446 to participate in the cache coherence protocol as a peer of the cores. Specifically, an interface 435 of the proxy circuit 425 provides connectivity via a high-speed interconnect 440 (e.g., a PCIe bus, NVLink, etc.), and an interface 437 connects the graphics accelerator 446 to the high-speed interconnect 440.

[0082] In one implementation, interface 437 is coupled to an accelerator integration circuit 436 that provides cache management, memory access, context management, and interrupt management services for graphics processing engines 431, 432, ..., N of graphics accelerator 446. Graphics processing engines 431, 432, ..., N may each comprise a separate graphics processing unit (GPU). Alternatively, graphics processing engines 431, 432, ..., N may comprise different types of graphics processing engines within a GPU, such as graphics execution units, media processing engines (e.g., video encoders / decoders), sampling units, and block image transfer (BLIT) engines. In other words, the graphics accelerator may comprise graphics processing engines 431, 432, ..., N of a single GPU, or the graphics processing engines 431, 432, ..., N may be associated with multiple GPUs integrated in a common package, on a common line card, or on a common chip. The graphics processing engines 431-432, N may be configured with any graphics processor or compute accelerator architecture described herein. The functions to be performed by the graphics processing engines 431, 432 may be specified via work descriptors that provide an indication of the functions to be performed by the graphics accelerator 446.

[0083] Accelerator integration circuitry 436 may include a memory management unit (MMU) 439 for performing various memory management functions such as virtual-to-physical memory translations (also referred to as effective-to-real memory translations) and memory access protocols for accessing system memory 441. MMU 439 may also include an address translation buffer (TLB) (not shown) for caching virtual / effective to physical / real address translations. In one implementation, a cache 438 stores instructions and data for efficient access by graphics processing engines 431, 432, ..., N. The data stored in cache 438 and graphics memories 433-434, ..., M may be kept coherent with core caches 462A-462D, shared cache(s) 456, and system memory 441.As mentioned, this may be achieved via a proxy circuit 425 that participates in the cache coherence mechanism on behalf of the cache 438 and the graphics memories 433, 434, ..., M (e.g., by sending updates to the cache 438 related to modifications / accesses of cache lines in the processor caches 462A to 462D and the shared caches 456, and receiving updates from the cache 438).

[0084] Registers 445 store context data for threads executed by the graphics processing engines 431-432, ..., N, and a context management circuit 448 manages the thread contexts. For example, the context management circuit 448 may perform save and restore operations to save and restore contexts of the various threads during context switches (e.g., when a first thread is saved and a second thread is restored so that the second thread can be executed by a graphics processing engine). For example, during a context switch, the context management circuit 448 may store current register values ​​in a designated area in memory (identified, e.g., by a context pointer). It may then restore the register values ​​when returning to the context.For example, an interrupt management circuit 447 may receive and process interrupts received from system devices.

[0085] In one implementation, virtual / effective addresses from a graphics processing engine are translated by the MMU 439 into real / physical addresses in system memory 441. Optionally, the accelerator integration circuit 436 supports multiple (e.g., 4, 8, 16) graphics accelerators 446 and / or other accelerator devices. The graphics accelerator module 446 may be dedicated to a single application executing in the processor 407, or it may be shared among multiple applications. Optionally, a virtualized graphics execution environment is provided in which the resources of the graphics processing engines 431, 432, ..., N are shared among multiple applications, virtual machines (VMs), or containers.The resources can be divided into "partitions," which are allocated to different VMs and / or applications based on processing requirements and priorities associated with the VMs and / or applications, or based on a predetermined partitioning profile for a graphics accelerator 446. VMs and containers can be used interchangeably here.

[0086] A virtual machine (VM) can be software that runs an operating system and one or more applications. A VM can be defined by a specification, configuration files, a virtual disk file, a non-volatile random access memory (NVRAM) settings file, and the log file, and is secured by the physical resources of a host computing platform. A VM can include an operating system (OS) or an application environment installed as software that mimics dedicated hardware. The end user has the same experience on a virtual machine as they would with dedicated hardware. Specialized software called a hypervisor fully emulates the CPU, memory, disk, network, and other hardware resources of the PC client or server, allowing virtual machines to share resources.The hypervisor can emulate multiple virtual hardware platforms that are isolated from each other, allowing virtual machines to run Linux®, Windows® Server, VMware ESXi, and other operating systems on the same underlying physical host.

[0087] A container can be a software package of applications, configurations, and dependencies, allowing the applications to run reliably from one computing environment to another. The containers can share an operating system installed on the server platform and run as isolated processes. A container can be a software package that contains everything the software uses to run, such as system tools, libraries, and settings. The containers are not installed like traditional software programs, which allows them to be isolated from other software and the actual operating system. The isolated nature of containers provides several advantages. First, the software in a container will run the container in different environments. For example, a container containing PHP and MySQL can run identically on both a Linux® computer and a Windows® machine.Second, containers provide additional security because the software will not interfere with the host operating system. While an installed application can change system settings and modify resources, such as the Windows Registry, a container can modify settings within the container.

[0088] Thus, accelerator integration circuitry 436 acts as a bridge to the system for graphics accelerator 446, providing address translation and system memory cache services. In one embodiment, to facilitate bridging functionality, accelerator integration circuitry 436 may also include shared I / O 497 (e.g., PCIe, USB, or other) and hardware to enable system control of voltage, clock, performance, thermal, and security. Shared I / O 497 may use separate physical connections or may traverse high-speed interconnect 440. Accelerator integration circuitry 436 may also provide virtualization facilities for the host processor to manage virtualization of the graphics processing engines, interrupts, and memory management.

[0089] Because hardware resources of graphics processing engines 431-432, N are explicitly mapped to the real address space seen by processor 407, each host processor can address these resources directly using an effective address value. An optional feature of accelerator integration circuitry 436 is the physical separation of graphics processing engines 431, 432, ..., N so that they appear to the system as independent entities. In one embodiment, accelerator integration circuitry 436 includes security circuitry 444 that enables configurable cryptographic isolation of data associated with each partition of resources. Different partitions may be associated with different security domains, so that data associated with the different security domains is encrypted using different cryptographic keys.In one embodiment, the security domains of graphics accelerator 446 may be integrated into trusted execution environments supported by processor 407. In one embodiment, secure I / O capabilities may be enabled that allow each security domain to be presented as a separate trusted I / O device, enabling support for trusted DMA and MMIO operations.

[0090] One or more graphics memories 433 - 434, ..., M may be coupled to each of the graphics processing engines 431 - 432, ..., N, respectively. The graphics memories 433, 434, ..., M store instructions and data processed by each of the graphics processing engines 431, 432, ..., N. The graphics memories 433, 434, ..., M may be volatile memories, such as DRAMs (including stacked DRAMs), GDDR memory (e.g., GDDR5, GDDR6), or HBM, and / or may be non-volatile memories, such as 3D XPoint / Optane, Samsung Z-NAND, or Nano-RAM.

[0091] To reduce data traffic over the high-speed link 440, biasing techniques can be used to ensure that the data stored in the graphics memories 433-434, ..., M is data most frequently used by the graphics processing engines 431-432, ..., N, and preferably not used by the cores 460A-460D (at least not frequently). Similarly, the biasing mechanism attempts to keep data used by the cores (and preferably not by the graphics processing engines 431-432, ..., N) in the caches 462A-462D, the one or more shared caches 456, and the system memory 441.

[0092] In an alternative variant, the accelerator integration circuit 436 is integrated within the processor 407, and the graphics processing engines 431-432, ..., N communicate with the accelerator integration circuit 436 via the high-speed link 440 via the interface 437 and the interface 435 (which may again utilize any form of bus, fabric, or interface protocol). In this variant, the accelerator integration circuit 436 performs the same operations as those described above.

[0093] The described embodiments may support various programming models, including a dedicated process programming model (no graphics accelerator virtualization) and shared programming models (with virtualization). The latter may include programming models controlled by the accelerator integration circuit 436 and programming models controlled by the graphics accelerator 446. In the dedicated process model embodiments, the graphics processing engines 431, 432, ... N may be dedicated to a single application or process under a single operating system. The single application may route other application requests to the graphics engines 431, 432, ... N and provide virtualization within a VM / partition. In the dedicated process programming models, the graphics processing engines 431, 432, ..., N are shared by multiple VM / application partitions. The shared models utilize a system hypervisor to virtualize the graphics processing engines 431-432, N to enable access by any operating system. For single-partition systems without a hypervisor, the graphics processing engines 431, 432, ..., N are owned by the operating system. In both cases, the operating system may virtualize the graphics processing engines 431, 432, ..., N to provide access to any process or application. For the shared programming model, the graphics accelerator 446 or a single graphics processing engine 431, 432, ..., N selects a process element using a process routine. The process elements may be stored in system memory 441 and may be addressable using the effective address to real address translation techniques described herein.The process handle may be an implementation-specific value provided to the host process when its context is registered in the graphics processing engines 431, 432, ..., N, which is performed by the host process by invoking system software to add the process element to the list of linked process elements. The lower 16 bits of the process routine may be the offset of the process element within the list of linked process elements.

[0094] Fig.4C illustrates an accelerator integration slice 490. As used herein, "partition" comprises a specific portion of the processing resources of the accelerator integration circuit 436. Process elements 483 are stored in the application address space 482 within the system memory 441. The process elements 483 may be stored in response to GPU calls 481 from an application 480 executing in the processor 407. A process element 483 contains the process state for the application 480. A work descriptor (WD) 484 contained in the process element 483 may be a single work job requested by an application or it may contain a pointer to a queue of work jobs. In the latter case, the WD 484 is a pointer to the work job request queue in the application address space 482.

[0095] The graphics accelerator 446 and / or the individual graphics processing engines 431-432, ..., N may be shared by all or a subset of the processes in the system. For example, the technologies described herein may include an infrastructure for establishing process state and sending a WD 484 to a graphics accelerator 446 to start a job in a virtualized environment.

[0096] In one implementation, the dedicated process programming model is implementation-specific. In this model, the graphics accelerator 446 or an individual graphics processing engine 431 is owned by a single process. Because a single process owns the graphics accelerator 446, the hypervisor initializes the accelerator integration circuit 436 for the owning partition, and the operating system initializes the accelerator integration circuit 436 for the owning process at the time the graphics accelerator 446 is allocated.

[0097] In operation, the fetch unit 491 in the accelerator integration slice 490 fetches a WD 484 to be processed. The WD 484 includes the indication of functions to be performed by one or more graphics processing engines of the graphics accelerator 446. Data from the WD 484 may be stored in registers 445 and used by the MMU 439, the interrupt management circuitry 447, and / or the context management circuitry 448 as illustrated. For example, the MMU 439 may include segment / page browsing circuitry for accessing segment / page tables 486 within the OS virtual address space 485. The interrupt management circuitry 447 may process interrupt events 492 received from the graphics accelerator 446. When performing graphics operations, an effective address 493 generated by a graphics processing engine 431-432, ..., N is translated into a real address by the MMU 439.

[0098] Registers 445 may be duplicated for each graphics processing engine 431-432, ..., N and / or the graphics accelerator 446 and may be initialized by the hypervisor or the operating system. Each of these duplicated registers may be included in an accelerator integration partition 490. In one embodiment, each graphics processing engine 431, 432, ..., N may be presented to the hypervisor 496 as a separate graphics processor device. Quality of service (QoS) settings may be configured for clients of a specific graphics processing engine 431, 432, ..., N. Cryptographic and physical data isolation between the clients of each engine may be enabled via isolated memory access paths and automatic data encryption for each client. Example registers that may be initialized by the hypervisor are shown in Table 1. Table 1 - Registers initialized by the hypervisor 1 Partition control register 2 Area pointer for scheduled processes with respect to real addresses (RA) 3 Overwrite register for authorization masks 4 Offset of table entries of interrupt vectors 5 Limit for table entries of interrupt vectors 6 Condition register 7 ID Logical Partition 8 Record pointer of hypervisor accelerator usages with respect to real addresses (RA) 9 Memory descriptor register

[0099] Example registers that can be initialized by the operating system are shown in Table 2. Table 2 - Registers initialized by an operating system 1 Process and thread identification 2 Effective Address (EA) Context Save / Restore Pointer 3 Record pointer of accelerator usages with respect to virtual addresses (VA) 4 Pointers to memory segment tables regarding virtual addresses (VA) 5 Authorization mask 6 Work descriptor

[0100] Each WD 484 may be specific to a particular graphics accelerator 446 and / or a graphics processing engine 431-432, ..., N. It contains all the information a graphics processing engine uses to do its work, or it may be a pointer to a memory location where the application has set up a command queue with work to be performed.

[0101] Fig.4D illustrates additional optional details of a shared model. This includes a real hypervisor address space 498 in which a process element list 499 is stored. The real hypervisor address space 498 is accessible via a hypervisor 496 that virtualizes the graphics accelerator engines for the operating system 495.

[0102] The shared programming models allow all or a subset of processes from all or a subset of partitions in the system to use a graphics accelerator 446. There are two programming models in which the graphics accelerator 446 is shared among multiple processes and partitions: time-split sharing and graphics-directed sharing.

[0103] In this model, the hypervisor 496 hosts the graphics accelerator 446 and makes its functionality available to all operating systems 495. For a graphics accelerator 446 to support virtualization through the hypervisor 496, the graphics accelerator 446 can meet the following requirements: 1) An application's job request should be autonomous (i.e., state does not need to be maintained between jobs), or the graphics accelerator 446 should provide a context save and restore mechanism. 2) An application's job request is guaranteed by the graphics accelerator 446 to complete in a specified amount of time, including any compilation errors, or the graphics accelerator 446 provides the ability to interrupt task processing.3) The graphics accelerator 446 must be guaranteed fairness between processes when operating in the directed shared programming model.

[0104] For the directed shared model, the application 480 may need to make a system call to the operating system 495 with a graphics accelerator 446 type, a work descriptor (WD), an Authority Mask Register (AMR) value, and a Context Save / Restore Area Pointer (CSRP). The graphics accelerator 446 type describes the targeted acceleration function for the system call. The graphics accelerator 446 type may be a system-specific value. The WD is formatted specifically for the graphics accelerator 446 and may be in the form of a graphics accelerator 446 instruction, an effective address pointer in a user-defined structure, an effective address pointer in a queue of instructions, or any other data structure for describing the functions to be performed by the graphics accelerator 446.In one embodiment, the AMR value is the AMR state to be used for the current process. The value passed to the operating system is similar to an application setting the AMR. If the implementations of the accelerator integration circuit 436 and the graphics accelerator 446 do not support a User Authority Mask Override Register (UAMOR), the operating system may apply the current UAMOR value to the AMR value before passing the AMR in the hypervisor call. The hypervisor 496 may optionally apply the current value in the Authority Mask Override Register (AMOR) before placing the AMR in the process element 483. The CSRP may be one of the registers 445 containing the effective address of a region in the application's address space 482 for the graphics accelerator 446 to save and restore context state.This pointer is optional if no status should be saved between work orders or when a work order is deferred. The context save / restore area can be a specified system memory location.

[0105] Upon receiving the system call, the operating system 495 can verify that the application 480 has been registered and has been granted permission to use the graphics accelerator 446. The operating system 495 then calls the hypervisor 496 with the information shown in Table 3. Table 3 - Call parameters from the OS to the hypervisor 1 A work descriptor (WD) 2 An authorization mask register (AMR) value (potentially masked). 3 A context save / restore range pointer (CSRP) with respect to effective addresses (EA) 4 A process ID (PID) and optional thread ID (TID) 5 An accelerator usage record pointer (AURP) with respect to virtual addresses (VA) 6 The virtual address of the memory segment table pointer (SSTP) 7 A logical interrupt service number (LISN)

[0106] Upon receiving the hypervisor call, the hypervisor 496 verifies that the operating system 495 has been registered and granted permission to use the graphics accelerator 446. The hypervisor 496 then inserts the process element 483 into the list of associated process elements for the corresponding type of graphics accelerator 446. The process element may include the information shown in Table 4. Table 4 - Process element information 1 A work descriptor (WD) 2 An authorization mask register (AMR) value (potentially masked). 3 A context save / restore range pointer (CSRP) with respect to effective addresses (EA) 4 A process ID (PID) and optional thread ID (TID) 5 An accelerator usage record pointer (AURP) with respect to virtual addresses (VA) 6 The virtual address of the memory segment table pointer (SSTP) 7 A logical interrupt service number (LISN) 8 Interrupt vector table derived from the hypervisor call parameters. 9 Ein Zustandsregisterwert (SR-Wert) 10 Eine logische Partitions-ID (LPID) 11 Ein Datensatzzeiger für Beschleunigernutzungen bezüglich realer Adressen(RA) 12 Das Speicherdeskriptorregister (SDR)

[0107] The hypervisor can initialize the registers 445 of the accelerator integration slice 490.

[0108] As in Fig. As illustrated in Figure 4E, in an optional implementation, a unified memory addressable via a common virtual memory address space used to access the physical processor memories 401-402 and GPU memories 420-423 is employed. In this implementation, operations performed in the GPUs 410-413 use the same virtual / effective memory address space to access the processor memories 401, 402 and vice versa, thereby simplifying programmability. A first portion of the virtual / effective memory address space may be assigned to the processor memory 401, a second portion may be assigned to the second processor memory 402, a third portion may be assigned to the GPU memory 420, etc.The total virtual / effective memory space (sometimes referred to as effective address space) can thereby be divided between each of the processor memories 401, 402 and the GPU memories 420 to 423, allowing any processor or GPU to access any physical memory with a virtual address assigned to that memory.

[0109] Bias / coherence management circuitry 494A-494E may be provided within one or more of the MMUs 439A-439E, ensuring cache coherence between the caches of the host processors (e.g., multi-core processor 405) and the GPUs 410-413, and implementing biasing techniques that specify the physical memories in which certain types of data should be stored. Although multiple instances of bias / coherence management circuitry 494A-494E may be provided in Fig. 4E, the bias / coherence circuitry may be implemented within the MMU of one or more host processors 405 and / or within the accelerator integration circuit 436.

[0110] The GPU-attached memory 420-423 can be mapped as part of the system memory and accessed using SVM (Shared Virtual Memory) technology, but without incurring the typical performance penalties associated with full system cache coherence. The ability to access the GPU-attached memory 420-423 as system memory without burdensome cache coherence overhead provides a favorable operating environment for GPU swapping. This arrangement allows the host processor software to set up operands and access computation results without the overhead of traditional I / O DMA data copies. These traditional copies include driver calls, interrupts, and memory-mapped I / O (MMIO) accesses, all of which are inefficient compared to simple memory accesses.At the same time, the ability to access the memory 420-423 attached to the GPU without cache coherence overhead can be relevant to the execution time of a paged computation. For example, in cases with significant streaming traffic when writing to the memory, the cache coherence overhead can significantly reduce the effective write bandwidth of a GPU 410-413. Operand setup efficiency, result access efficiency, and GPU computation efficiency all play a role in determining the effectiveness of GPU paged computation.

[0111] A choice between GPU bias and host processor bias can be controlled by a bias tracker data structure. For example, a bias table can be used, which can have a page-granular structure (i.e., controlled at the granularity of a memory page) comprising 1 or 2 bits per memory page attached to the GPU. The bias table can be implemented in a stolen memory area of ​​one or more GPU-attached memories 420 to 423, with or without a bias cache in the GPU 410 to 413 (e.g., to cache frequently / recently used bias table entries). Alternatively, the entire bias table can be maintained within the GPU.

[0112] In one implementation, the bias table entry associated with each access to GPU-attached memory 420-423 is accessed prior to the actual GPU memory access, causing the following operations. First, local requests from GPU 420-423 whose page is found in the GPU bias are forwarded directly to a corresponding GPU memory 410-413. Local requests from the GPU whose page is found in the host bias are forwarded to the processor (e.g., via a high-speed connection, as discussed above). Requests from host processor 405 that find the requested page in the host processor bias optionally complete the request like a normal memory read. Alternatively, requests directed to a GPU-biased page may be forwarded to GPU 410-413.The GPU can then pass the page to a host processor bias if it is not currently using the page. The bias state of a page can be changed either through a software-based mechanism, a hardware-assisted software-based mechanism, or, for a limited set of cases, a purely hardware-based mechanism. One mechanism for changing the bias state employs an API call (e.g., OpenCL), which in turn invokes the GPU's device driver, which in turn sends a message to the GPU (or queues a command descriptor) instructing it to change the bias state and, for some transitions, perform a cache flush operation on the host. The cache flush operation is used for a transition from host processor bias to GPU bias, but is not used for the reverse transition.

[0113] Cache coherence can be maintained by causing GPU-biased pages to be temporarily uncached by the host processor. To access these pages, the host processor can request access from GPU 410, which may or may not grant access immediately depending on the implementation. Thus, to reduce communication between the host processor and GPU 410, it is advantageous to ensure that GPU-biased pages are those used by the GPU but not the host processor, and vice versa. Graphics processing pipeline

[0114] Fig. 5 illustrates a graphics processing pipeline 500. A graphics multiprocessor, such as the graphics multiprocessor 234 as shown in Fig. 2D, or the graphics multiprocessor 235 of the Fig. 2E may implement the graphics processing pipeline 500. The graphics multiprocessor may be included in the parallel processing subsystems described herein, such as parallel processor 200 of the Fig. 2A, which is connected to the parallel processors 112 of the Fig. 1 and may be used instead thereof. The various parallel processing systems, as described herein, may implement the graphics processing pipeline 500 via one or more instances of the parallel processing unit (e.g., the parallel processing unit 202 of Fig. 2A). For example, a shader unit (e.g., the graphics multiprocessor 234 of the Fig. 2C) may be configured to perform the functions of one or more of a vertex processing unit 504, a tessellation control processing unit 508, a tessellation evaluation processing unit 512, a geometry processing unit 516, and a fragment / pixel processing unit 524. The functions of the data assembler 502, the primitive assemblers 506, 514, 518, the tessellation unit 510, the rasterization unit 522, and the raster operation unit 526 may also be performed by other processing engines within a processing cluster (e.g., processing cluster 214 of Fig. 2A) and a corresponding partition unit (e.g. partition unit 220A-220N from Fig. 2A). The graphics processing pipeline 500 may also be implemented using dedicated processing units for one or more functions. It is also possible for one or more sections of the graphics processing pipeline 500 to be performed by parallel processing logic within a general-purpose processor (e.g., a CPU). Optionally, one or more sections of the graphics processing pipeline 500 may be accessed via a memory interface 528, which may be an instance of the memory interface 218 of Fig. 2A, to a chip-internal memory (e.g. the parallel processor memory 222 as in Fig. 2A). The graphics processing pipeline 500 may also have a multi-core group 365A as in Fig. 3 be implemented.

[0115] The data assembler 502 is a processing unit that collects vertex data for surfaces and primitives. The data assembler 502 then outputs the vertex data, including the vertex attributes, to the vertex processing unit 504. The vertex processing unit 504 is a programmable execution unit that executes vertex shader programs, lighting, and transform vertex data specified by the vertex shader programs. The vertex processing unit 504 reads data stored in cache, local memory, or system memory for use in processing the vertex data and can be programmed to transform the vertex data from an object-based coordinate representation into a world-space coordinate space or a normalized device coordinate space.

[0116] A first instance of a primitive assembler 506 receives vertex attributes from the vertex processing unit 504. The primitive assembler 506 reads stored vertex attributes and constructs graphics primitives for processing by the tessellation control processing unit 508. The graphics primitives include triangles, line segments, points, patches, and so on, as supported by numerous application programming interfaces (APIs) of graphics processing.

[0117] The tessellation control processing unit 508 treats the input vertices as control points for a geometric patch. The control points are transformed from an input representation of the patch (e.g., the bases of the patch) into a representation suitable for use in surface evaluation by the tessellation evaluation processing unit 512. The tessellation control processing unit 508 may also calculate tessellation factors for edges of geometric patches. A tessellation factor applies to a single edge and quantifies a view-dependent level of detail associated with the edge. A tessellation unit 510 is configured to receive tessellation factors for edges of a patch and to tessellate the patch into a plurality of geometric primitives, such as a line, a triangle, or quadrilateral primitives, which are transmitted to a tessellation evaluation processing unit 512.The tessellation evaluation processing unit 512 operates on parameterized coordinates of the subdivided patch to generate a surface representation and vertex attributes for each vertex associated with the geometric primitives.

[0118] A second instance of a primitive assembler 514 receives vertex attributes from the tessellation evaluation processing unit 512, reads stored vertex attributes, and constructs graphics primitives for processing by the geometry processing unit 516. The geometry processing unit 516 is a programmable execution unit that executes geometry shader programs to transform graphics primitives received from the primitive assembler 514 as specified by the geometry shader programs. The geometry processing unit 516 may be programmed to partition the graphics primitives into one or more new graphics primitives and calculate parameters used to rasterize the new graphics primitives.

[0119] The geometry processing unit 516 may be capable of adding or deleting elements in the geometry stream. The geometry processing unit 516 outputs the parameters and vertices that specify new graphics primitives to the primitive assembler 518. The primitive assembler 518 receives the parameters and vertices from the geometry processing unit 516 and constructs graphics primitives for processing by a viewport scaling, suppression, and clipping unit 520. The geometry processing unit 516 reads data stored in the parallel processor memory or the system memory for use in processing the geometry data. The viewport scaling, suppression, and clipping unit 520 performs clipping, suppression, and scaling of the viewport and outputs processed graphics primitives to a rasterization unit 522.

[0120] Rasterizer 522 can perform depth culling and other depth-based optimizations. Rasterizer 522 also performs scan conversion on the new graphics primitives to generate fragments and outputs these fragments and associated coverage data to fragment / pixel processing unit 524. Fragment / pixel processing unit 524 is a programmable execution unit configured to execute fragment shader programs or pixel shader programs. Fragment / pixel processing unit 524 transforms fragments or pixels received from rasterizer 522 as specified by the fragment or pixel shader programs.For example, fragment / pixel processing unit 524 may be programmed to perform operations including, but not limited to, texture mapping, shading, blending, texture correction, and perspective correction to generate shaded fragments or pixels that are output to a raster operations unit 526. Fragment / pixel processing unit 524 may read data stored in either parallel processor memory or system memory for use when processing the fragment data. Fragment or pixel shader programs may be configured to shade samples, pixels, tiles, or other granularities depending on the sampling rate configured for the processing units.

[0121] The raster operation unit 526 is a processing unit that performs raster operations, including, but not limited to, stencil, z-test, blending, and the like, and outputs pixel data as processed graphics data stored in the graphics memory (e.g., parallel processor memory 222 as shown in Fig. 2A and / or system memory 104 as in Fig. 1) to be stored for display on the one or more display devices 110A-110B, or for further processing by one of the one or more processors 102 or parallel processors 112. The raster operation unit 526 may be configured to compress z or color data written to memory and decompress z or color data read from memory. Overview of machine learning

[0122] The architecture described above can be applied to perform training and inference operations using machine learning models. Machine learning has proven successful in solving many types of tasks. The computations involved in training and using machine learning algorithms (e.g., neural networks) lend themselves, by their very nature, to efficient parallel implementations. Accordingly, parallel processors such as general-purpose graphics processors (GPGPUs) have played a significant role in the practical implementation of deep neural networks. Parallel graphics processors with single-instruction multithreaded (SIMT) architecture are designed to maximize the amount of parallel processing in the graphics pipeline. In a SIMT architecture, groups of parallel threads attempt to execute program instructions synchronously together as often as possible to increase processing efficiency.The efficiency provided by implementations of parallel machine learning algorithms allows the use of high-capacity networks and enables these networks to be trained on larger datasets.

[0123] A machine learning algorithm is an algorithm that can learn based on a set of data. For example, a machine learning algorithm can be configured to model high-level abstractions within a dataset. For example, image recognition algorithms can be used to determine which of several categories belongs to which given input; regression algorithms can output a numerical value given an input; and pattern recognition algorithms can be used to generate translated text or to perform text-to-speech and / or speech recognition.

[0124] An example type of machine learning algorithm is a neural network. There are many types of neural networks; a simple type of neural network is a feedforward network. A feedforward network can be implemented as an acyclic graph in which the nodes are arranged in layers. Typically, the topology of a feedforward network comprises an input layer and an output layer separated by at least one hidden layer. The hidden layer transforms an input received at the input layer into a representation useful for generating an output at the output layer. The network nodes are fully connected by edges to the nodes in adjacent layers, but there are no edges between nodes within each layer.Data received at the nodes of an input layer of a feedforward network is propagated (i.e., "fedforward") to the nodes of the output layer via an activation function that computes the states of the nodes of each subsequent layer in the network based on coefficients ("weights") associated with each edge connecting the layers. Depending on the specific model represented by the algorithm being executed, the output from the neural network algorithm can take various forms.

[0125] Before a machine learning algorithm can be used to model a specific problem, the algorithm is trained using a training dataset. Training a neural network involves selecting a network topology, using a training dataset representing a problem modeled by the network, and adjusting the weights until the network model performs with minimal error for all instances of the training dataset.For example, during a supervised learning training process for a neural network, the output produced by the network in response to the input representing an instance in a training dataset is compared with the designated "correct" output for that instance. An error signal representing the difference between the output and the designated output is calculated. The weights associated with the connections are adjusted to minimize this error as the error signal is backpropagated through the layers of the network. The network is considered "trained" when the errors for each of the outputs produced from the instances in the training dataset are minimized.

[0126] The accuracy of a machine learning algorithm can be significantly influenced by the quality of the dataset used to train the algorithm. The training process can be computationally intensive and can consume a significant amount of time on a conventional general-purpose processor. Accordingly, parallel processing hardware is used to train many types of machine learning algorithms. This is particularly useful for optimizing the training of neural networks, as the computations performed when adjusting the coefficients in neural networks naturally lend themselves to parallel implementations. In particular, many machine learning algorithms and software applications have been adapted to utilize the parallel processing hardware in general-purpose graphics processing devices.

[0127] Fig. 6 is a generalized diagram of a machine learning software stack 600. A machine learning application 602 is any logic that may be configured to train a neural network using a training dataset or to use a trained deep neural network to implement machine intelligence. The machine learning application 602 may include neural network training and inference functionality and / or specialized software that may be used to train a neural network prior to deployment. The machine learning application 602 may implement any type of machine intelligence, including, but not limited to, image recognition, mapping, and localization, autonomous navigation, speech synthesis, medical imaging, or language translation.Example machine learning applications 602 include, but are not limited to, voice-based virtual assistants, image or facial recognition algorithms, autonomous navigation, and the software tools used to train the machine learning models used by the machine learning applications 602.

[0128] Hardware acceleration for the machine learning application 602 may be enabled via a machine learning framework 604. The machine learning framework 604 may provide a library of machine learning primitives. Machine learning primitives are basic operations typically performed by machine learning algorithms. Without the machine learning framework 604, developers of machine learning algorithms would have to create and optimize the core computational logic associated with the machine learning algorithm and then re-optimize the computational logic as new parallel processors are developed. Instead, the machine learning application may be configured to perform the computations using the primitives provided by the machine learning framework 604.Example primitives include tensor convolutions, activation functions, and pooling, which are computational operations performed during training of a convolutional neural network (CNN). The machine learning framework 604 may also provide primitives for implementing basic linear algebra subroutines performed by machine learning algorithms, such as matrix and vector operations. Examples of a machine learning framework 604 include, but are not limited to, TensorFlow, TensorRT, PyTorch, MXNet, Caffe, and other high-level machine learning frameworks.

[0129] The machine learning framework 604 may process input data received from the machine learning application 602 and generate the appropriate input to a compute framework 606. The compute framework 606 may abstract the underlying instructions provided to the GPGPU driver 608 to enable the machine learning framework 604 to utilize hardware acceleration via the GPGPU hardware 610 without requiring the machine learning framework 604 to have detailed knowledge of the architecture of the GPGPU hardware 610. Furthermore, the compute framework 606 may enable hardware acceleration for the machine learning framework 604 across a variety of types and generations of the GPGPU hardware 610. Example compute frameworks 606 may include the CUDA compute framework and associated machine learning libraries, such as the CUDA Deep Neural Network (cuDNN) library.The machine learning software stack 600 may also include communication libraries or frameworks to enable multi-GPU and multi-node computing. GPGPU acceleration for machine learning

[0130] Fig. Figure 7 illustrates a general-purpose graphics processing unit (GPGPU 700) that the parallel processor 200 of Fig. 2A or the one or more parallel processors 112 of Fig. 1. The general-purpose processing unit (GPGPU) 700 may be configured to provide support for hardware acceleration of primitives provided by a machine learning framework to accelerate the processing of the type of computational workloads associated with training deep neural networks. Additionally, the GPGPU 700 may be directly linked with other instances of the GPGPU to create a multi-GPU cluster to improve training speed for particularly deep neural networks. Primitives are also supported to accelerate inference operations for deployed neural networks.

[0131] The GPGPU 700 includes a host interface 702 for enabling a connection to a host processor. The host interface 702 may be a PCI Express interface. However, the host interface may also be a vendor-specific communications interface or communications fabric. The GPGPU 700 receives commands from the host processor and uses a global scheduler 704 to distribute execution threads associated with those commands among a number of processing clusters 706A-706H. The processing clusters 706A-706H share a cache 708. The cache 708 may serve as a higher-level cache for caches within the processing clusters 706A-706H. The depicted processing clusters 706A-706H may be interconnected with the processing clusters 214A-214N as shown in Fig. 2A.

[0132] The GPGPU 700 includes memory 714A-714B coupled to the processing clusters 706A-706H via a set of memory controllers 712A-712B. The memory 714A-714B may include various types of memory devices, including dynamic random access memory (DRAM) or graphics random access memory, such as synchronous graphics random access memory (SGRAM), including double data rate graphics memory (GDDR). The memory 714A-714B may also include 3D stacked memory, including but not limited to high bandwidth memory (HBM).

[0133] Each of the processing clusters 706A-706H may include a set of graphics multiprocessors, such as the graphics multiprocessor 234 of Fig. 2D, the graphics multiprocessor 235 from Fig. 2E, or can a multi-core group 365A-365N as in Fig. 3. The graphics multiprocessors of the compute cluster include multiple types of integer and floating-point logic units capable of performing computational operations in a range of precisions, including those suitable for machine learning computations. For example, at least a subset of the floating-point units in each of the processing clusters 706A-706H may be configured to perform 16-bit or 32-bit floating-point operations, while another subset of the floating-point units may be configured to perform 64-bit floating-point operations.

[0134] Multiple instances of GPGPU 700 may be configured to operate as a compute cluster. The communication mechanism used by the compute cluster for synchronization and data exchange varies between embodiments. For example, the multiple instances of GPGPU 700 communicate via host interface 702. In one embodiment, GPGPU 700 includes an I / O hub 709 that couples GPGPU 700 to a GPU link 710, which enables direct connection to other instances of the GPGPU. GPU link 710 may be coupled to a dedicated GPU-to-GPU bridge, which enables communication and synchronization between multiple instances of GPGPU 700. Optionally, GPU link 710 may be coupled to a high-speed interconnect to send and receive data to other GPGPUs or parallel processors.The multiple instances of GPGPU 700 may be located in separate computing systems and communicate via a network device accessible via host interface 702. GPU link 710 may be configured to enable connection to a host processor in addition to, or as an alternative to, host interface 702.

[0135] Although the illustrated configuration of GPGPU 700 may be configured for training neural networks, an alternative configuration of GPGPU 700 may be configured for use in a high-performance inference platform or a low-performance inference platform. In an inference configuration, GPGPU 700 includes fewer of the processing clusters 706A-706H relative to the training configuration. Additionally, the technology associated with memory 714A-714B may differ between the inference and training configurations. In one embodiment, the inference configuration of GPGPU 700 may support inference-specific instructions. For example, an inference configuration may provide support for one or more 8-bit integer or floating-point dot product instructions commonly used during inference operations for deployed neural networks.

[0136] Fig. 8 illustrates a multi-GPU computing system 800. The multi-GPU computing system 800 may include a processor 802 coupled to a plurality of GPGPUs 806A-806D via a host interface switch 804. The host interface switch 804 may be a PCI Express switch device that couples the processor 802 to a PCI Express bus over which the processor 802 can communicate with the set of GPGPUs 806A-806D. Each of the plurality of GPGPUs 806A-806D may be an instance of the GPGPU 700 of Fig. 7. The GPGPUs 806A-806D may be interconnected via a set of high-speed point-to-point (P2P) GPU links 816. The high-speed GPU-to-GPU links may connect to each of the GPGPUs 806A-806D via a dedicated GPU link, such as the GPU link 710 as shown in Fig. 7. The P2P GPU links 816 enable direct communication between each of the GPGPUs 806A-806D without requiring communication over the host interface bus to which the processor 802 is connected. When GPU-to-GPU traffic is routed to the P2P GPU links, the host interface bus remains available for system memory access or for communication with other instances of the multi-GPU computing system 800, for example, over one or more network devices. While in Fig. 8, the GPGPUs 806A-806D are connected to the processor 802 via the host interface switch 804, the processor 802 may alternatively include direct support for the P2P GPU links 816 and be directly connected to the GPGPUs 806A-806D. In one embodiment, the P2P GPU link 816 enables the multi-GPU computing system 800 to operate as a single logical GPU. Machine learning implementations of neural networks

[0137] The computational architecture described herein can be configured to perform the types of parallel processing particularly suitable for training and deploying neural networks for machine learning. A neural network can be generalized as a network of functions that are related in a graph. As is well known in the art, there are a variety of types of neural network implementations used in machine learning. One exemplary type of neural network is the feedforward network, as previously described. A second exemplary type of neural network is the convolutional neural network (CNN), while a third exemplary type of neural network is recurrent neural networks (RNNs).

[0138] A CNN is a specialized feedforward neural network for processing data with a known grid-like topology, such as image data. Accordingly, CNNs are typically used for computer vision and image recognition applications, but they can also be used for other types of pattern recognition, such as speech and language processing. The nodes of the CNN's input layer are organized into a set of "filters" (feature detectors inspired by the receptive fields found in the retina), and the output of each set of filters is propagated to nodes in successive layers of the network. The computations for a CNN involve applying the mathematical operation of convolution to each filter to produce that filter's output.Convolution is a specialized type of mathematical operation performed on two functions to produce a third function that is a modified version of one of the two original functions. In the terminology of a convolutional neural network, the first function of the convolution can be referred to as the input, while the second function can be referred to as the convolution kernel. The output can be referred to as the feature map. For example, the input to a convolutional layer can be a multidimensional array of data defining the various color components of an input image. The convolution kernel can be a multidimensional array of parameters, with the parameters being adjusted through the training process for the neural network.

[0139] RNNs are a family of feedforward neural networks that incorporate feedback connections between layers. RNNs enable the modeling of sequential data by sharing parameter data across different parts of the neural network. Cycles represent the influence of a present value of a variable on its own value at a future time, as at least a portion of the RNN's output data is used as feedback for processing following an input in a sequence. This property makes RNNs particularly suitable for language processing due to the variable nature of language data.

[0140] The figures described below depict exemplary feedforward, CNN, and RNN networks and describe a general process for training and deploying each of these types of networks, respectively. It should be understood that these descriptions are exemplary and not limiting of any specific embodiment described herein, and the concepts illustrated can be applied generally to deep neural networks and machine learning techniques in general.

[0141] Deep neural networks used in deep learning typically involve a front-end network for performing feature detection coupled with a back-end network representing a mathematical model that can perform operations (e.g., object classification, speech recognition, etc.) based on the feature representation provided to the model. Deep learning allows machine learning to be performed without the need for manual feature engineering of the model. Instead, deep neural networks can learn features based on a statistical structure or correlation within the input data. The learned features can be provided to a mathematical model that can map detected features to an output.The mathematical model used by the network is generally specialized for the specific task to be performed, and other models are used to perform other tasks.

[0142] Once the neural network is structured, a learning model can be applied to the network to train it to perform specific tasks. The learning model describes how to adjust the weights within the model to reduce the network's output error. Error backpropagation is a common technique used to train neural networks. An input vector is presented to the network for processing. The network's output is compared to the desired output using a loss function, and an error value is calculated for each of the neurons in the output layer. The error values ​​are then backpropagated until each neuron has an associated error value that roughly represents its contribution to the original output.The network can then learn from these errors using an algorithm, such as the stochastic gradient descent algorithm, to update the weights of the neural network.

[0143] Fig. Figures 9A-9B illustrate an example convolutional neural network. Fig. Figure 9A illustrates different layers within a CNN. As in Fig. As shown in Figure 9A, an exemplary CNN used to model image processing may receive an input 902 describing the red, green, and blue (RGB) components of an input image. The input 902 may be processed by multiple convolutional layers (e.g., first convolutional layer 904, second convolutional layer 906). The output from the multiple convolutional layers may optionally be processed by a set of fully connected layers 908. Neurons in a fully connected layer have full connections with all activations in the previous layer, as previously described for a feedforward network. The output from the fully connected layers 908 may be used to generate an output result of the network. The activations within the fully connected layers 908 may be computed using matrix multiplication instead of convolution.Not all CNN implementations use fully connected layers 908. For example, in some implementations, the second convolutional layer 906 may generate an output for the CNN.

[0144] The convolutional layers are sparsely connected, which differs from the conventional neural network configuration found in the fully connected layers 908. Conventional neural network layers are fully connected, so each output unit interacts with each input unit. However, the convolutional layers are sparsely connected because the output of the convolution of a field (rather than the respective state value of each of the nodes in the field) is input to the nodes of the subsequent layer, as illustrated. The kernels associated with the convolutional layers perform convolution operations, the result of which is passed on to the next layer. The dimensionality reduction performed within the convolutional layers is one aspect that allows the CNN to scale to process large images.

[0145] Fig. Figure 9B illustrates exemplary computation stages within a convolutional layer of a CNN. An input to a convolutional layer 912 of a CNN may be processed in three stages of a convolutional layer 914. The three stages may include a convolution stage 916, a detector stage 918, and a pooling stage 920. The convolutional layer 914 may then output data to a subsequent convolutional layer. The final convolutional layer of the network may generate output feature map data or provide input to a fully connected layer, for example, to generate a classification score for the input to the CNN.

[0146] In the convolution stage 916, multiple convolutions are performed in parallel to generate a set of linear activations. The convolution stage 916 may include an affine transformation, which is any transformation that can be specified as a linear transformation plus a translation. Affine transformations include rotations, translations, scaling, and combinations of these transformations. In the convolution stage, the output of functions (e.g., neurons) connected to specific regions in the input, which can be determined as the local region connected to the neuron, is computed. The neurons compute a dot product between the weights of the neurons and the region in the local input to which the neurons are connected. The output from the convolution stage 916 defines a set of linear activations that are processed by successive stages of the convolution layer 914.

[0147] The linear activations can be processed by a detector stage 918. In the detector stage 918, each linear activation is processed by a nonlinear activation function. The nonlinear activation function increases the nonlinear properties of the overall network without affecting the receptive fields of the convolutional layer. Various types of nonlinear activation functions can be used. One specific type is the rectified linear unit (ReLU), which uses an activation function defined as f(x) = max(0, x), such that the activation has a threshold of zero.

[0148] The pooling stage 920 uses a pooling function that replaces the output of the second convolutional layer 906 with a summary statistic of the nearby outputs. The pooling function can be used to introduce translation invariance into the neural network, so that small translations of the input do not change the pooled outputs. Invariance to local translation can be useful in scenarios where the presence of a feature in the input data is more useful than the exact location of the feature. Various types of pooling functions can be used during the pooling stage 920, including max pooling, average pooling, and L2 norm pooling. Additionally, some CNN implementations do not include a pooling stage. Instead, such implementations substitute an additional convolution stage with an increased stride size relative to previous convolution stages.

[0149] The output from the convolutional layer 914 can then be processed by the next layer 922. The next layer 922 can be an additional convolutional layer or one of the fully connected layers 908. The first convolutional layer 904 of Fig. 9A may output to the second convolutional layer 906, while the second convolutional layer may output to a first layer of the fully connected layers 908.

[0150] A variant of the CNN is a convolutional deep belief network, which has a similar structure to a CNN and is trained similarly to a deep belief network. A deep belief network (DBN) is a generative neural network composed of multiple layers of stochastic (random) variables. DBNs can be trained layer by layer using greedy unsupervised learning. The learned weights of the DBN can then be used to provide pre-training neural networks by determining an initial set of weights for the neural network.

[0151] Fig. 10A-10B illustrate example language models. Fig. Figure 10A illustrates a recurrent neural network (RNN). In a recurrent neural network (RNN), the previous state of the network influences the output of the current state of the network. RNNs can be constructed in many ways using a variety of features. Using RNNs generally involves using mathematical models to predict the future based on a previous sequence of inputs. For example, an RNN can be used to perform statistical language modeling to predict an upcoming word based on a previous sequence of words.The illustrated RNN 1000 can be described as having an input layer 1002 that receives an input vector, hidden layers 1004 for implementing a recurrent function, a feedback mechanism 1005 for enabling "memory" of previous states, and an output layer 1006 for outputting a result. The RNN 1000 operates on a time-step basis. The state of the RNN at a given time step is influenced based on the previous time step via the feedback mechanism 1005. For a given time step, the state of the hidden layers 1004 is defined by the previous state and the input at the current time step. An initial input (x1) at a first time step can be processed by the hidden layer 1004.A second input (x2) may be processed by hidden layer 1004 using state information determined during processing of the initial input (x1). A given state may be denoted as s. t = f(Ux t + Ws t-1), where U and W are parameter matrices. The function f is generally a nonlinearity, such as the hyperbolic tangent function (tanh) or a variant of the rectifier function f(x) = max(0, x). However, the specific mathematical function used in the hidden layers 1004 can vary depending on the specific implementation details of the RNN 1000. Acceleration for variations in RNN networks can also be enabled. An example RNN variant is the long short-term memory (LTSM) RNN. LSTM RNNs are capable of learning long-term dependencies that can be used to process longer speech sequences.

[0152] Fig. Figure 10B illustrates baseline components of a transformer model 1010. The transformer model 1010 addresses problems inherent in RNN models with long input sequences and enables greater parallelization. The transformer model 1010 and variants thereof are used to create language models (LLMs) that perform various tasks such as machine translation, automatic summarization, and dialogue management. A transformer model 1010 can also be configured to perform image generation tasks, such as text-to-image generation.

[0153] The transformer model 1010 includes multiple instances of an encoder 1016 and a decoder 1026. The encoders are stacked end-to-end, with the output of the last encoder routed as input to a multihead attention layer of each decoder 1026. An input entering the bottom-most instance of the encoder 1016 is processed by input embeddings 1012, which convert input tokens into vectors that can be processed by the encoder 1016. The output dictionary is vectorized by output embeddings 1022 before entering the bottom-most instance of the decoder 1026. Position encodings 1014, 1024 are added to input vectors and output vectors at the bottom of the encoder 1016 and decoder 1026 stacks.The position encodings 1014, 1024 inject information about the relative or absolute position of tokens in a sequence of tokens to be processed, since the transformer model 1010 does not naturally encode the order of the tokens.

[0154] The encoder 1016 of the transformer model 1010 analyzes the input text and creates a number of hidden states that protect the context and meaning of the text data. The layers of the encoder 1016 form part of the core of the transformer architecture, although variants of the transformer model 1010 are also possible for the decoder 1016. The encoder 1016 includes two sublayers: an MHA (multihead attention) sublayer and an FFN (feedforward network) layer. The MHA sublayer performs multiple concurrent self-attention operations to compute attention scores that enable the transformer model 1010 to context-awarely weight the importance and relative relationships of different tokens in the input sequence. The FFN sublayer is a positionally fully connected feedforward neural network.The output of each sublayer is processed by an addition and normalization (A&N) operation defined as LayerNorm(x + SubLayer(x)), where the output of the sublayer is added to the sublayer's input and normalized. In some implementations, the normalization operation may be performed before the sublayer, not after.

[0155] Decoder 1026 includes three sublayers: a masked MHA sublayer, an MHA sublayer, and an FFN sublayer. The masked MHA sublayer is similar to the MHA layer, except that it is masked to prevent the query position from being mapped to the keys of future positions. The MHA sublayer of decoder 1026 is similar to the MHA sublayer of the encoder, with an additional input from encoder 1016. The FFN of decoder 1026 is the same as the FFN of encoder 1016. The Linear and SoftMax block takes the output of the last instance of decoder 1026 of the decoder stack and generates a probability distribution representing output probabilities.

[0156] GPU acceleration is also available for variants of the Transformer Model 1010 that replace some or all of the FFN sublayers with sparse Mixture of Experts (MoE) layers. MoE layers each comprise a number of experts, with each expert being a neural network. The MoE layers can be FFNs themselves or MoEs, allowing for hierarchical MoE layers.

[0157] The training of a transformer model 1010 can be optimized through the use of adaptive precision logic, which adjusts the computational precision applied during training. The adaptive precision logic can attempt to use the smallest possible data type during training without significantly reducing training accuracy. For example, to minimize data loss due to the use of 16-bit, 8-bit, and 4-bit floating-point formats, dynamic scaling and casting can be applied during training based on the statistical analysis of the tensor data generated during training. Tensor data generated during training can be statistically analyzed using a variety of analysis techniques to determine a set of scaling factors to be applied to data blocks, such as the input to each layer of the transformer model 1010.For example, absolute minimum and / or maximum values ​​can be used to determine scaling factors that can prevent underflow or overflow of low-precision floating-point data types.

[0158] Fig. Figure 11 illustrates the training and deployment of a deep neural network. Once a given network has been structured for a task, the neural network is trained using a training dataset 1102. Various training frameworks 1104 have been developed to enable hardware acceleration of the training process. For example, the machine learning framework 604 of Fig. 6 may be configured as a training framework 1104. The training framework 1104 may hook into an untrained neural network 1106 and allow the untrained neural network to be trained using the parallel processing resources described herein to generate a trained neural network 1108. To begin the training process, the initial weights may be chosen randomly or by pretraining using a deep belief network. The training cycle may then be performed in either a supervised or unsupervised manner.

[0159] Supervised learning is a learning method in which training is performed as a mediated operation, e.g., when the training dataset 1102 contains inputs paired with the desired output for the input, or when the training dataset contains inputs with a known output and the output of the neural network is manually evaluated. The network processes the inputs and compares the resulting outputs with a set of expected or desired outputs. Errors are then backpropagated through the system. The training framework 1104 can adjust the weights that control the untrained neural network 1106. The training framework 1104 can provide tools to monitor how well the untrained neural network 1106 converges toward a model capable of generating correct answers based on known input data.The training process occurs repeatedly as the network's weights are adjusted to refine the output generated by the neural network. The training process can continue until the neural network reaches a statically desired accuracy, which is associated with a trained neural network 1108. The trained neural network 1108 can then be used to implement any number of machine learning operations to generate an inference result 1114 based on an input of new data 1112.

[0160] Unsupervised learning is a learning technique in which the network attempts to train itself using unlabeled data. Thus, the training dataset 1102 for unsupervised learning will include input data without associated output data. The untrained neural network 1106 can learn groupings within the unlabeled input and can determine how individual inputs relate to the overall dataset. Unsupervised training can be used to generate a self-organizing map, which is a type of trained neural network 1108 capable of performing operations useful for reducing the dimensionality of data. Unsupervised training can also be used to perform anomaly detection, which enables the identification of data points in an input dataset that deviate from the normal patterns of the data.

[0161] Variations of supervised and unsupervised training can also be used. Semi-supervised learning is a technique in which the training dataset 1102 includes a mixture of labeled and unlabeled data of the same distribution. Incremental learning is a variation of supervised learning in which input data is continuously used to further train the model. Incremental learning allows the trained neural network 1108 to adapt to new data 1112 without forgetting the knowledge imparted to the network during an initial training. Regardless of whether it is supervised or unsupervised, the training process for particularly deep neural networks can be too computationally intensive for a single compute node. Instead of using a single compute node, a distributed network of compute nodes can be used to accelerate the training process.

[0162] Fig. Figure 12A is a block diagram illustrating distributed learning. Distributed learning is a training model that uses multiple distributed compute nodes to perform supervised or unsupervised training of a neural network. The distributed compute nodes may each include one or more host processors and one or more of the general-purpose processing nodes, such as the highly parallel general-purpose graphics processing unit 700 as shown in Figure 700. As illustrated, distributed learning may be performed with model parallelism 1202, data parallelism 1204, or a combination of model and data parallelism 1206.

[0163] With model parallelism 1202, different compute nodes in a distributed system can perform training computations for different parts of a single network. For example, each layer of a neural network can be trained by a different processing node of the distributed system. The advantages of model parallelism include the ability to scale up to particularly large models. Sharing the computations associated with different layers of the neural network enables the training of very large neural networks in which the weights of all layers would not fit in the memory of a single compute node. In some cases, model parallelism can be particularly useful for performing unsupervised training of large neural networks. Some implementations of model parallelism 1202 may also be referred to as tensor parallelism.

[0164] With data parallelism 1204, the different nodes of the distributed network have a complete instance of the model, and each node receives a different portion of the data. The results from the different nodes are then combined. While different approaches to data parallelism are possible, all data-parallel training approaches utilize a technique for combining results and synchronizing model parameters between individual nodes. Example approaches to combining data include parameter averaging and update-based data parallelism. Parameter averaging trains each node on a subset of the training data and sets the global parameters (e.g., weights, biases) to the average of the parameters of each node. Parameter averaging uses a central parameter server that maintains the parameter data.Update-based data parallelism is similar to parameter averaging, except that instead of transmitting parameters from nodes to the parameter server, model updates are transmitted. Additionally, update-based data parallelism can be implemented in a decentralized manner, with updates compressed and transmitted between nodes.

[0165] For example, the combined model and data parallelism 1206 may be implemented in a distributed system where each compute node includes multiple GPUs. The combined model and data parallelism 1206 may also be referred to as hybrid parallelism. Each node may contain a complete instance of the model, with separate GPUs within each node being used to train different portions of the model. Distributed training has increased overhead compared to training on a single machine. However, the parallel processors and GPGPUs described herein may each implement various techniques for reducing the overhead of distributed learning, including techniques for enabling high-bandwidth data transfer between GPUs and accelerated remote data synchronization.Pipeline parallelism is a variant of combined model and data parallelism 1206, in which different nodes comprise less than the entire model but more than a single layer of the model. With pipeline parallelism, different groups of layers or submodels are distributed across the different processing nodes. Another variant is expert parallelism, which routes requests to specific experts within the model to different GPUs. Expert parallelism can be used, for example, in MoE transformer models.

[0166] Fig. 12B is a block diagram illustrating a programmable network interface 1210 and a data processing unit. The programmable network interface 1210 is a programmable network engine that can be used to accelerate network-based computing tasks within a distributed environment. The programmable network interface 1210 can be coupled to a host system via a host interface 1270. The programmable network interface 1210 can be used to accelerate network or memory operations for CPUs or GPUs of the host system. The host system can be, for example, a node of a distributed learning system used to perform distributed training, such as in Fig. 12A. The host system may also be a data center node within a data center.

[0167] In one embodiment, access to remote storage containing model data may be accelerated by the programmable network interface 1210. For example, the programmable network interface 1210 may be configured to present remote data storage devices to the host system as local data storage devices. The programmable network interface 1210 may also accelerate RDMA (remote direct memory access) operations performed between GPUs of the host system and GPUs of remote systems. In one embodiment, the programmable network interface 1210 may enable data storage functionality, such as, but not limited to, NVME-oF.The programmable network interface 1210 can also accelerate encryption, data integrity, compression, and other remote data storage operations on behalf of the host system, allowing remote data storage to approach the latencies of data storage devices directly attached to the host system.

[0168] The programmable network interface 1210 can also perform resource allocation and management for the host system. Data storage security operations can be offloaded to the programmable network interface 1210 and performed in conjunction with the allocation and management of remote data storage resources. Network-based operations for managing access to the remote data storage, which would otherwise be performed by a host system processor, can instead be performed by the programmable network interface 1210.

[0169] In one embodiment, network and / or data security operations may be offloaded from the host system to the programmable network interface 1210. Data center security policies for a data center node may be handled by the programmable network interface 1210 instead of the host system's processors. For example, the programmable network interface 1210 may detect and mitigate an attempted network-based attack (e.g., DDoS) on the host system, preventing the attack from impacting the host system's availability.

[0170] The programmable network interface 1210 may include a system-on-chip (SoC 1220) executing an operating system across multiple processor cores 1222. The processor cores 1222 may include general-purpose processor cores (e.g., CPU cores). In one embodiment, the processor cores 1222 may also include one or more GPU cores. The SoC 1220 may execute instructions stored in a memory device 1240. A storage device 1250 may store local operating system data. The storage device 1250 and the memory device 1240 may also be used to cache remote data for the host system. The network ports 1260A-1260B enable connection to a network or fabric and provide network access to the SoC 1220 and, via the host interface 1270, the host system. The programmable network interface 1210 may also include an I / O interface 1275, such as a USB interface.The I / O interface 1275 can be used to couple external devices to the programmable network interface 1210 or as a debug interface. The programmable network interface 1210 also includes a management interface 1230 that enables software on the host device to manage and configure the programmable network interface 1210 and / or the SoC 1220. In one embodiment, the programmable network interface 1210 can also include one or more accelerators or GPUs 1245 for accepting offloading of parallel computing tasks from the SoC 1220, the host system, or remote systems coupled via the network ports 1260A-1260B. Example machine learning applications

[0171] Machine learning can be applied to solve a variety of technological problems, including computer vision, autonomous driving and navigation, speech recognition, and natural language processing. Machine vision has traditionally been one of the most active research areas for machine learning applications. Applications of machine vision range from replicating human visual abilities, such as face recognition, to creating new categories of visual abilities. For example, machine vision applications can be configured to detect sound waves from the vibrations induced in objects visible in a video.Parallel processor-accelerated machine learning enables computer vision applications to be trained using significantly larger amounts of training data than was previously possible and enables inference systems to be deployed using low-performance parallel processors.

[0172] Parallel processor-accelerated machine learning has applications in autonomous driving, including lane and traffic sign recognition, obstacle avoidance, navigation, and driver control. Accelerated machine learning techniques can be used to train driving models based on data sets that define the appropriate responses to specific training inputs. The parallel processors described here can enable rapid training of the increasingly complex neural networks used for autonomous driving solutions and enable the use of low-power inference processors in a mobile platform suitable for integration into autonomous vehicles.

[0173] Parallel processor-accelerated deep neural networks have enabled machine learning approaches for automatic speech recognition (ASR). ASR involves generating a function that computes the most probable linguistic sequence given an acoustic input sequence. Accelerated machine learning using deep neural networks has enabled the replacement of the hidden Markov models (HMMs) and Gaussian mixture models (GMMs) previously used for ASR.

[0174] Parallel processor-accelerated machine learning can also be used to accelerate natural language processing. Automatic learning workflows can use statistical inference algorithms to generate models that are resilient to erroneous or unfamiliar inputs. Examples of applications for natural language processors include automatic machine translation between human languages.

[0175] The parallel processor platforms used for machine learning can be divided into training platforms and deployment platforms. Training platforms are generally highly parallel and include optimizations to accelerate multi-GPU single-node training and multi-node multi-GPU training. Examples of parallel processors suitable for training include the GPGPU 700 from Fig. 7 and the multi-GPU computing system 800 from Fig. 8. In contrast, deployed machine learning platforms generally include low-power parallel processors suitable for use in products such as cameras, autonomous robots, and autonomous vehicles.

[0176] Additionally, machine learning techniques can be applied to accelerate or enhance graphics processing activities. For example, a machine learning model can be trained to recognize an output generated by a GPU-accelerated application and generate an upscaled version of that output. Such techniques can be applied to accelerate the generation of high-resolution images for a gaming application. Various other graphics pipeline activities can benefit from the use of machine learning. For example, machine learning models can be trained to perform tessellation operations on geometry data to increase the complexity of geometric models, allowing high-detail geometry to be automatically generated from geometry of relatively lower detail.

[0177] Fig. 13 illustrates an exemplary inference system-on-chip (SOC 1300) suitable for performing inference using a trained model. The SOC 1300 may integrate processing components including a media processor 1302, an image processing processor 1304, a GPGPU 1306, and a multi-core processor 1308. The GPGPU 1306 may be a GPGPU as described herein, such as GPGPU 700, and the multi-core processor 1308 may be a multi-core processor described herein, such as multi-core processors 405-406. The SOC 1300 may additionally include an on-chip memory 1305, which may enable a shared on-chip data pool accessible by each of the processing components. The processing components can be optimized for low-power operation to enable use on a variety of machine learning platforms, including autonomous vehicles and autonomous robots.For example, an implementation of the SoC 1300 can be used as a section of the main control system for an autonomous vehicle. When the SOC 1300 is configured for use in autonomous vehicles, the SOC is designed and configured to comply with the relevant functional safety standards of the legislation of the area of ​​use.

[0178] During operation, media processor 1302 and vision processor 1304 may cooperate to accelerate computer vision operations. Media processor 1302 may enable low-latency decoding of multiple high-resolution video streams (e.g., 4K, 8K). The decoded video streams may be written to a buffer in on-chip memory 1305. Image processing processor 1304 may then parse the decoded video and perform preliminary processing operations on the frames of the decoded video in preparation for processing the frames using a trained image recognition model. For example, image processing processor 1304 may accelerate convolution operations for a CNN used to perform image recognition on the high-resolution video data, while the backend model computations are performed by GPGPU 1306.

[0179] The multi-core processor 1308 may include control logic that assists in the sequencing and synchronization of data transfers and shared memory operations performed by the media processor 1302 and the vision processor 1304. The multi-core processor 1308 may also serve as an application processor for executing software applications that can utilize the inference computing power of the GPGPU 1306. For example, at least some of the navigation and driving logic may be implemented in software executing on the multi-core processor 1308. Such software may issue computational workloads directly to the GPGPU 1306, or the computational workloads may be issued to the multi-core processor 1308, which may offload at least some of these operations to the GPGPU 1306.

[0180] The GPGPU 1306 may include compute clusters, such as a low-performance configuration of processing clusters 706A-706H within the GPGPU 700. The compute clusters within the GPGPU 1306 may support instructions specifically optimized to perform inference calculations on a trained neural network. For example, the GPGPU 1306 may support instructions to perform low-precision calculations, such as 8-bit and 4-bit integer vector operations. Overview of additional systems

[0181] Fig. 14 is a block diagram of a processing system 1400. The elements of Fig. 14 having the same or similar names as the elements of any other FIG. herein describe the same elements as in the other figures, may operate or function in a similar manner, may include the same components, and may be associated with entities other than those described elsewhere herein, but are not limited thereto. The processing system 1400 may be implemented in a single-processor desktop system, a multiprocessor workstation system, or a server system having one or more processors 1402 or processor cores 1407. The system 1400 may be a processing platform incorporated into a system-on-chip (SoC) integrated circuit for use in mobile, wearable, or embedded devices, such as Internet of Things (IoT) devices with wired or wireless connectivity to a local or wide area network.

[0182] The processing system 1400 may be a processing system having components similar to those of Fig. 1. In other configurations, for example, the one or more processors 1402 or processor cores 1407 may correspond to the one or more processors 102 of Fig. 1. The graphics processor(s) 1408 may be provided to the parallel processor(s) 112 of Fig. 1. The external graphics processor 1418 may be one of the add-in device(s) 120 from Fig. 1.

[0183] Processing system 1400 may include, be coupled to, or integrated with a server-based gaming platform; a game console, including a game and media console; a mobile gaming console, a handheld game console, or an online game console. System 1400 may be part of a mobile phone, smartphone, tablet computing device, or mobile internet-connected device, such as a laptop with low internal storage capacity.The processing system 1400 may also include, be coupled with, or integrated with: a wearable device, such as a smartwatch wearable device; smart glasses or smart clothing enhanced with augmented reality (AR) or virtual reality (VR) features to provide visual, auditory, or tactile output to supplement visual, auditory, or tactile real-world experiences, or otherwise provide text, audio, graphics, video, holographic images or video, or tactile feedback; another augmented reality (AR) device; or another virtual reality (VR) device. The processing system 1400 may include or be part of a television or set-top box device.System 1400 may include, be coupled to, or be integrated with a self-driving vehicle, such as a bus, tractor, passenger car, motorcycle or electric bicycle, aircraft, or glider (or any combination thereof). The self-driving vehicle may use processing system 1400 to process the environment perceived around the vehicle.

[0184] The one or more processors 1402 may include one or more instances of processor cores 1407 to process instructions that, when executed, perform operations for system and application software. At least one of the one or more processor cores 1407 may be configured to process a specific instruction set 1409. The instruction set 1409 may support complex instruction set computing (CISC), reduced instruction set computing (RISC), or extra-long instruction word (VLIW) computing. In one embodiment, one or more processor cores 1407 may process a different instruction set 1409, which may include instructions to support emulation of other instruction sets. The processor core 1407 may also include other processing devices, such as a digital signal processor (DSP).

[0185] The one or more processors 1402 may include cache memory 1404. Depending on the architecture, the processor 1402 may have a single internal cache or multiple levels of internal cache. In some embodiments, the cache memory is shared by various components of the processor(s) 1402. In some embodiments, the processor(s) 1402 also utilize an external cache (e.g., a Level 3 (L3) cache or a last level cache (LLC)) (not shown) that may be shared by processor cores 1407 using known cache coherence techniques. A register file 1406 may additionally be included in the processor(s) 1402 and may include various types of registers for storing different data types (e.g., integer registers, floating point registers, status registers, and an instruction pointer register).Some registers may be general purpose registers, while other registers may be specific to the design of the processor(s) 1402.

[0186] The one or more processors 1402 may be coupled to one or more interface buses 1410 for transmitting communication signals, such as address, data, or control signals, between the one or more processors 1402 and other components in the processing system 1400. The interface bus 1410, in one embodiment, may be a processor bus, such as a version of the DMI (Direct Media Interface) bus. However, processor buses are not limited to the DMI bus but may also include one or more PCI (Peripheral Component Interconnect) buses (e.g., PCI, PCI express), memory buses, or other types of interface buses. For example, the processor(s) 1402 may include an integrated memory controller 1416 and a platform control hub 1430.The memory controller 1416 enables communication between a memory device and other components of the processing system 1400, while the platform control hub 1430 provides connections to I / O devices via a local I / O bus.

[0187] The memory device 1420 may be a dynamic random access memory (DRAM) device, a static random access memory (SRAM) device, a flash memory device, a phase-change memory device, or other memory device having suitable performance to serve as process memory. For example, the memory device 1420 may serve as system memory for the processing system 1400 to store data 1422 and instructions 1421 used when the processor(s) 1402 execute an application or process. The memory controller 1416 may optionally be coupled to an external graphics processor 1418 that can communicate with the one or more graphics processors 1408 in the one or more processors 1402 to perform graphics and media operations.In some embodiments, graphics, media, or compute operations may be supported by an accelerator 1412, which is a coprocessor that can be configured to perform a specific set of graphics, media, or compute operations. For example, the accelerator 1412 may be a matrix multiplication accelerator used to optimize machine learning or compute operations. The accelerator 1412 may be a ray tracing accelerator that can be used to perform ray tracing operations in conjunction with the graphics processor 1408. The accelerator 1412 may also be an AI accelerator or NPU to accelerate neural network training or inference operations. In one embodiment, an external accelerator 1419 may be used instead of or in conjunction with the accelerator 1412.The accelerator 1412 and / or the external accelerator 1419 may have similar functionality to the accelerator(s) 130 of FIG. Fig. 1.

[0188] A display device 1411 may be provided that can be connected to the one or more processors 1402. The display device 1411 may be an internal display device, such as in a mobile electronic device or a laptop device, and / or an external display device connected via a display interface (e.g., DisplayPort, etc.). The display device 1411 may be a head-mounted display (HMD), such as a stereoscopic display device for use in virtual reality (VR) or augmented reality (AR) applications.

[0189] The platform control hub 1430 may enable peripherals to be connected to the storage device 1420 and the one or more processors 1402 via a high-speed I / O bus. The I / O peripherals include, among others, an audio controller 1446, a network controller 1434, a firmware interface 1428, a wireless transceiver 1426, touch sensors 1425, and a data storage device 1424 (e.g., non-volatile memory, volatile memory, hard disk drive, flash memory, NAND, 3D NAND, 3D XPoint / Optane, etc.). The data storage device 1424 may be connected via a storage interface (e.g., SATA) or via a peripheral bus, such as a peripheral component interconnect bus (e.g., PCI, PCI Express). The touch sensors 1425 may include touchscreen sensors, pressure sensors, or fingerprint sensors.The wireless transceiver 1426 may be a WiFi transceiver, a Bluetooth transceiver, or a cellular network transceiver, such as a 3G, 4G, 5G, or LTE (Long-Term Evolution) transceiver. The firmware interface 1428 enables communication with system firmware and may, for example, be a Unified Extensible Firmware Interface (UEFI). The network controller 1434 may enable network connection to a wired network. In some embodiments, a high-performance network controller (not shown) is coupled to the interface bus(es) 1410. The audio controller 1446 may be a multi-channel high-definition audio controller. In some of these embodiments, the processing system 1400 includes an optional legacy I / O controller 1440 to interface legacy devices (e.g., Personal System 2 (PS / 2)) to the system.The platform control hub 1430 may also be connected to one or more USB (Universal Serial Bus) controllers 1442 to connect to input devices such as keyboard and mouse combinations 1443, a camera 1444, or other USB input devices.

[0190] It should be understood that the processing system 1400 shown is exemplary and not limiting, as other types of data processing systems configured differently may also be used. For example, an instance of the memory controller 1416 and the platform control hub 1430 may be integrated into a discrete external graphics processor, such as the external graphics processor 1418. The platform control hub 1430 and / or the memory controller 1416 may be external to the one or more processors 1402. For example, the memory controller 1416 and the platform control hub 1430 may be external to the processing system 1400 and configured as a memory control hub and a peripheral control hub within a system chipset in communication with the processor(s) 1402.

[0191] For example, printed circuit boards (“sleds”) can be used to house components such as CPUs, memory, and other components designed for increased thermal performance. Processing components, such as processors, can be located on the top of a sled, while nearby memory, such as DIMMs, is located on the bottom of the sled. As a result of the improved airflow provided by this design, the components can operate at higher frequencies and power levels than in typical systems, thereby increasing performance. Furthermore, the sleds are configured to be blindly connected to power and data communication cables in a rack, enhancing their ability to be quickly removed, upgraded, reinstalled, and / or replaced.Likewise, individual components located on the sleds, such as processors, accelerators, memory, and storage drives, are configured to be easily upgraded thanks to their increased spacing. In the exemplary embodiment, the components additionally include hardware attestation features to prove their authenticity.

[0192] A data center can utilize a single network architecture ("fabric") that supports multiple other network architectures, including Ethernet and Omni-Path. The sleds can be coupled to switches using optical fibers that provide higher bandwidth and lower latency than typical twisted-pair cabling (e.g., Category 5, Category 5e, Category 6, Category 7, Category 8, etc.). Due to the high-bandwidth, low-latency connections and network architecture, the data center can pool resources, such as memory, accelerators (e.g., GPUs, graphics accelerators, FPGAs, ASICs, neural network and / or artificial intelligence accelerators, etc.), and data storage drives, which are physically separate, and expose them to compute resources (e.g., processors), allowing the compute resources to access the pooled resources as if they were local.

[0193] A power supply or source may provide voltage and / or current to the processing system 1400 or any component or system described herein. In one example, the power supply includes an AC-to-DC adapter for connection to an electrical outlet. This AC power may be sourced from renewable energy sources (e.g., solar power). In one example, the power source includes a DC power source, such as an external AC-to-DC converter. A power source or source may also include wireless charging hardware for charging via proximity to a charging pad. The power source may include an internal battery, an AC power supply, a motion-based power supply, a solar energy supply, or a fuel cell source.

[0194] Fig. Figures 15A-15C illustrate computing systems and graphics processors. The elements of Fig. 15A-15C having the same or similar names as the elements of any other figure herein describe the same elements as in the other figures, may operate or function in a similar manner, may include the same components, and may be associated with other entities such as, but are not limited to, those described elsewhere herein.

[0195] Fig. 15A is a block diagram of a processor 1500, which may be a variant of one of the one or more processors 1402 and may be used in place of one or more of them. Therefore, the disclosure of any features in combination with the processor 1500 herein also discloses, but is not limited to, a corresponding combination with the one or more processors 1402. The processor 1500 may include one or more processor cores 1502A-1502N, at least one memory controller 1514, and a graphics processor 1508. The graphics processor 1508 may be integrated into the processor 1500 or into a system chipset, or coupled via a system bus. The processor 1500 may include additional cores, up to and including the additional core 1502N, represented by the dashed boxes. Each of the processor cores 1502A-1502N includes one or more internal cache units 1504A-1504N.In some embodiments, each processor core 1502A-1502N also has access to one or more shared cache units 1506. The internal cache units 1504A-1504N and the one or more shared cache units 1506 represent a cache hierarchy within the processor 1500. The cache hierarchy may include at least one level of instruction and data cache within each processor core and one or more levels of mid-level shared cache, such as a Level 2 (L2), Level 3 (L3), Level 4 (L4), or other cache levels, with the highest cache level prior to external memory being classified as the LLC. In some embodiments, the cache coherency logic maintains coherency between the various cache units (e.g., the one or more shared cache units 1506 and the one or more internal cache units 1504A-1504N).

[0196] The processor 1500 may also include one or more bus control units 1516 and a system agent core 1510. The one or more bus control units 1516 manage a set of peripheral buses, such as one or more PCI or PCI Express buses. The one or more bus control units 1516 may also manage one or more memory buses to various external storage devices (not shown). The system agent core 1510 provides management functionality for the various processor components and may include the at least one memory controller 1514.

[0197] For example, one or more of the processor cores 1502A-1502N may include support for simultaneous multithreading. The system agent core 1510 includes components for coordinating and operating cores 1502A-1502N during multithreaded processing. The system agent core 1510 may additionally include a power control unit (PCU) that includes logic and components for regulating the power state of the processor cores 1502A-1502N and a graphics processor 1508.

[0198] Processor 1500 may additionally include a graphics processor 1508 for performing graphics processing operations. In some embodiments, graphics processor 1508 is coupled to the shared cache unit(s) 1506 and system agent core 1510, including at least one memory controller 1514. System agent core 1510 may also include a display controller 1511 for directing the graphics processor output to one or more coupled displays. Display controller 1511 may also be a separate module coupled to the graphics processor via at least one interconnect or may be integrated into graphics processor 1508.

[0199] A ring- or mesh-based interconnect 1512 may be used to couple the internal components of processor 1500. However, an alternative interconnect unit may also be used, such as a point-to-point interconnect, a switched interconnect, or other techniques, including techniques well known in the art. In some embodiments with a ring- or mesh-based interconnect 1512, graphics processor 1508 is coupled to ring- or mesh-based interconnect 1512 via an I / O connection 1513.

[0200] The example I / O link 1513 represents at least one of several variations of I / O interconnects, including a package I / O interconnect that enables communication between various processor components and a high-performance memory module 1518, such as an embedded DRAM (eDRAM) module or a high-bandwidth memory (HBM) module. Optionally, each of the processor cores 1502A-1502N and the graphics processor 1508 may use the high-performance memory module 1518 as unified memory and / or shared last-level cache when a DRAM memory system is also present. Optionally, the processor 1500 may also include one or more accelerators 1515, e.g., an NPU for accelerating certain neural network operations.The NPU may enable inference operations with lower power consumption compared to using the graphics processor 1508, or may cooperate with the graphics processor 1508 to enable higher inference performance compared to the graphics processor 1508 alone. In one embodiment, an NPU in the one or more accelerators 1515 may include matrix or tensor acceleration logic and be used to implement at least some of the computational operations described herein that are implementable via the graphics processor 1508.

[0201] For example, processor cores 1502A-1502N may be homogeneous cores executing the same instruction set architecture. Alternatively, processor cores 1502A-1502N are heterogeneous in instruction set architecture (ISA), where one or more of processor cores 1502A-1502N execute a first instruction set while at least one of the other cores executes a subset of the first instruction set or a different instruction set. Processor cores 1502A-1502N may be heterogeneous in microarchitecture, where one or more relatively higher-power cores are coupled with one or more lower-power performance cores. As another example, processor cores 1502A-1502N are heterogeneous in computational performance.Furthermore, the processor 1500 may be implemented on one or more chips or chiplets, or as an integrated circuit (SoC) including, among other components, the illustrated components. The integrated circuit (SoC) may be implemented with multiple chiplets.

[0202] Fig. 15B is a block diagram of hardware logic of a graphics processor core block 1519 according to some embodiments described herein. In some embodiments, elements of Fig. 15B having the same reference numerals (or names) as the elements of any other present figure operate or function in a similar manner to those described elsewhere herein. In one embodiment, graphics processor core block 1519 is an example of a partition of a graphics processor. Graphics processor core block 1519 may be comprised within graphics processor 1508. Fig. 15A or a discrete graphics processor, parallel processor, and / or compute accelerator. A graphics processor as described herein may include multiple graphics core blocks based on target power and performance envelopes. Each graphics processor core block 1519 may include a functional block 1530 coupled to multiple graphics cores 1521A-1521F, which may include modular blocks with fixed-function logic and programmable general-purpose logic. The graphics processor core block 1519 also includes a shared cache 1536 accessible by all graphics cores 1521A-1521F, rasterization logic 1537, and additional fixed-function logic 1538.

[0203] In some embodiments, functional block 1530 includes a geometry / fixed function pipeline 1531 that may be shared by all graphics cores in graphics processor core block 1519. In various embodiments, geometry / fixed function pipeline 1531 includes a 3D geometry pipeline, a video front-end unit, a thread spawner, a global thread dispatcher, and a unified return buffer manager that manages unified return buffers. In one embodiment, functional block 1530 further includes a graphics SoC interface 1532, a graphics microcontroller 1533, and a media pipeline 1534. Graphics SoC interface 1532 provides an interface between graphics processor core block 1519 and other core blocks in a graphics processor or compute accelerator SoC.Graphics microcontroller 1533 is a programmable subprocessor that is configurable to manage various functions of graphics processor core block 1519, including thread dispatch, scheduling, and preemption. Media pipeline 1534 includes logic to facilitate decoding, encoding, preprocessing, and / or post-processing of multimedia data, including image and video data. Media pipeline 1534 implements media operations via requests to compute or sampling logic in graphics cores 1521-1521F. One or more pixel backends 1535 may also be integrated into functional block 1530. The one or more pixel backends 1535 include a cache for storing pixel color values ​​and can perform blending operations and lossless color compression of rendered pixel data.

[0204] In one embodiment, graphics SoC interface 1532 enables graphics processor core block 1519 to communicate with general-purpose application processor cores (e.g., CPUs) and / or other components within a SoC or a system host CPU coupled to the SoC via a peripheral device interface. Graphics SoC interface 1532 also enables communication with off-chip memory hierarchy elements such as a shared last-level cache, system RAM, and / or embedded on-chip or on-package DRAM. Graphics SoC interface 1532 may also enable communication with fixed-function devices within the SoC, such as camera image pipelines, and enables the use of and / or implements global memory atomics that may be shared between graphics processor core block 1519 and CPUs within the SoC.The graphics SoC interface 1532 may also implement power management controls for the graphics processor core block 1519 and enable an interface between a clock domain of the graphics processor core block 1519 and other clock domains within the SoC. In one embodiment, the graphics SoC interface 1532 enables receipt of command buffers from a command streamer and a global thread dispatcher configured to deliver commands and instructions to each of one or more graphics cores within a graphics processor. The commands and instructions may be sent to the media pipeline 1534 when media operations are to be performed and to the geometry and fixed function pipeline 1531 when graphics processing operations are to be performed.When computational operations need to be performed, computational distribution logic can send the instructions to the 1521A-1521F graphics cores, bypassing the geometry and media pipelines.

[0205] The graphics microcontroller 1533 may be configured to perform various scheduling and management tasks for the graphics processor core block 1519. In one embodiment, the graphics microcontroller 1533 may perform graphics and / or compute load scheduling on the various vector engines 1522A-1522F, 1524A-1524F and matrix engines 1523A-1523F, 1525A-1525F within the graphics cores 1521A-1521F. In this scheduling model, host software executing on a CPU core of an SoC including the graphics processor core block 1519 may deliver workloads to one of multiple graphics processor doorbells, enabling a scheduling operation on the corresponding graphics engine.Scheduling operations include determining which workload to execute next, submitting a workload to a command streamer, preemption of existing workloads running on an engine, monitoring the progress of a workload, and notifying host software when a workload is complete. In one embodiment, graphics microcontroller 1533 may also enable low-power or idle states for graphics processor core block 1519 and provide graphics processor core block 1519 with the ability to save and restore registers within graphics processor core block 1519 across low-power state transitions independent of the operating system and / or graphics driver software on the system.

[0206] Graphics processor core block 1519 may include more or fewer than the illustrated graphics cores 1521A-1521F, up to N modular graphics cores. For each set of N graphics cores, graphics processor core block 1519 may further include shared memory / cache 1536, which may be configured as shared memory or cache, rasterization logic 1537, and additional fixed-function logic 1538 for accelerating various graphics and compute processing operations.

[0207] Within each graphics core 1521A-1521F, there is a set of execution resources that can be used to perform graphics, media, and compute operations in response to requests from the graphics pipeline, media pipeline, or shader programs. The graphics cores 1521A-1521F include multiple vector engines 1522A-1522F, 1524A-1524F, matrix acceleration units 1523A-1523F, 1525A-1525D, cache / shared local memory (SLM), a sampler 1526A-1526F, and a ray tracing unit 1527A-1527F.

[0208] The 1522A-1522F and 1524A-1524F vector engines are general-purpose graphics processing units capable of performing floating-point and integer / fixed-point logic operations in service of graphics, media, or compute operations, including graphics, media, or compute / GPGPU programs. The 1522A-1522F and 1524A-1524F vector engines can operate with variable vector width in SIMD, SIMT, or SIMT+SIMD execution modes. Matrix acceleration units 1523A-1523F, 1525A-1525D include matrix-to-matrix and matrix-to-vector acceleration logic that improves performance for matrix operations, particularly low- and mixed-precision matrix operations (e.g., INT8, FP16, BF16, FP8, FP4) used for machine learning. In one embodiment, matrix acceleration units 1523A-1523F, 1525A-1525D support microscale (MX) formats.In one embodiment, each of the matrix acceleration units 1523A-1523F, 1525A-1525D includes one or more systolic arrays of processing elements that can simultaneously perform matrix multiplication or dot product operations on matrix elements.

[0209] The sampler 1526A-1526F can read media or texture data into memory and can sample data differently based on a configured sampler state and the texture / media format being read. Threads executing on the vector engines 1522A-1522F, 1524A-1524F or matrix acceleration units 1523A-1523F, 1525A-1525D can utilize the cache / SLM 1528A-1528F within each of the graphics cores 1521A-1521F. The cache / SLM 1528A-1528F can be configured as a cache or as a pool of shared memory local to each of the respective graphics cores 1521A-1521F.Ray tracing units 1527A-1527F within graphics cores 1521A-1521F include ray traversal / ray intersection circuitry for performing ray traversal using bounding volume hierarchies (BVHs) and identifying intersections between rays and primitives contained within the BVH volumes. In one embodiment, ray tracing units 1527A-1527F include circuitry for performing depth probing and culling (e.g., using a depth buffer or similar arrangement). In one implementation, the ray tracing units 1527A-1527F perform traversal and intersection operations consistent with image denoising, at least a portion of which may be performed using an associated matrix acceleration unit 1523A-1523F, 1525A-1525D.

[0210] Fig. 15C is a block diagram of a general-purpose graphics processing unit (GPGPU 1570) that may be configured as a graphics processor, e.g., graphics processor 1508, and / or a compute accelerator, according to embodiments described herein. GPGPU 1570 may be connected to host processors (e.g., one or more CPUs 1546) and memory 1571, 1572 via one or more system and / or memory buses. Memory 1571 may be system memory shared with the one or more CPUs 1546, while memory 1572 is device memory dedicated to GPGPU 1570. For example, components within GPGPU 1570 and memory 1572 may be mapped to memory addresses accessible to the one or more CPUs 1546. Access to memories 1571 and 1572 can be facilitated via a memory controller 1568.The memory controller 1568 may include an internal DMA controller 1569 or may include logic for performing operations that would otherwise be performed by a DMA controller. In one embodiment, at least one of the one or more CPUs 1546 may include one or more accelerators 1545, including, but not limited to, neural network accelerators.

[0211] The GPGPU 1570 includes multiple global caches, including an L2 cache 1553, an LI cache 1554, an instruction cache 1555, and a shared memory 1556, at least a portion of which may also be partitioned as cache memory. The GPGPU 1570 also includes multiple compute units 1560A-1560N. Each compute unit 1560A-1560N includes a set of vector registers 1561, scalar registers 1562, vector logic units 1563, scalar logic units 1564, and a scheduler 1584. The compute units 1560A-1560N may also include a local shared memory 1565 and a local cache 1566. The compute units 1560A-1560N may be coupled to a constant cache 1567, which may be used to store constant data, which is data that does not change during the execution of the kernel or shader program on the GPGPU 1570.Constant cache 1567 may be a scalar data cache, and cached data may be fetched directly into scalar registers 1562. In one embodiment, compute units 1560A-1560N additionally include at least one matrix unit 1580 for accelerating matrix, tensor, or AI operations and at least one ray tracing unit (RT unit 1582) for accelerating ray tracing operations. The at least one matrix unit 1580 and the RT unit 1582 may have similar functionality to other matrix / tensor accelerators and ray tracing accelerators described herein.

[0212] During operation, the one or more CPUs 1546 may write instructions to registers or memory in the GPGPU 1570 that has been mapped into an accessible address space. The instruction processors 1557 may read the instructions from registers or memory and determine how to process those instructions within the GPGPU 1570. A thread dispatcher 1558 may then be used to dispatch threads to the compute units 1560A-1560N to execute those instructions. Each compute unit 1560A-1560N may execute threads independently of the other compute units. Additionally, each compute unit 1560A-1560N may be independently configured for a conditional computation and may conditionally output the results of the computation to memory. The instruction processors 1557 may interrupt the one or more CPUs 1546 when the dispatched instructions are complete.

[0213] Fig. 16 is a block diagram of a graphics processor 1600, which may be a discrete graphics processing unit or may be a graphics processor integrated with a plurality of processing cores or other semiconductor devices, such as, but not limited to, memory devices or network interfaces. Elements of graphics processor 1600 having the same or similar names as elements of any other present figure describe the same elements as in the other figures, may operate or function in a similar manner, may include the same components, and may be associated with entities other than those described elsewhere herein, but are not limited to them. For example, graphics processor 1600 may be a variant of graphics processor 1508 and may be used in place of graphics processor 1508.The graphics processor may communicate with registers on the graphics processor and with instructions placed in processor memory via a memory-mapped I / O interface. Graphics processor 1600 may include a memory interface 1614 for accessing local memory, one or more internal caches, one or more shared external caches, and / or system memory.

[0214] The graphics processor 1600 may include a display controller 1602 to control display output data to a display device 1618. The display controller 1602 includes hardware for one or more overlay layers for displaying and compositing multiple layers of video or user interface elements. The display device 1618 may be an internal or external display device. In one embodiment, the display device 1618 is a head-mounted display, such as a VR or AR display device. The graphics processor 1600 may include a video codec engine 1606 for encoding, decoding, or transcoding media to, from, or between one or more media coding formats, including, but not limited to, MPEG (Moving Picture Experts Group) formats such as MPEG-2, AVC (Advanced Video Coding) formats such as H.264 / MPEG-4 AVC, H.265 / HEVC, Alliance for Open Media (AOMedia) VP8, VP9 sowie die SMPTE- (Society of Motion Picture & Television Engineers) 421M / VC-1- und JPEG- (Joint Photographic Experts Group) Formate wie JPEG-, und Motion-JPEG- (MJPEG-) Formate.

[0215] The graphics processor 1600 may include a block image transfer (BLIT) engine 1603 for performing two-dimensional (2D) rasterizer operations, including, for example, bit boundary block transfers, or 2D graphics operations may be performed using one or more components of the graphics processing engine (GPE 1610). The GPE 1610 may include a 3D pipeline 1612 for performing 3D operations, such as rendering three-dimensional images and scenes using processing functions that act on 3D primitives (e.g., rectangle, triangle, etc.). The 3D pipeline 1612 includes programmable and fixed-function elements that perform various tasks within the element and / or spawn execution threads to a 3D / media subsystem 1615.While the 3D pipeline 1612 may be used to perform media operations, the GPE 1610 may also include a media pipeline 1616 specifically used to perform media operations such as image or video decoding, encoding, post-processing, and enhancement. The media pipeline 1616 may include fixed-function or programmable logic units to perform one or more specialized media operations, such as video decoding acceleration, video deinterleaving, and video encoding acceleration, instead of or on behalf of the video codec engine 1606. The media pipeline 1616 may also include a thread spawn unit to spawn threads for execution on the 3D / media subsystem 1615. The spawned threads perform computations for the media operations on one or more graphics execution units included in the 3D / media subsystem 1615.

[0216] Fig. 17A illustrates a graphics processor 1720 that is a variant of the graphics processor 1600 of Fig. 16 and may be used instead of graphics processor 1600, and vice versa. Therefore, the disclosure of any features in combination with graphics processor 1600 herein also discloses, but is not limited to, a corresponding combination with graphics processor 1720. Graphics processor 1720, according to embodiments described herein, has a tiled architecture. Graphics processor 1720 may include a graphics processing engine cluster 1722 having multiple graphics processing engines within a plurality of graphics engine tiles. Each graphics engine tile 1710A-1710D may be interconnected via a set of tile interconnects 1723A-1723F. Each graphics engine tile 1710A-1710D may also be connected to a memory module or memory devices 1726A-1726D via memory interconnects 1725A-1725D. The memory devices 1726A-1726D may use any graphics memory technology.The memory devices 1726A-1726D may be, for example, graphics double data rate (GDDR) memory. The memory devices 1726A-1726D may be high bandwidth memory (HBM) modules that may be located on a die with their respective graphics engine tile 1710A-1710D. The memory devices 1726A-1726D may be stacked memory devices that may be stacked on their respective graphics engine tile 1710A-1710D. Each graphics engine tile 1710A-1610D and the associated memory devices 1626A-1626D may be located on separate chiplets bonded to a base die or substrate, as shown in FIG. Fig. 25A-25B are described in more detail.

[0217] The graphics processor 1720 may be configured with a non-uniform memory access (NUMA) system in which the memory devices 1726A-1726D are coupled to associated graphics engine tiles 1710A-1710D. A particular memory device may be accessed from graphics engine tiles other than the tile to which it is directly connected. However, access latency to the memory devices 1726A-1726D may be lowest when accessing a local tile. In one embodiment, a cache-coherent NUMA (ccNUMA) system is enabled, which uses the tile interconnects 1723A-1723F to enable communication between cache controllers within the graphics engine tiles 1710A-1710D to maintain a consistent memory map when more than one cache stores the same memory location.

[0218] The graphics processing engine cluster 1722 may be connected to an interconnect fabric 1724, which may be connected to an on-chip or on-package fabric. In one embodiment, the interconnect fabric 1724 includes a network processor, a network on a chip (NoC), or other switching processor to enable the interconnect fabric 1724 to function as a packet-switched interconnect fabric that switches data packets between components of the graphics processor 1720. The interconnect fabric 1724 may support communication between the graphics engine tiles 1710A-1710D and components such as the video codec engine 1706 and one or more copy engines 1704. The one or more copy engines 1704 may be used to move data from, to, and between the storage devices 1726A-1726D and from, to, and between memory external to the graphics processor 1720 (e.g., system memory).The interconnect fabric 1724 may also be used to connect the graphics engine tiles 1710A-1710D. The graphics processor 1720 may optionally include a display controller 1702 to enable connection to a display device 1718. The graphics processor may also be configured as a graphics or compute accelerator. In the accelerator configuration, the display controller 1702 and the display device 1718 may be omitted.

[0219] The graphics processor 1720 may be connected to a host system via a host interface 1728. The host interface 1728 may enable communication between the graphics processor 1720, system memory, and / or other system components. The host interface 1728 may be, for example, a PCI Express bus or another type of host system interface. The host interface 1728 may be, for example, an NVLink or NVSwitch interface. The host interface 1728 and the interconnect fabric 1724 may cooperate to allow multiple instances of the graphics processor 1720 to act as a single logical device. The cooperation between the host interface 1728 and the interconnect fabric 1724 may also result in the individual graphics engine tiles 1710A-1710D being presented to the host system as different logical graphics devices.

[0220] Fig. 17B illustrates a computation accelerator 1730 according to embodiments described herein. The computation accelerator 1730 may have architectural similarities to the graphics processor 1720 of Fig. 17B and is optimized for compute acceleration. A compute engine cluster 1732 may include a set of compute engine tiles 1740A-1740D that include execution logic optimized for parallel or vector-based general-purpose compute operations. In one embodiment, compute accelerator 1730 may be configured as a Kr accelerator or an NPU. In such an embodiment, the execution logic of compute engine tiles 1740A-1740D may be primarily focused on matrix or tensor operations and may include the tensor cores or matrix engines described herein. Compute engine tiles 1740A-1740D may not include fixed-function graphics processing logic, although in some embodiments, one or more of compute engine tiles 1740A-1740D may include logic to perform media acceleration.Compute engine tiles 1740A-1740D may be connected to memory devices 1726A-1726D via memory interconnects 1725A-1725D. Memory devices 1726A-1726D and memory interconnects 1725A-1725D may be a similar technology to that used in graphics processor 1720 or a different technology. Compute engine tiles 1740A-1740D may also be interconnected via a set of tile interconnects 1723A-1723F and may be connected to and / or interconnected by interconnect fabric 1724. In one embodiment, compute accelerator 1730 includes a large L3 cache 1736, which may be configured as a device-wide cache. The compute accelerator 1730 may also be accessed via a host interface 1728 in a manner similar to the graphics processor 1720 of FIG. Fig. 17B be connected to a host processor and a memory.

[0221] The compute accelerator 1730 may also include an integrated network interface 1742. In one embodiment, the integrated network interface 1742 includes a network processor and control logic that enables the compute engine cluster 1732 to communicate over a physical layer interconnect 1744 without requiring data to traverse a host system's memory. In one embodiment, one of the compute engine tiles 1740A-1740D is replaced with network processor logic, and the data to be transmitted or received over the physical layer interconnect 1744 may be transmitted directly to or from the storage devices 1726A-1726D. Multiple instances of the compute accelerator 1730 may be interconnected into a single logical device over the physical layer interconnect 1744.Alternatively, the various compute engine tiles 1740A-1740D may also be represented as different network-accessible compute accelerator devices. Graphics processing resources

[0222] Fig. 18A-18C illustrate execution logic including an array of processing elements employed in a graphics processor, according to embodiments described herein. Fig. 18A illustrates a graphics core cluster according to one embodiment. Fig. 18B illustrates a vector engine of a graphics core according to one embodiment. Fig. 18C illustrates a matrix engine of a graphics core according to one embodiment. Elements of Fig. 18A-18C having the same reference numerals as the elements of any other present figure may operate or function in any manner similar to that described elsewhere herein, but are not limited thereto. For example, the elements of Fig. 18A-18C in the context of the graphics processor core block 1519 from Fig. 15B. In one embodiment, the elements of Fig. 18A-18C have similar functionality to corresponding components of the graphics processor 1508 from Fig. 15A or the GPGPU 1570 Fig. 15C.

[0223] As in Fig. 18A, in one embodiment, the graphics core cluster 1800 includes a graphics processor core block 1519, which may include any number of graphics cores (e.g., graphics core 1815A, graphics core 1815B, through graphics core 1815N). Multiple instances of the graphics processor core block 1519 may be included. In one embodiment, the elements of graphics cores 1815A-1815N have similar or equivalent functionality to the elements of graphics cores 1521A-1521F. Fig. 15B. In such an embodiment, the graphics cores 1815A-1815N each include circuitry including, but not limited to, vector engines 1802A-1802N, matrix engines 1803A-1803N, memory load / store units 1804A-1804N, instruction caches 1805A-1805N, data caches / shared local memory 1806A-1806N, ray tracing units 1808A-1808N, and samplers 1810A-1810N. The circuitry of the graphics cores 1815A-1815N may additionally include fixed function logic 1812A-1812N. The number of 1802A-1802N vector engines and 1803A-1803N matrix engines in a design's 1815A-1815N graphics cores can vary depending on the design's workload, performance, and power requirements.

[0224] With reference to the graphics core 1815A, the vector engine 1802A and the matrix engine 1803A are configurable to perform parallel computation operations on data in a variety of integer and floating-point data formats based on instructions associated with shader programs. Each vector engine 1802A and each matrix engine 1803A can function as a programmable general-purpose compute unit capable of executing multiple concurrent hardware threads, processing multiple data elements for each thread in parallel. The vector engine 1802A and the matrix engine 1803A support variable-width vector processing at various SIMD widths, including but not limited to SIMD8, SIMD16, and SIMD32.Input data elements can be stored as a packed data type in a register, and the 1802A vector engine and 1803A matrix engine can process the different elements based on the data size of the elements. For example, when processing a 256-bit wide vector, the bits of the vector are stored in a register, and the vector is processed as four separate packed 64-bit data elements (quad-word (QW) data elements), eight separate packed 32-bit data elements (double-word (DW) data elements), sixteen separate packed 16-bit data elements (word (W) data elements), or thirty-two separate 8-bit data elements (byte (B) data elements). However, other vector widths and register sizes are also possible. In one embodiment, the vector engine 1802A and the matrix engine 1803A are also configured to perform a SIMT operation on warps or thread groups of different sizes (e.g.,8, 16 or 32 threads) configurable.

[0225] Continuing with the graphics core 1815A, the memory load / store unit 1804A services memory access requests issued by the vector engine 1802A, the matrix engine 1803A, and / or other components of the graphics core 1815A that have access to memory. The memory access request may be processed by the memory load / store unit 1804A to load or store the requested data from the cache or memory into a register file associated with the vector engine 1802A and / or the matrix engine 1803A. The memory load / store unit 1804A may also perform prefetching operations. With additional reference to Fig. 19, the memory load / store unit 1804A is configured to provide SIMT scatter / gather prefetching or block prefetching for data from the memory 1910, from memory local to other tiles via the tile interconnect 1908, or from system memory. Prefetching may occur into a particular L1 cache (e.g., data cache / shared local memory 1806A), the L2 cache 1904, or the L3 cache 1906. In one embodiment, a prefetch into the L3 cache 1906 automatically results in the data being stored in the L2 cache 1904.

[0226] Instruction cache 1805A stores instructions to be executed by graphics core 1815A. In one embodiment, graphics core 1815A further includes instruction fetch and prefetch circuitry that fetches or prefetches instructions into instruction cache 1805A. Graphics core 1815A also includes instruction decode logic to decode instructions in instruction cache 1805A. Data cache / shared local memory 1806A may be configured as a data cache managed by a cache controller implementing a cache replacement policy and / or as explicitly managed shared memory. Ray tracing unit 1808A includes circuitry for accelerating ray tracing operations. Sampler 1810A provides texture sampling for 3D operations and media sampling for media operations.Fixed-function logic 1812A includes fixed-function circuitry shared by the various instances of vector engine 1802A and matrix engine 1803A. Graphics cores 1815B-1815N may operate in a similar manner to graphics core 1815A.

[0227] Functionality of instruction caches 1805A-1805N, data caches / shared local memory 1806A-1806N, ray tracing units 1808A-1808N, samplers 1810A-1810N, and fixed function logic 1812A-1812N corresponds to equivalent functionality in the graphics processor architectures described herein. For example, instruction caches 1805A-1805N can be accessed in a manner similar to instruction cache 1855 of Fig. 15C. The data caches / shared local memory 1806A-1806N, the ray tracing units 1808A-1808N, and the samplers 1810A-1810N may be configured in a similar manner to the cache / SLM 1528A-1528F, the ray tracing units 1527A-1527F, and the samplers 1526A-1526F of Fig. 15B. The fixed function logic 1812A-1812N may include elements of the geometry / fixed function pipeline 1531 and / or the additional fixed function logic 1538 from Fig. 15B. In one embodiment, the ray tracing units 1808A-1808N include circuitry for performing ray tracing acceleration operations performed by the ray tracing cores 372 of Fig. 3 be carried out.

[0228] As in Fig. 18B, in one embodiment, the vector engine 1802 includes an instruction fetch unit 1837, a general purpose register file array (GRF 1824), an architectural register file array (ARF 1826), a thread arbiter 1822, a dispatch unit 1830, a branch unit 1832, SIMD FPUs 1834, and, in one embodiment, SIMD ALUs 1835. The GRF 1824 and the ARF 1826 comprise the set of general purpose register files and architectural register files associated with each hardware thread that may be active in the vector engine 1802. In one embodiment, per-thread architectural state is maintained in the ARF 1826, while data used during thread execution is stored in the GRF 1824. The execution state of each thread, including the instruction pointers for each thread, can be held in thread-specific registers in the ARF 1826.Register renaming can be used to dynamically assign registers to hardware threads.

[0229] In one embodiment, the vector engine 1802 has an architecture that is a combination of simultaneous multi-threading (SMT) and fine-grained interleaved multi-threading (IMT). The architecture has a modular configuration that can be fine-tuned at design time based on the targeted number of concurrent threads and the number of registers per graphics core, with the graphics core resources being divided among the logic for executing multiple concurrent threads. The number of logical threads that can be executed by the vector engine 1802 is not limited to the number of hardware threads, and each hardware thread can be assigned multiple logical threads.

[0230] In one embodiment, the vector engine 1802 may issue multiple instructions together, each of which may be different instructions. The thread arbiter 1822 may forward the instructions to one of the dispatch unit 1830, the branch unit 1832, or the SIMD FPUs 1834 for execution. Each execution thread may access 128 general-purpose registers within the GRF 1824, where each register may store 32 bytes accessible as a variable-width vector of 32-bit data elements. In one embodiment, each thread has access to 4 KB within the GRF 1824, although embodiments are not so limited, and more or fewer register resources may be provided in other embodiments.In one embodiment, the vector engine 1802 is divided into seven hardware threads that can perform computational operations independently, although the number of threads per vector engine 1802 can vary depending on the embodiment. For example, in one embodiment, up to 16 hardware threads are supported. In an embodiment where seven threads can access 4 kilobytes, the GRF 1824 can store a total of 28 kilobytes. While 16 threads can access 4 kilobytes, the GRF 1824 can store a total of 64 kilobytes. Flexible addressing modes can allow registers to be addressed together to effectively build wider registers or to represent striped rectangular block data structures.

[0231] In one embodiment, memory operations, sampler operations, and other longer-latency system communications are dispatched via "send" instructions executed by message forwarding unit 1830. In one embodiment, branch instructions are dispatched to branch unit 1832 to enable SIMD divergence and eventual convergence.

[0232] In one embodiment, the SIMD FPUs 1834 of the vector engine 1802 perform floating-point operations. In one embodiment, the SIMD FPUs 1834 also support integer computation. In one embodiment, the SIMD FPUs 1834 can perform up to M 32-bit floating-point (or integer) operations or up to 2M 16-bit integer or 16-bit floating-point operations. In one embodiment, at least one of the FPUs provides enhanced computation capability to support high-throughput transcendental computation functions and 64-bit double-precision floating-point functions. In some embodiments, SIMD ALUs 1835 are also present, which are configured to perform 8-bit integer operations and may be specifically optimized to perform operations related to machine learning computations.In one embodiment, the SIMD ALUs 1835 are replaced by SIMD FPUs 1834, which can be configured to perform integer and floating-point operations. In one embodiment, the SIMD FPUs 1834 and SIMD ALUs 1835 are configurable to execute SIMT programs. In one embodiment, combined SIMD+SIMT operation is supported.

[0233] In one embodiment, arrays of multiple instances of the vector engine 1802 may be instantiated within a graphics core. To ensure scalability, product architects may choose the exact number of vector engines per graphics core grouping. In one embodiment, the vector engine 1802 may execute instructions across a plurality of execution channels. In another embodiment, each thread executed by the vector engine 1802 executes on a different channel.

[0234] As in Fig. 18C, in one embodiment, the matrix engine 1803 includes an array of processing elements configured to perform tensor operations, including, but not limited to, vector-matrix and matrix-matrix operations, such as matrix multiplication and / or dot product operations. The matrix engine 1803 is configured with M rows and N columns of processing elements 1852AA-1852MN, which include multiplier and adder circuits, organized in a pipeline. In one embodiment, the processing elements 1852AA-1852MN form the physical pipeline stages of an N-wide and M-deep systolic array that can be used to perform vector / matrix or matrix / matrix operations in a data-parallel manner, including matrix multiplication, fused multiplication-addition, dot product, or other general purpose matrix-matrix multiplication (GEMM) operations.In one embodiment, the matrix engine 1803 supports 16-bit and 8-bit floating-point operations, as well as 8-bit, 4-bit, 2-bit, and binary integer operations. The matrix engine 1803 may also be configured to accelerate certain machine learning operations. In such embodiments, the matrix engine 1803 may be configured to support the 16-bit bfloat (brain floating point) floating-point format or a 32-bit tensor float (TF32) floating-point format, which have a different number of mantissa and exponent bits compared to the IEEE (Institute of Electrical and Electronics Engineers) 754 formats.

[0235] In one embodiment, during each cycle, each stage may add the result of operations performed in that stage to the output of the previous stage. In other embodiments, the pattern of data movement between processing elements 1852AA-1852MN after a series of compute cycles may vary based on the instruction or macro operation being executed. For example, in one embodiment, partial sum feedback is enabled, and processing elements may instead add the output of a current cycle to the output generated in the previous cycle. In one embodiment, the last stage of the systolic array may be configured with feedback to the first stage of the systolic array. In such an embodiment, the number of physical pipeline stages may be decoupled from the number of logical pipeline stages supported by matrix engine 1803.For example, if the processing elements 1852AA-1852MN are configured as a systolic array of M physical stages, a flyback loop from stage M to the initial pipeline stage may enable the processing elements 1852AA-1852MN to operate as a systolic array of, for example, 2M, 3M, 4M, etc., logical pipeline stages.

[0236] In one embodiment, matrix engine 1803 includes memory 1841A-1841N, 1842A-1842M for storing input data in the form of row and column data for input matrices. Memories 1842A-1842M are configurable to store row elements (A0-Am) of a first input matrix, and memories 1841A-1841N are configurable to store column elements (B0-Bn) of a second input matrix. The row and column elements are provided to processing elements 1852AA-1852MN for processing. In one embodiment, the row and column elements of the input matrices may be stored in a systolic register file 1840 within matrix engine 1803 before these elements are provided to memory 1841A-1841N, 1842A-1842M. In one embodiment, the systolic register file 1840 is excluded and the memory 1841A-1841N, 1842A-1842M is generated from registers in an associated vector engine (e.g.,GRF 1824 of the vector engine 1802 from . Fig. 18B) or another memory of the graphics core that includes the matrix engine 1803 (e.g., a data cache / shared local memory 1806A for the matrix engine 1803A from Fig. 18A). Results generated by processing elements 1852AA-1852MN are then output to an output buffer and / or written to a register file (e.g., systolic register file 1840, GRF 1824, data cache / shared local memory 1806A-1806N) for further processing by other functional units of the graphics processor or for output to memory.

[0237] In some embodiments, the matrix engine 1803 is configured with support for input sparsity, where multiplication operations for sparse regions of input data can be bypassed by skipping multiplication operations that have a zero-valued operand. In one embodiment, the processing elements 1852AA-1852MN are configured to skip performing certain operations on zero-valued inputs. In one embodiment, sparse input matrices can be detected and operations with known zero output values ​​can be bypassed before being passed to the processing elements 1852AA-1852MN. Loading zero-valued operands into the processing elements can be bypassed, and the processing elements 1852AA-1852MN can be configured to perform multiplications on the non-zero-valued input elements.Matrix engine 1803 may also be configured to support sparse output, thus bypassing operations with results predetermined to be zero. In one embodiment, metadata is provided to processing elements 1852AA-1852MN to indicate, for a processing cycle, which processing elements and / or data channels should be active during that cycle.

[0238] In one embodiment, the matrix engine 1803 includes hardware to enable operations on sparse data with a compressed representation of a sparse matrix that stores non-zero values, and metadata that identifies the positions of the non-zero values ​​within the matrix. Example compressed representations include, but are not limited to, compressed tensor representations such as compressed sparse row (CSR), compressed sparse column (CSC), and compressed sparse fiber (CSF) representations. Support for compressed representations allows operations to be performed on an input in a compressed tensor format without requiring decompression or decoding of the compressed representation. In such an embodiment, operations may be performed on non-zero input values, and the resulting non-zero output values ​​may be mapped into an output matrix.In some embodiments, hardware support is also provided for machine-specific lossless data compression formats used when transferring data within the hardware or across system buses. Such data may be maintained in a compressed format for sparse input data, and the matrix engine 1803 may use the compression metadata for the compressed data to enable operations on non-zero values ​​or to bypass blocks of zero data input for multiplication operations.

[0239] In various embodiments, input data may be provided by a programmer in a compressed tensor representation, or a codec may compress input data into the compressed tensor representation or another sparse data encoding. In addition to supporting compressed tensor representations, streaming compression of sparse input data may be performed before the data is provided to processing elements 1852AA-1852MN. In one embodiment, compression is performed on data written to a cache associated with graphics core cluster 1800, where the compression is performed using an encoding supported by matrix engine 1803. In one embodiment, matrix engine 1803 supports structured sparse inputs, where a predetermined degree or pattern of sparsity is imposed on the input data.This data may be compressed to a known compression ratio, with the compressed data being processed by the processing elements 1852AA-1852MN according to metadata associated with the compressed data.

[0240] Fig. Figure 19 illustrates a tile 1900 of a multi-tile processor according to one embodiment. In one embodiment, tile 1900 represents one of the graphics engine tiles 1710A-1710D of Fig. 17A or compute engine tiles 1740A-1740D from Fig. 17B. The multi-tile graphics processor tile 1900 includes an array of graphics core clusters (e.g., graphics core cluster 1800A, graphics core cluster 1800B, through graphics core cluster 1800N), each graphics core cluster including an array of graphics cores 515A-515N. The tile 1900 also includes a global dispatcher 1902 that dispatches threads to processing resources of the tile 1900.

[0241] Tile 1900 may include or be coupled to an L3 cache 1906 and a memory 1910. In various embodiments, L3 cache 1906 may be omitted, or tile 1900 may include additional cache levels, such as an L4 cache. In one embodiment, each instance of tile 1900 in the multi-tile graphics processor is associated with memory 1910, as shown in Fig. 17A and Fig. 17B. In one embodiment, a multi-tile processor may be configured as a multi-chip module in which the L3 cache 1906 and / or the memory 1910 are located on different chiplets than the graphics core clusters 1800A-1800N. In this context, a chiplet is an at least partially packaged integrated circuit comprising individual logic units that may be assembled with other chiplets into a larger package. For example, the L3 cache 1906 may be included in a dedicated cache chiplet or may be located on the same chiplet as the graphics core clusters 1800A-1800N. In one embodiment, the L3 cache 1906 may be included in an active base die or an active interposer.

[0242] A memory fabric 1903 enables communication between the graphics core clusters 1800A-1800N, the L3 cache 1906, and the memory 1910. An L2 cache 1904 is coupled to the memory fabric 1903 and can be configured to cache transactions performed via the memory fabric 1903. A tile interconnect 1908 enables communication with other tiles on the graphics processors and can be assigned to one of the tile interconnects 1723A-1723F of Fig. 17A and Fig. 17B. In embodiments where L3 cache 1906 is omitted from tile 1900, L2 cache 1904 may be configured as a combined L2 / L3 cache. Memory fabric 1903 may be configured to forward data to L3 cache 1906 or to memory controllers associated with memory 1910, depending on whether or not L3 cache 1906 is present in a particular implementation. L3 cache 1906 may be configured as a per-tile cache dedicated to processing resources of tile 1900, or it may be a partition of a GPU-wide L3 cache.

[0243] Fig. 20 is a block diagram illustrating the instruction formats of a graphics processor 2000. The graphics processor execution units support an instruction set with instructions in multiple formats. The solid-line boxes illustrate the components generally included in an execution unit instruction, while the dashed lines contain components that are optional or included in a subset of the instructions. In some embodiments, the graphics processor instruction format 2000 described and depicted are macroinstructions, that is, instructions that are fed to the execution unit, as opposed to microoperations that result from instruction decode once the instruction has been processed. Thus, a single instruction can cause hardware to perform multiple microoperations.

[0244] The graphics processor execution units described herein may natively support instructions in a 128-bit instruction format 2010. A compressed 64-bit instruction format 2030 is available for some instructions based on the selected instruction, instruction options, and number of operands. The native 128-bit instruction format 2010 provides access to all instruction options, while some options and operations are restricted in the 64-bit instruction format 2030. The native instructions available in the 64-bit instruction format 2030 vary depending on the embodiment. The instruction is partially compressed using a set of index values ​​in an index field 2013. The execution unit hardware references a set of compression tables based on the index values ​​and uses the outputs of the compression tables to reconstruct a native instruction in the 128-bit instruction format 2010.Other command sizes and formats can also be used.

[0245] For each format, the instruction opcode 2012 defines the operation the execution unit is to perform. The execution units execute each instruction in parallel on the multiple data elements of each operand. For example, in response to an add instruction, the execution unit performs a simultaneous add operation for each color channel representing a texture element or a picture element. By default, the execution unit executes each instruction across all data channels of the operands. An instruction control field 2014 may provide control over certain execution options, such as channel selection (e.g., predication) and data channel ordering (e.g., swizzling). For instructions in the 128-bit instruction format 2010, an execution size field 2016 limits the number of data channels executed in parallel. An execution size field 2016 may not be available for use in the condensed 64-bit instruction format 2030.

[0246] Some execution unit instructions have up to three operands, including two source operands, src0 2020, src1 2022, and one destination operand (dest 2018). Other instructions, such as data manipulation instructions, dot product instructions, multiply-add instructions, or multiply-accumulate instructions, may have a third source operand (e.g., SRC2 2024). The instruction opcode 2012 determines the number of source operands. The last source operand of an instruction may be an immediate (e.g., hard-coded) value passed with the instruction. The execution units may also support multiple destination instructions, with one or more destinations being implied or implicit based on the instruction and / or the specified destination.

[0247] The 128-bit instruction format 2010 may include an access / address mode field 2026, which indicates, for example, whether direct register addressing mode or indirect register addressing mode is used. When direct register addressing mode is used, the register address of one or more operands is provided directly by bits in the instruction.

[0248] The 128-bit instruction format 2010 may also include an access / address mode field 2026 that specifies an address mode and / or an access mode for the instruction. The access mode may be used to define a data access alignment for the instruction. Access modes including an aligned 16-bit access mode and an aligned 1-byte access mode may be supported, with the byte alignment of the access mode determining the access alignment of the instruction operands. For example, when in a first mode, the instruction may use byte-aligned addressing for source and destination operands, and when in a second mode, the instruction may use aligned 16-byte addressing for all source and destination operands.

[0249] The address mode portion of the access / address mode field 2026 can determine whether the instruction should use direct or indirect addressing. When direct register addressing mode is used, bits in the instruction directly provide the register address of one or more operands. When indirect register addressing mode is used, the register address of one or more operands can be calculated based on an address register value and an immediate address field in the instruction.

[0250] Instructions may be grouped based on bit fields of the instruction opcode 2012 to simplify opcode decoding 2040. For an 8-bit opcode, bits 4, 5, and 6 allow the execution unit to determine the type of opcode. The exact opcode grouping shown is merely an example. A move-and-logic opcode group 2042 may include data movement and logic instructions (e.g., move (mov), compare (cmp). The move-and-logic opcode group 2042 may share the five least significant bits (LSB), where move(mov) instructions are of the form 0000xxxxb and logic instructions are of the form 0001xxxxb. A flow control instruction group 2044 (e.g., call, jump (jmp)) contains instructions in the form 0010xxxxb (e.g., 0x20). A miscellaneous instruction group 2046 contains a mix of instructions, including synchronization instructions (e.g., wait, send) in the form 0011xxxxb (e.g., 0x30).A parallel math instruction group 2048 includes component-wise arithmetic instructions (e.g., add, multiply (mul)) in the form 0100xxxxb (e.g., 0x40). The parallel math instruction group 2048 performs the arithmetic operations in parallel across data channels. The vector math group 2050 includes arithmetic instructions (e.g., dp4) in the form 0101xxxxb (e.g., 0x50). The vector math group performs arithmetic such as dot product calculations on vector operands. The illustrated opcode decoding 2040 may be used in one embodiment to determine which part of an execution unit is used to execute a decoded instruction. For example, some instructions may be referred to as systolic instructions, which are executed by a systolic array. Other instructions, such asRay tracing instructions (not shown) can be forwarded to a ray tracing core or ray tracing logic within a slice or partition of the execution logic. Graphics pipeline

[0251] Fig. 21 is a block diagram of a graphics processor 2100 according to another embodiment. The elements of Fig. 21 having the same or similar names as the elements of any other figure herein describe the same elements as in the other figures, may operate or function in a similar manner, may include the same components, and may be linked to other entities such as, but are not limited to, those described elsewhere herein.

[0252] The graphics processor 2100 may include various types of graphics processing pipelines, such as a geometry pipeline 2120, a media pipeline 2130, a display engine 2140, thread execution logic 2150, and a rendering output pipeline 2170. The graphics processor 2100 may be a graphics processor within a multi-core processing system that includes one or more general-purpose processing cores. The graphics processor may be controlled by register writes to one or more control registers (not shown) or by commands issued to the graphics processor 2100 via a ring or mesh interconnect 2102. The ring or mesh interconnect 2102 may couple the graphics processor 2100 to other processing components, such as other graphics processors or general-purpose processors.Commands from the ring or mesh interconnect 2102 are interpreted by a command streamer 2103, which delivers instructions to individual components of the geometry pipeline 2120 or the media pipeline 2130.

[0253] The command streamer 2103 may direct the operation of a vertex fetcher 2105, which reads vertex data from memory and executes vertex processing commands provided by the command streamer 2103. The vertex fetcher 2105 may provide vertex data to a vertex shader 2107, which performs coordinate space transformations and lighting operations on each vertex. The vertex fetcher 2105 and the vertex shader 2107 execute vertex processing instructions by distributing execution threads to the graphics cores 2152A-2152B via a thread dispatcher 2131.

[0254] Graphics cores 2152A-2152B may be an array of vector processors with an instruction set for performing graphics and media operations. Graphics cores 2152A-2152B may have an attached L1 cache 2151 that is specific to each array or shared among the arrays. The cache may be configured as a data cache, an instruction cache, or a single cache partitioned to contain data and instructions in different partitions.

[0255] A geometry pipeline 2120 may include tessellation components to perform hardware-accelerated tessellation of 3D objects. A programmable hull shader 2111 may configure the tessellation operations. A programmable domain shader 2117 may provide backend evaluation of the tessellation output. A tesselizer 2113 may operate under the instruction of the programmable hull shader 2111 and include specialized logic for generating a set of detailed geometric objects based on a coarse geometric model provided as input to the geometry pipeline 2120. Furthermore, if tessellation is not used, tessellation components (e.g., the programmable hull shader 2111, the tesselizer 2113, and the programmable domain shader 2117) may be bypassed. The tessellation components can operate based on the data received from the vertex shader 2107.

[0256] Entire geometric objects may be processed by a geometry shader 2119 via one or more threads dispatched to graphics cores 2152A-2152B, or they may proceed directly to clipper 2129. The geometry shader may operate on entire geometric objects rather than on vertices or vertex arrays, as in previous stages of the graphics pipeline. If tessellation is disabled, geometry shader 2119 receives input from vertex shader 2107. Geometry shader 2119 may be programmable by a geometry shader program to perform geometry tessellation if the tessellation units are disabled.

[0257] Before rasterization, a clipper 2129 processes vertex data. The clipper 2129 can be a fixed-function clipper or a programmable clipper that includes cropping and geometry shader functions. A rasterization and depth inspection component 2173 in the render output pipeline 2170 can dispatch pixel shaders to convert the geometric objects into pixel-wise representations. The pixel shader logic can be included in the thread execution logic 2150. Optionally, an application can bypass the rasterization and depth inspection component 2173 and access unrasterized vertex data via an output streaming unit 2123.

[0258] The graphics processor 2100 includes an interconnect bus, interconnect fabric, or other interconnect mechanism that enables data and messages to pass between the main components of the processor. In some embodiments, the graphics cores 2152A-2152B and associated logic units (e.g., L1 cache 2151, sampler 2154, texture cache 2158, etc.) are interconnected via a data port 2156 to perform memory access and communication with the processor's render output pipeline components. A sampler 2154, an L1 cache 2151, a texture cache 2158, and graphics cores 2152A-2152B may each have their own memory access paths. Optionally, the texture cache 2158 may also be configured as a sample cache.

[0259] The rendering output pipeline 2170 may include a rasterizer and depth check component 2173 that converts vertex-based objects into an associated pixel-based representation. The rasterization logic may include a windower / masker unit to perform fixed-function triangle and line rasterization. An associated render cache 2178 and depth cache 2179 are also available in some embodiments. A pixel operations component 2177 performs pixel-based operations on the data, although in some cases, pixel operations associated with 2D operations (e.g., bit-block image transfers with blending) are performed by the 2D engine 2141 or replaced at display time by the display controller 2143 using overlay display layers. A shared L3 cache 2175 may be available to all graphics components, enabling data sharing without the use of main system memory.

[0260] The media pipeline 2130 may include a media engine 2137 and a video frontend 2134. The video frontend 2134 may receive pipeline commands from the command streamer 2103. The media pipeline 2130 may include a separate command streamer. The video frontend 2134 may process media commands before sending the command to the media engine 2137. The media engine 2137 may include thread creation functionality to create threads for dispatch to the thread execution logic 2150 via the thread dispatcher 2131.

[0261] The graphics processor 2100 may include a display engine 2140. This display engine 2140 may be external to the processor 2100 and may be coupled to the graphics processor via the ring or mesh interconnect 2102 or another interconnect bus or fabric. The display engine 2140 may include a 2D engine 2141 and a display controller 2143. The display engine 2140 may include special-purpose logic capable of operating independently of the 3D pipeline. The display controller 2143 may couple to a display device (not shown), which may be a display device integrated into the system, such as in a laptop computer, or an external display device connected via a display device connector.

[0262] The geometry pipeline 2120 and the media pipeline 2130 may be configurable to perform operations based on multiple graphics and media programming interfaces and are not specific to any one application programming interface (API). Driver software for the graphics processor may translate API calls specific to a particular graphics or media library into commands that can be processed by the graphics processor. Support for the Open Graphics Library (OpenGL), Open Computing Language (OpenCL), and / or Vulkan graphics and compute API, all from the Khronos group, may be provided. Support may also be provided for the Direct3D library from Microsoft Corporation. A combination of these libraries may be supported. Support may also be provided for the open source Computer Vision Library (OpenCV).A future API with a compatible 3D pipeline would also be supported if a mapping can be performed from the pipeline of the future API to the pipeline of the graphics processor. Graphics pipeline programming

[0263] Fig. 22A is a block diagram illustrating a graphics processor instruction format 2200 used to program graphics processing pipelines, such as those described herein in connection with Fig. 16 and Fig. 21 described pipelines. Fig. Figure 22B is a block diagram illustrating a graphics processor command sequence 2210 according to one embodiment. The solid boxes in Fig. 22A illustrate the components generally included in a graphics instruction, while the dashed lines indicate components that are optional or included in a subset of the graphics instructions. The graphics processor instruction format 2200 of Fig. 22A includes fields to identify a client 2202, an instruction operation code (opcode 2204), and a data field 2206 for the instruction. A sub-opcode 2205 and an instruction size 2208 are also included in some instructions.

[0264] The client 2202 may specify the client unit of the graphics device that will process the command data. A graphics processor command parser may examine the client field of each command to condition further processing of the command and direct the command data to the appropriate client unit. The graphics processor client units may include a memory interface unit, a rendering unit, a 2D unit, a 3D unit, and a media unit. Each client unit may have a corresponding processing pipeline that processes the commands. Once the command is received from the client unit, the client unit reads the opcode 2204 and, if present, the sub-opcode 2205 to determine the operation to be performed. The client unit executes the command using information in the data field 2206. For some commands, an instruction size 2208 is expected to explicitly specify the size of the command.The instruction parser can automatically determine the size of at least some of the instructions based on the instruction opcode. Instructions can be aligned across multiples of a double word. Other instruction formats can also be used.

[0265] The flowchart in Fig. 22B illustrates an example graphics processor command sequence 2210. Software or firmware of a data processing system including an example graphics processor may use a version of the illustrated command sequence to set up, execute, and terminate a set of graphics operations. An example command sequence is shown and described for exemplary purposes and is not limited to these specific commands or to this command sequence. Furthermore, the commands may be issued as a batch of commands in a command sequence such that the graphics processor processes the sequence of commands at least partially concurrently.

[0266] The graphics processor instruction sequence 2210 may begin with a pipeline flush instruction 2212 to cause each active graphics pipeline to complete the currently pending instructions for the pipeline. Optionally, the 3D pipeline 2222 and media pipeline 2224 may not operate concurrently. The pipeline flush is performed to cause the active graphics pipeline to complete any pending instructions. In response to a pipeline flush, the instruction parser for the graphics processor suspends instruction processing until the active draw engines complete pending operations and the relevant read caches are invalidated. Optionally, any data in the render cache marked as "dirty" may be flushed to memory. A pipeline flush instruction 2212 may be used for pipeline synchronization or before placing the graphics processor into a low-power state.

[0267] A pipeline select instruction 2213 can be used when an instruction sequence uses the graphics processor to explicitly switch between pipelines. A pipeline select instruction 2213 can be used once in an execution context before issuing any pipeline instructions, unless the context issues instructions for both pipelines. A pipeline flush instruction 2212 can be used immediately before a pipeline switch via the pipeline select instruction 2213.

[0268] A pipeline control instruction 2214 can configure a graphics pipeline for operation and can be used to program the 3D pipeline 2222 and the media pipeline 2224. The pipeline control instruction 2214 can configure the pipeline state for the active pipeline. The pipeline control instruction 2214 can be used for pipeline synchronization and for flushing data from one or more caches within the active pipeline before processing a batch of instructions.

[0269] Commands related to return buffer state 2216 can be used to configure a set of return buffers for the respective pipelines to write data. Some pipeline operations cause the allocation, selection, or configuration of one or more return buffers to which the operations write intermediate data during processing. The graphics processor may also use one or more return buffers to store output data and perform cross-thread communication. Return buffer state 2216 may include selecting the size and number of return buffers to use for a set of pipeline operations.

[0270] The remaining instructions in the instruction sequence differ based on the active pipeline for operations. Based on a pipeline determination 2220, the instruction sequence is aligned to the 3D pipeline 2222 starting at the 3D pipeline state 2230 or the media pipeline 2224 starting at the media pipeline state 2240.

[0271] The 3D pipeline state configuration commands 2230 include 3D state setting commands for vertex buffer state, vertex element state, constant color state, depth buffer state, and other state variables to be configured before processing 3D primitive commands. The values ​​of these commands are determined at least in part based on the particular 3D API being used. The 3D pipeline state commands 2230 may also be capable of selectively disabling or bypassing certain pipeline elements if those elements are not used.

[0272] A 3D primitive instruction 2232 can be used to dispatch 3D primitives to be processed by the 3D pipeline. Instructions and associated parameters passed to the graphics processor via the 3D primitive instruction 2232 are passed to the vertex fetch function in the graphics pipeline. The vertex fetch function uses the data from the 3D primitive instruction 2232 to create vertex data structures. The vertex data structures are stored in one or more return buffers. The 3D primitive instruction 2232 can be used to perform vertex operations on 3D primitives via vertex shaders. To process vertex shaders, the 3D pipeline 2222 dispatches shader execution threads to graphics processor execution units.

[0273] The 3D pipeline 2222 can be triggered via an execute instruction 2234 or an execute event. A register can write trigger instruction executions. Execution can be triggered via a 'Go' or 'Kick' instruction in the instruction sequence. Instruction execution can be triggered using a pipeline synchronization instruction to flush the instruction sequence through the graphics pipeline. The 3D pipeline performs geometry processing on the 3D primitives. Once the operations are complete, the resulting geometric objects are rasterized, and the pixel engine colors the resulting pixels. Additional instructions to control pixel shading and pixel backend operations can also be included for these operations.

[0274] The graphics processor instruction sequence 2210 may follow the path of the media pipeline 2224 when performing media operations. In general, the specific usage and type of programming for the media pipeline 2224 depends on the media or computational operations to be performed. Specific media decoding operations may be offloaded to the media pipeline during media decoding. The media pipeline may also be bypassed, and media decoding may be performed in whole or in part using resources provided by one or more general-purpose processing cores. The media pipeline may also include elements for general-purpose graphics processing unit (GPGPU) operations, where the graphics processor is used to perform SIMD vector operations using computational shader programs not explicitly related to rendering graphics primitives.

[0275] The media pipeline 2224 may be configured in a similar manner to the 3D pipeline 2222. A set of commands for configuring the media pipeline state 2240 is dispatched or placed in a command queue before the media object commands 2242. Media pipeline state 2240 commands may include data for configuring the media pipeline elements used to process the media objects. This includes data for configuring the video decoding and video encoding logic within the media pipeline, such as the encoding or decoding format. Media pipeline state 2240 commands may also support the use of one or more pointers to "indirect" state elements that contain a batch of state settings.

[0276] Media object instructions 2242 may provide pointers to media objects for processing by the media pipeline. The media objects include memory buffers containing video data to be processed. Optionally, all media pipeline states should be valid before issuing a media object instruction 2242. Once the pipeline state is configured and the media object instructions 2242 are queued, the media pipeline 2224 is triggered via an execute instruction 2244 or an equivalent execution event (for example, a register write). Output from the media pipeline 2224 may then be post-processed by operations provided by the 3D pipeline 2222 or the media pipeline 2224. GPGPU operations may be configured and executed in a similar manner to media operations. Graphics software architecture

[0277] Fig. 23 illustrates an exemplary graphics software architecture for a data processing system 2300. Such a software architecture may include a 3D graphics application 2310, an operating system 2320, and a processor 2330. The processor 2330 may include a graphics processor 2332 and one or more general-purpose processor cores 2334. The processor 2330 may be a variant of one or more of the processors 1402 or any other of the processors described herein and may be used in place thereof. Therefore, the disclosure of any features in combination with the processor 1402 or any other of the processors described herein also discloses a corresponding combination with the graphics processor 2332, but is not limited thereto. Furthermore, the elements of Fig. 23 having the same or similar names as the elements of any other figure herein may represent the same elements as in the other figures, may operate or function in a similar manner, may include the same components, and may be associated with entities other than those described elsewhere herein, but are not limited thereto. The 3D graphics application 2310 and the operating system 2320 each execute in the system memory 2350 of the data processing system.

[0278] The 3D graphics application 2310 may include one or more shader programs including shader instructions 2312. The shader language instructions may be in a higher-level shader language, such as Direct3D's High-Level Shader Language (HLSL), OpenGL Shader Language (GLSL), and so on. The application may also include execution program instructions 2314 in a machine language suitable for execution by the general-purpose processor core 2334. The application may also include graphics objects 2316 defined by vertex data.

[0279] The operating system 2320 may be a Microsoft® Windows® operating system from Microsoft Corporation, a proprietary UNIX-like operating system, or an open-source UNIX-like operating system using a variant of the Linux kernel. The operating system 2320 may support a graphics API 2322, such as the Direct3D API, the OpenGL API, or the Vulkan API. When the Direct3D API is used, the operating system 2320 uses a front-end shader compiler 2324 to compile any shader instructions 2312 in HLSL into a lower-level shader language. The compilation may be just-in-time (JIT) compilation, or the application may perform shader precompilation. Higher-level shaders may be compiled into lower-level shaders during compilation of the 3D graphics application 2310.The shader instructions 2312 may be provided in an intermediate form, such as a version of the Standard Portable Intermediate Representation (SPIR) used by the Vulkan API.

[0280] A user-mode graphics driver 2326 may include a backend shader compiler 2327 to convert the shader instructions 2312 into a hardware-specific representation. When using the OpenGL API, shader instructions 2312 in the high-level GLSL language are passed to a user-mode graphics driver 2326 for compilation. The user-mode graphics driver 2326 may use operating system kernel-mode functions 2328 to communicate with a kernel-mode graphics driver 2329. The kernel-mode graphics driver 2329 may communicate with the graphics processor 2332 to dispatch commands and instructions. IP core implementations

[0281] One or more aspects may be implemented by representative code stored on a machine-readable medium that represents and / or defines logic in an integrated circuit, such as a processor. For example, the machine-readable medium may include instructions that represent various logic within the processor. When read by a machine, the instructions may cause the machine to generate the logic to perform the techniques described herein. Such representations, known as "IP cores," are reusable logic units for an integrated circuit that may be stored on a tangible, machine-readable medium as a hardware model describing the structure of the integrated circuit.The hardware model can be delivered to various customers or manufacturing facilities, which load the hardware model onto manufacturing machines that manufacture the integrated circuit. The integrated circuit can be manufactured such that the circuit performs operations described in connection with any of the presently described embodiments.

[0282] Fig. 24 is a block diagram illustrating an IP core development system 2400 that can be used to fabricate an integrated circuit to perform operations according to one embodiment. The IP core development system 2400 can be used to create modular, reusable designs that can be integrated into a larger design or used to build an entire integrated circuit (e.g., a SoC integrated circuit). A design facility 2430 can create a software simulation 2410 of an IP core design in a high-level programming language (e.g., C / C++). The software simulation 2410 can be used to design, test, and verify the behavior of the IP core using a simulation model 2412. The simulation model 2412 can include functional, behavioral, and / or timing simulations.A register transfer level (RTL) design 2415 can then be generated or synthesized from the simulation model 2412. The RTL design 2415 is an abstraction of the integrated circuit's behavior that models the flow of digital signals between hardware registers, including the associated logic performed using the modeled digital signals. In addition to an RTL design 2415, lower-level designs at the logic level or transistor level can also be generated, designed, or synthesized. Thus, the specific details of the initial design and simulation may vary.

[0283] The RTL design 2415 or equivalent may be further synthesized by the design facility into a hardware model 2420, which may be in a hardware description language (HDL) or other representation of physical design data. The HDL may be further simulated or tested to verify the IP core design. The IP core design may be stored for delivery to a manufacturing facility 2465 using non-volatile memory 2440 (e.g., hard disk, flash memory, or any non-volatile storage medium). The manufacturing facility 2465 may be a third-party manufacturing facility. Alternatively, the IP core design may be transmitted via a wired connection 2450 or wireless connection 2460 (e.g., over the Internet). The manufacturing facility 2465 may then fabricate an integrated circuit based at least in part on the IP core design.The fabricated integrated circuit may be configured to perform operations according to at least one embodiment described herein.

[0284] Fig. 25A illustrates a cross-sectional side view of an integrated circuit package assembly 2590 including multiple units of hardware logic chiplets coupled to a substrate 2580 (e.g., a base die). A graphics processing unit, a parallel processor, and / or a compute accelerator as described herein may be composed of various silicon chiplets that are separately manufactured. In this context, a chiplet is an at least partially packaged integrated circuit comprising individual logic units that may be assembled with other chiplets into a larger package. A plurality of chiplets with different IP core logic may be combined into a single device. Furthermore, the chiplets may be integrated into a base die or a base chiplet using active interposer technology.The concepts described here enable the interconnection and communication between the various forms of IP within the GPU. IP cores can be manufactured using different process technologies and assembled during manufacturing, avoiding the complexity of converging multiple IPs, especially in a large SoC with multiple IPs in different forms, using the same manufacturing process. The ability to use multiple process technologies improves time to market and provides a cost-effective way to create multiple product SKUs. Furthermore, the disaggregated IPs can be more easily powered independently. Components not used for a given workload can be turned off, reducing overall power consumption.

[0285] In various embodiments, a package assembly 2590 may include a smaller or larger number of components and chiplets interconnected by an interconnect fabric 2585 or a bridge structure 2587. The bridge structure 2587 may be used to establish a point-to-point interconnect between, for example, a logic or I / O chiplet 2574 and memory chiplets 2575. In some implementations, the bridge structure 2587 may also be embedded in the substrate 2580. The chiplets within the package assembly 2590 may have a 2.5D arrangement using a die-on-wafer-on-substrate stacking, where multiple dies are stacked side by side on a silicon interposer including through-silicon vias (TSVs) to couple the chiplets to the substrate 2580, which includes electrical connections to the package interconnect 2583.

[0286] In one embodiment, the silicon interposer is an active interposer 2589 that includes embedded logic in addition to TSVs. In such an embodiment, the chiplets in the package assembly 2590 are arranged on the active interposer 2589 using 3D face-to-face die stacking. The active interposer 2589 may include, in addition to the interconnect fabric 2585 and the bridge structure 2587, I / O hardware logic 2591, cache memory 2592, and other hardware logic 2593. The interconnect fabric 2585 enables communication between the various logic chiplets within the active interposer 2589. The interconnect fabric 2585 may be a NoC interconnect or another form of packet-switched fabric that switches data packets between components of the package assembly.For complex assemblies, the 2585 interconnect fabric can be a dedicated chiplet that enables communication between different hardware logic of the 2590 package assembly.

[0287] The hardware logic chiplets may include special-purpose hardware logic chiplets 2572, a logic or I / O chiplet 2574, and / or memory chiplets 2575. The special-purpose hardware logic chiplets 2572 and the logic or I / O chiplet 2574 may be implemented at least partially in configurable logic or in fixed-functionality logic hardware and may comprise one or more portions of one or more processor cores, graphics processors, parallel processors, or other acceleration devices described herein. The memory chiplets 2575 may be DRAM (e.g., GDDR, HBM) memory or cache (SRAM) memory. The cache memory 2592 in the active interposer 2589 (or substrate 2580) can serve as a global cache for the package assembly 2590, as part of a distributed global cache, or as a dedicated cache for the interconnect fabric 2585.

[0288] Each chiplet may be fabricated as a separate semiconductor die and coupled to a base die embedded in or coupled to the substrate 2580. Coupling to the substrate 2580 may be via an interconnect structure 2573. The interconnect structure 2573 may be configured to carry electrical signals between the various chiplets and the logic within the substrate 2580. The interconnect structure 2573 may include interconnects such as, but not limited to, bumps or pillars. In some embodiments, the interconnect structure 2573 may be configured to carry electrical signals such as input / output (I / O) signals and / or power or ground signals associated with the operation of the logic, I / O, and memory chiplets. In one embodiment, an additional interconnect structure couples the active interposer 2589 to the substrate 2580.

[0289] The substrate 2580 may be an epoxy-based laminate substrate and / or may include other suitable types of substrates. The package assembly 2590 may be connected to other electrical devices via a package interconnect 2583. The package interconnect 2583 may be coupled to a surface of the substrate 2580 to conduct electrical signals to other electrical devices, such as a motherboard, another chipset, or a multi-chip module.

[0290] A logic or I / O chiplet 2574 and a memory chiplet 2575 may be electrically coupled via a bridge structure 2587 configured to route electrical signals between the logic or I / O chiplet 2574 and a memory chiplet 2575. The bridge structure 2587 may be a dense interconnect structure that provides a route for electrical signals. The bridge structure 2587 may include a bridge substrate composed of glass or a suitable semiconductor material. Electrical routing features may be formed on the bridge substrate to provide a chip-to-chip connection between the logic or I / O chiplet 2574 and a memory chiplet 2575. The bridge structure 2587 may also be referred to as a silicon bridge or an interconnect bridge. For example, the bridge structure 2587 is an embedded multi-die interconnect bridge (EMIB).Alternatively, the 2587 bridge structure can simply be a direct connection from one chiplet to another chiplet.

[0291] Fig. Figure 25B illustrates a package assembly 2594 including replaceable chiplets 2595 according to one embodiment. The replaceable chiplets 2595 may be mounted in standardized chiplet slots or chiplet sockets on the base chiplets 2596, 2598. The base chiplets 2596, 2598 may be coupled via a bridge interconnect 2597, which may be similar to other bridge interconnects described herein, such as an EMIB. Memory chiplets may also be connected to logic or I / O chiplets via a bridge interconnect. I / O and logic chiplets may communicate via an interconnect fabric. The base chiplets may each support one or more slots in a standardized format for logic, I / O, or memory / cache.

[0292] SRAM and power delivery circuitry may be fabricated in one or more of the base chiplets 2596, 2598, which may be fabricated using a different process technology than the replaceable chiplets 2595 stacked on top of the base chiplets. For example, the base chiplets 2596, 2598 may be fabricated using a larger process technology, while the replaceable chiplets may be fabricated using a smaller process technology. One or more of the replaceable chiplets 2595 may be memory chiplets (e.g., DRAM). Different memory densities may be selected for the package assembly 2594, depending on the targeted power consumption and / or performance for the product using the package assembly 2594.Furthermore, logic chiplets with a different number or type of functional units can be selected during assembly, depending on the product's target power consumption and / or performance. Furthermore, chiplets containing IP logic cores of different types can be inserted into the interchangeable chiplet slots, enabling hybrid processor designs where IP blocks from different technologies can be mixed and matched. Example integrated system-on-chip circuit

[0293] Fig. Figure 26 illustrates an example integrated circuit that may be manufactured using one or more IP cores. In addition to what is illustrated, other logic and circuitry may be included, including additional graphics processors / cores, peripheral interface controllers, or general-purpose processor cores. The elements of Fig. 26 having the same or similar names as the elements of any other FIG. herein describe the same elements as in the other figures, may operate or function in a similar manner, may include the same components, and may be associated with entities other than, but are not limited to, those described elsewhere herein.

[0294] The integrated system-on-chip circuit 2600 includes one or more application processors 2605 (e.g., CPUs), a graphics processor 2610, which may be a variant of the one or more graphics processors 1408 or any graphics processor described herein, and may be used in place of any described graphics processors. Therefore, the disclosure of any features in combination with a graphics processor herein also discloses, but is not limited to, a corresponding combination with the graphics processor 2610. The integrated SoC circuit 2600 may additionally include an image processor 2615 and / or a video processor 2620, each of which may be a modular IP core from the same or several different design devices. The integrated SoC circuit 2600 may include peripheral or bus logic, including a USB controller 2625, a UART controller 2630, an SPI / SDIO controller 2635, and an I 2 S / I 2C controller 2640. Furthermore, the integrated circuit may include a display device 2645 coupled to one or more of an HDMI (High-Definition Multimedia Interface) controller 2650 and a RAS (Reliability, Availability, and Serviceability) engine 2655. The RAS engine 2655 is used to identify potential errors that may occur during device runtime in order to minimize the downtime that would result if these potential errors were to occur. Memory may be provided by a flash memory subsystem 2660, which includes flash memory and a flash memory controller. A memory interface may be provided via a memory controller 2665 for accessing SDRAM or SRAM memory devices.Some integrated circuits additionally include an embedded security engine 2670, which may include a randomness source such as Intel DRNG, secure time, secure storage, physically unclonable functions (PUF), one-time programmable (OTP) fuses, or OTP memory. Cross-die multicasting from high-bandwidth memory in a graphics processing environment

[0295] Parallel rendering graphics architectures are increasingly expected to deliver near-zero response time and excellent performance for large workloads and increased performance expectations. Examples of large workloads include neural networks, artificial intelligence (AI), machine learning, and more. Such workloads have become more common as they have been implemented in a variety of computing devices, such as personal computing devices, business computing devices, and more. Furthermore, with the increasing use of large workloads involving machine learning and neural networks, new silicon has been created aimed at running large workloads. Such new silicon includes dedicated hardware accelerators (e.g., graphics processing unit (GPU), field-programmable gate array (FPGA), vision processing unit (VPU), and more).) that are adapted to process data using data parallelism.

[0296] Artificial intelligence (AI), including machine learning (ML), deep learning (DL), neural networks, and / or other artificial machine-driven logic, enables machines (e.g., computers, logic circuits, etc.) to use a model to process input data to produce an output based on patterns and / or associations previously learned by the model through a training process. For example, the model may be trained on data to recognize patterns and / or associations and follow such patterns and / or associations when processing input data, so that other inputs result in outputs consistent with the recognized patterns and / or associations.

[0297] With regard to AI applications, parallel rendering graphics architectures have accelerated the processing and execution of AI applications. To achieve higher AI computational throughput, parallel rendering graphics architectures are being scaled to a multi-die design. In a multi-die system design, the question arises how to efficiently route data from the HBM (high-bandwidth memory) to the processing resources, such as execution units, in the processing cores of the dies. This is challenging because the HBM bandwidth increase lags far behind the computational performance increase provided by a multi-die design. This unbalanced scaling can limit the AI ​​system design in meeting AI computational requirements. On the other hand, matrix multiplication, which is an operation in AI workloads, offers a beneficial pattern for data reuse.Such data reuse can be realized through careful system design at different levels, e.g., execution units, local SRAM and memory subsystems, and explicit software management of data movement.

[0298] Some conventional approaches to improving HBM bandwidth in a multi-die graphics architecture design can rely on data reuse for matrix multiplication and employ a hierarchical cache system (e.g., L1$, L2$, and L3$). For example, a data load / store by processing resources (e.g., execution units) can cause the data to be cached in the L3 cache, L2 cache, and / or L1 cache if appropriate cache control is set. Subsequent data loads at the same address can hit the corresponding cache lines, skipping the HBM load and reducing traffic to the HBM. This conventional approach has been further improved by allowing processing resources (e.g., execution units) to pre-fetch data up to a certain level to reduce load time.

[0299] However, while the conventional prefetch mechanism approach described above works well with the cache hierarchy within a single-GPU die, in multi-die GPU systems, inter-die bandwidth is often limited, and using the cache mechanism to share data across multiple dies is often inefficient. For example, when a processing resource, such as an execution unit, from a first die loads data from HBM, it is often cached within a page of the cache (depending on whether it is a compute-side cache or a memory-side cache), typically in the cache located on the same die as the issuing execution units. When execution units from a second (distant) die request the same data, the data often must subsequently travel a long distance to the second die and retrieve the data.This long-distance access increases access latency but also consumes additional performance. Furthermore, there is no guaranteed data affinity across different levels of the cache hierarchy, causing additional latency in data movement and performance degradation.

[0300] Implementations of the disclosure address the aforementioned technical problems by providing cross-die multicasting from a high-bandwidth memory (HBM) in a graphics system. In one implementation, embodiments provide a scheme for loading common data across multiple dies in a graphics environment using explicit software management. This results in guaranteed data affinity and better latency coverage. Implementations herein provide an approach to enable processing resources (e.g., execution units) to prefetch / preload data from the HBM and place the data into the cache / SRAM of multiple dies in the graphics system. This data can then be subsequently consumed by processing resources across the multiple dies with low latency.

[0301] A technical advantage of implementations of the disclosure includes improving HBM bandwidth efficiency and reducing workload latency. This, in turn, improves workload performance in AI applications. Because data movement in the present implementations is explicitly controlled by shared cache (e.g., L2) memory, data affinity is enabled and the possibility of data thrashing is eliminated, resulting in improved software pipelining and latency coverage.

[0302] Fig. 27-30 provide further details on the approach to implementing cross-die multicasting from HBM in graphics processing environments.

[0303] Fig. 27 is a block diagram illustrating an exemplary multi-die GPU computing system 2700 for providing high-bandwidth cross-die multicasting from memory in a graphics processing environment, according to present implementations. The elements of Fig. 27 having the same or similar names as the elements of any other figure herein describe the same elements as in the other figures, may operate or function in a similar manner, may include the same components, and may be associated with other entities such as, but are not limited to, those described elsewhere herein. Therefore, the discussion of any features in combination with a graphics processor herein also discloses, but is not limited to, a corresponding combination with the example multi-die GPU computing system 2700.

[0304] In one implementation, the multi-die GPU computing system 2700 may include multiple dies, such as die 0 2710-0 and die 1 2710-1 (collectively referred to herein as die 2710). Each die 2710 may include a processing resource array 2720-0, 2720-1 (collectively referred to herein as processing resource array 2720), a cache 2730-0, 2730-1 (collectively referred to herein as cache 2730), and / or HBM 2740-0, 2740-1 (collectively referred to herein as HBM 2740).

[0305] In one implementation, the processing resource arrays may include a plurality of GPGPU and / or GPU processing cores, each housing GPU processing resources such as GPU execution units, etc. In one implementation, the plurality of GPGPUs / GPUs of the processing resource array 2720 may be the same as those in Fig. 8 described GPGPUs 806A-806D. Moreover, each of the multiple GPGPUs / GPUs of the processing resource array 2720 may include an instance of the GPGPU 700 of Fig. 7 be.

[0306] The processing resource array may be interconnected via a set of high-speed point-to-point GPU links. The high-speed GPU-to-GPU links may be connected to each of the GPGPUs via a dedicated GPU link, such as GPU link 710 as shown in Fig. 7, be connected.

[0307] As previously mentioned, parallel rendering graphics architectures have accelerated the processing and execution of AI applications. In the present implementations, the multi-die GPU compute system 2700 can be used to accelerate the processing and execution of AI applications by providing cross-die multicasting from HBM. In one implementation, embodiments provide cross-die multicast components 2735-0, 2735-1 (collectively referred to herein as cross-die multicast 2735) as part of the cache hierarchy to enable loading common data across multiple dies in the multi-die GPU compute system 2700. The cross-die multicast components 2735 can utilize explicit software management. This results in guaranteed data affinity and better latency coverage.

[0308] Implementations herein provide an approach that enables components of processing resource array 2720 (e.g., execution units) to prefetch / preload data from HBM 2740 and place the data into the cache (e.g., SRAM) 2730 of multiple dies in the graphics system via cross-die multicast components 2735. This data can then be subsequently consumed by processing resources (of processing resource array 2720) in the multiple dies 2710 with low latency. Further details of cross-die multicasting from HBM are described below with reference to Fig. 28-30 described in more detail.

[0309] Fig. 28 is a block diagram illustrating an exemplary multi-die GPU computing system 2800 for providing cross-die multicasting from HBM, according to present implementations. In one implementation, the multi-die GPU computing system 2800 may be the same as that shown in Fig. 27 described multi-die GPU computing system 2700. Thus, the elements from Fig. 28 with the same or similar names as the elements from Fig. 27, the same elements as in the other figures may operate or function in a similar manner, may include the same components, and may be associated with other entities such as, but are not limited to, those described elsewhere herein. Therefore, the discussion of any features in combination with a graphics processor herein also discloses, but is not limited to, a corresponding combination with the exemplary multi-die GPU computing system 2800.

[0310] In one implementation, the multi-die GPU computing system 2800 may include multiple dies, such as die 0 2810-0 and die 1 2810-1 (collectively referred to herein as die 2810). Each die 2810 may include a processing core array 2820-0, 2820-1 (collectively referred to herein as processing core array 2820), an L2 cache 2830-0, 2830-1 (collectively referred to herein as L2 cache 2830), a shared L2 memory (scratchpad) 2850-0, 2850-1 (collectively referred to herein as shared L2 memory (scratchpad) 2850), and / or HBM 2840-0, 2840-1 (collectively referred to herein as HBM 2840).

[0311] In one implementation, the processing core arrays 2820 may include a plurality of GPGPU and / or GPU processing cores, each housing GPU processing resources such as GPU execution units, etc. In one implementation, the plurality of GPGPUs / GPUs of the processing core array 2820 may be the same as those in Fig. 8 described GPGPUs 806A-806D. Moreover, each of the multiple GPGPUs / GPUs of the processing core array 2820 may include an instance of the GPGPU 700 of Fig. 7. As in Fig. 28, each processing core of the processing core array 2820 may include, but is not limited to, processing resources (e.g., execution units, etc.) 2812-0, 2812-1 (collectively referred to herein as processing resources 2812), matrix multiply-accumulate (MMA) 2814-0, 2814-1 (collectively referred to herein as MMA 2814), local direct memory access (DMA) 2816-0, 2816-1 (collectively referred to herein as local DMA 2816), and / or shared local memory (SLM) 2818-0, 2818-1 (collectively referred to herein as SLM 2818).

[0312] As mentioned above, the multi-die GPU computing system 2800 can be used to accelerate the processing and execution of AI applications by providing cross-die multicasting from the HBM 2840. In one implementation, embodiments provide cross-die multicast hardware components that enable common data to be loaded onto multiple dies in the multi-die GPU computing system 2800. In present implementations, the cross-die multicast hardware components use explicit software management to enable cross-die multicasting from the HBM 2840.

[0313] In particular, present implementations provide the following hardware components to enable cross-die multicasting from HBM 2840: A shared L2 memory (scratchpad) 2850 that shares the same memory space as L2 cache 2830 and is dynamically partitioned by L2 cache 2830 based on application requirements. A local DMA 2816 at each processing core of processing core array 2820 that is responsible for data movement between SLM 2818 and L2 cache 2830 / shared L2 memory 2850, and in some cases, HBM 2840. This local DMA 2816 can be addressed by processing resources 2812 (e.g., execution units) within the processing core.

[0314] The hardware components further include an L2 DMA 2855-0, 2855-1 (collectively referred to herein as L2 DMA 2855). The L2 DMA 2855 may be located in the shared L2 memory 2850 and is responsible for data movement between the shared L2 memory 2850 and the HBM 2840. In one implementation, the L2 DMA 2855 may be jointly controlled by the processing resources 2812 of the processing cores within the processing core array 2820 of the same die 2810. In present implementations, an L2 barrier is used to track a producer-consumer dependency of an associated L2 DMA copy.

[0315] In present implementations, the aforementioned hardware components of the multi-die GPU compute system 2800 may perform the following functions to enable cross-die multicasting from the HBM 2840. Data identified as common data to be shared across multiple dies (e.g., by an application, a programmer, etc.) from the associated HBM 2840 stacks into the shared L2 memory 2850. If multicast is enabled for the data, the copied data may be multicast into the shared L2 memory 2850 of the other dies 2810. In one implementation, it is the programmer's responsibility to ensure that the remote shared L2 memory 2850 is available before issuing an L2 DMA copy command.

[0316] As part of multicasting, the data may then be copied between the shared L2 memory 2850 across the dies 2810 using the L2 DMA 2855 hardware component. Once the data has been copied between the shared L2 memory 2850 of the multiple dies 2810, the data may then be copied from the shared L2 memory 2850 to the HBM 2840 of each die 2810. The L2 DMA component 2855 may also be responsible for copying the data between the shared L2 memory 2850 and the HBM 2840.

[0317] Present implementations also provide software programming support for the aforementioned hardware components, as described below. Within a die, an ND region is partitioned into multiple clusters. An ND region can refer to an index space or an overall execution region. Each cluster consists of multiple workgroups. The multiple clusters within a die are referred to as a region. A workgroup is placed in a processing core of the processing core array 2820. A workgroup is composed of multiple threads.

[0318] In present implementations, a cluster may have a distributed shared local memory 2818 connected via a dedicated interconnect fabric. Threads within a region may be started together and may be synchronized with each other through dedicated hardware barrier mechanisms.

[0319] In one implementation, a thread within a region can issue an L2 DMA copy command to move the data. To issue such a command, the issuing thread should wait on a region-by-region barrier to ensure that all threads within the region have consumed the data in the shared L2 memory 2850 and that the shared L2 memory 2850 is available for new data to be loaded.

[0320] In another implementation, an alternative approach is to connect a consumer barrier to the copy instruction. The L2 barrier can then track whether all threads within the region are signaling the availability of the memory space (i.e., whether the data has been consumed by any threads). As a result, the L2 DMA copy instruction can be issued without waiting, and the thread can move on to another task.

[0321] In present implementations, when the L2 DMA copy command is issued, the command is connected to another barrier, called a producer barrier. When the producer barrier is released, the L2 DMA 2855 can perform the data copy operation. Once the data copy is complete, the L2 DMA 2855 can notify the producer barrier. This producer barrier is either accessible to all threads in the region (via a polling mechanism) or can notify all threads in the region via a notification mechanism. When the thread is notified that the data in the shared L2 memory 2850 is ready for access, the thread can issue another local DMA copy to move the data to the shared local memory 2818, where it can be directly accessed by the processing resources 2812 or the MMA 2814.

[0322] Fig. 29 is a flowchart illustrating one embodiment of a method 2900 for providing hardware support for cross-die multicasting from high-bandwidth memory in a graphics processing environment. The method 2900 may be performed by processing logic that may include hardware (e.g., circuitry, dedicated logic, programmable logic, etc.), software (such as instructions executing on a processing device), or a combination thereof. The process of the method 2900 is illustrated in linear sequences for brevity and clarity of illustration; however, it is contemplated that any number of them may be performed in parallel, asynchronously, or in different orders. Further, for brevity, clarity, and ease of understanding, many of the components and processes described with respect to the Fig. 1-28 will not be repeated or discussed below. In one implementation, a processor, such as the processing die 2810 of Fig. 28, perform procedure 2900.

[0323] The method 2900 begins at processing block 2910, where the processor may copy data identified as shared data from an HBM of a first GPU die to a shared L2 cache of the first GPU die. Then, at block 2920, the processor may determine that multicast is enabled for the shared L2 cache on the first GPU die.

[0324] Thereafter, if multicast is enabled for the shared L2 cache, at block 2930, the processor may transfer the data to the remote shared L2 cache of one or more other GPU dies communicatively coupled to the first GPU die. Finally, at block 2940, the processor may copy the data from the remote shared L2 cache to the remote HBM stack of the one or more other GPU dies.

[0325] Fig. 30 is a flowchart illustrating one embodiment of a method 3000 for providing programmable support for cross-die multicasting from high-bandwidth memory in a graphics processing environment. The method 3000 may be performed by processing logic that may include hardware (e.g., circuitry, dedicated logic, programmable logic, etc.), software (such as instructions executing on a processing device), or a combination thereof. The process of the method 3000 is illustrated in linear sequences for brevity and clarity of illustration; however, it is contemplated that any number of them may be performed in parallel, asynchronously, or in different orders. Further, for brevity, clarity, and ease of understanding, many of the components and processes described with respect to the Fig. 1-29 will not be repeated or discussed below. In one implementation, a processor, such as the processing die 2810 of Fig. 28, perform procedure 3000.

[0326] The method 3000 begins with processing block 3010, where the processor may issue an L2 DMA copy instruction to move data through a thread of a region configured for a first GPU die. In one implementation, the instruction relies on a dedicated hardware barrier mechanism. Then, at block 3020, the processor may connect a producer barrier to the L2 DMA copy instruction.

[0327] Thereafter, if the producer barrier is released, in block 3030, the processor may perform a data copy operation through an L2 DMA component of the shared L2 memory of the first GPU die. In one implementation, the data copy operation causes a multicast of the data to a remote shared L2 memory of one or more other remote GPU dies communicatively coupled to the first GPU die. In block 3040, when the data copy operation is complete, the processor may notify the producer barrier of the completion of the data copy operation through the L2 DMA component.

[0328] Finally, in block 3050, if the thread determines via the producer barrier that the data in the shared L2 memory is ready to be accessed, the processor may issue a local DMA copy command to cause the data to be moved from the shared L2 memory to the shared local memory of a processor core containing the thread.

[0329] The following examples relate to further embodiments. Example 1 is an apparatus for enabling cross-die multicasting from high-bandwidth memory in a graphics processing environment. The apparatus of Example 1 includes a first processing die comprising: an array of processing cores, each comprising processing resources, shared local memory (SLM), and a local direct memory access (DMA) component;and a cache memory unit communicatively coupled to the array of processing cores, the cache memory unit partitioned with a shared memory cache communicatively coupled to a remote shared memory cache of one or more remote processing dies and comprising a shared memory DMA component configured to: copy data from a high bandwidth memory (HBM) of the device to the shared memory cache; determine that multicast is enabled for the shared memory cache; and if multicast is enabled for the shared memory cache, multicast the data to the remote shared memory cache of the one or more remote processing dies communicatively coupled to the first processing die.

[0330] In Example 2, the subject matter of Example 1 can optionally include where the cache memory unit comprises an L2 cache. In Example 3, the subject matter of any of Examples 1-2 can optionally include identifying the data as common data to be shared between the first processing die and the one or more remote processing dies. In Example 4, the subject matter of any of Examples 1-3 can optionally include a remote shared memory DMA component of the remote shared memory cache copying the data from the remote shared memory cache to remote HBMs of the one or more remote processing dies.

[0331] In Example 5, the subject matter of any of Examples 1-4 can optionally include the local DMA component copying the data to the SLM if the data is available in the shared memory cache. In Example 6, the subject matter of any of Examples 1-5 can optionally include a thread of a region configured for the first processing die issuing a DMA copy command to cause the data to be copied from the HBM, wherein a processing core of the array of processing cores executes the thread. In Example 7, the subject matter of any of Examples 1-6 can optionally include the DMA copy command utilizing a hardware barrier mechanism implemented by the first processing die and the one or more remote processing dies, and the thread connecting a producer barrier of the hardware barrier mechanism to the DMA copy command.

[0332] In Example 8, the subject matter of any of Examples 1-7 can optionally include the shared memory DMA component being further configured to: when the producer barrier is enabled, perform the multicasting of the data to the remote shared memory cache of the one or more remote processing dies; and, when the multicasting of the data is complete, notify the producer barrier of the completion of a data copy operation; wherein the thread issues a local DMA copy command to cause the data to be moved from the shared memory cache to the SLM based on a determination by the producer barrier that the data in the shared memory cache is ready for access.In Example 9, the subject matter of any of Examples 1-8 can optionally include wherein the first processing die and the one or more remote processing dies comprise graphics processing unit (GPU) dies.

[0333] Example 10 is a method for enabling cross-die multicasting from high-bandwidth memory in a graphics processing environment. The method of Example 10 may be performed by a shared memory DMA (direct memory access) component of a shared memory cache of a first processing die, copying data from high-bandwidth memory (HBM) of the first processor die to the shared memory cache, wherein the shared memory cache is partitioned from a cache unit of the first processing die comprising an array of processing cores, each comprising processing resources, shared local memory (SLM), and a local direct memory access (DMA) component; the shared memory DMA component determining that multicast is enabled for the shared memory cache;and if multicast is enabled for the shared memory cache, multicasting the data by the shared memory DMA component to remote shared memory of one or more remote processing dies communicatively coupled to the first processing die;

[0334] In Example 11, the subject matter of Example 10 can optionally include identifying the data as common data to be shared between the first processing die and the one or more remote processing dies. In Example 12, the subject matter of Examples 10-11 can optionally include a remote shared memory DMA component of the remote shared memory cache copying the data from the remote shared memory cache to remote HBMs of the one or more remote processing dies. In Example 13, the subject matter of Examples 10-12 can optionally include a thread of a region configured for the first processing die issuing a DMA copy command to cause the data to be copied from the HBM, wherein a processing core of the array of processing cores executes the thread.

[0335] In Example 14, the subject matter of Examples 10-13 can optionally include the DMA copy instruction utilizing a hardware barrier mechanism implemented by the first processing die and the one or more remote processing dies, and the thread connecting a producer barrier of the hardware barrier mechanism to the DMA copy instruction.In Example 15, the subject matter of Examples 10-14 can optionally further include: when the producer barrier is enabled, performing the multicasting of the data to the remote shared memory cache of the one or more remote processing dies; and when the multicasting of the data is complete, notifying the producer barrier of the completion of a data copy operation; wherein, based on a determination from the producer barrier that the data in the shared memory cache is ready for access, the thread issues a local DMA copy command to cause the data to be moved from the shared memory cache to the SLM.

[0336] Example 16 is a non-transient computer-readable storage medium for enabling cross-die multicasting from high-bandwidth memory in a graphics processing environment. The non-transient computer-readable storage medium of Example 16 has instructions stored thereon that, when executed by one or more processors, cause the one or more processors to perform operations comprising: copying, by a shared DMA (direct memory access) component of a shared memory cache of a first processing die, data from a high-bandwidth memory (HBM) of the first processing die to the shared memory cache, wherein the shared memory cache is partitioned from a cache memory unit of the first processing die comprising an array of processing cores, each having processing resources,shared local memory (SLM) and a local direct memory access (DMA) component; determining, by the shared memory DMA component, that multicast is enabled for the shared memory cache; and if multicast is enabled for the shared memory cache, multicasting, by the shared memory DMA component, the data to remote shared memory of the one or more remote processing dies communicatively coupled to the first processing die.

[0337] In Example 17, the subject matter of Example 16 can optionally include a remote shared memory DMA component of the remote shared memory cache copying the data from the remote shared memory cache to remote HBMs of the one or more remote processing dies. In Example 18, the subject matter of Examples 16-17 can optionally include a thread of a region configured for the first processing die issuing a DMA copy command to cause the data to be copied from the HBM, wherein a processing core of the array of processing cores executes the thread.

[0338] In Example 19, the subject matter of Examples 16-18 can optionally include the DMA copy instruction utilizing a hardware barrier mechanism implemented by the first processing die and the one or more remote processing dies, and the thread connecting a producer barrier of the hardware barrier mechanism to the DMA copy instruction.In Example 20, the subject matter of Examples 16-19 can optionally further comprise: when the producer barrier is enabled, performing the multicasting of the data to the remote shared memory cache of the one or more remote processing dies; and when the multicasting of the data is complete, notifying the producer barrier of the completion of a data copy operation; wherein, based on a determination from the producer barrier that the data in the shared memory cache is ready for access, the thread issues a local DMA copy command to cause the data to be moved from the shared memory cache to the SLM.

[0339] Example 21 is a system for enabling cross-die multicasting from high-bandwidth memory in a graphics processing environment. The system of Example 21 can optionally include a memory and a first processing die communicatively coupled to the memory and a second processing die, the first processing die comprising: an array of processing cores, each comprising processing resources, shared local memory (SLM), and a local direct memory access (DMA) component;and a cache memory unit communicatively coupled to the array of processing cores, the cache memory unit partitioned with a shared memory cache communicatively coupled to a remote shared memory cache of one or more remote processing dies and comprising a shared memory DMA component configured to: copy data from a high bandwidth memory (HBM) of the device to the shared memory cache; determine that multicast is enabled for the shared memory cache; and if multicast is enabled for the shared memory cache, multicast the data to the remote shared memory cache of the one or more remote processing dies communicatively coupled to the first processing die.

[0340] In Example 22, the subject matter of Example 21 can optionally include where the cache memory unit comprises an L2 cache. In Example 23, the subject matter of any of Examples 21-22 can optionally include identifying the data as common data to be shared between the first processing die and the one or more remote processing dies. In Example 24, the subject matter of any of Examples 21-23 can optionally include a remote shared memory DMA component of the remote shared memory cache copying the data from the remote shared memory cache to remote HBMs of the one or more remote processing dies.

[0341] In Example 25, the subject matter of any of Examples 21-24 can optionally include the local DMA component copying the data to the SLM if the data is available in the shared memory cache. In Example 26, the subject matter of any of Examples 21-25 can optionally include a thread of a region configured for the first processing die issuing a DMA copy command to cause the data to be copied from the HBM, wherein a processing core of the array of processing cores executes the thread. In Example 27, the subject matter of any of Examples 21-26 can optionally include the DMA copy command utilizing a hardware barrier mechanism implemented by the first processing die and the one or more remote processing dies, and the thread connecting a producer barrier of the hardware barrier mechanism to the DMA copy command.

[0342] In Example 28, the subject matter of any of Examples 21-27 can optionally include the shared memory DMA component being further configured to: when the producer barrier is enabled, perform the multicasting of the data to the remote shared memory cache of the one or more remote processing dies; and, when the multicasting of the data is complete, notify the producer barrier of the completion of a data copy operation; wherein the thread issues a local DMA copy command to cause the data to be moved from the shared memory cache to the SLM based on a determination by the producer barrier that the data in the shared memory cache is ready for access.In Example 29, the subject matter of any of Examples 21-28 can optionally include wherein the first processing die and the one or more remote processing dies comprise graphics processing unit (GPU) dies.

[0343] Example 30 is an apparatus for enabling cross-die multicasting from high-bandwidth memory in a graphics processing environment, comprising means for copying data from high-bandwidth memory (HBM) of the first processing die to the shared memory cache via a shared memory DMA (direct memory access) component of a first processing die, the shared memory cache being partitioned from a cache unit of the first processing die comprising an array of processing cores, each comprising processing resources, a shared local memory (SLM), and a local direct memory access (DMA) component; means for determining that multicast is enabled for the shared memory cache;and if multicast is enabled for the shared memory cache, means for multicasting the data to a remote shared memory cache of one or more remote processing dies communicatively coupled to the first processing die. In Example 31, the subject matter of Example 30 can optionally include the device being further configured to perform the method of any of Examples 11 to 15.

[0344] Example 32 is at least one machine-readable medium comprising a plurality of instructions that, in response to being executed on a computing device, cause the computing device to perform a method according to any of Examples 10-15. Example 33 is an apparatus for enabling cross-die multicasting from high-bandwidth memory in a graphics processing environment configured to perform the method of any of Examples 10 to 15. Example 34 is an apparatus for enabling cross-die multicasting from high-bandwidth memory in a graphics processing environment, comprising means for performing the method of any of Examples 10 to 15. The teachings in the examples can be used throughout one or more embodiments.

[0345] The foregoing description and drawings are to be considered in an illustrative and not restrictive sense. Those skilled in the art will understand that various modifications and changes may be made to the embodiments described herein without departing from the broader spirit and scope of the features as set forth in the appended claims. QUOTES CONTAINED IN THE DESCRIPTION

[0000] This list of documents submitted by the applicant was generated automatically and is included solely for the convenience of the reader. This list is not part of the German patent or utility model application. The DPMA assumes no liability for any errors or omissions. Cited non-patent literature

[0000] Shane Cook, CUDA Programming, Chapter 3, pages 37-51 (2013

[0003]

Claims

[1] Device comprising: a first processing die comprising: an array of processing cores, each comprising processing resources, shared local memory (SLM), and a local direct memory access (DMA) component; and a cache memory unit communicatively coupled to the array of processing cores, the cache memory unit partitioned with a shared memory cache communicatively coupled to a remote shared memory cache of one or more remote processing dies and comprising a shared memory DMA component configured to: Copying data from a high bandwidth memory (HBM) of the device to the shared memory cache; Determine that multicast is enabled for the shared memory cache; and if shared memory cache multicast is enabled, multicasting the data to the remote shared memory cache of the one or more remote processing dies communicatively coupled to the first processing die. [2] The apparatus of claim 1, wherein the cache memory unit comprises an L2 cache memory. [3] The apparatus of any of claims 1-2, wherein the data is identified as common data to be shared between the first processing die and the one or more remote processing dies. [4] The apparatus of any of claims 1-3, wherein a remote shared memory DMA component of the remote shared memory cache copies the data from the remote shared memory cache to remote HBM of the one or more remote processing dies. [5] The apparatus of any of claims 1-4, wherein the local DMA component copies the data to the SLM when the data is available in the shared memory cache. [6] The apparatus of any of claims 1-5, wherein a thread of a region configured for the first processing die issues a DMA copy command to cause the data to be copied from the HBM, wherein a processing core of the array of processing cores executes the thread. [7] The apparatus of any of claims 1-6, wherein the DMA copy instruction utilizes a hardware barrier mechanism implemented by the first processing die and the one or more remote processing dies, and the thread connects a producer barrier of the hardware barrier mechanism to the DMA copy instruction. [8] The apparatus of any of claims 1-7, wherein the shared memory DMA component is further configured to: if the producer barrier is released, multicasting the data to the remote shared memory cache of the one or more remote processing dies; and when the multicast sending of the data is completed, notifying the producer barrier of the completion of a data copy operation; wherein the thread issues a local DMA copy command to cause the data to be moved from the shared memory cache to the SLM based on a determination from the producer barrier that the data in the shared memory cache is ready for access. [9] The apparatus of any of claims 1-8, wherein the first processing die and the one or more remote processing dies comprise graphics processing unit (GPU) dies. [10] Method comprising: copying data from a high bandwidth memory (HBM) of the first processor die to the shared memory cache by a shared memory DMA (direct memory access) component of a shared memory cache of a first processing die, the shared memory cache being partitioned from a cache memory unit of the first processing die comprising an array of processing cores, each comprising processing resources, shared local memory (SLM), and a local direct memory access (DMA) component; determining, by the shared memory DMA component, that multicast is enabled for the shared memory cache; and if multicast is enabled for the shared memory cache, multicasting the data by the shared memory DMA component to remote shared memory of one or more remote processing dies communicatively coupled to the first processing die. [11] The method of claim 10, wherein the data is identified as common data to be shared between the first processing die and the one or more remote processing dies. [12] The method of any of claims 10-11, wherein a remote shared memory DMA component of the remote shared memory cache copies the data from the remote shared memory cache to remote HBM of the one or more remote processing dies. [13] The method of any of claims 10-12, wherein a thread of a region configured for the first processing die issues a DMA copy command to cause the data to be copied from the HBM, wherein a processing core of the array of processing cores executes the thread. [14] The method of any of claims 10-13, wherein the DMA copy instruction utilizes a hardware barrier mechanism implemented by the first processing die and the one or more remote processing dies, and the thread connects a producer barrier of the hardware barrier mechanism to the DMA copy instruction. [15] A method according to any one of claims 10-14, further comprising: if the producer barrier is released, multicasting the data to the remote shared memory cache of the one or more remote processing dies; and when the multicast sending of the data is completed, notifying the producer barrier of the completion of a data copy operation; wherein the thread issues a local DMA copy command to cause the data to be moved from the shared memory cache to the SLM based on a determination from the producer barrier that the data in the shared memory cache is ready for access. [16] A system for enabling cross-die multicasting from a high-bandwidth memory in a graphics processing environment, comprising: a memory; and a first processing die communicatively coupled to the memory and a second processing die, the first processing die comprising: an array of processing cores, each comprising processing resources, shared local memory (SLM), and a local direct memory access (DMA) component; and a cache memory unit communicatively coupled to the array of processing cores, the cache memory unit partitioned with a shared memory cache communicatively coupled to a remote shared memory cache of one or more remote processing dies and comprising a shared memory DMA component configured to: Copying data from a high bandwidth memory (HBM) of the device to the shared memory cache; Determine that multicast is enabled for the shared memory cache; and if shared memory cache multicast is enabled, multicasting the data to the remote shared memory cache of the one or more remote processing dies communicatively coupled to the first processing die. [17] The system of claim 16, wherein the cache unit comprises an L2 cache. [18] The system of any of claims 16-17, wherein the data is identified as common data to be shared between the first processing die and the one or more remote processing dies. [19] The system of any of claims 16-18, wherein a remote shared memory DMA component of the remote shared memory cache copies the data from the remote shared memory cache to remote HBM of the one or more remote processing dies. [20] The system of any of claims 16-19, wherein the local DMA component copies the data to the SLM when the data is available in the shared memory cache. [21] The system of any of claims 16-20, wherein a thread of a region configured for the first processing die issues a DMA copy command to cause the data to be copied from the HBM, wherein a processing core of the array of processing cores executes the thread. [22] The system of any of claims 16-21, wherein the DMA copy instruction utilizes a hardware barrier mechanism implemented by the first processing die and the one or more remote processing dies, and wherein the thread connects a producer barrier of the hardware barrier mechanism to the DMA copy instruction. [23] The system of any of claims 16-22, wherein the shared memory DMA component is further configured to: when the producer barrier is enabled, perform the multicasting of the data to the remote shared memory cache of the one or more remote processing dies; and, when the multicasting of the data is complete, notify the producer barrier of the completion of a data copy operation; wherein the thread issues a local DMA copy command to cause the data to be moved from the shared memory cache to the SLM based on a determination by the producer barrier that the data in the shared memory cache is ready for access. [24] Machine-readable medium(s) comprising a plurality of instructions which, when executed on a computing device, cause the computing device to perform a method according to any one of claims 10-15. [25] Apparatus for enabling cross-die multicasting from a high-bandwidth memory in a graphics processing environment, comprising means for performing the method of any of claims 10-15.