Decomposing the SOC Architecture
By integrating a GPU with SIMT architecture into the graphics processing system, the challenges of processing graphics data are addressed, resulting in improved efficiency and accelerated operations for both graphics and machine learning tasks.
Patent Information
- Application Number
- JP2024003662
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2019-03-15
- Filing Date
- 2024-01-15
- Publication Date
- 2025-05-14
- Estimated Expiration
- 2040-01-23
AI Technical Summary
Current parallel graphics data processing systems face challenges in efficiently processing graphics data due to limitations in fixed function computing units and the need for improved parallel processing techniques.
The implementation of a graphics processing unit (GPU) communicatively coupled to a host processor to accelerate graphics and machine learning operations, utilizing SIMT architecture and optimized circuitry for parallel processing.
This approach enhances processing efficiency by maximizing parallel processing in the graphics pipeline, allowing for accelerated graphics and machine learning operations, and improving overall system performance.
Smart Images

Figure 0007676599000008 
Figure 0007676599000009 
Figure 0007676599000010
Abstract
Description
[Technical field]
[0001] [Related Applications] This application claims the benefit of priority to U.S. Patent Application No. 16 / 355,377, filed March 15, 2019. The prior U.S. patent application is hereby incorporated by reference in its entirety.
[0002] [Field] Embodiments relate generally to the design and manufacture of general purpose graphics and parallel processing units. [Background technology]
[0003] Modern parallel graphics data processing includes systems and methods developed to perform specific operations on graphics data, such as, for example, linear interpolation, tessellation, rasterization, texture mapping, depth testing, etc. Traditionally, graphics processors used fixed function computation units to process the graphics data, but more recently, portions of graphics processors have become programmable, enabling such processors to support a greater variety of operations for processing vertex and fragment data.
[0004] To further improve performance, graphics processors typically implement processing techniques such as pipelining, which attempt to process as much graphics data as possible in parallel across different parts of the graphics pipeline. Parallel graphics processors with SIMT (single instruction, multiple thread) architectures are designed to maximize the amount of parallelism in the graphics pipeline. In a SIMT architecture, a group of parallel threads attempt to execute program instructions together simultaneously as often as possible to increase processing efficiency. An overview of the software and hardware for SIMT architectures can be found in Shane Cook, CUDA Programming Chapter 3, pages 37-51 (2013). [Brief description of the drawings]
[0005] So that the above-recited features of the present embodiments can be understood in detail, the embodiments briefly summarized above will now be more particularly described with reference to embodiments, some of which are illustrated in the accompanying drawings.
[0006] [Figure 1] FIG. 1 is a block diagram illustrating a computer system configured to implement one or more aspects of the embodiments described herein. [Figure 2A] 1 illustrates a parallel processor component according to an embodiment. [Figure 2B] 1 illustrates a parallel processor component according to an embodiment. [Figure 2C] 1 illustrates a parallel processor component according to an embodiment. [Figure 2D] 1 illustrates a parallel processor component according to an embodiment. [Figure 3A] FIG. 2 is a block diagram of a graphics multiprocessor and a multi-block-based GPU according to an embodiment. [Figure 3B] FIG. 2 is a block diagram of a graphics multiprocessor and a multi-block-based GPU according to an embodiment. [Figure 3C] FIG. 2 is a block diagram of a graphics multiprocessor and a multi-block-based GPU according to an embodiment. [Figure 4A] 1 depicts an exemplary architecture in which multiple GPUs are communicatively coupled to multiple multi-core processors. [Figure 4B] 1 depicts an exemplary architecture in which multiple GPUs are communicatively coupled to multiple multi-core processors. [Figure 4C] 1 depicts an exemplary architecture in which multiple GPUs are communicatively coupled to multiple multi-core processors. [Figure 4D] 1 depicts an exemplary architecture in which multiple GPUs are communicatively coupled to multiple multi-core processors. [Figure 4E] 1 depicts an exemplary architecture in which multiple GPUs are communicatively coupled to multiple multi-core processors. [Figure 4F] 1 depicts an exemplary architecture in which multiple GPUs are communicatively coupled to multiple multi-core processors. [Diagram 5] 1 illustrates a graphics processing pipeline according to an embodiment. [Figure 6] 1 illustrates a machine learning software stack according to an embodiment. [Figure 7] 1 illustrates a general purpose graphics processing unit according to an embodiment. [Figure 8] 1 illustrates a multi-GPU computing system according to an embodiment. [Figure 9A] 1 illustrates layers of an exemplary deep neural network. [Figure 9B] 1 illustrates layers of an exemplary deep neural network. [Figure 10] 1 illustrates an exemplary recurrent neural network. [Figure 11] Represents training and deploying deep neural networks. [Figure 12] FIG. 2 is a block diagram illustrating distributed learning. [Figure 13]1 illustrates an exemplary estimated SOC (system on chip) suitable for estimation using a trained model. [Figure 14] FIG. 2 is a block diagram of a processing system according to an embodiment. [Figure 15] FIG. 2 is a block diagram of a processor according to an embodiment. [Figure 16] FIG. 2 is a block diagram of a graphics processor according to an embodiment. [Figure 17] FIG. 2 is a block diagram of a graphics processing engine of a graphics processor according to some embodiments. [Figure 18] FIG. 2 is a block diagram of hardware logic of a graphics processor core according to some embodiments described herein. [Figure 19A] 1 illustrates thread execution logic including an array of processing elements used in a graphics processor according to an embodiment described herein. [Figure 19B] 1 illustrates thread execution logic including an array of processing elements used in a graphics processor according to an embodiment described herein. [Figure 20] FIG. 2 is a block diagram illustrating a graphics processor instruction format according to some embodiments. [Figure 21] FIG. 2 is a block diagram of a graphics processor according to another embodiment. [Figure 22A] 1 illustrates a graphics processor command format and format sequence according to some embodiments. [Figure 22B] 1 illustrates a graphics processor command format and format sequence according to some embodiments. [Diagram 23] 1 illustrates an exemplary graphics software architecture for a data processing system according to some embodiments. [Figure 24A] FIG. 1 is a block diagram illustrating an IP core development system according to an embodiment. [Figure 24B] 1 illustrates a cross-sectional side view of an integrated circuit package assembly according to some embodiments described herein. [Diagram 25] FIG. 1 is a block diagram illustrating an exemplary system on a chip integrated circuit according to an embodiment. [Figure 26A] FIG. 2 is a block diagram illustrating an example graphics processor for use within a SoC in accordance with an embodiment described herein. [Figure 26B] FIG. 2 is a block diagram illustrating an example graphics processor for use within a SoC in accordance with an embodiment described herein. [Figure 27] 1 illustrates a parallel computing system according to an embodiment. [Figure 28A] 1 illustrates a hybrid logical / physical view of a non-cohesive parallel processor according to an embodiment described herein. [Figure 28B] 1 illustrates a hybrid logical / physical view of a non-cohesive parallel processor according to an embodiment described herein. [Figure 29A] 2 illustrates a package diagram of a non-cohesive parallel processor according to an embodiment. [Figure 29B] 2 illustrates a package diagram of a non-cohesive parallel processor according to an embodiment. [Diagram 30] 1 illustrates a message transport system for an interconnect fabric according to an embodiment. [Diagram 31] It refers to the transmission of messages or signals between functional units across multiple physical links of the interconnect fabric. [Diagram 32] Represents the transmission of messages or signals of multiple functional units across a single physical link of the interconnect fabric. [Diagram 33] 1 depicts a method for configuring fabric connections for functional units in a non-cohesive parallel processor. [Diagram 34] 1 illustrates a method for relaying messages and / or signals across an interconnect fabric in a non-cohesive parallel processor. [Diagram 35] This shows how to power gate chiplets per workload. [Diagram 36] 1 depicts a parallel processor assembly that includes interchangeable chiplets. [Figure 37] 1 illustrates a replaceable chiplet system according to an embodiment. [Figure 38] 4 is an illustration of multiple traffic classes carried on a virtual channel according to an embodiment. [Figure 39] 1 illustrates a method of inter-slot agnostic data transmission for swappable chiplets, according to an embodiment. [Diagram 40] 1 illustrates a modular architecture for chiplets in a replaceable chiplet system according to an embodiment. [Diagram 41] Represents the use of a standardized chassis interface that is used to enable testing, validation, and integration of chiplets. [Diagram 42] 1 illustrates the use of individually binned chiplets to create different product tiers. [Diagram 43] 1 illustrates a method for enabling different product tiers based on chiplet configuration. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS
[0007] In some embodiments, a graphics processing unit (GPU) is communicatively coupled to a host / processor core to accelerate graphics operations, machine learning operations, pattern analysis operations, and various general-purpose GPU (GPGPU) functions. The GPU may be communicatively coupled to the host processor / core via a bus or other interconnect (e.g., a high-speed interconnect such as PCIe or NVLink). In other embodiments, the GPU may be integrated in the same package or chip as the core and communicatively coupled to the core via an internal processor bus / interconnect (i.e., within the package or chip). Regardless of the manner in which the GPU is connected, the processor core may assign work to the GPU in the form of a sequence of commands / instructions contained in a work descriptor. The GPU then uses dedicated circuitry / logic to efficiently process these commands / instructions.
[0008] In the following description, numerous specific details are set forth to provide a more thorough understanding. However, as will be apparent to one of ordinary skill in the art, the embodiments described herein may be practiced without one or more of these specific details. In other instances, well-known features have not been described so as not to obscure the details of the embodiments.
[0009] [System Overview] FIG. 1 is a block diagram illustrating a computing system 100 configured to implement one or more aspects of the embodiments described herein. The computing system 100 includes a processing subsystem 101 with one or more processors 102 and a system memory 104 that communicate via an interconnection path. The interconnection path may include a memory hub 105. The memory hub 105 may be a separate component in a chipset component or may be integrated within the one or more processors 102. The memory hub 105 couples to an I / O subsystem 111 via a communication link 106. The I / O subsystem 111 includes an I / O hub 107 that may enable the computing system 100 to receive input from one or more input devices 108. Furthermore, the I / O hub 107 may enable a display controller that may be included in the one or more processors 102 to provide output to one or more display devices 110A. In one embodiment, the one or more display devices 110A coupled to the I / O hub 107 may include local, built-in, or embedded display devices.
[0010] In one embodiment, the processing subsystem 101 includes one or more parallel processors 112 coupled to the memory hub 105 via a bus or other communication link 113. The communication link 113 may be one of any number of standard-based communication link technologies or protocols, such as, but not limited to, PCI Express, or may be a vendor-specific communication interface or fabric. In one embodiment, the one or more parallel processors 112 form a computationally intensive parallel or vector processing system that may include multiple processing cores and / or processing clusters, such as a many integrated core (MIC) processor. In one embodiment, the one or more parallel processors 112 form a graphics processing system that may output pixels to one of one or more display devices 110A coupled via the I / O hub 107. The one or more parallel processors 112 may also include a display controller and display interface (not shown) that allows for direct connection to one or more display devices 110B.
[0011] Within the I / O subsystem 111, a system storage unit 114 can be connected to the I / O subsystem 111 to provide a storage mechanism for the computing system 100. The I / O switch 116 can be used to provide an interface mechanism to enable connections between the I / O hub 107 and other components, such as a network adapter 118 and / or a wireless network adapter 119, which may be integrated into the platform, and various other devices that may be added via one or more add-in devices 120. The network adapter 118 can be an Ethernet adapter or other wired network adapter. The wireless network adapter 119 can include one or more of Wi-Fi, Bluetooth, Near Field Communication (NFC), or other network devices including one or more wireless radios.
[0012] Computing system 100 may include other components not expressly shown, including USB or other port connections, optical storage drives, video capture devices, etc. may also be connected to I / O hub 107. The communication paths interconnecting the various components in FIG. 1 may be implemented using any suitable protocol, such as a Peripheral Component Interconnect (PCI)-based protocol (e.g., PCI-Express), or any other bus or point-to-point communication interface and / or protocol, such as the NV-Link high-speed interconnect or interconnect protocols known in the art.
[0013] In one embodiment, one or more of the parallel processors 112 incorporate circuitry optimized for graphics and video processing, including, for example, video output circuitry, to form a graphics processing unit (GPU). In other embodiments, one or more of the parallel processors 112 incorporate circuitry optimized for general purpose processing while retaining the underlying computational architecture, as described in further detail herein. In yet other embodiments, the components of the computing system 100 may be integrated with one or more other system elements on a single integrated circuit. For example, one or more of the parallel processors 112, the memory hub 105, the processor 102, and the I / O hub 107 may be integrated into a system-on-chip (SoC) integrated circuit. Alternatively, the components of the computing system 100 may be integrated into a single package to form a system-in-package (SIP) configuration. In one embodiment, at least some of the components of the computing system 100 may be integrated into a multi-chip module (MCM) that may be interconnected with other multi-chip modules in a modular computing system.
[0014] It will be understood that the computing system 100 shown herein is illustrative and that variations and modifications are possible. The connection topology, including the number and arrangement of bridges, the number of processors 102, and the number of parallel processors 112, may be modified as desired. For example, in some embodiments, the system memory 104 is connected directly to the processors 102 rather than through a bridge, while other devices communicate with the system memory 104 via the memory hub 105 and the processors 102. In other alternative topologies, the parallel processors 112 are connected directly to the I / O hub 107 or to one of the one or more processors 102, rather than to the memory hub 105. In other embodiments, the I / O hub 107 and the memory hub 105 may be integrated into a single chip. Some embodiments may include two or more sets of processors 102 attached via multiple sockets that can be coupled with two or more instances of the parallel processors 112.
[0015] Some of the specific components illustrated herein are optional and may not be included in all implementations of computing system 100. For example, any number of add-in cards or peripherals may be supported, or some components may be omitted. Furthermore, some architectures may use different terminology for components similar to those depicted in FIG. 1. For example, memory hub 105 may be referred to as a northbridge in some architectures, while I / O hub 107 may be referred to as a southbridge.
[0016] 2A illustrates a parallel processor 200 according to an embodiment. Various components of the parallel processor 200 may be implemented using one or more integrated circuit devices, such as a programmable processor, an application specific integrated circuit (ASIC), or a field programmable gate array (FPGA). The illustrated parallel processor 200 is a variation of one or more of the parallel processors 112 shown in FIG. 1 according to an embodiment.
[0017] In one embodiment, parallel processor 200 includes parallel processing units 202. Parallel processing units 202 include I / O units 204 that enable communication with other devices, including other instances of parallel processing units 202. I / O units 204 may be directly connected to other devices. In one embodiment, I / O units 204 connect to other devices through the use of a hub or switch interface, such as memory hub 105. The connection between memory hub 105 and I / O units 204 forms communication link 113. Within parallel processing units 202, I / O units 204 connect to host interface 206 and memory crossbar 216, where host interface 206 receives commands directed to perform processing operations and memory crossbar 216 receives commands directed to perform memory operations.
[0018] When the host interface 206 receives command buffers via the I / O unit 204, the host interface 206 can direct work operations to the front end 208 to execute those commands. In one embodiment, the front end 208 couples to a scheduler 210, which is configured to distribute commands or other work items to the processing cluster array 212. In one embodiment, the scheduler 210 ensures that the processing cluster array 212 is properly configured and in a valid state before tasks are distributed to the processing clusters of the processing cluster array 212. In one embodiment, the scheduler 210 is implemented by firmware logic executing on a microcontroller. The microcontroller-implemented scheduler 210 is configured to perform complex scheduling and work distribution operations at both coarse and fine granularity while allowing rapid preemption and context switching of threads executing on the processing array 212. In one embodiment, the host software can present workloads for scheduling to the processing array 212 through one of a number of graphics processing doorbells. The workload may then be automatically distributed across the processing array 212 by scheduler 210 logic within the scheduler microcontroller.
[0019] The processing cluster array 212 may include up to "N" processing clusters (e.g., cluster 214A, cluster 214B, through cluster 214N). Each cluster 214A-214N of the processing cluster array 212 may execute multiple concurrent threads. The scheduler 210 may assign work to the clusters 214A-214N of the processing cluster array 212 using various scheduling and / or work distribution algorithms that may vary depending on the workload generated for each type of program or computation. Scheduling may be handled dynamically by the scheduler 210 or may be partially assisted by compiler logic during compilation of program logic configured for execution by the processing cluster array 212. In one embodiment, different clusters 214A-214N of the processing cluster array 212 may be assigned to process different types of programs and perform different types of computations.
[0020] The processing cluster array 212 may be configured to perform various types of parallel processing operations. In one embodiment, the processing cluster array 212 is configured to perform general-purpose parallel computing operations. For example, the processing cluster array 212 may include logic for performing processing tasks including filtering video and / or audio data, performing modeling operations including physics operations, and performing data transformations.
[0021] In one embodiment, the processing cluster array 212 is configured to perform parallel graphics processing operations. In embodiments in which the parallel processor 200 is configured to perform graphics processing operations, the processing cluster array 212 may include additional logic to support the execution of such graphics processing operations, including, but not limited to, texture sampling logic to perform texture operations along with tessellation logic and other vertex processing logic. Additionally, the processing cluster array 212 may be configured to execute graphics processing related shader programs, such as, but not limited to, vertex shaders, tessellation shaders, geometry shaders, and pixel shaders. The parallel processing units 202 may transfer data from the system memory for processing via the I / O units 204. During processing, the transferred data may be stored in an on-chip memory (e.g., the parallel processor memory 222) during processing and then written to the system memory.
[0022] In one embodiment, when the parallel processing unit 202 is used to perform graphics processing, the scheduler 210 may be configured to divide the processing workload into tasks of approximately equal size to better enable distribution of graphics processing operations to the multiple clusters 214A-214N of the processing cluster array 212. In some embodiments, portions of the processing cluster array 212 may be configured to perform different types of processing. For example, to generate a rendered image for display, a first portion may be configured to perform vertex shading and topology generation, a second portion may be configured to perform tessellation and geometry shading, and a third portion may be configured to perform pixel shading or other screen space operations. Intermediate data generated by one or more of the clusters 214A-214N may be stored in a buffer to allow the intermediate data to be transmitted between the clusters 214A-214N for further processing.
[0023] During operation, the processing cluster array 212 may receive processing tasks to be performed via the scheduler 210, which receives commands defining the processing tasks from the front end 208. For graphics processing operations, the processing tasks may include indices of data to be processed, e.g., surface (patch) data, primitive data, vertex data, and pixel data, along with state parameters and commands that define how the data should be processed (e.g., what program should be executed). The scheduler 210 may be configured to fetch the index corresponding to the task, or may receive the index from the front end 208. The front end 208 may be configured to ensure that the processing cluster array 212 is set to a valid state before a workload specified by an incoming command buffer (e.g., batch buffer, push buffer, etc.) is initiated.
[0024] Each of the one or more instances of the parallel processing unit 202 can be coupled to a parallel processor memory 222. The parallel processor memory 222 can be accessed via a memory crossbar 216. The memory crossbar 216 can receive memory requests from the processing cluster array 212 and the I / O unit 204. The memory crossbar 216 can access the parallel processor memory 222 via a memory interface 218. The memory interface 218 can include a number of partition units (e.g., partition unit 220A, partition unit 220B, through partition unit 220N), each of which can be coupled to a portion (e.g., a memory unit) of the parallel processor memory 222. In one implementation, the number of partition units 220A-220N is configured to be equal to the number of memory units, such that the first partition unit 220A has a corresponding first memory unit 224A, the second partition unit 220B has a corresponding memory unit 224B, and the Nth partition unit 220N has a corresponding Nth memory unit 224N. In other embodiments, the number of partition units 220A-220N may not be equal to the number of memory devices.
[0025] In various embodiments, the memory units 224A-224N may include various types of memory devices including dynamic random access memory (DRAM) or graphics random access memory, such as synchronous graphics random access memory (SGRAM), including graphics double data rate (GDDR) memory. In one embodiment, the memory units 224A-224N may also include 3D stacked memory, including but not limited to high bandwidth memory (HBM). As will be apparent to one skilled in the art, the memory units 224A-224N may vary and be selected from one of a variety of conventional designs. Render targets, such as frame buffers or texture maps, may be stored across the memory units 224A-224N, allowing the partition units 220A-220N to write portions of each render target to efficiently use the available bandwidth of the parallel processor memory 222. In some embodiments, the local instances of the parallel processor memory 222 may be eliminated in favor of a unified memory design that utilizes system memory along with local cache memory.
[0026] In one embodiment, any one of the clusters 214A-214N of the processing cluster array 212 can process data that is to be written to any one of the memory units 224A-224N in the parallel processor memory 222. The memory crossbar 216 can be configured to forward the output of each cluster 214A-214N to any of the partition units 220A-220N or to other clusters 214A-214N that can perform additional processing operations on the output. Each cluster 214A-214N can communicate with a memory interface 218 through the memory crossbar 216 to read from or write to various external memory devices. In one embodiment, the memory crossbar 216 has connections to the memory interface 218 for communicating with the I / O unit 204 and to a local instance of the parallel processor memory 222, allowing processing units in different processing clusters 214A-214N to communicate with system memory or other memory that is not local to the parallel processing units 202. In one embodiment, the memory crossbar 216 can use virtual channels to separate traffic streams between the clusters 214A-214N and the partition units 220A-220N.
[0027] Although a single instance of parallel processing unit 202 is depicted in parallel processor 200, any number of instances of parallel processing unit 202 may be included. For example, multiple instances of parallel processing unit 202 may be provided on a single add-in card, or multiple add-in cards may be interconnected. Different instances of parallel processing unit 202 may be configured to interoperate even if the different instances have different numbers of processing cores, different amounts of local parallel processor memory, and / or other configuration differences. For example, in one embodiment, some instances of parallel processing unit 202 may include higher precision floating point units than other instances. Systems incorporating one or more instances of parallel processing unit 202 or parallel processor 200 may be implemented in a variety of configurations and form factors, including, but not limited to, desktop, laptop, or portable personal computers, servers, workstations, gaming consoles, and / or embedded systems.
[0028] FIG. 2B is a block diagram of partition unit 220 according to an embodiment. In one embodiment, partition unit 220 is an instance of one of partition units 220A-220N of FIG. 2A. As depicted, partition unit 220 includes L2 cache 221, frame buffer interface 225, and raster operations unit (ROP) 226. L2 cache 221 is a read / write cache configured to execute load and store operations received from memory crossbar 216 and ROP 226. Read misses and urgent writeback surfaces are output by L2 cache 221 to frame buffer interface 225 for processing. Updates may also be sent to the frame buffer via frame buffer interface 225 for processing. In one embodiment, frame buffer interface 225 interfaces with one of the memory units in a parallel processor memory, such as memory units 224A-224N (e.g., in parallel processor memory 222) of FIG. 2A.
[0029] In graphics applications, ROP 226 is a processing unit that performs raster operations such as stencil, z-test, blending, etc. ROP 226 then outputs processed graphics data, which is stored in graphics memory. In some embodiments, ROP 226 includes compression logic that compresses depth or color data that is written to memory and decompresses depth or color data that is read from memory. The compression logic can be lossless compression logic that uses one or more of a number of compression algorithms. The type of compression performed by ROP 226 can vary based on the statistical characteristics of the data to be compressed. For example, in one embodiment, delta color compression is performed on depth and color data for each tile.
[0030] In some embodiments, ROP 226 is included within each processing cluster (e.g., clusters 214A-214N of FIG. 2A) rather than within partition unit 220. In such embodiments, read and write requests for pixel data are transmitted through memory crossbar 216 instead of pixel fragment data. The processed graphics data may be displayed on a display device, such as one of one or more display devices 110 of FIG. 1, routed for further processing by processor 102, or routed for further processing by one of the processing entities in parallel processor memory 222 of FIG. 2A.
[0031] FIG. 2C is a block diagram of a processing cluster 214 in a parallel processing unit, according to an embodiment. In one embodiment, the processing cluster is an instance of one of the processing clusters 214A-214N of FIG. 2A. The processing cluster 214 may be configured to execute many threads simultaneously, where the term "thread" refers to an instance of a particular program executing on a particular set of input data. In some embodiments, a single-instruction, multiple-data (SIMD) instruction issue technique is used to support the parallel execution of many threads without providing many independent instruction units. In other embodiments, a single-instruction, multiple-thread (SIMT) technique is used to support the parallel execution of many roughly synchronized threads using a common instruction unit configured to issue instructions to a set of processing engines in each one of the processing clusters. Unlike a SIMD execution regime, where all processing engines typically execute the same instructions, SIMT execution allows different threads to more easily follow different execution paths through a given thread program. Those skilled in the art will appreciate that the SIMD processing regime represents a functional subset of the SIMT processing regime.
[0032] The operation of the processing clusters 214 may be controlled through a pipeline manager 232 that distributes processing tasks to the SIMT parallel processors. The pipeline manager 232 receives instructions from the scheduler 210 of FIG. 2 and manages the execution of these instructions by the graphics multiprocessor 234 and / or the texture unit 236. The depicted graphics multiprocessor 234 is an exemplary instance of a SIMT parallel processor. However, various types of SIMT parallel processors of different architectures may be included in the processing clusters 214. One or more instances of the graphics multiprocessor 234 may be included in the processing clusters 214. The graphics multiprocessor 234 may process data, and the data crossbar 240 may be used to distribute the processed data to one of multiple possible destinations, including other shader units. The pipeline manager 232 may facilitate the distribution of the processed data by specifying the destinations for the processed data to be distributed via the data crossbar 240.
[0033] Each graphics multiprocessor 234 in a processing cluster 214 may include the same set of function execution logic (e.g., arithmetic logic units, load-store units, etc.). The function execution logic may be configured in a pipelined manner where new instructions may be issued before previous instructions are completed. The function execution logic supports a variety of operations including integer and floating point operations, comparison operations, Boolean operations, bit shifts, and calculation of various algebraic functions. In one embodiment, the same functional unit hardware is available to perform the various operations, and any combination of functional units may be present.
[0034] An instruction transmitted to a processing cluster 214 constitutes a thread. A set of threads executing across parallel processing engines is a thread group. A thread group executes the same program on different input data. Each thread in a thread group may be assigned to a different processing engine in the graphics multiprocessor 234. A thread group may contain fewer threads than the number of processing engines in the graphics multiprocessor 234. When a thread group contains fewer threads than the number of processing engines, one or more of the processing engines may be idle during a cycle in which a thread group is processing. A thread group may also contain more threads than the number of processing engines in the graphics multiprocessor 234. When a thread group contains more threads than the number of processing engines in the graphics multiprocessor 234, processing may be executed over consecutive clock cycles. In one embodiment, multiple thread groups may be executed simultaneously in the graphics multiprocessor 234.
[0035] In one embodiment, the graphics multiprocessor 234 includes an internal cache memory to perform load and store operations. In one embodiment, the graphics multiprocessor 234 can defers the internal cache and use a cache memory (e.g., L1 cache 248) in the processing cluster 214. Each graphics multiprocessor 234 also has access to an L2 cache in a partition unit (e.g., partition units 220A-220N of FIG. 2A) that may be shared across all processing clusters 214 and used to transfer data between threads. The graphics multiprocessor 234 may also access off-chip global memory. The off-chip global memory may include one or more of the local parallel processor memory and / or the system memory. Any memory outside of the parallel processing unit 202 may be used as global memory, and embodiments in which the processing cluster 214 includes multiple instances of the graphics multiprocessor 234 can share common instructions and data that may be stored in the L1 cache 248.
[0036] Each processing cluster 214 may include a memory management unit (MMU) 245 configured to map virtual addresses to physical addresses. In other embodiments, one or more instances of MMU 245 may reside in memory interface 218 of FIG. 2A. MMU 245 includes a set of page table entries (PTEs) used to map virtual addresses to physical addresses of tiles and, optionally, cache line indices. MMU 245 may include an address translation lookaside buffer (TLB) or cache that may reside in graphics multiprocessor 234 or L1 cache 248 or processing cluster 214. The physical addresses are processed to distribute the locality of surface data accesses to enable efficient request interleaving among partition units. The cache line index may be used to determine whether a request for a cache line is a hit or a miss.
[0037] For graphics and computing applications, processing clusters 214 may be configured such that each graphics multiprocessor 234 is coupled to a texture unit 236 to perform texture mapping operations, such as determining texture sample locations, retrieving texture data, and filtering the texture data. Texture data is retrieved from an internal texture L1 cache (not shown), or in some embodiments, from an L1 cache within the graphics multiprocessor 234, and fetched as needed from an L2 cache, local parallel processor memory, or system memory. Each graphics multiprocessor 234 outputs processed tasks to data crossbar 240 to provide the processed tasks to other processing clusters 214 for further processing, or to store the processed tasks in an L2 cache, local parallel processor memory, or system memory via memory crossbar 216. A pre-raster operations unit (preROP) 242 is configured to receive data from the graphics multiprocessor 234 and direct the data to the ROP units, which may be located by partition units (e.g., partition units 220A-220N of FIG. 2A) described herein. The preROP 242 unit may perform optimizations for color mixing, organize pixel color data, and perform address conversion.
[0038] It will be understood that the core architecture described herein is illustrative and that modifications and variations are possible. Any number of processing units may be included in the processing clusters 214, such as the graphics multiprocessor 234, texture unit 236, preROP 242, etc. Additionally, while only one processing unit 214 is shown, the parallel processing units described herein may include any number of instances of the processing clusters 214. In one embodiment, each processing cluster 214 may be configured to operate independently of the other processing clusters 214 using separate and distinct processing units, L1 caches, etc.
[0039] 2D illustrates a graphics multiprocessor 234 according to one embodiment. In such an embodiment, the graphics multiprocessor 234 couples with the pipeline manager 232 of the processing cluster 214. The graphics multiprocessor 234 comprises an execution pipeline including, but not limited to, an instruction cache 252, an instruction unit 254, an address mapping unit 256, a register file 258, one or more general purpose graphics processing unit (GPGPU) cores 262, and one or more load / store units 266. The GPGPU cores 262 and the load / store units 266 are coupled to a cache memory 272 and a shared memory 270 via a memory and cache interconnect 268. In one embodiment, the graphics multiprocessor 234 further includes a tensor and / or ray tracing core 263 that includes hardware logic for accelerating matrix and / or ray tracing operations.
[0040] In one embodiment, instruction cache 252 receives a stream of instructions to execute from pipeline manager 232. The instructions are cached in instruction cache 252 and dispatched for execution by instruction unit 254. Instruction unit 254 may dispatch instructions as thread groups (e.g., warps), with each thread of a thread group being assigned to a different execution unit within GPGPU core 262. Instructions may access either local, shared, or global address spaces by specifying addresses within the unified address space. Address mapping unit 256 may be used to translate addresses in the unified address space to different memory addresses that may be accessed by load / store unit 266.
[0041] Register file 258 provides a set of registers for the functional units of graphics multiprocessor 234. Register file 258 provides temporary storage for operands connected to the data paths of the functional units (e.g., GPGPU cores 262, load / store unit 266) of graphics multiprocessor 234. In one embodiment, register file 258 is split among each of the functional units such that each functional unit is assigned a dedicated portion of register file 258. In one embodiment, register file 258 is split among the different warps executed by graphics multiprocessor 234.
[0042] GPGPU cores 262 may each include a floating point unit (FPU) and / or an integer arithmetic logic unit (ALU) used to execute instructions for graphics multiprocessor 234. GPGPU cores 262 may be similar in architecture or may differ in architecture depending on the embodiment. For example, in one embodiment, a first portion of GPGPU cores 262 includes a single precision FPU and an integer ALU, while a second portion of GPGPU cores 262 includes a double precision FPU. In one embodiment, the FPU implements IEEE 754-2008 for floating point arithmetic or may enable variable precision floating point arithmetic. Graphics multiprocessor 234 may further include one or more fixed functions or special functions to perform specific functions, such as copy rectangles or pixel blending operations. In one embodiment, one or more of the GPGPU cores may also include fixed or special function logic.
[0043] In one embodiment, GPGPU core 262 includes SIMD logic that can execute a single instruction on multiple sets of data. In one embodiment, GPGPU core 262 can physically execute SIMD4, SIMD8, and SIMD16 instructions and logically execute SIMD1, SIMD2, and SIMD32 instructions. SIMD instructions for the GPGPU cores can be generated at compile time by a shader compiler or automatically when executing a program written and compiled for a single program multiple data (SPMD) or SIMT architecture. Multiple threads of a program configured for a SIMT execution model can be executed by a single SIMD instruction. For example, in one embodiment, eight SIMD threads performing the same or similar operations can be executed in parallel by a single SIMD8 logic unit.
[0044] Memory and cache interconnect 268 is an interconnect network that connects each of the functional units of graphics multiprocessor 234 to register file 258 and to shared memory 270. In one embodiment, memory and cache interconnect 268 is a crossbar interconnect that allows load / store unit 266 to implement load and store operations between shared memory 270 and register file 258. Because register file 258 can operate at the same frequency as GPGPU core 262, data transfers between GPGPU core 262 and register file 258 have very low latency. Shared memory 270 may be used to enable communication between threads executing on functional units within graphics multiprocessor 234. Cache memory 272 may be used as a data cache, for example, to cache texture data communicated between the functional units and texture unit 236. Shared memory 270 may also be used as a managed program cache. Threads executing on GPGPU cores 262 can programmably store data in the shared memory in addition to the automatically cached data stored in cache memory 272 .
[0045] Figures 3A-3C depict additional graphics multiprocessors according to embodiments. Figures 3A-3B depict graphics multiprocessors 325, 350 that are variations of graphics multiprocessor 234 of Figure 2C. Figure 3C depicts a graphics processing unit (GPU) 380 that includes a dedicated set of graphics processing resources arranged into multicore groups 365A-365N. The depicted graphics multiprocessors 325, 350 and multicore groups 365A-365N may be streaming multiprocessors (SM) capable of simultaneous execution of multiple execution threads.
[0046] 3A illustrates a graphics multiprocessor 325 according to a further embodiment. The graphics multiprocessor 325 includes multiple additional instances of execution resource units relative to the graphics multiprocessor 234 of FIG. 2D. For example, the graphics multiprocessor 325 may include multiple instances of instruction units 332A-332B, register files 334A-334B, and texture units 344A-344B. The graphics multiprocessor 325 also includes multiple sets of graphics or compute execution units (e.g., GPGPU cores 336A-336B, tensor cores 337A-337B, ray tracing cores 338A-338B) and multiple sets of load / store units 340A-340B. In one embodiment, the execution resource units include a common instruction cache 330, texture and / or data cache memory 342, and a shared memory 346.
[0047] The various components may communicate via interconnect fabric 327. In one embodiment, interconnect fabric 327 includes one or more crossbar switches to enable communication between the various components of graphics multiprocessor 325. In one embodiment, interconnect fabric 327 is a separate high-speed network fabric layer on which each component of graphics multiprocessor 325 is stacked. The components of graphics multiprocessor 325 communicate with remote components via interconnect fabric 327. For example, GPGPU cores 336A-336B, 337A-337B, and 338A-338B may each communicate with shared memory 346 via interconnect fabric 327. Interconnect fabric 327 may arbitrate communications within graphics multiprocessor 325 to ensure fair bandwidth allocation between components.
[0048] FIG. 3B illustrates a graphics multiprocessor 350 according to a further embodiment. The graphics multiprocessor 350 includes multiple sets of execution resources 356A-356D, each set including multiple instruction units, register files, GPGPU cores, and load / store units as depicted in FIG. 2D and FIG. 3A. The execution resources 356A-356D can operate in concert with texture units 360-360D for texture operations while sharing an instruction cache 354 and a shared memory 353. In one embodiment, the execution resources 356A-356D can share the instruction cache 354 and the shared memory 353 along with multiple instances of texture and / or data cache memories 358A-358B. The various components can communicate via an interconnect fabric 352 similar to the interconnect fabric 327 of FIG. 3A.
[0049] Those skilled in the art will appreciate that the architectures shown in Figures 1, 2A-2D, and 3A-3B are illustrative and not limiting on the scope of the embodiments. Thus, the techniques described herein may be implemented in any suitably configured processing unit, including, without limitation, one or more mobile application processors, one or more desktop or server central processing units including multi-core CPUs, one or more parallel processing units such as parallel processing unit 202 of Figure 2A, and one or more graphics processors or special purpose processing units, without departing from the scope of the embodiments described herein.
[0050] In some embodiments, a parallel processor or GPGPU as described herein is communicatively coupled to a host / processor core to accelerate graphics operations, machine learning operations, pattern analysis operations, and various general-purpose GPU (GPGPU) functions. The GPU may be communicatively coupled to the host processor / core via a bus or other interconnect (e.g., a high-speed interconnect such as PCIe or NVLink). In other embodiments, the GPU may be integrated in the same package or chip as the core and communicatively coupled to the core via an internal processor bus / interconnect (i.e., in a package or chip). Regardless of the manner in which the GPU is connected, the processor core may assign work to the GPU in the form of a sequence of commands / instructions contained in a work description. The GPU then uses dedicated circuitry / logic to efficiently process those commands / instructions.
[0051] 3C depicts a graphics processing unit (GPU) 380 that includes a dedicated set of graphics processing resources allocated to multi-core groups 365A-N. Although details of only one multi-core group 365A are given, it will be understood that the other multi-core groups 365A-N may be equipped with the same or similar sets of graphics processing resources.
[0052] As depicted, multicore group 365A may include a set of graphics cores 370, a set of tensor cores 371, and a set of ray tracing cores 372. Scheduler / dispatcher 368 schedules and dispatches the graphics thread to execute on the various cores 370, 371, and 372. A set of register files 369 stores operand values used by cores 370, 371, and 372 when executing the graphics thread. These may include, for example, integer registers to store integer values, floating point registers to store floating point values, vector registers to store packed data elements (integer and / or floating point data elements), and tile registers to store tensor / matrix values. In one embodiment, the tile registers are implemented as a combined set of vector registers.
[0053] One or more combined level 1 (L1) cache and shared memory units 373 store graphics data, such as texture data, vertex data, pixel data, ray data, bounding volume data, etc., locally within each multicore group 365A. One or more texture units 374 may also be used to perform texturing operations, such as texture mapping and sampling. A level 2 (L2) cache 375, shared by all or a subset of the multicore groups 365A-365N, stores graphics data and / or instructions for multiple concurrent graphics threads. As shown, the L2 cache 375 may be shared across multiple multicore groups 365A-365N. One or more memory controllers 367 couple the GPU 380 to memory 366. The memory 366 may be system memory (e.g., DRAM) and / or dedicated graphics memory (e.g., GDDR6 memory).
[0054] Input / output (I / O) circuitry 363 couples GPU 380 to one or more I / O devices 362, such as a digital signal processor (DSP), a network controller, or a user input device. An on-chip interconnect may be used to couple I / O devices 362 to GPU 380 and memory 366. One or more I / O memory management units (IOMMUs) 364 of I / O circuitry 363 couple I / O devices 362 directly to system memory 366. In one embodiment, IOMMU 364 manages sets of page tables to map virtual addresses to physical addresses in system memory 366. In this embodiment, I / O devices 362, CPU 361, and GPU 380 may share the same virtual address space.
[0055] In one implementation, IOMMU 364 supports virtualization. In this case, it may manage a first set of page tables that map guest / graphics virtual addresses to guest / graphics physical addresses and a second set of page tables that map guest / graphics physical addresses to system / host physical addresses (e.g., in system memory 366). The base addresses of each of the first and second sets of page tables may be stored in control registers and swapped out at context switch (e.g., so that a new context is given access to the associated set of page tables). Although not shown in FIG. 3C, each of cores 370, 371, 372 and / or multicore groups 365A-365N may include a translation lookaside buffer (TLB) to cache guest virtual to guest physical, guest physical to host physical, and guest virtual to host physical translations.
[0056] In one embodiment, CPU 361, GPU 380, and I / O devices 362 are integrated on a single semiconductor chip and / or chip package. Memory 366 is depicted and may be integrated on the same chip or may be coupled to memory controller 367 via an off-chip interface. In one implementation, memory 366 comprises GDDR6 memory that shares the same virtual address space with other physical system level memory, although the underlying principles of the invention are not limited to this specific implementation.
[0057] In one embodiment, tensor core 371 includes multiple execution units specifically designed to perform matrix operations, which are the basic computational operations used to perform deep learning operations. For example, simultaneous matrix multiplication operations may be used for neural network training and inference. Tensor core 371 may perform matrix processing using a variety of operand precisions, including single precision floating point (e.g., 32 bits), half precision floating point (e.g., 16 bits), integer word (16 bits), byte (8 bits), and half byte (4 bits). In one embodiment, the neural network implementation extracts features of each rendered scene, potentially combining details from multiple frames to compose a high quality final image.
[0058] In a deep learning implementation, parallel matrix multiplication operations may be scheduled for execution on tensor cores 371. Training neural networks, in particular, requires a significant number of matrix dot product operations. To process the dot product formulation of N×N×N matrix multiplication, tensor cores 371 may include at least N dot product processing elements. Before the matrix multiplication begins, an entire matrix is loaded into a tile register, and at least one column of a second matrix is loaded in each of the N cycles. There are N dot products processed per cycle.
[0059] Matrix elements may be stored in different precisions depending on the particular implementation, including 16-bit words, 8-bit bytes (e.g., INT8), and 4-bit half-bytes (e.g., INT4). Different precision modes may be specified for a set of tensor cores 371 to ensure that the most effective precision is used for different workloads (e.g., inferring workloads that can tolerate quantization to bytes and half-bytes, etc.).
[0060] In one embodiment, ray tracing core 372 accelerates ray tracing computations for both real-time and non-real-time ray tracing implementations. In particular, ray tracing core 372 includes ray traverse / intersection circuitry to perform ray traversal using a bounding volume hierarchy (BVH) and identify intersections between rays and primitives contained within the BVH volume. Ray tracing core 372 may also include circuitry for performing depth testing and selection (e.g., using a Z-buffer or similar arrangement). In one implementation, ray tracing core 372 performs traverse and intersection operations in response to image denoising techniques described herein, at least some of which may be performed in tensor core 371. For example, in one embodiment, tensor core 371 implements a deep learning neural network to perform denoising of frames generated by ray tracing core 372. However, the CPU 361, the graphics core 370, and / or the ray tracing core 372 may also implement all or part of the denoising and / or deep learning algorithms.
[0061] Additionally, as described above, a distributed approach to denoising may be used, with GPU 380 in a computing device coupled to other computing devices via a network or high-speed interconnect. In this embodiment, the interconnected computing devices share neural network learning / training data to improve the speed at which the entire system learns to perform denoising for different types of image frames and / or different graphics applications.
[0062] In one embodiment, the ray tracing core 372 handles all BVH traversals and ray primitive intersections to prevent the graphics core 370 from being overloaded with thousands of instructions per ray. In one embodiment, each ray tracing core 372 includes a first set of dedicated circuitry to perform bounding box tests (e.g., for traverse operations) and a second set of dedicated circuitry to perform ray-triangle intersection tests (e.g., which intersecting rays are being traversed). Thus, in one embodiment, the multi-core group 365A can simply initiate a ray probe, and the ray tracing core 372 performs the ray traversal and intersection independently, returning hit data (e.g., hit, no hit, many hits, etc.) to the thread context. While the ray tracing core 372 performs the traverse and intersection operations, the other cores 370, 371 are free to perform other graphics and computational work.
[0063] In one embodiment, each ray tracing core 372 includes a traverse unit that performs BVH test operations and an intersection unit that performs ray-primitive intersection tests. The intersection unit generates a "hit," "no hit," or "many hit" response and provides it to the appropriate thread. During traverse and intersection operations, the execution resources of other cores (e.g., graphics core 370 and tensor core 371) are free to perform other forms of graphics work.
[0064] In one particular embodiment described below, a hybrid rasterization / ray tracing approach is used, with work distributed between the graphics core 370 and the ray tracing core 372.
[0065] In one embodiment, the ray tracing core 372 (and / or other cores 370, 371) includes hardware support for ray tracing instructions such as Microsoft's DirectX Ray Tracing (DXR), including the DispatchRays command, and ray generation, closest hit, any hit, and miss shaders that allow for the assignment of a unique set of shaders and textures per object. Another ray tracing platform that may be supported by the ray tracing core 372, graphics core 370, and tensor core 371 is Vulkab 1.1.85. However, it should be noted that the underlying principles of the present invention are not limited to any particular ray tracing ISA.
[0066] In general, the various cores 372, 371, 370 may support a ray tracing instruction set that includes instructions / functions for ray generation, closest hit, ray-primitive intersection, per-primitive hierarchical bounding box construction, miss, visit, and exceptions. More specifically, one embodiment includes ray tracing instructions for performing the following functions:
[0067] Ray Generation - Ray generation instructions may be executed for each pixel, sample, or other user-defined work allocation.
[0068] Closest Hit - The Closest Hit command may be executed to find the closest intersection of a ray with a primitive in the scene.
[0069] Any Hit - The Any Hit command identifies multiple intersections between rays and primitives in the scene to potentially identify a new closest intersection point.
[0070] Intersection - The intersection instruction performs a ray-primitive intersection test and outputs the result.
[0071] Per-primitive Bounding box Construction - This instruction creates a bounding box around a given primitive or group of primitives (e.g., when creating a new BVH or other acceleration data structure).
[0072] Miss - indicates that the ray misses all geometry in the scene, or a specified region of the scene.
[0073] Visit - indicates the children volumes that the ray is traversing.
[0074] Exceptions - Contains various types of exception handlers (eg, called for various error conditions).
[0075] [Technology in which the GPU hosts the processor interconnect] 4A depicts an exemplary architecture in which multiple GPUs 410-413 are communicatively coupled to multiple multi-core processors 405-406 via high-speed links 440A-440D (e.g., buses, point-to-point interconnects, etc.). In one embodiment, the high-speed links 440A-440D support a communications throughput of 4 GB / s, 30 GB / s, 80 GB / s, or more, depending on the implementation. A variety of interconnect protocols may be used, including, but not limited to, PCIe 4.0 or 5.0 and NVLink 2.0. However, the underlying principles of the present invention are not limited to any particular communications protocol or throughput.
[0076] Further, in one embodiment, two or more of the GPUs 410-413 are interconnected via high speed links 442A-442B, which may be implemented using the same or different protocol / link as used for the high speed links 440A-440D. Similarly, two or more of the multi-core processors 405-406 may be connected via high speed link 443. The high speed link 443 may be a symmetric multi-processor (SMP) bus operating at 20 GB / s, 30 GB / s, 120 GB / s or more. Alternatively, all communications between the various system components shown in FIG. 4A may be achieved using the same protocol / link (e.g., via a communications interconnection fabric). As stated, however, the underlying principles of the present invention are not limited to any particular type of interconnect technology.
[0077] In one embodiment, each multi-core processor 405-406 is communicatively coupled to a processor memory 401-402 via a memory interconnect 430A-430B, respectively, and each GPU 410-413 is communicatively coupled to a GPU memory 420-423 via a GPU memory interconnect 450A-450D, respectively. The memory interconnects 430A-430B and 450A-450D may utilize the same or different memory access technologies. By way of example, and not by way of limitation, the processor memory 401-402 and the GPU memory 420-423 may be volatile memory, such as dynamic random access memory (DRAM) (including stacked DRAM), graphics DDR SDRAM (GDDR) (e.g., GDDR5, GDDR6), high bandwidth memory (HBM), and / or non-volatile memory, such as 3D XPoint or Nano-Ram. In one embodiment, some portions of the memory may be volatile memory and other portions may be non-volatile memory (eg, using a two-level memory (2LM) hierarchy).
[0078] As described below, the various processors 405-406 and GPUs 410-413 may each be physically coupled to a particular memory 401-402, 420-423, but a unified memory architecture may be implemented in which the same virtual system address space (also called the "effective address" space) is distributed among all of the various physical memories. For example, the processor memories 401-402 may each have a system memory address space of 64 GB, and the GPU memories 420-423 may each have a system memory address space of 32 GB (resulting in a total of 254 GB of addressable memory in this example).
[0079] 4B depicts further details of the interconnection between multi-core processor 407 and graphics acceleration module 446, according to one embodiment. Graphics acceleration module 446 may include one or more GPU chips integrated on a line card that is coupled to processor 407 via high-speed link 440. Alternatively, graphics acceleration module 446 may be integrated on the same package or chip as processor 407.
[0080] The depicted processor 407 includes multiple cores 460A-460D, each having a translation lookaside buffer 461A-461D and one or more caches 462A-462D. The cores may include various other components for executing instructions and processing data, but these components are not depicted so as not to obscure the underlying principles of the invention (e.g., instruction fetch units, branch prediction units, decoders, execution units, reordering buffers, etc.). The caches 462A-462D may include a level 1 (L1) and a level 2 (L2). Additionally, one or more shared caches 456 may be included in the cache hierarchy and shared by the set of cores 460A-460D. For example, one embodiment of the processor 407 includes 24 cores, each having its own L1 cache, 12 shared L2 caches, and 12 shared L3 caches. In this embodiment, one of the L2 and L3 caches is shared by two adjacent cores. The processor 407 and the graphics acceleration module 446 connect to a system memory 441, which may include the processor memories 401-402.
[0081] Data and instructions stored in the various caches 462A-462D, 456 and system memory 441 are kept coherent via intercore communication on a coherence bus 464. For example, each cache may have cache coherency logic / circuitry associated with them that communicate on the coherence bus 464 in response to detecting a read or write to a particular cache line. In one implementation, a cache snooping protocol is implemented on the coherence bus 464 to snoop cache accesses. Cache snooping / coherency techniques are well understood by those skilled in the art and will not be described in detail herein so as not to obscure the underlying principles of the present invention.
[0082] In one embodiment, proxy circuit 425 communicatively couples graphics acceleration module 446 to coherence bus 464 to enable graphics acceleration module 446 to participate in cache coherence protocols as a peer of the cores. In particular, interface 435 provides a connection to proxy circuit 425 via high-speed link 440 (e.g., PCIe bus, NVLink, etc.), and interface 437 connects graphics acceleration module 446 to high-speed link 440.
[0083] In one implementation, the accelerator integrated circuit 436 provides cache management, memory access, context management, and interrupt management services on behalf of the multiple graphics processing engines 431, 432, N of the graphics acceleration module 446. The graphics processing engines 431, 432, N may each comprise a separate graphics processing unit (GPU). Alternatively, the graphics processing engines 431, 432, N may comprise different types of graphics processing engines within a GPU, such as graphics execution units, media processing engines (e.g., video encoders / decoders), samplers, and bit engines. In other words, the graphics acceleration module may be a GPU with multiple graphics processing engines 431-432, N, or the graphics processing engines 431-432, N may be individual GPUs integrated on a common package, line card, or chip.
[0084] In one embodiment, accelerator integrated circuit 436 includes a memory management unit (MMU) 439 that performs various memory management functions such as virtual-to-physical memory translation (also referred to as effective-to-real memory translation) and memory access protocols for accessing system memory 441. MMU 439 may also include a translation lookaside buffer (TLB) (not shown) for caching virtual / effective-to-physical / real address translations. In one implementation, cache 438 stores commands and data for effective access by graphics processing engines 431, 432, N. In one embodiment, data stored in cache 438 and graphics memories 433-434, M are kept coherent with core caches 462A-462D, 456 and system memory 441. As discussed, this may be accomplished via a proxy circuit 425 which participates in cache coherency mechanisms on behalf of cache 438 and memory 433-434, M (e.g., sending updates to and receiving updates from cache 438 related to modifications / accesses of cache lines on processor caches 462A-462D, 456).
[0085] A set of registers 445 stores context data for threads executed by the graphics processing engines 431-432,N, and a context management circuit 448 manages thread contexts. For example, the context management circuit 448 may perform save and restore operations to save and restore the context of various threads during a context switch (e.g., a first thread is saved and a second thread is stored so that the second thread can be executed by the graphics processing engine). For example, upon a context switch, the context management circuit 448 may store current register values in a designated area in memory (e.g., identified by a context pointer). It may then restore the register values upon returning to the context. In one embodiment, the interrupt management circuit 447 receives and processes interrupts received from system devices.
[0086] In one implementation, virtual / effective addresses from the graphics processing engine 431 are translated by the MMU 439 to real / physical addresses in the system memory 441. One embodiment of the accelerator integrated circuit 436 supports multiple (e.g., 4, 8, 16) graphics accelerator modules 446 and / or other accelerator devices. The graphics accelerator modules 446 may be dedicated to a single application executing on the processor 407 or may be shared among multiple applications. In one embodiment, a virtualized graphics execution environment is provided in which the resources of the graphics processing engines 431-432, N are shared with multiple applications or virtual machines (VMs). The resources may be subdivided into "slices" that are allocated to different VMs and / or applications based on the processing requirements and priorities associated with the VMs and / or applications.
[0087] Thus, the accelerator integrated circuitry acts as a bridge to the system for the graphics acceleration module 446 and provides address translation and system memory caching services. Additionally, the accelerator integrated circuitry 436 may provide a virtualization facility for the host processor to manage graphics processing engine virtualization, interrupts, and memory management.
[0088] The hardware resources of the graphics processing engines 431-432, N are explicitly mapped into the virtual address space seen by the host processor 407 so that any host processor can directly address those resources using effective address values. One feature of the accelerator integrated circuit 436, in one embodiment, is the physical separation of the graphics engines 431-432, N so that they appear to the system as independent units.
[0089] As mentioned, in the illustrated embodiment, one or more graphics memories 433-434, M are respectively coupled to each of the graphics processing engines 431-432, N. The graphics memories 433-434, M store instructions and data that are processed by each of the graphics processing engines 431-432, N. The graphics memories 433-434, M may be volatile memory such as DRAM (including stacked DRAM), GDDR memory (e.g., GDDR5, GDDR6), or HBM, and / or may be non-volatile memory such as 3D XPoint or Nano-Ram.
[0090] In one embodiment, to reduce data traffic on the high-speed link 440, biasing techniques are used to ensure that the data stored in the graphics memories 433-434, M is the data that is used most frequently by the graphics processing engines 431-432, N, and preferably is not used (at least not frequently) by the cores 460A-460D. Similarly, the biasing mechanism attempts to keep in the cores' caches 462A-462D, 456 and system memory 441 data needed by the cores (and preferably not needed by the graphics processing engines 431-432, N).
[0091] 4C illustrates another embodiment in which accelerator integrated circuit 436 is integrated within processor 407. In this embodiment, graphics processing engines 431-432,N communicate directly with accelerator integrated circuit 436 over high speed link 440 via interface 437 and interface 435 (which again may utilize any type of bus or interface protocol). Accelerator integrated circuit 436 may perform the same operations as described with respect to FIG. 4B, but potentially at a higher throughput given its close proximity to coherency bus 464 and caches 462A-462D, 456.
[0092] One embodiment supports different programming models, including a dedicated process programming model (without graphics acceleration module virtualization) and a shared programming model (with virtualization), which may include a programming model controlled by the accelerator integrated circuitry 436 and a programming model controlled by the graphics acceleration module 446.
[0093] In one embodiment of the dedicated process model, the graphics processing engines 431-432,N are dedicated to a single application or process under a single operating system. A single application can direct other application requests to the graphics engines 431-432,N to provide virtualization within a VM / partition.
[0094] In a dedicated process programming model, the graphics processing engines 431-432,N may be shared by multiple VM / application partitions. The shared model requires the system hypervisor to virtualize the graphics processing engines 431-432,N to allow access by each operating system. For a single partition system without a hypervisor, the graphics processing engines 431-432,N are owned by the operating system. In either case, the operating system can virtualize the graphics processing engines 431-432,N to provide access to each process or application.
[0095] For a shared programming model, the graphics acceleration module 446 or the individual graphics processing engines 431-432,N select a process element using a process handle. In one embodiment, the process element is stored in system memory 441 and is addressable using the effective to real address translation techniques described herein. The process handle may be an implementation-specific value provided to a host process when it registers its context with a graphics processing engine 431-432,N (i.e., calls system software that adds the process element to its process element linked list). The lower 16 bits of the process handle may be the offset of the process element in the process element linked list.
[0096] 4D illustrates an example accelerator integration slice 490. As used herein, a "slice" comprises a designated portion of the processing resources of the accelerator integration circuitry 436. An application effective address space 482 in the system memory 441 stores a process element 483. In one embodiment, the process element 483 is stored in response to a GPU invocation 481 from an application 480 running on the processor 407. The process element 483 includes a process state for the corresponding application 480. A work descriptor (WD) 484 included in the process element 483 can be a single job requested by the application or may include a pointer to a queue of jobs. In the latter case, the WD 484 is a pointer to a job request queue in the application's address space 482.
[0097] The graphics acceleration module 446 and / or the individual graphics processing engines 431-432,N may be shared by all or a subset of the processes in the system. An embodiment of the invention includes an infrastructure to set up process state and send WD484 to the graphics acceleration module 446 to start a job in a virtualized environment.
[0098] In one implementation, a dedicated process programming model is implementation specific, in which a single process owns the graphics acceleration module 446 or the individual graphics processing engine 431. Because the graphics acceleration module 446 is owned by a single process, the hypervisor initializes the accelerator integrated circuit 436 for the owning partition, and the operating system initializes the accelerator integrated circuit 436 for the owning partition at the time the graphics acceleration module 446 is allocated.
[0099] During operation, a WD fetch unit 491 in the accelerator integrated circuit 436 fetches the next WD 484, which contains instructions for work to be performed by one of the graphics processing engines of the graphics acceleration module 446. Data from the WD 484 may be stored in registers 445 and used by the MMU 439, the interrupt management circuit 447, and / or the context management circuit 448, as shown. For example, one embodiment of the MMU 439 includes a segment / page walk circuit for accessing a segment / page table 486 in the OS virtual address space 485. The interrupt management circuit 447 may process interrupt events 492 received from the graphics acceleration module 446. When performing graphics operations, effective addresses 493 generated by the graphics processing engines 431-432,N are converted to real addresses by the MMU 439.
[0100] In one embodiment, the same set of registers 445 may be replicated for each graphics processing engine 431-432, N and / or graphics acceleration module 446 and initialized by the hypervisor or operating system. Each of those replicated registers may be included in the accelerator integration slice 490. Exemplary registers that may be initialized by the hypervisor are shown in Table 1. [Table 1]
[0101] Exemplary registers that may be initialized by the operating system are shown in Table 2. [Table 2]
[0102] In one embodiment, each WD 484 is specific to a particular graphics acceleration module 446 and / or graphics processing engines 431-432, N. It contains all the information that the graphics processing engines 431-432, N need to do its work, or it can be a pointer to a memory location where the application has set up a command queue of work to be completed.
[0103] 4E depicts further details of one embodiment of the sharing model, which includes a hypervisor real address space 498 in which a process element list 499 is stored. The hypervisor real address space 498 is accessible via a hypervisor 496, which virtualizes the graphics acceleration module engine for an operating system 495.
[0104] The shared programming model allows all or a subset of processes from all or a subset of the partitions in the system to use the graphics acceleration module 446. There are two programming models in which the graphics acceleration module 446 is shared by multiple processes and partitions: timeslice sharing and graphics-oriented sharing.
[0105] In this model, the system hypervisor 496 owns the graphics acceleration module 446 and makes its functionality available to all operating systems 495. In order for the graphics acceleration module 446 to support virtualization by the system hypervisor 496, the graphics acceleration module 446 may adhere to the following requirements: 1) the application's job requests must be autonomous (i.e., state does not need to be maintained between jobs) or the graphics acceleration module 446 must provide a context save and restore mechanism; 2) the application's job requests must be guaranteed by the graphics acceleration module 446 to be completed in a specified amount of time including any conversion errors or the graphics acceleration module 446 must provide the ability to preempt the processing of jobs; and 3) the graphics acceleration module 446 must ensure fairness between processes when operating in a directional sharing programming model.
[0106] In one embodiment, for the shared model, application 480 is required to make a system call to operating system 495 with the graphics acceleration module 446 type, a working descriptor (WD), an authority mask register (AMR) value, and a context save / restore area pointer (CSRP). The graphics acceleration module 446 type describes the acceleration function that the system call is directed to. The graphics acceleration module 446 type may be a system specific value. The WD is formatted specifically for graphics acceleration module 446 and may take the form of a graphics acceleration module 446 command, an effective address pointer to a user defined structure, an effective address pointer to a queue of commands, or some other data structure that describes the work to be done by graphics acceleration module 446. In one embodiment, the AMR value is the AMR state to be used for the current process. The value passed to the operating system is the same as the application setting the AMR. If the accelerator integrated circuit 436 and graphics acceleration module 446 implementations do not support a User Authority Mask Override Register (UAMOR), the operating system may apply the current UAMOR value to the AMR value before passing the AMR in the hypervisor call. The hypervisor 496 may optionally apply the current Authority Mask Override Register (AMOR) value before placing the AMR in the process element 483. In one embodiment, the CSRP is one of the registers 445 that contains the effective address of an area in the application's address space 482 for the graphics acceleration module 446 to save and restore context state.This pointer is optional in case state does not need to be saved between jobs or if a job is preempted. The context save / restore area may be pinned system memory.
[0107] Upon receiving the system call, operating system 495 may verify that application 480 is registered and authorized to use graphics acceleration module 446. Operating system 495 then calls hypervisor 496 with the information shown in Table 3. [Table 3]
[0108] Upon receiving the hypervisor call, the hypervisor 496 verifies that the operating system 495 is registered and authorized to use the graphics acceleration module 446. The hypervisor 496 then places the process element 483 in a process element linked list for the corresponding type of graphics acceleration module 446. The process element may include the information shown in Table 4. [Table 4]
[0109] In one embodiment, the hypervisor initializes slice 445 of multiple accelerator integration slices 490.
[0110] As depicted in FIG. 4F, one embodiment of the present invention employs a unified memory addressable via a common virtual memory address space that is used to access the physical processor memories 401-402 and the GPU memories 420-423. In this implementation, operations performed on the GPUs 410-413 utilize the same virtual / effective memory address space to access the processor memories 401-402, and vice versa, thereby simplifying programmability. In one embodiment, a first portion of the virtual / effective address space is assigned to the processor memory 401, a second portion is assigned to the second processor memory 402, a third portion is assigned to the GPU memory 420, and so on. The entire virtual / effective memory space (sometimes referred to as the effective address space) is thereby distributed across each of the processor memories 401-402 and the GPU memories 420-423, allowing any processor or GPU to access any physical memory using the virtual addresses that are mapped to that memory.
[0111] In one embodiment, bias / coherence management circuits 494A-494E in one or more of the MMUs 439A-439E ensure cache coherence between the caches of the host processors (e.g., 405) and the GPUs 410-413 and implement biasing techniques that indicate the physical memory in which particular types of data should be stored. While multiple instances of bias / coherence management circuits 494A-494E are depicted in FIG. 4F, bias / coherence circuits may be implemented within the MMUs of one or more of the host processors 405 and / or within the accelerator integrated circuitry 436.
[0112] One embodiment allows the GPU-attached memory 420-423 to be mapped as a portion of system memory and accessed using shared virtual memory (SVM) techniques, but without suffering the typical performance drawbacks associated with full system cache coherence. The ability of the GPU-attached memory 420-423 to be accessed as system memory without cumbersome cache coherence overhead provides a favorable operating environment for GPU offload. This arrangement allows software on the host processor 405 to set up operands and access computation results without the overhead of traditional I / O DMA data copies. Such traditional copies require driver calls, interrupts, and memory mapped I / O (MMIO) accesses, all of which are inefficient for simple memory accesses. At the same time, the ability to access the GPU-attached memory 420-423 without cache coherence overhead can be critical to the execution time of the offloaded computation. When streaming write memory traffic is substantial, for example, cache coherence overhead can significantly reduce the effective write bandwidth seen by GPUs 410-413. The efficiency of operand setup, the efficiency of result access, and the efficiency of GPU computation all play a role in determining the effectiveness of GPU offload.
[0113] In one implementation, the selection between GPU bias and host processor bias is driven by a bias tracker data structure. For example, a bias table may be used. The bias table may be a page granular structure containing one or two bits per GPU attached memory page (i.e., controlled at the granularity of a memory page). The bias table may be implemented in stolen memory ranges of one or more GPU attached memories 420-423, with or without a bias cache in the GPU 410-413 (e.g., for caching frequently / recently used entries of the bias table). Alternatively, the entry bias table may be maintained within the GPU.
[0114] In one implementation, the bias table entry associated with each access to the GPU attached memory 420-423 is accessed before the actual access to the GPU memory, causing the following actions: First, local requests from the GPUs 410-413 to find their pages in the GPU bias are forwarded directly to the corresponding GPU memory 420-423. Local requests from the GPUs to find their pages in the host bias are forwarded to the processor 405 (e.g., over the high-speed link described above). In one embodiment, a request from the processor 405 to find the requested page in the host processor bias completes the request like a normal memory read. Alternatively, requests directed to the GPU bias page may be forwarded to the GPUs 410-413. The GPU may then migrate the page to the host processor if the host processor is not currently using the page.
[0115] The bias state of a page can be changed by a software-based mechanism or a hardware-assisted software-based mechanism, or for a limited set of cases, purely by a hardware-based mechanism.
[0116] One mechanism for changing the bias state uses an API call (e.g., OpenCL), which then calls the GPU's device driver, which then sends a message (or enqueues a command descriptor) to the GPU instructing it to change the bias state, and for some transitions, performs a cache flush operation in the host. A cache flush operation is required for transitions from host processor 405 bias to GPU bias, but not for the reverse transition.
[0117] In one embodiment, cache coherency is maintained by making GPU bias pages temporarily non-cacheable by host processor 405. To access those pages, processor 405 may request an access from GPU 410 which may or may not immediately grant the access, depending on the implementation. Thus, to reduce communication between host processor 405 and GPU 410, it is beneficial to ensure that GPU bias pages are pages that are needed by the GPU but not by host processor 405, and vice versa.
[0118] [Graphics Processing Pipeline] FIG. 5 illustrates a graphics processing pipeline 500 according to an embodiment. In one embodiment, a graphics processor can implement the illustrated graphics processing pipeline 500. The graphics processor can be included within a parallel processing subsystem described herein, such as parallel processor 200 of FIG. 2A, which in one embodiment is a variation of parallel processor 112 of FIG. 1. Various parallel processing systems can implement graphics processing pipeline 500 via one or more instances of a parallel processing unit as described herein (e.g., parallel processing unit 202 of FIG. 2A). For example, a shader unit (e.g., graphics multiprocessor 234 of FIG. 2C) can be configured to perform the functions of one or more of vertex processing unit 504, tessellation control processing unit 508, tessellation evaluation processing unit 512, geometry processing unit 516, and fragment / pixel processing unit 524. The functions of data assembler 502, primitive assemblers 506, 514, 518, tessellation unit 510, rasterizer 522, and raster operation unit 526 may also be performed by other processing engines within a processing cluster (e.g., processing cluster 214 of FIG. 2A) and corresponding partition unit (e.g., partition units 220A-220N of FIG. 2A). Graphics processing pipeline 500 may also be implemented using dedicated processing units for one or more functions. In one embodiment, one or more portions of graphics processing pipeline 500 may be performed by parallel processing logic within a general-purpose processor (e.g., a CPU). In one embodiment, one or more portions of graphics processing pipeline 500 may access on-chip memory (e.g., parallel processor memory as seen in FIG. 2A) via memory interface 528, which may be an instance of memory interface 218 of FIG. 2A.
[0119] In one embodiment, data assembler 502 is a processing unit that assembles vertex data for surfaces and primitives. Data assembler 502 then outputs vertex data, including vertex attributes, to vertex processing unit 504. Vertex processing unit 504 is a programmable execution unit that runs vertex shader programs, and lights and transforms the vertex data specified by the vertex shader programs. Vertex processing unit 504 may be programmed to read data stored in cache, local, or system memory for use in processing the vertex data, and to transform the vertex data from an object-based coordinate representation to a world space coordinate space or a normalized device coordinate space.
[0120] A first instance of primitive assembler 506 receives vertex attributes from vertex processing unit 504. Primitive assembler 506 reads stored vertex attributes as necessary and constructs graphics primitives for processing by tessellation control processing unit 508. Graphics primitives include triangles, line segments, points, patches, etc. that are supported by various graphics processing application programming interfaces (APIs).
[0121] The tessellation control processing unit 508 treats the input vertices as control points for the geometric patch. The control points are converted from an input representation from the patch (e.g., patch-based) to a representation suitable for use in appropriate evaluation by the tessellation evaluation processing unit 512. The tessellation control processing unit 508 may also calculate tessellation coefficients for edges of the geometric patch. The tessellation coefficients are applied to a single edge to quantify a view-dependent level of detail associated with the edge. The tessellation unit 510 is configured to receive the tessellation coefficients for the edges of the patch and tessellate the patch into a number of geometric primitives, such as line, triangle, or quadrilateral primitives. The number of geometric primitives are sent to the tessellation evaluation processing unit 512. The tessellation evaluation processing unit 512 operates on the parameterized coordinates of the subdivided patch to generate a surface representation and vertex attributes for each vertex associated with the geometric primitive.
[0122] A second instance of primitive assembler 514 receives vertex attributes from tessellation evaluation processing unit 512, which reads stored vertex attributes as necessary, and constructs graphics primitives for processing by geometry processing unit 516. Geometry processing unit 516 is a programmable execution unit that executes a geometry shader program to transform the graphics primitives received from primitive assembler 514 as specified by the geometry shader program. In one embodiment, geometry processing unit 516 is programmed to subdivide the graphics primitives into one or more new graphics primitives and calculate parameters used to rasterize the new graphics primitives.
[0123] In some embodiments, geometry processing unit 516 can add or remove elements in the geometry stream. Geometry processing unit 516 outputs parameters and vertices that specify new graphics primitives to primitive assembler 518. Primitive assembler 518 receives parameters and vertices from geometry processing unit 516 and constructs graphics primitives for processing by viewport scale, cull, and clip unit 520. Geometry processing unit 516 reads data stored in parallel processor memory or system memory for use in processing the geometry data. Viewport scale, cull, and clip unit 520 performs clipping, culling, and viewport scaling and outputs processed graphics primitives to rasterizer 522.
[0124] The rasterizer 522 can perform depth sculling and other depth-based optimizations. The rasterizer 522 also performs scan conversion on new graphics primitives to generate fragments and output these fragments and associated coverage data to the fragment / pixel processing unit 524. The fragment / pixel processing unit 524 is a programmable execution unit configured to execute a fragment shader program or a pixel shader program. The fragment / pixel processing unit 524 transforms fragments or pixels received from the rasterizer 522 as specified by the fragment or pixel shader program. For example, the fragment / pixel processing unit 524 may be programmed to perform operations including, but not limited to, texture mapping, shading, blending, texture correction, and perspective correction to generate shaded fragments or pixels that are output to the raster operation unit 526. The fragment / pixel processing unit 524 can read data stored in either the parallel processor memory or the system memory for use when processing the fragment data. The fragment or pixel shader programs may be configured to perform shading at a sample, pixel, tile, or other granularity depending on the sampling rate configured for the processing unit.
[0125] Raster operations unit 526 is a processing unit that performs raster operations, including but not limited to stencil, z-test, blending, etc., and outputs pixel data as processed graphics data to be stored in a graphics memory (e.g., parallel processor memory 222 of FIG. 2A and / or system memory 104 of FIG. 1), to be displayed on one or more display devices 110, or for further processing by one or more of processors 102 or parallel processors 112. In some embodiments, raster operations unit 526 is configured to compress z or color data that is written to memory and to decompress z or color data that is read from memory.
[0126] [Machine Learning Overview] The above architectures can be applied to perform training and inference operations using machine learning models. Machine learning has been successful in solving many types of tasks. The computations that occur when training and using machine learning algorithms (e.g., neural networks) naturally lend themselves to efficient parallel implementation. Thus, parallel processors such as general-purpose graphic processing units (GPGPUs) play an important role in the practical implementation of deep neural networks. Parallel graphics processors with single instruction, multiple thread (SIMT) architecture are designed to maximize the amount of parallelism in the graphics pipeline. In a SIMT architecture, a group of parallel threads attempts to execute program instructions synchronously together as often as possible to increase processing efficiency. The efficiency provided by parallel machine learning algorithm implementations allows the use of high-capacity networks, allowing those networks to be trained on larger datasets.
[0127] Machine learning algorithms can learn based on a set of data. An embodiment of a machine learning algorithm can be designed to model high-level abstractions within a data set. For example, an image recognition algorithm can be used to determine which of several categories a given input belongs to, a regression algorithm can output a numerical value given an input, and a pattern recognition algorithm can be used to generate transformed text or perform speech recognition and / or speech recognition from text.
[0128] An exemplary type of machine learning algorithm is a neural network. There are many kinds of neural networks, and a simple type of neural network is a feedforward network. A feedforward network may be implemented as an acyclic graph with nodes arranged in layers. Typically, a feedforward network topology includes an input layer and an output layer separated by at least one hidden layer. The hidden layer converts the input received by the input layer into a representation that is useful for generating an output at the output layer. The network nodes are fully connected via edges to nodes in adjacent layers, but there are no edges between nodes in each layer. Data received at nodes in the input layer of a feedforward network is propagated (i.e., sent forward) to nodes in the output layer by an activation function that calculates the state of the nodes in each successive layer in the network based on coefficients ("weights") respectively associated with each of the edges connecting the layers. Depending on the specific model represented by the algorithm being executed, the output from a neural network algorithm can take a variety of forms.
[0129] Before a machine learning algorithm can be used to model a particular problem, the algorithm is trained using a training data set. Training a neural network requires selecting a network topology, using a set of training data that represents the problem to be modeled by the network, and adjusting the weights until the network model performs with minimal error for all instances of the training data set. For example, during a supervised learning training process for a neural network, the output generated by the network in response to an input representing an instance in the training data set is compared with an output labeled as "correct" for that instance, an error signal representing the difference between the output and the labeled output is calculated, and the weights associated with the connections are adjusted to minimize that error, with the error signal propagating backward through the network layers. The network is considered "trained" when the error for each of the outputs generated from the instances of the training data set is minimized.
[0130] The accuracy of machine learning algorithms can be greatly affected by the quality of the data set used to train the algorithm. The training process is computationally intensive and can require a significant amount of time on a conventional general-purpose processor. Therefore, parallel processing hardware is used to train many types of machine learning algorithms. This is particularly useful for optimizing the training of neural networks, since the calculations performed in adjusting coefficients in neural networks naturally lend themselves to parallel implementation. In particular, many machine learning algorithms and software applications are adapted to use parallel processing hardware in general-purpose graphics processing devices.
[0131] 6 is a generalized diagram of a machine learning software stack 600. A machine learning application 602 may be configured to train a neural network using a training dataset or to implement machine learning intelligence using a trained deep neural network. The machine learning application 602 may include training and inference functions for the neural network and / or specialized software that may be used to train the neural network prior to deployment. The machine learning application 602 may implement any type of machine intelligence, including, but not limited to, image recognition, mapping and localization, autonomous navigation, speech synthesis, medical imaging, or language conversion.
[0132] Hardware acceleration for the machine learning application 602 can be enabled through the machine learning framework 604. The machine learning framework 604 can provide a library of machine learning primitives. Machine learning primitives are basic operations typically performed by machine learning algorithms. Without the machine learning framework 604, developers of machine learning algorithms would be required to create and optimize the main computation logic associated with the machine learning algorithm and then reoptimize the computation logic when new parallel processors are developed. Instead, machine learning applications can be configured to perform the necessary computations using primitives provided by the machine learning framework 604. Exemplary primitives include tensor convolution, activation functions, and pooling, which are computation operations performed while training a convolutional neural network (CNN). The machine learning framework 604 can also provide primitives to implement basic linear algebra subprograms performed by many machine learning algorithms, such as matrix and vector operations.
[0133] The machine learning framework 604 can process input data received from the machine learning application 602 and generate appropriate inputs to the computation framework 606. The computation framework 606 can abstract the underlying instructions provided to the GPGPU driver 608 to enable the machine learning framework 604 to take advantage of hardware acceleration via the GPGPU hardware 610 without requiring the machine learning framework 604 to have in-depth knowledge of the architecture of the GPGPU hardware 610. Furthermore, the computation framework 606 can enable hardware acceleration for the machine learning framework 604 across different types and generations of GPGPU hardware 610.
[0134] [GPGPU Machine Learning Acceleration] 7 illustrates a general purpose graphics processing unit 700, according to an embodiment. In one embodiment, the general purpose processing unit (GPGPU) 700 may be configured to be particularly efficient at processing the types of computational workloads associated with training deep neural networks. Moreover, the GPGPU 700 may be directly linked to other instances of GPGPUs to create multi-GPU clusters to improve training speeds, particularly for deep neural networks.
[0135] The GPGPU 700 includes a host interface 702 that allows for connection to a host processor. In one embodiment, the host interface 702 is a PCI Express interface. However, the host interface can also be a vendor-specific communication interface or fabric. The GPGPU 700 receives commands from the host processor and uses a global scheduler 704 to distribute execution threads associated with those commands to a set of compute clusters 706A-706H. The compute clusters 706A-706H share a cache memory 708. The cache memory 708 can act as a higher level cache for the cache memories in the compute clusters 706A-706H.
[0136] The GPGPU 700 includes memories 714A-B coupled to the compute clusters 706A-H via a set of memory controllers 712A-B. In various embodiments, the memories 714A-B can include various types of memory devices including dynamic random access memory (DRAM) or graphics random access memory, such as synchronous graphics random access memory (SGRAM), including graphics double data rate (GDDR) memory. In one embodiment, the memories 714A-714B can also include 3D stacked memory, including but not limited to high bandwidth memory (HBM).
[0137] In one embodiment, each of the compute clusters 706A-H includes a set of graphics multiprocessors, such as the graphics multiprocessor 400 of FIG. 4A. The graphics multiprocessors of a compute cluster include multiple types of integer and floating point logic units capable of performing computational operations with a range of precisions, including those suitable for machine learning computations. For example, in one embodiment, at least a subset of the floating point units in each of the compute clusters 706A-H may be configured to perform 16-bit or 32-bit floating point operations, while another subset of the floating point units may be configured to perform 64-bit floating point operations.
[0138] Multiple instances of GPGPU 700 may be configured to operate as a compute cluster. The communication mechanisms used by the compute cluster for synchronization and data exchange vary from embodiment to embodiment. In one embodiment, multiple instances of GPGPU 700 communicate through host interface 702. In one embodiment, GPGPU 700 includes an I / O hub 709 that couples GPGPU 700 with GPU link 710 that allows direct connection to other instances of GPGPU. In one embodiment, GPU link 710 is coupled to a dedicated GPU-to-GPU bridge that allows communication and synchronization between multiple instances of GPGPU 700. In one embodiment, GPU link 710 couples to a high-speed interconnect to send and receive data to other GPGPUs or parallel processors. In one embodiment, multiple instances of GPGPU 700 communicate through a network device located in another data processing system and accessible by host interface 702. In one embodiment, GPU link 710 may be configured to allow connection to a host processor in addition to or instead of host interface 702.
[0139] While the depicted configuration of GPGPU 700 may be configured to train a neural network, one embodiment provides an alternative configuration of GPGPU 700 that may be configured for deployment in a high performance or low power inference platform. In the inference configuration, GPGPU 700 includes fewer compute clusters 706A-706H as compared to the training configuration. Furthermore, the memory technology associated with memories 714A-714B may differ between the inference and training configurations. In one embodiment, the inference configuration of GPGPU 700 may support inferring certain instructions. For example, the inference configuration may support one or more 8-bit integer dot product instructions that are commonly used during inference operations for deployed neural networks.
[0140] FIG. 8 illustrates a multi-GPU computing system 800 according to an embodiment. The multi-GPU computing system 800 can include a processor 802 coupled to multiple GPGPUs 806A-806B via a host interface switch 804. The host interface switch 804, in one embodiment, is a PCI Express switch device that couples the processor 802 to a PCI Express bus. Through the PCI Express bus, the processor 802 can communicate with a set of GPGPUs 806A-806D. Each of the multiple GPGPUs 806A-806D can be an instance of GPGPU 700 of FIG. 7. The GPGPUs 806A-806D can be interconnected via a set of high-speed point-to-point inter-GPU links 816. The high-speed inter-GPU links can be connected to each of the GPGPUs 806A-806D via a dedicated GPU link, such as GPU link 710 of FIG. 7. The P2P GPU link 816 enables direct communication between each of the GPGPUs 806A-806D without requiring communication over the host interface bus to which the processor 802 is connected. With inter-GPU traffic directed to the P2P GPU link, the host interface bus remains available for system memory access or to communicate with other instances of the multi-GPU computing system 800, for example, via one or more network devices. While the depicted embodiment of the GPGPUs 806A-D connect to the processor 802 through the host interface switch 804, in one embodiment, the processor 802 includes direct support for the P2P GPU link 816 and can connect directly to the GPGPUs 806-806D.
[0141] [Machine learning neural network implementation] The computing architecture provided by the embodiments described herein can be configured to perform a type of parallel processing that is particularly suitable for training and deploying neural networks for machine learning. A neural network can be generalized as a network of functions with graph relationships. As is well known in the art, there are various types of neural network implementations used in machine learning. One exemplary type of neural network is a feedforward network, as described above.
[0142] A second exemplary type of neural network is the convolutional neural network (CNN). CNNs are specialized feed-forward neural networks for processing data with a known grid-like topology, such as image data. Thus, while CNNs are widely used for computer vision and image recognition applications, they may also be used for other types of pattern recognition, such as speech and language processing. Nodes in a CNN input layer are organized into sets of "filters" (feature detectors evoked by receptive fields in the retina), and the output of each set of filters is propagated to nodes in successive layers of the network. A CNN's computation involves applying a convolution mathematical operation to each filter to generate the filter's output. Convolution is a specialized type of mathematical operation performed by two functions to generate a third function that is a modified version of one of the two original functions. In convolutional network terminology, the first function of the convolution may be called the input, while the second function may be called the convolution kernel. The output may be called a feature map. For example, the input to a convolutional layer may be a multidimensional array of data that defines the various color components of the input image. The convolution kernel can be a multidimensional array of parameters, where the parameters are adapted by a training process for the neural network.
[0143] Recurrent neural networks (RNNs) are a family of feedforward neural networks that contain feedback connections between layers. RNNs allow for the modeling of sequential data by sharing parameter data across different parts of the neural network. The architecture of an RNN contains cycles, which represent the effect that a variable's current value has on its own value at a future time, such that at least a portion of the output data from the RNN is used as feedback to process subsequent inputs in sequence. This feature makes RNNs particularly useful for language processing due to the mutable nature in which language data can be constructed.
[0144] The figures below present exemplary feedforward, CNN, and RNN networks and describe the general process for training and deploying each of these types of networks, respectively. These descriptions are exemplary and non-limiting with respect to any specific embodiments described herein, and the concepts presented may be generally applied to deep neural networks and machine learning techniques in general.
[0145] The above exemplary neural network can be used to perform deep learning. Deep learning is machine learning that uses deep neural networks. The deep neural networks used in deep learning are artificial neural networks that consist of multiple hidden layers, as opposed to shallow neural networks that only contain a single hidden layer. Deeper neural networks generally require more computational load to train. However, the additional hidden layers of the network enable multi-step pattern recognition that results in reduced output error relative to shallow machine learning techniques.
[0146] A deep neural network used in deep learning typically includes a front-end network for performing feature recognition coupled to a back-end network representing a mathematical model that can perform operations (e.g., object classification, speech recognition, etc.) based on feature representations given to the model. Deep learning allows machine learning to be performed without the need for hand-crafted feature engineering to be performed on the model. Instead, deep neural networks can learn features based on statistical structures or correlations in the input data. The learned features can be fed into a mathematical model that can map the detected features to an output. The mathematical model used by the network will generally be specialized for a particular task to be performed, and different models will be used to perform different tasks.
[0147] Once a neural network is structured, a learning model can be applied to the network to train it to perform a particular task. The learning model describes how weights in the model should be adjusted to reduce the network's output error. Backward error propagation is a common model used to train neural networks. An input bell is given to the network for processing. The output of the network is compared to the desired output using a loss function, and an error value is calculated for each of the neurons in the output layer. The error values are then propagated backwards until each neuron has an associated error value that roughly represents its contribution to the original output. The network can then learn from those errors using algorithms such as the stochastic gradient descent algorithm to update the neural network weights.
[0148] 9A-9B depict an exemplary convolutional neural network. FIG. 9A depicts various layers within a CNN. As shown in FIG. 9A, an exemplary CNN used to model image processing can receive an input 902 describing the red, green, and blue (RGB) components of an input image. The input 902 can be processed by multiple convolutional layers (e.g., convolutional layer 904, convolutional layer 906). The output from the multiple convolutional layers can optionally be processed by a set of fully connected layers 908. Neurons in a fully connected layer have full connections to all activations in the previous layer, as described above for feedforward networks. The output from the fully connected layer 908 can be used to generate an output result from the network. The activations in the fully connected layer 908 can be calculated using matrix multiplication instead of convolution. Not all CNN implementations utilize fully connected layers 908. For example, in some implementations, the convolutional layers 906 can generate the output of the CNN.
[0149] Convolutional layers are loosely connected, which differs from traditional neural network configurations found in the fully connected layer 908. Traditional neural network layers are fully connected, whereby every output unit interacts with every input unit. However, convolutional layers are loosely connected because, as depicted, the output of the convolution of the field is input to the nodes of the subsequent layer (instead of the state values of each of the nodes in the field). The kernels associated with the convolutional layer perform the convolution operation, and the output is sent to the next layer. The dimensionality reduction performed in the convolutional layers is one aspect that allows CNNs to scale to process large images.
[0150] 9B illustrates exemplary computational stages within a convolutional layer of a CNN. An input 912 to a convolutional layer of a CNN may be processed by three stages of convolutional layers 914. The three stages may include a convolutional stage 916, a detector stage 918, and a pooling stage 920. The convolutional layers 914 may output data to successive convolutional layers. The final convolutional layer of the network may generate output feature map data or may provide input to a fully connected layer to generate, for example, a classification value for the input to the CNN.
[0151] The convolution stage 916 performs several convolutions in parallel to generate a set of linear activations. The convolution stage 916 can include an affine transformation. An affine transformation is any transformation that can be specified as a linear transformation plus a translation. Affine transformations include rotation, translation, scaling, and combinations of these transformations. The convolution stage calculates the output of a function (e.g., a neuron) that is connected to a particular region in the input, which can be determined as a local region associated with the neuron. The neuron calculates a dot product between the weight of the neuron and the region in the local input to which the neuron is connected. The output from the convolution stage 916 defines a set of linear activations that are processed by subsequent stages of the convolution layer 914.
[0152] The linear activations may be processed by a detector stage 918. In the detector stage 918, the nonlinear activations are processed by a nonlinear activation function. The nonlinear activation function enhances the nonlinear characteristics of the entire network without affecting the fields of each of the convolutional layers. Several types of nonlinear activation functions may be used. One particular type is the rectified linear unit (ReLU), which uses an activation function defined as f(x)=max(0,x) such that the activation is thresholded at zero.
[0153] The pooling stage 920 uses a pooling function that replaces the output of the convolutional layer 906 with a summary statistic of nearby outputs. The pooling function may be used to introduce translation invariance into the neural network so that small translations to the input do not change the pooled output. Invariance to local translations may be useful in scenarios where the presence of a feature in the input data is more important than the exact location of the feature. Various types of pooling functions may be used during the pooling stage 920, including max pooling, average pooling, and 12-norm pooling. Furthermore, some CNN implementations do not include a pooling stage. Instead, such implementations substitute an additional convolutional stage with an increased stride relative to the previous convolutional layer.
[0154] The output from the convolutional layer 914 may then be processed by a next layer 922, which may be an additional convolutional layer or one of the fully connected layers 908. For example, the first convolutional layer 904 in FIG. 9A may output to a second convolutional layer 906, which in turn may output to the first of the fully connected layers 908.
[0155] FIG. 10 depicts an example recurrent neural network 100. In a recurrent neural network (RNN), the previous state of the network influences the output of the current state of the network. RNNs can be constructed in a variety of ways with a variety of functions. The use of RNNs generally revolves around using mathematical models to predict the future based on previous input sequences. For example, RNNs may be used to perform statistical language modeling to predict upcoming words given previous word sequences. The depicted RNN 1000 can be described as having an input layer 1002 that receives input vectors, a hidden layer 1004 that implements a regression function, a feedback mechanism 1005 that allows 'memory' of previous states, and an output layer 1006 that outputs the results. The RNN 1000 operates on a time-step basis. The state of the RNN at a given time-step is influenced based on the previous time-step via the feedback mechanism 1005. For a given time-step, the state of the hidden layer 1004 is defined by the previous state and the input at the current time-step. A first input (x1) at a first time step may be processed by the hidden layer 1004. A second input (x2) may be processed by the hidden layer 1004 using state information determined during the processing of the first input (x1). A given state is s t =f(Ux t +Ws t-1 ), where U and W are parameter matrices. The function f is typically a nonlinearity, such as a hyperbolic tangent function (Tanh) or a rectified linear function f(x)=max(0,x). However, the specific mathematical functions used in the hidden layer 1004 can vary depending on the specific implementation details of the RNN 1000.
[0156] In addition to the basic CNN and RNN networks described, variations on those networks may be possible. One example of a variation of an RNN is the long short term memory (LSTM) RNN. LSTM RNNs are capable of learning long-term dependencies that may be necessary to process longer language sequences. A variation of a CNN is the convolutional deep belief network, which has a similar structure to a CNN and is trained in a similar manner to a deep belief network. A deep belief network (DBN) is a generative neural network consisting of multiple layers of probabilistic (random) variables. DBNs may be trained layer-by-layer using unsupervised greedy learning. The learned weights of the DBN may then be used to provide a pre-trained neural network by determining an optimal initial set of weights for the neural network.
[0157] FIG. 11 illustrates the training and deployment of a deep neural network. Once a given network is structured for a task, the neural network is trained using a training dataset 1102. Various training frameworks 1104 have been developed to enable hardware acceleration of the training process. For example, the machine learning framework 604 of FIG. 6 may be configured as the training framework 1104. The training framework 1104 can lead to an untrained neural network 1106 and enable the untrained neural network 1106 to be trained using parallel processing resources as described herein to generate a trained neural network 1108.
[0158] To begin the training process, initial weights may be selected randomly or by pre-training with a deep belief network. Training cycles are then performed in either a supervised or unsupervised manner.
[0159] Supervised learning is a learning method in which training is performed as a mediated operation, for example when the training data set 1112 includes inputs paired with desired outputs for the inputs, or when the training data set includes inputs with known outputs and the neural network's output is manually graded. The network processes the inputs and compares the resulting output to a set of expected or desired outputs. Errors are then propagated back through the system. The training framework 1104 can adjust the weights that control the untrained neural network 1106. The training framework 1104 can provide tools to monitor how well the untrained neural network 1106 is converging towards a model suitable for generating correct answers based on known input data. The training process is iterative, such that the network weights are adjusted to refine the outputs generated by the neural network. The training process can continue until the neural network reaches a statistically desired accuracy associated with the trained neural network 1108. The trained neural network 1108 may then be deployed to implement any number of machine learning operations to generate inference results 1114 based on the input of new data 1102.
[0160] Unsupervised learning is a learning model in which the network attempts to train itself using unlabeled data. Thus, for unsupervised learning, the training dataset 1112 will include inputs without any associated output data. The untrained neural network 1106 can learn groupings within the unlabeled inputs and determine how individual inputs relate to the overall dataset. Unsupervised training can be used to generate self-organizing maps, which are a type of trained neural network 1108 that can perform operations useful for reducing the dimensionality of data. Unsupervised training can also be used to perform anomaly detection, which allows for the identification of data points within an input dataset that fall outside of normal patterns of data.
[0161] Variations of supervised and unsupervised training may also be used. Semi-supervised learning is a technique in which the training dataset 1112 contains a mixture of labeled and unlabeled data of the same distribution. Incremental learning is a variation of supervised learning in which input data may be used successively to further train the model. Incremental learning allows the trained neural network 1108 to adapt to new data 1112 without forgetting the knowledge implanted in the network during the initial training.
[0162] The training process, especially for deep neural networks, whether supervised or unsupervised, can be computationally too demanding for a single computing node. Instead of using a single computing node, a distributed network of computing nodes can be used to accelerate the training process.
[0163] FIG. 12 is a block diagram illustrating distributed learning. Distributed learning is a training model that uses multiple distributed computing nodes to perform supervised or unsupervised training of neural networks. The distributed computing nodes may each include one or more host processors and one or more general purpose processing nodes, such as the highly parallel general purpose graphics processing unit 700 seen in FIG. 7. As depicted, distributed learning may perform model parallelism 1202, data parallelism 1204, or a combination of model and data parallelism 1206.
[0164] In model parallelism 1202, different computational nodes in a distributed system can perform training computations for different parts of a single network. For example, each layer of a neural network can be trained by a different processing node of the distributed system. Advantages of model parallelism include the ability to scale to particularly large models. Separating the computations associated with different layers of a neural network enables the training of very large neural networks where not all layer weights fit into the memory of a single computational node. In some cases, model parallelism can be particularly useful for performing unsupervised training of large neural networks.
[0165] In data parallelism 1204, different nodes of the distributed network have a complete instance of the model, and each node receives a different portion of the data. The results from the different nodes are then combined. While different approaches to data parallelism are possible, all data parallel training approaches require techniques to combine the results and synchronize the model parameters between each node. Examples of approaches to data combination include parameter averaging and update-based data parallelism. Parameter averaging trains each node on a subset of the training data and sets the global parameters (e.g., weights, biases) to the average of the parameters from each node. Parameter averaging uses a central parameter server that holds the parameter data. Update-based data parallelism is similar to parameter averaging, except that instead of transferring parameters from the nodes to the parameter server, updates to the model are transferred. Furthermore, update-based data parallelism can be performed in a decentralized manner, and updates are compressed and transferred between nodes.
[0166] Complex model and data parallelism 1206 may be implemented, for example, in a distributed system where each compute node contains multiple GPUs, each node has a complete instance of the model, and separate GPUs within each node are used to train different parts of the model.
[0167] Distributed training incurs increased overhead compared to training on a single machine, but the parallel processors and GPGPUs described herein can each implement various techniques to reduce the overhead of distributed training, including techniques that enable high-bandwidth inter-GPU data transfers and accelerated remote data synchronization.
[0168] [Examples of machine learning applications] Machine learning can be applied to solve a variety of technical problems, including but not limited to computer vision, autonomous driving and navigation, speech recognition, and language processing. Computer vision has traditionally been one of the most active research areas for machine learning applications. Computer vision applications range from replicating human visual capabilities, such as face recognition, to creating new categories of visual capabilities. For example, a computer vision application may be configured to recognize sound waves from vibrations induced in objects visible in a video. Machine learning accelerated by parallel processors allows computer vision applications to be trained with significantly larger training datasets than was previously feasible, and allows inference systems to be deployed with lower power parallel processors.
[0169] Machine learning accelerated by parallel processors has applications in autonomous driving, including lane and road sign recognition, obstacle avoidance, navigation, and driving control. Accelerated machine learning techniques can be used to train driving models based on data sets that define appropriate responses to specific training inputs. The parallel processors described herein can enable rapid training of increasingly complex neural networks used for autonomous driving solutions, and enable the deployment of low-power inference processors on mobile platforms suitable for incorporation into autonomous vehicles.
[0170] Deep neural networks accelerated by parallel processors have enabled machine learning approaches to automatic speech recognition (ASR). ASR involves the generation of a function that computes the most likely language sequence given an input acoustic sequence. Accelerated machine learning using deep neural networks has enabled the replacement of hidden Markov models (HMMs) and Gaussian mixture models (GMMs) previously used for ASR.
[0171] Machine learning accelerated by parallel processors can also be used to accelerate natural language processing. Automated learning procedures can utilize statistical inference algorithms to generate models that are robust to erroneous or unfamiliar input. An example natural language processor application is automatic machine translation between human languages.
[0172] Parallel processing platforms used for machine learning can be divided into training platforms and deployment platforms. Training platforms are generally highly parallel and include optimizations to accelerate multi-GPU single-node training and multi-node multi-GPU training. Examples of parallel processors suitable for training include the general-purpose graphics processing unit 700 of FIG. 7 and the multi-GPU computing system 800 of FIG. 8. In contrast, deployed machine learning platforms generally include lower power parallel processors suitable for use in products such as cameras, autonomous robots, and autonomous cars.
[0173] FIG. 13 illustrates an exemplary inference system-on-a-chip (SOC) 1300 suitable for performing inference using a trained model. The SOC 1300 can incorporate processing components including a media processor 1302, a vision processor 1304, a GPGPU 1306, and a multi-core processor 1308. The SOC 1300 can further include an on-chip memory 1305 that can enable a shared on-chip data pool that is accessible by each of the processing components. The processing components can be optimized for low power operation to enable deployment in various machine learning platforms, including autonomous vehicles and autonomous robots. For example, one implementation of the SOC 1300 can be used as part of a master control system for an autonomous vehicle. When the SOC 1300 is configured for use in an autonomous vehicle, the SOC is designed and configured to comply with the relevant functional safety standards of the deployment authority.
[0174] In operation, the media processor 1302 and the vision processor 1304 can cooperate to accelerate computer vision operations. The media processor 1302 can enable low latency decoding of multiple high resolution (e.g., 4K, 8K) video streams. The decoded video streams can be written to a buffer in the on-chip memory 1305. The vision processor 1304 can then parse the decoded video and perform pre-processing operations on frames of the decoded video in preparation for processing the frames with a trained image recognition model. For example, the vision processor 1304 can accelerate the convolution operations of a CNN used to perform image recognition on high resolution video data, while back-end model computations are performed by the GPGPU 1306.
[0175] The multi-core processor 1308 may include control logic that assists in ordering and synchronizing data transfer and shared memory operations performed by the media processor 1302 and the vision processor 1304. The multi-core processor 1308 may also operate as an application processor to execute software applications that can utilize the inference computational power of the GPGPU 1306. For example, at least a portion of the navigation and driving logic may be implemented in software that runs on the multi-core processor 1308. Such software may issue computational workloads directly to the GPGPU 1306, or computational workloads may be issued to the multi-core processor 1308, which may offload at least a portion of these operations to the GPGPU 1306.
[0176] GPGPU 1306 may include a compute cluster, such as a low-power configuration of compute clusters 706A-706H in general-purpose graphics processing unit 700. The compute cluster in GPGPU 1306 may support instructions that are specifically optimized to perform inference calculations on trained neural networks. For example, GPGPU 1306 may support instructions for performing low-precision calculations, such as 8-bit and 4-bit integer vector operations.
[0177] [Further Example Graphics Processing Systems] Details of the above-described embodiments may be incorporated into the graphics processing systems and devices described below. The graphics processing systems and devices of Figures 14-26 represent alternative systems and graphics processing hardware capable of implementing any and all of the above techniques.
[0178] 14 is a block diagram of a processing system 1400 according to an embodiment. System 1400 may be used in a single processor desktop system, a multiprocessor workstation system, or a server system with multiple processors 1402 or processor cores 1407. In one embodiment, system 1400 is a processing platform embedded in a system-on-a-chip (SoC) integrated circuit for use in portable, handheld, embedded devices, such as in Internet of Things (IoT) devices with wired or wireless connections to local or wide area networks.
[0179] In one embodiment, the system 1400 can include, be coupled to, or be incorporated within a server-based gaming platform; a gaming console, including a game and media console; a portable gaming console, a handheld gaming machine, or an online gaming machine. In some embodiments, the system 1400 is part of a mobile phone, a smart phone, a tablet computing device, or a mobile Internet-connected device, such as a laptop with low internal storage capacity. The processing system 1400 can also include, be coupled to, or be incorporated within a wearable device, such as a smart watch wearable device; smart eyewear or clothing enhanced with augmented reality (AR) or virtual reality (VR) capabilities that provide visual, audio, or haptic output to complement a real-world visual, audio, or haptic experience, or that provides text, audio, video, holographic images or video, or haptic feedback; or other virtual reality (VR) devices. In some embodiments, the processing system 1400 can include, be coupled to, or be incorporated within a television or set-top box device.
[0180] In some embodiments, system 1400 may include, be coupled to, or be incorporated within an autonomous vehicle, such as a bus, a trailer truck, an automobile, a motorized or electric bicycle, an airplane, or a glider (or any combination thereof), which may use system 1400 to process the sensed environment around the vehicle.
[0181] In some embodiments, the one or more processors 1402 each include one or more processor cores 1407 that process instructions that, when executed, perform operations for system or user software. In some embodiments, at least one of the one or more processor cores 1407 is configured to process a particular instruction set 1409. In some embodiments, the instruction set 1409 may facilitate Complex Instruction Set Computing (CISC), Reduced Instruction Set Computing (RISC), or Very Long Instruction Word (VLIW) computing. One or more processor cores 1407 may process different instruction sets 1409. The different instruction sets 1409 may include instructions that facilitate emulation of other instruction sets. The processor cores 1407 may also include other processing devices, such as a Digital Signal Processor (DSP).
[0182] In some embodiments, the processor 1402 includes a cache memory 1404. Depending on the architecture, the processor 1402 may have a single internal cache, or multiple levels of internal cache. In some embodiments, the cache memory is shared among various components of the processor 1402. In some embodiments, the processor 1402 also uses an external cache (e.g., a level-3 (L3) cache or a last level cache (LLC)) (not shown). The external cache may be shared among the processor cores 1407 using known cache coherence techniques. The processor 1402 may further include a register file 1406. The register file 1406 may include different types of registers (e.g., integer registers, floating point registers, status registers, and instruction pointer registers) that store different types of data. Some registers may be general purpose registers, while other registers may be specific to the design of the processor 1402.
[0183] In some embodiments, the one or more processors 1402 are coupled to one or more interface buses 1410 to carry communication signals, such as address, data, or control signals, between the processors 1402 and other components in the system 1400. The interface bus 1410, in one embodiment, can be a processor bus, such as a variation of a Direct Media Interface (DMI) bus. However, the processor bus is not limited to a DMI bus and may include one or more Peripheral Component Interconnect (e.g., PCI, PCI Express) buses, memory buses, or other types of interface buses. In one embodiment, the processor 1402 includes an integrated memory controller 1416 and a platform controller hub 1430. The memory controller 1416 facilitates communication between memory devices and other components of the system 1400, while the platform controller hub (PCH) 1430 provides connectivity to I / O devices via a local I / O bus.
[0184] The memory device 1420 can be a dynamic random access memory (DRAM) device, a static random access memory (SRAM) device, a flash memory device, a phase change memory device, or other memory device with suitable performance to function as a process memory. In one embodiment, the memory device 1420 can operate as a system memory for the system 1400 to store data 1422 and instructions 1421 used when the one or more processors 1402 execute applications or processes. The memory controller 1416 also couples to an optional external graphics processor 1418. The external graphics processor 1418 can communicate with one or more graphics processors 1408 in the processor 1402 to perform graphics and media operations. In some embodiments, graphics, media, and / or computation operations can be assisted by an accelerator 1412, which is a co-processor that can be configured to perform a specialized set of graphics, media, or computation operations. For example, in one embodiment, the accelerator 1412 is a matrix multiplication accelerator used to optimize machine learning or computational operations. In one embodiment, the accelerator 1412 is a ray tracing accelerator that may be used to perform ray tracing operations in cooperation with the graphics processor 1408. In some embodiments, a display device 1411 may be connected to the processor 1402. The display device 1411 may be one or more of an internal display device such as found in a mobile electronic device or laptop device, or an external display device attached via a display interface (e.g., DisplayPort, etc.). In one embodiment, the display device 1411 may be a head mounted display (HMD), such as a stereoscopic display device used in virtual reality (VR) or augmented reality (AR) applications.
[0185] In some embodiments, the platform controller hub 1430 allows peripherals to connect to the memory device 1420 and the processor 1402 via a high-speed I / O bus. The I / O peripherals include, but are not limited to, an audio controller 1446, a network device 1434, a firmware interface 1428, a wireless transceiver 1426, a touch sensor 1425, and a data storage device 1424 (e.g., non-volatile memory, volatile memory, hard disk drive, flash memory, NAND, 3D NAND, 3D XPoint, etc.). The data storage device 1424 can be connected via a peripheral bus such as a Peripheral Component Interconnect (e.g., PCI, PCI Express) bus or via a storage interface (e.g., SATA). The touch sensor 1425 can include a touch screen sensor, a pressure sensor, or a fingerprint sensor. The wireless transceiver 1426 can be a Wi-Fi transceiver, a Bluetooth transceiver, or a mobile network transceiver, e.g., 3G, 4G, 5G, or a Long Term Evolution (LTE) transceiver. The firmware interface 1428 enables communication with system firmware and can be, for example, a unified extensible firmware interface (UEFI). The network controller 1434 can enable network connectivity to a wired network. In some embodiments, a high performance network controller (not shown) couples to the interface bus 1410. The audio controller 1446, in one embodiment, is a multi-channel high definition audio controller. In one embodiment, the system 1400 includes an optional legacy I / O controller 1440 for coupling legacy (e.g., Personal System 2 (PS / 2)) devices to the system.The platform controller hub 1430 may also be connected to one or more universal serial bus (USB) controllers 1442 for connecting input devices, such as a keyboard and mouse 1443 combination, a camera 1444, or other USB input devices.
[0186] It should be appreciated that the illustrated system 1400 is by way of example and not limitation, and other types of data processing systems configured otherwise may be used. For example, instances of memory controller 1416 and platform controller hub 1430 may be incorporated into a separate external graphics processor, such as external graphics processor 1418. In one embodiment, platform controller hub 1430 and / or memory controller 1416 may be external to one or more processors 1402. For example, system 1400 may include external memory controller 1416 and platform controller hub 1430, which may be configured as a memory controller hub and peripheral controller hub within a system chipset that communicates with processor 1402.
[0187] For example, circuit boards ("sleds") are available on which components such as CPUs, memory, and other components are mounted and are designed to improve thermal performance. In some examples, processing components such as processors are placed on the top surface of the sled, while nearby memory such as DIMMs are placed on the bottom surface of the sled. As a result of the enhanced airflow provided by this design, the components may operate at higher frequencies and power levels than typical systems to improve performance. Additionally, the sleds are configured to blindly mate with power and data communication cables in the rack, thereby allowing for rapid removal, upgrade, reinstallation, and / or replacement. Similarly, the individual components mounted on the sleds, such as the processors, accelerators, memory, and data storage drives, are configured to be easily upgraded by being spaced apart from one another. In an illustrative embodiment, the components further include hardware authentication features to prove their authenticity.
[0188] Data centers can utilize a single network architecture ("fabric") that supports multiple other network architectures, including Ethernet and Omni-Path. The threads can be coupled to switches via optical fiber, which provides higher bandwidth and lower latency than typical twisted pair cabling (Cat 5, Cat 5e, Cat 6, etc.). The high bandwidth, low latency interconnect and network architecture allows data centers to pool resources such as memory, accelerators (e.g., GPUs, graphics accelerators, FPGAs, ASICs, neural network and / or artificial intelligence accelerators, etc.), and data storage drives that are physically separated into components during use, and provide them to computational resources (e.g., processors) as needed, allowing the computational resources to access the pooled resources as if they were local.
[0189] A power supply or source can provide voltage and / or current to system 1400 or any component or system described herein. In one example, the power supply includes an AC-DC (alternating current to direct current) adapter to plug into a wall outlet. Such AC power can be a renewable energy (e.g., solar) source. In one example, the power supply includes a DC source, such as an external AC-DC converter. In one example, the power supply or source includes wireless charging hardware that charges by proximity to a charging field. In one example, the power source can include an internal battery, an AC source, a motion-based source, a solar source, or a fuel cell.
[0190] FIG. 15 is a block diagram of an embodiment of a processor 1500 with one or more processor cores 1502A-1502N, an integrated memory controller 1514, and an integrated graphics processor 1508. Those elements of FIG. 15 having the same reference numbers (or names) as elements of any other figure herein may operate or function in the same manner as described elsewhere herein, but are not limited to such. The processor 1500 may include additional cores, up to additional core 1502N, represented by a dashed box. Each of the processor cores 1502A-1502N includes one or more internal cache units 1504A-1504N. In some embodiments, each processor core also has access to one or more shared cache units 1506.
[0191] The internal cache units 1504A-1504N and the shared cache unit 1506 represent a cache memory hierarchy within the processor 1500. The cache memory hierarchy may include at least one level of instruction and data cache in each processor core, and one or more levels of shared mid-level cache, such as a level 2 (L2), level 3 (L3), level 4 (L4), or other level cache, with the highest level cache before external memory being classified as an LLC. In some embodiments, cache coherency logic maintains coherency between the various cache units 1506 and 1504A-1504N.
[0192] In some embodiments, processor 1500 may also include a set of one or more bus controller units 1516 and a system agent core 1510. The one or more bus controller units 1516 manage a set of peripheral buses, such as one or more PCI or PCI Express buses. System agent core 1510 provides management functionality for various processor components. In some embodiments, system agent core 1510 includes one or more integrated memory controllers 1514 that manage access to various external memory devices (not shown).
[0193] In some embodiments, one or more of the processor cores 1502A-1502N include support for simultaneous multithreading. In such embodiments, the system agent core 1510 includes components for coordinating and operating the cores 1502A-1502N during multithreaded processing. The system agent core 1510 may further include a power control unit (PCU) that includes logic and components for adjusting the power state of the processor cores 1502A-1502N and the graphics processor 1508.
[0194] In some embodiments, the processor 1500 further includes a graphics processor 1508 that performs graphics processing operations. In some embodiments, the graphics processor 1508 couples to a system agent core 1510 that includes a set of shared cache units 1506 and one or more integrated memory controllers 1514. In some embodiments, the system agent core 1510 also includes a display controller 1511 that drives the graphics processor output to one or more coupled displays. In some embodiments, the display controller 1511 may also be a separate module coupled to the graphics processor via at least one interconnect or may be incorporated within the graphics processor 1508.
[0195] In some embodiments, a ring-based interconnect unit 1512 is used to couple the internal components of the processor 1500. However, alternative interconnect units may be used, such as a point-to-point interconnect, a switched interconnect, or other technologies, including those well known in the art. In some embodiments, the graphics processor 1508 couples to the ring interconnect 1512 via an I / O link 1513.
[0196] The example I / O link 1513 represents at least one of a wide variety of I / O interconnects, including a package I / O interconnect that facilitates communication between various processor components and a high performance embedded memory module 1518, such as an eDRAM module. In some embodiments, each of the processor cores 1502A-1502N and the graphics processor 1508 can use the embedded memory module 1518 as a shared last level cache.
[0197] In some embodiments, the processor cores 1502A-1502N are homogenous cores that execute the same instruction set architecture. In other embodiments, the processor cores 1502A-1502N are heterogeneous with respect to instruction set architecture (ISA), where one or more of the processor cores 1502A-1502N execute a first instruction set, while at least one of the remaining cores executes a subset of the first instruction set or a different instruction set. In one embodiment, the processor cores 1502A-1502N are heterogeneous with respect to microarchitecture, where one or more cores having relatively higher power consumption are combined with one or more power cores having relatively lower power consumption. In one embodiment, the processor cores 1502A-1502N are heterogeneous with respect to computational capability. Additionally, the processor 1500 may be implemented on one or more chips or as a SoC integrated circuit with the depicted components in addition to other components.
[0198] 16 is a block diagram of a graphics processor 1600, which may be a separate graphics processing unit or may be a graphics processor integrated with multiple processing cores or other semiconductor devices including, but not limited to, memory devices or network interfaces. In some embodiments, the graphics processor communicates with registers on the graphics processor and with commands placed in the processor memory via a memory-mapped I / O interface. In some embodiments, the graphics processor 1600 includes a memory interface 1614 for accessing memory. The memory interface 1614 can be an interface to local memory, one or more internal caches, one or more shared external caches, and / or system memory.
[0199] In some embodiments, graphics processor 1600 also includes a display controller 1602 that drives display output data to a display device 1618. Display controller 1602 includes hardware for one or more overlay planes for display and compositing multiple layers of video or user interface elements. Display device 1618 can be an internal or external display device. In one embodiment, display device 1618 is a head-mounted display device, such as a virtual reality (VR) display device or an augmented reality (AR) display device. In some embodiments, graphics processor 1600 includes a video codec engine 1606 that encodes, decodes, or transcodes media to, from, or between one or more media encoding formats, including but not limited to Moving Picture Experts Group (MPEG) formats such as MPEG-2, Advanced Video Coding (AVC) formats such as H.265 / MPEG-4 AVC, H.265 / HEVC, Alliance for Open Media (AOMedia) VP8, VP9, and Society of Motion Picture & Television Engineers (SMPTE) 421M / VC-1, and Joint Photographic Experts Group (JPEG) formats such as JPEG and Motion JPEG (MJPEG) formats.
[0200] In some embodiments, graphics processor 1600 includes a block image transfer (BLIT) engine 1604 that performs two-dimensional (2D) rasterizer operations, including, for example, bit-boundary block transfers. However, in one embodiment, 2D graphics operations are performed using one or more components of a graphics processing engine (GPE) 1610. In some embodiments, GPE 1610 is a computation engine that performs graphics operations, including three-dimensional (3D) graphics operations and media operations.
[0201] In some embodiments, GPE 1610 includes a 3D pipeline 1612 that performs 3D operations such as rendering three-dimensional images and scenes using processing functions that operate on 3D primitive shapes (e.g., rectangles, triangles, etc.). The 3D pipeline 1612 includes programmable fixed functions that perform various tasks within elements and / or spawn execution threads for the 3D / media subsystem 1615. While the 3D pipeline 1612 may be used to perform media operations, embodiments of GPE 1610 also include a media pipeline 1616 that is used specifically to perform media operations such as video post-processing and image enhancement.
[0202] In some embodiments, the media pipeline 1616 includes fixed function or programmable logic units that perform one or more specialized media operations, such as video decode acceleration, video deinterlacing, and video encode acceleration, in place of or on behalf of the video codec engine 1606. In some embodiments, the media pipeline 1616 further includes a thread generation unit that generates threads to execute in the 3D / media subsystem 1615. The generated threads perform calculations for the media operations in one or more graphics execution units included in the 3D / media subsystem 1615.
[0203] In some embodiments, the 3D / Media subsystem 1615 includes logic for executing threads created by the 3D pipeline 1612 and the media pipeline 1616. In one embodiment, these pipelines send thread execution requests to the 3D / Media subsystem 1615. The 3D / Media subsystem 1615 includes thread dispatch logic that arbitrates and dispatches the various requests to available thread execution resources. The execution resources include an array of graphics execution units that process the 3D and media threads. In some embodiments, the 3D / Media subsystem 1615 includes one or more internal caches for thread instructions and data. In some embodiments, the subsystem also includes shared memory, including registers and addressable memory, for sharing data between threads and storing output data.
[0204] [Graphics processing engine] FIG. 17 is a block diagram of a graphics processing engine 1710 of a graphics processor according to some embodiments. In one embodiment, the graphics processing engine (GPE) 1710 is a variation of the GPE 1610 shown in FIG. 16. Elements of FIG. 17 having the same reference numbers (or names) as elements of other figures herein can operate or function in the same manner as described elsewhere herein, but are not limited to such. For example, the 3D pipeline 1612 and the media pipeline 1616 of FIG. 16 are depicted. The media pipeline 1616 is optional in some embodiments of the GPE 1710 and may not be explicitly included within the GPE 1710. For example, in at least one embodiment, a separate media and / or image processor is coupled to the GPE 1710.
[0205] In some embodiments, the GPE 1710 is coupled to or includes a command streamer 1703 that provides a command stream to the 3D pipeline 1612 and / or the media pipeline 1616. In some embodiments, the command streamer 1703 is coupled to a memory, which may be a system memory or one or more of an internal cache memory and a shared cache memory. In some embodiments, the command streamer 1703 receives commands from memory and sends the commands to the 3D pipeline 1612 and / or the media pipeline 1616. The commands are instructions fetched from a ring buffer that stores commands for the 3D pipeline 1612 and the media pipeline 1616. In one embodiment, the ring buffer may further include a batch command buffer that stores a batch of multiple commands. The commands for the 3D pipeline 1612 may also include references to data stored in memory, such as, but not limited to, vertex and geometry data for the 3D pipeline 1612 and / or image data and memory objects for the media pipeline 1616. The 3D pipeline 1612 and the media pipeline 1616 process commands and data by executing operations through logic within the respective pipelines or by dispatching one or more execution threads to the graphics core array 1714. In one embodiment, the graphics core array 1714 includes one or more blocks of graphics cores (e.g., graphics core 1715A, graphics core 1715B), each block including one or more graphics cores. Each graphics core includes a set of graphics execution resources including general-purpose and graphics-specific execution logic that performs graphics and computation operations, and fixed-function texture processing and / or machine learning and artificial intelligence acceleration logic.
[0206] In various embodiments, the 3D pipeline 1612 may include fixed function and programmable logic that processes one or more shader programs, such as vertex shaders, geometry shaders, pixel shaders, fragment shaders, compute shaders, or other shader programs, by processing instructions and dispatching execution threads to the graphics core array 1714. The graphics core array 1714 provides a unified block of execution resources used in processing these shader programs. General purpose execution logic (e.g., execution units) within the graphics cores 1715A-1715B of the graphics core array 1714 includes support for a variety of 3D API shader languages and can execute multiple concurrent threads of execution associated with multiple shaders.
[0207] In some embodiments, graphics core array 1714 includes execution logic to perform media functions such as video and / or image processing. In one embodiment, the execution unit includes general purpose logic that is programmable to perform parallel general purpose computing operations in addition to graphics processing operations. The general purpose logic can perform processing operations in parallel or in conjunction with general purpose logic in processor core 1407 of FIG. 14 or cores 1502A-1502N of FIG. 15.
[0208] Output data generated by threads executing in graphics core array 1714 can output data to memory in unified return buffer (URB) 1718. URB 1718 can store data for multiple threads. In some embodiments, URB 1718 can be used to route data between different threads executing in graphics core array 1714. In some embodiments, URB 1718 can also be used for synchronization between threads in graphics core array and fixed function logic in shared function logic 1720.
[0209] In some embodiments, graphics core array 1714 is scalable such that the array includes a variable number of graphics cores, each having a variable number of execution units based on the target power and performance levels of GPE 1710. In one embodiment, the execution resources are dynamically scalable such that execution resources can be enabled or disabled as needed.
[0210] The graphics core array 1714 couples to shared function logic 1720, which includes multiple resources shared among the graphics cores in the graphics core array. The shared functions in the shared function logic 1720 are hardware logic units that provide specialized auxiliary functions to the graphics core array 1714. In various embodiments, the shared function logic 1720 includes, but is not limited to, sampler 1721, math 1722, and inter-thread communication (ITC) 1723 logic. Additionally, some embodiments implement one or more caches 1725 in the shared function logic 1720.
[0211] Shared functionality is implemented at least when the demand for a given specialized functionality is insufficient to justify inclusion within the graphics core array 1714. Instead, a single instantiation of this specialized functionality is implemented as a stand-alone entity in the shared functionality logic 1720 and shared among the execution resources within the graphics core array 1714. The exact set of functionality shared among and included within the graphics core array 1714 varies from embodiment to embodiment. In some embodiments, certain shared functionality within the shared functionality logic 1720 that is used extensively by the graphics core array 1714 may be included within the shared functionality logic 1716 within the graphics core array 1714. In various embodiments, the shared functionality logic 1716 within the graphics core array 1714 may include some or all of the logic within the shared functionality logic 1720. In one embodiment, all logic elements within the shared functionality logic 1720 may be duplicated within the shared functionality logic 1716 of the graphics core array 1714. In one embodiment, the shared functionality logic 1720 is omitted in favor of the shared functionality logic 1716 in the graphics core array 1714 .
[0212] FIG. 18 is a block diagram of hardware logic of a graphics processor core 1800 according to some embodiments described herein. Elements of FIG. 18 having the same reference numbers (or names) as elements of any other figures herein may operate or function in the same manner as described elsewhere herein, but are not limited to such. The depicted graphics processor core 1800, in some embodiments, is included within the graphics core array 1714 of FIG. 17. The graphics processor core 1800, sometimes referred to as a core slice, can be one or more graphics cores within a modular graphics processor. The graphics processor core 1800 is an example of one graphics core slice, and the graphics processors described herein may include multiple graphics core slices based on target power and performance envelopes. Each graphics processor core 1800 may include a fixed function block 1830 coupled with multiple sub-cores 1801A-1801F, also referred to as sub-slices, that include modular blocks of general purpose and fixed function logic.
[0213] In some embodiments, fixed function block 1830 includes a geometry / fixed function pipeline 1836 that may be shared by all sub-cores in graphics processor core 1800, for example in lower performance and / or lower power graphics processor implementations. In various embodiments, geometry / fixed function pipeline 1836 includes a 3D fixed function pipeline (e.g., 3D pipeline 1612 seen in FIGS. 16 and 17), a video front end unit, a thread spawner and thread dispatcher, and a unified return buffer manager that manages a unified return buffer, such as unified return buffer 1718 of FIG. 17.
[0214] In one embodiment, the fixed function block 1830 also includes a graphics SoC interface 1837, a graphics microcontroller 1838, and a media pipeline 1839. The graphics SoC interface 1837 provides an interface between the graphics processor core 1800 and other processors in the SoC integrated circuit. The graphics microcontroller 1838 is a programmable sub-processor that is configurable to manage various functions of the graphics processor core 1800, including thread dispatch, scheduling, and preemption. The media pipeline 1839 (e.g., the media pipeline 1616 of Figures 16 and 17) includes logic that aids in decoding, encoding, pre-processing, and / or post-processing of multimedia data, including image and video data. The media pipeline 1839 implements media operations by requesting computation or sampling logic in the sub-cores 1801A-1801F.
[0215] In one embodiment, the SoC interface 1837 enables the graphics processor core 1800 to communicate with other components in the SoC, including a general-purpose application processor core (e.g., a CPU) and / or memory hierarchy elements such as a shared last-level cache memory, system RAM, and / or embedded on-chip or on-package DRAM. The SoC interface 1837 may also enable communication with fixed function devices in the SoC, such as a camera imaging pipeline, and enable and / or implement the use of global memory atomics that may be shared between the graphics processor core 1800 and CPUs in the SoC. The SoC interface 1837 may also implement power management controls for the graphics processor core 1800 and enable an interface between the clock domain of the graphics processor core 1800 and other clock domains in the SoC. In one embodiment, the SoC interface 1837 enables receipt of command buffers from a command streamer and a global thread dispatcher configured to provide commands and instructions to each of one or more graphics cores in the graphics processor. Commands and instructions may be dispatched to the media pipeline 1839 if media operations are to be performed, or to the geometry and fixed function pipelines (e.g., geometry and fixed function pipelines 1836, geometry and fixed function pipelines 1814) if graphics processing operations are to be performed.
[0216] The graphics microcontroller 1838 may be configured to perform various scheduling and management tasks for the graphics processor core 1800. In one embodiment, the graphics microcontroller 1838 may perform graphics and / or compute workload scaling to various graphics parallel engines in the execution unit (EU) arrays 1802A-1802F, 1804A-1804F in the sub-cores 1801A-1801F. In this scheduling model, host software executing on a CPU core of an SoC including the graphics processor core 1800 may issue a workload to one of a number of graphics processor doorbells which invokes a scaling operation on the appropriate graphics engine. Scheduling operations include determining which workload to run next, issuing the workload to the command streamer, preempting existing workloads running on the engines, managing the progress of the workloads, and notifying the host software when the workloads are completed. In one embodiment, graphics microcontroller 1838 can also facilitate low power or idle states of graphics processor core 1800, providing graphics processor core 1800 with the ability to save and restore registers within graphics processor core 1800 across low power state transitions independent of the operating system and / or graphics driver software on the system.
[0217] Graphics processor core 1800 may have more or less modular sub-cores than the depicted sub-cores 1801A-1801F, up to N. For each set of N sub-cores, graphics processor core 1800 may also include shared function logic 1810, shared and / or cache memory 1812, geometry / fixed function pipeline 1814, and additional fixed function logic 1816 to accelerate various graphics and computational processing operations. Shared function logic 1810 may include logic units associated with shared function logic 1720 of FIG. 17 (e.g., sampler, math, and / or inter-thread communication logic) that may be shared by each of the N sub-cores in graphics processor core 1800. Shared and / or cache memory 1812 may be a last level cache for the set of N sub-cores 1801A-1801F in graphics processor core 1800 and may also act as a shared memory that is accessible by multiple sub-cores. The geometry / fixed function pipeline 1814 may be included in place of the geometry / fixed function pipeline 1836 in the fixed function block 1830 and may include the same or similar logic units.
[0218] In one embodiment, graphics processor core 1800 includes additional fixed function logic 1816, which may include various fixed function acceleration logic used by graphics processor core 1800. In one embodiment, additional fixed function logic 1816 includes an additional geometry pipeline used in position only shading. In position only shading, there are two geometry pipelines: a full geometry pipeline in geometry / fixed function pipelines 1814, 1836, and a cull pipeline, which is an additional geometry pipeline that may be included in additional fixed function logic 1816. In one embodiment, the cull pipeline is a scaled-down version of the full geometry pipeline. The full pipeline and the cull pipeline may run different instances of the same application, with each instance having a separate context. Position only shading may hide long cull runs for discarded triangles, allowing shading to be completed sooner in some instances. For example, in one embodiment, the cull pipeline logic in additional fixed function logic 1816 can execute position shaders in parallel with the main application and generally produce critical results faster than the full pipeline because the cull pipeline fetches and shades only the position attributes of vertices without performing rasterization and rendering of pixels to the frame buffer. The cull pipeline can use the generated deterministic results to calculate visibility information for all triangles, regardless of whether those triangles are culled or not.The full pipeline (which in this instance may be called the replay pipeline) can consume the visibility information to skip over the culled triangles in order to shade only the visible triangles that are eventually passed to the rasterization phase.
[0219] In some embodiments, the additional fixed function logic 1816 may also include machine learning acceleration logic, such as fixed function matrix multiplication, for implementations that include optimizations for machine learning training or inference.
[0220] Within each graphics sub-core 1801A-1801F is a set of execution resources that can be used to perform graphics, media, and compute operations in response to requests by a graphics pipeline, a media pipeline, or a shader program. The graphics sub-cores 1801A-1801F include a number of EU arrays 1802A-1802F, 1804A-1804F, thread dispatch and inter-thread communication (TD / IC) logic 1803A-1803F, 3D (e.g., texture) samplers 1805A-1805F, media samplers 1806A-1806F, shader processors 1807A-1807F, and shared local memories (SLMs) 1808A-1808F. The EU arrays 1802A-1802F, 1804A-1804F each include a number of execution units that are general purpose graphics processing units capable of performing floating point and integer / fixed point logic operations in service of graphics, media, or compute operations including graphics, media, and compute shader programs. The TD / IC logic 1803A-1803F performs local thread dispatch and thread control operations for the execution units within the sub-cores and facilitates communication between threads executing on the execution units of the sub-cores. The 3D samplers 1805A-1805F can load textures or other 3D graphics related data into memory. The 3D samplers can load texture data differently based on the texture format and configured sample state associated with a given texture. The media samplers 1806A-1806F can perform similar load operations based on the type and format associated with the media data. In one embodiment, each graphics sub-core 1801A-1801F may alternatively include an integrated 3D and media sampler. Threads executing on execution units within each of the sub-cores 1801A-1801F may utilize shared local memory 1808A-1808F within each sub-core to allow threads executing within a thread group to execute using a common pool of on-chip memory.
[0221] [Execution unit] 19A-19B depict thread execution logic 1900 including an array of processing elements for use in a graphics processor core, according to embodiments described herein. Elements of FIG. 19A-19B having the same reference numbers (or names) as elements of any other figure of the present application may operate or function in the same manner as described elsewhere in the present application, but are not limited to such. FIG. 19A depicts an overview of thread execution logic 1900, which may include variations of the hardware logic represented by each of sub-cores 1801A-1801F of FIG. 18. FIG. 19B depicts an example of internal details of an execution unit.
[0222] As depicted in Figure 19A, in some embodiments, thread execution logic 1900 includes a shader processor 1902, a thread dispatcher 1904, an instruction cache 1906, a scalable execution unit array including multiple execution units 1908A-1908N, a sampler 1910, a data cache 1912, and a data port 1914. In one embodiment, the scalable execution unit array is dynamically scalable by enabling or disabling one or more execution units (e.g., any of execution units 1908A, 1908B, 1908C, 1908D, through 1908N-1 and 1908N) based on the computational requirements of a workload. In one embodiment, the included components are interconnected via an interconnect fabric that provides links to each of the components. In some embodiments, the thread execution logic 1900 includes one or more connections to memory, such as system memory or cache memory, through one or more of an instruction cache 1906, a data port 1914, a sampler 1910, and execution units 1908A-1908N. In some embodiments, each execution unit (e.g., 1908A) is a standalone programmable general-purpose computational unit capable of executing multiple simultaneous hardware threads, processing multiple data elements in parallel for each thread. In various embodiments, the array of execution units 1908A-1908N is scalable to include any number of individual execution units.
[0223] In some embodiments, the execution units 1908A-1908N are primarily used to execute shader programs. The shader processor 1902 processes various shader programs and may dispatch execution threads associated with the shader programs via a thread dispatcher 1904. In one embodiment, the thread dispatcher includes logic to arbitrate thread initiation requests from the graphics and media pipelines and instantiate the requested threads on one or more of the execution units 1908A-1908N. For example, the geometry pipeline may dispatch vertex, tessellation, or geometry shaders to thread execution logic for processing. In some embodiments, the thread dispatcher 1904 may also handle run-time thread creation requests from executing shader programs.
[0224] In some embodiments, the execution units 1908A-1908N support an instruction set that includes native support for many standard 3D graphics shader instructions so that shader programs from graphics libraries (e.g., Direct 3D and OpenGL) are executed with minimal translation. The execution units support vertex and geometry processing (e.g., vertex programs, geometry programs, vertex shaders), pixel processing (e.g., pixel shaders, fragment shaders), and general-purpose processing (e.g., compute and media shaders). Each of the execution units 1908A-1908N is capable of multi-issue single instruction multiple data (SIMD) execution, and multi-threaded operation enables an efficient execution environment in the face of higher latency memory accesses. Each hardware thread within each execution unit has a dedicated high-bandwidth register file and associated independent thread state. Execution multiplexes every clock into a pipeline capable of integer, single and double precision floating point operations, SIMD branch functions, logical operations, transcendental operations, and other miscellaneous operations. While waiting for data from memory or one of the shared functions, dependency logic within the execution units 1908A-1908N causes the waiting thread to sleep until the requested data is returned. While the waiting thread is sleeping, hardware resources may be dedicated to processing other threads. For example, during a delay associated with a vertex shader operation, the execution unit may execute operations for other types of shader programs, including pixel shaders, fragment shaders, or other vertex shaders. Various embodiments are applicable to use execution through the use of single instruction multiple threads (SIMT) as an alternative to or in addition to the use of SIMD. Reference to a SIMD core or operation may also apply to SIMT or to SIMD combined with SIMT.
[0225] Each execution unit in execution units 1908A-1908N operates on an array of data elements. The number of data elements is the "execution size" for an instruction, or the number of channels. An execution channel is a logic unit of execution for data element access, masking, and flow control within an instruction. The number of channels may be independent of the number of physical arithmetic logic units (ALUs) or floating point units (FPUs) for a particular graphics processor. In some embodiments, execution units 1908A-1908N support integer and floating point data types.
[0226] The execution unit instruction set includes SIMD instructions. Various data elements can be stored in registers as packed data types, and the execution unit will process the various elements based on the data size of the elements. For example, when operating on a 256-bit wide vector, 256 bits of the vector are stored in registers, and the execution unit operates on the vector as four separate 64-bit packed data elements (Quad-Word (QW) size data elements), eight separate 32-bit packed data elements (Double-Word (DW) size data elements), sixteen separate 16-bit packed data elements (Word (W) size data elements), or thirty-two separate 8-bit data elements (Byte (B) size data elements). However, different vector widths and register sizes are possible.
[0227] In one embodiment, one or more execution units may be grouped into a fused execution unit 1909A-1909N with thread control logic (1907A-1907N) common to the fused EUs. Multiple EUs may be fused into an EU group. Each EU in a fused EU group may be configured to execute a separate SIMD hardware thread. The number of EUs in a fused EU group may vary depending on the embodiment. Furthermore, various SIMD widths may be executed per EU, including but not limited to SIMD8, SIMD16, and SIMD32. Each of the fused graphics execution units 1909A-1909N includes at least two execution units. For example, the fused execution unit 1909A includes a first EU 1908A, a second EU 1908B, and a third control logic 1907A common to the first EU 1908A and the second EU 1908B. Thread control logic 1907A controls the threads executing in the fused graphics execution unit 1909A, allowing each EU in the fused execution units 1909A-1909N to execute using a common instruction pointer register.
[0228] One or more internal instruction caches (e.g., 1906) are included in the thread execution logic 1900 to cache thread instructions for the execution units. In some embodiments, one or more data caches (e.g., 1912) are included to cache thread data during thread execution. In some embodiments, a sampler 1910 is included to provide texture sampling for 3D operations and a media sampler for media operations. In some embodiments, the sampler 1910 includes specialized texture or media sampling functions that process texture or media data during the sampling process before providing the sampled data to the execution units.
[0229] During execution, the graphics and media pipeline sends thread start requests to the thread execution logic 1900 via thread creation and dispatch logic. Once a group of geometric objects has been processed and rasterized into pixel data, pixel processor logic (e.g., pixel shader logic, fragment shader logic, etc.) within the shader processor 1902 is invoked to further compute output information and write the results to an output surface (e.g., color buffer, depth buffer, stencil buffer, etc.). In some embodiments, the pixel shader or fragment shader computes values for various vertex attributes that are to be interpolated across the rasterized objects. In some embodiments, the pixel processor logic within the shader processor 1902 then executes pixel or fragment shader programs provided by an application programming interface (API). To execute the shader programs, the shader processor 1902 dispatches threads to the execution units (e.g., 1908A) via the thread dispatcher 1904. In some embodiments, the shader processor 1902 uses texture sampling logic in the sampler 1910 to access texture data in a texture map stored in memory. Arithmetic operations on the texture data and the input geometry data calculate pixel color data for each geometric fragment or discard one or more pixels from further processing.
[0230] In some embodiments, the data port 1914 provides a memory access mechanism for the thread execution logic 1900 to output processed data to memory for further processing in the graphics processor output pipeline. In some embodiments, the data port 1914 includes or is coupled to one or more cache memories (e.g., data cache 1912) that cache data for memory access by the data port.
[0231] As depicted in FIG. 19B , the graphics execution unit 1908 may include an instruction fetch unit 1937, a general purpose register file array (GFR) 1924, an architectural register file array (ARF) 1926, a thread arbiter 1922, a send unit 1930, a branch unit 1932, a SIMD floating point unit (FPU) 1934, and, in one embodiment, a set of dedicated integer SIMD ALUs 1935. The GRF 1924 and ARF 1926 include a set of general purpose register files and architectural register files associated with each concurrent hardware thread that may be active in the graphics execution unit 1908. In one embodiment, per-thread architectural state is maintained in the ARF 1926, while data used during thread execution is stored in the GRF 1924. The execution state of each thread, including the instruction pointer for each thread, may be maintained in thread-specific registers within the ARF 1926.
[0232] In one embodiment, the graphics execution unit 1908 has an architecture that is a combination of simultaneous multi-threading (SMT) and fine-grained interleaved multi-threading (IMT). The architecture has a modular organization that can be tuned at design time based on the target number of simultaneous threads and the number of registers per execution unit, and execution unit resources are divided across the logic used to execute multiple simultaneous threads.
[0233] In one embodiment, the graphics execution unit 1908 can co-issue multiple instructions, each of which may be a different instruction. The thread arbiter 1922 of the graphics execution unit 1908 can dispatch instructions to one of the transmit unit 1930, the branch unit 1932, or the SIMD FPU 1934 for execution. Each execution thread can access 128 general purpose registers in the GRF 1924, each of which can store 32 bytes accessible as an 8-element vector of 32-bit data elements. In one embodiment, each execution unit thread has access to 4K bytes in the GRF 1924, although embodiments are not so limited and more or less register resources may be provided in other embodiments. In one embodiment, a maximum of seven threads can execute simultaneously, although the number of threads per execution unit can also vary depending on the embodiment. In an embodiment in which seven threads can access 4K bytes, the GRF 1924 can store a total of 28K bytes. Flexible addressing modes can allow registers to be addressed together to effectively configure wider registers, or to represent strided rectangular block data structures.
[0234] In one embodiment, memory operations, sampler operations, and other longer latency system communications are dispatched by "send" instructions executed by a message passing send unit 1930. In one embodiment, branch instructions are dispatched to a dedicated branch unit 1932 to facilitate SIMD divergence and resulting convergence.
[0235] In one embodiment, the graphics execution unit 1908 includes one or more SIMD floating point units (FPUs) 1934 to perform floating point operations. In one embodiment, the FPUs 1934 also support integer calculations. In one embodiment, the FPUs 1934 can SIMD execute up to M 32-bit floating point (or integer) operations or SIMD execute up to 2M 16-bit integer or 16-bit floating point operations. In one embodiment, at least one of the FPUs provides extended math capabilities to support high throughput transcendental math functions and double precision 64-bit floating point. In some embodiments, a set of 8-bit integer SIMD ALUs 1935 are also present and may be specifically optimized to perform operations related to machine learning calculations.
[0236] In one embodiment, an array of multiple instances of graphics execution unit 1908 may be instantiated in a graphics sub-core grouping (e.g., a sub-slice). For scalability, product inventors can select the exact number of execution units per sub-core grouping. In one embodiment, execution unit 1908 can execute instructions across multiple execution channels. In a further embodiment, each thread executing in graphics execution unit 1908 executes on a different channel.
[0237] 20 is a block diagram illustrating a graphics processor instruction format 2000 according to some embodiments. In one or more embodiments, the graphics processor execution unit supports an instruction set having instructions in multiple formats. The solid lined boxes represent components that are typically included in the execution unit instructions, while the dashed lines include components that are optional or that are only included in a subset of the instructions. In some embodiments, the instruction format 2000 described and illustrated are macro-instructions, in that they are instructions that are supplied to the execution units as the instructions are processed, as opposed to micro-operations that result from instruction decoding.
[0238] In some embodiments, the graphics processor execution units natively support instructions in a 128-bit instruction format 2010. A 64-bit compact instruction format 2030 is available for some instructions based on the selected instruction, instruction options, and number of operands. The original 128-bit instruction format 2010 provides access to all instruction options, while in the 64-bit format 2030 some options and operations are restricted. The original instructions available in the 64-bit format 2030 vary depending on the embodiment. In some embodiments, the instructions are partially compressed using a set of index values in index field 2013. The execution unit hardware looks up a set of compression tables based on the index values and uses the compression table output to reconstruct the original instruction in the 128-bit instruction format 2010. Instructions of other sizes and formats are available.
[0239] For each format, the instruction opcode 2012 defines the operation that the execution unit should perform. The execution unit executes each instruction in parallel across multiple data elements of each operand. For example, in response to an add instruction, the execution unit performs a simultaneous add operation across each color channel representing a texture or picture element. By default, the execution unit executes each instruction across all data channels of an operand. In some embodiments, the instruction control field 2014 allows control over certain execution options such as channel selection (e.g., predication) and data channel order (swizzle). For instructions in the 128-bit instruction format 2010, the execution size field 2016 limits the number of data channels that are executed in parallel. In some embodiments, the execution size field 2016 is not available in the 64-bit compact instruction format 2030.
[0240] Some execution unit instructions have up to three operands, including two source operands src0 2020 and src1 2022, and one destination 2018. In some embodiments, the execution units support dual destination instructions, where one of the destinations is implicit. Data manipulation instructions may have a third source operand (e.g., SRC2 2024), and the instruction opcode 2012 determines the number of source operands. The last source operand of an instruction may be an immediate (e.g., hard-coded) value passed with the instruction.
[0241] In some embodiments, the 128-bit instruction format 2010 includes an access / address mode field 2026 that specifies, for example, whether direct or indirect register addressing mode is used. When direct register addressing mode is used, the register addresses of one or more operands are provided directly by bits in the instruction.
[0242] In some embodiments, the 128-bit instruction format 2010 includes an access / address mode field 2026 that specifies the address mode and / or access mode of the instruction. In one embodiment, the access mode is used to define the data access alignment for the instruction. Some embodiments support access modes including a 16-byte aligned access mode and a 1-byte aligned access mode, where the byte alignment of the access mode determines the access alignment of the instruction operands. For example, when in a first mode, the instruction may use byte-aligned addressing for source and destination operands, and when in a second mode, the instruction may use 16-byte aligned addressing for all source and destination operands.
[0243] In one embodiment, the address mode portion of the access / address mode field 2026 determines whether the instruction should use direct or indirect addressing. If the direct register addressing mode is used, bits in the instruction directly provide the register addresses of one or more operands. If the indirect register addressing mode is used, the register addresses of one or more operands may be calculated based on an address intermediate field in the instruction and an address register value.
[0244] In some embodiments, instructions are grouped based on a bit field in the opcode 2012 to simplify the opcode code 2040. For an 8-bit opcode, bits 4, 5, and 6 allow the execution unit to determine the type of opcode. The exact opcode groupings shown are only examples. In some embodiments, the move and logical opcode group 2042 includes data movement and logical instructions (e.g., move (mov), compare (cmp)). In some embodiments, the move and logical group 2042 shares the five most significant bits (MSBs), with move (mov) instructions taking the form 0000xxxxb and logical instructions taking the form 0001xxxxb. The flow control instructions 2044 (e.g., call, jump (jmp)) include instructions taking the form 0010xxxxb (e.g., 0x20). The miscellaneous instruction group 2046 includes a mix of instructions including synchronization instructions (e.g., wait, send) that take the form of 0011xxxxb (e.g., 0x30). The parallel arithmetic instruction group 2048 includes component-wise arithmetic instructions (e.g., add, multiply (mul)) that take the form of 0100xxxxb (e.g., 0x40). The parallel arithmetic group 2048 performs arithmetic operations in parallel across data channels. The vector arithmetic group 2050 includes arithmetic instructions (e.g., dp4) that take the form of 0101xxxxb (e.g., 0x50). The vector arithmetic group performs arithmetic operations such as dot product calculations on vector operands.
[0245] [Graphics Pipeline] Figure 21 is a block diagram of another embodiment of a graphics processor 2100. Elements of Figure 21 having the same reference numbers (or names) as elements of any other figure in this application can operate or function in the same manner as described elsewhere in this application, but are not limited to such.
[0246] In some embodiments, the graphics processor 2100 includes a geometry pipeline 2120, a media pipeline 2130, a display engine 2140, thread execution logic 2150, and a render output pipeline 2170. In some embodiments, the graphics processor 2100 is a graphics processor in a multi-core processing system that includes one or more general-purpose processing cores. The graphics processor is controlled by commands issued to the graphics processor 2100 via a ring interconnect 2102 or by register writes to one or more control registers (not shown). In some embodiments, the ring interconnect 2102 couples the graphics processor 2100 to other processing components, such as other graphics processors or general-purpose processors. Commands from the ring interconnect 2102 are interpreted by a command streamer 2103, which provides instructions to individual components of the geometry pipeline 2120 or the media pipeline 2130.
[0247] In some embodiments, command streamer 2103 directs the operation of vertex fetcher 2105, which reads vertex data from memory and executes vertex processing commands provided by command streamer 2103. In some embodiments, vertex fetcher 2105 provides vertex data to vertex shader 2107, which performs coordinate space transformations and lighting operations on each vertex. In some embodiments, vertex fetcher 2105 and vertex shader 2107 execute vertex processing instructions by dispatching execution threads to execution units 2152A-2152B via thread dispatcher 2131.
[0248] In some embodiments, the execution units 2152A-B are an array of vector processors with an instruction set for performing graphics and media operations. In some embodiments, the execution units 2152A-B have an associated L1 cache 2151 that is specific to each array or shared between the arrays. The cache can be configured as a data cache, an instruction cache, or a single cache that is partitioned to contain data and instructions in different partitions.
[0249] In some embodiments, the geometry pipeline 2120 includes a tessellation component that performs hardware accelerated tessellation of 3D objects. In some embodiments, a programmable hull shader 2111 configures the tessellation operations. A programmable domain shader 2117 provides back-end evaluation of the tessellation output. A tessellator 2113 operates at the direction of the hull shader 2111 and includes dedicated logic to generate a set of detailed geometric objects based on a coarse geometric model provided as input to the geometry pipeline 2120. In some embodiments, the tessellation components (e.g., the hull shader 2111, the tessellator 2113, and the domain shader 2117) may be bypassed if tessellation is not used.
[0250] In some embodiments, a complete geometric object may be processed by the geometry shader 2119 via one or more threads dispatched to the execution units 2152A-2152B, or may proceed directly to the clipper 2129. In some embodiments, the geometry shader 2119 operates on entire geometric objects, rather than vertices or patches of vertices as found in previous stages of the graphics pipeline. The geometry shader 2119 receives input from the vertex shader 2107 if tessellation is disabled. In some embodiments, the geometry shader 2119 is programmable by a geometry shader program to perform geometry tessellation if the tessellation units are disabled.
[0251] Prior to rasterization, a clipper 2129 processes the vertex data. The clipper 2129 may be a fixed function clipper or a programmable clipper with clipping and geometry shader functionality. In some embodiments, a rasterizer and depth test component 2173 in the render output pipeline 2170 dispatches pixel shaders to convert geometric objects into per-pixel representations. In some embodiments, pixel shader logic is included in the thread execution logic 2150. In some embodiments, an application can bypass the rasterizer and depth test component 2173 and access un-rasterized vertex data via the stream output unit 2123.
[0252] The graphics processor 2100 includes an interconnect bus, interconnect fabric, or other interconnect mechanism that allows data and message passing between the major components of the processor. In some embodiments, the execution units 2152A-2152B and associated logic units (e.g., L1 cache 2151, sampler 2154, texture cache 2158, etc.) interconnect via data ports 2156 to perform memory accesses and communicate with the processor's render output pipeline components. In some embodiments, the sampler 2154, L1 cache 2151, texture cache 2158, and the execution units 2152A-2152B each have a separate memory access path. In one embodiment, the texture cache 2158 may also be configured as a sampler cache.
[0253] In some embodiments, the render output pipeline 2170 includes a rasterizer and depth test component 2173 that converts vertex-based objects into an associated pixel-based representation. In some embodiments, the rasterizer logic includes a windower / masker unit that performs fixed function triangle and line rasterization. An associated render cache 2178 and depth cache 2179 are also available in some embodiments. A pixel manipulation component 2177 performs pixel-based manipulations on the data, although in some instances pixel operations associated with 2D operations (e.g., bit-block image transfer with blending) are performed by the 2D engine 2141 or are replaced at display time by the display controller 2143 using an overlay display surface. In some embodiments, a shared L3 cache 2175 is available to all graphics components, allowing sharing of data without the use of main system memory.
[0254] In some embodiments, the graphics processor media pipeline 2130 includes a media engine 2137 and a video front end 2134. In some embodiments, the video front end 2134 receives pipeline commands from the command streamer 2103. In some embodiments, the media pipeline 2130 includes a separate command streamer. In some embodiments, the video front end 2134 processes the media commands before sending the commands to the media engine 2137. In some embodiments, the media engine 2137 includes a thread spawning function to spawn threads for dispatch by the thread dispatcher 2131 to the thread execution logic 2150.
[0255] In some embodiments, the graphics processor 2100 includes a display engine 2140. In some embodiments, the display engine 2140 is external to the processor 2100 and couples to the graphics processor via a ring interconnect 2102 or other interconnect bus or fabric. In some embodiments, the display engine 2140 includes a 2D engine 2141 and a display controller 2143. In some embodiments, the display engine 2140 includes dedicated logic capable of operating independently from the 3D pipeline. In some embodiments, the display controller 2143 couples to a display device (not shown), which may be a built-in display device, such as found on a laptop computer, or an external display device attached via a display device connector.
[0256] In some embodiments, the geometry pipeline 2120 and the media pipeline 2130 are configurable to perform operations based on multiple graphics and media programming interfaces and are not specific to any one application programming interface (API). In some embodiments, driver software for the graphics processor translates API calls that are specific to a particular graphics or media library into commands that can be processed by the graphics processor. In some embodiments, support is provided for the Open Graphics Library (OpenGL), Open Computing Language (OpenCL), and / or Vulkan graphics and compute APIs, all from the Khronos Group. In some embodiments, support may also be provided for the Direct3D library from Microsoft Corporation. In some embodiments, a combination of these libraries may be supported. Support may also be provided for the Open Source Computer Vision Library (OpenCV). Future APIs with compatible 3D pipelines will also be supported if a mapping can be made from the future API's pipeline to the graphics processor's pipeline.
[0257] [Graphics Pipeline Programming] Figure 22A is a block diagram illustrating a graphics processor command format 2200 according to some embodiments. Figure 22B is a block diagram illustrating a graphics processor command sequence 2210 according to an embodiment. The solid lined boxes in Figure 22A represent components that are typically included in a graphics command, while the dashed lines include components that are optional or that are only included in a subset of the graphics commands. The example graphics processor command format 2200 in Figure 22A includes data fields that identify a client 2202, a command operation code (opcode) 2204, and data 2206 for the command. A sub-opcode 2205 and a command size 2208 are also included in some commands.
[0258] In some embodiments, the client 2202 specifies a client unit of the graphics device that will process the command data. In some embodiments, the graphics processor command parser examines the client field of each command to condition further processing of the command and route the command data to the appropriate client unit. In some embodiments, the graphics processor client units include a memory interface unit, a render unit, a 2D unit, a 3D unit, and a media unit. Each client unit has a corresponding processing pipeline that processes the command. When a command is received by a client unit, the client unit reads the opcode 2204 and, if present, the sub-opcode 2205 to determine the operation to perform. The client unit executes the command using the information in the data field 2206. For some commands, an explicit command size 2208 is expected to specify the size of the command. In some embodiments, the command parser automatically determines the size of at least a portion of the command based on the command opcode. In some embodiments, the commands are aligned by a multiple of a doubleword. Other command formats are possible.
[0259] The flow diagram of Figure 22B depicts an example graphics processor command sequence 2210. In some embodiments, the data processing system software or firmware featuring an embodiment of a graphics processor uses variations of the command sequences shown to set up, execute, and terminate a set of graphics operations. The sample command sequences are shown and described for illustrative purposes only, as embodiments are not limited to these specific commands or to this command sequence. Furthermore, commands may be issued as a batch of commands in the command sequence such that the graphics processor processes the sequence of commands at least partially concurrently.
[0260] In some embodiments, the graphics processor command sequence 2210 may begin with a pipeline flush command 2212 to force any active graphics pipeline to complete any currently pending commands for that pipeline. In some embodiments, the 3D pipeline 2222 and the media pipeline 2224 do not operate simultaneously. A pipeline flush is performed to force the active graphics pipeline to complete any pending commands. In response to the pipeline flush, the graphics processor's command parser pauses command processing until the active drawing engine completes its pending operations and the associated read cache is invalidated. Optionally, any data in the render cache that is marked as "dirty" may be flushed to memory. In some embodiments, the pipeline flush command 2212 may be used before placing the graphics processor in a low power state or for pipeline synchronization.
[0261] In some embodiments, the pipeline select command 2213 is used when a command sequence requires the graphics processor to explicitly switch between pipelines. In some embodiments, the pipeline select command 2213 is only needed once in an execution context before issuing a pipeline command unless the context should issue commands for both pipelines. In some embodiments, a pipeline flush command 2212 is needed immediately before a pipeline switch due to the pipeline select command 2213.
[0262] In some embodiments, pipeline control commands 2214 are used to set up the graphics pipeline for operation and to program the 3D pipeline 2222 and the media pipeline 2224. In some embodiments, pipeline control commands 2214 set the pipeline state for an active pipeline. In one embodiment, pipeline control commands 2214 are used for pipeline synchronization and to clear data from one or more cache memories in an active pipeline before processing a batch of commands.
[0263] In some embodiments, commands to set return buffer state 2216 are used to set a set of return buffers for each pipeline to write data to. Some pipeline operations require the allocation, selection, or configuration of one or more return buffers to which the operation writes intermediate data during processing. In some embodiments, the graphics processor also uses one or more return buffers to store output data and to perform cross-thread communication. In some embodiments, return buffer state 2216 includes selecting the size and number of return buffers to use for a set of pipeline operations.
[0264] The remaining commands in the command sequence differ based on the active pipeline for the operation. Based on the pipeline decision 2220, the command sequence is tailored to the 3D pipeline 2222, starting at the 3D pipeline state 2230, or the media pipeline 2224, starting at the media pipeline state 2240.
[0265] Commands to set the 3D pipeline state 2230 include set 3D state commands for vertex buffer states, vertex element states, constant color states, depth buffer states, and other state variables that should be set before a 3D primitive command is processed. The values of these commands are determined at least in part based on the particular 3D API being used. In some embodiments, the 3D pipeline state 2230 commands can also selectively disable or bypass certain pipeline elements if those elements are not used.
[0266] In some embodiments, the 3D primitive 2232 command is used to submit a 3D primitive to be processed by the 3D pipeline. The command and associated parameters passed to the graphics processor by the 3D primitive 2232 command are forwarded to a vertex fetch function in the graphics pipeline. The vertex fetch function uses the 3D primitive 2232 command data to generate a vertex data structure. The vertex data structure is stored in one or more return buffers. In some embodiments, the 3D primitive 2232 command is used to perform vertex operations on the 3D primitives by vertex shaders. To process the vertex shaders, the 3D pipeline 2222 dispatches shader execution threads to the graphics processor execution units.
[0267] In some embodiments, the 3D pipeline 2222 is triggered by an execute 2234 command or event. In some embodiments, a register write triggers command execution. In some embodiments, execution is triggered by a "go" or "kick" command in a command sequence. In one embodiment, command execution is triggered using a pipeline synchronization command to flush the command sequence through the graphics pipeline. The 3D pipeline performs geometry operations on 3D primitives. Once the operations are complete, the resulting geometric objects are rasterized and the pixel engine colors the resulting pixels. Additional commands to control pixel shading and pixel backend operations may also be included for these operations.
[0268] In some embodiments, the graphics processor command sequence 2210 follows the media pipeline 2224 path when performing a media operation. In general, the specific use and manner of programming for the media pipeline 2224 depends on the media or computational operations to be performed. Certain media decoding operations may be offloaded to the media pipeline during media decoding. In some embodiments, the media pipeline may also be bypassed and media decoding may be performed in whole or in part using resources provided by one or more general purpose processing cores. In one embodiment, the media pipeline also includes elements for general purpose graphics processor unit (GPGPU) operation, where the graphics processor is used to perform SIMD vector operations using computational shader programs that are not explicitly related to rendering graphics primitives.
[0269] In some embodiments, the media pipeline 2224 is configured similarly to the 3D pipeline 2222. A set of commands to set the media pipeline state 2240 are dispatched or inserted into the command queue before the media object commands 2242. In some embodiments, the commands for the media pipeline state 2240 contain data to configure the media pipeline elements used to process the media object. This includes data that configures the video decoding and video encoding logic in the media pipeline, such as the encoding or decoding format. In some embodiments, the commands for the media pipeline state 2240 also support the use of one or more pointers to "indirect" state elements that contain a batch of state settings.
[0270] In some embodiments, the media object command 2242 provides a pointer to a media object for processing by the media pipeline. A media object includes a memory buffer containing the video data to be processed. In some embodiments, all media pipeline state must be valid before issuing the media object command 2242. Once the pipeline state is set and the media object command 2242 is queued, the media pipeline 2224 is triggered by an execute command 2244 or equivalent execution event (e.g., a register write). The output from the media pipeline 2224 may then be post-processed by operations provided by the 3D pipeline 2222 or the media pipeline 2224. In some embodiments, GPGPU operations are set up and executed similarly to media operations.
[0271] [Graphics Software Architecture] 23 depicts an exemplary graphics software architecture for a data processing system 2300 according to some embodiments. In some embodiments, the software architecture includes a 3D graphics application 2310, an operating system 2320, and at least one processor 2330. In some embodiments, the processor 2330 includes a graphics processor 2332 and one or more general-purpose processor cores 2334. The graphics application 2310 and the operating system 2320 each execute in the system memory 2350 of the data processing system.
[0272] In some embodiments, the 3D graphics application 2310 includes one or more shader programs that include shader instructions 2312. The shader language instructions may be in a high-level shader language, such as Direct3D's High-Level Shader Language (HLSL), OpenGL Shader Language (GLSL), etc. The application also includes executable instructions 2314 in a machine language suitable for execution by a general-purpose processor core 2334. The application also includes graphics objects 2316 defined by vertex data.
[0273] In some embodiments, the operating system 2320 is a Microsoft® Windows® operating system from Microsoft Corporation, a proprietary Unix-like operating system, or an open source UNIX-like operating system using a variant of the Linux® kernel. The operating system 2320 can support a graphics API 2322, such as a Direct3D API, an OpenGL API, or a Vulkan API. When the Direct3D API is in use, the operating system 2320 uses a front-end shader compiler 2324 to compile any shader instructions 2312 in HLSL into a lower-level shader language. The compilation can be a just-in-time (JIT) compilation, or the application can perform shader precompilation. In some embodiments, the higher-level shaders are compiled into lower-level shaders during compilation of the 3D graphics application 2310. In some embodiments, the shader instructions 2312 are provided in an intermediate form, such as a variant of the Standard Portable Intermediate Representation (SPIR) used by the Vulkan API.
[0274] In some embodiments, user mode graphics driver 2326 includes a backend shader compiler 2327 to convert shader instructions 2312 into a hardware-specific representation. When the OpenGL API is in use, shader instructions 2312 in the GLSL high level language are passed to user mode graphics driver 2326 for compilation. In some embodiments, user mode graphics driver 2326 uses operating system kernel mode functions 2328 to communicate with kernel mode graphics driver 2329. In some embodiments, kernel mode graphics driver 2329 communicates with graphics processor 2332 to dispatch commands and instructions.
[0275] [IP core implementation] One or more aspects of at least one embodiment may be implemented by representative code stored on a machine-readable medium that represents and / or defines logic within an integrated circuit, such as a processor. For example, the machine-readable medium may include instructions that represent various logic within a processor. When read by a machine, the instructions may cause the machine to assemble the logic to perform the techniques described herein. Such representations, known as "IP cores," are reusable units of logic for an integrated circuit that may be stored in a tangible machine-readable medium as a hardware model that describes the structure of the integrated circuit. The hardware model may be provided to various customers or manufacturing facilities that load the hardware model into an assembly machine that produces the integrated circuit. The integrated circuit may be assembled such that the circuit performs the operations described in connection with any of the embodiments described herein.
[0276] FIG. 24A is a block diagram illustrating an IP core development system 2400 that may be used to fabricate an integrated circuit to perform operations according to an embodiment. The IP core development system 2400 may be used to generate modular, reusable designs that may be incorporated into a larger design or may be used to construct an entire integrated circuit (e.g., a SOC integrated circuit). A design engine 2430 may generate a software simulation 2410 of the IP core design in a high-level programming language (e.g., C / C++). The software simulation 2410 may be used to design, test, and verify the behavior of the IP core using a simulation model 2412. The simulation model 2412 may include functional, behavioral, and / or timing simulation. A register transfer level (RTL) design 2415 is an abstraction of the behavior of an integrated circuit that models the flow of digital signals between hardware registers, including associated logic that is implemented using the modeled digital signals. In addition to the RTL design 2415, lower-level designs at the logic level or transistor level may also be generated, designed, or synthesized. Thus, the specific details of the initial design and simulation may vary.
[0277] The RTL design 2415 or equivalent may be further synthesized by a design facility into a hardware model, which may be a hardware description language (HDL) or other representation of physical design data. The HDL may be further simulated or tested to verify the IP core design. The IP core design may be stored using non-volatile memory 2440 (e.g., a hard disk, a flash memory, or any non-volatile storage medium) for delivery to a third party assembly facility 2465. Alternatively, the IP core design may be transmitted over a wired connection 2450 or a wireless connection 2460 (e.g., via the Internet). The assembly facility 2465 may then assemble an integrated circuit based at least in part on the IP core design. The assembled integrated circuit may be configured to perform operations according to at least one embodiment described herein.
[0278] 24B illustrates a side cross-sectional view of an integrated circuit package assembly 2470 according to some embodiments described herein. The integrated circuit package assembly 2470 represents an implementation of one or more processor or accelerator devices described herein. The package assembly 2470 includes multiple units 2672, 2674 of hardware logic coupled to a substrate 2480. The logic 2672, 2674 may be implemented at least partially in configurable logic or fixed function logic hardware and may include one or more portions of any of the processor cores, graphics processors, or other accelerator devices described herein. Each unit of logic 2672, 2674 may be implemented in a semiconductor die and coupled to the substrate 2480 via an interconnect structure 2473. The interconnect structure 2473 may be configured to convey electrical signals between the logic 2672, 2674 and the substrate 2480 and may include interconnects such as, but not limited to, bumps or pillars. In some embodiments, the interconnect structure 2473 may be configured to carry electrical signals, such as, for example, input / output (I / O) signals and / or power or ground signals associated with the operation of the logic 2672, 2674. In some embodiments, the substrate 2480 is an epoxy-based laminate substrate. The substrate 2480 may include other suitable types of substrates in other embodiments. The package assembly 2470 may be connected to other electrical devices via the package interconnect 2483. The package interconnect 2483 may be coupled to a surface of the substrate 2480 to carry electrical signals to other electrical devices, such as a motherboard, another chipset, or a multi-chip module.
[0279] In some embodiments, the units of logic 2672, 2674 are electrically coupled to a bridge 2482 configured to convey electrical signals between the logic 2672, 2674. The bridge 2482 may be a dense interconnect structure that provides a route for the electrical signals. The bridge 2482 may include a bridge substrate made of glass or a suitable semiconductor material. Electrical routing structures may be formed on the bridge substrate to provide chip-to-chip connections between the logic 2672, 2674.
[0280] Although two units of logic 2672, 2674 and bridge 2482 are shown, the embodiments described herein may include more or less logic units on one or more dies. One or more dies may be connected by zero or more bridges, such that bridge 2482 may be omitted if the logic is contained on a single die. Alternatively, multiple dies or units of logic may be connected by one or more bridges. Additionally, multiple logic units, dies, and bridges may be linked in other possible configurations, including three-dimensional configurations.
[0281] [Example of a System-on-a-Chip Integrated Circuit] 25-26 depict exemplary integrated circuits and associated graphics processors that may be assembled using one or more IP cores in accordance with various embodiments described herein. In addition to what is depicted, other logic and circuitry may be included, including additional graphics processors / cores, processors / cores, peripheral interface controllers, or general-purpose processor cores.
[0282] 25 is a block diagram illustrating an exemplary system-on-a-chip integrated circuit 2500 that may be assembled using one or more IP cores, according to an embodiment. The exemplary integrated circuit 2500 includes one or more application processors 2505 (e.g., a CPU) and at least one graphics processor 2510, and may further include an image processor 2515 and / or a video processor 2520, any of which may be modular IP cores from the same or multiple different design facilities. The integrated circuit 2500 includes a USB controller 2525, a UART controller 2530, an SPI / SDIO controller 2535, and an I / O controller 2540. 2 S / I 2 The integrated circuit 2500 may further include peripheral or bus logic including a C controller 2540. Additionally, the integrated circuit 2500 may include a display device 2545 coupled to one or more of a High Definition Multimedia Interface (HDMI®) controller 2550 and a Mobile Industry Processor Interface (MIPI) display interface 2555. Storage may be provided by a flash memory subsystem 2560 including a flash memory and a flash memory controller. Memory interface may be provided via a memory controller 2565 for access to a SCRAM or SRAM memory device. Some integrated circuits further include an embedded security engine 2570.
[0283] 26A-26B are block diagrams illustrating example graphics processors for use within an SoC in accordance with embodiments described herein. FIG. 26A illustrates an example graphics processor 2610 of an SoC integrated circuit that may be fabricated using one or more IP cores in accordance with embodiments. FIG. 26B illustrates a further example graphics processor 2640 of an SoC integrated circuit that may be fabricated using one or more IP cores in accordance with embodiments. The graphics processor 2610 of FIG. 26A is an example of a low-power graphics processor core. The graphics processor 2640 of FIG. 26B is an example of a higher performance graphics processor core. Each of the graphics processors 2610, 2640 may be a variation of the graphics processor 2510 of FIG. 25.
[0284] As shown in FIG. 26A, the graphics processor 2610 includes a vertex processor 2605 and one or more fragment processors 2615A-2615N (e.g., 2615A, 2615B, 2615C, 2615d, through 2615N-1, and 2615N). The graphics processor 2610 can execute different shader programs through separate logic such that the vertex processor 2605 is optimized to execute the operations of a vertex shader program, while the one or more fragment processors 2615A-2615N execute the fragment (e.g., pixel) shading operations of a fragment or pixel shader program. The vertex processor 2605 executes the vertex processing stage of the 3D graphics pipeline, generating primitive and vertex data. The fragment processors 2615A-2615N use the primitive and vertex data generated by the vertex processor 2605 to generate a frame buffer that is displayed on a display device. In one embodiment, fragment processors 2615A-2615N are optimized to execute fragment shader programs such as those provided in the OpenGL API, which can be used to perform operations similar to pixel shader programs such as those provided in the Direct3D API.
[0285] Graphics processor 2610 further includes one or more memory management units (MMUs) 2620A-2620B, caches 2625A-2625B, and circuit interconnects 2630A-2630B. The one or more MMUs 2620A-2620B provide virtual-to-physical address mapping for graphics processor 2610, which includes vertex processor 2605 and / or fragment processors 2615A-2615N, which may reference vertices or image / texture data stored in memory in addition to vertices or image / texture data stored in one or more caches 2625A-2625B. In one embodiment, one or more MMUs 2620A-2620B may synchronize with other MMUs in the system, including one or more MMUs associated with one or more of application processors 2505, image processor 2515, and / or video processor 2520 of FIG. 25, such that each processor 2505-2520 can participate in a shared or unified virtual memory system. One or more circuit interconnects 2630A-2630B enable graphics processor 2610 to interface with other IP cores in the SoC via the SoC's internal bus or via direct connections, according to an embodiment.
[0286] As shown in FIG. 26B, the graphics processor 2640 includes one or more MMUs 2620A-2620B, caches 2625A-2625B, and circuit interconnects 2630A-2630B of the graphics processor 2610 of FIG. 26A. The graphics processor 2640 includes one or more shader cores 2655A-2655N (e.g., 2655A, 2655B, 2655C, 2655D, 2655E, 2655F, through 2655N-1, and 2655N) that provide a unified shader core architecture in which a single core or type or cores can execute all types of programmable shader code, including shader program code to implement vertex shaders, fragment shaders, and / or compute shaders. The exact number of shader cores can vary depending on the embodiment and implementation. Additionally, the graphics processor 2640 includes an inter-core task manager 2645 that acts as a thread dispatcher that dispatches execution threads to one or more shader cores 2655A-2655N and tiling unit 2658 to accelerate tiling operations for tile-based rendering in which rendering operations on a scene are subdivided in image space, for example to exploit local spatial coherence within the scene or to optimize use of internal caches.
[0287] [SoC architecture decomposition] Constructing ever larger silicon dies is difficult for a variety of reasons. As silicon dies get larger, manufacturing yields become smaller and process technology requirements of different components may diverge. On the other hand, to have a high performance system, critical components should be interconnected by high speed, high bandwidth, low latency interfaces. These conflicting needs pose a challenge to the development of high performance chips.
[0288] The embodiments described herein provide a technique to split the architecture of a SoC circuit into different chiplets that can be packaged on a common chassis. In one embodiment, a graphics processing unit or parallel processor is composed of different silicon chiplets that are manufactured separately. A chiplet is an at least partially packaged integrated circuit that includes different logic units that can be collected together with other chiplets into a larger package. Various sets of chiplets that include different IP core logic can be assembled into a single device. Furthermore, chiplets can be integrated into a base die or base chiplet using active interposer technology. The concepts described herein allow interconnection or communication between different forms of IP within a GPU. Development of IPs in different processes can be intermixed. This avoids the complexity of converging multiple IPs into the same process, especially in large SoCs with several flavors of IP.
[0289] Enabling the use of multiple process technologies improves time to market and provides a cost-effective way to create multiple product SKUs. For customers, this means getting products that better meet their requirements in a cost-effective and timely manner. Additionally, decomposed IP is more amenable to being independently power-gated, allowing components that are not being used in a given workload to be powered off, reducing overall power consumption.
[0290] FIG. 27 illustrates a parallel computing system 2700 according to an embodiment. In one embodiment, the parallel computing system 2700 includes a parallel processor 2720, which can be a graphics processor or a computation accelerator as described herein. The parallel processor 2720 includes a global logic 2701 unit, an interface 2702, a thread dispatcher 2703, a media unit 2704, a set of computing units 2705A-2705H, and a cache / memory unit 2706. The global logic unit 2701 includes global functionality for the parallel processor 2720, which in one embodiment includes device configuration registers, a global scheduler, power management logic, etc. The interface 2702 can include a front-end interface for the parallel processor 2720. The thread dispatcher 2703 can receive a workload from the interface 2702 and dispatch threads of the workload to the computing units 2705A-2705H. If the workload includes any media operations, at least some of those operations may be performed by media unit 2704. Media unit 2704 also offloads some operations to compute units 2705A-2705H. Cache / memory unit 2706 may include cache memory (e.g., L3 cache) and local memory (e.g., HBM, GDDR) for parallel processor 2720.
[0291] 28A-28B depict a hybrid logical / physical view of a decomposed parallel processor according to embodiments described herein. Figure 28A depicts a decomposed parallel computing system 2800. Figure 28B depicts a chiplet 2830 of the decomposed parallel computing system 2800.
[0292] As shown in FIG. 28A, a decomposed computing system 2800 can include a parallel processor 2820 in which various components of the parallel processor SOC are distributed across multiple chiplets. Each chiplet can be a distinct IP core that is independently designed and configured to communicate with other chiplets through one or more common interfaces. The chiplets include, but are not limited to, compute chiplets 2805, media chiplets 2804, and memory chiplets 2806. Each chiplet can be separately manufactured using different process technologies. For example, the compute chiplets 2805 may be manufactured using the smallest or most advanced process technology available at the time of manufacture, while the memory chiplets 2806 or other chiplets (e.g., I / O, networking, etc.) may be manufactured using larger or less advanced process technologies.
[0293] The various chiplets may be attached to a base die 2810 and configured to communicate with each other and with logic within the base die 2810 via an interconnect layer 2812. In one embodiment, the base die 2810 may include global logic 2801, which may include scheduler 2811 and power management 2821 logic units, an interface 2802, a dispatch unit 2803, and an interconnect fabric module 2808 coupled or integrated with one or more L3 cache banks 2809A-2809N. The interconnect fabric module 2808 may be an inter-chiplet fabric that is embedded within the base die 2810. The logic chiplets may use the fabric 2808 to relay messages between the various chiplets. Additionally, L3 cache banks 2809A-2809N in the base die and / or L3 cache banks in memory chiplets 2806 can cache data read from DRAM chiplets in memory chiplets 2806 and data transmitted to the DRAM chiplets and host system memory.
[0294] In one embodiment, global logic 2801 is a microcontroller that can execute firmware to perform scheduler 2811 and power management 2821 functions for parallel processors 2820. The microcontroller that executes the global logic may be tailored to the target use case of the parallel processors 2820. Scheduler 2811 may perform global scheduling operations for the parallel processors 2820. Power management 2821 functions may be used to enable or disable individual chiplets in the parallel processor when those chiplets are not in use.
[0295] The various chiplets of the parallel processor 2820 may be designed to perform specific functions that would otherwise be integrated onto a single die in existing designs. The set of compute chiplets 2805 may include a cluster of compute units (e.g., execution units, streaming multiprocessors, etc.) that include programmable logic to execute computational or graphics shader instructions. The media chiplets 2804 may include hardware logic to accelerate media encoding and decoding operations. The memory chiplets 2806 may include volatile memory (e.g., DRAM) and one or more SRAM cache memory banks (e.g., L3 banks).
[0296] As shown in FIG. 28B, each chiplet 2830 may include common components and application-specific components. Chiplet logic 2836 within the chiplet 2830 may include chiplet specific components, such as a streaming multiprocessor, compute units, or an array of execution units as described herein. The chiplet logic 2836 may be coupled with any cache or shared local memory 2838 or may include a cache or shared local memory within the chiplet logic 2836. The chiplet 2830 may include a fabric interconnect node 2842 that receives commands via the inter-chiplet fabric. Commands and data received via the fabric interconnect node 2842 may be temporarily stored in an interconnect buffer 2839. Data transmitted to and received from the fabric interconnect node 2842 may be stored in an interconnect cache 2840. Power control 2832 and clock control 2834 logic may also be included within the chiplet. The power control 2832 and clock control 2834 logic can receive configuration commands via the fabric that can configure dynamic voltage and frequency scaling for the chiplets 2830. In one embodiment, each chiplet can have independent clock and power domains and can be clock gated and power gated independently from other chiplets.
[0297] At least some of the components in the depicted chiplets 2830 may also be included in the logic embedded in the base die 2810 of Figure 28A. For example, logic in the base die that communicates with the fabric may include variations of fabric interconnect nodes 2842. Base die logic that may be independently clocked or power gated may include variations of power control 2832 and / or clock control 2834 logic.
[0298] 29A-29B depict package views of a decomposed parallel processor according to an embodiment. Figure 29A depicts the physical layout of a package assembly 2920. Figure 29B depicts the interconnects between multiple chiplets 2904, 2906 and an interconnect fabric 2940.
[0299] As shown in FIG. 29A, package assembly 2920 can include multiple units of hardware logic chiplets connected to substrate 2910 (e.g., base die). Hardware logic chiplets can include dedicated hardware logic chiplets 2902, logic or I / O chiplets 2904, and / or memory chiplets 2905. Hardware logic chiplets 2902 and logic or I / O chiplets 2904 can be at least partially implemented in configurable logic or fixed function logic hardware and can include one or more portions of any of the processor cores, graphics processors, parallel processors, or other accelerator devices described herein. Memory chiplets 2905 can be DRAM (e.g., GDDR, HBM) memory or cache (SRAM) memory as described and illustrated herein.
[0300] Each chiplet may be fabricated as a separate semiconductor die and coupled to substrate 2910 via interconnect structures 2903. Interconnect structures 2903 may be configured to carry electrical signals between the various chiplets and logic within substrate 2910. Interconnect structures 2903 may include interconnects such as, for example, but not limited to, bumps or pillars. In some embodiments, interconnect structures 2903 may be configured to carry electrical signals such as, for example, input / output (I / O) signals and / or power or ground signals associated with operation of the logic, I / O, and memory chiplets.
[0301] In some embodiments, the substrate 2910 is an epoxy-based laminate substrate. The substrate 2910 may include other suitable types of substrates in other embodiments. The package assembly 2920 may be connected to other electrical devices via a package interconnect 2914. The package interconnect 2914 may be coupled to a surface of the substrate 2910 to conduct electrical signals to other electrical devices, such as a motherboard, another chipset, or a multi-chip module.
[0302] In some embodiments, logic or I / O chiplets 2904 and memory chiplets 2905 may be electrically coupled through a bridge 2917 configured to convey electrical signals between the logic or I / O chiplets 2904 and the memory chiplets 2905. The bridge 2917 may be a dense interconnect structure that provides a route for the electrical signals. The bridge 2917 may include a bridge structure made of glass or a suitable semiconductor material. Electrical routing functions may be formed in the bridge structure to provide inter-chip connections between the logic or I / O chiplets 2904 and the memory chiplets 2905. The bridge 2917 may also be referred to as a silicon bridge or an interconnect bridge. For example, the bridge 2917 is an Embedded Multi-die Interconnect Bridge (EMIB) in some embodiments. In some embodiments, the bridge 2917 may simply be a direct connection from one chiplet to another chiplet.
[0303] The substrate 2910 may include hardware components for I / O 2911, cache memory 2912, and other hardware logic 2913. A fabric 2915 may be embedded in the substrate 2910 to enable communication between the various logic chiplets and the logic 2911, 2913 within the substrate 2910.
[0304] In various embodiments, package assembly 2920 can include fewer or more components and chiplets interconnected by fabric 2915 or one or more bridges 2917. The chiplets in package assembly 2920 can be arranged in a 3D or 2.5D arrangement. In general, bridge structures 2917 can be used to facilitate point-to-point interconnects between, for example, logic or I / O chiplets 2904 and memory chiplets 2905. Fabric 2915 can be used to interconnect various logic and / or I / O chiplets (e.g., chiplets 2902, 2904, 2911, 2913) with other logic and / or I / O chiplets. In one embodiment, cache memory 2912 in substrate 2910 can operate as a global cache for package assembly 2920, or as part of a distributed global cache, or as a dedicated cache for fabric 2915.
[0305] As shown in Figure 29B, memory chiplets 2906 can connect with logic or I / O chiplets 2904 via chiplet interconnects 2935 routed through interconnect bridges 2947. Interconnect bridges 2947 can be a variation of bridges 2917 embedded in substrate 2910 of package assembly 2920 shown in Figure 29A. Bridges or I / O chiplets 2904 can communicate with other chiplets via interconnect fabric 2940, which is a variation of fabric 2915 of Figure 29A.
[0306] In one embodiment, memory chiplet 2906 includes a set of memory banks 2931 corresponding to the memory technology provided by the chiplet. Memory banks 2931 can include any of the types of memory described herein, including but not limited to DRAM, SRAM, or flash, or 3D XPoint memory. Memory control protocol layer 2932 can enable control of memory banks 2931 and can include logic for one or more memory controllers. Interconnect bridge protocol layer 2933 can relay messages between memory control protocol layer 2932 and interconnect bridge I / O layer 2934. Interconnect bridge I / O layer 2934 can communicate with interconnect bridge I / O layer 2936 via chiplet interconnect 2935. Interconnect bridge I / O layer 2934, 2936 can represent physical layers that transmit signals to or receive signals from corresponding interconnects via chiplet interconnect 2935. The physical I / O layer can include circuitry that drives signals on and / or receives signals from the chiplet interconnect 2935. An interconnect bridge protocol layer 2937 in the logic or I / O chiplet 2904 can convert signals from the interconnect bridge I / O layer 2936 into messages or signals that can be communicated to the computation or I / O logic 2939. In one embodiment, a digital adapter layer 2938 can be used to facilitate the conversion of signals into messages or signals used by the computation or I / O logic 2939.
[0307] The compute or I / O logic 2939 can communicate with other logic or I / O chiplets via interconnect fabric 2940. The compute or I / O logic 2939, in one embodiment, includes embedded fabric node logic 2939 that can communicate with interconnect fabric 2940, such as fabric interconnect node 2842 in FIG.
[0308] In one embodiment, control layer 2986 in memory chiplet 2906 can communicate with control layer 2970 in logic or I / O chiplet 2904. These control layers 2968, 2970 can be used to propagate or transmit certain control signals in an out-of-band manner, for example, to send power and configuration messages between interface bus protocol layer 2937 of logic or I / O chiplet 2904 and interface bus protocol layer 2933, memory control protocol layer 2932, and / or memory banks 2931 of memory chiplet 2906.
[0309] FIG. 30 illustrates a message transport system 300 for an interconnect fabric, according to an embodiment. The message transport system 300 may be configured to handle traffic at different rates depending on the available interface width. The specific interface width may vary from chiplet to chiplet, or the speed or configuration of the fabric may be adjusted to allow data to be transmitted at a rate appropriate for the various functional units being interconnected. The transport layer may also cross one or more clock domains through the use of clock domain crossing FIFOs at the boundaries between clock domains to allow data to be transferred between clock domains. The transport layer may also be divided into one or more sublayers, each sublayer including one or more clock domains. Although a transport layer is described and illustrated, in some embodiments, the operations depicted may be performed in the data link layer of the interconnect fabric.
[0310] In one embodiment, a first functional unit 3001A in an origin layer 3010 can communicate with a second functional unit 3001B in a destination layer 3013 via one or more transport layers 3011, 3012. The origin layer 3010 can be logic or I / O within a chiplet or within a substrate of a package assembly, such as substrate 2910 of package assembly 2920 as seen in FIG. 29A. The destination layer 3013 can also be logic or I / O within a chiplet or within a substrate. In one embodiment, the origin layer 3010 and / or destination layer 3013 can be associated with cache memory within a chiplet or substrate / base die.
[0311] One or more transport layers 3011, 3012 can be in separate clock domains. For example, transport layer 3011 can be in a first clock domain while transport layer 3012 can be in a second clock domain. The separate clock domains can operate at different frequencies. Data can be transferred between the clock domains via a clock crossing module 3003 in the transport layer. In one embodiment, a first buffer or high speed memory module 3004 in the clock crossing module 3003 can buffer data that is relayed to a second buffer or high speed memory module 3006 via a cross crossing FIFO 3005. The first buffer / high speed memory module 3004 can be in a first clock domain while the second buffer / high speed memory module 3006 can be in a second clock domain.
[0312] Functional unit 3001A and functional unit 3001B can transmit and receive messages to and from the fabric via respective fabric interfaces 3002A, 3002B. Fabric interfaces 3002A, 3002B can dynamically configure the width of the connections used within the fabric to relay messages and signals across the transport layer, as shown below in Figures 31-32.
[0313] 31 depicts the transmission of messages or signals between functional units across multiple physical links of an interconnect fabric. In one embodiment, when a communication channel between a set of functional units has bandwidth requirements that exceed what can be provided through a single physical link, multiple links may be used to facilitate communication between the functional units.
[0314] In one embodiment, the first functional unit 3101A can send messages or signals to the first fabric interface 3102A. The first fabric interface 3102A can branch messages or signals and send messages or signals across multiple physical links as a single virtual channel. For example, multiple physical links can be assigned to the same virtual channel, each capable of carrying messages or signals for the virtual channel.
[0315] Data may be transferred between clock domains via a clock crossing module 3103 in one or more transport layers 3011, 3012. The clock crossing module 3103 transfers channel messages or signals across multiple physical links via multiple buffers or high speed memories 3104A-3104B, 3106A-3106B (and clock domain crossing FIFOs) in a similar manner as shown in Figure 30. The multiple physical links may be converged at a second fabric interface 3102B before the messages or signals are provided to a second functional unit 3101B.
[0316] 32 illustrates the transmission of messages or signals of multiple functional units across a single channel of the interconnect fabric. In one embodiment, multiple virtual channels may be transported across a physical link when a communication channel between a set of functional units has bandwidth requirements that do not utilize all of the available bandwidth of the physical link. The virtual channels may be switched over time along the physical link, and specific sets of data lines within the physical link may be assigned to specific functional units.
[0317] In one embodiment, a first set of functional units 3201A, 3211A can communicate with a second set of functional units 3201B, 3211B over a single physical link of the interconnect fabric. The functional units 3201A, 3201B can be associated with a first virtual channel, while the functional units 3211A, 3211B can be associated with a second virtual channel. The first and second virtual channels can be converged at a first fabric interface 3202A. Messages or signals can be relayed across one or more transport layers 3011, 3012. Data can be transmitted between clock domains via a clock crossing module 3203 in one or more transport layers 3011, 3012 via multiple buffers or high-speed memories 3204, 3206 (and clock domain crossing FIFOs) in a similar manner as shown in FIG. 30. Multiple virtual channels may be branched at the second fabric interface 3202B before messages or signals are provided to a second set of functional units 3201B, 3211B.
[0318] 33 illustrates a method 3300 for configuring fabric connections for functional units in a decomposed parallel processor. The interconnect fabric of the decomposed parallel processor described herein is configurable to carry messages and / or signals for a variety of different components, which may be IP cores having different designers and / or manufacturers. The specific types of links used by each functional unit within a chiplet or base die logic component are configurable. In one embodiment, the links are configurable at the fabric interconnect node used by the functional unit.
[0319] In one embodiment, the fabric interconnect node may receive bandwidth configuration data for functional units that are to be configured to communicate over the interconnect fabric in the parallel processor package (block 3302). The configuration data may be statically provided during initial assembly and provisioning of the decomposed parallel processor or may be dynamically configured during initialization of the decomposed parallel processor. For dynamic initialization, the fabric interconnect node may receive a bandwidth configuration request from the functional unit that specifies the physical width and frequency of the interconnect between the functional unit and the fabric interconnect node, as well as the bandwidth requirements of the functional unit.
[0320] The fabric interconnect node may then analyze the configured interconnect width and frequency for the functional unit (block 3304). The fabric interconnect node may then configure converging and / or dropping links for the functional unit (block 3306). Once configured, the fabric interconnect node may relay messages and / or signals of the functional unit across the configured links (block 3308).
[0321] 34 illustrates a method 3400 for relaying messages and / or signals across an interconnect fabric in a decomposed parallel processor. The interconnect fabric of the decomposed parallel processors described herein can relay messages and / or signals across one or more layers of the decomposed parallel processor while traversing multiple clock domains.
[0322] In one embodiment, a first functional unit in a chiplet or base die of a processor may generate data in the form of a message or signal to be transmitted (block 3402). The first functional unit may transmit the message or signal to the interconnect fabric via a first fabric interface node (block 3404). The fabric interface node may converge or drop virtual channels to physical transport links (block 3406), as shown in Figures 31 and 32. The message or signal to be transmitted may be associated with a virtual channel and transmitted over the associated virtual channel. How the fabric interconnect performs forwarding and / or switching operations of the message or signal may be influenced by the virtual channel assigned to the message or signal. Furthermore, multiple virtual channels may be aggregated into a single physical link, or a virtual channel may be carried by multiple physical links.
[0323] In one embodiment, the interconnect fabric can carry messages or signals across multiple clock domains in one or more transport and / or data link layers (block 3408). One or more clock crossing modules including high speed memory and domain crossing FIFOs can be used to cross the multiple clock domains. In one embodiment, each chiplet can have a separate clock domain with respect to the interconnect fabric. The interconnect fabric can also have multiple clock domains. Transmitting messages or signals across multiple clock domains can include switching messages or signals with switching logic in the interconnect fabric.
[0324] A second fabric interface node may receive the message or signal (block 3410). The second fabric interface node may then drop or converge the virtual channel from the physical transport link at the second fabric interface node (block 3412). Multiple virtual channels may be dropped from a single physical link, or alternatively, virtual channels may be converged from multiple physical links. A second functional unit within the chiplet or base die of the processor may then receive the data in the form of a message or signal at second hardware logic (block 3414). The second functional unit may then perform an operation based on the received data.
[0325] 35 illustrates a method 3500 for power gating chiplets per workload. In one embodiment, power control logic in a decomposed parallel processor can determine which chiplets or logic units should be powered while executing a workload based on the requirements of the workload. In one embodiment, the power control logic can cooperate with other global logic, such as a global scheduler or front-end interface, to determine which components will be used to process a workload.
[0326] The method 3500 includes receiving a command buffer for a workload to be executed on a parallel processor (block 3502). The command buffer may be received, for example, at a global scheduler or a front-end interface. The method 3500 further includes determining a set of chiplets to be used to execute the workload (block 3504). This determination may be performed by determining a global configuration and / or type of functional units to be used to execute the commands in the command buffer.
[0327] The method 3500 further includes determining whether any functional chiplets to be used to process the workload are present in the power-gated chiplets or powering up the chiplets to be used if they are not already powered (block 3506). Additionally, functional units not used to process the workload may be determined. If not all functional units in the chiplet are used to process the workload, the power control logic may power down (power gate) the chiplets not used to process the workload (block 3508). To power up or down the chiplets, the global power control logic may signal the local power control logic in the chiplet. The local power control logic in the chiplet may then execute the appropriate power-down sequence of the chiplets. The decomposed parallel processor may then execute the workload using the powered-up (e.g., active) chiplets (block 3510).
[0328] [Enabling product SKUs based on chiplet configuration] Semiconductor dies are tested during manufacturing to evaluate the integrated circuits formed on the die. Standard tests for overall functionality may be performed by probe testing the die on a wafer. Burn-in tests may be performed after the die are separated and packaged, or with a test harness for bare dies. Defective dies may be discarded. However, dies that pass initial testing but fail subsequent tests most often may be less likely to be able to operate correctly. This binning-out process selects low or high performance dies and can target those dies for higher or lower performance products in various stock keeping units (SKUs). For monolithic SoC integrated circuits, the binning-out process is a coarse-grained process. Processors that contain some defective computational or graphics cores may be down-binned, but there should be a minimum number of defect-free components to meet minimum product requirements.
[0329] With the decomposed SoC architecture described herein, individual chiplets can be tested and binned at the chiplet level, and the SKU level can be determined for a product during assembly based on the demand for a given product SKU. During assembly, different product SKUs with different configurations can be assembled with different amounts of memory, different functionality, and different performance by specifying a specific chiplet or different bins of the same chiplet design.
[0330] The disassembled processor package can be configured to accept interchangeable chiplets. Interchangeability is enabled by specifying a standard physical interconnect for the chiplets that can enable the chiplets to interface with a fabric or bridge interconnect. Chilets from different IP designers can be fitted with a common interconnect, allowing such chiplets to be interchangeable during assembly. The fabric and bridge interconnect logic on the chiplets can then be configured to match the actual interconnect layout of the chiplet's on-board logic. Furthermore, data from the chiplets can be transmitted across the inter-chiplet fabric using encapsulation such that the actual data being transferred is opaque to the fabric, further enabling interchangeability of individual chiplets. With such an interchangeable design, higher or lower density memories can be inserted into memory chiplet slots, while higher or lower core count compute or graphics chiplets can be inserted into logic chiplets.
[0331] Functionality may also be determined during assembly. For example, media chiplets may be added or removed based on product usage and demand. In some products, on-package networking or other communication chiplets may be added. In some products, different types of host connections may be enabled using different chiplets. For example, host interconnect version changes involve modifications to the interconnect logic without a change in physical form factor, and upgrading to a new interconnect version may be performed by changing the host interconnect chiplets during assembly without requiring a redesign of the SoC to insert new interconnect logic into the monolithic die.
[0332] Chiplet binning may be further enabled by the provision of a chiplet test harness that fits into a standardized chassis interface, which may enable rapid testing and binning of chiplets for different SKUs.
[0333] FIG. 36 depicts a parallel processor assembly 3600 including a swappable chiplet 3602. The swappable chiplet 3602 may be assembled into a standardized slot on one or more base chiplets 3604, 3608. The base chiplets 3604, 3608 may be coupled through a bridge interconnect 3606, which may be similar to other bridge interconnects described herein. The memory chiplets may be connected to logic or I / O chiplets through the bridge interconnect. The I / O and logic chiplets may communicate through an interconnect fabric. The base chiplets may support one or more slots with a standardized format for one of logic or I / O or memory / cache, respectively. Different memory densities may be incorporated into the chiplet slots based on the target SKU of the product. Furthermore, logic chiplets with different numbers of types of functional units may be selected at assembly time based on the target SKU of the product. Furthermore, chiplets containing different types of IP logic cores may be inserted into the swappable chiplet slots.
[0334] FIG. 37 illustrates a swappable chiplet system 3700, according to an embodiment. In one embodiment, the swappable chiplet system 3700 includes at least one base chiplet 3710 that includes a plurality of memory chiplet slots 3701A-3701F and a plurality of logic chiplet slots 3702A-3702F. The logic chiplet slots (e.g., 3702A) and the memory chiplet slots (e.g., 3701A) may be connected by an interconnect bridge 3735, which may be similar to other interconnect bridges described herein. The logic chiplet slots 3702A-3702F may be interconnected via a fabric interconnect 3708. The fabric interconnect 3708 includes switching logic 3718, which may be configured to enable relaying of data packets between the logic chiplet slots in a data agnostic manner by encapsulating the data in fabric packets. The fabric packets may then be switched to a destination slot within the fabric interconnect 3708.
[0335] The fabric interconnect 3708 may include one or more physical data channels. One or more programmable virtual channels may be carried by each physical channel. The virtual channels may be independently arbitrated with channel access negotiated separately for each virtual channel. Traffic on the virtual channels may be classified into one or more traffic classes. In one embodiment, a prioritization system allows the virtual channels and traffic classes to be assigned relative priorities for arbitration. In one embodiment, a traffic balancing algorithm operates to maintain approximately equal bandwidth and throughput for each node coupled to the fabric. In one embodiment, the fabric interconnect logic operates at a higher clock rate than the nodes coupled to the fabric to allow for reduced interconnect width while maintaining bandwidth requirements between the nodes. In the event that a particular node requires higher bandwidth, multiple physical links may be combined to carry a single virtual channel, as seen above in FIG. 31. In one embodiment, each physical link is clock gated separately when idle. An early indication of the next operation may be used as a trigger to wake up the physical link before data should be transmitted.
[0336] FIG. 38 is an illustration of multiple traffic classes carried on virtual channels, according to an embodiment. A first fabric connector 3802 and a second fabric connector 3804 facilitate communication over a fabric channel 3806 having up to 'M' virtual channels 3806A-3806M. The virtual channels allow for the transfer of variable length information over a fixed set of physical channels. The virtual channels may be permanent virtual channels, or they may be dynamically enabled or disabled based on system configuration. The use of permanent virtual channels allows for fixed channel IDs, minimizing virtual channel management overhead. Dynamically configuring channels allows for greater design flexibility at the expense of increased channel management overhead.
[0337] Each virtual channel may be assigned multiple traffic classes. A traffic class is a classification of traffic that is relevant for arbitration. Each virtual channel may carry up to 'N' traffic classes. Each class of traffic is assigned to a particular virtual channel through programming (fuses, configuration registers, etc.). Up to 'L' classes of traffic classes may be assigned to a given virtual channel. [Table 5]
[0338] Table 5 above illustrates an example traffic class virtual channel assignment as depicted in FIG. 38. The fabric interconnect may include logic to classify each unit of incoming traffic and ensure that the incoming unit moves within its assigned virtual channel. In one embodiment, data transmission on a channel occurs in first-in-first-out (FIFO) order, and channel arbitration is based on the virtual channel. Traffic within a virtual channel may block the transmission of additional traffic on the same virtual channel. However, a given virtual channel does not block different virtual channels. Thus, traffic on different virtual channels is arbitrated independently.
[0339] Coherency is maintained during data transmission between fabric interconnect nodes. In one embodiment, data of outgoing threads on a GPGPU or parallel processor is routed within the same traffic class, and the traffic class is assigned to a particular virtual channel. Data within a single traffic class on a single virtual channel is transmitted in FIFO order. Thus, data from a single thread is strictly ordered as it is transmitted through the fabric, and per-thread coherency is maintained to avoid read-after-write or write-after-read data hazards. In one embodiment, thread group coherency is maintained by a global synchronization mechanism between resource nodes. [Table 6]
[0340] Table 6 above shows an example traffic class prioritization. A priority algorithm may be programmed to determine the priority to assign to each of the traffic classes. Programmable traffic class prioritization allows traffic classes to be used as an arbitrary traffic grouping mechanism. Then, traffic may simply be grouped within a class to maintain coherency, or certain traffic may be assigned a high priority and used exclusively for high priority data. For example, classes 1 and 4, each assigned to virtual channel 1 3806B, may be assigned a priority of 2. Classes 2 and 5, each assigned to virtual channel 1 3806A, may be assigned a priority of 1. Traffic class 'N' may be assigned to virtual channel 3 3806C, which has a priority of 3. Class 2 traffic may be delay-sensitive data that should be transmitted as soon as possible or not blocked by other traffic classes, while class 1 traffic may be moderately delay-sensitive traffic from a single thread that is grouped to maintain coherency.
[0341] Traffic classes may be assigned a priority relative to all traffic classes or relative to the priority of traffic classes on the same virtual channel. In one embodiment, a priority scheme is designed by assigning weights to traffic classes, where a higher weight indicates a higher priority. A fair prioritization algorithm may be used, where each participant is guaranteed a minimum amount of bandwidth to prevent starvation. In one embodiment, an absolute priority algorithm is used under certain circumstances, where higher priority traffic always blocks lower priority traffic.
[0342] Additional algorithms are implemented to prevent communication deadlocks when absolute priority is used. The combined use of virtual channels and traffic classes reduces the likelihood of deadlocks because a single traffic class having absolute priority on a given virtual channel will not block traffic on another virtual channel. In one embodiment, if a starvation condition or potential deadlock is detected on one virtual channel, the blocked traffic class may be reassigned to another virtual channel. [Table 7]
[0343] Table 7 above shows an example virtual channel prioritization. Like traffic classes, each virtual channel may also receive a priority, and channel arbitration may take into account the relative priority of the virtual channels. For example, data traffic on virtual channel 2 may have a higher relative priority than data on other virtual channels. A weighted priority system may be used by the virtual channel prioritization, where a higher weight indicates a higher priority. A fair priority system or an absolute priority system may be used.
[0344] FIG. 39 illustrates a method 3900 of agnostic data transmission between slots of a replaceable chiplet, according to an embodiment. The method 3900 may be performed by hardware logic within the fabric interconnect and fabric interconnect nodes described herein. In one embodiment, the method 3900 includes a first fabric interface node receiving data from a first chiplet logic slot (block 3902). The first fabric interface node may encapsulate the data into a fabric packet (block 3904). The first fabric interface node may transmit the packet to a second fabric interface node via switching logic within the fabric interconnect (block 3906). The second fabric interface node may receive the packet (block 3908) and depacketize the data from the packet (3910). The second fabric interface node may then transmit the decapsulated data from the packet to a second chiplet logic slot (block 3912).
[0345] FIG. 40 illustrates a modular architecture for interchangeable chiplets, according to an embodiment. In one embodiment, a chiplet design 4030 can be made interchangeable by adapting the chiplet logic 4002 for interoperation with an interface template 4008. The interface template 4008 can include standardized logic, such as power control 2832 and clock control 2834 logic, interconnect buffers 2839, interconnect cache 2840, and fabric interconnect nodes 2842, as seen in the chiplet 2830 of FIG. 28B. An IP designer can then provide chiplet logic 4002 that is designed to interface with the interface template 4008. The specifications of the chiplet logic 4002 can vary and can include execution units, computational units, or streaming multiprocessors as described herein. The chiplet logic 4002 can also include media encoding and / or decoding logic, matrix acceleration logic, or ray tracing logic. For memory chiplets, chiplet logic 4002 can be replaced by memory cells and fabric interconnect nodes can be replaced by interconnect bridge I / O circuitry, such as that depicted in memory chiplet 2906 in FIG. 29B.
[0346] FIG. 41 depicts the use of a standardized chassis interface used in chiplet testing, validation, and integration. Chiplet 4130 can include a logic layer 4110 and an interface layer 4112, similar to chiplet 4030 of FIG. 40. Interface layer 4112 can be a standardized interface that can communicate with temporary interconnect 4114 that allows the chiplet to be removably coupled to a test harness 4116. Test harness 4116 can communicate with a test host 4118. Test harness 4116, under communication with test host 4118, can run a series of tests on individual chiplets 4130 during an initial testing or binning out process to identify defects in logic layer 4110 and determine performance or functionality bins of chiplets 4130. For example, logic layer 4110 can be tested to determine the number of defective and non-defective functional units and whether a threshold number of specific functional units (e.g., matrix accelerators, ray tracing cores, etc.) are valid. The logic layer 4110 may also be tested to determine whether the internal logic is capable of operating at the target frequency.
[0347] 42 depicts the use of individually binned chiplets to create various product tiers. A set of untested chiplets 4202 may be tested and binned into a set of performance bins 4204, mainstream bins 4206, and economic bins 4208 depending on whether the individual chiplets fit a particular performance or functionality tier. Performance bins 4204 may include chiplets that exceed the performance (e.g., stable frequency) of mainstream bin 4206, while economic bin 4208 may include chiplets that are useful but below mainstream bin 4206 performance.
[0348] Since chiplets may be arranged interchangeably during assembly, different product tiers may be assembled based on selected sets of chiplets. Tier 1 products 4212 may be assembled only from chiplets in performance bins 4204, while tier 2 products 4214 may include a selection of chiplets from performance bins 4204 and other chiplets from mainstream bins 4206. For example, tier 2 products 4214 designed for workloads requiring high bandwidth, low latency memory may use high performance memory chiplets from performance bins 4204 while using compute, graphics, or media chiplets from mainstream bins 4206. Furthermore, tier 3 products 4216 may be assembled using mainstream compute chiplets from mainstream bins 4206 and memory from economic bins 4208 when such products are tailored for workloads without high memory bandwidth requirements. Tier 4 products 4218 may be assembled from useful but lower performance chiplets in economic bins 4208.
[0349] FIG. 43 illustrates a method 4300 for enabling different product tiers based on chiplet configuration. The method 4300 includes packaging chiplet dies into test packaging (block 4302). The chiplets may then be tested to bin out the chiplets based on frequency and / or number of functional units (block 4304). A decomposed parallel processor may then be assembled using chiplets from one or more bins based on product requirements (block 4306). Additional chiplets (e.g., media, ray tracing, etc.) may also be added based on functional requirements (block 4308).
[0350] The following notes and / or examples relate to specific embodiments or examples thereof. Certain in the examples may be used anywhere in one or more embodiments. Various features of different embodiments or examples may be combined in various ways, with some features included and others excluded, to suit a variety of different applications. Examples may include methods according to the embodiments and examples described herein, means for performing the operations of the methods, at least one machine-readable medium including instructions that, when executed by a machine, cause the machine to perform the operations of the methods, or objects such as devices or systems. Various components may be means for performing the operations or functions described.
[0351] The embodiments described herein provide techniques to decompose the architecture of an SoC integrated circuit into multiple different chiplets that can be packaged on a common chassis. In one embodiment, a graphics processing unit or parallel processor is composed of multiple different silicon chiplets that are manufactured separately. A chiplet is an at least partially packaged integrated circuit that contains different logic units that can be assembled with other chiplets into a larger package. A multiple set of chiplets with different IP core logic can be assembled into a single device.
[0352] One embodiment provides a general purpose graphics processor having a base die including an interconnect fabric and one or more chiplets coupled to the base die and the interconnect fabric by an interconnect structure, the interconnect structure enabling electrical communication between the one or more chiplets and the interconnect fabric. The one or more chiplets may include a first chiplet and a second chiplet, the first chiplet coupled to the base die and connected to the interconnect fabric by a first interconnect structure, and the second chiplet coupled to the base die and connected to the interconnect fabric by a second interconnect structure. A chiplet may include functional units configured to perform general purpose graphics processing operations, media encoding or decoding operations, matrix operation acceleration, and / or ray tracing. In one embodiment, a chiplet includes a network processor and a physical network interface (e.g., a network port, a wireless radio, etc.). A chiplet may further include memory, which may be a cache memory or a DRAM. Each chiplet may be separately and independently power gated. Additionally, logic or memory may be included in the base die. In one embodiment, the base die includes a cache memory. The cache memory in the base die may be a processor-wide cache. The base die cache memory may be configured to cooperate with the cache memory in the chiplets.
[0353] One embodiment provides a data processing system having a general purpose graphics processor including a base die including an interconnect fabric and a plurality of chiplets coupled to the base die and the interconnect fabric by a plurality of interconnect structures, the plurality of interconnect structures enabling electrical communication between the plurality of chiplets and the interconnect fabric, the interconnect fabric receiving messages or signals from a first fabric interface node associated with a first chiplet in the plurality of chiplets and relaying the messages or signals to a second fabric interface node associated with a second chiplet in the plurality of chiplets. The interconnect fabric can transmit messages or signals through multiple virtual channels on multiple physical links of the interconnect fabric. In one embodiment, multiple virtual channels can be transmitted over a single physical link. In one embodiment, a single virtual channel can be transmitted over multiple physical links. Idle physical links can be separately power gated.
[0354] One embodiment provides a method comprising generating data at a first functional unit in a chiplet or base die of a processor, transmitting the data via a first fabric interface node to an interconnect fabric, conveying the data across multiple clock domains in the processor, receiving the data at a second fabric interface node, transmitting the data to a second functional unit in the chiplet or base die of the processor, and performing an operation at the second functional unit based on the received data. The method further comprises associating the data with a virtual channel of the interconnect fabric and forwarding or switching the data based on the virtual channel. In a further embodiment, the method includes branching the virtual channel at the first fabric interface node, conveying data of the virtual channel across the multiple clock domains using multiple physical links, and converging the virtual channel at the second fabric interface node. In yet another embodiment, the virtual channel is a first virtual channel, and the method further includes converging the first virtual channel with a second virtual channel at the first fabric interface node, carrying the first virtual channel and the second virtual channel across the multiple clock domains using a single physical link, and branching the first virtual channel and the second virtual channel at the second fabric interface node.
[0355] One embodiment is a non-transitory machine-readable medium storing firmware for a microcontroller in a processor having a decomposed architecture, the firmware including instructions that cause the microcontroller to perform operations including receiving a command buffer for a workload to be executed on the processor, determining a set of chiplets on the processor that include functional units to be used to execute the workload, power gating one or more chiplets that do not include functional units to be used to execute the word, and executing the workload with powered chiplets. Operations may further include determining whether a chiplet that includes functional units to be used to execute the workload is powered and powering the chiplet when the chiplet is powered.
[0356] Further embodiments provide a disassembled processor package that can be configured to accept interchangeable chiplets. Interchangeability is enabled by specifying a standard physical interconnect for the chiplets that can allow the chiplets to interface with a fabric or bridge interconnect. Chilets from different IP designers can fit into a common interconnect, allowing such chiplets to be interchangeable during assembly. The fabric and bridge interconnect logic on the chiplets can then be configured to match the actual interconnect layout of the chiplet's on-board logic. Furthermore, data from the chiplets can be transmitted across the inter-chiplet fabric using encapsulation such that the actual data being transferred is opaque to the fabric, further enabling interchangeability of individual chiplets. With such an interchangeable design, higher or lower density memories can be inserted into memory chiplet slots, while higher or lower core count compute or graphics chiplets can be inserted into logic chiplets.
[0357] One embodiment provides a general-purpose graphics processor including a base die including an interconnect fabric and one or more chiplets coupled to the base die and the interconnect fabric by an interconnect structure, the interconnect structure enabling electrical communication between the one or more chiplets and the interconnect fabric, the one or more chiplets being replaceable during assembly of the general-purpose graphics processor. The one or more chiplets include a memory chiplet having memory cells associated with a memory device. The memory chiplet is coupled to a first chiplet slot. The one or more chiplets may further include a first logic chiplet and a second logic chiplet. The first logic chiplet is coupled to the base die and connected to the interconnect fabric by a first interconnect structure. The first interconnect structure is affixed to a first logic chiplet slot. The second logic chiplet is coupled to the base die and connected to the interconnect fabric by a second interconnect structure. The second interconnect structure is secured to a second logic chiplet slot.
[0358] In one embodiment, the first logic chiplet slot is configured to receive the first logic chiplet or a third logic chiplet. The first logic chiplet includes functional units (e.g., execution units, compute units, streaming multiprocessors, etc.) configured to perform general-purpose graphics processing operations, while the third logic chiplet includes functional units configured to perform matrix acceleration operations, such as tensor cores. The second logic chiplet slot may be configured to receive the second logic chiplet or a fourth logic chiplet. The second logic chiplet includes functional units configured to perform media operations to encode, decode, or transcode media to or from one or more media encoding formats described herein. The fourth logic chiplet may alternatively include a network processor and a physical network interface.
[0359] In one embodiment, a logic chiplet as described herein includes a first layer including functional units and a second layer including fabric interconnect nodes. A memory chiplet as described herein can include a first layer including banks of memory cells and a second layer including I / O circuitry associated with interconnect bridges between the memory chiplets and the logic chiplets. In a further embodiment, the base die is a first base die that couples to a second base die via an interconnect bridge.
[0360] One embodiment provides a data processing system having a general purpose graphics processor having a base die including an interconnect fabric and a plurality of chiplets coupled to the base die and the interconnect fabric by a plurality of interconnect structures. The plurality of interconnect structures enable electrical communication between the plurality of chiplets and the interconnect fabric. The interconnect fabric can receive a fabric packet from a first fabric interface node associated with a first logic chiplet in the plurality of chiplets and relay the fabric packet to a second fabric interface node associated with a second logic chiplet in the plurality of chiplets. The interconnect fabric can transmit the fabric packet via a plurality of virtual channels on a plurality of physical links of the interconnect fabric. The interconnect fabric can transmit fabric packets associated with a single virtual channel across a plurality of physical links of the interconnect fabric. The interconnect fabric can also transmit fabric packets of a plurality of virtual channels across a single physical link of the interconnect fabric. A physical link of the plurality of physical links can be power gated when the physical link is idle. In one embodiment, fabric packets of a virtual channel may be associated with one or more traffic classes, and each virtual channel and traffic class may have an associated priority.
[0361] One embodiment provides a method comprising receiving data from a first logic chiplet slot at a first fabric interface node, encapsulating the data into a fabric packet at the first fabric interface node, transmitting the fabric packet through switching logic to a second fabric interface node, receiving the packet at the second fabric interface node, decapsulating data from the fabric packet at the second fabric interface node, and transmitting the data from the fabric packet from the second fabric interface node to a second logic chiplet slot, where the fabric packet may traverse multiple clock domains between the first fabric interface node and the second fabric interface node.
[0362] Of course, one or more parts of the embodiments may be implemented using different combinations of software, firmware, and / or hardware. Throughout this detailed description, for purposes of explanation, numerous specific details are set forth in order to provide a thorough understanding of the present invention. However, it will be apparent to one skilled in the art that the embodiments may be practiced without some of these specific details. In certain instances, well-known structures and functions have not been detailed in order not to obscure the inventive subject matter of the embodiments. Thus, the scope and spirit of the present invention should be judged in terms of the following claims.
Claims
1. a first base chiplet including a first interconnect fabric and a first plurality of L3 cache banks coupled to or integrated with the first interconnect fabric; a first logic chiplet including a cluster of compute units for performing parallel execution of compute or graphics shader instructions, stacked on the first base chiplet; a first interconnect structure coupling the cluster of computing units to the first interconnect fabric; a second base chiplet including a second interconnect fabric and a second plurality of L3 cache banks coupled to or integrated with the second interconnect fabric, the second base chiplet being coupled to the first base chiplet by a second interconnect structure; a second logic chiplet including a plurality of processor cores and stacked on the second base chiplet; a third interconnect structure coupling the second logic chiplet to the second interconnect fabric; and having the first logic chiplet is manufactured using a process technology different from that used to manufacture the first base chiplet and the second base chiplet. Device.
2. At least one of the first base chiplet and the second base chiplet is a fourth interconnect structure coupling the at least one of the first base chiplet and the second base chiplet to a memory.
2. The apparatus of claim 1.
3. the cluster of computing units and the plurality of processor cores access the memory via the first base chiplet and the second base chiplet, respectively.
3. The apparatus of claim 2.
4. The memory comprises a high bandwidth memory (HBM).
4. The apparatus of claim 3.
5. a packaging structure coupled to the first base chiplet and the second base chiplet; a package interconnect for electrically coupling the package structure to one or more devices; 5. The apparatus of claim 1 , further comprising:
6. each of the first base chiplet and the second base chiplet and the first logic chiplet and the second logic chiplet is in an independent clock domain and power domain and may be clock gated and power gated independently of other chiplets; 6. Apparatus according to any one of claims 1 to 5.
7. 1. A method of manufacturing a package assembly, comprising: The package assembly comprises: a first base chiplet including a first interconnect fabric and a first plurality of L3 cache banks coupled to or integrated with the first interconnect fabric; a first logic chiplet including a cluster of compute units for performing parallel execution of compute or graphics shader instructions, stacked on the first base chiplet; a first interconnect structure coupling the cluster of computing units to the first interconnect fabric; a second base chiplet including a second interconnect fabric and a second plurality of L3 cache banks coupled to or integrated with the second interconnect fabric, the second base chiplet being coupled to the first base chiplet by a second interconnect structure; a second logic chiplet including a plurality of processor cores and stacked on the second base chiplet; a third interconnect structure coupling the second logic chiplet to the second interconnect fabric; and The method of claim 1, further comprising: fabricating the first logic chiplet by using a process technology different from that used to fabricate the first base chiplet and the second base chiplet; fabricating the first base chiplet and the second base chiplet by using a process technology different from that used to fabricate the first logic chiplet; The method according to claim 1,
8. At least one of the first base chiplet and the second base chiplet is a fourth interconnect structure coupling the at least one of the first base chiplet and the second base chiplet to a memory. The method of claim 7.
9. the cluster of computing units and the plurality of processor cores access the memory via the first base chiplet and the second base chiplet, respectively. The method according to claim 8.
10. The memory comprises a high bandwidth memory (HBM).
10. The method of claim 9.
11. The package assembly comprises: a packaging structure coupled to the first base chiplet and the second base chiplet; a package interconnect for electrically coupling the package structure to one or more devices; Further comprising 11. The method according to any one of claims 7 to 10.
12. power and clock control for the first and second base chiplets and the first and second logic chiplets; each of the first base chiplet and the second base chiplet and the first logic chiplet and the second logic chiplet is in an independent clock domain and power domain and may be clock gated and power gated independently of other chiplets; 12. The method according to any one of claims 7 to 11.
Citation Information
Patent Citations
Chiplet display with multi-passive matrix controller
JP2013546012A
Architecture for on-die interconnect
JP2016085733A
Cobalt-based interconnects and methods of making them
JP2016541113A
Display pipe request aggregation
US20140071140A1
Three-dimensional computer processor systems having multiple local power and cooling layers and a global interconnection structure
US20140281378A1