Apparatus and method for manageable segmented acceleration structures
By adopting segmented acceleration structure and BVH compression technology in ray tracing, the problem of low ray tracing performance in the existing technology is solved, and more efficient visibility query and ray-scene cross-processing are achieved.
Patent Information
- Application Number
- CN202411591074.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2024-03-28
- Filing Date
- 2024-11-08
- Publication Date
- 2025-06-13
AI Technical Summary
The prior art is difficult to efficiently handle visibility queries and ray-scene crossovers in ray tracing, resulting in resource intensiveness and poor performance.
The parallel processing and data compression of ray tracing operations are optimized through multi-core groups and dedicated graphics processing resource collections.
Improve the real-time performance of ray tracing, reduce resource usage, and achieve more efficient visibility query and ray-scene cross-processing.
Smart Images

Figure CN120147100A_ABST
Abstract
Description
Technical Field
[0001] The present invention generally relates to the field of graphics processors. More particularly, the present invention relates to apparatuses and methods for a manageable segmented acceleration structure. Background Art
[0002] Ray tracing is a technique in which light transport is simulated through physically based rendering. Widely used in movie rendering, until a few years ago it was considered too resource-intensive for real-time performance. One of the key operations in ray tracing is processing visibility queries for ray-scene intersections (referred to as "ray traversal"), and ray-scene intersections are calculated by traversing and intersecting nodes in a bounding volume hierarchy (BVH).
[0003] Rasterization is a technique in which a screen object is created from a 3D model of an object created from a triangle mesh. The vertices of each triangle intersect with the vertices of other triangles of different shapes and sizes. Each vertex has a spatial position and information about color, texture, and its normal, and this information is used to determine the way the surface of the object faces. The rasterization unit converts the triangles of the 3D model into pixels in 2D screen space and can assign an initial color value to each pixel based on the vertex data. Brief Description of the Drawings
[0004] A better understanding of the present invention can be obtained from the following detailed description in conjunction with the following drawings, in which:
[0005] Figure 1 is a block diagram of a processing system according to an embodiment.
[0006] Figure 2A is a block diagram of an embodiment of a processor having one or more processor cores, an integrated memory controller, and an integrated graphics processor.
[0007] Figure 2B is a block diagram of the hardware logic of a graphics processor core block according to some embodiments described herein.
[0008] Figure 2C illustrates a graphics processing unit (GPU) including a collection of dedicated graphics processing resources arranged in a multi-core group.
[0009] Figure 2D is a block diagram of a general-purpose graphics processing unit (GPGPU) that can be configured as a graphics processor and / or a computing accelerator according to embodiments described herein.
[0010] Figure 3Ais a block diagram of a graphics processor, which can be a discrete graphics processing unit or a graphics processor integrated with multiple processing cores or other semiconductor devices such as, but not limited to, memory devices or network interfaces.
[0011] Figure 3B Illustrates a graphics processor with a tiled architecture according to embodiments described herein.
[0012] Figure 3C Illustrates a computing accelerator according to embodiments described herein.
[0013] Figure 4 is a block diagram of a graphics processing engine of a graphics processor according to some embodiments.
[0014] Figure 5A Illustrates a graphics core cluster according to an embodiment.
[0015] Figure 5B Illustrates a vector engine of a graphics core according to an embodiment.
[0016] Figure 5C Illustrates a matrix engine of a graphics core according to an embodiment.
[0017] Figure 6 Illustrates a tile of a multi-tile processor according to an embodiment.
[0018] Figure 7 is a block diagram illustrating a graphics processor instruction format according to some embodiments.
[0019] Figure 8 is a block diagram of another embodiment of a graphics processor.
[0020] Figure 9A is a block diagram illustrating a graphics processor command format that can be used to program a graphics processing pipeline according to some embodiments.
[0021] Figure 9B is a block diagram illustrating a graphics processor command sequence according to an embodiment.
[0022] Figure 10 Illustrates an exemplary graphics software architecture of a data processing system according to some embodiments.
[0023] Figure 11A is a block diagram illustrating an IP core development system that can be used to fabricate an integrated circuit to perform operations according to an embodiment.
[0024] Figure 11B Illustrates a cross-sectional side view of an integrated circuit package assembly according to some embodiments described herein.
[0025] Figure 11CIllustrated is a packaged assembly including a plurality of hardware logic die units connected to a substrate.
[0026] Figure 11D Illustrated is a packaged assembly including interchangeable dies according to an embodiment.
[0027] Figure 12 Is a block diagram of an exemplary system-on-chip integrated circuit that can be fabricated using one or more IP cores according to an embodiment.
[0028] Figure 13 Illustrated is an exemplary graphics processor of a system-on-chip integrated circuit that can be fabricated using one or more IP cores;
[0029] Figure 14 Illustrated is an additional exemplary graphics processor of a system-on-chip integrated circuit that can be fabricated using one or more IP cores;
[0030] Figure 15 Illustrated is a processing architecture including a ray tracing core and a tensor core;
[0031] Figure 16 Illustrated is an exemplary hybrid ray tracing device;
[0032] Figure 17 Illustrated is a stack for ray tracing operations;
[0033] Figure 18 Illustrated are additional details of the hybrid ray tracing device;
[0034] Figure 19 Illustrated is a bounding volume hierarchy;
[0035] Figure 20 Illustrated is a call stack and a traversal state storage device;
[0036] Figure 21 Illustrated is the operation flow of a programmable ray tracing pipeline;
[0037] Figures 22A - 22B Illustrated is how multiple dispatch cycles are required to execute certain shaders;
[0038] Figure 23 Illustrated is how a single dispatch cycle executes multiple shaders;
[0039] Figure 24 Illustrated is how a single dispatch cycle executes multiple shaders;
[0040] Figure 25 Illustrated is an architecture for executing ray tracing instructions;
[0041] Figure 26Illustrates a method for executing ray tracing instructions within a thread;
[0042] Figure 27 Illustrates an embodiment of an architecture for asynchronous ray tracing;
[0043] Figure 28A Illustrates a displacement function applied to a mesh;
[0044] Figure 28B Illustrates an embodiment of a compression circuit module for compressing a mesh or a meshlet;
[0045] Figure 29A Illustrates a displacement map on a base subdivision surface;
[0046] Figures 29B - 29C Illustrates a difference vector relative to a coarse base mesh;
[0047] Figure 30 Illustrates a method according to an embodiment of the present invention;
[0048] Figures 31 - 33 Illustrates a mesh including a plurality of interconnected vertices;
[0049] Figure 34 Illustrates an embodiment of a refiner for generating a mesh;
[0050] Figures 35 - 36 Illustrates an embodiment in which an enclosing volume is formed based on a mesh;
[0051] Figure 37 Illustrates an embodiment of a mesh sharing overlapping vertices;
[0052] Figure 38 Illustrates a mesh having shared edges between triangles;
[0053] Figure 39 Illustrates a ray tracing engine according to an embodiment;
[0054] Figure 40 Illustrates a BVH compressor according to an embodiment;
[0055] Figure 41A Illustrates an embodiment of a ray tracing architecture;
[0056] Figure 41B Illustrates an embodiment including meshlet compression;
[0057] Figure 42 Illustrates a plurality of threads, including synchronous threads, divergent spawned threads, regular spawned threads, and convergent spawned threads;
[0058] Figure 43 Illustrates an embodiment of a ray tracing architecture with an unbounded thread dispatcher;
[0059] Figure 44 Illustrates level-of-detail (LoD) selection within a BVH during traversal according to an embodiment of the present invention;
[0060] Figure 45 Illustrates a method for level-of-detail (LoD) selection within a BVH during traversal according to an embodiment of the present invention;
[0061] Figure 46 Illustrates a hierarchical arrangement of an acceleration structure (AS) according to some embodiments of the present invention;
[0062] Figure 47 Illustrates an example of a linked acceleration structure with an internal node referencing another AS; and
[0063] Figure 48 Illustrates another example of a linked acceleration structure, where an internal node references a sub-region containing a reference to another AS. Detailed Description
[0064] In the following description, for purposes of explanation, numerous specific details are set forth in order to provide a thorough understanding of embodiments of the invention described below. However, one of ordinary skill in the art will appreciate that embodiments of the invention may be practiced without some of these specific details. In other instances, well-known structures and devices are shown in block diagram form to avoid obscuring the underlying principles of embodiments of the invention. Exemplary Graphics Processor Architecture and Data Types
[0065] System Overview
[0066] Figure 1 is a block diagram of a processing system 100 according to an embodiment. The processing system 100 can be used in a single-processor desktop computer system, a multi-processor workstation system, or a server system having a large number of processors 102 or processor cores 107. In one embodiment, the processing system 100 is a processing platform incorporated within a system-on-chip (SoC) integrated circuit for use in mobile, handheld, or embedded devices (such as within an Internet of Things (IoT) device with wired or wireless connectivity to a local area network or a wide area network).
[0067] In one embodiment, the processing system 100 can include, be coupled to, or be integrated within: a server-based gaming platform; a game console, including a gaming and media console; a mobile gaming console, a handheld gaming console, or an online gaming console. In some embodiments, the processing system 100 is part of: a mobile phone, a smart phone, a tablet computing device, or a mobile Internet-connected device, such as a laptop computer with a low internal storage capacity. The processing system 100 can also include, be coupled to, or be integrated within: a wearable device, such as a smartwatch wearable device; smart glasses or clothing that are enhanced with augmented reality (AR) or virtual reality (VR) features to provide visual, audio, or tactile output to supplement a real-world visual, audio, or tactile experience or otherwise provide text, audio, graphics, video, holographic images or video, or tactile feedback; other augmented reality (AR) devices; or other virtual reality (VR) devices. In some embodiments, the processing system 100 includes a television or set-top box device, or is part of a television or set-top box device. In one embodiment, the processing system 100 can include, be coupled to, or be integrated within: an autonomous vehicle, such as a bus, a tractor-trailer, a car, a motorcycle, or an electric bicycle, an airplane, or a glider (or any combination thereof). The autonomous vehicle can use the processing system 100 to process the environment sensed around the vehicle.
[0068] In some embodiments, one or more processors 102 each include one or more processor cores 107 to process instructions that, when executed, perform operations for system or user software. In some embodiments, at least one of the one or more processor cores 107 is configured to process a particular instruction set 109. In some embodiments, the instruction set 109 can facilitate complex instruction set computing (CISC), reduced instruction set computing (RISC), or computing via very long instruction words (VLIW). The one or more processor cores 107 can process different instruction sets 109, which can include instructions to facilitate the emulation of other instruction sets. The processor cores 107 can also include other processing devices, such as a digital signal processor (DSP).
[0069] In some embodiments, the processor 102 includes a cache memory 104. Depending on the architecture, the processor 102 can have a single internal cache or multiple levels of internal caches. In some embodiments, the cache memory is shared among the various components of the processor 102. In some embodiments, the processor 102 also uses an external cache (e.g., a level 3 (L3) cache or a last-level cache (LLC)) (not shown), which can be shared among the processor cores 107 using known cache coherence techniques. A register file 106 can additionally be included in the processor 102 and can include different types of registers for storing different types of data (e.g., integer registers, floating-point registers, status registers, and instruction pointer registers). Some registers can be general-purpose registers, while other registers can be specific to the design of the processor 102.
[0070] In some embodiments, one or more processors 102 are coupled to one or more interface buses 110 to transfer communication signals, such as address, data, or control signals, between the processor 102 and other components in the processing system 100. The interface bus 110 can be a processor bus, such as a certain version of the direct media interface (DMI) bus, in one embodiment. However, the processor bus is not limited to the DMI bus and can include one or more peripheral component interconnect buses (e.g., PCI, PCI Express), a memory bus, or other types of interface buses. In one embodiment, the (one or more) processors 102 include a memory controller 116 and a platform controller hub 130. The memory controller 116 facilitates communication between the memory device and other components of the processing system 100, while the platform controller hub (PCH) 130 provides connections to I / O devices via a local I / O bus.
[0071] The memory device 120 can be a dynamic random access memory (DRAM) device, a static random access memory (SRAM) device, a flash memory device, a phase change memory device, or some other memory device having suitable performance to act as a process memory. In one embodiment, the memory device 120 can operate as a system memory for the processing system 100 to store data 122 and instructions 121 for use when one or more processors 102 execute an application or process. The memory controller 116 is also coupled to an optional external graphics processor 118, which can communicate with one or more of the graphics processors 108 in the processor 102 to perform graphics and media operations. In some embodiments, the graphics, media, and / or computing operations can be assisted by an accelerator 112, which is a coprocessor that can be configured to perform a specialized set of graphics, media, or computing operations. For example, in one embodiment, the accelerator 112 is a matrix multiplication accelerator used to optimize machine learning or computing operations. In one embodiment, the accelerator 112 is a ray tracing accelerator that can be used to perform ray tracing operations in cooperation with the graphics processor 108. In one embodiment, an external accelerator 119 can be used in place of or in cooperation with the accelerator 112.
[0072] In some embodiments, the display device 111 can be connected to the processor(s) 102. The display device 111 can be one or more of an internal display device such as in a mobile electronic device or a laptop device or an external display device attached via a display interface (e.g., DisplayPort, etc.). In one embodiment, the display device 111 can be a head-mounted display (HMD), such as a stereoscopic display device for use in virtual reality (VR) applications or augmented reality (AR) applications.
[0073] In some embodiments, the platform controller hub 130 enables peripherals to be connected to the memory device 120 and the processor 102 via a high-speed I / O bus. The I / O peripherals include, but are not limited to, an audio controller 146, a network controller 134, a firmware interface 128, a wireless transceiver 126, a touch sensor 125, and a data storage device 124 (e.g., non-volatile memory, volatile memory, hard disk drive, flash memory, NAND, 3D NAND, 3D XPoint, etc.). The data storage device 124 can be connected via a storage interface (e.g., SATA) or via a peripheral bus such as a Peripheral Component Interconnect bus (e.g., PCI, PCI Express). The touch sensor 125 can include a touch screen sensor, a pressure sensor, or a fingerprint sensor. The wireless transceiver 126 can be a Wi-Fi transceiver, a Bluetooth transceiver, or a mobile network transceiver such as a 3G, 4G, 5G, or Long Term Evolution (LTE) transceiver. The firmware interface 128 enables communication with the system firmware and can be, for example, a Unified Extensible Firmware Interface (UEFI). The network controller 134 can enable a network connection to a wired network. In some embodiments, a high-performance network controller (not shown) is coupled to the interface bus 110. The audio controller 146 is a multi-channel high-definition audio controller in one embodiment. In one embodiment, the processing system 100 includes an optional legacy I / O controller 140 for coupling legacy (e.g., Personal System 2 (PS / 2)) devices to the system. The platform controller hub 130 can also be connected to one or more Universal Serial Bus (USB) controllers 142 to connect to input devices such as a keyboard and mouse 143 combination, a camera 144, or other USB input devices.
[0074] It will be appreciated that the illustrated processing system 100 is exemplary and not restrictive, as other types of data processing systems configured in different ways can also be used. For example, instances of the memory controller 116 and the platform controller hub 130 can be integrated into a discrete external graphics processor such as the external graphics processor 118. In one embodiment, the platform controller hub 130 and / or the memory controller 116 can be external to one or more of the processors 102 and reside in a system chipset that communicates with the (one or more) processors 102.
[0075] For example, a circuit board (“sled”) can be used, on which components such as a CPU, memory, and other components are placed and are designed for increased thermal performance. In some embodiments, a processing component such as a processor is located on the top side of the sled, while near-memory such as DIMMs is located on the bottom side of the sled. As a result of the enhanced air flow provided by this design, the components can operate at higher frequencies and power levels than in a typical system, thereby increasing performance. In addition, the sled is configured to blindly mate with power and data communication cables in a rack, thus enhancing their ability to be quickly removed, upgraded, reinstalled, and / or replaced. Similarly, the individual components located on the sled (such as processors, accelerators, memory, and data storage drives) are configured to be easily upgraded due to their increased spacing from each other. In an illustrative embodiment, the components additionally include hardware attestation features to verify their authenticity.
[0076] A data center can utilize a single network architecture (“fabric”) that supports multiple other network architectures including Ethernet and Omni-Path. The sled can be coupled to a switch via optical fibers that provide higher bandwidth and lower latency than typical twisted pair cables (e.g., Category 5, Category 5e, Category 6, etc.). Due to the high-bandwidth, low-latency interconnects and network architecture, the data center can use physically disaggregated pool resources such as memory, accelerators (e.g., GPUs, graphics accelerators, FPGAs, ASICs, neural network, and / or artificial intelligence accelerators, etc.), and data storage drives, and provide them to computing resources (e.g., processors) on a demand basis such that the computing resources can access the pooled resources as if they were local.
[0077] A power supply or power source can provide voltage and / or current to the processing system 100 or any component or system described herein. In one example, the power supply includes an AC-to-DC adapter for plugging into a wall outlet. Such AC power can be a renewable energy (e.g., solar) power source. In one example, the power source includes a DC power source such as an external AC-to-DC converter. In one example, the power source or power supply includes wireless charging hardware for charging via a proximity charging field. In one example, the power source can include an internal battery, an AC power supply, a motion-based power supply, a solar power supply, or a fuel cell source.
[0078] Figures 2A - 2D Illustrated is a computing system and a graphics processor provided by the embodiments described herein. Those having the same reference numeral (or name) as an element in any other figure herein Figures 2A - 2DThe component can operate or function in any manner similar to the manner described elsewhere in this document, but is not limited thereto.
[0079] Figure 2A FIG. is a block diagram of an embodiment of a processor 200 having one or more processor cores 202A - 202N, an integrated memory controller 214, and an integrated graphics processor 208. The processor 200 can include additional cores, up to and including additional cores 202N represented by the dashed boxes. Each of the processor cores 202A - 202N includes one or more internal cache units 204A - 204N. In some embodiments, each processor core may also have access to one or more shared cache units 206. The internal cache units 204A - 204N and the shared cache units 206 represent the cache memory hierarchy within the processor 200. The cache memory hierarchy can include at least one level of instruction and data cache within each processor core, and one or more levels of shared mid - level cache, such as level 2 (L2), level 3 (L3), level 4 (L4), or other levels of cache, where the highest - level cache before the external memory is classified as the LLC. In some embodiments, cache coherence logic maintains coherence between the various cache units 206 and 204A - 204N.
[0080] In some embodiments, the processor 200 may also include a set of one or more bus controller units 216 and a system agent core 210. The one or more bus controller units 216 manage a set of peripheral buses, such as one or more PCI or Express PCI buses. The system agent core 210 provides management functions for the various processor components. In some embodiments, the system agent core 210 includes one or more integrated memory controllers 214 to manage access to various external memory devices (not shown).
[0081] In some embodiments, one or more of the processor cores 202A - 202N include support for simultaneous multithreading. In such embodiments, the system agent core 210 includes components for coordinating and operating the cores 202A - 202N during multithreaded processing. The system agent core 210 may additionally include a power control unit (PCU), which includes logic and components for regulating the power states of the processor cores 202A - 202N and the graphics processor 208.
[0082] In some embodiments, the processor 200 further includes a graphics processor 208 for performing graphics processing operations. In some embodiments, the graphics processor 208 is coupled to a set of shared cache units 206 and a system agent core 210 including one or more integrated memory controllers 214. In some embodiments, the system agent core 210 further includes a display controller 211 for driving the graphics processor output to one or more coupled displays. In some embodiments, the display controller 211 can also be a separate module coupled to the graphics processor via at least one interconnect, or can be integrated within the graphics processor 208.
[0083] In some embodiments, a ring-based interconnect 212 is used to couple the internal components of the processor 200. However, alternative interconnect units can be used, such as point-to-point interconnects, switched interconnects, mesh interconnects, or other techniques, including those well known in the art. In some embodiments, the graphics processor 208 is coupled to the ring-based interconnect 212 via an I / O link 213.
[0084] Exemplary I / O link 213 represents at least one of multiple types of I / O interconnects, including an on-package I / O interconnect that facilitates communication between various processor components and high-performance embedded memory modules 218 such as eDRAM modules or high-bandwidth memory (HBM) modules. In some embodiments, each of the processor cores 202A - 202N and the graphics processor 208 can use the embedded memory module 218 as a shared last-level cache.
[0085] In some embodiments, the processor cores 202A - 202N are homogeneous cores that execute the same instruction set architecture. In another embodiment, the processor cores 202A - 202N are heterogeneous in terms of instruction set architecture (ISA), where one or more of the processor cores 202A - 202N execute a first instruction set, while at least one of the other cores executes a subset of the first instruction set or a different instruction set. In one embodiment, the processor cores 202A - 202N are heterogeneous in terms of microarchitecture, where one or more cores with relatively high power consumption are coupled to one or more power-efficient cores with lower power consumption. In one embodiment, the processor cores 202A - 202N are heterogeneous in terms of computing capabilities. Additionally, the processor 200 can be implemented on one or more chips, or as a SoC integrated circuit having the illustrated components among other components.
[0086] Figure 2B is a block diagram of the hardware logic of a graphics processor core block 219 according to some embodiments described herein. In some embodiments, elements having the same reference numeral (or name) as elements in any other figure herein Figure 2BThe components can operate or function in a manner similar to that described elsewhere in this document. The graphics processor core block 219 is an example of a partition of a graphics processor. The graphics processor core block 219 can be included in Figure 2A the integrated graphics processor 208 or a discrete graphics processor, parallel processor, and / or computing accelerator. The graphics processor as described herein can include multiple graphics core blocks based on the target power and performance envelope. Each graphics processor core block 219 can include a functional block 230 coupled to multiple graphics cores 221A - 221F, which include modular blocks of fixed function logic and general - purpose programmable logic. The graphics processor core block 219 also includes a shared / cache memory 236 accessible by all of the graphics cores 221A - 221F, rasterizer logic 237, and additional fixed function logic 238.
[0087] In some embodiments, the functional block 230 includes a geometry / fixed function pipeline 231 that can be shared by all of the graphics cores in the graphics processor core block 219. In various embodiments, the geometry / fixed function pipeline 231 includes a 3D geometry pipeline, a video front - end unit, a thread spawner and a global thread dispatcher, and a unified return buffer manager that manages a unified return buffer. In one embodiment, the functional block 230 also includes a graphics SoC interface 232, a graphics microcontroller 233, and a media pipeline 234. The graphics SoC interface 232 provides an interface between the graphics processor core block 219 and other core blocks within the graphics processor or computing accelerator SoC. The graphics microcontroller 233 is a programmable sub - processor that can be configured to manage various functions of the graphics processor core block 219, including thread dispatching, scheduling, and pre - emption. The media pipeline 234 includes logic for facilitating the decoding, encoding, pre - processing, and / or post - processing of multimedia data including image and video data. The media pipeline 234 implements media operations via requests to the computation or sampling logic within the graphics cores 221A - 221F. One or more pixel back - ends 235 can also be included within the functional block 230. The pixel back - end 235 includes a cache memory for storing pixel color values and can perform blending operations and lossless color compression on the rendered pixel data.
[0088] In one embodiment, the graphics SoC interface 232 enables the graphics processor core block 219 to communicate with a general-purpose application processor core (e.g., a CPU) and / or other components within the SoC or a system host CPU coupled to the SoC via a peripheral interface. The graphics SoC interface 232 also enables communication with off-chip memory hierarchy elements such as a shared last-level cache memory, system RAM, and / or embedded on-chip or on-package DRAM. The graphics SoC interface 232 also enables communication with fixed-function devices within the SoC (such as a camera imaging pipeline), and enables the use and / or implementation of global memory atoms, which can be shared between the graphics processor core block 219 and the CPU within the SoC. The graphics SoC interface 232 is also capable of implementing power management control for the graphics processor core block 219, and enables an interface between the clock domain of the graphics processor core block 219 and other clock domains within the SoC. In one embodiment, the graphics SoC interface 232 enables receipt of command buffers from a command streamer and a global thread dispatcher, the command buffers being configured to provide commands and instructions to each of one or more graphics cores within the graphics processor. The commands and instructions can be dispatched to the media pipeline 234 when a media operation is to be performed, and to the geometry and fixed-function pipeline 231 when a graphics processing operation is to be performed. When a computing operation is to be performed, the compute dispatch logic can dispatch commands to the graphics cores 221A - 221F, bypassing the geometry and media pipelines.
[0089] The graphics microcontroller 233 can be configured to perform various scheduling and management tasks for the graphics processor core block 219. In one embodiment, the graphics microcontroller 233 can execute graphics and / or compute workload scheduling on various vector engines 222A-222F, 224A-224F and matrix engines 223A-223F, 225A-225F within the graphics cores 221A-221F. In this scheduling model, host software executing on the CPU core of the SoC including the graphics processor core block 219 can submit a workload to one of multiple graphics processor doorbells, which invokes a scheduling operation on the appropriate graphics engine. Scheduling operations include determining which workload to run next, submitting the workload to the command stream converter, preempting existing workloads running on the engine, monitoring the progress of the workload, and notifying the host software when the workload is complete. In one embodiment, the graphics microcontroller 233 is also capable of facilitating a low-power or idle state for the graphics processor core block 219, thereby providing the ability to save and restore registers within the graphics processor core block 219 across low-power state transitions independent of the operating system and / or graphics driver software on the system.
[0090] The graphics processor core block 219 can have up to N modular graphics cores more or less than the illustrated graphics cores 221A-221F. For each set of N graphics cores, the graphics processor core block 219 can also include a shared / cache memory 236, rasterizer logic 237, and additional fixed function logic 238 that can be configured to be shared memory or cache memory to accelerate various graphics and compute processing operations.
[0091] Within each graphics core 221A-221F is a set of execution resources that can be used to perform graphics, media, and compute operations in response to requests from the graphics pipeline, media pipeline, or shader program. The graphics cores 221A-221F include multiple vector engines 222A-222F, 224A-224F, matrix acceleration units 223A-223F, 225A-225D, cache / shared local memory (SLM), samplers 226A-226F, and ray tracing units 227A-227F.
[0092] The vector engines 222A - 222F, 224A - 224F are general - purpose graphics processing units capable of performing floating - point and integer / fixed - point logic operations in the service of graphics, media, or computing operations, including graphics, media, or computing / GPGPU programs. The vector engines 222A - 222F, 224A - 224F are capable of operating with variable vector widths using SIMD, SIMT, or SIMT + SIMD execution modes. The matrix acceleration units 223A - 223F, 225A - 225D include matrix - matrix and matrix - vector acceleration logic, which improves performance regarding matrix operations, especially low - precision and mixed - precision (e.g., INT8, FP16, BF16) matrix operations for machine learning. In one embodiment, each of the matrix acceleration units 223A - 223F, 225A - 225D includes a systolic array of one or more processing elements capable of performing concurrent matrix multiplication or dot - product operations on matrix elements.
[0093] The samplers 226A - 226F can read media or texture data into memory and can sample the data differently based on the configured sampler state and the texture / media format being read. Threads executing on the vector engines 222A - 222F, 224A - 224F or the matrix acceleration units 223A - 223F, 225A - 225D can utilize the cache / SLM 228A - 228F within each execution core. The cache / SLM 228A - 228F can be configured as a cache memory or as a shared memory pool local to each of the respective graphics cores 221A - 221F. The ray - tracing units 227A - 227F within the graphics cores 221A - 221F include ray - traversal / intersection circuit modules for performing ray traversal using a bounding - volume hierarchy (BVH) and identifying intersections between rays and primitives enclosed within the BVH volume. In one embodiment, the ray - tracing units 227A - 227F include circuit modules for performing depth testing and culling (e.g., using a depth buffer or similar arrangement). In one implementation, the ray - tracing units 227A - 227F perform traversal and intersection operations in concert with image denoising, at least a portion of which can be performed using the associated matrix acceleration units 223A - 223F, 225A - 225D.
[0094] Figure 2C A graphics processing unit (GPU) 239 is illustrated that includes a collection of dedicated graphics - processing resources arranged in multi - core groups 240A - 240N. Details of the multi - core group 240A are illustrated. The multi - core groups 240B - 240N can be equipped with the same or similar collections of graphics - processing resources.
[0095] As illustrated, the multi-core group 240A may include a set of graphics cores 243, a set of tensor cores 244, and a set of ray tracing cores 245. The scheduler / dispatcher 241 schedules and dispatches graphics threads for execution on the respective cores 243, 244, 245. In one embodiment, the tensor core 244 is a sparse tensor core having hardware that enables bypassing multiplication operations with zero-valued inputs. Figure 2C The graphics cores 243 of the GPU 239 are different from Figure 2B the graphics cores 221A - 221F at a hierarchical level of abstraction, Figure 2B and the graphics cores 221A - 221F are similar to Figure 2C the multi-core groups 240A - 240N. Figure 2C The graphics cores 243, tensor cores 244, and ray tracing cores 245 are respectively similar to Figure 2B the vector engines 222A - 222F, 224A - 224F, matrix engines 223A - 223F, 225A - 225F, and ray tracing units 227A - 227F.
[0096] The set of register files 242 is capable of storing operand values used by the cores 243, 244, 245 when executing graphics threads. These registers may include, for example, integer registers for storing integer values, floating-point registers for storing floating-point values, vector registers for storing packed data elements (integer and / or floating-point data elements), and tile registers for storing tensor / matrix values. In one embodiment, the tile registers are implemented as a combined set of vector registers.
[0097] One or more combined level 1 (L1) caches and shared memory units 247 locally store graphics data, such as texture data, vertex data, pixel data, ray data, bounding volume data, etc., within each multi-core group 240A. One or more texture units 247 are also capable of performing texture operations, such as texture mapping and sampling. The level 2 (L2) cache 253, shared by all or a subset of the multi-core groups 240A - 240N, stores graphics data and / or instructions for multiple concurrent graphics threads. As illustrated, the L2 cache 253 may be shared across multiple multi-core groups 240A - 240N. One or more memory controllers 248 couple the GPU 239 to the memory 249, which may be system memory (e.g., DRAM) and / or dedicated graphics memory (e.g., GDDR6 memory).
[0098] The input / output (I / O) circuit module 250 couples the GPU 239 to one or more I / O devices 252, such as a digital signal processor (DSP), a network controller, or a user input device. An on-chip interconnect may be used to couple the I / O device 252 to the GPU 239 and the memory 249. One or more I / O memory management units (IOMMUs) 251 of the I / O circuit module 250 couple the I / O device 252 directly to the memory 249. In one embodiment, the IOMMU 251 manages multiple sets of page tables to map virtual addresses to physical addresses in the memory 249. In this embodiment, the I / O device 252, the (one or more) CPUs 246, and the GPU 239 may share the same virtual address space.
[0099] In one implementation, the IOMMU 251 supports virtualization. In this case, it may manage a first set of page tables to map guest / graphics virtual addresses to guest / graphics physical addresses, and manage a second set of page tables to map guest / graphics physical addresses to system / host physical addresses (e.g., within the memory 249). The base addresses of each of the first and second sets of page tables may be stored in control registers and swapped out during a context switch (e.g., to provide access to the relevant set of page tables for the new context). Although not illustrated in Figure 2C , each of the cores 243, 244, 245, and / or the multi-core groups 240A - 240N may include a translation lookaside buffer (TLB) to cache guest virtual to guest physical translations, guest physical to host physical translations, and guest virtual to host physical translations.
[0100] In one embodiment, the CPU 246, the GPU 239, and the I / O device 252 are integrated on a single semiconductor chip and / or chip package. The memory 249 may be integrated on the same chip or may be coupled to the memory controller 248 via an off-chip interface. In one implementation, the memory 249 includes GDDR6 memory, which shares the same virtual address space as other physical system-level memories, although the underlying principles of the embodiments described herein are not limited to this particular implementation.
[0101] In one embodiment, the tensor core 244 includes a plurality of functional units specifically designed to perform matrix operations, which are the basic computational operations used to perform deep learning operations. For example, simultaneous matrix multiplication operations can be used for neural network training and inference. The tensor core 244 can perform matrix processing using various operand precisions, including single-precision floating point (e.g., 32 bits), half-precision floating point (e.g., 16 bits), integer (16 bits), byte (8 bits), and nibble (4 bits). In one embodiment, the neural network implementation obtains the features of each rendered scene, potentially combining details from multiple frames to construct a high-quality final image.
[0102] In a deep learning implementation, parallel matrix multiplication work can be scheduled for execution on the tensor core 244. Training of neural networks particularly requires a large number of matrix dot product operations. To handle the inner product formula for N×N×N matrix multiplication, the tensor core 244 can include at least N dot product processing elements. Before matrix multiplication begins, a complete matrix is loaded into the tile register, and in each of the N cycles, at least one column of the second matrix is loaded. In each cycle, there are N dot products being processed.
[0103] Depending on the specific implementation, matrix elements can be stored in different precisions, including 16-bit words, 8-bit bytes (e.g., INT8), and 4-bit nibbles (e.g., INT4). Different precision modes can be specified for the tensor core 244 to ensure that the most efficient precision is used for different workloads (e.g., inference workloads that can tolerate quantization to bytes and nibbles).
[0104] In one embodiment, the ray tracing core 245 accelerates ray tracing operations for both real-time and non-real-time ray tracing implementations. Specifically, the ray tracing core 245 includes a ray traversal / intersection circuit module for performing ray traversal using a bounding volume hierarchy (BVH) and identifying intersections between the primitives enclosed within the BVH volume and the ray. The ray tracing core 245 can also include a circuit module for performing depth testing and culling (e.g., using a Z-buffer or similar arrangement). In one implementation, the ray tracing core 245 cooperates with the image denoising techniques described herein to perform traversal and intersection operations, at least a portion of which can be executed on the tensor core 244. For example, in one embodiment, the tensor core 244 implements a deep learning neural network to perform denoising of the frames generated by the ray tracing core 245. However, the (one or more) CPUs 246, the graphics core 243, and / or the ray tracing core 245 can also implement all or a portion of the denoising and / or deep learning algorithms.
[0105] Additionally, as described above, a distributed approach for denoising can be employed where the GPU 239 is in a computing device coupled to other computing devices via a network or high-speed interconnect. In this embodiment, the interconnected computing devices share neural network learning / training data to improve the speed at which the overall system learns to perform denoising for different types of image frames and / or different graphics applications.
[0106] In one embodiment, the ray tracing core 245 handles all BVH traversals and ray-primitive intersections, thereby freeing the graphics core 243 from being overloaded with thousands of instructions per ray. In one embodiment, each ray tracing core 245 includes a first set of dedicated circuit modules for performing bounding box tests (e.g., for traversal operations) and a second set of dedicated circuit modules for performing ray-triangle intersection tests (e.g., intersecting the traversed rays). Thus, in one embodiment, the multi-core group 240A is able to simply launch a ray probe, and the ray tracing core 245 independently performs ray traversal and intersection and returns hit data (e.g., hit, miss, multiple hits, etc.) to the thread context. While the ray tracing core 245 performs traversal and intersection operations, the other cores 243, 244 are freed up to perform other graphics or computing work.
[0107] In one embodiment, each ray tracing core 245 includes a traversal unit for performing BVH test operations and an intersection unit for performing ray-primitive intersection tests. The intersection unit generates "hit", "miss", or "multiple hits" responses, and the intersection unit provides the responses to the appropriate thread. During traversal and intersection operations, the execution resources of other cores (e.g., the graphics core 243 and the tensor core 244) are freed up to perform other forms of graphics work.
[0108] In one particular embodiment described below, a hybrid rasterization / ray tracing method is used where the work is distributed between the graphics core 243 and the ray tracing core 245.
[0109] In one embodiment, the ray tracing core 245 (and / or other cores 243, 244) includes hardware support for a ray tracing instruction set such as Microsoft's DirectX Ray Tracing (DXR) as well as ray generation, closest hit, any hit, and miss shaders (which enable the assignment of a unique set of textures and shaders to each object), where the DXR includes a DispatchRays command. Another ray tracing platform that can be supported by the ray tracing core 245, the graphics core 243, and the tensor core 244 is Vulkan 1.1.85. However, note that the underlying principles of the embodiments described herein are not limited to any particular ray tracing ISA.
[0110] Generally, various cores 245, 244, 243 may support a ray tracing instruction set, and the ray tracing instruction set includes instructions / functions for ray generation, closest hit, any hit, ray-primitive intersection, per-primitive and hierarchical bounding box construction, miss, traversal, and exception. More specifically, one embodiment includes ray tracing instructions for performing the following functions:
[0111] Ray generation – Ray generation instructions can be executed for each pixel, sample, or other user-defined work assignment.
[0112] Closest hit – The closest hit instruction can be executed to locate the closest intersection of a ray with a primitive in the scene.
[0113] Any hit – The any hit instruction identifies multiple intersections between primitives in the scene and a ray, potentially identifying a new closest intersection.
[0114] Intersection – The intersection instruction performs a ray-primitive intersection test and outputs the result.
[0115] Per-primitive bounding box construction – This instruction constructs a bounding box around a given primitive or group of primitives (e.g., when building a new BVH or other acceleration data structure).
[0116] Miss – Indicates that the ray misses all geometry in the scene or a specified region of the scene.
[0117] Traversal – Indicates the child volume that the ray will traverse.
[0118] Exception – Includes various types of exception handlers (e.g., invoked for various error conditions).
[0119] In one embodiment, the ray tracing core 245 may be adapted to accelerate general computing operations, and the general computing operations may be accelerated using computational techniques similar to ray intersection tests. A computational framework may be provided that enables shader programs to be compiled into low-level instructions and / or primitives for performing general computing operations via the ray tracing core. Exemplary computational problems that may benefit from computational operations performed on the ray tracing core 245 include computations involving the propagation of beams, waves, rays, or particles within a coordinate space. Interactions associated with the propagation may be computed relative to geometries or meshes within the coordinate space. For example, computations associated with the propagation of electromagnetic signals through an environment may be accelerated via the use of instructions or primitives executed via the ray tracing core. The diffraction and reflection of signals by objects in the environment may be computed as direct ray tracing analogies.
[0120] The ray tracing core 245 can also be used to perform calculations that are not directly analogous to ray tracing. For example, the ray tracing core 245 can be used to accelerate mesh projection, mesh refinement, and volume sampling calculations. General coordinate space calculations, such as nearest neighbor calculations, can also be performed. For example, a set of points near a given point can be discovered by defining a bounding box in the coordinate space around that point. The BVH and ray probing logic within the ray tracing core 245 can then be used to determine a set of point intersections within the bounding box. The intersections constitute the origin and the nearest neighbor to that origin. The calculations performed using the ray tracing core 245 can be executed in parallel with the calculations performed on the graphics core 243 and the tensor core 244. The shader compiler can be configured to compile compute shaders or other general-purpose graphics programs into low-level primitives that can be parallelized across the graphics core 243, the tensor core 244, and the ray tracing core 245.
[0121] Figure 2D is a block diagram of a general-purpose graphics processing unit (GPGPU) 270 that can be configured as a graphics processor and / or a compute accelerator according to embodiments described herein. The GPGPU 270 can be interconnected with a host processor (e.g., one or more CPUs 246) and memories 271, 272 via one or more system and / or memory buses. In one embodiment, the memory 271 is a system memory that can be shared with one or more CPUs 246, while the memory 272 is a device memory dedicated to the GPGPU 270. In one embodiment, the memory 272 and the components within the GPGPU 270 can be mapped to memory addresses accessible by one or more CPUs 246. Access to the memories 271 and 272 can be facilitated via a memory controller 268. In one embodiment, the memory controller 268 includes an internal direct memory access (DMA) controller 269, or can include logic to perform operations that would otherwise be performed by a DMA controller.
[0122] The GPGPU 270 includes multiple caches, including an L2 cache 253, an L1 cache 254, an instruction cache 255, and a shared memory 256, at least a portion of which can also be partitioned as a cache. The GPGPU 270 also includes multiple compute units 260A - 260N, which represent graphics cores 221A - 221F similar to Figure 2B and Figure 2CThe hierarchical abstraction level of the multi-core groups 240A - 240N. Each computing unit 260A - 260N includes a set of vector registers 261, scalar registers 262, vector logic units 263, and scalar logic units 264. The computing units 260A - 260N can also include local shared memory 265 and program counters 266. The computing units 260A - 260N can be coupled to a constant cache 267, which can be used to store constant data, i.e., data that will not change during the execution of a kernel or shader program on the GPGPU 270. In one embodiment, the constant cache 267 is a scalar data cache, and the cached data can be directly fetched into the scalar registers 262.
[0123] During operation, one or more CPUs 246 can write commands into registers or memory in the GPGPU 270 that have been mapped to an accessible address space. The command processor 257 can read commands from the registers or memory and determine how those commands will be processed within the GPGPU 270. The thread dispatcher 258 can then be used to dispatch threads to the computing units 260A - 260N to execute those commands. Each computing unit 260A - 260N can execute threads independently of other computing units. Additionally, each computing unit 260A - 260N can be independently configured for conditional computing and can conditionally output the results of the computation to memory. When the submitted commands are completed, the command processor 257 can interrupt one or more CPUs 246.
[0124] Figures 3A - 3C A block diagram illustrating additional graphics processor and computing accelerator architectures provided by the embodiments described herein. Elements having the same reference numeral (or name) as elements in any other figure herein Figures 3A - 3C can operate or function in any manner similar to the ways described elsewhere herein, but are not limited thereto.
[0125] Figure 3A is a block diagram of a graphics processor 300, which can be a discrete graphics processing unit, or can be a graphics processor integrated with multiple processing cores, or other semiconductor devices such as, but not limited to, memory devices or network interfaces. In some embodiments, the graphics processor communicates via a memory - mapped I / O interface to registers on the graphics processor and uses commands placed in the processor memory. In some embodiments, the graphics processor 300 includes a memory interface 314 for accessing memory. The memory interface 314 can be an interface to local memory, one or more internal caches, one or more shared external caches, and / or to system memory.
[0126] In some embodiments, the graphics processor 300 further includes a display controller 302 for driving display output data to a display device 318. The display controller 302 includes hardware for one or more overlay planes for displaying and combining multiple layers of video or user interface elements. The display device 318 can be an internal or external display device. In one embodiment, the display device 318 is a head-mounted display device, such as a virtual reality (VR) display device or an augmented reality (AR) display device. In some embodiments, the graphics processor 300 includes a video codec engine 306 for encoding media into one or more media encoding formats, decoding media from one or more media encoding formats, or transcoding media between one or more media encoding formats, the media encoding formats including but not limited to Moving Picture Experts Group (MPEG) formats (such as MPEG-2), Advanced Video Coding (AVC) formats (such as H.264 / MPEG-4 AVC), H.265 / HEVC, Alliance for Open Media (AOMedia) VP8, VP9, and Society of Motion Picture and Television Engineers (SMPTE) 421M / VC-1 and Joint Photographic Experts Group (JPEG) formats (such as JPEG and Motion JPEG (MJPEG) formats).
[0127] In some embodiments, the graphics processor 300 includes a block image transfer (BLIT) engine for performing two-dimensional (2D) rasterizer operations (including, for example, bit boundary block transfer). However, in one embodiment, one or more components of the graphics processing engine (GPE) 310 are used to perform 2D graphics operations. In some embodiments, the GPE 310 is a computing engine for performing graphics operations including three-dimensional (3D) graphics operations and media operations.
[0128] In some embodiments, the GPE 310 includes a 3D pipeline 312 for performing 3D operations, such as rendering three-dimensional images and scenes using processing functions for 3D primitive shapes (e.g., rectangles, triangles, etc.). The 3D pipeline 312 includes programmable and fixed-function elements that perform various tasks within the element and / or spawn execution threads to the 3D / media subsystem 315. Although the 3D pipeline 312 can be used to perform media operations, embodiments of the GPE 310 also include a media pipeline 316 specifically for performing media operations (such as video post-processing and image enhancement).
[0129] In some embodiments, the media pipeline 316 includes fixed function or programmable logic units to perform one or more specialized media operations, such as video decode acceleration, video deinterlacing, and video encode acceleration, in place of or on behalf of the video codec engine 306. In some embodiments, the media pipeline 316 further includes a thread spawning unit to spawn threads for execution on the 3D / media subsystem 315. The spawned threads execute computations for media operations on one or more graphics cores included in the 3D / media subsystem 315.
[0130] In some embodiments, the 3D / media subsystem 315 includes logic for executing threads spawned by the 3D pipeline 312 and the media pipeline 316. In one embodiment, the pipeline sends thread execution requests to the 3D / media subsystem 315, which includes thread dispatch logic for arbitrating and dispatching various requests to available thread execution resources. The execution resources include an array of graphics cores for processing 3D and media threads. In some embodiments, the 3D / media subsystem 315 includes one or more internal caches for thread instructions and data. In some embodiments, the subsystem further includes a shared memory that includes registers and addressable memory for sharing data between threads and storing output data.
[0131] Figure 3B Illustrated is a graphics processor 320 having a tiled architecture in accordance with embodiments described herein. In one embodiment, the graphics processor 320 includes a graphics processing engine cluster 322 having multiple instances of the graphics processing engine 310 within the graphics engine tiles 310A - 310D. Figure 3A Each graphics engine tile 310A - 310D can be interconnected via a set of tile interconnects 323A - 323F. Each graphics engine tile 310A - 310D can also be connected to a memory module or memory device 326A - 326D via a memory interconnect 325A - 325D. The memory devices 326A - 326D can use any graphics memory technology. For example, the memory devices 326A - 326D can be graphics double data rate (GDDR) memories. The memory devices 326A - 326D are in one embodiment HBM modules that can be on - die with their respective graphics engine tiles 310A - 310D. In one embodiment, the memory devices 326A - 326D are stacked memory devices that can be stacked on top of their respective graphics engine tiles 310A - 310D. In one embodiment, as Figures 11B - 11DAs described in further detail below, each graphics engine die 310A - 310D and associated memory 326A - 326D reside on separate chiplets that are bonded to a base die or base substrate.
[0132] The graphics processing unit 320 may be configured with a non - uniform memory access (NUMA) system where the memory devices 326A - 326D are coupled with the associated graphics engine die 310A - 310D. A given memory device may be accessed by a graphics engine die other than the die to which it is directly connected. However, the access latency to the memory devices 326A - 326D may be lowest when accessing the local die. In one embodiment, a cache - coherent NUMA (ccNUMA) system is enabled that uses the die interconnects 323A - 323F to enable communication between cache controllers within the graphics engine die 310A - 310D to maintain a consistent memory image when more than one cache stores the same memory location.
[0133] The graphics processing engine cluster 322 may be connected to an on - die or on - package fabric interconnect 324. In one embodiment, the fabric interconnect 324 includes a network processor, a network - on - chip (NoC), or another switching processor to enable the fabric interconnect 324 to act as a packet - switched fabric interconnect for exchanging data packets between components of the graphics processing unit 320. The fabric interconnect 324 may enable communication between the graphics engine die 310A - 310D and components such as the video codec engine 306 and one or more copy engines 304. The copy engines 304 may be used to move data out of, into the memory devices 326A - 326D and memory external to the graphics processing unit 320 (e.g., system memory) and move between them. The fabric interconnect 324 may also be coupled to one or more of the die interconnects 323A - 323F to facilitate or enhance the interconnect between the graphics engine die 310A - 310D. The fabric interconnect 324 may also be configured to (e.g., via the host interface 328) interconnect multiple instances of the graphics processing unit 320, thereby enabling die - to - die communication between the graphics engine die 310A - 310D of multiple GPUs. In one embodiment, the graphics engine die 310A - 310D of multiple GPUs may be presented to the host system as a single logical device.
[0134] The graphics processing unit 320 may optionally include a display controller 302 to enable connection to a display device 318. The graphics processing unit may also be configured as a graphics or computing accelerator. In the accelerator configuration, the display controller 302 and the display device 318 may be omitted.
[0135] The graphics processor 320 may be connected to a host system via a host interface 328. The host interface 328 may enable communication between the graphics processor 320, system memory, and / or other system components. The host interface 328 may be, for example, a fast PCI bus or another type of host system interface. For example, the host interface 328 may be an NVLink or NVSwitch interface. The host interface 328 and the fabric interconnect 324 may cooperate to enable multiple instances of the graphics processor 320 to act as a single logical device. The cooperation between the host interface 328 and the fabric interconnect 324 may also cause the respective graphics engine tiles 310A - 310D to appear as different logical graphics devices to the host system.
[0136] Figure 3C Illustrated is a computing accelerator 330 according to an embodiment described herein. The computing accelerator 330 can include an architecture similarity to the Figure 3B graphics processor 320 and is optimized for computing acceleration. The compute engine cluster 332 can include a set of compute engine tiles 340A - 340D that include execution logic optimized for the execution of parallel or vector-based general computing operations. In some embodiments, the compute engine tiles 340A - 340D do not include fixed-function graphics processing logic, although in one embodiment, one or more of the compute engine tiles 340A - 340D can include logic for performing media acceleration. The compute engine tiles 340A - 340D can be connected to memories 326A - 326D via memory interconnects 325A - 325D. The memories 326A - 326D and the memory interconnects 325A - 325D can be of a similar technology as in the graphics processor 320 or can be different technologies. The compute engine tiles 340A - 340D can also be interconnected via a set of tile interconnects 323A - 323F and can be connected to and / or interconnected through the fabric interconnect 324. Communication across tiles can be facilitated via the fabric interconnect 324. The fabric interconnect 324 (e.g., via the host interface 328) can also facilitate communication between the compute engine tiles 340A - 340D of multiple instances of the computing accelerator 330. In one embodiment, the computing accelerator 330 includes a large L3 cache 336 that can be configured as a device-wide cache. The computing accelerator 330 can also be connected to a host processor and memory via the host interface 328 in a manner similar to the Figure 3B graphics processor 320.
[0137] The compute accelerator 330 may also include an integrated network interface 342. In one embodiment, the network interface 342 includes network processors and controller logic that enable the compute engine cluster 332 to communicate via the physical layer interconnect 344 without data traversing the memory of the host system. In one embodiment, one of the compute engine tiles 340A - 340D is replaced with network processor logic, and data to be transmitted or received via the physical layer interconnect 344 can be directly transmitted to or from the memories 326A - 326D. Multiple instances of the compute accelerator 330 can be joined into a single logical device via the physical layer interconnect 344. Alternatively, the various compute engine tiles 340A - 340D can be presented as different network - accessible compute accelerator devices.
[0138] Graphics processing engine
[0139] Figure 4 is a block diagram of a graphics processing engine 410 of a graphics processor according to some embodiments. In one embodiment, the graphics processing engine (GPE) 410 is Figure 3A a certain version of the GPE 310 shown in and may also represent Figure 3B the graphics engine tiles 310A - 310D. Elements with the same reference numeral (or name) as elements in any other figure in this document Figure 4 are capable of operating or functioning in any manner similar to the ways described elsewhere in this document, but are not limited to such. For example, the Figure 3A 3D pipeline 312 and media pipeline 316 are illustrated. The media pipeline 316 is optional in some embodiments of the GPE 410 and may not be explicitly included within the GPE 410. For example and in at least one embodiment, a separate media and / or image processor is coupled to the GPE 410.
[0140] In some embodiments, GPE 410 is coupled to, or includes, command stream converter 403 that provides a command stream to 3D pipeline 312 and / or media pipeline 316. Alternatively or additionally, command stream converter 403 may be directly coupled to unified return buffer 418. Unified return buffer 418 may be communicatively coupled to graphics core cluster 414. In some embodiments, command stream converter 403 is coupled to a memory, which can be system memory, or one or more of an internal cache memory and a shared cache memory. In some embodiments, command stream converter 403 receives commands from the memory and sends the commands to 3D pipeline 312 and / or media pipeline 316. The commands are indications fetched from a ring buffer that stores commands for 3D pipeline 312 and media pipeline 316. In one embodiment, the ring buffer can additionally include a batch command buffer that stores a batch of multiple commands. Commands for 3D pipeline 312 can also include references to data stored in the memory, such as, but not limited to, vertex and geometry data for 3D pipeline 312 and / or image data and memory objects for media pipeline 316. 3D pipeline 312 and media pipeline 316 process commands and data by performing operations via logic within the respective pipelines or by dispatching one or more execution threads to graphics core cluster 414. In one embodiment, graphics core cluster 414 includes one or more blocks of graphics cores (e.g., graphics core blocks 415A, 415B), each block including one or more graphics cores. Each graphics core includes: a set of graphics execution resources that includes general-purpose and graphics-specific execution logic for performing graphics and computing operations; and fixed-function texture processing and / or machine learning and artificial intelligence acceleration logic, such as matrix or AI acceleration logic.
[0141] In various embodiments, 3D pipeline 312 can include fixed-function and programmable logic for processing one or more shader programs (such as vertex shaders, geometry shaders, pixel shaders, fragment shaders, compute shaders, or other shaders and / or GPGPU programs) by processing instructions and dispatching execution threads to graphics core cluster 414. Graphics core cluster 414 provides a unified block of execution resources for use in processing these shader programs. The multi-purpose execution logic within graphics core blocks 415A - 415B of graphics core cluster 414 includes support for various 3D API shader languages and can execute multiple simultaneous execution threads associated with multiple shaders.
[0142] In some embodiments, the graphics core cluster 414 includes execution logic for performing media functions such as video and / or image processing. In one embodiment, the graphics core includes general-purpose logic that is programmable to perform parallel general-purpose computing operations in addition to graphics processing operations. The general-purpose logic is capable of performing processing operations in parallel with or in conjunction with the (one or more) processor cores 107 in Figure 1 or the general-purpose logic within the core 202A - 202N in Figure 2A .
[0143] Output data generated by threads executing on the graphics core cluster 414 can output data to a memory in the unified return buffer (URB) 418. The URB 418 is capable of storing data for multiple threads. In some embodiments, the URB 418 can be used to send data between different threads executing on the graphics core cluster 414. In some embodiments, the URB 418 can additionally be used for synchronization between threads on the graphics core array and fixed-function logic within the shared function logic 420.
[0144] In some embodiments, the graphics core cluster 414 is scalable such that the cluster includes a variable number of graphics cores, each having a variable number of graphics cores based on the target power and performance levels of the GPE 410. In one embodiment, the execution resources are dynamically scalable such that execution resources can be enabled or disabled as needed.
[0145] The graphics core cluster 414 is coupled to the shared function logic 420, which includes a plurality of resources shared among the graphics cores in the graphics core array. The shared functions within the shared function logic 420 are hardware logic units that provide dedicated supplementary functions to the graphics core cluster 414. In various embodiments, the shared function logic 420 can include, but is not limited to, sampler 421, math 422, and inter-thread communication (ITC) 423 logic. Additionally, some embodiments implement one or more caches 425 within the shared function logic 420. The shared function logic 420 can implement functions that are the same as or similar to the Figure 2B additional fixed-function logic 238.
[0146] Implement shared functionality at least in cases where the demand for a given dedicated function is not sufficient to be included within the graphics core cluster 414. A single instantiation of the dedicated function is instead implemented as a stand-alone entity in the shared functionality logic 420 and shared among the execution resources within the graphics core cluster 414. The exact set of functions that are shared among the graphics core clusters 414 and included within the graphics core clusters 414 varies across embodiments. In some embodiments, certain shared functions within the shared functionality logic 420 that are widely used by the graphics core cluster 414 may be included within the shared functionality logic 416 within the graphics core cluster 414. In various embodiments, the shared functionality logic 416 within the graphics core cluster 414 can include some or all of the logic within the shared functionality logic 420. In one embodiment, all of the logic elements within the shared functionality logic 420 may be replicated within the shared functionality logic 416 of the graphics core cluster 414. In one embodiment, the shared functionality logic 420 is excluded in favor of the shared functionality logic 416 within the graphics core cluster 414.
[0147] Graphics processing resources
[0148] Figures 5A - 5C Illustrates execution logic including an array of processing elements employed in a graphics processor according to embodiments described herein. Figure 5A Illustrates a graphics core cluster according to an embodiment. Figure 5B Illustrates a vector engine of a graphics core according to an embodiment. Figure 5C Illustrates a matrix engine of a graphics core according to an embodiment. Figures 5A - 5C Elements having the same reference numerals as elements in any other figure herein may operate or function in any manner similar to that described elsewhere herein, but are not limited thereto. For example, Figures 5A - 5C the elements may be considered in the context of Figure 2B the graphics processor core block 219 and / or Figure 4 the graphics core blocks 415A - 415B. In one embodiment, Figures 5A - 5C the elements have functions similar to equivalent components of Figure 2A the graphics processor 208, Figure 2C the GPU 239, or Figure 2D the GPGPU 270.
[0149] As Figure 5A shown in, in one embodiment, the graphics core cluster 414 includes graphics core blocks 415, which may be Figure 4graphics core block 415A or graphics core block 415B. The graphics core block 415 may include any number of graphics cores (e.g., graphics core 515A, graphics core 515B to graphics core 515N). Multiple instances of the graphics core block 415 may be included. In one embodiment, the elements of the graphics cores 515A-515N have functions similar or equivalent to those of Figure 2B the elements of the graphics cores 221A-221F. In such an embodiment, each of the graphics cores 515A-515N includes circuit modules, which include but are not limited to vector engines 502A-502N, matrix engines 503A-503N, memory load / store units 504A-504N, instruction caches 505A-505N, data caches / shared local memories 506A-506N, ray tracing units 508A-508N, samplers 510A-510N. The circuit modules of the graphics cores 515A-515N may additionally include fixed function logic 512A-512N. The number of vector engines 502A-502N and matrix engines 503A-503N within the designed graphics cores 515A-515N may vary based on the designed workload, performance, and power goals.
[0150] Referring to graphics core 515A, the vector engine 502A and the matrix engine 503A can be configured to perform parallel computing operations on data in various integer and floating-point data formats based on instructions associated with a shader program. Each of the vector engine 502A and the matrix engine 503A can act as a programmable general-purpose computing unit that is capable of executing multiple hardware threads simultaneously while processing multiple data elements in parallel for each thread. The vector engine 502A and the matrix engine 503A support processing variable-width vectors in various SIMD widths, including but not limited to SIMD8, SIMD16, and SIMD32. Input data elements can be stored in registers as compressed data types, and the vector engine 502A and the matrix engine 503A can process various elements based on the data size of the elements. For example, when operating on a 256-bit-wide vector, the 256 bits of the vector are stored in a register, and the vector is processed as four independent 64-bit compressed data elements (quad-word (QW) sized data elements), eight independent 32-bit compressed data elements (double-word (DW) sized data elements), sixteen independent 16-bit compressed data elements (word (W) sized data elements), or thirty-two independent 8-bit data elements (byte (B) sized data elements). However, different vector widths and register sizes are possible. In one embodiment, the vector engine 502A and the matrix engine 503A can also be configured for SIMT operations on thread groups of various sizes (e.g., 8, 16, or 32 threads) or warps.
[0151] Continuing with graphics core 515A, the memory load / store unit 504A services memory access requests issued by the vector engine 502A, matrix engine 503A, and / or other components of the graphics core 515A that can access memory. The memory access requests can be processed by the memory load / store unit 504A to load or store the requested data to or from the cache or memory into the register file associated with the vector engine 502A and / or matrix engine 503A. The memory load / store unit 504A can also perform prefetch operations. In one embodiment, the memory load / store unit 504A is configured to provide SIMT scatter / gather prefetch or block prefetch of data stored in memory 610 from memory local to other tiles or from system memory via the tile interconnect 608. Prefetch can be performed on a particular L1 cache (e.g., data cache / shared local memory 506A), L2 cache 604, or L3 cache 606. In one embodiment, prefetching of the L3 cache 606 automatically causes the data to be stored in the L2 cache 604.
[0152] The instruction cache 505A stores instructions to be executed by the graphics core 515A. In one embodiment, the graphics core 515A also includes instruction fetch and prefetch circuitry that fetches or prefetches instructions into the instruction cache 505A. The graphics core 515A also includes instruction decoding logic to decode the instructions within the instruction cache 505A. The data cache / shared local memory 506A can be configured as a data cache managed by a cache controller implementing a cache replacement policy and / or as explicitly managed shared memory. The ray tracing unit 508A includes circuitry to accelerate ray tracing operations. The sampler 510A provides texture sampling for 3D operations and media sampling for media operations. The fixed function logic 512A includes fixed function circuitry shared among various instances of the vector engine 502A and matrix engine 503A. The graphics cores 515B - 515N can operate in a manner similar to the graphics core 515A.
[0153] The functions of the instruction caches 505A - 505N, data caches / shared local memories 506A - 506N, ray tracing units 508A - 508N, samplers 510A - 510N, and fixed function logics 512A - 512N correspond to the equivalent functions in the graphics processor architectures described herein. For example, the instruction caches 505A - 505N can operate in a manner similar to Figure 2D the instruction cache 255. The data caches / shared local memories 506A - 506N, ray tracing units 508A - 508N, and samplers 510A - 510N can operate in a manner similar to Figure 2BThe cache / SLM 228A-228F, ray tracing units 227A-227F, and samplers 226A-226F operate in a similar manner. The fixed function logic 512A-512N may include Figure 2B Elements of the geometry / fixed function pipeline 231 and / or additional fixed function logic 238 of. In one embodiment, the ray tracing units 508A-508N include circuitry for performing ray tracing acceleration operations performed by Figure 2C The ray tracing core 245 of.
[0154] As Figure 5B Shown, in one embodiment, the vector engine 502 includes an instruction fetch unit 537, a general register file array (GRF) 524, an architectural register file array (ARF) 526, a thread arbiter 522, a dispatch unit 530, a branch unit 532, a set of SIMD floating point units (FPU) 534, and in one embodiment, a set of integer SIMD ALUs 535. The GRF 524 and ARF 526 include a set of general register files and architectural register files associated with each hardware thread that may be active in the vector engine 502. In one embodiment, the architectural state of each thread is maintained in the ARF 526, while the data used during thread execution is stored in the GRF 524. The execution state of each thread, including the instruction pointer of each thread, may be saved in a thread-specific register in the ARF 526.
[0155] In one embodiment, the vector engine 502 has an architecture that is a combination of simultaneous multithreading (SMT) and fine-grained interleaved multithreading (IMT). The architecture has a modular configuration that can be fine-tuned at design time based on the number of registers per graphics core and the target number of simultaneous threads, where the graphics core resources are divided across the logic used to execute multiple simultaneous threads. The number of logical threads that can be executed by the vector engine 502 is not limited to the number of hardware threads, and multiple logical threads can be assigned to each hardware thread.
[0156] In one embodiment, the vector engine 502 is capable of co - issuing multiple instructions, each of which can be a different instruction. The thread arbiter 522 is capable of dispatching instructions to one of the send unit 530, the branch unit 532, or the (one or more) SIMD FPUs 534 for execution. Each execution thread can access 128 general - purpose registers within the GRF 524, where each register can store 32 bytes, which can be accessed as a variable - width vector of 32 - bit data elements. In one embodiment, each thread can access 4 kilobytes within the GRF 524, although the embodiment is not limited thereto, and more or fewer register resources can be provided in other embodiments. In one embodiment, the vector engine 502 is partitioned into seven hardware threads capable of independently performing computational operations, although the number of threads per vector engine 502 can also vary according to the embodiment. For example, up to 16 hardware threads are supported in one embodiment. In an embodiment where seven threads can access 4 kilobytes, the GRF 524 can store a total of 28 kilobytes. In the case where 16 threads can access 4 kilobytes, the GRF 524 can store a total of 64 kilobytes. Flexible addressing modes can allow registers to be addressed together to efficiently construct wider registers or represent strided rectangular block data structures.
[0157] In one embodiment, memory operations, sampler operations, and other longer - latency system communications are dispatched via "send" instructions executed by the messaging send unit 530. In one embodiment, branch instructions are dispatched to a dedicated branch unit 532 to facilitate SIMD divergence and eventual convergence.
[0158] In one embodiment, the vector engine 502 includes one or more SIMD floating-point units ((one or more) FPUs) 534 to perform floating-point operations. In one embodiment, the (one or more) FPUs 534 also support integer computations. In one embodiment, the (one or more) FPUs 534 are capable of performing up to M 32-bit floating-point (or integer) operations, or up to 2M 16-bit integer or 16-bit floating-point operations. In one embodiment, at least one of the (one or more) FPUs provides extended math capabilities to support high-throughput transcendental math functions and double-precision 64-bit floating point. In some embodiments, there is also a set of 8-bit integer SIMD ALUs 535, and the set of 8-bit integer SIMD ALUs 535 can be specifically optimized to perform operations associated with machine learning computations. In one embodiment, the SIMD ALUs are replaced by a set of additional SIMD FPUs 534 configurable to perform integer and floating-point operations. In one embodiment, the SIMD FPUs 534 and SIMD ALUs 535 are configurable to execute SIMT programs. In one embodiment, combined SIMD+SIMT operations are supported.
[0159] In one embodiment, an array of multiple instances of the vector engine 502 can be instantiated in a graphics core. For scalability, a product architect can select the exact number of vector engines grouped per graphics core. In one embodiment, the vector engine 502 is capable of executing instructions across multiple execution lanes. In additional embodiments, each thread executed on the vector engine 502 is executed on a different lane.
[0160] As Figure 5CAs shown, in one embodiment, matrix engine 503 includes an array of processing elements configured to perform tensor operations, including vector / matrix and matrix / matrix operations such as, but not limited to, matrix multiplication and / or dot product operations. Matrix engine 503 is configured with M rows and N columns of processing elements (552AA - 552MN), which include multiplier and adder circuits organized in a pipelined fashion. In one embodiment, processing elements 552AA - 552MN form the physical pipeline stage of an N-wide M-deep systolic array that can be used to perform vector / matrix or matrix / matrix operations including matrix multiplication, fused multiply-add, dot product, or other general matrix-matrix multiplication (GEMM) operations in a data parallel manner. In one embodiment, matrix engine 503 supports 16-bit floating-point operations, as well as 8-bit, 4-bit, 2-bit, and binary integer operations. Matrix engine 503 can also be configured to accelerate specific machine learning operations. In such an embodiment, matrix engine 503 can be configured with support for the bfloat (brain floating point) 16-bit floating-point format or the tensor floating point 32-bit floating-point format (TF32) that have a different number of mantissa and exponent bits relative to the Institute of Electrical and Electronics Engineers (IEEE) 754 format.
[0161] In one embodiment, during each cycle, each stage can add the result of the operation performed at that stage to the output of the previous stage. In other embodiments, after a group of computational cycles, the data movement pattern among processing elements 552AA - 552MN can vary based on the instruction or macro operation being executed. For example, in one embodiment, partial sum feedback is enabled, and the processing element can alternatively add the output of the current cycle to the output generated in the previous cycle. In one embodiment, the last stage of the systolic array can be configured with feedback to the initial stage of the systolic array. In such an embodiment, the number of physical pipeline stages can be decoupled from the number of logical pipeline stages supported by matrix engine 503. For example, in the case where processing elements 552AA - 552MN are configured as an M-physical-stage systolic array, the feedback from stage M to the initial pipeline stage can enable processing elements 552AA - 552MN to operate as a systolic array with, for example, 2M, 3M, 4M, etc. logical pipeline stages.
[0162] In one embodiment, matrix engine 503 includes memories 541A - 541N, 542A - 542M to store input data in the form of row and column data of an input matrix. Memories 542A - 542M can be configured to store row elements (A0 - Am) of a first input matrix, while memories 541A - 541N can be configured to store column elements (B0 - Bn) of a second input matrix. The row and column elements are provided as inputs to processing elements 552AA - 552MN for processing. In one embodiment, before the row and column elements of the input matrix are provided to memories 541A - 541N, 542A - 542M, those elements can be stored in the systolic register file 540 within matrix engine 503. In one embodiment, the systolic register file 540 is excluded, and memories 541A - 541N, 542A - 542M are loaded from registers in an associated vector engine (e.g., Figure 5B the GRF 524 of vector engine 502) or other memories of a graphics core that includes matrix engine 503 (e.g., Figure 5A the data cache / shared local memory 506A of matrix engine 503A). Results generated by processing elements 552AA - 552MN are then output to an output buffer and / or written to a register file (e.g., systolic register file 540, GRF 524, data cache / shared local memory 506A - 506N) for further processing by other functional units of the graphics processor or for output to memory.
[0163] In some embodiments, matrix engine 503 is configured with support for input sparsity, where multiplication operations in sparse regions of the input data can be bypassed by skipping multiplication operations with zero - valued operands. In one embodiment, processing elements 552AA - 552MN are configured to skip the execution of certain operations with zero - valued inputs. In one embodiment, sparsity within the input matrix can be detected, and operations with known zero output values can be bypassed before being submitted to processing elements 552AA - 552MN. Loading zero - valued operands into the processing elements can be bypassed, and processing elements 552AA - 552MN can be configured to perform multiplication on non - zero - valued input elements. Matrix engine 503 can also be configured with support for output sparsity such that operations with results pre - determined to be zero are bypassed. For input sparsity and / or output sparsity, in one embodiment, metadata is provided to processing elements 552AA - 552MN to indicate which processing elements and / or data channels are to be active during a processing cycle.
[0164] In one embodiment, matrix engine 503 includes hardware that enables operations on sparse data having a compressed representation of a sparse matrix storing non-zero values and metadata identifying the positions of the non-zero values within the matrix. Exemplary compressed representations include, but are not limited to, compressed tensor representations such as compressed sparse row (CSR), compressed sparse column (CSC), compressed sparse fiber (CSF) representations. Support for the compressed representation enables operations to be performed on inputs in a compressed tensor format without the need to decompress or decode the compressed representation. In such an embodiment, operations can be performed only on the non-zero input values, and the resulting non-zero output values can be mapped into the output matrix. In some embodiments, hardware support is also provided for a machine-specific lossless data compression format used when transferring data within the hardware or across system buses. For sparse input data, such data can be retained in a compressed format, and matrix engine 503 can use the compressed metadata for the compressed data to enable operations to be performed only on non-zero values or to enable bypassing of zero data input blocks for multiplication operations.
[0165] In various embodiments, the input data can be provided by a programmer in a compressed tensor representation, or a codec can compress the input data into a compressed tensor representation or another sparse data encoding. In addition to support for compressed tensor representations, stream compression of sparse input data can be performed before the data is provided to processing elements 552AA - 552MN. In one embodiment, compression is performed on data written to the cache associated with graphics core cluster 414, where the compression is performed using an encoding supported by matrix engine 503. In one embodiment, matrix engine 503 includes support for inputs having a structured sparsity, in which a predetermined sparsity level or pattern is imposed on the input data. This data can be compressed to a known compression ratio, where the compressed data is processed by processing elements 552AA - 552MN according to the metadata associated with the compressed data.
[0166] Figure 6 Illustrated is tile 600 of a multi-tile processor according to an embodiment. In one embodiment, tile 600 represents Figure 3B one of graphics engine tiles 310A - 310D or Figure 3C one of compute engine tiles 340A - 340D of a multi-tile graphics processor. Tile 600 of the multi-tile graphics processor includes an array of graphics core clusters (e.g., graphics core clusters 414A, graphics core clusters 414B through graphics core clusters 414N), where each graphics core cluster has an array of graphics cores 515A - 515N. Tile 600 also includes a global dispatcher 602 for dispatching threads to the processing resources of tile 600.
[0167] The tile 600 may include or be coupled to an L3 cache 606 and a memory 610. In various embodiments, the L3 cache 606 may be excluded, or the tile 600 may include additional levels of cache, such as an L4 cache. In one embodiment, each instance of the tile 600 in a multi-tile graphics processor has an associated memory 610, such as in Figure 3B and Figure 3C . In one embodiment, the multi-tile processor may be configured as a multi-chip module, where the L3 cache 606 and / or the memory 610 reside on separate dies different from the graphics core clusters 414A - 414N. In this context, a die is an at least partially packaged integrated circuit that includes different logic units that can be assembled with other dies into a larger package. For example, the L3 cache 606 may be included in a dedicated cache die, or may reside on the same die as the graphics core clusters 414A - 414N. In one embodiment, the L3 cache 606 may be included in an active substrate die or an active interposer, as illustrated in Figure 11C .
[0168] The memory fabric 603 enables communication between the graphics core clusters 414A - 414N, the L3 cache 606, and the memory 610. The L2 cache 604 is coupled to the memory fabric 603 and may be configured to cache transactions executed via the memory fabric 603. The tile interconnect 608 enables communication with other tiles on the graphics processor and may be one of the tile interconnects 323A - 323F of Figure 3B and Figure 3C . In embodiments where the L3 cache 606 is excluded from the tile 600, the L2 cache 604 may be configured as a combined L2 / L3 cache. In a particular implementation, the memory fabric 603 may be configured to route data to the L3 cache 606 or the memory controller associated with the memory 610 based on the presence or absence of the L3 cache 606. The L3 cache 606 may be configured as a per-tile cache for the processing resources dedicated to the tile 600, or may be a partition of the GPU-wide L3 cache.
[0169] Figure 7FIG. 0 is a block diagram of a graphics processor instruction format 700 according to some embodiments. In one or more embodiments, a graphics processor core supports an instruction set having instructions in multiple formats. The solid boxes illustrate components generally included in a graphics core instruction, while the dashed boxes include optional or components only included in a subset of the instructions. In some embodiments, the graphics processor instruction format 700 described and illustrated is a macro-instruction, as it is an instruction supplied to the graphics core, as opposed to micro-operations generated by instruction decoding once the instruction is processed. Thus, a single instruction can cause the hardware to execute multiple micro-operations.
[0170] In some embodiments, the graphics processor natively supports instructions in a 128-bit instruction format 710. Based on the selected instruction, instruction options, and number of operands, a 64-bit compressed instruction format 730 can be used for some instructions. The native 128-bit instruction format 710 provides access to all instruction options, while some options and operations are restricted in the 64-bit format 730. The available native instructions in 64-bit format 730 vary by embodiment. In some embodiments, a set of index values in an index field 713 are used to partially compress the instruction. The graphics core hardware references a set of compression tables based on the index values and uses the compressed table output to reconstruct the native instruction in 128-bit instruction format 710. Other sizes and formats of instructions can be used.
[0171] For each format, an instruction opcode 712 defines the operation for the graphics core to perform. The graphics core executes each instruction in parallel across multiple data elements of each operand. For example, in response to an add instruction, the graphics core performs a simultaneous add operation across each color channel representing a texture element or picture element. By default, the graphics core executes each instruction across all data channels of the operand. In some embodiments, an instruction control field 714 enables control of certain execution options such as channel selection (e.g., predication) and data channel order (e.g., swizzle). For instructions in 128-bit instruction format 710, an execution size field 716 limits the number of data channels to be executed in parallel. In some embodiments, the execution size field 716 is not available for use in the 64-bit compact instruction format 730.
[0172] Some graphics core instructions have up to three operands, which include two source operands, src0 720, src1 722, and one destination 718. In some embodiments, the graphics core supports dual-destination instructions, where one of the destinations is implicit. Data manipulation instructions can have a third source operand (e.g., SRC2 724), where the instruction opcode 712 determines the number of source operands. The last source operand of the instruction can be an immediate (e.g., hard-coded) value passed with the instruction.
[0173] In some embodiments, the 128-bit instruction format 710 includes an access / address mode field 726 that specifies, for example, whether to use direct register addressing mode or indirect register addressing mode. When using direct register addressing mode, the register addresses of one or more operands are directly provided by bits in the instruction.
[0174] In some embodiments, the 128-bit instruction format 710 includes an access / address mode field 726 that specifies the address mode and / or access mode of the instruction. In one embodiment, the access mode is used to define the data access alignment of the instruction. Some embodiments support access modes including a 16-byte alignment access mode and a 1-byte alignment access mode, where the byte alignment of the access mode determines the access alignment of the instruction operands. For example, when in a first mode, the instruction may use byte-aligned addressing for source and destination operands, and when in a second mode, the instruction may use 16-byte-aligned addressing for all source and destination operands.
[0175] In one embodiment, the address mode portion of the access / address mode field 726 determines whether the instruction will use direct addressing or indirect addressing. When using direct register addressing mode, the bits in the instruction directly provide the register addresses of one or more operands. When using indirect register addressing mode, the register addresses of one or more operands can be calculated based on the address register value and the address immediate field in the instruction.
[0176] In some embodiments, instructions are grouped based on the 712-bit opcode field to simplify opcode decoding 740. For an 8-bit opcode, bits 4, 5, and 6 allow the graphics core to determine the type of opcode. The exact opcode grouping shown is merely an example. In some embodiments, the move and logic opcode group 742 includes data move and logic instructions (e.g., move (mov), compare (cmp)). In some embodiments, the move and logic group 742 shares the five most significant bits (MSB), where the move (mov) instruction takes the form 0000xxxxb and the logic instruction takes the form 0001xxxxb. The flow control instruction group 744 (e.g., call, jump (jmp)) includes instructions taking the form 0010xxxxb (e.g., 0x20). The miscellaneous instruction group 746 includes a mix of instructions, including synchronization instructions (e.g., wait, send) taking the form 0011xxxxb (e.g., 0x30). The parallel math instruction group 748 includes per-component arithmetic instructions (e.g., add, multiply (mul)) taking the form 0100xxxxb (e.g., 0x40). The parallel math group 748 performs arithmetic operations in parallel across data channels. The vector math group 750 includes arithmetic instructions (e.g., dp4) taking the form 0101xxxxb (e.g., 0x50). The vector math group performs arithmetic operations such as dot product calculations on vector operands. The illustrated opcode decoding 740 can be used in one embodiment to determine which part of the graphics core will be used to execute the decoded instruction. For example, some instructions can be designated as systolic instructions to be executed by the systolic array. Other instructions such as ray tracing instructions (not shown) can be routed to the ray tracing core or ray tracing logic within a slice or partition of the execution logic.
[0177] Graphics pipeline
[0178] Figure 8 is a block diagram of another embodiment of the graphics processor 800. Elements with the same reference numeral (or name) as elements in any other figure in this document Figure 8 can operate or function in any way similar to the ways described elsewhere in this document, but are not limited to such.
[0179] In some embodiments, the graphics processor 800 includes a geometry pipeline 820, a media pipeline 830, a display engine 840, thread execution logic 850, and a render output pipeline 870. In some embodiments, the graphics processor 800 is a graphics processor within a multi-core processing system that includes one or more general-purpose processing cores. The graphics processor is controlled by register writes to one or more control registers (not shown) or via commands issued to the graphics processor 800 over a ring interconnect 802. In some embodiments, the ring interconnect 802 couples the graphics processor 800 to other processing components, such as other graphics processors or general-purpose processors. Commands from the ring interconnect 802 are interpreted by a command stream converter 803 that supplies instructions to the various components of the geometry pipeline 820 or the media pipeline 830.
[0180] In some embodiments, the command stream converter 803 directs the operation of a vertex fetcher 805 that reads vertex data from memory and executes vertex processing commands provided by the command stream converter 803. In some embodiments, the vertex fetcher 805 provides vertex data to a vertex shader 807 that performs coordinate space transformations and lighting operations on each vertex. In some embodiments, the vertex fetcher 805 and the vertex shader 807 execute vertex processing instructions by dispatching execution threads to the graphics cores 852A - 852B via a thread dispatcher 831.
[0181] In some embodiments, the graphics cores 852A - 852B are an array of vector processors having instruction sets for performing graphics and media operations. In some embodiments, the graphics cores 852A - 852B have attached L1 caches 851 that are specific to each array or shared between the arrays. The cache can be configured as a data cache, an instruction cache, or a single cache partitioned to hold data and instructions in different partitions.
[0182] In some embodiments, the geometry pipeline 820 includes tessellation components to perform hardware-accelerated tessellation of 3D objects. In some embodiments, a programmable hull shader 811 configures the tessellation operation. A programmable domain shader 817 provides backend evaluation of the tessellation output. The tessellator 813 operates under the guidance of the hull shader 811 and includes specialized logic to generate a set of detailed geometric objects based on a coarse geometric model provided as input to the geometry pipeline 820. In some embodiments, if tessellation is not used, the tessellation components (e.g., the hull shader 811, the tessellator 813, and the domain shader 817) can be bypassed. The tessellation components can operate based on data received from the vertex shader 807.
[0183] In some embodiments, the complete geometric object can be processed by the geometry shader 819 via one or more threads dispatched to the graphics cores 852A - 852B, or can proceed directly to the clipper 829. In some embodiments, the geometry shader operates on the entire geometric object, rather than on vertices or vertex patches as in previous stages of the graphics pipeline. If tessellation is disabled, the geometry shader 819 receives input from the vertex shader 807. In some embodiments, the geometry shader 819 can be programmed by a geometry shader program to perform geometric tessellation when the tessellation unit is disabled.
[0184] Before rasterization, the clipper 829 processes the vertex data. The clipper 829 can be a programmable clipper or a fixed - function clipper with clipping and geometry shader functionality. In some embodiments, the rasterizer and depth test component 873 in the render output pipeline 870 dispatches pixel shaders to convert the geometric object into a per - pixel representation. In some embodiments, the pixel shader logic is included in the thread execution logic 850. In some embodiments, the application can bypass the rasterizer and depth test component 873 and access the un - rasterized vertex data via the stream output unit 823.
[0185] The graphics processor 800 has an interconnect bus, interconnect fabric, or some other interconnect mechanism that allows data and messages to be passed between the major components of the processor. In some embodiments, the graphics cores 852A - 852B and associated logic units (e.g., L1 cache 851, sampler 854, texture cache 858, etc.) are interconnected via data ports 856 to perform memory accesses and communicate with the processor's render output pipeline components. In some embodiments, the sampler 854, caches 851, 858, and graphics cores 852A - 852B each have separate memory access paths. In one embodiment, the texture cache 858 can also be configured as a sampler cache.
[0186] In some embodiments, the rendering output pipeline 870 includes a rasterizer and depth test component 873 that converts vertex-based objects into associated pixel-based representations. In some embodiments, the rasterizer logic includes a windower / masker unit for performing fixed-function triangle and line rasterization. Associated render cache 878 and depth cache 879 are also available in some embodiments. The pixel operation component 877 performs pixel-based operations on the data, although in some instances, pixel operations associated with 2D operations (e.g., bit-block image transfer with blending) are performed by the 2D engine 841 or, at display time, by the display controller 843 using an overlay display plane instead. In some embodiments, the shared L3 cache 875 is available to all graphics components, allowing data to be shared without using the main system memory.
[0187] In some embodiments, the media pipeline 830 includes a media engine 837 and a video front end 834. In some embodiments, the video front end 834 receives pipeline commands from the command stream converter 803. In some embodiments, the media pipeline 830 includes a separate command stream converter. In some embodiments, the video front end 834 processes media commands before sending them to the media engine 837. In some embodiments, the media engine 837 includes a thread spawning function to spawn threads for dispatch to the thread execution logic 850 via the thread dispatcher 831.
[0188] In some embodiments, the graphics processor 800 includes a display engine 840. In some embodiments, the display engine 840 is external to the processor 800 and is coupled to the graphics processor via the ring interconnect 802 or some other interconnect bus or fabric. In some embodiments, the display engine 840 includes a 2D engine 841 and a display controller 843. In some embodiments, the display engine 840 contains dedicated logic that can operate independently of the 3D pipeline. In some embodiments, the display controller 843 is coupled to a display device (not shown), which may be a system-integrated display device (such as in a laptop computer) or an external display device attached via a display device connector.
[0189] In some embodiments, the geometry pipeline 820 and the media pipeline 830 may be configured to perform operations based on multiple graphics and media programming interfaces and are not specific to any one application programming interface (API). In some embodiments, the driver software for the graphics processor translates API calls specific to a particular graphics or media library into commands that can be processed by the graphics processor. In some embodiments, support is provided for all of the Open Graphics Library (OpenGL), Open Computing Language (OpenCL), and / or Vulkan graphics and compute APIs from the Khronos Group. In some embodiments, support may also be provided for the Direct3D library from Microsoft Corporation. In some embodiments, combinations of these libraries may be supported. Support may also be provided for the Open Source Computer Vision Library (OpenCV). Future APIs with compatible 3D pipelines will also be supported if a mapping from the future API's pipeline to the graphics processor's pipeline can be made.
[0190] Graphics Pipeline Programming
[0191] Figure 9A is a block diagram illustrating a graphics processor command format 900 that can be used to program a graphics processing pipeline according to some embodiments. Figure 9B is a block diagram illustrating a graphics processor command sequence 910 according to an embodiment. Figure 9A The solid boxes in illustrate the components that are generally included in a graphics command, while the dashed lines include optional or components that are only included in a subset of graphics commands. Figure 9A An exemplary graphics processor command format 900 of includes data fields for a client 902, a command operation code (opcode) 904, and a data field 906 that are used to identify the command. Some commands also include a sub-opcode 905 and a command size 908.
[0192] In some embodiments, client 902 designates the client unit of the graphics device that processes command data. In some embodiments, the graphics processor command parser examines the client fields of each command to condition further processing of the command and routes the command data to the appropriate client unit. In some embodiments, the graphics processor client units include a memory interface unit, a rendering unit, a 2D unit, a 3D unit, and a media unit. Each client unit has a corresponding processing pipeline for processing commands. Once a client unit receives a command, the client unit reads the opcode 904 and sub-opcode 905 (if the sub-opcode 905 exists) to determine the operation to be performed. The client unit uses the information in the data field 906 to execute the command. For some commands, an explicit command size 908 is expected to specify the size of the command. In some embodiments, the command parser automatically determines the size of at least some of the commands in the command based on the command opcode. In some embodiments, commands are aligned via multiples of a double word. Other command formats can be used.
[0193] Figure 9B The flowchart in illustrates an exemplary graphics processor command sequence 910. In some embodiments, software or firmware of a data processing system characterized by an embodiment of the graphics processor uses a version of the illustrated command sequence to set up, execute, and terminate a set of graphics operations. The sample command sequence is shown and described for illustrative purposes only as embodiments are not limited to these particular commands or this command sequence. Additionally, commands can be issued as a batch of commands in a command sequence such that the graphics processor will process the sequence of commands at least partially concurrently.
[0194] In some embodiments, the graphics processor command sequence 910 can begin with a pipeline flush command 912 to cause any active graphics pipeline to complete the current outstanding commands of that pipeline. In some embodiments, the 3D pipeline 922 and the media pipeline 924 do not operate concurrently. A pipeline flush is performed to cause the active graphics pipeline to complete any outstanding commands. In response to the pipeline flush, the command parser for the graphics processor will suspend command processing until the active drawing engine has completed the outstanding operations and the associated read caches are invalidated. Optionally, any data marked "dirty" in the render cache can be flushed to memory. In some embodiments, the pipeline flush command 912 can be used for pipeline synchronization or before placing the graphics processor in a low power state.
[0195] In some embodiments, the pipeline select command 913 is used when a command sequence requires the graphics processor to explicitly switch between pipelines. In some embodiments, unless the context will issue commands for both pipelines, the pipeline select command 913 is only required once within an execution context before issuing a pipeline command. In some embodiments, immediately before a pipeline switch via the pipeline select command 913, the pipeline dump clear command 912 is required.
[0196] In some embodiments, the pipeline control command 914 configures the graphics pipeline for operation and is used to program the 3D pipeline 922 and the media pipeline 924. In some embodiments, the pipeline control command 914 configures the pipeline state for the active pipeline. In one embodiment, the pipeline control command 914 is used for pipeline synchronization and clears data from one or more caches within the active pipeline before processing a batch of commands.
[0197] In some embodiments, the command 916 related to the return buffer state is used to configure a set of return buffers for writing data for the corresponding pipeline. Some pipeline operations require the allocation, selection, or configuration of one or more return buffers into which intermediate data is written during processing. In some embodiments, the graphics processor also uses one or more return buffers to store output data and perform cross-thread communication. In some embodiments, the return buffer state 916 includes selecting the size and number of return buffers to be used for a set of pipeline operations.
[0198] The remaining commands in the command sequence vary based on the active pipeline used for operation. Based on the pipeline determination 920, the command sequence is customized to the 3D pipeline 922 starting with the 3D pipeline state 930 or the media pipeline 924 starting with the media pipeline state 940.
[0199] The commands used to configure the 3D pipeline state 930 include 3D state setting commands for vertex buffer state, vertex element state, constant color state, depth buffer state, and other state variables to be configured before processing 3D primitive commands. The values of these commands are determined at least in part based on the specific 3D API in use. In some embodiments, the 3D pipeline state 930 commands can also selectively disable or bypass those elements if certain pipeline elements will not be used.
[0200] In some embodiments, 3D primitive 932 commands are used to submit 3D primitives to be processed by the 3D pipeline. The commands and associated parameters passed to the graphics processor via the 3D primitive 932 commands are forwarded to the vertex fetch function in the graphics pipeline. The vertex fetch function uses the 3D primitive 932 command data to generate vertex data structures. The vertex data structures are stored in one or more return buffers. In some embodiments, 3D primitive 932 commands are used to perform vertex operations on 3D primitives via a vertex shader. To process the vertex shader, the 3D pipeline 922 dispatches the shader program to the graphics core.
[0201] In some embodiments, the 3D pipeline 922 is triggered via execution of a 934 command or event. In some embodiments, a register write triggers command execution. In some embodiments, execution is triggered via a "go" or "kick" command in a command sequence. In one embodiment, a pipeline synchronization command used to dump and clear a command sequence passing through the graphics pipeline is used to trigger command execution. The 3D pipeline will perform geometric processing for 3D primitives. Once the operations are complete, the resulting geometric objects are rasterized, and the pixel engine colors the resulting pixels. For those operations, additional commands may also be included to control pixel shading and pixel backend operations.
[0202] In some embodiments, when performing media operations, the graphics processor command sequence 910 follows the media pipeline 924 path. Generally, the specific use and manner of programming for the media pipeline 924 depends on the media or compute operation to be performed. Specific media decoding operations can be offloaded to the media pipeline during media decoding. In some embodiments, the media pipeline can also be bypassed, and resources provided by one or more general-purpose processing cores can be used to perform media decoding in whole or in part. In one embodiment, the media pipeline also includes elements for general-purpose graphics processing unit (GPGPU) operations, where the graphics processor is used to perform SIMD vector operations using a compute shader program that is not explicitly related to the rendering of graphics primitives.
[0203] In some embodiments, the media pipeline 924 is configured in a manner similar to the 3D pipeline 922. A set of commands used to configure the media pipeline state 940 are dispatched or placed in a command queue before the media object command 942. In some embodiments, the commands for the media pipeline state 940 include data used to configure the media pipeline elements that will be used to process media objects. This includes data used to configure video decoding and video encoding logic within the media pipeline, such as encoding and decoding formats. In some embodiments, the commands for the media pipeline state 940 also support the use of one or more pointers to "indirect" state elements containing a batch of state settings.
[0204] In some embodiments, the media object command 942 supplies a pointer to a media object for processing by the media pipeline. The media object includes a memory buffer that contains video data to be processed. In some embodiments, all media pipeline states must be valid before issuing the media object command 942. Once the pipeline state is configured and the media object command 942 is queued, the media pipeline 924 is triggered via an execute command 944 or an equivalent execution event (e.g., a register write). The output from the media pipeline 924 can then be post-processed by operations provided by the 3D pipeline 922 or the media pipeline 924. In some embodiments, GPGPU operations are configured and executed in a manner similar to media operations.
[0205] Graphics Software Architecture
[0206] Figure 10 Illustrated is an exemplary graphics software architecture for a data processing system 1000 in accordance with some embodiments. In some embodiments, the software architecture includes a 3D graphics application 1010, an operating system 1020, and at least one processor 1030. In some embodiments, the processor 1030 includes a graphics processor 1032 and one or more general-purpose processor cores 1034. The graphics application 1010 and the operating system 1020 each execute in the system memory 1050 of the data processing system.
[0207] In some embodiments, the 3D graphics application 1010 contains one or more shader programs that include shader instructions 1012. The shader language instructions can be in a high-level shader language such as the High-Level Shader Language (HLSL) of Direct3D or the OpenGL Shader Language (GLSL), etc. The application also includes executable instructions 1014 in a machine language suitable for execution by the general-purpose processor cores 1034. The application also includes a graphics object 1016 defined by vertex data.
[0208] In some embodiments, the operating system 1020 is from Microsoft Corporation An operating system, a proprietary UNIX-like operating system, or an open-source UNIX-like operating system using a Linux kernel variant. The operating system 1020 is capable of supporting a graphics API 1022, such as the Direct3D API, the OpenGL API, or the Vulkan API. When the Direct3D API is in use, the operating system 1020 uses a front-end shader compiler 1024 to compile any shader instructions 1012 in HLSL into a lower-level shader language. The compilation can be just-in-time (JIT) compilation or the application can perform shader pre-compilation. In some embodiments, the high-level shader is compiled into a low-level shader during the compilation of the 3D graphics application 1010. In some embodiments, the shader instructions 1012 are provided in an intermediate form, such as a certain version of the standard portable intermediate representation (SPIR) used by the Vulkan API.
[0209] In some embodiments, the user-mode graphics driver 1026 contains a backend shader compiler 1027 for converting the shader instructions 1012 into a hardware-specific representation. When the OpenGL API is in use, the shader instructions 1012 in the GLSL high-level language are passed to the user-mode graphics driver 1026 for compilation. In some embodiments, the user-mode graphics driver 1026 uses the operating system kernel-mode function 1028 to communicate with the kernel-mode graphics driver 1029. In some embodiments, the kernel-mode graphics driver 1029 communicates with the graphics processor 1032 to dispatch commands and instructions.
[0210] IP core implementation
[0211] One or more aspects of at least one embodiment can be implemented by representative code stored on a machine-readable medium, the representative code representing and / or defining logic within an integrated circuit such as a processor. For example, the machine-readable medium can include instructions representing various logics within the processor. When read by the machine, the instructions can cause the machine to fabricate the logic to perform the techniques described herein. Such a representation, known as an "IP core", is a reusable unit of logic for an integrated circuit and can be stored on a tangible machine-readable medium as a hardware model describing the structure of the integrated circuit. The hardware model can be supplied to various customers or manufacturing facilities, which load the hardware model onto a fabrication machine for manufacturing the integrated circuit. The integrated circuit can be fabricated such that the circuit performs the described operations associated with any of the embodiments described herein.
[0212] Figure 11AFIG. is a block diagram of an IP core development system 1100 that can be used to fabricate integrated circuits to perform operations. The IP core development system 1100 can be used to generate modular, reusable designs that can be incorporated into larger designs or used to construct an entire integrated circuit (e.g., a SOC integrated circuit). A design facility 1130 can generate a software simulation 1110 of an IP core design in a high-level programming language (e.g., C / C++). The software simulation 1110 can be used to design, test, and verify the behavior of the IP core using a simulation model 1112. The simulation model 1112 can include functional, behavioral, and / or timing simulations. A register transfer level (RTL) design 1115 can then be created or synthesized from the simulation model 1112. The RTL design 1115 is an abstraction of the behavior of an integrated circuit that models the digital signal flow between hardware registers, including the associated logic performed using the modeled digital signals. In addition to the RTL design 1115, lower-level designs at the logic level or transistor level can also be created, designed, or synthesized. Thus, the specific details of the initial design and simulation can vary.
[0213] The RTL design 1115 or equivalent can be further synthesized by the design facility into a hardware model 1120, which can be in a hardware description language (HDL) or some other representation of physical design data. The HDL can be further simulated or tested to verify the IP core design. A non-volatile memory 1140 (e.g., a hard disk, flash memory, or any non-volatile storage medium) can be used to store the IP core design for delivery to a third-party fabrication facility 1165. Alternatively, the IP core design can be transmitted via a wired connection 1150 or a wireless connection 1160 (e.g., via the Internet). The fabrication facility 1165 can then fabricate an integrated circuit based at least in part on the IP core design. The fabricated integrated circuit can be configured to perform operations in accordance with at least one embodiment described herein.
[0214] Figure 11BFIG. illustrates a cross-sectional side view of an integrated circuit package assembly 1170 in accordance with some embodiments described herein. The integrated circuit package assembly 1170 illustrates the implementation of one or more processor or accelerator devices as described herein. The package assembly 1170 includes a plurality of hardware logic units 1172, 1174 connected to a substrate 1180. The logic 1172, 1174 may be implemented at least partially in configurable logic or fixed-function logic hardware and can include any one or more portions of the (one or more) processor cores, (one or more) graphics processors, or other accelerator devices described herein. Each unit of the logic 1172, 1174 can be implemented within a semiconductor die and is coupled to the substrate 1180 via an interconnect structure 1173. The interconnect structure 1173 may be configured to route electrical signals between the logic 1172, 1174 and the substrate 1180 and can include interconnects such as, but not limited to, bumps or posts. In some embodiments, the interconnect structure 1173 may be configured to route electrical signals, e.g., input / output (I / O) signals and / or power or ground signals associated with the operation of the logic 1172, 1174. In some embodiments, the substrate 1180 is an epoxy-based laminate substrate. In other embodiments, the substrate 1180 may include other suitable types of substrates. The package assembly 1170 can be connected to other electrical devices via package interconnects 1183. The package interconnects 1183 may be coupled to the surface of the substrate 1180 to route electrical signals to other electrical devices such as a motherboard, other chip sets, or multi-chip modules.
[0215] In some embodiments, the logic units 1172, 1174 are electrically coupled to a bridge 1182 that is configured to route electrical signals between the logic 1172, 1174. The bridge 1182 may be a dense interconnect structure that provides routing for electrical signals. The bridge 1182 may include a bridge substrate made of glass or a suitable semiconductor material. Circuit routing features can be formed on the bridge substrate to provide chip-to-chip connections between the logic 1172, 1174.
[0216] Although two logic units 1172, 1174 and a bridge 1182 are illustrated, the embodiments described herein can include more or fewer logic units on one or more dies. Since the bridge 1182 can be excluded when the logic is included on a single die, one or more dies can be connected by zero or more than zero bridges. Alternatively, multiple dies or logic units can be connected by one or more bridges. Additionally, multiple logic units, dies, and bridges can be connected together in other possible configurations, including three-dimensional configurations.
[0217] Figure 11CFIG. illustrates a packaged assembly 1190 including a plurality of hardware logic die units coupled to a substrate 1180. Graphics processing units, parallel processors, and / or compute accelerators as described herein can be composed of diverse silicon dies fabricated separately. A diverse collection of dies with different IP core logics can be assembled into a single device. Additionally, die integration into a base die or base die can be enabled using active interposer technology. The concepts described herein enable interconnection and communication between different forms of IP within a GPU. Different process technologies can be used to fabricate and form IP cores during fabrication, which avoids the complexity of aggregating multiple IPs (especially on large SoCs with several flavors of IP) into the same fabrication process. Enabling the use of multiple process technologies improves time to market and provides a cost-effective way to create multiple product SKUs. Additionally, disaggregated IP is more easily power-gated independently, and components not in use for a given workload can be powered off, thereby reducing overall power consumption.
[0218] In various embodiments, the packaged assembly 1190 can include components and dies interconnected by a fabric 1185 and / or one or more bridges 1187. Dies within the packaged assembly 1190 can have a 2.5D arrangement using chip-on-wafer-on-substrate stacking, in which multiple dies are stacked side by side on a silicon interposer 1189 that couples the dies to the substrate 1180. The substrate 1180 includes electrical connections to the package interconnects 1183. In one embodiment, the silicon interposer 1189 is a passive interposer that includes through-silicon vias (TSVs) to electrically couple the dies within the packaged assembly 1190 to the substrate 1180. In one embodiment, the silicon interposer 1189 is an active interposer that includes embedded logic in addition to the TSVs. In such an embodiment, the dies within the packaged assembly 1190 are arranged on top of the active interposer 1189 using a 3D face-to-face die stacking arrangement. In addition to the interconnect fabric 1185 and the silicon bridges 1187, the active interposer 1189 can include hardware logic 1191 for I / O, a cache memory 1192, and other hardware logic 1193. The fabric 1185 enables communication between the various logic dies 1172, 1174 and the logic 1191, 1193 within the active interposer 1189. The fabric 1185 can be a NoC interconnect or another form of packet-switching fabric that exchanges data packets between components of the packaged assembly. For complex assemblies, the fabric 1185 can be a dedicated die such that communication can occur between the various hardware logics of the packaged assembly 1190.
[0219] The bridge structure 1187 within the active interposer 1189 can be used to facilitate point-to-point interconnects, for example, between the logic or I / O die 1174 and the memory die 1175. In some implementations, the bridge structure 1187 can also be embedded within the substrate 1180. The hardware logic die can include the dedicated hardware logic die 1172, the logic or I / O die 1174, and / or the memory die 1175. The hardware logic die 1172 and the logic or I / O die 1174 can be implemented at least in part in configurable logic or fixed-function logic hardware and can include any one or more portions of the (one or more) processor cores, (one or more) graphics processors, parallel processors, or other accelerator devices described herein. The memory die 1175 can be DRAM (e.g., GDDR, HBM) memory or cache (SRAM) memory. The cache memory 1192 within the active interposer 1189 (or substrate 1180) can act as a global cache for the package assembly 1190, a part of a distributed global cache, or a dedicated cache for the fabric 1185.
[0220] Each die can be fabricated as a separate semiconductor die and coupled to a base die embedded within or coupled to the substrate 1180. The coupling to the substrate 1180 can be performed via the interconnect structure 1173. The interconnect structure 1173 can be configured to route electrical signals between the various dies and the logic within the substrate 1180. The interconnect structure 1173 can include interconnects such as, but not limited to, bumps or pillars. In some embodiments, the interconnect structure 1173 can be configured to route electrical signals, e.g., input / output (I / O) signals and / or power or ground signals associated with the operation of the logic, I / O, and memory dies. In one embodiment, an additional interconnect structure couples the active interposer 1189 to the substrate 1180.
[0221] In some embodiments, the substrate 1180 is an epoxy-based laminate substrate. In other embodiments, the substrate 1180 can include other suitable types of substrates. The package assembly 1190 can be connected to other electrical devices via the package interconnect 1183. The package interconnect 1183 can be coupled to the surface of the substrate 1180 to route electrical signals to other electrical devices such as a motherboard, other chip sets, or multi-chip modules.
[0222] In some embodiments, the logic or I / O die 1174 and the memory die 1175 can be electrically coupled via a bridge 1187, which is configured to route electrical signals between the logic or I / O die 1174 and the memory die 1175. The bridge 1187 can be a dense interconnect structure that provides routing for electrical signals. The bridge 1187 can include a bridge substrate made of glass or a suitable semiconductor material. Circuit routing features can be formed on the bridge substrate to provide die-to-die connections between the logic or I / O die 1174 and the memory die 1175. The bridge 1187 can also be referred to as a silicon bridge or an interconnect bridge. For example, in some embodiments, the bridge 1187 is an Embedded Multi-Die Interconnect Bridge (EMIB). In some embodiments, the bridge 1187 can be just a direct connection from one die to another die.
[0223] Figure 11D Illustrated is a package assembly 1194 that includes interchangeable dies 1195 according to an embodiment. The interchangeable dies 1195 can be assembled into standardized slots on one or more base dies 1196, 1198. The base dies 1196, 1198 can be coupled via a bridge interconnect 1197, which can be similar to other bridge interconnects described herein and can be, for example, an EMIB. The memory die can also be connected to the logic or I / O die via a bridge interconnect. The I / O and logic dies can communicate via an interconnect fabric. The base dies can each support one or more slots in a standardized format for one of logic or I / O or memory / cache.
[0224] In one embodiment, SRAM and power delivery circuitry can be fabricated into one or more of the base dies 1196, 1198, and different process technologies can be used to fabricate the base dies 1196, 1198 relative to the interchangeable dies 1195 stacked on top of the base dies. For example, a larger process technology can be used to fabricate the base dies 1196, 1198, while a smaller process technology can be used to fabricate the interchangeable dies. One or more of the interchangeable dies 1195 can be memory (e.g., DRAM) dies. Different memory densities can be selected for the package assembly 1194 based on the power and / or performance targeted for the product using the package assembly 1194. Additionally, logic dies with different numbers of types of functional units can be selected at assembly based on the power and / or performance targeted for the product. Additionally, dies containing different types of IP logic cores can be inserted into the interchangeable die slots, enabling a hybrid processor design that can mix and match different technology IP blocks.
[0225] Exemplary System-on-Chip Integrated Circuit
[0226] Figures 12 - 13 Illustrated are exemplary integrated circuits and associated graphics processors that can be fabricated using one or more IP cores in accordance with various embodiments described herein. In addition to what is illustrated, other logic and circuitry can be included, including additional graphics processors / cores, peripheral interface controllers, or general-purpose processor cores.
[0227] Figure 12 is a block diagram illustrating an exemplary system-on-chip integrated circuit 1200 that can be fabricated using one or more IP cores in accordance with an embodiment. The exemplary integrated circuit 1200 includes one or more application processors 1205 (e.g., CPUs), at least one graphics processor 1210, and can additionally include an image processor 1215 and / or a video processor 1220, any of which processors can be modular IP cores from the same or multiple different design facilities. The integrated circuit 1200 includes peripheral or bus logic that includes a USB controller 1225, a UART controller 1230, an SPI / SDIO controller 1235, and an I2S / I2C controller 1240. Additionally, the integrated circuit can include a display device 1245 coupled to one or more of a high-definition multimedia interface (HDMI) controller 1250 and a mobile industry processor interface (MIPI) display interface 1255. Storage can be provided by a flash memory subsystem 1260 that includes flash memory and a flash memory controller. A memory interface can be provided via a memory controller 1265 to access SDRAM or SRAM memory devices. Some integrated circuits additionally include an embedded security engine 1270. As Figure 13As shown in, graphics processor 1310 includes vertex processor 1305 and one or more fragment processors 1315A - 1315N (e.g., 1315A, 1315B, 1315C, 1315D up to 1315N - 1 and 1315N). Graphics processor 1310 is capable of executing different shader programs via separate logic such that vertex processor 1305 is optimized to perform operations for vertex shader programs while one or more fragment processors 1315A - 1315N perform fragment (e.g., pixel) shading operations for fragment or pixel shader programs. Vertex processor 1305 executes the vertex processing stage of the 3D graphics pipeline and generates primitives and vertex data. The one or more fragment processors 1315A - 1315N use the primitives and vertex data generated by vertex processor 1305 to produce a frame buffer to be displayed on a display device. In one embodiment, the one or more fragment processors 1315A - 1315N are optimized to execute fragment shader programs as provided for in the OpenGL API, and the fragment shader programs can be used to perform operations similar to those of pixel shader programs as provided for in the Direct 3D API.
[0228] Graphics processor 1310 further includes one or more memory management units (MMUs) 1320A - 1320B, one or more caches 1325A - 1325B, and one or more circuit interconnects 1330A - 1330B. The one or more MMUs 1320A - 1320B provide virtual address to physical address mapping for graphics processor 1310 (including for vertex processor 1305 and / or one or more fragment processors 1315A - 1315N), and these processors can reference vertex or image / texture data stored in memory in addition to vertex or image / texture data stored in one or more caches 1325A - 1325B. In one embodiment, one or more MMUs 1320A - 1320B can be synchronized with other MMUs within the system, and the other MMUs include one or more MMUs associated with Figure 12 one or more application processors 1205, image processor 1215, and / or video processor 1220 such that each processor 1205 - 1220 can participate in a shared or unified virtual memory system. According to an embodiment, one or more circuit interconnects 1330A - 1330B enable graphics processor 1310 to interface with other IP cores within the SoC via the internal bus of the SoC or via a direct connection.
[0229] As Figure 14 shown, graphics processor 1340 includes Figure 13One or more MMUs 1320A-1320B, (one or more) caches 1325A-1325B, and (one or more) circuit interconnections 1330A-1330B of the graphics processor 1310. The graphics processor 1340 includes one or more shader cores 1355A-1355N (e.g., 1355A, 1355B, 1355C, 1355D, 1355E, 1355F up to 1355N-1 and 1355N) providing a unified shader core architecture in which a single core or a single type of core can execute all types of programmable shader code, including shader program code for implementing vertex shaders, fragment shaders, and / or compute shaders. The exact number of shader cores present can vary between embodiments and implementations. Additionally, the graphics processor 1340 includes: an inter-core task manager 1345 that acts as a thread dispatcher for dispatching execution threads to one or more shader cores 1355A-1355N; and a tiling unit 1358 for accelerating tiling operations for patch-based rendering, in which rendering operations for a scene are subdivided in image space, e.g., to take advantage of local spatial coherence within the scene or to optimize the use of internal caches. Ray tracing architecture
[0230] In one implementation, the graphics processor includes circuit modules and / or program code for performing real-time ray tracing. A dedicated set of ray tracing cores may be included in the graphics processor to perform the various ray tracing operations described herein, including ray traversal and / or ray intersection operations. In addition to the ray tracing cores, multiple sets of graphics processing cores for performing programmable shading operations and multiple sets of tensor cores for performing matrix operations on tensor data may also be included.
[0231] Figure 15 An exemplary portion of one such graphics processing unit (GPU) 1505 is illustrated that includes a dedicated set of graphics processing resources arranged in multi-core groups 1500A-1500N. The graphics processing unit (GPU) 1505 may be a variant of the following disclosed herein: graphics processor 300, GPGPU 1340, and / or any other graphics processor. Thus, the disclosure of any feature of a graphics processor also discloses the corresponding combination with the GPU 1505, but is not limited thereto. Additionally, having the same or similar names as the elements of any other figure herein Figure 15The components described herein are the same as those in other figures, can operate or function in a manner similar to the same components in other figures, can include the same components, and can be linked to other entities (such as those described elsewhere herein, but not limited thereto). Although details are provided for only a single multi-core group 1500A, it will be appreciated that other multi-core groups 1500B - 1500N can be equipped with the same or similar sets of graphics processing resources.
[0232] As illustrated, the multi-core group 1500A can include a set of graphics cores 1530, a set of tensor cores 1540, and a set of ray tracing cores 1550. The scheduler / dispatcher 1510 schedules and dispatches graphics threads for execution on the various cores 1530, 1540, 1550. A set of register files 1520 stores the operand values used by the cores 1530, 1540, 1550 when executing graphics threads. These registers can include, for example, integer registers for storing integer values, floating-point registers for storing floating-point values, vector registers for storing packed data elements (integer and / or floating-point data elements), and tile registers for storing tensor / matrix values. The tile registers can be implemented as a combined set of vector registers.
[0233] One or more level 1 (L1) caches and texture units 1560 locally store graphics data, such as texture data, vertex data, pixel data, ray data, bounding volume data, etc., within each multi-core group 1500A. A level 2 (L2) cache 1580, shared by all or a subset of the multi-core groups 1500A - 1500N, stores graphics data and / or instructions for multiple concurrent graphics threads. As illustrated, the L2 cache 1580 can be shared across multiple multi-core groups 1500A - 1500N. One or more memory controllers 1570 couple the GPU 1505 to a memory 1598, which can be system memory (e.g., DRAM) and / or local graphics memory (e.g., GDDR6 memory).
[0234] The Input / Output (IO) circuit module 1595 couples the GPU 1505 to one or more IO devices 1595, such as a Digital Signal Processor (DSP), a network controller, or a user input device. An on-chip interconnect can be used to couple the I / O devices 1590 to the GPU 1505 and the memory 1598. One or more IO Memory Management Units (IOMMUs) 1575 of the IO circuit module 1595 couple the IO devices 1590 directly to the system memory 1598. The IOMMU 1575 can manage multiple sets of page tables to map virtual addresses to physical addresses in the system memory 1598. Additionally, the IO devices 1590, the (one or more) CPUs 1599, and the (one or more) GPUs 1505 can share the same virtual address space.
[0235] The IOMMU 1575 can also support virtualization. In this case, it can manage a first set of page tables to map guest / graphics virtual addresses to guest / graphics physical addresses, and manage a second set of page tables to map guest / graphics physical addresses to system / host physical addresses (e.g., within the system memory 1598). The base addresses of each of the first and second sets of page tables can be stored in control registers and swapped out during a context switch (e.g., to provide access to the relevant set of page tables for the new context). Although not illustrated in Figure 15 each of the cores 1530, 1540, 1550, and / or the multi-core groups 1500A - 1500N can include a Translation Lookaside Buffer (TLB) to cache guest virtual to guest physical translations, guest physical to host physical translations, and guest virtual to host physical translations.
[0236] The CPU 1599, the GPU 1505, and the IO devices 1590 can be integrated on a single semiconductor chip and / or chip package. The illustrated memory 1598 can be integrated on the same chip or can be coupled to the memory controller 1570 via an off-chip interface. In one implementation, the memory 1598 includes GDDR6 memory that shares the same virtual address space as other physical system-level memories, although the basic principles of the present invention are not limited to this particular implementation.
[0237] The tensor core 1540 may include multiple execution circuits specifically designed to perform matrix operations, which are fundamental computational operations for performing deep learning operations. For example, simultaneous matrix multiplication operations can be used for neural network training and inference. The tensor core 1540 can perform matrix processing using various operand precisions, including single-precision floating point (e.g., 32 bits), half-precision floating point (e.g., 16 bits), integer (16 bits), byte (8 bits), and nibble (4 bits). Neural network implementations can also capture the characteristics of each rendered scene, potentially combining details from multiple frames to construct a high-quality final image.
[0238] In deep learning implementations, parallel matrix multiplication work can be scheduled for execution on the tensor core 1540. Training of neural networks particularly requires a large number of matrix dot product operations. To handle the inner product formula for N×N×N matrix multiplication, the tensor core 1540 can include at least N dot product processing elements. Before matrix multiplication begins, a complete matrix is loaded into the tile register, and in each of the N cycles, at least one column of the second matrix is loaded. In each cycle, there are N dot products being processed.
[0239] Depending on the specific implementation, matrix elements can be stored at different precisions, including 16-bit words, 8-bit bytes (e.g., INT8), and 4-bit nibbles (e.g., INT4). Different precision modes can be specified for the tensor core 1540 to ensure that the most efficient precision is used for different workloads (e.g., inference workloads that can tolerate quantization to bytes and nibbles).
[0240] The ray tracing core 1550 can be used to accelerate ray tracing operations for both real-time and non-real-time ray tracing implementations. In particular, the ray tracing core 1550 can include a ray traversal / intersection circuit module for performing ray traversal using a bounding volume hierarchy (BVH) and identifying intersections between primitives enclosed within the BVH volume and the ray. The ray tracing core 1550 can also include a circuit module for performing depth testing and culling (e.g., using a Z-buffer or similar arrangement). In one implementation, the ray tracing core 1550 works in conjunction with the image denoising techniques described herein to perform traversal and intersection operations, at least a portion of which can be performed on the tensor core 1540. For example, the tensor core 1540 can implement a deep learning neural network to perform denoising of frames generated by the ray tracing core 1550. However, the (one or more) CPUs 1599, graphics core 1530, and / or ray tracing core 1550 can also implement all or a portion of the denoising and / or deep learning algorithms.
[0241] Additionally, as described above, a distributed approach for denoising can be employed, where the GPU 1505 is in a computing device coupled to other computing devices via a network or high-speed interconnect. Additionally, the interconnected computing devices can share neural network learning / training data to improve the speed at which the overall system learns to perform denoising for different types of image frames and / or different graphics applications.
[0242] The ray tracing core 1550 can handle all BVH traversals and ray-primitive intersections, thus freeing the graphics core 1530 from being overloaded with thousands of instructions per ray. Each ray tracing core 1550 can include a first set of dedicated circuit modules for performing bounding box tests (e.g., for traversal operations) and a second set of dedicated circuit modules for performing ray-triangle intersection tests (e.g., for intersections of traversed rays). Thus, the multi-core group 1500A is able to simply launch a ray probe, and the ray tracing core 1550 independently performs ray traversal and intersection and returns hit data (e.g., hit, miss, multiple hits, etc.) to the thread context. While the ray tracing core 1550 performs traversal and intersection operations, the other cores 1530, 1540 can be freed up to perform other graphics or computing work.
[0243] Each ray tracing core 1550 can include a traversal unit for performing BVH test operations and an intersection unit for performing ray-primitive intersection tests. The intersection unit can then generate a "hit", "miss", or "multiple hits" response, and the intersection unit provides the response to the appropriate thread. During traversal and intersection operations, the execution resources of other cores (e.g., the graphics core 1530 and the tensor core 1540) can be freed up to perform other forms of graphics work.
[0244] A hybrid rasterization / ray tracing method can also be used, where work is distributed between the graphics core 1530 and the ray tracing core 1550.
[0245] The ray tracing core 1550 (and / or other cores 1530, 1540) can include hardware support for ray tracing instruction sets such as Microsoft's DirectX Ray Tracing (DXR) and ray generation, closest hit, any hit, and miss shaders (which enable the assignment of a unique set of textures and shaders to each object), where the DXR includes a DispatchRays command. Another ray tracing platform that can be supported by the ray tracing core 1550, the graphics core 1530, and the tensor core 1540 is Vulkan 1.1.85. However, note that the underlying principles of the present invention are not limited to any particular ray tracing ISA.
[0246] Generally, various cores 1550, 1540, 1530 can support a ray tracing instruction set, and the ray tracing instruction set includes instructions / functions for ray generation, closest hit, any hit, ray-primitive intersection, per-primitive and hierarchical bounding box construction, miss, access, and exception. More specifically, ray tracing instructions may be included to perform the following functions:
[0247] Ray generation – Ray generation instructions can be executed for each pixel, sample, or other user-defined work assignment.
[0248] Closest hit – The closest hit instruction can be executed to locate the closest intersection of a ray with a primitive in the scene.
[0249] Any hit – The any hit instruction identifies multiple intersections between primitives in the scene and a ray, potentially identifying a new closest intersection.
[0250] Intersection – The intersection instruction performs a ray-primitive intersection test and outputs the result.
[0251] Per-primitive bounding box construction – This instruction constructs a bounding box around a given primitive or group of primitives (e.g., when building a new BVH or other acceleration data structure).
[0252] Miss – Indicates that the ray misses all geometries in the scene or a specified region of the scene.
[0253] Access – Indicates the sub-volumes that the ray will traverse.
[0254] Exception – Includes various types of exception handlers (e.g., invoked for various error conditions). Graphics processor with hardware-accelerated hybrid ray tracing
[0255] Next, a hybrid rendering pipeline is presented, which performs rasterization on the graphics core 1530 and ray tracing operations on the ray tracing cores 1550, graphics core 1530, and / or CPU 1599 cores. For example, rasterization and depth testing can be performed on the graphics core 1530 to replace the primary ray casting stage. The ray tracing core 1550 can then generate secondary rays for ray reflection, refraction, and shadow. Additionally, certain regions of the scene will be selected in which the ray tracing core 1550 will perform ray tracing operations (e.g., based on a material property threshold such as a high reflectivity level), while other regions of the scene will be rendered using rasterization on the graphics core 1530. This hybrid implementation can be used for real-time ray tracing applications – where latency is a critical issue.
[0256] The ray traversal architecture described below can perform programmable shading and control of ray traversal using, for example, existing single instruction multiple data (SIMD) and / or single instruction multiple thread (SIMT) graphics processors, while using dedicated hardware to accelerate key functions such as BVH traversal and / or intersection. The SIMD occupancy for incoherent paths can be increased by regrouping derived shaders at specific points during traversal and before shading. This is achieved using dedicated hardware for dynamic classification of shaders on-chip. Recursion is managed by splitting functions into continuations that execute upon return and regrouping the continuations before execution to increase SIMD occupancy.
[0257] Programmable control of ray traversal / intersection is achieved by decomposing the traversal function into an inner traversal that can be implemented as fixed function hardware and an outer traversal that executes on the GPU processor and enables programmable control via user-defined traversal shaders. The cost of transferring traversal context between hardware and software is reduced by conservatively truncating the inner traversal state during the transition between inner and outer traversals.
[0258] Programmable control of ray tracing can be expressed by different shader types listed in Table A below. There can be multiple shaders for each type. For example, each material can have a different hit shader. Shader Type Function Main Launch Primary Ray Hit Bidirectional Reflectance Distribution Function (BRDF) Sampling, Launch Secondary Ray Any Hit Calculate Transmittance of Alpha Texture Geometry Miss Calculate Radiation from Light Source Cross Cross Custom Shape Traverse Instance Selection and Transformation Callable General Function Table A
[0259] Recursive ray tracing can be initiated by an API function that commands the graphics processor to start a set of primary shaders or intersection circuit modules that can derive ray-scene intersections for primary rays. This in turn derives other shaders such as traversal, hit shaders, or miss shaders. The shader that derives a child shader can also receive a return value from that child shader. Callable shaders are general functions that can be directly derived by another shader and can also return a value to the calling shader.
[0260] Figure 16 A graphics processing architecture is illustrated that includes a shader execution circuit module 1600 and a fixed function circuit module 1610. The general execution hardware subsystem includes multiple single instruction multiple data (SIMD) and / or single instruction multiple thread (SIMT) cores 1601 (i.e., each core can include a group of execution circuit modules), one or more samplers 1602, and a level 1 (L1) cache 1603 or other form of local memory. The fixed function hardware subsystem 1610 includes a message unit 1604, a scheduler 1607, a ray-BVH traversal / intersection circuit module 1605, a classification circuit module 1608, and a local L1 cache 1606.
[0261] In operation, the main dispatcher 1609 dispatches a set of primary rays to the scheduler 1607, which dispatches work to be executed on the SIMD / SIMT graphics processor core blocks 1601. The SIMD graphics processor core blocks 1601 can be the ray tracing cores 1550 and / or the graphics cores 1530 described above. Execution of the main shader spawns additional work to be executed (e.g., by one or more sub-shaders and / or fixed function hardware). The message unit 1604 distributes the work spawned by the SIMD graphics processor core blocks 1601 to the scheduler 1607, the binning circuit module 1608, or the ray-BVH intersection circuit module 1605 as needed to access the free stack pool. If the additional work is sent to the scheduler 1607, it is scheduled for processing on the SIMD / SIMT graphics processor core blocks 1601. Before scheduling, the binning circuit module 1608 can bin the rays into groups or bins (e.g., group rays with similar characteristics) as described herein. The ray-BVH intersection circuit module 1605 performs ray intersection tests using the BVH volume. For example, the ray-BVH intersection circuit module 1605 can compare the ray coordinates to each level of the BVH to identify the volumes intersected by the ray.
[0262] Shaders can be referenced using shader records, user-assigned structures (which include pointers to entry functions), vendor-specific metadata, and global arguments to the shaders executed by the SIMD graphics processor core blocks 1601. Each execution instance of a shader is associated with a call stack, which can be used to store arguments passed between the parent shader and the sub-shaders. The call stack can also store references to continuation functions to be executed upon call return.
[0263] Figure 17 An example set of assigned stacks 1701 is illustrated, which includes a main shader stack, a hit shader stack, a traversal shader stack, a continuation function stack, and a ray-BVH intersection stack (which can be executed by the fixed function hardware 1610 as described). New shader invocations can implement new stacks from the free stack pool 1702. The call stacks (e.g., the stacks included in the set of assigned stacks) can be cached in the local L1 caches 1603, 1606 to reduce access latency.
[0264] There can be a finite number of call stacks, each having a fixed maximum size of "Sstack" allocated in a contiguous region of memory. Thus, the base address of a stack can be directly calculated from the stack index (SID) as base address = SID * Sstack. The stack ID can be allocated and deallocated by the scheduler 1607 when dispatching work to the SIMD graphics processor core blocks 1601.
[0265] The main dispatcher 1609 may include a graphics processor command processor that dispatches main shaders in response to dispatch commands from a host (e.g., a CPU). If the scheduler 1607 can assign stack IDs to each SIMD channel, the scheduler 1607 may receive these dispatch requests and initiate main shaders on SIMD processor threads. The stack IDs may be allocated from the free stack pool 1702, which is initialized at the start of the dispatch command.
[0266] An execution shader may spawn a child shader by sending a spawn message to the message passing unit 1604. The command includes the stack ID associated with the shader and also includes pointers to the child shader records for each active SIMD channel. The parent shader may post this message only once for the active channels. After sending the spawn messages to all relevant channels, the parent shader may terminate.
[0267] Shaders executing on the SIMD graphics processor core block 1601 may also spawn fixed function tasks, such as ray-BVH intersections, using a spawn message with a shader record pointer reserved for fixed function hardware. As mentioned, the message passing unit 1604 sends the spawned ray-BVH intersection work to the fixed function ray-BVH intersection circuit module 1605 and sends callable shaders directly to the sorting circuit module 1608. The sorting circuit module may export SIMD batches with similar characteristics by grouping shaders via shader record pointers. Thus, stack IDs from different parent shaders may be grouped by the sorting circuit module 1608 in the same batch. The sorting circuit module 1608 sends the grouped batches to the scheduler 1607, which accesses the shader records from the graphics memory 1611 or the last level cache (LLC) 1620 and initiates the shaders on processor threads.
[0268] A continuation may be considered a callable shader and may also be referenced via a shader record. When a child shader is spawned and returns a value to the parent shader, a pointer to the continuation shader record may be pushed onto the call stack 1701. When the child shader returns, the continuation shader record may then be popped from the call stack 1701 and the continuation shader may be spawned. Optionally, the spawned continuation may pass through a sorting unit similar to the callable shaders and be initiated on a processor thread.
[0269] As in Figure 18As shown in the figure, the classification circuit module 1608 groups the derived tasks through shader record pointers 1801A, 1801B, 1801n to create SIMD batches for shading. The stack ID or context ID in the classified batches can be grouped according to different dispatches and different input SIMD channels. The grouping circuit module 1810 can perform classification using a content addressable memory (CAM) structure 1801, which includes a plurality of entries, and each entry is identified by a label 1801. As mentioned, the label 1801 can be the corresponding shader record pointer 1801A, 1801B, 1801n. The CAM structure 1801 can store a limited number of labels (e.g., 32, 64, 128, etc.), and each label is associated with an incomplete SIMD batch corresponding to the shader record pointer.
[0270] For an incoming derived command, each SIMD channel has a corresponding stack ID (shown as 16 context IDs 0 - 15 in each CAM entry) and shader record pointers 1801A - B, …, n (serving as label values). The grouping circuit module 1810 can compare the shader record pointer of each channel with the label 1801 in the CAM structure 1801 to find a matching batch. If a matching batch is found, the stack ID / context ID can be added to the batch. Otherwise, a new entry with a new shader record pointer label can be created, and the older entry may be evicted with the incomplete batch.
[0271] When the call stack is empty, the execution shader can deallocate the call stack by sending a deallocation message to the message unit. The deallocation message is relayed to the scheduler, and the scheduler returns the stack ID / context ID to the free pool for the active SIMD channels.
[0272] A hybrid method for ray traversal operations using a combination of fixed - function ray traversal and software ray traversal is presented. Thus, it provides the flexibility of software traversal while maintaining the efficiency of fixed - function traversal. Figure 19 An acceleration structure that can be used for hybrid traversal is shown. The acceleration structure is a two - level tree with a single top - level BVH 1900 and several bottom - level BVHs 1901 and 1902. Graphic elements are shown on the right to indicate the inner traversal path 1903, the outer traversal path 1904, traversal nodes 1905, leaf nodes 1906 with triangles, and leaf nodes 1907 with custom primitives.
[0273] A leaf node 1906 with a triangle in the top-level BVH 1900 can reference a triangle, a cross-shader record for a custom primitive, or a traversal shader record. A leaf node 1906 with a triangle in the bottom-level BVHs 1901 - 1902 can only reference a triangle and a cross-shader record for a custom primitive. The type of reference is encoded within the leaf node 1906. An inner traversal 1903 refers to the traversal within each BVH 1900 - 1902. Inner traversal operations include the calculation of ray-BVH intersections and the traversal across the BVH structures 1900 - 1902 is called an outer traversal. Inner traversal operations can be efficiently implemented in fixed-function hardware, while outer traversal operations can be performed with acceptable performance using programmable shaders. Thus, a fixed-function circuit module 1610 can be used to perform inner traversal operations and a shader execution circuit module 1600 can be used to perform outer traversal operations, where the shader circuit module 1600 includes a SIMD / SIMT graphics processor core block 1601 for executing programmable shaders.
[0274] Note that, for simplicity, the SIMD / SIMT graphics processor core block 1601 is sometimes referred to herein simply as a "core", "SIMD core", "graphics core", or "SIMD processor". Similarly, the ray-BVH traversal / crossing circuit module 1605 is sometimes referred to simply as a "traversal unit", "traversal / crossing unit", or "traversal / crossing circuit module". When using the replacement terms, the specific name used to denote the corresponding circuit module / logic does not change the underlying function performed by the circuit module / logic, as described herein.
[0275] Furthermore, although illustrated as a single component in Figure 16 for explanatory purposes, the traversal / crossing unit 1605 can include different traversal units and separate crossing units, each of which can be implemented in a circuit module and / or logic as described herein.
[0276] When a ray intersects a traversal node during an inner traversal, a traversal shader can be derived. The classification circuit module 1608 can group these shaders into SIMD batches via shader record pointers 1801A - B, …, n, which are launched by a scheduler 1607 for SIMD execution on the graphics SIMD graphics processor core block 1601. The traversal shaders can modify the traversal in several ways to enable a wide range of applications. For example, the traversal shaders can select a BVH at a coarser level of detail (LOD) or transform the ray to enable rigid body transformation. The traversal shaders can then derive an inner traversal for the selected BVH.
[0277] The inner traversal calculates ray - BVH intersections by traversing the BVH and computing ray - box and ray - triangle intersections. The inner traversal is spawned in the same way as the shader by sending a message to the message - passing circuit module 1604, which relays the corresponding spawned message to the ray - BVH intersection circuit module 1605 that computes the ray - BVH intersection.
[0278] The stack of the inner traversal can be locally stored in the fixed - function circuit module 1610 (e.g., within the L1 cache 1606). When the ray intersects a leaf node corresponding to a traversal shader or an intersection shader, the inner traversal can be terminated and the inner stack truncated. The truncated stack, along with pointers to the ray and the BVH, can be written to memory at a location specified by the calling shader and then the corresponding traversal shader or intersection shader can be spawned. If the ray intersects any triangle during the inner traversal, the corresponding hit information can be provided as input arguments to these shaders, as shown in the code below. These spawned shaders can be grouped by the classification circuit module 1608 to create SIMD batches for execution.
[0279] Truncating the inner traversal stack reduces the cost of overflowing the inner traversal stack into memory. The method described in RestartTrail for Stackless BVH Traversal, High Performance Graphics (2010), pages 107 - 111 can be applied to truncate the stack into a small number of entries at the top of the stack, a 42 - bit restart trail and a 6 - bit depth value. The restart trail indicates the branches that have been taken inside the BVH and the depth value indicates the traversal depth corresponding to the last stack entry. This is sufficient information to resume the inner traversal at a later time.
[0280] The inner traversal is completed when the internal stack is empty and there are no more BVH nodes to test. In this case, an outer stack handler is spawned, which pops the top of the outer stack and resumes the traversal if the outer stack is not empty.
[0281] The outer traversal can execute the main traversal state machine and can be implemented in the program code executed by the shader execution circuit module 1600. It can spawn an inner traversal query under the following conditions: (1) when a new ray is spawned by a hit shader or the main shader; (2) when a traversal shader selects a BVH for traversal; and (3) when the outer stack handler resumes the inner traversal of a BVH.
[0282] As in Figure 20As shown in the figure, before traversing within the derivative, space is allocated on the call stack 2005 for the fixed-function circuit module 1610 to store the truncated inner stack 2010. The offsets 2003 - 2004 to the tops of the call stack and the inner stack are maintained in the traversal state 2000, which is also stored in the memory 1611. The traversal state 2000 also includes the rays in the world space 2001 and the object space 2002, as well as the hit information of the closest intersection primitive.
[0283] The traversal shader, the intersection shader, and the external stack handler are all derived from the ray - BVH intersection circuit module 1605. The traversal shader is allocated on the call stack 2005 before initiating a new inner traversal of the second-level BVH. The external stack handler is the shader responsible for updating the hit information and resuming any outstanding inner traversal tasks. The external stack handler is also responsible for spawning the hit or miss shaders when the traversal is complete. The traversal is complete when there are no outstanding inner traversal queries to be derived. When the traversal is complete and an intersection is found, the hit shader is spawned; otherwise, the miss shader is spawned.
[0284] Although the hybrid traversal scheme described above uses a two-level BVH hierarchy, any number of BVH levels with corresponding changes in the outer traversal implementation can also be implemented.
[0285] In addition, although the fixed-function circuit module 1610 is described above for performing ray - BVH intersection, other system components can also be implemented in the fixed-function circuit module. For example, the external stack handler described above can be an internal (not visible to the user) shader that can potentially be implemented in the fixed-function BVH traversal / intersection circuit module 1605. This implementation can be used to reduce the round trips between the fixed-function intersection hardware 1605 and the processor and the number of dispatched shader levels.
[0286] The examples described herein use user-defined functions that can be executed with greater SIMD efficiency on existing and future GPU processors to enable programmable shading and ray traversal control. The programmable control of ray traversal enables several important features, such as procedural instancing, random level-of-detail selection, custom primitive intersection, and lazy BVH update.
[0287] A programmable multiple instruction multiple data (MIMD) ray tracing architecture that supports speculative execution of the hit shader and the intersection shader is also provided. In particular, the architecture focuses on reducing the Figure 16The scheduling and communication overhead between the programmable SIMD / SIMT core 1601 (e.g., graphics core 1530) and the fixed-function MIMD traversal / crossing unit 1605 in the hybrid ray tracing architecture described above. Multiple speculative execution schemes for hit shaders and cross shaders are described below, which can be dispatched in a single batch from the traversal hardware, avoiding several traversal and shading round trips. Dedicated circuit modules for implementing these techniques can be used.
[0288] Embodiments of the present invention are particularly beneficial in use cases where it is desirable to execute multiple hit shaders or cross shaders based on a ray traversal query (which would incur significant overhead when implemented without dedicated hardware support). These include, but are not limited to, nearest k-hit queries (hit shaders are launched for the k closest intersections) and multiple programmable cross shaders.
[0289] The techniques described herein can be implemented as Figure 16 Figure (and about Figures 16 - 20 In particular, the present embodiment of the invention builds upon that architecture with enhancements to improve the performance of the use cases mentioned above.
[0290] The performance limitation of the hybrid ray tracing traversal architecture is the overhead of launching the traversal query from the graphics core and the overhead of calling the programmable shader from the ray tracing hardware. This overhead generates an "execution round trip" between the programmable core 1601 and the traversal / crossing unit 1605 when multiple hit shaders or cross shaders are called during the traversal of the same ray. This also puts additional pressure on the classification unit 1608 that needs to call from each shader to obtain SIMD / SIMT coherency.
[0291] Several aspects of ray tracing require programmable control, which can be expressed through the different shader types listed in Table A above (i.e., primary shader, hit shader, any hit shader, miss shader, cross shader, traversal shader, and callable shader). There can be multiple shaders of each type. For example, each material can have a different hit shader. In the current Some of these shader types are defined in the Ray Tracing API.
[0292] As a brief review, recursive ray tracing is initiated via an API function that commands a GPU to launch a set of primary shaders capable of deriving ray-scene intersections (implemented in hardware and / or software) for primary rays. This in turn can derive other shaders (such as traversal, hit, or miss shaders). The shader that derives the child shaders is also capable of receiving return values from that shader. Callable shaders are general-purpose functions that can be directly derived by another shader and can also return values to the calling shader.
[0293] Ray traversal computes ray-scene intersections by traversing and intersecting nodes in a bounding volume hierarchy (BVH). Recent research has shown that techniques better suited for fixed-function hardware (such as reduced-precision arithmetic, BVH compression, per-ray state machines, dedicated intersection pipelines, and custom caches) can be used to improve the efficiency of computing ray-scene intersections by more than an order of magnitude.
[0294] Figure 16 The architecture shown includes a system in which an array of SIMD / SIMT cores 1601 interacts with a fixed-function ray tracing / intersection unit 1605 to perform programmable ray tracing. Programmable shaders are mapped to SIMD / SIMT threads on the graphics core 1601, where SIMD / SIMT utilization, execution, and data coherence are key for optimal performance. Ray queries typically break coherence for various reasons, such as: · Traversal Divergence : The duration of BVH traversal varies highly. · Tends to asynchronous ray processing between rays. · Execution Divergence : Rays derived from different lanes of the same SIMD / SIMT thread may result in different shader invocations. · Data Access Divergence : For example, rays hitting different surfaces sample different BVH nodes and primitives, and shaders access different textures. Various other scenarios may cause data access divergence.
[0295] Figure 21 Illustrates the operation flow of a programmable ray tracing pipeline. Shading elements including traversal 2102 and intersection 2103 can be implemented in fixed-function circuit modules, while the remaining elements can be implemented using programmable graphics processor core blocks.
[0296] The primary ray shader 2101 sends work to a traversal circuit module 2102 that traverses the current ray(s) through a BVH (or other acceleration structure). When a leaf node is reached, the traversal circuit module calls an intersection circuit module 2103, which invokes any hit shader 2104 when a ray-triangle intersection is identified (as indicated, the any hit shader can provide results back to the traversal circuit module).
[0297] Alternatively, the traversal can terminate before reaching a leaf node and the closest hit shader 2107 (if a hit is recorded) or miss shader 2106 (in the case of a miss) called at 2107.
[0298] As indicated at 2105, if the traversal circuit module reaches a custom primitive leaf node, a cross shader can be invoked. The custom primitive can be any non-triangle primitive (such as a polygon or polyhedron (e.g., tetrahedron, voxel, hexahedron, wedge, pyramid, or other "unstructured" volume)). The cross shader 2105 identifies any intersection between the ray and the custom primitive to any hit shader 2104 that implements any hit processing.
[0299] When the hardware traversal 2102 reaches the programmable stage, the traversal / intersection unit 1605 can generate shader dispatch messages for the relevant shaders 2105 - 2107 that correspond to a single SIMD lane of a graphics processor core block for executing the shaders. Since the dispatches occur in an arbitrary order per ray and are divergent within the called program, the sorting unit 1608 can accumulate multiple dispatch calls to obtain coherent SIMD batches. The updated traversal state and optional shader arguments can be written by the traversal / intersection unit 1605 to the memory 1611.
[0300] In the k-closest intersection problem, the closest hit shader 2107 is executed for the first k intersections. In a conventional manner, this would mean that the ray traversal ends when the closest intersection is found, the hit shader is invoked, and a new ray is derived from the hit shader to find the next closest intersection (with a ray origin offset so the same intersection does not occur again). It is easy to see that this implementation would require k ray derivations for a single ray. Another implementation operates using any hit shader 2104 that is invoked for all intersections and maintains a global list of the closest intersections, using an insertion sort operation. The main problem with this method is that there is no upper bound on the number of any hit shader invocations.
[0301] As mentioned, the intersection shader 2105 can be invoked on non-triangle (custom) primitives. Depending on the results of the intersection tests and traversal state (pending nodes and primitive intersections), traversal of the same ray can continue after execution of the intersection shader 2105. Thus, finding the closest hit can require several round trips to the graphics processor core block.
[0302] Focus can also be placed on reducing the SIMD-MIMD context switches for the intersection shader 2105 and the hit shaders 2104, 2107 by changing the traversal hardware and shader scheduling model. First, the ray traversal circuit module 1605 reduces shader invocation latency by accumulating multiple potential invocations and dispatching them in larger batches. Additionally, certain invocations that prove unnecessary can be culled at this stage. Further, the shader scheduler 1607 can aggregate multiple shader invocations from the same traversal context into a single SIMD batch, which results in a single-ray spawn message. In an exemplary implementation, the traversal hardware 1605 pauses the traversal thread and waits for the results of multiple shader invocations. Since this mode of operation allows multiple shaders to be dispatched, some of these shaders may not be invoked when using sequential invocations, and thus this mode of operation is referred to herein as "speculative" shader execution.
[0303] Figure 22A Illustrates an example where the traversal operation encounters multiple custom primitives 2250 in a subtree, and Figure 22B Illustrates how this can be resolved using three intersection dispatch cycles C1 - C3. In particular, the scheduler 1607 may require three cycles to submit work to the SIMD processor 1601, and the traversal circuit module 1605 requires three cycles to provide the results to the classification unit 4008. The traversal state 2201 required by the traversal circuit module 1605 can be stored in a memory (such as, a local cache (e.g., L1 cache and / or L2 cache)). A. Delayed Ray Tracing Shader Invocation
[0304] The manner in which the hardware traversal state 2201 is managed can also be modified to allow multiple potential intersection or hit invocations to accumulate in a list. At a given time during traversal, each entry in the list can be used to generate a shader invocation. For example, k-closest intersections can be accumulated on the traversal hardware 1605 and / or in the traversal state 2201 in memory, and if traversal is completed, the hit shader can be invoked for each element. For the hit shader, multiple potential intersections can be accumulated for a subtree in the BVH.
[0305] For the nearest-k use case, the benefit of this method is that instead of k-1 round trips and k-1 new ray derivation messages to the SIMD core / graphics core 1601, all hit shaders are fetched from the same traversal thread during a single traversal operation on the traversal circuit module 1605. The challenge for a potential implementation is that ensuring the execution order of the hit shaders is not straightforward (the standard "round-trip" method ensures that the hit shader closest to the intersection is executed first, etc.). This can be addressed by relaxing the synchronization or ordering of the hit shaders.
[0306] For the intersection shader use case, the traversal circuit module 1605 does not know in advance whether a given shader will return a positive intersection test. However, it is possible to speculatively execute multiple intersection shaders and, if at least one returns a positive hit result, merge it into the global nearest hit. The specific implementation needs to find the optimal number of deferred intersection tests to reduce the number of dispatch calls, but avoid calling too many redundant intersection shaders. B. Aggregate Shader Invocations from Traversal Circuit Module
[0307] When dispatching multiple shaders from the same ray derivation on the traversal circuit module 1605, branches can be created in the flow of the ray traversal algorithm. This can be problematic for intersection shaders because the remainder of the BVH traversal depends on the results of all dispatched intersection tests. This means that synchronization operations are necessary to wait for the results of the shader fetches, which can be challenging on asynchronous hardware.
[0308] Two points where the results of shader calls can be merged are: the SIMD processor 1601 and the traversal circuit module 1605. Regarding the SIMD processor 1601, multiple shaders can use the standard programming model to synchronize and aggregate their results. A relatively simple way to do this is to use global atoms and aggregate the results in a shared data structure in memory where the intersection results of multiple shaders can be stored. Then, the last shader can parse the data structure and call back to the traversal circuit module 1605 to continue the traversal.
[0309] A more efficient method can also be implemented to limit the execution of multiple shader fetches to the lanes of the same SIMD thread on the SIMD processor 1601. Then, SIMD / SIMT reduction operations (instead of relying on global atoms) are used to locally reduce the intersection tests. This implementation may rely on a new circuit module within the classification unit 1608 to keep small batches of shader fetches within the same SIMD batch.
[0310] The execution of the traversal thread on the traversal circuit module 1605 can also be paused. Using a conventional execution model, when a shader is dispatched during traversal, the traversal thread is terminated and the ray traversal state is saved to memory to allow other ray dispatch commands to be executed while the graphics processor core block 1601 processes the shader. If only the traversal thread is paused, the traversal state does not need to be stored and can wait for each shader result individually. This implementation can include circuit modules for avoiding deadlocks and providing sufficient hardware utilization.
[0311] Figures 23 - 24 The figure illustrates an example of a latency model that uses three shaders 2301 to fetch the latency of a single shader fetch on the SIMD core 1601. All cross-tests are evaluated within the same SIMD / SIMT group when saved. As a result, the nearest cross can also be calculated on the programmable core 1601.
[0312] As mentioned, all or part of the shader aggregation and / or latency can be performed by the traversal / crossover circuit module 1605 and / or the core / graphics core scheduler 1607. Figure 23 The figure illustrates how the shader latency / aggregator circuit module 2306 within the scheduler 1607 can delay the scheduling of a shader associated with a particular SIMD / SIMT thread / channel until a specified trigger event has occurred. When the trigger event is detected, the scheduler 1607 dispatches multiple aggregated shaders in a single SIMD / SIMT batch to the graphics processor core block 1601.
[0313] Figure 24 The figure illustrates how the shader latency / aggregator circuit module 2405 within the traversal / crossover circuit module 1605 can delay the scheduling of a shader associated with a particular SIMD thread / channel until a specified trigger event has occurred. When the trigger event is detected, the traversal / crossover circuit module 1605 submits the aggregated shaders to the sorting unit 1608 in a single SIMD / SIMT batch.
[0314] However, note that the shader latency and aggregation techniques can be implemented within various other components (such as the sorting unit 1608) or can be distributed across multiple components. For example, the traversal / crossover circuit module 1605 can perform a first set of shader aggregation operations, and the scheduler 1607 can perform a second set of shader aggregation operations to ensure efficient scheduling of shaders for SIMD threads on the graphics processor core block 1601.
[0315] The "trigger event" that causes the aggregation shader to be dispatched to a graphics processor core block can be a processing event, such as the minimum latency associated with a particular thread or a particular number of accumulated shaders. Alternatively or additionally, the trigger event can be a time event, such as a certain duration since the latency of the first shader or a particular number of processor cycles. Other variables (such as the current workload on the graphics processor core block 1601 and the traversal / intersection unit 1605) can also be evaluated by the scheduler 1607 to determine when to dispatch a SIMD / SIMT batch of shaders.
[0316] Different embodiments of the present invention can be implemented using different combinations of the above methods based on the specific system architecture used and the requirements of the application. Ray tracing instructions
[0317] The following ray tracing instructions are included in the instruction set architecture (ISA) supported by the CPU 1599 and / or the GPU 1505. If executed by the CPU, single instruction multiple data (SIMD) instructions can use vector / packed source and destination registers to perform the described operations and can be decoded and executed by a CPU core. If executed by the GPU 1505, the instructions can be executed by the graphics core 1530. For example, any of the above-described graphics processor core blocks 1601 can execute the instructions. Alternatively or additionally, the instructions can be executed by the execution circuit modules on the ray tracing core 1550 and / or the tensor core 1540.
[0318] Figure 25 An architecture for executing ray tracing instructions is illustrated. The illustrated architecture can be integrated within one or more of the cores 1530, 1540, 1550 described above that can be included in different processor architectures (see, for example Figure 15 and the associated text).
[0319] In operation, the instruction fetch unit 2503 fetches the ray tracing instruction 2500 from the memory 1598, and the decoder 2595 decodes the instruction. In one implementation, the decoder 2595 decodes the instruction to generate executable operations (e.g., micro-operations or uops in a microcoded core). Alternatively, some or all of the ray tracing instructions 2500 can be executed without decoding, and in such a case, the decoder 2504 is not required.
[0320] In any implementation, the scheduler / dispatcher 2505 schedules and dispatches instructions (or operations) across a set of functional units (FUs) 2510 - 2512. The illustrated implementation includes: a vector FU 2510 for executing single instruction multiple data (SIMD) instructions that operate on multiple packed data elements stored in vector register 2515, and a scalar FU 2511 for operating on scalar values stored in one or more scalar registers 2516. An optional ray tracing FU 2512 can operate on the packed data values stored in vector register 2515 and / or the scalar values stored in scalar register 2516. In implementations without the dedicated FU 2512, the vector FU 2510 and possibly the scalar FU 2511 can execute the ray tracing instructions described below.
[0321] The various FUs 2510 - 2512 access the ray tracing data 2502 (e.g., traversal / intersection data) required to execute the ray tracing instructions 2500 from the vector register 2515, scalar register 2516, and / or the local cache subsystem 2508 (e.g., L1 cache). The FUs 2510 - 2512 can also perform accesses to the memory 1598 via load and store operations, and the cache subsystem 2508 can operate independently to cache data locally.
[0322] While ray tracing instructions can be used to improve the performance of ray traversal / intersection and BVH construction, ray tracing instructions can also be applicable to other areas such as high performance computing (HPC) and general purpose GPU (GPGPU) implementations.
[0323] In the following description, the term double word is sometimes abbreviated as dw, and unsigned byte is abbreviated as ub. Additionally, the source and destination registers mentioned below (e.g., src0, src1, dest, etc.) can refer to the vector register 2515, or in some cases a combination of the vector register 2515 and scalar register 2516. Generally, if the source or destination values used by an instruction include packed data elements (e.g., where the source or destination stores N data elements), the vector register 2515 is used. Other values can use the scalar register 2516 or the vector register 2515. Dequantization
[0324] An example of a dequantization instruction "dequantizes" a previously quantized value. For example, in a ray tracing implementation, certain BVH subtrees can be quantized to reduce storage and bandwidth requirements. The dequantization instruction can take the form dequantize dest src0src1 src2, where source register src0 stores N unsigned bytes, source register src1 stores 1 unsigned byte, source register src2 stores 1 floating-point value, and destination register dest stores N floating-point values. All of these registers can be vector registers 2515. Alternatively, src0 and dest can be vector registers 2515, and src1 and src2 can be scalar registers 2516.
[0325] The following code sequence defines a specific implementation of the dequantization instruction: In this example, ldexp multiplies a double-precision floating-point value by a specified integer power of two (i.e., ldexp(x,exp) = x * 2 exp )). In the code above, if the execution mask value (execMask[i]) associated with the current SIMD data element is set to 1, the SIMD data element at position i in src0 is converted to a floating-point value, multiplied by the integer power of the value in src1 (2 src1 value ), and the value is added to the corresponding SIMD data element in src2. Selective Minimum or Maximum
[0326] The selective minimum or maximum instruction can perform a minimum or maximum operation per channel as indicated by the bits in a bitmask (i.e., return the minimum or maximum of a set of values). The bitmask can utilize a separate set of vector registers 2515, scalar registers 2516, or mask registers (not shown). The following code sequence defines a specific implementation of the minimum / maximum instruction: sel_min_max dest src0 src1 src2, where src0 stores N doublewords, src1 stores N doublewords, src2 stores one doubleword, and the destination register stores N doublewords.
[0327] The following code sequence defines a specific implementation of the selective minimum / maximum instruction: In this example, the value of (1<<i)&src2 (a 1 left-shifted by i ANDed with src2) is used to select the minimum or maximum of the i-th data element in src0 and src1. This operation is performed for the i-th data element only if the execution mask value (execMask[i]) associated with the current SIMD data element is set to 1. Shuffle Index Instruction
[0328] The shuffle index instruction can copy any set of input channels to output channels. For a SIMD width of 32, this instruction can be executed with a low throughput. This instruction takes the following form: shuffle_index dest src0 src1 <optional flags>, where src0 stores N doublewords, src1 stores N unsigned bytes (i.e., index values), and dest stores N doublewords.
[0329] The following code sequence defines a specific implementation of the shuffle index instruction:
[0330] In the code above, the index in src1 identifies the current channel. If the i-th value in the execution mask is set to 1, a check is performed to ensure that the source channel is in the range 0 to the SIMD width. If so, the flag is set to (srcLaneMod), and the i-th data element in the destination is set equal to the i-th data element in src0. If the channel is in range (i.e., is valid), the index value from src1 (srcLane 0) is used as an index into src0 (dst[i] = src0[srcLane]). Immediate Value Shuffle Up / Dn / XOR Instruction
[0331] The immediate value shuffle instruction can shuffle input data elements / channels based on the immediate value of the instruction. The immediate value can specify shifting the input channels by 1, 2, 4, 8, or 16 positions based on the value of the immediate value. Optionally, an additional scalar source register can be specified as a fill value. When the source channel index is invalid, the fill value (if provided) is stored into the data element position in the destination. If no fill value is provided, all the data element positions are set to zero.
[0332] The flag register can be used as a source mask. If the flag bit of the source channel is set to 1, the source channel can be marked as invalid, and the instruction can proceed.
[0333] The following are examples of different implementations of the immediate value shuffle instruction: shuffle_ <up dn xor>_<1 / 2 / 4 / 8 / 16>dest src0<optional src1> <optionalflag> shuffle_ <up dn xor>_<1 / 2 / 4 / 8 / 16>dest src0<optional src1> <optionalflag> In this implementation, src0 stores N doublewords, src1 stores one doubleword for the fill value (if present), and dest stores N doublewords including the result.
[0334] The following code sequence defines a specific implementation of the immediate value shuffle instruction:
[0335] Here, the input data elements / channels are displaced 1, 2, 4, 8, or 16 positions based on the value of the immediate value. Register src1 is an additional scalar source register that is used as a fill value for the data element position stored into the destination when the source channel index is invalid. If no fill value is provided and the source channel index is invalid, the data element position in the destination is set to 0. The flag register (FLAG) is used as a source mask. As described above, if the flag bit for a source channel is set to 1, the source channel is marked as invalid and the instruction proceeds. Indirect Shuffle Up / Dn / XOR Instruction
[0336] The indirect shuffle instruction has a source operand (src1) that controls the mapping from source channels to destination channels. The indirect shuffle instruction can take the following form: shuffle_ <up dn xor>dest src0 src1<optional flag> Among them, src0 stores N doublewords, src1 stores 1 doubleword, and dest stores N doublewords.
[0337] The following code sequence defines a specific implementation of the immediate value shuffle instruction:
[0338] Thus, the indirect shuffle instruction operates in a similar manner to the immediate value shuffle instruction above, but the mapping from source channels to destination channels is controlled by the source register src1 rather than by an immediate value. Cross-Channel Minimum / Maximum Instruction
[0339] Cross-channel minimum / maximum instructions are supported for both floating-point and integer data types. The cross-channel minimum instruction can take the form of lane_min dest src0, and the cross-channel maximum instruction can take the form of lane_max dest src0, where src0 stores N doublewords and dest stores 1 doubleword.
[0340] For example, the following code sequence defines a specific implementation of the cross-channel minimum: dst = src[0]; for (int i = 1; i < SIMD_WIDTH) { if (execMask[i]) { dst = min(dst, src[i]); } } In this example, the doubleword value in the data element position i of the source register is compared with the data element in the destination register, and the minimum of these two values is copied to the destination register. The cross-channel maximum instruction operates in essentially the same way, with the only difference being that the maximum of the data element in position i and the destination value is selected. Cross-Channel Minimum / Maximum Index Instruction
[0341] The cross-channel minimum index instruction can take the form of lane_min_index dest src0, and the cross-channel maximum index instruction can take the form of lane_max_index dest src0, where src0 stores N doublewords and dest stores 1 doubleword.
[0342] For example, the following code sequence defines a specific implementation of the cross-channel minimum index instruction: In this example, the destination index increments from 0 to the SIMD width across the destination registers. If the execution mask bit is set, the data element at position i in the source register is copied to the temporary storage location (tmp), and the destination index is set to the data element position i. Cross-Channel Sorting Network Instruction
[0343] The cross-channel sorting network instruction can use a (stable) sorting network of width N to sort all N input elements either in ascending order (sortnet_min) or in descending order (sortnet_max). The minimum / maximum versions of the instruction can take the form of sortnet_min dest src0 and sortnet_max dest src0, respectively. In one implementation, src0 and dest store N doublewords. The minimum / maximum sorting is performed on the N doublewords in src0, and the elements in ascending order (for minimum) or in descending order (for maximum) are stored in dest in their respective sorted order. An example of the code sequence defining the instruction is: dst = apply_N_wide_sorting_network_min / max(src0). Cross-Channel Sorting Network Index Instruction
[0344] The cross-channel sorting network index instruction can use a (stable) sorting network of width N to sort all N input elements, but returns the indices of the sequence changes either in ascending order (sortnet_min) or in descending order (sortnet_max). The minimum / maximum versions of the instruction can take the form of sortnet_min_index dest src0 and sortnet_max_index dest src0, where src0 and dest each store N doublewords. An example of the code sequence defining the instruction is dst = apply_N_wide_sorting_network_min / max_index(src0).
[0345] The method for executing any of the above instructions is illustrated in Figure 26 It can be implemented on the specific processor architectures described above, but is not limited to any particular processor or system architecture.
[0346] At 2601, the instructions of the main graphics thread are executed on the processor core. This can include, for example, any of the cores described above (e.g., graphics core 1530). When it is determined at 2602 that ray tracing work is reached within the main graphics thread, the ray tracing instructions are offloaded to the ray tracing execution circuit module, which can take, for example, the form of the one described above with respect to Figure 25 The described functional unit (FU) may be in the form of or may be located in a dedicated ray tracing core 1550 as described with respect to Figure 15 the dedicated ray tracing core 1550 described.
[0347] At 2603, the ray tracing instructions are decoded and fetched from memory, and at 2605, the instructions are decoded into executable operations (e.g., in embodiments that require a decoder). At 2604, the ray tracing instructions are scheduled and dispatched for execution by the ray tracing circuit module. At 2605, the ray tracing instructions are executed by the ray tracing circuit module. For example, the instructions may be dispatched and executed on the above-described FUs (e.g., vector FU 2510, ray tracing FU 2512, etc.) and / or on the graphics core 1530 or ray tracing core 1550.
[0348] Upon completion of the execution of the ray tracing instructions, at 2606, the results are stored (e.g., stored back to memory 1598), and at 2607, the main graphics thread is notified. At 2608, the ray tracing results are processed within the context of the main thread (e.g., read from memory and integrated into the graphics rendering results).
[0349] In an embodiment, the terms "engine" or "module" or "logic" may refer to, be part of, or include the following: an application specific integrated circuit (ASIC), an electronic circuit, a processor (shared, dedicated, or grouped), and / or a memory (shared, dedicated, or grouped) that executes one or more software or firmware programs, combinational logic circuits, and / or other suitable components that provide the described functionality. In an embodiment, the engine, module, or logic may be implemented in firmware, hardware, software, or any combination of firmware, hardware, and software. Apparatus and method for asynchronous ray tracing
[0350] Embodiments of the present invention include a combination of fixed function acceleration circuit modules and general purpose processing circuit modules to perform ray tracing. For example, certain operations related to ray traversal and intersection testing of a bounding volume hierarchy (BVH) may be performed by the fixed function acceleration circuit modules, while multiple execution circuits execute various forms of ray tracing shaders (e.g., any hit shader, intersection shader, miss shader, etc.). One embodiment includes a dual high bandwidth memory bank that includes multiple entries for storing rays and corresponding dual stacks for storing BVH nodes. In this embodiment, the traversal circuit module alternates between the dual ray bank and the stacks to process rays on each clock cycle. Additionally, one embodiment includes a priority selection circuit module / logic that differentiates between internal nodes, non-internal nodes, and primitives, and uses this information to intelligently prioritize the processing of BVH nodes and primitives bounded by BVH nodes.
[0351] One particular embodiment uses a short stack to store a limited number of BVH nodes during traversal operations to reduce the high-speed memory required for traversal. This embodiment includes stack management circuitry / logic to efficiently push entries onto and pop entries from the short stack to ensure that the required BVH nodes are available. Additionally, the traversal operation is traced by performing updates to a tracing data structure. When the traversal circuitry / logic is paused, it can query the tracing data structure to start the traversal operation at the same location within the BVH where it stopped. And the tracing data maintained in the data structure tracing is executed such that the traversal circuitry / logic can be restarted.
[0352] Figure 27 An embodiment is illustrated that includes a shader execution circuitry module 1600 for executing shader program code and processing associated ray tracing data 2502 (e.g., BVH node data and ray data), a ray tracing acceleration circuitry module 2710 for performing traversal and intersection operations, and a memory 1598 for storing the program code and associated data processed by the RT acceleration circuitry module 2710 and the shader execution circuitry module 1600.
[0353] In one embodiment, the shader execution circuitry module 1600 includes a plurality of graphics processor core blocks 1601 that execute shader program code to perform various forms of data parallel operations. For example, in one embodiment, the graphics processor core blocks 1601 can execute a single instruction across multiple channels, where each instance of the instruction operates on data stored in different channels. For example, in a SIMT implementation, each instance of the instruction is associated with a different thread. During execution, the L1 cache stores certain ray tracing data (e.g., recently or frequently accessed data) for efficient access.
[0354] A set of primary rays can be dispatched to a scheduler 1607, which schedules work to the shaders executed by the graphics processor core blocks 1601. The graphics processor core blocks 1601 can be ray tracing cores 1526, graphics cores 1530, CPU cores 1599, or other types of circuitry modules that can execute shader program code. One or more primary ray shaders 2701 process the primary rays and derive additional work to be executed by the ray tracing acceleration circuitry module 2710 and / or the graphics processor core blocks 1601 (e.g., to be executed by one or more sub-shaders). The new work derived by the primary ray shaders 2701 or other shaders (executed by the graphics processor core blocks 1601) can be distributed to a classification circuitry module 1608, which classifies the rays into groups or bins as described herein (e.g., groups rays with similar characteristics). The scheduler 1607 then schedules the new work on the graphics processor core blocks 1601.
[0355] Other shaders that can be executed include any hit shader 2114 and closest hit shader 2107 that process hit results (e.g., identifying any hit or closest hit of a given ray, respectively) as described above. The miss shader 2106 processes ray misses (e.g., cases where the ray does not intersect a node / primitive). As described above, shader records can be used to reference various shaders, and the shader records can include one or more pointers, vendor-specific metadata, and global variables. In one embodiment, the shader record is identified by a shader record identifier (SRI). In one embodiment, each execution instance of a shader is associated with a call stack 4303 that stores variables passed between the parent shader and the child shader. The call stack 2721 can also store a reference to a continuation function to be executed upon return from the call.
[0356] The ray traversal circuit module 2702 traverses each ray through the nodes of the BVH, working down the BVH hierarchy (e.g., through parent nodes, child nodes, and leaf nodes) to identify the nodes / primitives traversed by the ray. The ray-BVH intersection circuit module 2703 performs intersection tests on the ray, determines the hit points on the primitive, and generates results in response to a hit. The traversal circuit module 2702 and the intersection circuit module 2703 can retrieve work from one or more call stacks 2721. Within the ray tracing acceleration circuit module 2710, the call stack 2721 and the associated ray tracing data 2502 can be stored in a local ray tracing cache (RTC) 2707 or other local storage device for efficient access by the traversal circuit module 2702 and the intersection circuit module 2703. One particular embodiment described below includes a high-bandwidth ray library (e.g., see Figure 41A ).
[0357] The ray tracing acceleration circuit module 2710 can be a variant of the various traversal / intersection circuits described herein, which include the ray-BVH traversal / intersection circuit 1605, the traversal circuit 2102, and the intersection circuit 2103, as well as the ray tracing core 1550. The ray tracing acceleration circuit module 2710 can be used in place of the ray-BVH traversal / intersection circuit 1605, the traversal circuit 2102, and the intersection circuit 2103, as well as the ray tracing core 1550, or any other circuit module / logic for processing the BVH stack and / or performing traversal / intersection. Thus, the disclosure of any feature in combination with the ray-BVH traversal / intersection circuit 1605, the traversal circuit 2102, and the intersection circuit 2103, as well as the ray tracing core 1550 described herein also discloses the corresponding combination with the ray tracing acceleration circuit module 2710, but is not limited thereto. Apparatus and method for displacement mesh compression
[0358] One embodiment of the present invention uses ray tracing for visibility queries to perform path tracing to render photo-realistic images. In this implementation, rays are cast from a virtual camera and traced through the simulated scene. Then, random sampling is performed in order to incrementally compute the final image. The random sampling in path tracing results in noise in the rendered image, which can be removed by allowing more samples to be generated. The samples in this implementation can be color values produced by a single ray.
[0359] In one embodiment, the ray tracing operations for visibility queries rely on a bounding volume hierarchy (BVH) (or other 3D hierarchical arrangement) generated on scene primitives (e.g., triangles, quadrilaterals, etc.) during a preprocessing stage. Using the BVH, the renderer can quickly determine the closest intersection between a ray and a primitive.
[0360] When accelerating these ray queries in hardware (e.g., such as using the traversal / intersection circuit modules described herein), memory bandwidth issues may arise due to the amount of triangle data fetched. Fortunately, much of the complexity in the modeled scene is created by displacement mapping, in which a smooth base surface representation (such as a subdivision surface) is refined using subdivision rules to generate a refined mesh 2891 as shown in Figure 28A FIG. The displacement function 2892 is applied to each vertex of the refined mesh, which typically either displaces only along the geometric normal of the base surface or displaces in an arbitrary direction to generate the displaced mesh 2893. The amount of displacement added to the surface is limited in range; thus, cases where a large displacement is generated from the base surface are rare.
[0361] One embodiment of the present invention uses lossy waterproof compression to effectively compress the displacement mapping mesh. In particular, this implementation quantifies the displacement relative to a coarse base mesh, which may match the base subdivision mesh. In one embodiment, the original quadrilaterals of the base subdivision mesh can be subdivided into a grid with the same precision as the displacement mapping using bilinear interpolation.
[0362] Figure 28B FIG. illustrates a compression circuit module / logic 2800 that compresses the displacement mapping mesh 2802 to generate a compressed displacement mesh 2810 according to the embodiments described herein. In the illustrated embodiment, the displacement mapping circuit module / logic 2811 generates the displacement mapping mesh 2802 from the base subdivision surface. Figure 29A FIG. illustrates an example in which a primitive surface 2900 is refined to generate a base subdivision surface 2901. A displacement function is applied to the vertices of the base subdivision surface 2901 to create a displacement mapping 2902.
[0363] Returning to Figure 28B , in one embodiment, the quantizer 2812 quantizes the displacement map grid 2802 relative to the coarse base grid 2803 to generate a compressed displacement grid 2810, the compressed displacement grid 2810 including a 3D displacement array 2804 and base coordinates 2805 associated with the coarse base grid 2803. By way of example and not limitation, Figure 29B Illustrates a set of difference vectors d1-d4 2922, each difference vector associated with a different displacement vertex v1-v4.
[0364] In one embodiment, the coarse base grid 2903 is the base subdivision grid 2901. Alternatively, the interpolator 2821 uses bilinear interpolation to subdivide the original quadrilaterals of the base subdivision grid into a grid having the same precision as the displacement map.
[0365] The quantizer 2812 determines the difference vectors d1-d4 2922 from each coarse base vertex to the corresponding displacement vertices v1-v4 and combines the difference vectors 2922 in the 3D displacement array 2804. In this way, the displacement grid is defined using only the coordinates of the quadrilaterals (base coordinates 2805) and the 3D displacement vector array 2804. Note that these 3D displacement vectors 2804 do not necessarily match the displacement vectors used to compute the original displacement 2902, because modeling tools will generally not use bilinear interpolation to subdivide quadrilaterals, but rather apply more complex subdivision rules to create a smooth surface to be displaced.
[0366] As Figure 29C illustrated, the grids of two adjacent quadrilaterals 2990-2991 will stitch together seamlessly because along the boundary 2992, the two quadrilaterals 2990-2991 will evaluate to exactly the same vertex positions v5-v8. Since the displacements stored along the edge 2992 of the adjacent quadrilaterals 2990-2991 are also the same, the displaced surface will not have any cracks. This property is important because it particularly means that for the entire grid, the precision of the stored displacements can be arbitrarily reduced, resulting in a lower quality connected displacement grid.
[0367] In one embodiment, half-precision floating point numbers are used to encode displacements (e.g., 16-bit floating point values). Alternatively or additionally, a shared exponent representation is used, which stores only one exponent for all three vertex components and three mantissas. Further, since the range of displacements is typically well-bounded, the displacements of a grid can be encoded using fixed-point coordinates scaled by some constant to obtain sufficient range to encode all displacements. While one embodiment of the present invention uses bilinear patches as the basic primitive, using only planar triangles, another embodiment uses triangle pairs to process each quadrilateral.
[0368] Figure 30 The figure illustrates a method according to an embodiment of the present invention. The method can be implemented on the architectures described herein, but is not limited to any particular processor or system architecture.
[0369] At 3001, a displacement map grid is generated from a base subdivision surface. For example, a primitive surface can be finely refined to generate a base subdivision surface. At 3002, a base grid (e.g., such as a base subdivision grid in one embodiment) is generated or identified.
[0370] At 3003, a displacement function is applied to the vertices of the base subdivision surface to create a 3D displacement array of difference vectors. At 3004, base coordinates associated with the base grid are generated. As described above, the base coordinates can be used in combination with the difference vectors to reconstruct the displaced grid. At 3005, a compressed displacement grid is stored, including the 3D displacement array and the base coordinates.
[0371] At the next time a primitive is read from a storage device or memory, determined at 3006, a displaced grid is generated from the compressed displacement grid at 3003. For example, the 3D displacement array can be applied to the base coordinates to reconstruct the displaced grid. Enhanced Lossy Displacement Grid Compression and Hardware BVH Traversal / Intersection of Lossy Mesh Primitives
[0372] Complex dynamic scenes are challenging for real-time ray tracing implementations. Procedural surfaces, skinned animations, etc. require triangulation and update of the acceleration structure in each frame, even before the first ray is emitted.
[0373] One embodiment of the present invention does not merely use bilinear patches as base primitives, but extends the method to support bicubic quadrilateral or triangular patches, which need to be evaluated in a watertight manner at the patch boundaries. In one implementation, a bit field is added to the lossy raster primitive indicating whether the implicit triangle is valid. One embodiment also includes a modified hardware block that extends an existing refiner to directly produce a lossy displacement grid (e.g., as described above with reference to Figures 28A - 30 ), which is then stored in memory.
[0374] In one implementation, a hardware extension to the BVH traversal unit takes a lossy raster primitive as input and dynamically obtains bounding boxes for a subset of the implicitly referenced triangles / quadrilaterals. The obtained bounding boxes are in a format compatible with the ray / box traversal unit 4130 of the BVH traversal unit's ray-box test circuit module (e.g., described below). The result of the ray-crossing test with the dynamically generated bounding boxes is passed to the ray quadrilateral / triangle intersection unit 4140, which obtains the relevant triangles contained in the bounding boxes and intersects them.
[0375] One implementation also includes an extension to lossy mesh primitives using indirectly referenced vertex data (similar to other embodiments), thereby reducing memory consumption by sharing vertex data across adjacent raster primitives. In one embodiment, a modified version of the hardware BVH triangle intersection block is made aware that the input is from a triangle of a lossy displacement mesh, allowing it to reuse edge calculations for adjacent triangles. Lossy displacement mesh compression also adds extensions to handle motion-blurred geometry.
[0376] As described above, assuming the input is a raster grid of arbitrary size, this input raster grid is first subdivided into smaller sub-grids of a fixed resolution, such as the 4×4 vertices illustrated in Figure 31 Figure.
[0377] As Figure 32 shown in
[0378] struct GridPrim
[0379] In one implementation, these operations consume 100 bytes: 18 bits can be reserved from PrimLeafDesc to disable individual triangles. For example, a bitmask of 000000000100000000b (in top-down, left-to-right order) will disable Figure 33 the highlighted triangle 3301 shown in
[0380] Implicit triangles can be 3×3 quads (4×4 vertices) or more triangles. Many of these are stitched together to form a mesh. The mask tells us whether we want to intersect with the triangle. If a hole is reached, a single triangle of each 4×4 raster is deactivated. This enables higher precision and significantly reduces memory usage: approximately 5.5 bytes per triangle, which is a very compact representation. In contrast, if stored in a linear array at full precision, each triangle occupies 48 and 64 bytes.
[0381] As Figure 34 As shown, the hardware tessellator 3450 tessellates patches into triangles in units of 4×4 and stores them outwards into memory, so that a BVH can be built on them and ray tracing can be performed on them. In this embodiment, the hardware tessellator 3450 is modified to directly support lossy displacement raster primitives. The hardware tessellation unit 3450 can directly generate lossy raster primitives and store them outwards into memory, instead of generating individual triangles and passing them to the rasterization unit.
[0382] An extension to the hardware BVH traversal unit 3450 takes lossy raster primitives as input and dynamically obtains the bounding boxes of subsets of implicitly referenced triangles / quadrilaterals. In Figure 35 the example shown, nine bounding boxes 3501A-I are obtained from the lossy raster, one for each quadrilateral, and passed to the hardware BVH traversal unit 3450 as special nine-wide BVH nodes to perform ray-box intersection.
[0383] Testing all 18 triangles one by one is very expensive. Referring to Figure 36 , one embodiment obtains a bounding box 3501A-I for each quadrilateral (although this is just an example; any number of triangles can be obtained). When a subset of triangles is read and the bounding boxes are calculated, N-wide BVH nodes 3600 are generated - one child node 3501A-I for each quadrilateral. Then, this structure is passed to the hardware traversal unit 3610, which traverses the ray through the newly constructed BVH. Thus, in this embodiment, the raster primitives are used as implicit BVH nodes from which the bounding boxes can be determined. When the bounding boxes are generated, it is known that they contain two triangles. When the hardware traversal unit 3610 determines that the ray traverses one of the bounding boxes 3501A-I, the same structure is passed to the ray-triangle intersecter 3615 to determine which bounding box has been hit. That is, if the bounding box has been hit, an intersection test is performed on the triangles contained in the bounding box.
[0384] In one embodiment of the present invention, these techniques are used as a pre-culling step for the ray-triangle traversal 3610 and intersection unit 3615. The intersection test is much cheaper when the triangles can be inferred using only the BVH node processing unit. For each crossed bounding box 3501A-I, the two corresponding triangles are passed to the ray tracing triangle / quadrilateral intersection unit 3615 to perform the ray-triangle intersection test.
[0385] The raster primitive and implicit BVH node handling techniques described above can be integrated within any traversal / intersection unit (e.g., such as the ray / box traversal unit 4130 described below) described herein, or used as a preprocessing step therefor.
[0386] In one embodiment, an extension of such 4×4 lossy raster primitives is used to support motion blur processing with two time steps. An example is provided in the code sequence below:
[0387] The motion blur operation is similar to the shutter time in a simulated camera. To ray trace this effect, moving from t0 to t1, there are two representations of the triangle, one for t0 and one for t1. In one embodiment, interpolation is performed between them (e.g., linearly interpolating the primitive representation at 0.5 at each of the two time points).
[0388] The disadvantages of acceleration structures such as bounding volume hierarchies (BVHs) and k-d trees are that they require both time and memory to build and store. One way to reduce this overhead is to employ some compression and / or quantization of the acceleration data structure, which is particularly effective for BVHs, which naturally leads to conservative delta coding. On the plus side, this can significantly reduce the size of the acceleration structure, typically halving the size of BVH nodes. On the minus side, compressing BVH nodes also incurs overhead, which can fall into different categories. First, there is the obvious cost of decompressing each BVH node during traversal; second, especially for hierarchical coding schemes, the need to keep track of parent information makes stack operations slightly more complex; and third, conservatively quantifying the bounds means that the bounding boxes are not as tight as the uncompressed bounding boxes, resulting in a significant increase in the number of nodes and primitives that must be traversed and intersected separately.
[0389] Compressing BVHs through local quantization is a known method of reducing their size. An n-wide BVH node contains axis-aligned bounding boxes (AABBs) of its "n" child nodes in single-precision floating-point format. Local quantization represents the "n" child AABBs relative to the AABB of the parent node and stores these values in a quantized, e.g., 8-bit format, thus reducing the size of the BVH node.
[0390] The local quantization of the entire BVH introduces multiple overhead factors because (a) the de-quantized AABB is coarser than the original single-precision floating-point AABB, introducing additional traversal and intersection steps for each ray, and (b) the de-quantization operation itself is expensive, increasing the overhead of each ray traversal step. Due to these drawbacks, compressed BVHs are only used in specific application scenarios and have not been widely adopted.
[0391] One embodiment of the present invention employs techniques to compress the leaf nodes of hair primitives in a bounding volume hierarchy, as described in the co-pending application entitled "Apparatus and Method for Compressing Leaf Nodes of Bounding Volume Hierarchies", Serial No. 16 / 236,185, filed on December 28, 2018, which is assigned to the assignee of the present application. In particular, as described in the co-pending application, several groups of oriented primitives are stored together with the parent bounding box, eliminating the storage of child pointers in the leaf nodes. Then, using 16-bit coordinates quantized relative to the corners of the parent box, an oriented bounding box is stored for each primitive. Finally, quantized normals are stored for each group of primitives to indicate direction. This approach can result in a significant reduction in the bandwidth and memory footprint of BVH hair primitives.
[0392] In some embodiments, BVH nodes (e.g., for an 8-wide BVH) are compressed by storing the parent bounding box and encoding N sub-bounding boxes (e.g., 8 sub-bounding boxes) relative to that parent bounding box using a lower precision. The drawback of applying this idea to each node of the BVH is that when traversing a ray through this structure, some decompression overhead is introduced at each node, which can degrade performance.
[0393] To address this issue, one embodiment of the present invention uses compressed nodes only at the lowest level of the BVH. This provides the advantage of a higher BVH level that runs with optimal performance (i.e., the compressed nodes are touched as frequently as when the boxes are large, but there are very few compressed nodes), and compression at the lower / lowest levels is also very effective because most of the data in the BVH is in the (one or more) lowest levels.
[0394] In addition, in one embodiment, quantization is also applied to BVH nodes that store oriented bounding boxes. As discussed below, this operation is somewhat more complex than axis-aligned bounding boxes. In one implementation, the use of compressed BVH nodes with oriented bounding boxes is combined with using compressed nodes only at the lowest level (or lower levels) of the BVH.
[0395] Thus, one embodiment improves the fully compressed BVH by introducing a single dedicated layer for compressed leaf nodes while using regular uncompressed BVH nodes for internal nodes. One motivation behind this approach is that almost all of the compression savings come from the lowest levels of the BVH (especially for 4-wide and 8-wide BVHs, which make up the vast majority of all nodes), while most of the overhead comes from internal nodes. Thus, introducing a single layer of dedicated "compressed leaf nodes" gives nearly the same (and in some cases even better) compression gain as the fully compressed BVH while maintaining nearly the same traversal performance as the uncompressed BVH.
[0396] Figure 39 Illustrated is an exemplary ray tracing engine 3900 that performs the leaf node compression and decompression operations described herein. In one embodiment, the ray tracing engine 3900 includes circuitry for one or more of the ray tracing cores described above. Alternatively, the ray tracing engine 3900 may be implemented on a core of a CPU or on other types of graphics cores (e.g., Gfx cores, tensor cores, etc.).
[0397] In one embodiment, a ray generator 3902 generates rays and a traversal / intersection unit 3903 traces the rays through a scene that includes a plurality of input primitives 3906. For example, an application such as a virtual reality game may generate a command stream from which the input primitives 3906 are generated. The traversal / intersection unit 3903 traverses the rays through a BVH 3905 generated by a BVH builder 3907 and identifies hit points where the rays intersect one or more of the primitives 3906. Although illustrated as a single unit, the traversal / intersection unit 3903 may include traversal units coupled to different intersection units. These units may be implemented in circuitry, software / commands executed by a GPU or CPU, or any combination thereof.
[0398] In one embodiment, the BVH processing circuitry / logic 3904 includes a BVH builder 3907 that generates a BVH 3905 as described herein based on the spatial relationships between primitives 3906 in a scene. In addition, the BVH processing circuitry / logic 3904 includes a BVH compressor 3909 and a BVH decompressor 3908 for compressing and decompressing leaf nodes, respectively, as described herein. For illustrative purposes, the following description will focus on an 8-wide BVH (BVH8).
[0399] As Figure 40 As illustrated, an embodiment of a single 8-wide BVH node 4000A includes eight bounding boxes 4001 - 4008 and eight (64-bit) sub-pointers / references 4010 to the bounding box / leaf data 4001 - 4008. In one embodiment, the BVH compressor 3925 performs encoding where the eight sub-bounding boxes 4001A - 4008A are expressed relative to the parent bounding box 4000A and quantized to 8-bit uniform values as shown by the bounding box leaf data 4001B - 4008B. The quantized 8-wide BVH QBVH 8 node 4000B is encoded by the BVH compression 4025 using start and range values stored as two 3D single-precision vectors (2 x 12 bytes). The eight quantized sub-bounding boxes 4001B - 4008B are stored as 2 x 8 bytes (48 bytes total) of lower and upper bounds per dimension of the bounding box. Note that this layout is different from existing implementations as the range is stored in full precision, which generally provides tighter bounds but requires more space.
[0400] In one embodiment, the BVH decompressor 3926 decompresses the QBVH8 node 4000B as follows. The decompressed lower bound for dimension i can be calculated on the CPU 4099 by QBVH8.start i +(byte-to-float)QBVH8.lower i *QBVH8.extend i which requires five instructions per dimension and box on the CPU 4099. The five instructions are: 2 loads (start, extend), byte-to-int load + upconversion, int-to-float conversion, and one multiply-add. In one embodiment, all eight quantized sub-bounding boxes 4001B - 4008B are decompressed in parallel using SIMD instructions, which adds an overhead of approximately 10 instructions to the ray-node intersection test, making it at least more than twice as expensive as the standard uncompressed node case. In one embodiment, these instructions are executed on the cores of the CPU 4099. Alternatively, a comparable set of instructions is executed by the ray tracing core 4050.
[0401] In the case without pointers, the QBVH8 node requires 72 bytes while the uncompressed BVH8 node requires 192 bytes, resulting in a reduction factor of 2.66. In the case with eight (64-bit) pointers, the reduction factor drops to 1.88, which makes it necessary to address the storage cost of disposing of leaf pointers.
[0402] In one embodiment, when only the leaf level of the BVH8 nodes is compressed into QBVH8 nodes, all the child pointers of the 8 child nodes 4001 - 4008 will only reference leaf raw data. In one implementation, this fact is exploited by storing all the referenced raw data directly after the QBVH8 node 4000B itself, as Figure 40 illustrated. This allows reducing the full 64-bit child pointer 4010 of the QBVH8 to just an 8-bit offset 4022. In one embodiment, if the primitive data is of a fixed size, the offset 4022 is skipped altogether because the offset can be directly calculated from the index of the bounding volume hierarchy and the pointer to the QBVH8 node 4000B itself.
[0403] When using a top-down BVH8 builder, compressing only the BVH8 leaf level requires only minor modifications to the build process. In one embodiment, these build modifications are implemented in the BVH builder 3907. During the recursive build phase, the BVH builder 3907 keeps track of whether the current number of primitives is below a certain threshold. In one implementation, N×M is the threshold, where N refers to the width of the BVH and M is the number of primitives within a BVH leaf. For a BVH8 node and for example four triangles per leaf, the threshold is 32. Thus, for all subtrees with fewer than 32 primitives, the BVH processing circuit module / logic 3904 will enter a special code path where it will continue the surface area heuristic (SAH)-based splitting process but create a single QBVH8 node 4000B. When the QBVH8 node 4000B is finally created, the BVH compressor 3909 then collects all the referenced primitive data and copies it just after the QBVH8 node.
[0404] The actual BVH8 traversal performed by the ray tracing kernel 4050 or the CPU 4099 is only slightly affected by the leaf level compression. Essentially, the leaf level QBVH8 node 4000B is treated as an extended leaf type (e.g., it is marked as a leaf). This means that the regular BVH8 top-down traversal continues until the QBVH node 4000B is reached. At this point, a single ray-QBVH node intersection is performed, and for all the children in its intersecting children 4001B - 4008B, the corresponding leaf pointers are reconstructed, and the regular ray-primitive intersections are performed. Interestingly, sorting the intersecting children 4001B - 4008B of the QBVH based on the intersection distance may not provide any significant benefit because in most cases, only a single child is intersected by the ray anyway.
[0405] One embodiment of the leaf-level compression scheme even allows for lossless compression of the actual primitive leaf data by obtaining common features. For example, triangles within a compressed leaf BVH (CLBVH) node are likely to share vertices / vertex indices and attributes, just like the same object ID. By storing these shared attributes only once per CLBVH node and using small local byte-sized indices within the primitive, the memory consumption is further reduced.
[0406] In one embodiment, the techniques for leveraging common spatially coherent geometric features within BVH leaves are also used for other more complex primitive types. Primitives such as hair segments may share a common orientation for each BVH leaf. In one embodiment, the BVH compressor 3909 implements a compression scheme that takes into account this common orientation attribute to efficiently compress the oriented bounding box (OBB), which has been shown to be very useful for the boundary long diagonal primitive type.
[0407] The leaf-level compressed BVH described herein introduces BVH node quantization only at the lowest BVH level and thus allows for additional memory reduction optimizations while maintaining the traversal performance of the uncompressed BVH. Since only the lowest-level BVH nodes are quantized, all of their child nodes point to leaf data 4001B - 4008B, which can be stored continuously in a memory block or one or more cache lines 3998.
[0408] This idea can also be applied to hierarchies using oriented bounding boxes (OBBs), which are commonly used to accelerate the rendering of hair primitives. To illustrate a particular embodiment, the memory reduction in the typical case of a standard 8-wide BVH on triangles will be evaluated.
[0409] The layout of the 8-wide BVH node 4000 is represented by the following code sequence: And requires 276 memory bytes. The layout of the standard 8-wide quantized node can be defined as: And requires 136 bytes.
[0440] Since only the quantized BVH nodes are used at the leaf level, all child node pointers will actually point to leaf data 4001A - 4008A. In one embodiment, by storing the quantized node 4000B and all of the leaf data 4001B - 4008B pointed to by its child nodes in a single contiguous block of memory 3998, the 8 child node pointers in the quantized BVH node 4000B are removed. Saving the child node pointers reduces the quantized node layout to: It only requires 72 bytes. Due to the contiguous layout in the memory / cache 3998, the child pointer of the i-th child node can now be simply calculated as: childPtr(i) = addr(QBVH8NodeLeaf) + sizeof(QBVH8NodeLeaf) + i * sizeof(LeafDataType).
[0411] Since the nodes at the lowest level of the BVH account for more than half of the entire size of the BVH, the leaf-level-only compression described herein provides a reduction of 0.5 + 0.5 * 72 / 256 = 0.64 times the original size.
[0412] In addition, the overhead of having coarser boundaries and the cost of decompressing the BVH nodes themselves only occur at the BVH leaf level (as opposed to all levels when the entire BVH is quantized). Thus, the typically rather large traversal and intersection overheads due to the coarser boundaries (introduced by quantization) are largely avoided.
[0413] Another benefit of embodiments of the present invention is improved hardware and software prefetch efficiency. This stems from the fact that all leaf data is stored in relatively small contiguous memory blocks or cache lines.
[0414] Since the geometry of the BVH leaf level is spatially consistent, all the primitives referenced by the QBVH8NodeLeaf nodes are likely to share common attributes / features, such as objectID, one or more vertices, etc. Thus, an embodiment of the present invention further reduces storage by removing primitive data duplicates. For example, each QBVH8NodeLeaf node can store the primitive and associated data only once, thereby further reducing the memory consumption of the leaf data.
[0415] The effective boundaries of hair primitives are described below as an example of significant memory reduction achieved by leveraging the common geometric attributes of the BVH leaf level. To precisely constrain hair primitives (which are long and thin structures oriented in space), a well-known method is to calculate an oriented bounding box to tightly constrain the geometry. First, a coordinate space aligned with the hair direction is calculated. For example, the z-axis can be determined to point in the hair direction, while the x-axis and y-axis are perpendicular to the z-axis. Using this oriented space, a standard AABB can now be used to tightly constrain the hair primitive. Making a ray cross such an oriented boundary requires first transforming the ray into the oriented space and then performing a standard ray / box intersection test.
[0416] The problem with this method is its memory usage. The transformation to the orientation space requires 9 floating-point values, and storing the bounding box requires another 6 floating-point values, resulting in a total of 60 bytes.
[0417] In one embodiment of the present invention, the BVH compressor 3925 compresses this orientation space and the bounding box for a plurality of hair primitives that are spatially close together. Then, these compressed boundaries can be stored inside the compressed leaf level to tightly constrain the hair primitives stored inside the leaf. In one embodiment, the following method is used to compress the orientation boundaries. The orientation space can be represented by three mutually orthogonal normalized vectors v x 、v y and v z . By projecting the point p onto these axes, the point p is transformed into this space: p x =dot(v x ,p) p y =dot(v y ,p) p z =dot(v z ,p)
[0418] When the vectors v x 、v y and v z are normalized, their components are in the range [-1, 1]. Therefore, 8-bit signed fixed-point numbers are used instead of 8-bit signed integers and constant scaling to quantize these vectors. This method generates the quantized v x 、v y and v z . This method reduces the memory required to encode the orientation space from 36 bytes (9 floating-point values) to only 9 bytes (9 fixed-point numbers, where each fixed-point number is 1 byte).
[0419] In one embodiment, by taking advantage of the fact that all the vectors are orthogonal to each other, the memory consumption of the orientation space is further reduced. Therefore, one only needs to store two vectors (e.g., p y and p z ), and p x =cross(p y ,p z ) can be calculated, further reducing the required storage space to only six bytes.
[0420] All that remains is to quantize the AABB inside the quantized orientation space. The problem here is to project the point p onto the compressed coordinate axes of this space (e.g., by calculating dot(v x , p) can produce potentially large ranges of values (since value p is typically encoded as a floating point number). For this reason, one would need to use floating point numbers to encode the bounds, reducing the potential savings.
[0421] To address this issue, an embodiment of the present invention first transforms a plurality of hair primitives into a space where its coordinates are within the range [0, 1 / √3]. This can be accomplished by determining the world space axis-aligned bounding box b of the plurality of hair primitives and using a transformation T that first translates b.lower to the left and then scales by 1 / max(b.size.x, b.size.y, b.size.z) in each coordinate:
[0422] An embodiment ensures that this transformed geometry remains within the range [0, 1 / √3] because the projection of the transformed points on the quantization vector p x , p y ’ or p z ’ remains within the range [1, 1]. This means that when transformed using T, the AABB of the curve geometry can be quantized and then transformed into the quantized oriented space. In one embodiment, 8-bit signed fixed-point arithmetic is used. However, for accuracy reasons, 16-bit signed fixed-point numbers can be used (e.g., using 16-bit signed integers and constant scaling encoding). This reduces the memory requirement for encoding the axis-aligned bounding box from 24 bytes (6 floating point values) to only 12 bytes (6 words) plus the offset b.lower (3 floating points) and scale (1 floating point) shared for the plurality of hair primitives.
[0423] For example, for 8 hair primitives to be constrained, this embodiment reduces the memory consumption from 8 * 60 bytes = 480 bytes to only 8 * (6 + 12) + 3 * 4 + 4 = 160 bytes, which is reduced to 1 / 3. Intersecting the ray with these quantized oriented bounds works as follows: First, the ray is transformed using transformation T, and then the quantized v x , v y and v z project the ray. Finally, the ray intersects the quantized AABB.
[0424] The fat leaf method described above provides even more opportunities for compression. Assuming there is an implicit single floating point 3-pointer in the fat BVH leaf pointing to the shared vertex data of multiple adjacent GridPrims, the vertices in each grid primitive can be indirectly addressed by a byte-sized index ("vertex_index_*"), thus leveraging vertex sharing. In Figure 37 Among them, vertices 3701 - 3702 are shared and stored in full precision. In this embodiment, the shared vertices 3701 - 3702 are stored only once, and the indices pointing to the array containing the unique vertices are stored. Therefore, only 4 bytes are stored for each timestamp instead of 48 bytes. The indices in the following code sequence are used to identify the shared vertices.
[0425] In one embodiment, the shared edges of the primitives are evaluated only once to save processing resources. For example, in Figure 38 Assume that the bounding box consists of the highlighted quadrilaterals. One embodiment of the present invention performs a ray-edge calculation once for each of the three shared edges instead of intersecting with all triangles separately. Thus, the results of the three ray-edge calculations are shared across four triangles (i.e., the ray-edge calculation is performed only once for each shared edge). Additionally, in one embodiment, the results are stored in on-chip memory (e.g., scratchpad memory / cache directly accessible across the intersection units). Apparatus and method for box - box testing and accelerated collision detection for ray tracing
[0426] Figures 41A - 41B Illustrates a ray tracing architecture according to an embodiment of the present invention. Multiple graphics processor core blocks 4110 execute shaders and other program code related to ray tracing operations. The "Traceray" function executed on one of the graphics processor core blocks 4110 triggers the ray state initializer 4120 to initialize the state required to trace the current ray (identified via the ray ID / descriptor), and the current ray passes through a bounding volume hierarchy (BVH) (e.g., stored in stack 2721 in memory buffer 4118, or other data structures in local or system memory 1598).
[0427] In one embodiment, if the Traceray function identifies a ray for which a previous traversal operation is partially completed, the state initializer 4120 uses the unique ray ID to load the associated ray tracing data 2502 and / or stack 2721 from one or more buffers 4118 in memory 1598. As mentioned above, memory 1598 can be on-chip / local memory or cache and / or a system-level memory device.
[0428] As discussed with respect to other embodiments, a trace array 4149 can be maintained to store the traversal progress of each ray. If the current ray has partially traversed the BVH, the state initializer 4120 can use the trace array 4149 to determine the BVH level / node to restart from.
[0429] The traversal and ray box test unit 4130 traverses rays through the BVH. When a primitive has been identified within a leaf node of the BVH, the instance / quad intersection tester 4140 tests the rays for intersection with the primitive (e.g., one or more primitive quads) and retrieves the associated ray / shader record from the ray tracing cache 4160 integrated within the cache hierarchy of the graphics processor (shown here coupled to the L1 cache 4170). The instance / quad intersection tester 4140 is sometimes referred to herein simply as the intersection unit (e.g., Figure 39 the intersection unit 3903 in
[0430] The ray / shader record is provided to the thread dispatcher 4150, which dispatches new threads to the graphics processor core block 4110 at least in part using the unbounded thread dispatching techniques described herein. In one embodiment, the ray / box traversal unit 4130 includes the traversal / stack trace logic 4348 described above, which tracks and stores the traversal progress of each ray within the trace array 4149.
[0431] A class of problems in rendering can be mapped to testing the collision of a test box with other bounding volumes or bounding boxes (e.g., due to overlap). Such box queries can be used to enumerate the geometry inside the query bounding box for various applications. For example, box queries can be used to collect photons during photon mapping, enumerate all light sources that may affect a query point (or query region), and / or search for the closest surface point to a certain query point. In one embodiment, box queries operate on the same BVH structure as ray queries; thus, a user can trace rays through a scene and perform box queries on the same scene.
[0432] In one embodiment of the present invention, with respect to ray tracing hardware / software, box queries are processed similarly to ray queries, where the ray / box traversal unit 4130 performs traversal using box / box operations instead of ray / box operations. In one embodiment, the traversal unit 4130 can use the same set of features for box / box operations as for ray / box operations, including but not limited to motion blur, masks, flags, nearest hit shaders, any hit shaders, miss shaders, and traversal shaders. One embodiment of the present invention adds one bit to each ray tracing message or instruction (e.g., TraceRay as described herein) to indicate that the message / instruction is associated with a BoxQuery operation. In one implementation, BoxQuery is enabled in both synchronous and asynchronous ray tracing modes (e.g., using standard dispatching and unbounded thread dispatching operations, respectively).
[0433] In one embodiment, once the bit is set to BoxQuery mode, ray tracing hardware / software (e.g., traversal unit 4130, instance / quad intersection tester 4140, etc.) interprets data associated with ray tracing messages / instructions as box data (e.g., minimum / maximum values in three dimensions). In one embodiment, the traversal acceleration structure is generated and maintained as described previously, but boxes are initialized for each primary stack ID instead of rays.
[0434] In one embodiment, no hardware instantiation is performed for box queries. However, instantiation can be emulated in software using a traversal shader. Thus, when an instance node is reached during a box query, the hardware can treat the instance node as a program node. Since the headers of the two structures are the same, this means the hardware will call the shader stored in the header of the instance node and then it can continue with point queries inside the instance.
[0435] In one embodiment, a ray flag is set to indicate that the instance / quad intersection tester 4140 will accept the first hit and end the search (e.g., the ACCEPT_FIRST_HIT_AND_END_SEARCH flag). When this ray flag is not set, similar to a ray query, the intersecting child nodes are input from front to back based on their distance to the query box. This traversal order significantly improves performance when searching for the geometry closest to a point, as in the case of a ray query.
[0436] One embodiment of the present invention uses any-hit shaders to filter out false positive hits. For example, although the hardware may not perform accurate box / triangle tests at the leaf level, it will conservatively report all triangles that hit the leaf node. Additionally, when the search box is shrunk by any-hit shaders, the hardware may return the primitives of the popped leaf node as hits, even if the leaf node box may no longer overlap with the shrunk query box.
[0437] As Figure 41A indicated, a box query can be issued by sending a message / command (i.e., Traceray) to the hardware through the graphics processor core block 4110. Then, the processing proceeds as described above, i.e., through the state initializer 4120, ray / box traversal logic 4130, instance / quad intersection tester 4140, and unbounded thread dispatcher 4150.
[0438] In one embodiment, the box query reuses the MemRay data layout used for ray queries by storing the lower bound of the query box in the same location as the ray origin, the upper bound in the same location as the ray direction, and the query radius in the far value.
[0439] Using this MemBox layout, the hardware uses box[lower-radius, upper+radius] to perform queries. Thus, the boundaries of the storage are extended by a certain radius in each dimension using the L0 norm. This query radius may be useful for easily narrowing down the search area, such as for nearest point search.
[0440] Since the MemBox layout only reuses the ray origin, ray direction, and T far members of the MemRay layout, data management in the hardware does not need to be changed for ray queries. Instead, the data is stored in internal storage devices (e.g., ray tracing cache 4160 and L1 cache 4170) like ray data and will only be interpreted differently for box / box tests.
[0441] In one embodiment, the following operations are performed by the ray / state initialization unit 4120 and the ray / box traversal unit 4130. The additional bit "BoxQueryEnable" from the TraceRay message is pipelined in the state initializer 4120 (affecting its compression across messages), providing an indication of the BoxQueryEnable setting to each ray / box traversal unit 4130.
[0442] The ray / box traversal unit 4130 stores the "BoxQueryEnable" for each ray and sends this bit as a tag along with the initial ray load request. When the requested ray data is returned from the memory interface, in the case where BoxQueryEnable is set, the reciprocal calculation is bypassed, and instead different configurations (i.e., according to the box rather than the ray) are loaded for all components in RayStore.
[0443] The ray / box traversal unit 4130 pipelines the BoxQueryEnable bit to the underlying test logic. In one embodiment, the ray box data path is modified according to the following configuration settings. If BoxQueryEnable == 1, the plane of the box does not change because it changes based on the signs of the x, y, and z components of the ray direction. Bypass the checks performed on the ray that are unnecessary for the ray box. For example, assume that the query box has no INF or NAN, so these checks are bypassed in the data path.
[0444] In one embodiment, before being processed by the hit determination logic, another addition operation is performed to determine the values lower+radius (essentially the t value from the hit) and upper-radius. Additionally, when hitting an "instance node" (in a hardware instantiation implementation), it does not calculate any transforms but instead uses the shader ID in the instance node to start the cross shader.
[0445] In one embodiment, when BoxQueryEnable is set, the ray / box traversal unit 4130 does not perform an empty shader lookup for any hit shaders. Additionally, when BoxQueryEnable is set, when the valid node is a quadrilateral meshlet type, the ray / box traversal unit 4130 calls the intersection shader as if it would call any hit shader after updating the potential hit information in memory.
[0446] In one embodiment, a set of independent and various components as illustrated in Figure 41A are provided in each multi-core group 1500A (e.g., within the ray tracing core 1550). In this implementation, each multi-core group 1500A can operate on different sets of ray data and / or box data in parallel to perform traversal and intersection operations as described herein. Apparatus and method for meshlet compression and decompression for ray tracing
[0447] As described above, a "meshlet" is a subset of a mesh created by geometric partitioning, which includes a certain number of vertices (e.g., 16, 32, 64, 256, etc.) based on the number of associated attributes. Meshlets can be designed to share as many vertices as possible to allow vertex reuse during rendering. This partitioning can be pre-computed to avoid runtime processing, or can be dynamically performed at runtime each time a mesh is drawn.
[0448] One embodiment of the present invention performs meshlet compression to reduce the storage requirements of the underlying acceleration structure (BLAS). This embodiment takes advantage of the fact that meshlets represent a small portion of a larger mesh with similar vertices to allow for efficient compression within a 128B data block. However, note that the underlying principles of the present invention are not limited to any specific block size.
[0449] Meshlet compression can be performed when building the corresponding bounding volume hierarchy (BVH) and decompressed when the BVH is consumed (e.g., by the ray tracing hardware block). In certain embodiments described below, meshlet decompression is performed between the L1 cache (sometimes the "LSC unit") and the ray tracing cache (sometimes the "RTC unit"). As described herein, the ray tracing cache is a high-speed local cache used by the ray traversal / intersection hardware.
[0450] In one embodiment, meshlet compression is accelerated in hardware. For example, if the graphics processor core block path supports decompression (e.g., potentially supports traversal shader execution), meshlet decompression can be integrated into the common path outside the L1 cache.
[0451] In one embodiment, a message is used to initiate the compression of a small grid in 128B chunks in memory. For example, a 4×64B message input can be compressed into a 128B chunk output to the shader. In this implementation, additional node types are added in the BVH to indicate the association with the compressed small grid.
[0452] Figure 41B FIG. illustrates a particular implementation of small grid compression, including a small grid compression block (RTMC) 4230 and a small grid decompression block (RTMD) 4290 integrated within a ray tracing cluster. When a new message is transferred from the graphics processor core block 4110 executing the shader to the ray tracing cluster (e.g., within the ray tracing core 1550), the small grid compression 4230 is invoked. In one embodiment, the message includes four 64B phases and a 128B write address. The message from the graphics core 4110 indicates to the small grid compression block 4131 where in the local memory 1598 (and / or system memory, depending on the implementation) to locate the vertices and associated small grid data. The small grid compression block 4131 then performs the small grid compression as described herein. The compressed small grid data can then be stored in the local memory 1598 and / or the ray tracing cache 4160 via the memory interface 4133 and accessed by the instance / quad cross-tester 4140 and / or the traversal / cross-shader.
[0453] In Figure 41B the small grid collection and decompression block 4190 can collect the compressed data of the small grid and decompress the data into multiple 64B chunks. In one implementation, only the decompressed small grid data is stored within the L1 cache 4170. In one embodiment, the small grid decompression is activated when obtaining BVH node data based on the node type (e.g., leaf node, compressed) and the primitive ID. The traversal shader can also access the compressed small grid using the same semantics as the rest of the ray tracing implementation.
[0454] In one embodiment, the small grid compression block 4131 accepts an input triangle array from the graphics core 4110 and produces a compressed 128B small grid leaf structure. A pair of consecutive triangles in this structure forms a quadrilateral. In one implementation, the graphics core message includes up to 14 vertices and triangles as indicated in the following code sequence. The compressed small grid is written to memory at the address provided in the message via the memory interface 4133.
[0455] In one embodiment, the shader calculates the bit budget for the set of small grids and thus provides an address such that occupancy space compression is possible. These messages are only initiated for compressible small grids.
[0456] In one embodiment, the meshlet decompression block 4190 decompresses two consecutive quadrilaterals (128B) from a 128B meshlet and stores the decompressed data in the L1 cache 4170. The tags in the L1 cache 4170 track the index of each decompressed quadrilateral (including triangle indices) and the meshlet address. The ray tracing cache 4160 and the graphics core 4110 can fetch 64B decompressed quadrilaterals from the L1 cache 4170. In one embodiment, the graphics core 4110 fetches the decompressed quadrilaterals by issuing a MeshletQuadFetch message to the L1 cache 4160, as shown below. Separate messages can be issued to fetch the first 32 bytes and the last 32 bytes of the quadrilateral.
[0457] The shader can access triangle vertices from the quadrilateral structure, as shown below. In one embodiment, the "if" statement is replaced by a "sel" instruction.
[0458] / / Assume vertex i is a constant determined by the compiler
[0459] In one embodiment, the ray tracing cache 4160 can directly fetch the decompressed quadrilaterals from the L1 cache 4170 bank by providing the meshlet address and the quadrilateral index. Small Mesh Compression Process
[0460] After allocating bits for fixed overheads such as geometric attributes (e.g., flags and masks), the meshlet data is added to the compression block while calculating the remaining bit budget based on the difference between (pos.x, pos.y, pos.z) and (base.x, base.y, base.z), where the base values include the position of the first vertex in the list. Similarly, the prim-ID (raster ID) difference can also be calculated. Since the differences are compared to the first vertex, it is cheaper to decompress with low latency. The base position and primID are part of the constant overhead in the data structure, along with the width of the difference bits. For the remaining vertices of even triangles, the position differences and prim-ID differences are stored on different 64B blocks to pack them in parallel.
[0461] Using these techniques, the BVH construction operation consumes lower memory bandwidth when writing compressed data via the memory interface 4133. Additionally, in one embodiment, storing compressed small meshes in the L3 cache allows for storing more BVH data with the same L3 cache size. In a viable implementation, more than 50% of the small meshes are compressed 2:1. When using a BVH with compressed small meshes, the bandwidth savings at the memory result in power savings. Apparatus and method for unbounded thread dispatch and workgroup / thread preemption in a compute and ray tracing pipeline
[0462] As described above, unbounded thread dispatch (BTD) is a way to address the SIMD divergence problem in ray tracing in implementations that do not support shared local memory (SLM) or memory barriers. Embodiments of the present invention include support for a generalized BTD that can be used to address SIMD divergence in various compute models. In one embodiment, any compute dispatch with a thread group barrier and SLM can spawn unbounded child threads, and all threads in the thread can be regrouped and dispatched via BTD to improve efficiency. In one implementation, each parent thread allows one unbounded child thread at a time, and the original thread is allowed to share its SLM space with the unbounded child thread. Both the SLM and the barrier are released only when the ultimately converging parent thread terminates (i.e., executes EOT). A particular embodiment allows amplification within a callable mode, thus allowing a tree traversal scenario where more than one child thread is spawned.
[0463] Figure 42 An initial thread group 4200 that can be synchronously processed by a SIMD pipeline is graphically illustrated. For example, the threads 4200 can be dispatched and synchronously executed as a workgroup. However, in this embodiment, the initial synchronized thread group 4200 can generate multiple divergent spawned threads 4201, and the spawned threads 4201 can generate other spawned threads 4211 within the asynchronous ray tracing architecture described herein. Ultimately, the converging spawned threads 4221 return to the original thread group 4200, which can then continue synchronous execution, restoring context as needed according to the tracing array 4149.
[0464] In one embodiment, the unbounded thread dispatch (BTD) function supports SIMD16 and SIMD32 modes, variable general-purpose register (GPR) usage, shared local memory (SLM), and BTD barriers by persisting in the restoration of the parent thread after execution and completion (post-divergence and then convergence of spawned threads). One embodiment of the present invention includes an implementation of hardware management for restoring the parent thread and dereferencing of software management of SLM and barrier resources.
[0465] In one embodiment of the present invention, the following terms have the following meanings:
[0466] Callable Mode : Threads derived from unbound thread dispatches are in the "callable mode". These threads can access the inherited shared local memory space and can optionally spawn threads per thread in the callable mode. In this mode, the threads cannot access the workgroup-level barrier.
[0467] Workgroup (WG) Mode : When threads are executing in the same way as the SIMD lanes composed of those dispatched by standard thread dispatches, they are defined to be in the workgroup mode. In this mode, the threads can access the workgroup-level barrier as well as the shared local memory. In one embodiment, in response to issuing a "compute walker" command for a compute-only context, a thread dispatch is initiated.
[0468] Normal Derivation : Also known as the regular spawned thread 4211( Figure 42 ), ordinary spawn is initiated whenever one callable program calls another. Threads spawned in this way are considered to be in the callable mode.
[0469] Divergent Derivation : As Figure 42 shown, when a thread transitions from the workgroup mode to the callable mode, the divergent spawn thread 4201 is triggered. The arguments for the divergent spawn are the SIMD width and the fixed function thread ID (FFTID), which are subgroup-uniform.
[0470] Convergent Derivation : When a thread transitions back from the callable mode to the workgroup mode, the convergent spawn thread 4221 is executed. The arguments for the convergent spawn are the FFTID per lane and a mask indicating whether the stack for the lane is empty. This mask must be dynamically calculated by checking the value of the per-lane stack pointer at the return site. The compiler must calculate this mask because these callable threads may call each other recursively. Lanes in a convergent spawn without the convergence bit set will behave like an ordinary spawn.
[0471] In some implementations where sharing local memory or barrier operations are not allowed, unbounded thread dispatch solves the SIMD divergence problem in ray tracing. Additionally, in one embodiment of the present invention, using various computational models, BTD is used to solve SIMD divergence. In particular, any computational dispatch with thread group barriers and shared local memory can spawn unbounded child threads (e.g., one child thread at a time per parent thread), and all identical threads can be regrouped and dispatched by BTD for better efficiency. This embodiment allows the original threads to share their shared local memory space with their child threads. The shared local memory allocation and barriers are released only when the ultimately converging parent threads terminate (as indicated by the end-of-thread (EOT) indicator). One embodiment of the present invention also provides amplification in the callable mode, allowing tree traversal cases where more than one child thread is spawned.
[0472] Although not limited to this, one embodiment of the present invention is implemented on a system in which no SIMD lane provides support for amplification (i.e., only a single prominent SIMD lane in the form of divergent or convergent spawned threads is allowed). Additionally, in one implementation, when dispatching threads, 32b of (FFTID, BARRIER_ID, SLM_ID) is sent to the dispatcher 4150 enabling BTD. In one embodiment, all of these spaces are released before starting the threads and sending this information to the unbounded thread dispatcher 4150. In one implementation, only a single context is active at a time. Thus, a rogue kernel cannot access the address space of another context even after tempering the FFTID.
[0473] In one embodiment, if StackID (stack ID) allocation is enabled, the shared local memory and barriers are no longer dereferenced when a thread terminates. Instead, they are dereferenced only when all associated StackIDs have been released when the thread terminates. One embodiment prevents fixed-function thread ID (FFTID) leaks by ensuring that StackIDs are properly released.
[0474] In one embodiment, the barrier message is specified to explicitly obtain the barrier ID from the sending thread. This is necessary to enable barrier / SLM usage after an unbounded thread dispatch call.
[0475] Figure 43 Illustrates an embodiment of an architecture for performing unbounded thread dispatch and thread / workgroup preemption as described herein. The graphics processor core block 4110 of this embodiment supports direct manipulation of thread execution masks 4350 - 4353, and each BTD-derived message supports an FFTID reference count for re-deriving the parent thread after the convergence derivation 4221 is completed. Thus, the ray tracing circuit module described herein supports additional message variants for BTD-derived and TraceRay messages. In one embodiment, the BTD-enabled dispatcher 4150 maintains a per-FFTID (as assigned by thread dispatch) count of the original SIMD lanes on the divergent derived threads 4201 and counts down for the convergent derived threads 4221 to initiate the resumption of the parent thread 4200.
[0476] During execution, various events can be counted, including but not limited to regular derived execution 4211; divergent derived execution 4201; convergent derived events 4221; an FFTID counter reaching a minimum threshold (e.g., 0); and loads performed for (FFTID, BARRIER_ID, SLM_ID).
[0477] In one embodiment, BTD-enabled threads are allowed to share local memory (SLM) and barrier allocations (i.e., comply with ThreadGroup semantics). The BTD-enabled thread dispatcher 4150 decouples the FFTID release and barrier ID release from the end-of-thread (EOT) indication (e.g., via a specific message).
[0478] In one embodiment, to support callable shaders from compute threads, a driver-managed buffer 4370 is used to store workgroup information across unbounded thread dispatches. In a particular implementation, the driver-managed buffer 4370 includes multiple entries, where each entry is associated with a different FFTID.
[0479] In one embodiment, within the state initializer 4120, two bits are allocated to indicate the pipeline-derived types considered for message compression. For divergent messages, the state initializer 4120 also takes into account the FFTID from the pipeline with each SIMD channel and the message to the ray / box traversal block 4130 or the unbounded thread dispatcher 4150. For convergent derivation 4221, for the ray / box traversal unit 4130 or the unbounded thread dispatcher 4150, there is an FFTID for each SIMD channel in the message, and each SIMD channel has a pipeline FFTID. In one embodiment, the ray / box traversal unit 4130 also pipelines the derivation types, including convergent derivation 4221. In particular, in one embodiment, the ray / box traversal unit 4130 pipelines and stores the FFTID for each ray convergent derivation 4221 of the TraceRay message.
[0480] In one embodiment, the thread dispatcher 4150 has a dedicated interface to provide the following data structures to prepare for dispatching new threads with the unbounded thread dispatch enable bit set:
[0481] The unbounded thread dispatcher 4150 also processes the thread end (EOT) message with three additional bits: Release_FFTID, Release_BARRIER_ID, Release_SLM_ID. As mentioned, the thread end (EOT) message does not necessarily release / dereference all ID-associated allocations, but only those with the release bit set. A typical use case is when a divergent derivation 4201 is initiated, the derived thread generates an EOT message, but the release bit is not set. Its continuation after the convergent derivation 4221 will generate another EOT message, but this time the release bit is set. Only at this stage will all per-thread resources be reclaimed.
[0482] In one embodiment, the unbounded thread dispatcher 4150 implements a new interface to load the FFTID, BARRIER_ID, SLM_ID, and channel count. It stores all this information in the FFTID-addressable storage device 4321, which has a certain number of entry depths (in one embodiment, the max_fftid depth is 144 entries). In one implementation, the BTD-enabled dispatcher 4150, in response to any regular derivation 4211 or convergent derivation 4201, uses this identification information for each SIMD channel, performs a query on the FFTID-addressable storage device 4321 on a per-FFTID basis, and stores the thread data in the sorting buffer, as described above (e.g., see Figure 18 The content in (addressable memory 1801). This results in storing an additional amount of data (e.g., 24 bits) per SIMD channel in the sorting buffer 1801.
[0483] Upon receiving a convergent derived message, for each SIMD channel from the state initializer 4120 or the ray / box traversal block 4130 to the unbounded thread dispatcher 4150, the count is decremented per FFTID. When the FFTID counter of a given parent thread becomes zero, the entire thread is scheduled with the original execution mask 4350 - 4353, where the continuation shader record 1801 is provided by the convergent derived message in the sorting circuit module 4008.
[0484] Different embodiments of the present invention can operate according to different configurations. For example, in one embodiment, all divergent derivatives 4201 executed by a thread must have a matching SIMD width. Additionally, in one embodiment, a SIMD channel shall not execute a convergent derivative 4221 in which the convergence mask bit is set within the relevant execution mask 4350 - 4353, unless an earlier thread has executed a divergent derivative with the same FFTID. If a divergent derivative 4201 is executed with a given StackID, the convergent derivative 4221 must occur before the next divergent derivative.
[0485] If any SIMD channel in a thread executes a divergent derivative, then all channels must eventually execute a divergent derivative. A thread that has executed a divergent derivative may not execute a barrier, or a deadlock will occur. This restriction is necessary to enable derivatives within a divergent control flow. A parent subgroup cannot be re-derived until all channels have diverged and re-converged.
[0486] A thread must eventually terminate after executing any derivative to ensure forward progress. If multiple derivatives are executed before the thread terminates, a deadlock may occur. In a particular embodiment, the following invariants are followed, although the underlying principles of the present invention are not limited to this: · All divergent derivatives executed by a thread must have a matching SIMD width. · A SIMD channel shall not execute a convergent derivative in which the ConvergenceMask bit is set within the relevant execution mask 4350 - 4353, unless an earlier thread has executed a divergent derivative with the same FFTID. · If a divergent derivative is executed with a given stackID, the convergent derivative must occur before the next divergent derivative. · If any SIMD lanes in a thread perform a divergent spawn, then all lanes must eventually perform a divergent spawn. A thread that has performed a divergent spawn may not execute a barrier or a deadlock will occur. This restriction allows spawning within a divergent control flow. A parent-child group cannot be re-spawned until all lanes have diverged and re-converged. · A thread must eventually terminate after any spawns are executed to guarantee forward progress. If multiple spawns are executed before the thread terminates, a deadlock may occur.
[0487] In one embodiment, the BTD-enabled dispatcher 4150 includes thread preemption logic 4320 to preempt the execution of certain types of workloads / threads, thereby freeing resources for the execution of other types of workloads / threads. For example, the various embodiments described herein can execute both compute workloads and graphics workloads (including ray tracing workloads) that can run at different priorities and / or have different latency requirements. To meet the requirements of each workload / thread, one embodiment of the present invention pauses the ray traversal operation to free execution resources for a higher-priority workload / thread or a workload / thread that would otherwise not meet the specified latency requirements.
[0488] One embodiment uses short stacks 4303-4304 to store a limited number of BVH nodes during the traversal operation to reduce the storage requirements of the traversal. These techniques can be used by Figure 43 the embodiments in which the ray / box traversal unit 4130 efficiently pushes entries onto and pops entries from the short stacks 4303-4304 to ensure that the required BVH nodes 5290-5291 are available. Additionally, when the traversal operation is executed, the traversal / stack tracker 4348 updates the tracking data structure, herein referred to as the tracking array 4149, and the associated stacks 4303-4304 and ray tracing data 2502. Using these techniques, when the traversal of a ray is paused and restarted, the traversal circuit module / logic 4130 can consult the tracking data structure 4149 and access the associated stacks 4303-4304 and ray tracing data 2502 to start the traversal operation of the ray at the same location where the ray stopped within the BVH.
[0489] In one embodiment, the thread preemption logic 4320 determines when a set of traversal threads (or other thread types) are to be preempted, as described herein (e.g., to free resources for a higher-priority workload / thread), and notifies the ray / box traversal unit 4130 so that it can pause processing one of the current threads to free resources for processing a higher-priority thread. In one embodiment, the "notification" is simply performed by dispatching instructions for the new thread before the traversal on the old thread is complete.
[0490] Accordingly, one embodiment of the present invention includes hardware support for both synchronous ray tracing operating in workgroup mode (i.e., where all threads of a workgroup are executed synchronously) and asynchronous ray tracing using unbounded thread dispatching as described herein. These techniques greatly improve performance compared to current systems that require all threads in a workgroup to complete before preemption. In contrast, the embodiments described herein can perform stack-level and thread-level preemption by closely tracking traversal operations, storing only the data required for restart, and using short stacks when appropriate. These techniques are at least partially possible because the ray tracing acceleration hardware and the graphics processor core block 4110 communicate via the persistent memory structure 1598 managed at each ray level and each BVH level.
[0491] When a Traceray message is generated as described above and there is a preemption request, the ray traversal operation can be preempted at various stages, including (1) not yet started, (2) partially completed and preempted, (3) traversal completed without unbounded thread dispatching, and (4) traversal completed but with unbounded thread dispatching. If the traversal has not yet started, no additional data from the trace array 4149 is required when the Traceray message resumes. If the traversal is partially completed, then the traversal / stack tracer 4348 will read the trace array 4149 to determine where to resume the traversal using the ray tracing data 2502 and the stack 5121 as needed. It can query the trace array 4149 using the unique ID assigned to each ray.
[0492] If the traversal is completed and there is no unbounded thread dispatching, any hit information stored in the trace array 4149 (and / or other data structures 2502, 5121) can be used to schedule unbounded thread dispatching. If the traversal is completed and there is unbounded thread dispatching, the unbounded thread is resumed and execution continues until completion.
[0493] In one embodiment, the trace array 4149 includes an entry for each unique ray ID of the rays in flight, and each entry can include one of the execution masks 4350 - 4353 of the corresponding thread. Alternatively, the execution masks 4350 - 4353 can be stored in a separate data structure. In either implementation, each entry in the trace array 4149 can include a 1-bit value or be associated with a 1-bit value to indicate whether the corresponding ray needs to be resubmitted when the ray / box traversal unit 4130 resumes operation after preemption. In one implementation, this 1-bit value is managed within a thread group (i.e., a workgroup). This bit can be set to 1 at the start of ray traversal and can be reset back to 0 when ray traversal is completed.
[0494] The techniques described herein allow traversal threads associated with ray traversal to be preempted by other threads (e.g., compute threads) without waiting for the traversal thread and / or the entire workgroup to complete, thereby improving performance associated with high-priority and / or low-latency threads. Additionally, due to the techniques described herein for tracking traversal progress, traversal threads can resume where they left off, saving significant processing cycles and resource usage. Further, the embodiments described above allow workgroup threads to spawn unbound threads and provide a mechanism for re-converging to return to the original SIMD architecture state. These techniques effectively improve the performance of ray tracing and compute threads by an order of magnitude. Apparatus and method for level-of-detail selection within a BVH
[0495] Embodiments of the present invention include a multi-LoD traversal mechanism and node layout within a BVH that allow multiple in-mesh LoD levels to be rendered using fixed-function traversal hardware without the need for additional programmable shaders. These embodiments reduce the frequency of BVH reconstruction across LoD changes and allow for efficient random LoD transitions.
[0496] As described throughout this specification, ray tracing architectures typically rely on a bounding volume hierarchy (BVH) to perform ray traversal and intersection. The BVH is constructed around the objects in a graphics scene, and then each ray traverses through the nodes of the BVH to efficiently identify the objects that may be intersected by the ray.
[0497] Modern real-time graphics APIs define the acceleration structure (AS) as opaque as possible to allow hardware-vendor-specific implementations. However, this AS generalization limits the complexity of the data structures that can be used and the programmability of the traversal operations.
[0498] Level-of-detail (LoD) techniques are commonly used in most real-time rendering systems to limit the memory footprint and traversal cost of scene surfaces and volumes and push the limits of visual complexity. Using LoD techniques, the complexity of the 3D model representation decreases as instances of the 3D model become farther from the observer. Until recently, LoD techniques were implemented at the per-instance level, where the LoD was adjusted based on the distance of the object instance from the observer.
[0499] Most real-time LoD methods operate at the per-instance level, replacing complex surfaces or volumes with coarser representations based on heuristic rules. In ray-tracing APIs, this involves replacing references to the bottom-level acceleration structure (BLAS) in the top-level acceleration structure (TLAS). However, such sudden replacements often result in visually distracting "popping" artifacts. Traversal shaders are programmable mechanisms that allow per-ray BLAS selection and can also implement stochastic LoD transitions. An alternative approach is to use instance masks.
[0500] A new emerging trend is to use hierarchical LoD structures to directly render large-scale micro-polygonal surfaces. A well-known example of adaptive micro-polygonal LoD rendering is Nanite in the Unreal Engine, which generates a hierarchy of micro-polygonal clusters in a preprocessing step, forming a directed acyclic graph (DAG) structure. See, e.g., Brian Karis et al., Nanite, A Deep Dive, Advances in Real-Time Rendering Course, Siggraph (2021). Before rendering, LoD selection identifies view-dependent cuts in the DAG, and micro-polygonal clusters at the same DAG level form a seamless continuous surface. Other in-mesh dynamic LoD techniques, such as adaptive refinement, are less flexible and require higher-level representations, such as parametric surfaces.
[0501] In-mesh micro-polygonal LoD requires frequent changes to the "active" micro-polygonal clusters, resulting in the reconstruction of the TLAS and BLAS. They are not suitable for current API and HW ray-tracing implementations and are limited to specialized software rasterization. These in-mesh LoD techniques do not prevent "popping" artifacts during LoD transitions. Instead, such artifacts are mitigated by using sub-pixel-sized polygons and temporal filtering.
[0502] Instance-based LoD solutions, such as traversal shaders or instance mask tests, are not suitable for in-mesh LoD. Although different clusters can technically be treated as independent instances, this would be an inefficient solution, resulting in tiny bottom-level acceleration structures in a very large monolithic TLAS. Additionally, treating clusters as instances would limit the applicability of these techniques (e.g., LoD transitions would require more than two levels of acceleration structures). Furthermore, the standard two-level hierarchy of TLAS and BLAS is not suitable for LoD schemes of complex surfaces, as any topological change would require a complete reconstruction of the BLAS of the affected geometry.
[0503] To address these limitations, embodiments of the present invention extend the concept of instance masks to the level of multi-LoD interior nodes in a BVH. Compared to standard interior BVH nodes, these novel node types occupy a larger memory footprint, but allow for efficient routing of rays between two child nodes of the same bounds based on per-ray bitmask comparison. This mechanism allows for random LoD selection within a BLAS or TLAS without the need for programmable shaders.
[0504] Some embodiments use a new interior LoD node type (referred to herein as a multi-LoD node or dual child) during ray tracing to dynamically select LoD within a mesh. Just like a regular BVH interior node, it has a set of bounding boxes defined for its child nodes, and ray traversal enters a child node when it intersects with its bounding box. However, a multi-LoD node is a bounding box with two corresponding child nodes. After a successful intersection test, a binary selection mechanism determines which child node for a given ray traversal.
[0505] In terms of the ray traversal mechanism, one embodiment associates a bitmask with each interior LoD node, and this bitmask is compared with a per-ray bitmask to perform a binary selection of one of the multi-LoD nodes. As long as the first half of the multi-LoD node always corresponds to the coarser LoD of the subtree, this single comparison will result in a consistent LoD selection for all child nodes of the same node.
[0506] Figure 44 An example BVH 4400 is illustrated having two multi-LoD nodes 4410-4411 located below a parent node 4400. Each multi-LoD node 4410-4411 includes a plurality of child nodes 4431-4432, each child node associated with a different LoD. In particular, a child node 4431 associated with a first LoD and a child node 4432 associated with a second LoD (LoD2) are included in the multi-LoD node 4411 in the hierarchy. Similarly, a child node 4433 associated with the first LoD and a child node 4434 associated with the second LoD (LoD2) are included in the multi-LoD node 4410 in the hierarchy. In one embodiment, the first LoD (LoD1) includes a relatively coarser LoD compared to the second LoD (LoD2), which provides greater precision when performing BVH traversal operations (e.g., where its bounding box is subdivided into a larger number of sub-bounding boxes).
[0507] In operation, if the traversal unit 4453 determines that the multi-LoD node 4411 is to be traversed by a ray (e.g., based on a ray / box test as described herein), then the LoD node bitmask 4416 is used to select one of the two child nodes 4431 - 4432. In particular, based on the comparison of the associated LoD node bitmask 4416 with each ray bitmask 4422, one of the two child nodes 4431 - 4432 of the multi-LoD node 4411 is selected for traversal. Similarly, if the traversal unit 4453 determines that the multi-LoD node 4410 is to be traversed by a ray, then based on the comparison of the associated LoD node bitmask 4415 with each ray bitmask 4422, only one of the two child nodes 4432 - 4433 of the multi-LoD node 4410 is selected for traversal.
[0508] In operation, each ray 4452 generated by the ray generation logic 4450 has an associated per-ray bitmask 4422. During ray traversal through the BVH by the traversal unit 4453, the comparison logic 4458 compares the LoD node bitmask 4416 with the per-ray bitmask 4422 to determine which child node under the multi-LoD node 4410 or multi-LoD node 4411 to select for further traversal. For example, if the ray 4452 is determined to pass through the bounding box of the multi-LoD node 4411, then the LoD node bitmask 4416 is compared with the per-ray bitmask 4422 to determine whether to continue traversing the first child node 4431 (at LoD1) or the second child node 4432 (at LoD2).
[0509] Various types of comparison operations can be performed by the comparison logic 4458, including but not limited to less_equal and greater. If the comparison operation returns true (i.e., the comparison requirement is satisfied), then the first child node 4431 is used to continue traversal at LoD1; otherwise, the second child node 4432 is used for traversal at LoD2.
[0510] Traversal continues normally until an exit condition is reached. For example, subsequent traversal / intersection operations can be performed as described with respect to Figures 40 - 48 The underlying principles of the present invention are not limited to any particular subsequent traversal / intersection operation.
[0511] In some implementations, the micro-polygon mesh can have multiple associated LoDs, where the current LoD can be selected based on bitmask comparison operations such as those described above and / or based on other variables (e.g., such as the current distance from the observer).
[0512] Figure 45 A method according to an embodiment of the present invention is shown. The method can be implemented on various architectures described herein, but is not limited to any particular architecture.
[0513] At 4501, a BVH including one or more multi-LoD nodes is constructed based on the primitives / meshes of the current graphics scene. At 4502, rays for traversing through the BVH are generated, and at 4502, the rays traverse through the BVH nodes in the hierarchy. When reaching the multi-LoD node determined at 4503, at 4504, each per-ray bitmask associated with the ray is compared with the multi-LoD node bitmask associated with the LoD node to select the child node and the corresponding LoD to be used. At 4505, the traversal continues through the selected child node / LoD and potentially other nodes of the BVH until an exit condition is reached.
[0514] Once the child node has been traversed, if it is determined at 4506 that the traversal is not complete, the process returns to 4502, and the traversal continues through the next BVH node. When the traversal is complete, at 4507, the intersection with the mesh / primitive is determined (or if the ray misses the mesh / primitive, it is determined that there is no intersection), and the process returns to 4502.
[0515] Embodiments of the present invention may depend on different memory layout configurations for implementation. To support a cluster hierarchy without always copying primitive leaves, one embodiment allows a primitive leaf to be referenced from multiple parent LoD nodes (e.g., two child nodes with different LoDs). In these implementations, the micro-polygon cluster hierarchy is a directed acyclic graph (DAG), which is the main mechanism for high-quality LoD clustering and seamless surface refinement when rendering micro-polygon clusters from cuts of the DAG. However, this requires storing pointers or at least integer offsets for all child nodes in the multi-LoD nodes 4410 - 4411, which makes their traversal more expensive and requires a larger memory footprint compared to standard internal nodes.
[0516] One embodiment improves the rendering performance of the nanocrystalline micro-polygon cluster hierarchy in two key ways. First, the frequency of BLAS reconstruction is reduced. Since multiple LoDs exist within the same BLAS (e.g., child nodes 4431 - 4434), once a new LoD (e.g., LoD3) is needed, only the BLAS needs to be reconstructed. This reduces the BVH construction cost but also increases some overhead in the traversal itself because larger LoD nodes occupy more memory bandwidth and cache. To address this issue, the presence of such LoD nodes is limited to only 2 - 3 adjacent LoD levels within the entire BVH and cluster hierarchy.
[0517] Second, "popping" artifacts are eliminated through random LoD transitions. In particular, these embodiments provide a pop-free random transition between cluster levels, the quality of which depends only on the number of bits allowed in the masks 4415 - 4416 of the multi-LoD interior nodes. This can not only improve the quality of surface rendering, but also produce better performance by allowing larger polygons with fewer noticeable differences.
[0518] Regarding BVH construction considerations, API extensions are provided to allow developers to utilize the multi-LoD nodes 4410 - 4411 as described herein. In some implementations, the BVH builder 4490 receives a list of primitives and their bounding boxes and uses heuristics such as SAH to construct the BVH. In the case of micropolygon clusters, the BVH generator 4490 considers the bounding boxes of the coarsest "parent" primitives on which a classical BVH structure can be built. Each of these parent clusters is used as a coarser child node (e.g., child node 4431) within the aforementioned multi-LoD nodes 4410 - 4411, and their finer LoD sibling nodes (e.g., child node 4432) are interior nodes that continue traversal to the finer clusters of each "parent" primitive. Thus, in one embodiment, the BVH builder 4490 is configured with the relationship between the parent and child clusters to achieve an efficient memory layout and also ignores the finer LoD primitives when constructing the initial standard BVH.
[0519] While embodiments of the present invention have been described above in the context of micropolygon meshes, the basic principles of the present invention are not so limited. For example, the techniques described above can be used in other hierarchical LoD selection schemes, including but not limited to sparse volumes. Apparatus and method for a manageable segmented acceleration structure
[0520] Acceleration structures (AS) for real-time ray tracing are mainly implemented as bounding volume hierarchies (BVH) and are managed by a two-level macrostructure. The underlying AS (BLAS) is referenced in the leaves of the top-level AS (TLAS). This reference allows rays to be transformed from world space (TLAS space) to object space (BLAS space).
[0521] This pattern is challenged by two trends. First, applications need to maintain geometric structures in fragments and expect the fragments to quickly transition to higher or lower levels of detail (LOD). To reflect this, applications currently need to reconstruct the entire BLAS even when only fragments of its geometry are replaced whenever the number of primitives or connectivity (e.g., the way vertices are shared across triangles) changes. Additionally, as the TLAS grows larger, changes that only affect fragments of the scene require the AS to be reconstructed or at least updated. These issues become even more challenging when applications treat both the BLAS and TLAS as single opaque entities that can only be reconstructed or updated as a whole.
[0522] One solution to the first problem is to build separate BLASs for each LOD fragment, which places a great deal of pressure on the TLAS builder (and its performance requirements) since the TLAS is built on the leaves of a larger group. A solution to the second problem is to maintain two TLASs, one for dynamic objects and another for static objects. To obtain the closest hit for a given ray, the application must trace both TLASs (a second pass can cull intersections, searching a second part of the TLAS that is farther than the hit found in the first pass), select the closer intersection, and then somehow manually execute the closest hit shader. Another possible solution would be to use multi-level instancing or programmable instancing.
[0523] However, these solutions all come at a cost. Multi-level instancing is difficult because it requires maintaining the state of each transformed ray for each level of instancing traversing the input. It also requires complex calculations to obtain the object-to-world matrix in the shader. The object-to-world matrix is the superposition of all the inverse transformation matrices of all the instances along the traversal path of the ray. Such a matrix will likely have to be calculated in the shader at a cost of computation and bandwidth utilization. Programmable instancing requires emitting shaders at each entry of the programmable instance, which consumes a great deal of time.
[0524] Embodiments of the present invention solve this problem by representing the BVH as separate parts that can be treated by the application as separate smaller objects. These embodiments include improvements to traditional TLAS or BLAS systems and have fewer problems than multi-level instancing. As used herein, an "AS-let", an acceleration structure that can be used as a leaf of a higher-level acceleration structure, is referred to as an "acceleration structure link" or "AS link". An AS link can also be an AS-let, so multi-level linking is possible. An AS link can also directly have leaves that are not ASs, such as primitives or instances.
[0525] In Figure 46 An example arrangement is illustrated, including a link node 4601 at the top / root of a hierarchy, which has two direct child nodes: a link AS-let node 4602 and an AS-let node 4603. The link AS-let node 4602 forms links to two additional AS-let nodes 4604, while the AS-let node 4603 is a leaf node.
[0526] In some embodiments of the present invention, the acceleration structure is a BVH having internal nodes that allow full memory jumps for each of the child nodes and / or allow limited jumps of relative offsets for child nodes that are closely placed together within a defined memory region. These embodiments solve the above-mentioned manageability problem by providing a way to express acceleration structure fragments. In addition, these em...
Claims
1. A device comprising: An acceleration structure construction logic is configured to construct an acceleration structure (AS) including a multi-level linked hierarchy structure having different types of AS fragments, the different types of AS fragments including a first type of AS fragments having leaves, a second type of AS fragments including AS links, and a third type of AS fragments including both leaves and AS links, wherein to construct the AS fragments, the acceleration structure construction logic is to: Evaluate multiple entity references, Determine whether each primitive reference refers to a primitive or an AS fragment, and If the primitive reference indicates an AS fragment, encoding a pointer or offset directly or indirectly into a bounding volume hierarchy (BVH) of the AS fragment; as well as Traversal hardware logic for traversing light rays through the AS.
2. The device according to claim 1, wherein: A primitive reference indicating that a primitive is to refer to a triangle, instance, or procedural primitive.
3. The device according to claim 1 or 2, wherein: Each AS fragment of the second type comprises one or more links to other AS fragments and / or to objects.
4. The device according to any one of claims 1 to 3, wherein: Each AS fragment of the third type comprises one or more links to AS fragments of the first type.
5. The device according to any one of claims 1 to 4, wherein: To construct the AS, the AS construction hardware logic generates internal nodes of the AS fragment of the second type or the third type from a root node of any type of AS fragment.
6. The device according to claim 5, wherein: The AS construction hardware logic is to generate an internal node of the AS fragment, wherein the internal node includes a jump reference to one or more BVH nodes of the AS fragment.
7. The device according to claim 5, wherein: The generated internal node is to be stored in a sub-area contained in the AS fragment of the second type or the third type adjacent to one or more other internal nodes, primitives or instances.
8. A method comprising: Building an acceleration structure (AS), the acceleration structure including a multi-level linked hierarchy having different types of AS fragments, the different types of AS fragments including a first type of AS fragment having leaves, a second type of AS fragment including AS links, and a third type of AS fragment including both leaves and AS links, the AS fragments to be built by performing the following operations: Evaluate multiple entity references, Determine whether each primitive reference refers to a primitive or an AS fragment, and If the primitive reference indicates an AS fragment, encoding a pointer or offset directly or indirectly into a bounding volume hierarchy (BVH) of the AS fragment; and traversing the ray through the AS.
9. The method of claim 8, wherein: A primitive reference indicating that a primitive is to refer to a triangle, instance, or procedural primitive.
10. The method according to claim 8 or 9, wherein: Each AS fragment of the second type comprises one or more links to other AS fragments and / or to objects.
11. The method according to any one of claims 8 to 10, wherein: Each AS fragment of the third type comprises one or more links to AS fragments of the first type.
12. The method according to any one of claims 8 to 11, wherein: To construct the AS, internal nodes of the AS fragment of the second type or the third type are generated from a root node of an AS fragment of any type.
13. The method according to any one of claims 8 to 12, wherein: To construct the AS, internal nodes of the AS fragment are generated, the internal nodes including jump references to one or more BVH nodes of the AS fragment.
14. The method of claim 12, wherein: The generated internal node is to be stored in a sub-area contained in the AS fragment of the second type or the third type adjacent to one or more other internal nodes, primitives or instances.
15. A machine-readable medium having program code stored thereon, the program code, when executed by a machine, causing the machine to perform the following operations: Building an acceleration structure (AS), the acceleration structure including a multi-level linked hierarchy having different types of AS fragments, the different types of AS fragments including a first type of AS fragment having leaves, a second type of AS fragment including AS links, and a third type of AS fragment including both leaves and AS links, the AS fragments to be built by performing the following operations: Evaluate multiple entity references, Determine whether each primitive reference refers to a primitive or an AS fragment, and If the primitive reference indicates an AS fragment, encoding a pointer or offset directly or indirectly into a bounding volume hierarchy (BVH) of the AS fragment; and Traverse the ray through the AS.
16. The machine-readable medium of claim 15, wherein: A primitive reference indicating that a primitive is to refer to a triangle, instance, or procedural primitive.
17. The machine-readable medium of claim 15 or 16, wherein: Each AS fragment of the second type comprises one or more links to other AS fragments and / or to objects.
18. The machine-readable medium of any one of claims 15 to 17, wherein: Each AS fragment of the third type comprises one or more links to AS fragments of the first type.
19. The machine-readable medium of any one of claims 15 to 18, wherein: To construct the AS, internal nodes of the AS fragment of the second type or the third type are generated from a root node of an AS fragment of any type.
20. The machine-readable medium of any one of claims 15 to 19, wherein: To construct the AS, internal nodes of the AS fragment are generated, the internal nodes including jump references to one or more BVH nodes of the AS fragment.
Citation Information
Patent Citations
Apparatus and method for compressing leaf nodes of a bounding volume hierarchy (BVH)
US20190318445A1