Apparatus and method for displacement mesh compression
The method addresses the challenge of compressing displacement meshes by employing a graphics processor with integrated compression circuits and algorithms, resulting in improved performance and memory efficiency for rendering complex scenes.
Patent Information
- Application Number
- JP2020212137
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2020-03-15
- Filing Date
- 2020-12-22
- Publication Date
- 2025-05-26
- Estimated Expiration
- 2040-12-22
AI Technical Summary
Current graphics processing technologies face challenges in efficiently compressing displacement meshes, which are essential for rendering photo-realistic images using path tracing techniques.
The proposed solution involves an apparatus and method for displacement mesh compression that utilizes specific technological measures, including the use of a graphics processor with integrated compression circuits and algorithms tailored for displacement mapping, to reduce the complexity and size of displacement meshes.
This approach significantly enhances the performance of graphics processing by reducing memory usage and computational load, thereby enabling faster rendering of complex scenes with high visual fidelity.
Smart Images

Figure 0007682625000008 
Figure 0007682625000009 
Figure 0007682625000010
Abstract
Description
Technical Field
[0001] The present invention generally relates to the field of graphics processors. More particularly, the present invention relates to an apparatus and method for displacement mesh compression.
Background Art
[0002] Path tracing is a technique for rendering photo-realistic images for movies, animated movies, and special effects in professional visualization. Generating these realistic images requires a physical simulation of light propagation in a virtual 3D scene using ray tracing for visibility queries. A high-performance implementation of these visibility queries requires the construction of a 3D hierarchy (typically a bounding volume hierarchy) on scene primitives (typically triangles) in a preprocessing stage. The hierarchy enables the renderer to quickly determine the closest intersection between a ray and a primitive (triangle).
Brief Description of the Drawings
[0003] A better understanding of the present invention can be obtained from the following detailed description in conjunction with the following drawings.
Figure 1
Figure 2A
Figure 2B
Figure 2C
Figure 2D
Figure 3A
Figure 3B
Figure 3C
Figure 4
Figure 5A
Figure 5B
Figure 6
Figure 7
Figure 8
Figure 9A
Figure 9B
Figure 10
Figure 11A
Figure 11B
Figure 11C
Figure 11D
Figure 12
Figure 13
Figure 14
Figure 15
Figure 16
Figure 17
Figure 18A
Figure 18B
Figure 19
Figure 20
Figure 21
Figure 22
Figure 23
Figure 24
Figure 25
Figure 26
Figure 27
Figure 28
Figure 29
Figure 30A
Figure 30B
Figure 31
Figure 32
Figure 33
Figure 34
Figure 35
Figure 36
Figure 37
Figure 38
Figure 39
Figure 40
Figure 41
Figure 42
Figure 43
Figure 44
Figure 45
Figure 46A
Figure 46B
Figure 47
Figure 48
Figure 49
Figure 50
Figure 51
Figure 52A
Figure 52B
Figure 53
Figure 54
Figure 55
Figure 56
Figure 57
Figure 58
Figure 59
Figure 60
Figure 61
Figure 62
Figure 63
Figure 64A
Figure 64B
Figure 64C
Figure 65
DETAILED DESCRIPTION OF THE INVENTION
[0004] In the following description, for purposes of explanation, in order to provide a thorough understanding of the embodiments of the present invention described below, a plurality of specific details are set forth. However, it will be apparent to one of ordinary skill in the art that the embodiments of the present invention may be practiced without some of these specific details. In other instances, well-known structures and devices are shown in block diagram form in order to avoid obscuring the underlying principles of the embodiments of the present invention.
[0005] [Exemplary Graphics Processor Architecture and Data Types] (System Overview) FIG. 1 is a block diagram of a processing system 100 according to an embodiment. The system 100 may be used in a single-processor desktop system, a multi-processor workstation system, or a server system having a number of processors 102 or processor cores 107. In one embodiment, the system 100 is a processing platform incorporated within a system-on-a-chip (SoC) integrated circuit for use in mobile, handheld, or embedded devices, such as within an Internet-of-things (IoT) device having a wired or wireless connection to a local or wide area network.
[0006] In one embodiment, system 100 can include a server-based game platform, a game console including a game and media console, a mobile game console, a handheld game console, or an online game console, and can be coupled to or integrated within these. In some embodiments, system 100 is part of a mobile Internet-connected device such as a cell phone, smartphone, tablet computing device, or laptop having a small internal memory capacity. Processing system 100 can also include wearable devices such as smartwatch wearable devices; smart eyewear or clothing enhanced with augmented reality (AR) or virtual reality (VR) capabilities to provide visual, audio, tactile outputs to supplement real-world visual, audio, tactile experiences, or to provide text, audio, graphics, video, holographic images or video, or tactile feedback; other augmented reality (AR) devices; or other virtual reality (VR) devices, and can be coupled to or integrated within these. In some embodiments, processing system 100 includes or is part of a television or set-top box device. In one embodiment, system 100 can include, be coupled to, or be integrated within an autonomous vehicle such as a bus, tractor trailer, automobile, motor or power cycle, airplane or glider (or any combination thereof). The autonomous vehicle may use system 100 to process the environment sensed around the vehicle.
[0007] In some embodiments, one or more processors 102 each include one or more processor cores 107 for processing instructions that, when executed, perform operations for system or user software. In some embodiments, at least one of the one or more processor cores 107 is configured to process a particular instruction set 109. In some embodiments, the instruction set 109 may implement computing via a Complex Instruction Set Computing (CISC), Reduced Instruction Set Computing (RISC), or Very Long Instruction Word (VLIW). The one or more processor cores 107 may process different instruction sets 109, and the different instruction sets 109 may include instructions for implementing emulation of other instruction sets. The processor core 107 may also include other processing devices such as a Digital Signal Processor (DSP).
[0008] In some embodiments, processor 102 includes cache memory 104. Depending on the architecture, processor 102 can have a single internal cache or multiple levels of internal caches. In some embodiments, the cache memory is shared among various components of processor 102. In some embodiments, processor 102 also uses an external cache (e.g., a Level-3 (L3) cache or a Last Level Cache (LLC)) (not shown), which may be shared among processor cores 107 using known cache coherency techniques. Register file 106 can further be included in processor 102 and may include different types of registers for storing different types of data (e.g., integer registers, floating-point registers, status registers, and instruction pointer registers). Some of the registers may be general-purpose registers, and other registers may be specific to the design of processor 102.
[0009] In some embodiments, one or more processors 102 are coupled to one or more interface buses 110 to transmit communication signals, such as address, data, or control signals, between the processor 102 and other components within the system 100. In one embodiment, the interface bus 110 can be a processor bus, such as a version of the Direct Media Interface (DMI) bus. However, the processor bus is not limited to the DMI bus and may include one or more peripheral component interconnect buses (e.g., PCI, PCI express), memory buses, or other types of interface buses. In one embodiment, the processor 102 includes an integrated memory controller 116 and a platform controller hub 130. The memory controller 116 enables communication between the memory devices and other components of the system 100, and the platform controller hub (PCH) 130 provides connections to I / O devices via a local I / O bus.
[0010] The memory device 120 can be a dynamic random access memory (DRAM) device, a static random access memory (SRAM) device, a flash memory device, a phase change memory device, or other memory devices having performance suitable for functioning as a process memory. In one embodiment, the memory device 120 can operate as a system memory for the system 100 to store data 122 and instructions 121 for use when one or more processors 102 execute an application or process. The memory controller 116 can also be coupled to an optional external graphics processor 118, and the external graphics processor 118 may communicate with one or more graphics processors 108 within the processor 102 to execute graphics and media operations. In some embodiments, the graphics, media, or compute operations may be assisted by an accelerator 112, which is a coprocessor configured to execute a special set of graphics, media, or compute operations. For example, in one embodiment, the accelerator 112 is a matrix multiplication accelerator used to optimize machine learning or compute operations. In one embodiment, the accelerator 112 is a ray tracing accelerator that can be used to execute ray tracing operations in cooperation with the graphics processor 108. In one embodiment, an external accelerator 119 may be used instead of or in cooperation with the accelerator 112.
[0011] In some embodiments, the display device 111 can be connected to the processor 102. The display device 111 can be one or more internal display devices, such as those of a mobile electronic device or a laptop device, or one or more of external display devices attached via a display interface (e.g., DisplayPort, etc.). In one embodiment, the display device 111 can be a head mounted display (HMD), such as a stereoscopic display device for use in a virtual reality (VR) application or an augmented reality (AR) application.
[0012] In some embodiments, the platform controller hub 130 enables peripheral devices to be connected to the memory device 120 and the processor 102 via a high-speed I / O bus. The I / O peripheral devices include, but are not limited to, an audio controller 146, a network controller 134, a firmware interface 128, a wireless transceiver 126, a touch sensor 125, and a data storage device 124 (e.g., non-volatile memory, volatile memory, hard disk drive, flash memory, NAND, 3D NAND, 3D XPoint, etc.). The data storage device 124 can be connected via a peripheral device bus such as a storage interface (e.g., SATA) or a Peripheral Component Interconnect bus (e.g., PCI, PCI express). The touch sensor 125 can include a touch screen sensor, a pressure sensor, or a fingerprint sensor. The wireless transceiver 126 can be a Wi-Fi transceiver, a Bluetooth transceiver, or a mobile network transceiver such as a 3G, 4G, 5G, or Long-Term Evolution (LTE) transceiver. The firmware interface 128 enables communication with system firmware and can be, for example, a unified extensible firmware interface (UEFI). The network controller 134 can enable a network connection to a wired network. In some embodiments, a high-performance network controller (not shown) is coupled to the interface bus 110. In one embodiment, the audio controller 146 is a multi-channel high-definition audio controller. In one embodiment, the system 100 includes an optional legacy I / O controller 140 for coupling legacy (e.g., Personal System 2 (PS / 2)) devices to the system.The platform controller hub 130 can also be connected to one or more Universal Serial Bus (USB) controllers 142, and can connect input devices such as a combination of a keyboard and a mouse 143, a camera 144, or other USB input devices.
[0013] It is recognized that the illustrated system 100 is exemplary and not limiting. This is because other types of data processing systems configured differently may also be used. For example, instances of the memory controller 116 and the platform controller hub 130 may be integrated into an individual external graphics processor such as the external graphics processor 118. In one embodiment, the platform controller hub 130 and / or the memory controller 116 may be external to one or more processors 102. For example, the system 100 can include an external memory controller 116 and a platform controller hub 130, which may be configured as a memory controller hub and a peripheral device controller hub in a system chipset that communicates with the processor 102.
[0014] For example, a circuit board ("sled") on which components such as CPUs, memories, and other components designed for further thermal performance are placed can be used. In some examples, processing components such as processors are placed on the top surface of the sled, and near memories such as DIMMs are placed on the bottom surface of the sled. As a result of the increased airflow provided by this design, the components may operate at higher frequencies and power levels than in a typical system, thereby increasing performance. Further, the sled is configured to blindly mate with the power and data communication cables within the rack, thereby enhancing the ability to quickly remove, upgrade, reinstall, and / or replace. Similarly, individual components placed on the sled, such as processors, accelerators, memories, and data storage drives, are configured to be easily upgradable because the spacing between them is increased. In an exemplary embodiment, the components further include a hardware authentication function to prove their authenticity.
[0015] The data center can utilize a single network architecture ("fabric") that supports multiple other network architectures including Ethernet and Omni-Path. The sled can be coupled to the switch via an optical fiber that provides higher bandwidth and lower latency than a typical twisted pair cable (e.g., Category 5, Category 5e, Category 6, etc.). Due to the high bandwidth, low latency interconnectivity and network architecture, the data center can pool resources such as physically disaggregated memories, accelerators (e.g., GPUs, graphics accelerators, FPGAs, ASICs, neural networks, and / or artificial intelligence accelerators, etc.) and data storage drives during use and provide them to computing resources (e.g., processors) as needed, enabling the computing resources to access the pooled resources as if they were local.
[0016] Power supply or power source can provide voltage and / or current to system 100 or any component or system described herein. In one example, the power supply includes an AC-DC (alternating current - direct current) adapter for plugging into an outlet. Such AC power can be from a renewable energy source (e.g., solar power). In one example, the power source includes a DC power source such as an external AC-DC converter. In one example, the power source or power supply includes wireless charging hardware for charging via the vicinity of a charging field. In one example, the power source can include an internal battery, an alternating current power source, a motion-based power source, a solar power source, or a fuel cell power source.
[0017] Figures 2A - 2D illustrate a computing system and a graphics processor provided according to the embodiments described herein. Elements of Figures 2A - 2D having the same reference numerals (or names) as elements in any other drawing herein can operate or function in a similar manner as those described elsewhere herein, but are not limited thereto.
[0018] FIG. 2A is a block diagram of an embodiment of a processor 200 having one or more processor cores 202A-202N, an integrated memory controller 214, and an integrated graphics processor 208. The processor 200 can include additional cores, including an additional core 202N represented by the dashed box. Each of the processor cores 202A-202N includes one or more internal cache units 204A-204N. In some embodiments, each processor core also has access to one or more shared cache units 206. The internal cache units 204A-204N and the shared cache unit 206 represent the cache memory hierarchy within the processor 200. The cache memory hierarchy can include at least one level of instruction and data cache within each processor core and one or more levels of shared intermediate-level cache, such as level 2 (L2), level 3 (L3), level 4 (L4), or other levels of cache, where the highest level of cache before external memory is classified as the LLC. In some embodiments, cache coherence logic maintains coherence among the various cache units 206 and 204A-204N.
[0019] In some embodiments, the processor 200 can include a set of one or more bus controller units 216 and a system agent core 210. The one or more bus controller units 216 manage a set of peripheral device buses, such as one or more PCI or PCI Express buses. The system agent core 210 provides management functions for the various processor components. In some embodiments, the system agent core 210 includes one or more integrated memory controllers 214 to manage access to various external memory devices (not shown).
[0020] In some embodiments, one or more of the processor cores 202A-202N include support for simultaneous multithreading. In such embodiments, the system agent core 210 includes components for coordinating the operation of cores 202A-202N during multithreaded processing. The system agent core 210 may further include a power control unit (PCU), and the power control unit (PCU) includes logic and components for adjusting the power states of the processor cores 202A-202N and the graphics processor 208.
[0021] In some embodiments, the processor 200 further includes a graphics processor 208 for performing graphics processing operations. In some embodiments, the graphics processor 208 couples to a set of shared cache units 206 and system agent cores 210, including one or more integrated memory controllers 214. In some embodiments, the system agent core 210 also includes a display controller 211 for driving the graphics processor output to one or more attached displays. In some embodiments, the display controller 211 may also be a separate module coupled to the graphics processor via at least one interconnect, or may be integrated within the graphics processor 208.
[0022] In some embodiments, a ring-based interconnect unit 212 is used to couple the internal components of the processor 200. However, alternative interconnect units such as point-to-point interconnects, switched interconnects, or other techniques well known in the art may be used. In some embodiments, the graphics processor 208 couples to the ring interconnect 212 via an I / O link 213.
[0023] The exemplary I / O link 213 represents at least one of multiple types of I / O interconnects, including an on-package I / O interconnect that enables communication between various processor components and high-performance embedded memory modules 218 such as eDRAM modules. In some embodiments, each of the processor cores 202A-202N and the graphics processor 208 can use the embedded memory module 218 as a shared last-level cache.
[0024] In some embodiments, the processor cores 202A-202N are homogeneous cores that execute the same instruction set architecture. In other embodiments, the processor cores 202A-202N are heterogeneous with respect to instruction set architecture (ISA), and one or more of the processor cores 202A-202N execute a first instruction set while at least one of the other cores executes a subset of the first instruction set or a different instruction set. In one embodiment, the processor cores 202A-202N are heterogeneous with respect to microarchitecture, and one or more cores with relatively high power consumption are coupled with one or more power cores with lower power consumption. In one embodiment, the processor cores 202A-202N are heterogeneous with respect to computing power. Further, the processor 200 can be implemented on one or more chips having the illustrated components in addition to other components or as a SoC integrated circuit.
[0025] FIG. 2B is a block diagram of the hardware logic of a graphics processor core 219 according to some embodiments described herein. Elements of FIG. 2B having the same reference numerals (or names) as elements in any other drawing herein may operate or function in a similar manner as those described elsewhere herein, but are not limited thereto. The graphics processor core 219, sometimes referred to as a core slice, may be one or more graphics cores within a modular graphics processor. The graphics processor core 219 is an example of one graphics core slice, and the graphics processors described herein may include multiple graphics core slices based on a target power and performance envelope. Each graphics processor core 219 may include a fixed function block 230 coupled to a plurality of sub-cores 221A - 221F, also referred to as sub-slices, where the plurality of sub-cores 221A - 221F include modular blocks of general purpose and fixed function logic.
[0026] In some embodiments, the fixed function block 230 may include, for example, a geometry / fixed function pipeline 231 that can be shared by all sub-cores within the graphics processor core 219 in a low performance and / or low power graphics processor implementation. In various embodiments, the geometry / fixed function pipeline 231 includes a 3D fixed function pipeline (e.g., a 3D pipeline 312 as in FIGS. 3 and 4 described below), a video front end unit, a thread generator and thread dispatcher, and a unified return buffer manager that manages a unified return buffer (e.g., unified return buffer 418 in FIG. 4 described below).
[0027] In one embodiment, the fixed function block 230 also includes a graphics SoC interface 232, a graphics microcontroller 233, and a media pipeline 234. The graphics SoC interface 232 provides an interface between the graphics processor core 219 and other processor cores within the system-on-chip integrated circuit. The graphics microcontroller 233 is a programmable sub-processor configurable to manage various functions of the graphics processor core 219, including thread dispatch, scheduling, and preemption. The media pipeline 234 (e.g., the media pipeline 316 of FIGS. 3 and 4) includes logic for performing decoding, encoding, pre-processing, and / or post-processing of multimedia data, including image and video data. The media pipeline 234 executes media operations via requests to compute or sample logic within the sub-cores 221-221F.
[0028] In one embodiment, the SoC interface 232 enables the graphics processor core 219 to communicate with general-purpose application processor cores (e.g., CPUs) and / or other components within the SoC that include memory hierarchy elements such as shared last-level cache memory, system RAM, and / or embedded on-chip or on-package DRAM. The SoC interface 232 can also enable communication with fixed-function devices within the SoC, such as a camera image processing pipeline, and enable and / or implement the use of global memory atomics that can be shared between the graphics processor core 219 and the CPU within the SoC. The SoC interface 232 also implements power management control for the graphics processor core 219 and can enable an interface between the clock domain of the graphics core 219 and other clock domains within the SoC. In one embodiment, the SoC interface 232 enables receipt of command buffers from a command streamer and a global thread dispatcher configured to provide commands and instructions to each of one or more graphics cores within the graphics processor. The commands and instructions can be dispatched to the media pipeline 234 when media operations are executed and to the geometry and fixed-function pipelines (e.g., geometry and fixed-function pipeline 231, geometry and fixed-function pipeline 237) when graphics processing operations are executed.
[0029] The graphics microcontroller 233 can be configured to perform various scheduling and management tasks for the graphics processor core 219. In one embodiment, the graphics microcontroller 233 can perform graphics and / or compute workload scheduling on the execution unit (EU) arrays 222A - 222F, 224A - 224F within the sub - cores 221A - 221F on various graphics parallel engines. In this scheduling model, host software running on the CPU core of the SoC that includes the graphics processor core 219 can submit a workload to one of the doorbells of the multiple graphics processors, which invokes a scheduling operation on the appropriate graphics engine. The scheduling operation includes determining which workload to execute next, submitting the workload to the command streamer, pre - empting existing workloads running on the engine, monitoring the progress of the workload, and notifying the host software when the workload is complete. In one embodiment, the graphics microcontroller 233 can also implement a low - power or idle state of the graphics processor core 219 and provide the graphics processor core 219 with the ability to save and restore registers within the graphics processor core 219 during low - power state transitions, independent of the operating system and / or graphics driver software on the system.
[0030] The graphics processor core 219 may have sub-cores 221A-221F larger or smaller than those illustrated, and may have up to N modular sub-cores. For each set of N sub-cores, the graphics processor core 219 may also include shared function logic 235, shared and / or cache memory 236, geometry / fixed function pipeline 237, and further fixed function logic 238 for accelerating various graphics and compute processing operations. The shared function logic 235 may include logic units associated with the shared function logic 420 of FIG. 4 (e.g., sampler, arithmetic, and / or inter-thread communication logic) that can be shared by each of the N sub-cores within the graphics processor core 219. The shared and / or cache memory 236 may serve as a last-level cache for the set of N sub-cores 221A-221F within the graphics processor core 219, and may also function as shared memory accessible by multiple sub-cores. The geometry / fixed function pipeline 237 may be included instead of the geometry / fixed function pipeline 231 within the fixed function block 230, and may include the same or similar logic units.
[0031] In one embodiment, the graphics processor core 219 includes further fixed function logic 238 that can include various fixed function acceleration logic used by the graphics processor core 219. In one embodiment, the further fixed function logic 238 includes a further geometry pipeline for use in position only shading. In position only shading, there are two geometry pipelines, namely, the cull pipeline, which may be included within the geometry / fixed function pipelines 238, 231, all the geometry pipelines within the fixed function pipelines 238, 231, and a further geometry pipeline within the further fixed function logic 238. In one embodiment, the cull pipeline is a trimmed version of all the geometry pipelines. The full pipeline and the cull pipeline can execute different instances of the same application, and each instance has a separate context. Position only shading can hide the long culling execution of discarded triangles and, in some cases, enable shading to be completed faster. For example, in one embodiment, the culling pipeline logic within the further fixed function logic 238 can execute a position shader in parallel with the main application and generally generates critical results faster than the full pipeline because the culling pipeline fetches and shades only the vertex position attributes without performing pixel rasterization and rendering to the frame buffer. The culling pipeline can use the generated critical results to calculate the visible information for all triangles regardless of whether the triangles are culled. The full pipeline (which may be referred to as the playback pipeline in this case) can consume the visible information to skip culled triangles that are shaded only with visible triangles that are ultimately passed to the rasterization stage.
[0032] In one embodiment, the further fixed function logic 238 can also include machine learning acceleration logic, such as fixed function matrix multiplication logic, for an implementation that includes optimizations for machine learning training or inference.
[0033] Within each of the graphics sub-cores 221A - 221F, there is a set of execution resources that can be used to execute graphics, media, and computational operations in response to requests by a graphics pipeline, media pipeline, or shader program. The graphics sub-cores 221A - 221F include a plurality of EU arrays 222A - 222F, 224A - 224F, thread dispatch and inter-thread communication (TD / IC) logic 223A - 223F, 3D (e.g., texture) samplers 225A - 225F, media samplers 226A - 226F, shader processors 227A - 227F, and shared local memory (SLM) 228A - 228F. The EU arrays 222A - 222F, 224A - 224F each include a plurality of execution units, which are general-purpose graphics processing units capable of performing floating-point and integer / fixed-point logical operations in the service of graphics, media, or computational operations including graphics, media, and compute shader programs. The TD / IC logic 223A - 223F performs local thread dispatch and thread control operations for the execution units within the sub-core and enables communication between the threads executing on the execution units of the sub-core. The 3D samplers 225A - 225F can read texture or other 3D graphics-related data into memory. The 3D samplers can read texture data differently based on the configured sample state and the texture format associated with a given texture. The media samplers 226A - 226F can perform similar read operations based on the type and format associated with media data. In one embodiment, each of the graphics sub-cores 221A - 221F can alternately include unified 3D and media samplers. Threads executing on the execution units within each of the sub-cores 221A - 221F can utilize the shared local memory 228A - 228F within each sub-core to enable threads executing within a thread group to execute using a common pool of on-chip memory.
[0034] FIG. 2C shows a graphics processing unit (GPU) 239 that includes a dedicated set of graphics processing resources disposed in multi-core groups 240A-240N. Although details of only a single multi-core group 240A are provided, it is recognized that the other multi-core groups 240B-240N may comprise the same or similar sets of graphics processing resources.
[0035] As shown, multi-core group 240A may include a set of graphics cores 243, a set of tensor cores 244, and a set of ray tracing cores 245. A scheduler / dispatcher 241 schedules and dispatches graphics threads for execution on the various cores 243, 244, 245. A set of register files 242 stores operand values used by the cores 243, 244, 245 when executing graphics threads. These may include, for example, integer registers for storing integer values, floating-point registers for storing floating-point values, vector registers for storing packed data elements (integer and / or floating-point data elements), and tile registers for storing tensor / matrix values. In one embodiment, the tile registers are implemented as a combined set of vector registers.
[0036] One or more combined level 1 (L1) caches and shared memory unit 247 locally store graphics data such as texture data, vertex data, pixel data, ray data, bounding volume data, etc. within each multicore group 240A. One or more texture units 247 can also be used to perform texture operations such as texture mapping and sampling. Level 2 (L2) cache 253 is shared by all or a subset of multicore groups 240A - 240N and stores graphics data and / or instructions for multiple simultaneous graphics threads. As shown, L2 cache 253 may be shared among multiple multicore groups 240A - 240N. One or more memory controllers 248 couple GPU 239 to a memory 249 which may be system memory (e.g., DRAM) and / or dedicated graphics memory (e.g., GDDR6 memory).
[0037] Input / output (I / O) circuitry 250 couples GPU 239 to one or more I / O devices such as a digital signal processor (DSP), network controller, or user input device at 252. On-chip interconnects may be used to couple I / O device 252 to GPU 239 and memory 249. One or more I / O memory management units (IOMMUs) 251 of I / O circuitry 250 directly couple I / O device 252 to system memory 249. In one embodiment, IOMMU 251 manages multiple sets of page tables to map virtual addresses to physical addresses within system memory 249. In this embodiment, I / O device 252, CPU 246, and GPU 239 may share the same virtual address space.
[0038] In one implementation, the IOMMU 251 supports virtualization. In this case, the IOMMU 251 may manage a first set of page tables for mapping guest / graphics virtual addresses to guest / graphics physical addresses (e.g., within the system memory 249), and a second set of page tables for mapping guest / graphics physical addresses to system / host physical addresses. The base address of each of the first and second sets of page tables may be stored in a control register and may be swapped during a context switch (e.g., thereby providing access to the set of page tables associated with the new context). Although not shown in FIG. 2C, each of the cores 243, 244, 245 and / or the multicore groups 240A-240N may include a translation lookaside buffer (TLB) for caching guest virtual-to-guest physical translations, guest physical-to-host physical translations, and guest virtual-to-host physical translations.
[0039] In one embodiment, the CPU 246, the GPU 239, and the I / O device 252 are integrated on a single semiconductor chip and / or chip package. The illustrated memory 249 may be integrated on the same chip or may be coupled to the memory controller 248 via an off-chip interface. In one implementation, the memory 249 includes GDDR6 memory that shares the same virtual address space as other physical system-level memories, but the underlying principles of the present invention are not limited to this particular implementation.
[0040] In one embodiment, the tensor core 244 includes a plurality of execution units specifically designed to perform matrix operations, which are the basic computational operations used to perform deep learning operations. For example, a simultaneous matrix multiplication operation may be used for neural network training and inference. The tensor core 244 may perform matrix processing using various operand precisions including single-precision floating point (e.g., 32 bits), half-precision floating point (e.g., 16 bits), integer words (16 bits), bytes (8 bits), and nibbles (4 bits). In one embodiment, the implementation of the neural network potentially combines details from multiple frames to extract features of each rendered scene in order to construct a high-quality final image.
[0041] In the implementation of deep learning, parallel matrix multiplication operations may be scheduled to be executed on the tensor core 244. Neural network training, in particular, requires a fairly large number of matrix dot product operations. To process the inner product formula of an N×N×N matrix multiplication, the tensor core 244 may include at least N dot product processing elements. Before the matrix multiplication begins, one entire matrix is loaded into the tile register, and at least one column of the second matrix is loaded every N cycles. There are N dot products to be processed every cycle.
[0042] Matrix elements may be stored at different precisions depending on the particular implementation, including 16-bit words, 8-bit bytes (e.g., INT8), and 4-bit nibbles (e.g., INT4). Different precision modes may be specified for the tensor core 244 to ensure that the most efficient precision is used for different workloads (e.g., inference workloads that can tolerate quantization to bytes and nibbles, etc.).
[0043] In one embodiment, the ray tracing core 245 accelerates ray tracing operations for both real-time and non-real-time ray tracing implementations. In particular, the ray tracing core 245 uses a bounding volume hierarchy (BVH) to perform ray traversal and includes a ray traversal / intersection circuit for identifying intersections between rays and primitives enclosed within the BVH volumes. The ray tracing core 245 may also include circuitry for performing depth testing and culling (e.g., using a Z-buffer or similar configuration). In one embodiment, the ray tracing core 245 performs traversal and intersection operations in cooperation with the image noise reduction techniques described herein, at least a portion of which may be executed on the tensor core 244. For example, in one embodiment, the tensor core 244 implements a deep learning neural network to perform noise reduction on frames generated by the ray tracing core 245. However, the CPU 246, the graphics core 243, and / or the ray tracing core 245 may also implement all or part of the noise reduction and / or deep learning algorithms.
[0044] Furthermore, as described above, a distributed approach for noise reduction may be used, where the GPU 239 is within a computing device coupled to other computing devices over a network or high-speed interconnect. In this embodiment, the interconnected computing devices share neural network learning / training data to improve the rate at which the overall system learns to perform noise reduction for different types of image frames and / or different graphics applications.
[0045] In one embodiment, the ray tracing core 245 processes all BVH traversals and ray-primitive intersections, preventing the graphics core 243 from being overloaded with thousands of instructions per ray. In one embodiment, each ray tracing core 245 includes a first set of special circuits for performing bounding box tests (e.g., for traversal operations) and a second set of special circuits for performing ray-triangle intersection tests (e.g., for traversed intersecting rays). Thus, in one embodiment, the multi-core group 240A can simply initiate a ray probe, and the ray tracing core 245 independently performs ray traversal and intersection and returns hit data (e.g., hit, no hit, multiple hits, etc.) to the thread context. The other cores 243, 244 are freed to perform other graphics or compute work while the ray tracing core 245 performs traversal and intersection operations.
[0046] In one embodiment, each ray tracing core 245 includes a traversal unit for performing BVH test operations and an intersection unit for performing ray-primitive intersection tests. The intersection unit generates a "hit", "no hit", or "multi-hit" response and provides it to the appropriate thread. During traversal and intersection operations, the execution resources of other cores (e.g., the graphics core 243 and the tensor core 244) are freed to perform other forms of graphics work.
[0047] In one particular embodiment described below, a hybrid rasterization / ray tracing approach is used in which work is distributed between the graphics core 243 and the ray tracing core 245.
[0048] In one embodiment, the ray tracing core 245 (and / or other cores 243, 244) includes a ray tracing instruction set such as Microsoft's DXR (DirectX Ray Tracing) that includes DispatchRays commands, as well as hardware support for ray generation shaders, closest-hit shaders, any-hit shaders, and miss shaders, enabling the assignment of a unique set of shaders and textures for each object. Other ray tracing platforms that may be supported by the ray tracing core 245, the graphics core 243, and the tensor core 244 are Vulkan 1.1.85. However, it should be noted that the underlying principles of the present invention are not limited to a specific ray tracing ISA.
[0049] Generally, the various cores 245, 244, 243 may support a ray tracing instruction set that includes instructions / functions for ray generation, closest-hit, any-hit, ray-primitive intersection, per-primitive and hierarchical bounding box construction, miss, visit, and exception. More specifically, one embodiment includes ray tracing instructions for performing the following functions.
[0050] Ray Generation - Ray generation instructions may be executed for each pixel, sample, or other user-defined work assignment.
[0051] Closest-Hit - Closest-hit instructions may be executed to find the closest intersection of a ray with a primitive in the scene.
[0052] Any-Hit - Any-hit instructions identify multiple intersections between a ray and primitives in the scene to potentially identify a new closest intersection.
[0053] Intersection - The Intersection command performs a ray-primitive intersection test and outputs the result.
[0054] Per-primitive Bounding box Construction - This command constructs a bounding box around a given primitive or group of primitives (e.g., when constructing a new BVH or other acceleration data structure).
[0055] Miss - Indicates that the ray misses all the geometry in the scene or a specified region of the scene.
[0056] Visit - Indicates the child volume through which the ray passes.
[0057] Exceptions - Includes various types of exception handlers (e.g., called on various error conditions).
[0058] FIG. 2D is a block diagram of a general purpose graphics processing unit (GPGPU) 270 that can be configured as a graphics processor and / or a computer accelerator according to the embodiments described herein. The GPGPU 270 can be interconnected with a host processor (e.g., one or more CPUs 246) and memories 271, 272 via one or more system and / or memory buses. In one embodiment, the memory 271 is a system memory that may be shared with one or more CPUs 246, and the memory 272 is a device memory dedicated to the GPGPU 270. In one embodiment, the components within the GPGPU 270 and the device memory 272 may be mapped to memory addresses accessible by one or more CPUs 246. Access to the memories 271 and 272 may be realized via a memory controller 268. In one embodiment, the memory controller 268 can include an internal direct memory access (DMA) controller 269 or, alternatively, can include logic for performing operations executed by the DMA controller.
[0059] The GPGPU 270 includes a plurality of cache memories including an L2 cache 253, an L1 cache 254, an instruction cache 255, and a shared memory 256, and at least some of these may be partitioned as cache memories. The GPGPU 270 also includes a plurality of computing units 260A - 260N. Each computing unit 260A - 260N includes a set of a vector register 261, a scalar register 262, a vector logic unit 263, and a scalar logic unit 264. The computing units 260A - 260N may also include a local shared memory 265 and a program counter 266. The computing units 260A - 260N can be coupled to a constant cache 267, and the constant cache 267 can be used to store constant data, which is data that does not change during the execution of a kernel or shader program executed on the GPGPU 270. In one embodiment, the constant cache 267 is a scalar data cache, and the cached data can be directly fetched into the scalar register 262.
[0060] During operation, one or more CPUs 246 can write commands to registers or memories within the GPGPU 270 that are mapped to an accessible address space. The command processor 257 can read commands from the registers or memories and determine how these commands are to be processed within the GPGPU 270. The thread dispatcher 258 can then dispatch threads to the computing units 260A - 260N for use in executing these commands. Each computing unit 260A - 260N can execute threads independently of other computing units. Further, each computing unit 260A - 260N can be independently configured for conditional computing and can conditionally output the results of the computation to memory. The command processor 257 can interrupt one or more CPUs 246 when the submitted commands are completed.
[0061] Figures 3A-3C show block diagrams of a further graphics processor and compute accelerator architecture provided by the embodiments described herein. Elements of Figures 3A-3C having the same reference numerals (or names) as elements in any of the other drawings herein may operate or function in a similar manner as those described elsewhere herein, but are not limited to such.
[0062] Figure 3A is a block diagram of a graphics processor 300, which may be an individual graphics processing unit, or a graphics processor integrated with a plurality of processing cores or other semiconductor devices such as, but not limited to, memory devices or network interfaces. In some embodiments, the graphics processor communicates with commands located in processor memory via a memory-mapped I / O interface to registers on the graphics processor. In some embodiments, the graphics processor 300 includes a memory interface 314 for accessing memory. The memory interface 314 can interface to local memory, one or more internal caches, one or more shared external caches, and / or system memory.
[0063] In some embodiments, the graphics processor 300 also includes a display controller 302 for driving display output data for the display device 318. The display controller 302 includes hardware for one or more overlay planes for the display and composition of multiple layers of video or user interface elements. The display device 318 can be an internal or external display device. In one embodiment, the display device 318 is a head-mounted display device such as a virtual reality (VR) display device or an augmented reality (AR) display device. In some examples, the graphics processor 300 includes a video codec engine 306 for encoding media into, decoding media from, or transcoding between one or more media encoding formats including, but not limited to, MPEG (Moving Picture Experts Group) formats such as MPEG-2, AVC (Advanced Video Coding) formats such as H.264 / MPEG-4 AVC, H.265 / HEVC, AOMedia (Alliance for Open Media) VP8, VP9, and JPEG (Joint Photographic Experts Group) formats such as SMPTE (Society of Motion Picture & Television Engineers) 421M / VC-1 and JPEG and Motion JPEG (MJPEG).
[0064] In some embodiments, the graphics processor 300 includes, for example, a block image transfer (BLIT) engine 304 for performing two-dimensional (2D) rasterization operations including, for example, bit-boundary block transfer. However, in one embodiment, 2D graphics operations are performed using one or more components of the graphics processing engine (GPE) 310. In some embodiments, the GPE 310 is a computational engine for performing graphics operations including three-dimensional (3D) graphics operations and media operations.
[0065] In some embodiments, the GPE 310 includes a 3D pipeline 312 for performing 3D operations such as rendering three-dimensional images and scenes using processing functions that act on 3D primitive shapes (e.g., rectangles, triangles, etc.). The 3D pipeline 312 includes programmable and fixed function elements that perform various tasks within the element and / or generate execution threads for the 3D / media subsystem 315. While the 3D pipeline 312 can be used to perform media operations, embodiments of the GPE 310 also include a media pipeline 316 that is specifically used to perform media operations such as video post-processing and image enhancement.
[0066] In some embodiments, media pipeline 316 includes fixed function or programmable logic units for performing one or more special media operations, such as video decode acceleration, video de-interlacing, and video encode acceleration, instead of or in addition to video codec engine 306. In some embodiments, media pipeline 316 further includes a thread generation unit that generates threads for execution on 3D / media subsystem 315. The generated threads execute calculations for media operations on one or more graphics execution units included in 3D / media subsystem 315.
[0067] In some embodiments, 3D / media subsystem 315 includes logic for executing threads generated by 3D pipeline 312 and media pipeline 316. In one embodiment, the pipeline sends thread execution requests to 3D / media subsystem 315, and 3D / media subsystem 315 includes thread dispatch logic for arbitrating various requests and dispatching them to available thread execution resources. The execution resources include an array of graphics execution units for processing 3D and media threads. In some embodiments, 3D / media subsystem 315 includes one or more internal caches for thread instructions and data. In some embodiments, the subsystem also includes shared memory that includes registers and addressable memory for sharing data between threads and storing output data.
[0068] Figure 3B shows a graphics processor 320 having a tiled architecture according to the embodiments described herein. In one embodiment, the graphics processor 320 includes a graphics processing engine cluster 322 having a plurality of instances of the graphics processing engine 310 of FIG. 3A within graphics engine tiles 310A-310D. Each of the graphics engine tiles 310A-310D can be interconnected via a set of tile interconnects 323A-323F. Each of the graphics engine tiles 310A-310D can also be connected to memory modules or memory devices 326A-326D via memory interconnects 325A-325D. The memory devices 326A-326D can use any graphics memory technology. For example, the memory devices 326A-326D may be graphics double data rate (GDDR) memory. In one embodiment, the memory devices 326A-326D are high-bandwidth memory (HBM) modules that can be on-die with their respective graphics engine tiles 310A-310D. In one embodiment, the memory devices 326A-326D are stacked memory devices that can be stacked on top of their respective graphics engine tiles 310A-310D. In one embodiment, each of the graphics engine tiles 310A-310D and associated memories 326A-326D are present on separate chiplets coupled to a base die or base substrate as described in more detail in FIGS. 11B-11D.
[0069] The graphics processing engine cluster 322 can be connected to an on-chip or on-package fabric interconnect 324. The fabric interconnect 324 can enable communication between the graphics engine tiles 310A - 310D and components such as the video codec 306 and one or more copy engines 304. The copy engine 304 can be used to move data from, to, or between the memory devices 326A - 326D and memory external to the graphics processor 320 (e.g., system memory). The fabric interconnect 324 can also be used to interconnect the graphics engine tiles 310A - 310D. The graphics processor 320 may optionally include a display controller 302 to enable connection to an external display device 318. The graphics processor may also be configured as a graphics or compute accelerator. In the accelerator configuration, the display controller 302 and the display device 318 may be omitted.
[0070] The graphics processor 320 can be connected to a host system via a host interface 328. The host interface 328 can enable communication between the graphics processor 320, the system memory, and / or other system components. The host interface 328 can be, for example, a PCI Express bus or other type of host system interface.
[0071] FIG. 3C shows a compute accelerator 330 according to the embodiments described herein. The compute accelerator 330 can include architectural similarities with the graphics processor 320 of FIG. 3B and is optimized for compute acceleration. The compute engine cluster 332 can include a set 340A - 340D of compute engine tiles that include execution logic optimized for parallel or vector-based general-purpose compute operations. In some embodiments, the compute engine tiles 340A - 340D do not include fixed-function graphics processing logic, although in one embodiment, one or more of the compute engine tiles 340A - 340D can include logic for performing media acceleration. The compute engine tiles 340A - 340D can be connected to memories 326A - 326D via memory interconnects 325A - 325D. The memories 326A - 326D and the memory interconnects 325A - 325D can be of the same technology as or different from the graphics processor 320 of FIG. 3B. The graphics compute engine tiles 340A - 340D can also be interconnected via a set 323A - 323F of tile interconnects, can be connected to the fabric interconnect 324, and / or can be interconnected by the fabric interconnect 324. In one embodiment, the compute accelerator 330 includes a large L3 cache 336 that can be configured as a device-wide cache. The compute accelerator 330 can also be connected to the host processor and memory via a host interface 328, similar to the graphics processor 320 of FIG. 3B.
[0072] (Graphics Processing Engine) FIG. 4 is a block diagram of a graphics processing engine 410 of a graphics processor according to some embodiments. In one embodiment, the graphics processing engine (GPE) 410 is a version of the GPE 310 shown in FIG. 3A and may also represent the graphics engine tiles 310A-310D of FIG. 3B. Elements of FIG. 4 having the same reference numerals (or names) as elements in any of the other drawings herein may operate or function in a similar manner as those described elsewhere herein, but are not limited thereto. For example, the 3D pipeline 312 and the media pipeline 316 of FIG. 3A are shown. The media pipeline 316 is optional in some embodiments of the GPE 410 and may not be explicitly included within the GPE 410. For example, in at least one embodiment, a separate media and / or image processor is coupled to the GPE 410.
[0073] In some embodiments, GPE410 is coupled to or includes a command streamer 403 that provides a command stream to the 3D pipeline 312 and / or the media pipeline 316. In some embodiments, the command streamer 403 is coupled to a memory, which can be a system memory, or one or more of an internal cache memory and a shared cache memory. In some embodiments, the command streamer 403 receives commands from the memory and transmits the commands to the 3D pipeline 312 and / or the media pipeline 316. The commands are instructions fetched from a ring buffer that stores commands for the 3D pipeline 312 and the media pipeline 316. In one embodiment, the ring buffer can further include a batch command buffer that stores batches of multiple commands. The commands for the 3D pipeline 312 can also include references to data stored in memory, such as, but not limited to, vertex and geometry data for the 3D pipeline 312 and / or image data and memory objects for the media pipeline 316. The 3D pipeline 312 and the media pipeline 316 process commands and data by executing operations via logic within each pipeline or by dispatching one or more execution threads to the graphics core array 414. In one embodiment, the graphics core array 414 includes one or more blocks of graphics cores (e.g., graphics core 415A, graphics core 415B), and each block includes one or more graphics cores. Each graphics core includes a set of graphics execution resources that includes general-purpose and graphics-specific execution logic for executing graphics and computational operations, and fixed-function texture processing and / or machine learning and artificial intelligence acceleration logic.
[0074] In various embodiments, the 3D pipeline 312 can include fixed function and programmable logic for processing one or more shader programs, such as vertex shaders, geometry shaders, pixel shaders, fragment shaders, compute shaders, or other shader programs, by processing instructions and dispatching execution threads to the graphics core array 414. The graphics core array 414 provides a unified block of execution resources for use in processing these shader programs. The general-purpose execution logic (e.g., execution units) within the graphics cores 415A - 414B of the graphics core array 414 includes support for various 3D API shader languages and can execute multiple concurrent execution threads associated with multiple shaders.
[0075] In some embodiments, the graphics core array 414 includes execution logic for performing media functions such as video and / or image processing. In one embodiment, the execution units include general-purpose logic programmable to execute parallel general-purpose computing operations in addition to graphics processing operations. The general-purpose logic can execute processing operations in parallel with or in relation to the general-purpose logic within the processor core 107 of FIG. 1 or cores 202A - 202N such as in FIG. 2A.
[0076] Output data generated by threads executing on the graphics core array 414 can output data to memory within a unified return buffer (URB) 418. The URB 418 can store data from multiple threads. In some embodiments, the URB 418 may be used to transmit data between different threads executing on the graphics core array 414. In some embodiments, the URB 418 may be further used for synchronization between threads on the graphics core array and fixed function logic within the shared function logic 420.
[0077] In some embodiments, the graphics core array 414 is scalable such that the array includes a variable number of graphics cores, each having a variable number of execution units based on the target power and performance levels of the GPE 410. In one embodiment, the execution resources are dynamically scalable such that the execution resources may be enabled or disabled as needed.
[0078] The graphics core array 414 is coupled to shared function logic 420 that includes a plurality of resources shared among the graphics cores within the graphics core array. The shared functions within the shared function logic 420 are hardware logic units that provide special supplemental functions to the graphics core array 414. In various embodiments, the shared function logic 420 includes, but is not limited to, sampler 421, numerical operations 422, and inter-thread communication (ITC) 423 logic. Additionally, some embodiments implement one or more caches 425 within the shared function logic 420.
[0079] The shared functions are implemented at least when the demand for a given special function is insufficiently included within the graphics core array 414. Instead, a single instantiation of that special function is implemented as a stand-alone entity within the shared function logic 420 and shared among the execution resources within the graphics core array 414. The exact set of functions that are shared among the graphics core arrays 414 and included within the graphics core arrays 414 varies by embodiment. In some embodiments, certain shared functions within the shared function logic 420 that are more widely used by the graphics core array 414 may be included within the shared function logic 416 within the graphics core array 414. In various embodiments, the shared function logic 416 within the graphics core array 414 can include some or all of the logic within the shared function logic 420. In one embodiment, all of the logical elements within the shared function logic 420 may be replicated within the shared function logic 416 of the graphics core array 414. In one embodiment, the shared function logic 420 is excluded for the shared function logic 416 within the graphics core array 414.
[0080] (Execution Unit) FIGS. 5A-5B illustrate a thread execution logic 500 including an array of processing elements used in a graphics processor core according to the embodiments described herein. Elements of FIGS. 5A-5B having the same reference numerals (or names) as elements in any other drawing here may operate or function in a similar manner as those described elsewhere herein, but are not limited to such. FIGS. 5A-5B illustrate an overview of a thread execution logic 600, which may represent the hardware logic shown in each sub-core 221A-221F of FIG. 2B. FIG. 5A represents an execution unit within a general-purpose graphics processor, and FIG. 5B represents an execution unit that may be used within a compute accelerator.
[0081] As shown in FIG. 5A, in some embodiments, the thread execution logic 500 includes a shader processor 502, a thread dispatcher 504, an instruction cache 506, a scalable execution unit array including a plurality of execution units 508A-508N, a sampler 510, a shared local memory 511, a data cache 512, and a data port 514. In one embodiment, the scalable execution unit array can be dynamically scaled by enabling or disabling one or more execution units (e.g., any of execution units 508A, 508B, 508C, 508D-508N-1, and 508N) based on the computational requirements of the workload. In one embodiment, the components included are interconnected via an interconnect fabric that couples to each of the components. In some embodiments, the thread execution logic 500 includes one or more connections to memory such as system memory or cache memory through one or more of the instruction cache 506, the data port 514, the sampler 510, and the execution units 508A-508N. In some embodiments, each execution unit (e.g., 508A) is a stand-alone programmable general-purpose computing unit that can execute a plurality of simultaneous hardware threads while processing a plurality of data elements in parallel for each thread. In various embodiments, the array of execution units 508A-508N is scalable to include any number of individual execution units.
[0082] In some embodiments, execution units 508A - 508N are mainly used to execute shader programs. Shader processor 502 can process various shader programs and dispatch execution threads related to the shader programs via thread dispatcher 504. In one embodiment, the thread dispatcher arbitrates thread start requests from the graphics and media pipeline and includes logic for instantiating the requested threads on one or more of the execution units 508A - 508N. For example, the geometry pipeline can dispatch vertex, tessellation, or geometry shaders for processing to the thread execution logic. In some embodiments, thread dispatcher 504 can also process runtime thread generation requests from the executing shader program.
[0083] In some embodiments, execution units 508A - 508N support an instruction set that includes native support for many standard 3D graphics shader instructions so that shader programs from a graphics library (e.g., Direct3D and OpenGL) can be executed with minimal translation. The execution units support vertex and geometry processing (e.g., vertex programs, geometry programs, vertex shaders), pixel processing (e.g., pixel shaders, fragment shaders), and general - purpose processing (e.g., compute and media shaders). Each of execution units 508A - 508N is capable of multi - issue single instruction multiple data (SIMD) execution, and multithreaded operations enable an efficient execution environment when facing higher - latency memory accesses. Each hardware thread within each execution unit has a dedicated wide - bandwidth register file and associated independent thread state. Execution is multi - issue per clock to a pipeline capable of integer, single - precision, and double - precision floating - point operations, SIMD branch capabilities, logical operations, transcendental operations, and other diverse operations. While waiting for data from one of memory or shared functions, the dependent logic within execution units 508A - 508N puts the waiting thread to sleep until the requested data is returned. While the waiting thread is sleeping, hardware resources may be allocated to process other threads. For example, during the latency associated with vertex shader operations, the execution unit can execute operations of a pixel shader, fragment shader, or other type of shader program (including a different vertex shader). Various embodiments can be applied to use execution by single instruction multiple thread (SIMT) as an alternative to or in addition to the use of SIMD. References to SIMD cores or operations can also apply to SIMT or can apply to SIMD in combination with SIMT.
[0084] Each execution unit in execution units 508A-508N operates on an array of data elements. The number of data elements is the "execution size" or number of channels of an instruction. An execution channel is a logical unit of execution for data element access, masking, and flow control within an instruction. The number of channels may be independent of the number of physical Arithmetic Logic Units (ALUs) or Floating Point Units (FPUs) for a particular graphics processor. In some embodiments, execution units 508A-508N support integer and floating point data types.
[0085] The execution unit instruction set includes SIMD instructions. Various data elements can be stored as packed data types in registers, and the execution unit processes the various elements based on the data size of the elements. For example, when operating on a 256-bit wide vector, the 256 bits of the vector are stored in a register, and the execution unit operates on the vector as four separate 54-bit packed data elements (QW (Quad-Word) sized data elements), eight separate 32-bit packed data elements (DW (Double Word) sized data elements), sixteen separate 16-bit packed data elements (W (Word) sized data elements), or thirty-two separate 8-bit data elements (B (byte) sized data elements). However, different vector widths and register sizes are possible.
[0086] In one embodiment, one or more execution units can be coupled to fused execution units 509A - 509N having thread control logic (507A - 507N) common to the fused EUs. Multiple EUs can be fused into EU groups. Each EU within a fused EU group can be configured to execute separate SIMD hardware threads. The number of EUs within a fused EU group can vary according to the embodiment. Further, various SIMD widths, including but not limited to SIMD8, SIMD16, and SIMD32, can be executed per EU. Each fused graphics execution unit 509A - 509N includes at least two execution units. For example, fused execution unit 509A includes a first EU 508A, a second EU 508B, and thread control logic 507A common to the first EU 508A and the second EU 508B. Thread control logic 507A controls the threads executed on fused graphics execution unit 509A and enables each EU within fused execution units 509A - 509N to execute using a common instruction pointer register.
[0087] One or more internal instruction caches (e.g., 506) are included in thread execution logic 500 to cache thread instructions for the execution units. In some embodiments, one or more data caches (e.g., 512) are included to cache thread data during thread execution. Threads executing on execution logic 500 can also store explicitly managed data in shared local memory 511. In some embodiments, sampler 510 is included to provide texture sampling for 3D operations and media sampling for media operations. In some embodiments, sampler 510 includes special texture or media sampling functions for processing texture or media data during the sampling process before providing the sampled data to the execution units.
[0088] During execution, the graphics and media pipeline sends thread start requests to the thread execution logic 500 via thread generation and dispatch logic. When a group of geometric objects is processed and rasterized into pixel data, the pixel processor logic (e.g., pixel shader logic, fragment shader logic, etc.) within the shader processor 502 is called to further compute output information and cause the results to be written to an output surface (e.g., a color buffer, a depth buffer, a stencil buffer, etc.). In some embodiments, the pixel shader or fragment shader computes values for various vertex attributes to be interpolated between the rasterized objects. In some embodiments, the pixel processor logic within the shader processor 502 then executes a pixel or fragment shader program supplied by an application programming interface (API). To execute the shader program, the shader processor 502 dispatches threads to execution units (e.g., 508A) via a thread dispatcher 504. In some embodiments, the shader processor 502 uses texture sampling logic within a sampler 510 to access texture data within a texture map stored in memory. Arithmetic operations on the texture data and input geometry data either compute pixel color data for each geometric fragment or discard one or more pixels from further processing.
[0089] In some embodiments, the data port 514 provides a memory access mechanism for the thread execution logic 500 to output processed data to memory for further processing on the graphics processor output pipeline. In some embodiments, the data port 514 includes or is coupled to one or more cache memories (e.g., a data cache 512) to cache data for memory access via the data port.
[0090] In one embodiment, the execution logic 500 can also include a ray tracer 505 that can provide a ray tracing acceleration function. The ray tracer 505 can support a ray tracing instruction set that includes instructions / functions for ray generation. The ray tracing instruction set can be the same as the ray tracing instruction set supported by the ray tracing core 245 in FIG. 2C, or it can be different.
[0091] FIG. 5B shows exemplary internal details of the execution unit 508 according to an embodiment. The graphics execution unit 508 can include an instruction fetch unit 537, a general register file array (GRF) 524, an architectural register file array (ARF) 526, a thread arbiter 522, a dispatch unit 530, a branch unit 532, a set of SIMD floating-point units 534, and in one embodiment, a set of dedicated integer SIMD ALUs 535. The GRF 524 and the ARF 526 include a set of general register files and architectural register files associated with each simultaneous hardware thread that can be active in the graphics execution unit 508. In one embodiment, the architectural state per thread is maintained in the ARF 526, while the data used during thread execution is stored in the GRF 524. The execution state of each thread, including the instruction pointer for each thread, can be held in thread-specific registers within the ARF 526.
[0092] In one embodiment, the graphics execution unit 508 has an architecture that is a combination of Simultaneous Multi-Threading (SMT) and Interleaved Multi-Threading (IMT). The architecture has a modular configuration that can be fine-tuned at design time based on the target number of simultaneous threads and the number of registers per execution unit, and the execution unit resources are divided among the logic used to execute multiple simultaneous threads. The number of logical threads that can be executed by the graphics execution unit 508 is not limited to the number of hardware threads, and multiple logical threads can be assigned to each hardware thread.
[0093] In one embodiment, the graphics execution unit 508 can simultaneously issue a plurality of instructions, which may be different instructions. The thread arbiter 522 of the graphics execution unit thread 508 can dispatch an instruction to one of the transmission unit 530, the branch unit 532, or the SIMD FPU 534 for execution. Each execution thread can access 128 general-purpose registers in the GRF 524. Each register can store 32 bytes and can be accessed as an 8-element vector of SIMD of 32-bit data elements. In one embodiment, each execution unit thread has access to 4K bytes in the GRF 524, but the embodiments are not limited thereto, and in other embodiments, larger or smaller register resources may be provided. In one embodiment, the graphics processing unit 508 is divided into seven hardware threads that can independently execute computational operations, but the number of threads per execution unit can also vary according to the embodiment. For example, in one embodiment, up to 16 hardware threads are supported. In an embodiment where seven threads can access 4K bytes, the GRF 524 can store a total of 28K bytes. When 16 threads can access 4K bytes, the GRF 524 can store a total of 64K bytes. The flexible addressing mode can enable registers to be addressed together to effectively construct wider registers or represent a strided rectangular block data structure.
[0094] In one embodiment, memory operations, sampler operations, and other longer-latency system communications are dispatched via "send" instructions executed by messages passing through the transmission unit 530. In one embodiment, branch instructions are dispatched to a dedicated branch unit 532 to achieve SIMD divergence and final convergence.
[0095] In one embodiment, the graphics execution unit 508 includes one or more SIMD floating point units (FPUs) 534 for performing floating point operations. In one embodiment, the FPU 534 also supports integer calculations. In one embodiment, the FPU 534 can perform up to M 32-bit floating point (or integer) operations in SIMD, or up to 2M 16-bit integer or 16-bit floating point operations in SIMD. In one embodiment, at least one of the FPUs provides extended numerical operation capabilities that support high throughput transcendental numerical operation functions and double precision 54-bit floating point. In some embodiments, there is also a set of 8-bit integer SIMD ALUs 535, which may be particularly optimized for performing operations related to machine learning calculations.
[0096] In one embodiment, an array of multiple instances of the graphics execution unit 508 can be instantiated with graphics sub-core grouping (e.g., sub-slicing). For scalability, the product architect can select the exact number of execution units per sub-core group. In one embodiment, the execution unit 508 can execute instructions among multiple execution channels. In a further embodiment, each thread executed on the graphics execution unit 508 is executed on a different channel.
[0097] FIG. 6 shows a further execution unit 600 according to an embodiment. The execution unit 600 may be, for example, a computation-optimized execution unit for use in computation engine tiles 340A-340D such as those in FIG. 3C, but is not limited to such. Also, in graphics engine tiles 310A-310D such as those in FIG. 3B, a variation of the execution unit 600 may be used. In one embodiment, the execution unit 600 includes a thread control unit 601, a thread state unit 602, an instruction fetch / prefetch unit 603, and an instruction decode unit 604. The execution unit 600 further includes a register file 606 that stores registers that can be assigned to hardware threads within the execution unit. The execution unit 600 further includes a send unit 607 and a branch unit 608. In one embodiment, the send unit 607 and the branch unit 608 can operate in a manner similar to the send unit 530 and the branch unit 532 of the graphics execution unit 508 in FIG. 5B.
[0098] The execution unit 600 also includes a computing unit 610 that includes a plurality of different types of functional units. In one embodiment, the computing unit 610 includes an ALU unit 611 that includes an array of arithmetic logic units. The ALU unit 611 can be configured to perform 64-bit, 32-bit, and 16-bit integer and floating-point operations. The integer and floating-point operations may be performed simultaneously. The computing unit 610 can also include a systolic array 612 and a numerical operation unit 613. The systolic array 612 includes a network of data processing units of width W and depth D that can be used to perform systolic vector or other data parallel operations. In one embodiment, the systolic array 612 can be configured to perform matrix operations such as matrix dot product operations. In one embodiment, the systolic array 612 supports 16-bit floating-point operations and 8-bit and 4-bit integer operations. In one embodiment, the systolic array 612 can be configured to accelerate machine learning operations. In such an embodiment, the systolic array 612 can be configured to support the bfloat 16-bit floating-point format. In one embodiment, the numerical operation unit 613 can be included to perform a specific subset of numerical operations in a more efficient and lower power manner than the ALU unit 611. The numerical operation unit 613 can include a variation of the numerical operation logic that can be found within the shared function logic of a graphics processing engine provided by other embodiments (e.g., the numerical operation logic 422 of the shared function logic 420 of FIG. 4). In one embodiment, the numerical operation unit 613 can be configured to perform 32-bit and 64-bit floating-point operations.
[0099] The thread control unit 601 includes logic for controlling the execution of threads within the execution unit. The thread control unit 601 can include thread arbitration logic for starting, stopping, and pre-empting the execution of threads within the execution unit 600. The thread state unit 602 can be used to store the thread state for threads assigned to execute on the execution unit 600. Storing the thread state within the execution unit 600 enables rapid pre-emption of threads when these threads become blocked or idle. The instruction fetch / prefetch unit 603 can fetch instructions from the instruction cache of a higher-level execution logic (e.g., instruction cache 506 as shown in FIG. 5A). The instruction fetch / prefetch unit 603 can also issue prefetch requests for instructions to be loaded into the instruction cache based on the analysis of the currently executing threads. The instruction decoding unit 604 can be used to decode instructions to be executed by the computing unit. In one embodiment, the instruction decoding unit 604 can be used as a secondary decoder for decoding complex instructions into component micro-operations.
[0100] The execution unit 600 further includes a register file 606 that can be used by the hardware threads executing on the execution unit 600. The registers within the register file 606 can be partitioned among the logic used to execute multiple concurrent threads within the computing unit 610 of the execution unit 600. The number of logical threads that can be executed by the graphics execution unit 600 is not limited to the number of hardware threads, and multiple logical threads can be assigned to each hardware thread. The size of the register file 606 can vary among embodiments based on the number of supported hardware threads. In one embodiment, register renaming may be used to dynamically assign registers to hardware threads.
[0101] FIG. 7 is a block diagram showing a graphics processor instruction format 700 according to some embodiments. In one or more embodiments, a graphics processor execution unit supports an instruction set having instructions of multiple formats. The solid box indicates components generally included in the execution unit instructions, and the dashed lines include components that are optional or included only in a subset of the instructions. In some embodiments, the instruction format 700 described and illustrated is a macro instruction in that it is an instruction supplied to the execution unit as opposed to the micro-operations resulting from instruction decoding when the instruction is processed.
[0102] In some embodiments, a graphics processor execution unit natively supports instructions of a 128-bit instruction format 710. A 64-bit compact instruction format 730 is available for some instructions based on the selected instructions, instruction options, and number of operands. The native 128-bit instruction format 710 provides access to all instruction options, while some options and operations are restricted in the 64-bit format 730. The native instructions available in the 64-bit format 730 vary by embodiment. In some embodiments, instructions are compacted using a subset of the index values in the index field 713. The execution unit hardware references a set of compacting tables based on the index values and uses the output of the compacting tables to reconstruct the native instructions of the 128-bit instruction format 710. Other sizes and formats of instructions may also be used.
[0103] For each format, the instruction opcode 712 defines the operation to be performed by the execution unit. The execution unit executes each instruction in parallel among multiple data elements of each operand. For example, in response to an add instruction, the execution unit performs a simultaneous addition operation among each color channel representing a texture element or a picture element. By default, the execution unit executes each instruction among all data channels of the operand. In some embodiments, the instruction control field 714 enables control of specific execution options such as channel selection (e.g., prediction) and data channel order (e.g., swizzle). For instructions in the 128-bit instruction format 710, the exec size field 716 limits the number of data channels to be executed in parallel. In some embodiments, the exec size field 716 is not available for use in the 64-bit compact instruction format 730.
[0104] Some execution unit instructions have up to three operands, namely two source operands, i.e., src0 720, src1 722, and one destination 718. In some embodiments, the execution unit supports dual destination instructions, with one of the destinations being implied. Data manipulation instructions can have a third source operand (e.g., SRC2 724), and the instruction opcode 712 determines the number of source operands. The last source operand of the instruction can be a value of an immediate value (e.g., hard-coded) passed with the instruction.
[0105] In some embodiments, the 128-bit instruction format 710 includes an access / address mode field 726 that specifies, for example, whether a direct register addressing mode or an indirect register addressing mode is used. When the direct register addressing mode is used, the register addresses of one or more operands are provided directly by bits within the instruction.
[0106] In some embodiments, the 128-bit instruction format 710 includes an access / address mode field 726 that specifies the instruction's address mode and / or access mode. In one embodiment, the access mode is used to define the data access alignment for the instruction. Some embodiments support access modes that include a 16-byte aligned access mode and a 1-byte aligned access mode, and the byte alignment of the access mode determines the access alignment of the instruction operands. For example, in a first mode, the instruction may use byte-aligned addressing for source and destination operands, and in a second mode, the instruction may use 16-byte-aligned addressing for all source and destination operands.
[0107] In one embodiment, the address mode portion of the access / address mode field 726 determines whether the instruction uses direct addressing or indirect addressing. When the direct register addressing mode is used, bits within the instruction directly provide the register addresses of one or more operands. When the indirect register addressing mode is used, the register addresses of one or more operands may be calculated based on the address register value and the address immediate field within the instruction.
[0108] In some embodiments, the instructions are grouped based on bit fields of the opcode 712 to simplify opcode decoding 740. For 8-bit opcodes, bits 4, 5, and 6 enable the execution unit to determine the type of the opcode. The exact opcode grouping shown is merely an example. In some embodiments, the move and logic opcode group 742 includes data movement and logic instructions (e.g., move (mov), compare (cmp)). In some embodiments, the move and logic group 742 shares the five most significant bits (MSB), the move (mov) instruction is in the form of 0000xxxxb, and the logic instruction is in the form of 0001xxxb. The flow control instruction group 744 (e.g., call, jump (jmp)) includes instructions in the form of 0010xxxxb (e.g., 0x20). The various instruction group 746 includes a mix of instructions including synchronization instructions (e.g., wait, send) in the form of 0011xxxxb (e.g., 0x30). The parallel numerical arithmetic instruction group 748 includes per-component arithmetic instructions (e.g., add, multiply (mul)) in the form of 0100xxxxb (e.g., 0x40). The parallel numerical arithmetic group 748 executes arithmetic operations in parallel between data channels. The vector numerical arithmetic group 750 includes arithmetic instructions (e.g., dp4) in the form of 0101xxxxb (e.g., 0x50). The vector numerical arithmetic group executes arithmetic such as dot product calculations on vector operands. In one embodiment, the illustrated opcode decoding 740 can be used to determine which part of the execution unit is used to execute the decoded instructions. For example, some instructions may be designated as systolic instructions to be executed by a systolic array. Other instructions, such as ray tracing instructions (not shown), can be routed to a ray tracing core or ray tracing logic within a slice or partition of the execution logic.
[0109] (Graphics Pipeline) FIG. 8 is a block diagram of another embodiment of the graphics processor 800. Elements of FIG. 8 that have the same reference numerals (or names) as elements in any of the other drawings can operate or function in a similar manner as those described elsewhere herein, but are not limited thereto.
[0110] In some embodiments, the graphics processor 800 includes a geometry pipeline 820, a media pipeline 830, a display engine 840, a thread execution logic 850, and a rendering output pipeline 870. In some embodiments, the graphics processor 800 is a graphics processor within a multi-core processing system that includes one or more general-purpose processing cores. The graphics processor is controlled by register writes to one or more control registers (not shown) or via commands issued to the graphics processor 800 via the ring interconnect 802. In some embodiments, the ring interconnect 802 couples the graphics processor 800 to other processing components such as other graphics processors or general-purpose processors. Commands from the ring interconnect 802 are interpreted by a command streamer 803 that supplies instructions to individual components of the geometry pipeline 820 or the media pipeline 830.
[0111] In some embodiments, the command streamer 803 instructs the operation of a vertex fetcher 805 that reads vertex data from memory and executes vertex processing commands provided by the command streamer 803. In some embodiments, the vertex fetcher 805 provides the vertex data to a vertex shader 807, and the vertex shader 807 performs coordinate space transformation and lighting calculations for each vertex. In some embodiments, the vertex fetcher 805 and the vertex shader 807 execute vertex processing instructions by dispatching execution threads to execution units 852A - 852B via a thread dispatcher 831.
[0112] In some embodiments, execution units 852A - 852B are an array of vector processors having an instruction set for performing graphics and media operations. In some embodiments, execution units 852A - 852B have attached L1 caches 851 that are specific to each array or shared among the arrays. The cache can be configured as a data cache, an instruction cache, or a single cache partitioned to contain data and instructions.
[0113] In some embodiments, geometry pipeline 820 includes a tessellation component for performing hardware - accelerated tessellation of 3D objects. In some embodiments, a programmable hull shader 811 configures the tessellation operation. A programmable domain shader 817 provides back - end evaluation of the tessellation output. The tessellator 813 operates in the direction of the hull shader 811 and includes special - purpose logic for generating a set of detailed geometric objects based on a coarse geometric model provided as an input to geometry pipeline 820. In some embodiments, if tessellation is not used, the tessellation components (e.g., hull shader 811, tessellator 813, and domain shader 817) can be bypassed.
[0114] In some embodiments, a complete geometric object can be processed by the geometry shader 819 via one or more threads dispatched to execution units 852A - 852B, or can proceed directly to the clipper 829. In some embodiments, the geometry shader operates on entire geometric objects rather than vertices or vertex patches as in previous stages of the graphics pipeline. When tessellation is disabled, the geometry shader 819 receives input from the vertex shader 807. In some embodiments, the geometry shader 819 is programmable by a geometry shader program to perform geometry tessellation when the tessellation unit is disabled.
[0115] Before rasterization, the clipper 829 processes vertex data. The clipper 829 may be a fixed - function clipper or a programmable clipper having clipping and geometry shader functionality. In some embodiments, the rasterizer and depth test component 873 within the rendering output pipeline 870 dispatches pixel shaders to convert geometric objects into per - pixel representations. In some embodiments, the pixel shader logic is included in the thread execution logic 850. In some embodiments, an application can bypass the rasterizer and depth test component 873 and access the non - rasterized vertex data via the stream output unit 823.
[0116] The graphics processor 800 has an interconnect bus, an interconnect fabric, or some other interconnect mechanism that passes data and messages among the major components of the processor. In some embodiments, execution units 852A - 852B and associated logic units (e.g., L1 cache 851, sampler 854, texture cache 858, etc.) are interconnected via data port 856 to perform memory accesses and communicate with the rendering output pipeline components of the processor. In some embodiments, sampler 854, caches 851, 858, and execution units 852A - 852B each have separate memory access paths. In one embodiment, texture cache 858 can also be configured as a sampler cache.
[0117] In some embodiments, the rendering output pipeline 870 includes a rasterizer and depth test component 873 that converts vertex - based objects to their associated pixel - based representations. In some embodiments, the rasterizer logic includes a window / mask unit for performing fixed - function triangle and line rasterization. Associated rendering cache 878 and depth cache 879 are also available in some embodiments. Pixel operation component 877 performs pixel - based operations on the data, although in some instances, pixel operations related to 2D operations (e.g., bit - block image transfer using blend) are performed by 2D engine 841 or replaced by display controller 843 using an overlay display surface at display time. In some embodiments, shared L3 cache 875 is available to all graphics components, enabling sharing of data without using the main system memory.
[0118] In some embodiments, the graphics processor media pipeline 830 includes a media engine 837 and a video front end 834. In some embodiments, the video front end 834 receives pipeline commands from the command streamer 803. In some embodiments, the media pipeline 830 includes a separate command streamer. In some embodiments, the video front end 834 processes media commands before sending the commands to the media engine 837. In some embodiments, the media engine 837 includes a thread generation function for generating threads that are dispatched to the thread execution logic 850 via the thread dispatcher 831.
[0119] In some embodiments, the graphics processor 800 includes a display engine 840. In some embodiments, the display engine 840 is external to the processor 800 and is coupled to the graphics processor via the ring interconnect 802 or other interconnect bus or fabric. In some embodiments, the display engine 840 includes a 2D engine 841 and a display controller 843. In some embodiments, the display engine 840 includes special purpose logic that is operable independent of the 3D pipeline. In some embodiments, the display controller 843 is coupled to a display device (not shown), which may be a system integrated display device such as in a laptop computer, or an external display device attached via a display device connector.
[0120] In some embodiments, geometry pipeline 820 and media pipeline 830 are configurable to execute operations based on multiple graphics and media programming interfaces and are not specific to any one application programming interface (API). In some embodiments, driver software for a graphics processor converts API calls specific to a particular graphics or media library into commands that can be processed by the graphics processor. In some embodiments, support is provided for OpenGL (Open Graphics Library), OpenCL (Open Computing Language), and / or Vulkan graphics and compute APIs from the Khronos Group. In some embodiments, support for the Direct3D library from Microsoft Corporation may also be provided. In some embodiments, combinations of these libraries may be supported. Support for OpenCV (Open Source Computer Vision Library) may also be provided. Future APIs with compatible 3D pipelines are also supported if they can be mapped from the future API pipeline to the graphics processor pipeline.
[0121] (Programming of Graphics Pipeline) FIG. 9A is a block diagram showing a graphics processor command format 900 according to some embodiments. FIG. 9B is a block diagram showing a graphics processor command sequence 910 according to an embodiment. The solid box in FIG. 9A indicates components generally included in a graphics command, and the dashed lines include components that are optional or included only in a subset of graphics commands. The exemplary graphics processor command format 900 of FIG. 9A includes data fields for identifying a client 902, a command operation code (opcode) 904, and command data 906. A sub-opcode 905 and a command size 908 are also included in some commands.
[0122] In some embodiments, the client 902 specifies a client unit of a graphics device that processes command data. In some embodiments, a graphics processor command parser examines the client field of each command to route command data to an appropriate client unit conditional on further processing of the command. In some embodiments, the graphics processor client units include a memory interface unit, a rendering unit, a 2D unit, a 3D unit, and a media unit. Each client unit has a corresponding processing pipeline for processing commands. When a command is received by a client unit, the client unit reads the opcode 904 to determine the operation to be performed and, if present, the sub-opcode 905. The client unit executes the command using the information in the data field 906. For some commands, an explicit command size 908 is assumed to specify the size of the command. In some embodiments, the command parser automatically determines the size of at least some of the commands based on the command opcode. In some embodiments, commands are aligned via multiples of a double word. Other command formats can also be used.
[0123] The flowchart in FIG. 9B shows an exemplary graphics processor command sequence 910. In some embodiments, the software or firmware of a data processing system characterized by an embodiment of a graphics processor uses a version of a command sequence that is shown to set up, execute, and then end a set of graphics operations. A sample command sequence is illustrated and described for example purposes only, and embodiments are not limited to these specific commands or this command sequence. Further, commands may be issued as a batch of commands within a command sequence such that the graphics processor processes the sequence of commands at least partially simultaneously.
[0124] In some embodiments, the graphics processor command sequence 910 begins with a pipeline flush command 912 that may complete any commands currently pending for the pipeline in any active graphics pipeline. In some embodiments, the 3D pipeline 922 and the media pipeline 924 do not operate simultaneously. The pipeline flush is executed to complete the commands pending in the active graphics pipeline. In response to the pipeline flush, the command parser of the graphics processor halts command processing until the active rendering engine has completed the pending operations and the associated read cache has been invalidated. Optionally, any data in the rendering cache marked as "dirty" can be flushed to memory. In some embodiments, the pipeline flush command 912 can be used for pipeline synchronization or before putting the graphics processor into a low power state.
[0125] In some embodiments, the pipeline selection command 913 is used when the command sequence requires the graphics processor to explicitly switch between pipelines. In some embodiments, the pipeline selection command 913 is only required once within the execution context before issuing pipeline commands, unless the context issues commands for both pipelines. In some embodiments, the pipeline flush command 912 is required immediately before switching pipelines via the pipeline selection command 913.
[0126] In some embodiments, the pipeline control command 914 is used to configure the graphics pipeline for operations and to program the 3D pipeline 922 and the media pipeline 924. In some embodiments, the pipeline control command 914 configures the pipeline state for the active pipeline. In one embodiment, the pipeline control command 914 is used for pipeline synchronization and to clear data from one or more cache memories within the active pipeline before processing a batch of commands.
[0127] In some embodiments, the return buffer state command 916 is used to configure a set of return buffers for each pipeline for writing data. Some pipeline operations require the allocation, selection, or configuration of one or more return buffers into which the operation writes intermediate data during processing. In some embodiments, the graphics processor also uses one or more return buffers to store output data and perform inter-thread communication. In some embodiments, the return buffer state 916 includes selecting the size and number of return buffers to use for a set of pipeline operations.
[0128] The remaining commands within the command sequence are different based on the active pipeline for the operation. Based on the pipeline determination 920, the command sequence is adjusted to a 3D pipeline 922 starting with a 3D pipeline state 930 or a media pipeline 924 starting with a media pipeline state 940.
[0129] Commands for configuring the 3D pipeline state 930 include 3D state setting commands for vertex buffer state, vertex element state, constant color state, depth buffer state, and other state variables that should be configured before 3D primitive commands are processed. The values of these commands are determined at least in part based on the specific 3D API in use. In some embodiments, the 3D pipeline state 930 commands can also selectively disable or bypass a particular pipeline element if that particular pipeline element is not used.
[0130] In some embodiments, the 3D primitive 932 commands are used to submit 3D primitives to be processed by the 3D pipeline. The commands and associated parameters passed to the graphics processor via the 3D primitive 932 commands are transferred to the vertex fetch function within the graphics pipeline. The vertex fetch function uses the 3D primitive 932 command data to generate a vertex data structure. The vertex data structure is stored in one or more return buffers. In some embodiments, the 3D primitive 932 commands are used to perform vertex operations on the 3D primitive via a vertex shader. To process the vertex shader, the 3D pipeline 922 dispatches shader execution threads to the graphics processor execution units.
[0131] In some embodiments, the 3D pipeline 922 is triggered via an execute 934 command or event. In some embodiments, a register write triggers command execution. In some embodiments, execution is triggered via a "go" or "kick" command within a command sequence. In one embodiment, command execution is triggered using a pipeline synchronization command to flush a command sequence through a graphics pipeline. The 3D pipeline performs geometry processing on 3D primitives. When the operations are complete, the resulting geometric object is rasterized and the pixel engine colors the resulting pixels. Additional commands for controlling pixel shading and pixel backend operations may also be included for these operations.
[0132] In some embodiments, the graphics processor command sequence 910 follows the media pipeline 924 path when performing media operations. Generally, the specific use and manner of programming for the media pipeline 924 depends on the media or compute operations being performed. Certain media decode operations may be offloaded to the media pipeline during media decoding. In some embodiments, the media pipeline can also be bypassed and media decoding can be performed entirely or in part using resources provided by one or more general-purpose processing cores. In one embodiment, the media pipeline also includes elements for general-purpose graphics processor unit (GPGPU) operations and the graphics processor is used to perform SIMD vector operations using a compute shader program that is not explicitly related to the rendering of graphics primitives.
[0133] In some embodiments, the media pipeline 924 is configured in a manner similar to the 3D pipeline 922. A set of commands for configuring the media pipeline state 940 is dispatched or placed in a command queue before the media object commands 942. In some embodiments, the commands for the media pipeline state 940 include data for configuring the media pipeline elements used to process media objects. This includes data for configuring video decoding and video encoding logic within the media pipeline, such as an encoding or decoding format. In some embodiments, the commands for the media pipeline state 940 also support the use of one or more pointers to "indirect" state elements that include a batch of state settings.
[0134] In some embodiments, the media object commands 942 supply pointers to media objects for processing by the media pipeline. A media object includes a memory buffer that contains video data to be processed. In some embodiments, all media pipeline states must be valid before issuing the media object commands 942. Once the pipeline state is configured and the media object commands 942 are queued, the media pipeline 924 is triggered via an execute command 944 or an equivalent execution event (e.g., a register write). The output from the media pipeline 924 may then be post-processed by operations provided by the 3D pipeline 922 or the media pipeline 924. In some embodiments, GPGPU operations are configured and executed in a manner similar to media operations.
[0135] (Graphics Software Architecture) Figure 10 shows an exemplary graphics software architecture for a data processing system 1000 according to some embodiments. In some embodiments, the software architecture includes a 3D graphics application 1010, an operating system 1020, and at least one processor 1030. In some embodiments, the processor 1030 includes a graphics processor 1032 and one or more general-purpose processor cores 1034. The graphics application 1010 and the operating system 1020 each execute within the system memory 1050 of the data processing system.
[0136] In some embodiments, the 3D graphics application 1010 includes one or more shader programs that include shader instructions 1012. The instructions of the shader language may be a high-level shader language such as the High-Level Shader Language (HLSL) of Direct3D, the OpenGL Shader Language (GLSL), etc. The application also includes machine language executable instructions 1014 suitable for execution by the general-purpose processor core 1034. The application also includes a graphics object 1016 defined by vertex data.
[0137] In some embodiments, the operating system 1020 is a Microsoft(R) Windows(R) operating system from Microsoft Corporation, an operating system such as a UNIX (registered trademark) with proprietary specifications, or an open-source UNIX (registered trademark) operating system that uses a variant of the Linux (registered trademark) kernel. The operating system 1020 can support a graphics API 1022 such as the Direct3D API, the OpenGL API, or the Vulkan API. When the Direct3D API is used, the operating system 1020 uses a front-end shader compiler 1024 to compile any shader instruction 1012 of HLSL into a low-level shader language. The compilation may be just-in-time (JIT) compilation, or alternatively, the application can perform pre-compilation of the shader. In some embodiments, the high-level shader is compiled into a low-level shader during compilation of the 3D graphics application 1010. In some embodiments, the shader instruction 1012 is provided in an intermediate format such as a version of SPIR (Standard Portable Intermediate Representation) used by the Vulkan API.
[0138] In some embodiments, the user-mode graphics driver 1026 includes a back-end shader compiler 1027 for converting shader instructions 1012 into hardware-specific representations. When the OpenGL API is used, the shader instructions 1012 in the GLSL high-level language are passed to the user-mode graphics driver 1026 for compilation. In some embodiments, the user-mode graphics driver 1026 uses the operating system kernel-mode function 1028 to communicate with the kernel-mode graphics driver 1029. In some embodiments, the kernel-mode graphics driver 1029 communicates with the graphics processor 1032 to dispatch commands and instructions.
[0139] (Implementation of IP Core) One or more aspects of at least one embodiment may be implemented by an expression code stored on a machine-readable medium that represents and / or defines logic within an integrated circuit such as a processor. For example, the machine-readable medium may include instructions that represent various logic within the processor. When read by a machine, the instructions may cause the machine to manufacture logic for performing the techniques described herein. Such an expression is known as an "IP core" and is a reusable unit of logic for an integrated circuit that can be stored on a tangible machine-readable medium as a hardware model that describes the structure of the integrated circuit. The hardware model may be supplied to various customers or manufacturing facilities, and the various customers or manufacturing facilities may load the hardware model into manufacturing machines that manufacture integrated circuits. The integrated circuit may be manufactured to perform the operations described in connection with any of the embodiments herein.
[0140] FIG. 11A is a block diagram showing an IP core development system 1100 that can be used to manufacture an integrated circuit to execute operations according to an embodiment. The IP core development system 1100 may be used to generate a modular reusable design that can be incorporated into a larger design or used to construct an entire integrated circuit (e.g., an SOC integrated circuit). The design facility 1130 can generate a software simulation 1110 of the IP core design in a high-level programming language (e.g., C / C++). The software simulation 1110 can be used to design, test, and verify the behavior of the IP core using a simulation model 1112. The simulation model 1112 may include functional simulation, behavioral simulation, and / or timing simulation. Next, a register transfer level (RTL) design 1115 can be created or synthesized from the simulation model 1112. The RTL design 1115 is an abstraction of the behavior of an integrated circuit that models the flow of digital signals between hardware registers and includes the associated logic executed using the modeled digital signals. In addition to the RTL design 1115, a low-level design at the logic level or transistor level may also be created, designed, or synthesized. Thus, the specific details of the initial design and simulation may vary.
[0141] The RTL design 1115 or equivalent may be further synthesized into the hardware model 1120 by a design function, and the hardware model 1120 may be in a hardware description language (HDL) or other representation of physical design data. The HDL may be further simulated or tested to verify the IP core design. The IP core design can be stored for delivery to a third-party manufacturing facility 1165 using a non-volatile memory 1140 (e.g., a hard disk, flash memory, or any non-volatile storage medium). Alternatively, the IP core design may be transmitted over a wired connection 1150 or a wireless connection 1160 (e.g., via the Internet). The manufacturing facility 165 may then manufacture an integrated circuit based at least in part on the IP core design. The manufactured integrated circuit can be configured to perform operations according to at least one embodiment described herein.
[0142] FIG. 11B shows a cross-sectional side view of an integrated circuit package assembly 1170 according to some embodiments described herein. The integrated circuit package assembly 1170 shows the implementation of one or more of the processors or accelerator devices described herein. The package assembly 1170 includes a plurality of units of hardware logic 1172, 1174 connected to a substrate 1180. The logic 1172, 1174 may be at least partially implemented in hardware of configurable logic or fixed-function logic and may include one or more portions of any of the processor cores, graphics processors, or other accelerator devices described herein. Each unit of the logic 1172, 1174 is implemented within a semiconductor die and can be coupled to the substrate 1180 via an interconnect structure 1173. The interconnect structure 1173 may be configured to route electrical signals between the logic 1172, 1174 and the substrate 1180 and can include interconnects such as bumps or pillars, but is not limited thereto. In some embodiments, the interconnect structure 1173 may be configured to route electrical signals such as input / output (I / O) signals and / or power or ground signals related to the operation of the logic 1172, 1174. In some embodiments, the substrate 1180 is an epoxy-based laminate substrate. In other embodiments, the package substrate 1180 may include other suitable types of substrates. The package assembly 1170 can be connected to other electrical devices via package interconnects 1183. The package interconnects 1183 may be coupled to the surface of the substrate 1180 to route electrical signals to other electrical devices such as a motherboard, other chip sets, or multi-chip modules.
[0143] In some embodiments, the units of logic 1172, 1174 are electrically coupled to a bridge 1182 configured to route electrical signals between the logic 1172, 1174. The bridge 1182 may be a high-density interconnect structure that provides a path for the electrical signals. The bridge 1182 may include a bridge substrate composed of glass or a suitable semiconductor material. Electrical routing features may be formed on the bridge substrate to provide an inter-chip connection between the logic 1172, 1174.
[0144] Although two units of logic 1172, 1174 and bridge 1182 are shown, the embodiments described herein may include more or fewer logic units on one or more dies. Since the bridge 1182 may be excluded when the logic is included on a single die, one or more dies may be connected by zero or more bridges. Alternatively, multiple dies or logic units may be connected by one or more bridges. Further, multiple logic units, dies, and bridges may be connected to each other in other possible configurations including three-dimensional configurations.
[0145] FIG. 11C shows a package assembly 1190 that includes a plurality of unit hardware logic chiplets connected to a substrate 1180 (e.g., a base die). The graphics processing units, parallel processors, and / or compute accelerators described herein can be composed of a variety of separately manufactured silicon chiplets. In this context, a chiplet is an at least partially packaged integrated circuit that includes a separate unit of logic that can be assembled with other chiplets into a larger package. A variety of sets of chiplets having different IP core logics can be assembled into a single device. Further, the chiplets can be integrated into the base die or base chiplet using active interposer technology. The concepts described herein enable interconnection and communication between different forms of IP within a GPU. The IP cores can be manufactured using different process technologies and configured during manufacturing. This avoids the complexity of converging multiple IPs into the same manufacturing process, especially on a large SoC having multiple flavors of IP. Enabling the use of multiple process technologies improves time-to-market and provides a cost-effective way to create multiple product SKUs. Further, the decomposed IP is more easily power-gated independently, and components not used for a given workload can be powered off, reducing overall power consumption.
[0146] The hardware logic chiplets can include application-specific hardware logic chiplets 1172, logic or I / O chiplets 1174, and / or memory chiplets 1175. The hardware logic chiplets 1172 and the logic or I / O chiplets 1174 may be at least partially implemented in configurable logic or fixed-function logic hardware and can include one or more portions of any of the processor cores, graphics processors, parallel processors, or other accelerator devices described herein. The memory chiplets 1175 can be DRAM (e.g., GDDR, HBM) memory or cache (SRAM) memory.
[0147] Each chiplet can be manufactured as a separate semiconductor die and coupled to the substrate 1180 via the interconnect structure 1173. The interconnect structure 1173 may be configured to route electrical signals between various chiplets and logic within the substrate 1180. The interconnect structure 1173 can include, but is not limited to, interconnects such as bumps or pillars. In some embodiments, the interconnect structure 1173 may be configured to route electrical signals such as input / output (I / O) signals and / or power or ground signals related to the operation of logic, I / O, and memory chiplets, for example.
[0148] In some embodiments, the substrate 1180 is an epoxy-based laminate substrate. In other embodiments, the substrate 1180 may include other suitable types of substrates. The package assembly 1190 can be connected to other electrical devices via the package interconnect 1183. The package interconnect 1183 may be coupled to the surface of the substrate 1180 to route electrical signals to other electrical devices such as a motherboard, other chip sets, or multi-chip modules.
[0149] In some embodiments, the logic or I / O chiplet 1174 and the memory chiplet 1175 can be electrically coupled via a bridge 1187 configured to route electrical signals between the logic or I / O chiplet 1174 and the memory chiplet 1175. The bridge 1187 may be a high-density interconnect structure that provides a path for electrical signals. The bridge 1187 may include a bridge substrate composed of glass or a suitable semiconductor material. Electrical routing features can be formed on the bridge substrate to provide an inter-chip connection between the logic or I / O chiplet 1174 and the memory chiplet 1175. The bridge 1187 may also be referred to as a silicon bridge or an interconnect bridge. For example, in some embodiments, the bridge 1187 is an Embedded Multi-die Interconnect Bridge (EMIB). In some embodiments, the bridge 1187 may simply be a direct connection from one chiplet to another.
[0150] The substrate 1180 can include hardware components for I / O 1191, cache memory 1192, and other hardware logic 1193. A fabric 1185 can be embedded within the substrate 1180 to enable communication between various logic chiplets within the substrate 1180 and the logic 1191, 1193. In one embodiment, the I / O 1191, the fabric 1185, the cache, the bridge, and the other hardware logic 1193 can be integrated into a base die stacked on top of the substrate 1180.
[0151] In various embodiments, the package assembly 1190 can include a fewer or greater number of components and die interconnected by the fabric 1185 or one or more bridges 1187. The die within the package assembly 1190 may be arranged in a 3D or 2.5D configuration. Generally, the bridge structure 1187 may be used, for example, to effect point-to-point interconnects between logic or I / O die and memory die. The fabric 1185 can be used to interconnect various logic and / or I / O die (e.g., die 1172, 1174, 1191, 1193) to other logic and / or I / O die. In one embodiment, the cache memory 1192 within the substrate can function as a global cache for the package assembly 1190, a part of a distributed global cache, or a dedicated cache for the fabric 1185.
[0152] FIG. 11D shows a package assembly 1194 including an interchangeable die 1195 according to an embodiment. The interchangeable die 1195 can be assembled into standardized slots on one or more base die 1196, 1198. The base die 1196, 1198 can be coupled via a bridge interconnect 1197, which can be similar to other bridge interconnects described herein, e.g., EMIB. Memory die can also be connected to logic or I / O die via a bridge interconnect. The I / O and logic die can communicate via an interconnect fabric. The base die can support one or more slots in a standardized format for one of logic or I / O or memory / cache.
[0153] In one embodiment, the SRAM and the power delivery circuit can be fabricated on one or more of the base chiplets 1196, 1198, which can be fabricated using different process technologies with respect to the replaceable chiplet 1195 stacked on the base chiplet. For example, the base chiplets 1196, 1198 can be fabricated using a larger process technology, and the replaceable chiplet can be fabricated using a smaller process technology. One or more of the replaceable chiplets 1195 may be memory (e.g., DRAM) chiplets. For the package assembly 1194, different memory densities can be selected based on the power and / or performance of the product using the package assembly 1194. Further, logic chiplets having different numbers of types of functional units can be selected at the time of assembly based on the power and / or performance of the product. Further, chiplets including different types of IP logic cores can be inserted into the slots of the replaceable chiplets, enabling a hybrid processor design that can mix and match IP blocks of different technologies.
[0154] (Exemplary System-on-Chip Integrated Circuit) FIGS. 12-14 show exemplary integrated circuits and related graphics processors that can be fabricated using one or more IP cores according to the various embodiments described herein. In addition to those shown, other logic and circuits may be included, including additional graphics processors / cores, peripheral interface controllers, or general-purpose processor cores.
[0155] FIG. 12 is a block diagram showing an exemplary system-on-chip integrated circuit 1200 that may be manufactured using one or more IP cores according to an embodiment. The exemplary integrated circuit 1200 includes one or more application processors 1205 (e.g., CPUs) and at least one graphics processor 1210, and may further include an image processor 1215 and / or a video processor 1220, any of which may be modular IP cores from the same or multiple different design facilities. The integrated circuit 1200 includes peripheral devices or bus logic including a USB controller 1225, a UART controller 1230, an SPI / SDIO controller 1235, and an I 2 S / I 2 2C controller 1240. Further, the integrated circuit may include a display device 1245 coupled to one or more of an HDMI (registered trademark) (high-definition multimedia interface) controller 1250 and a MIPI (mobile industry processor interface) display interface 1255. Storage may be provided by a flash memory subsystem 1260 including a flash memory and a flash memory controller. The memory interface may be provided via a memory controller 1265 for access to SDRAM or SRAM memory devices. Some integrated circuits further include an embedded security engine 1270.
[0156] Figures 13 to 14 are block diagrams showing exemplary graphics processors used within an SoC according to the embodiments described herein. FIG. 13A shows an exemplary graphics processor 1310 of a system-on-chip integrated circuit that can be manufactured using one or more IP cores according to an embodiment. FIG. 13B shows a further exemplary graphics processor 1340 of a system-on-chip integrated circuit that can be manufactured using one or more IP cores according to an embodiment. The graphics processor 1310 of FIG. 13A is an example of a low-power graphics processor core. The graphics processor 1340 of FIG. 13B is an example of a high-performance graphics processor core. Each of the graphics processors 1310, 1340 can be a variation of the graphics processor 1210 of FIG. 12.
[0157] As shown in FIG. 13, the graphics processor 1310 includes a vertex processor 1305 and one or more fragment processors 1315A to 1315N (e.g., 1315A, 1315B, 1315C, 1315D to 1315N - 1, and 1315N). The graphics processor 1310 is optimized such that the vertex processor 1305 executes operations for a vertex shader program, while the one or more fragment processors 1315A to 1315N can execute different shader programs via separate logic so as to execute fragment (e.g., pixel) shader operations for a fragment or pixel shader program. The vertex processor 1305 executes the vertex processing stage of the 3D graphics pipeline and generates primitives and vertex data. The fragment processors 1315A to 1315N use the primitives and vertex data generated by the vertex processor 1305 to generate a frame buffer to be displayed on a display device. In one embodiment, the fragment processors 1315A to 1315N are optimized to execute a fragment shader program as provided in the OpenGL API, and the fragment shader program may be used to perform operations similar to a pixel shader program as provided in the Direct 3D API.
[0158] The graphics processor 1310 further includes one or more memory management units (MMUs) 1320A - 1320B, caches 1325A - 1325B, and circuit interconnects 1330A - 1330B. The one or more MMUs 1320A - 1320B provide virtual - to - physical address mapping for the graphics processor 1310, which includes the vertex processor 1305 and / or the fragment processors 1315A - 1315N, and may reference vertex or image / texture data stored in memory in addition to vertex or image / texture data stored in the one or more caches 1325A - 1325B. In one embodiment, the one or more MMUs 1320A - 1320B may be synchronized with other MMUs in the system, including one or more MMUs associated with the one or more application processors 1205, image processors 1215, and / or video processors 1220 of FIG. 12, such that each processor 1205 - 1220 can participate in a shared or unified virtual memory system. The one or more circuit interconnects 1330A - 1330B enable the graphics processor 1310 to interface with other IP cores within the SoC via the internal bus or direct connection of the SoC, according to an embodiment.
[0159] As shown in FIG. 14, the graphics processor 1340 includes one or more MMUs 1320A-1320B, caches 1325A-1325B, and circuit interconnects 1330A-1330B of the graphics processor 1310 of FIG. 13. The graphics processor 1340 includes one or more shader cores 1355A-1355N (e.g., 1355A, 1355B, 1355C, 1355D, 1355E, 1355F-1355N-1, and 1355N), which provide a unified shader core architecture that can execute all types of programmable shader code, including shader program code for a single core or type of core to implement vertex shaders, fragment shaders, and / or compute shaders. The exact number of shader cores present can vary between embodiments and implementations. Further, the graphics processor 1340 includes an inter-core task manager 1345 that functions as a thread dispatcher for dispatching execution threads to one or more shader cores 1355A-1355N, and a tile unit 1358 that accelerates tile operations for tile-based rendering, where the rendering operations of a scene are subdivided in image space, for example, to utilize local spatial coherence within the scene or to optimize the use of internal caches.
[0160] [Ray Tracing by Machine Learning] As described above, ray tracing is a graphics processing technique in which light propagation is simulated through physically based rendering. One of the important operations in ray tracing is to process visibility queries that require traversal and intersection testing of nodes within a bounding volume hierarchy (BVH).
[0161] Techniques based on ray tracing and path tracing calculate an image by tracing rays and paths through each pixel and using random sampling to compute advanced effects such as shadows, glossiness, indirect illumination, etc. Using only a few samples is fast but produces a noisy image, while using many samples produces a high-quality image but at a very high cost.
[0162] Machine learning includes any circuit, program code, or combination thereof that can innovatively improve the performance of a specified task or can make an innovatively more accurate prediction or judgment. Some machine learning engines can perform these tasks or make these predictions / decisions without being explicitly programmed to perform the task or make the prediction / judgment. There are various machine learning techniques including, but not limited to, supervised learning, semi-supervised learning, unsupervised learning, and reinforcement learning.
[0163] Over the past few years, an epoch-making solution for real-time ray tracing / path tracing has emerged in the form of "denoising". This is a process that uses image processing techniques to generate a high-quality filtered / denoised image from a noisy low-sample count input. The most effective denoising techniques rely on machine learning techniques where a machine learning engine learns how a noisy image would look if it were computed with more samples. In one particular embodiment, the machine learning is performed by a convolutional neural network (CNN), but the underlying principles of the present invention are not limited to the implementation of a CNN. In such an implementation, the training data is generated with low-sample count inputs and ground-truth. The CNN is trained to predict the converged pixels from the vicinity of the noisy pixel inputs around the pixel in question.
[0164] Although not perfect, this AI-based noise removal technology has proven to be surprisingly effective. However, it is important to note that good training data is required. The reason for this is that otherwise the network may predict incorrect results. For example, if an animation movie studio trains a noise removal CNN on past movies with on-land scenes and then tries to use the trained CNN to remove noise from frames from a new movie set on water, the noise removal operation will be executed sub-optimally.
[0165] To address this issue, learning data can be dynamically collected during rendering and a machine learning engine such as a CNN may be continuously trained based on the data currently being executed, thus continuously improving the machine learning engine for the current task. Therefore, the training phase may still be executed before runtime, but the machine learning weights are continuously adjusted as needed during runtime. Thereby, the high cost of calculating the reference data required for training is avoided by restricting the generation of learning data to sub-regions of the image every frame or every N frames. In particular, noisy inputs of frames are generated to remove noise from the entire frame with the current network. Additionally, as described below, small regions of reference pixels are generated and used for continuous training.
[0166] Here, the implementation of the CNN will be described, but any form of machine learning engine may be used, including but not limited to systems that perform supervised learning (e.g., constructing a numerical operation model of a set of data including both inputs and desired outputs), unsupervised learning (e.g., evaluating input data for a particular type of structure), and / or a combination of supervised learning and unsupervised learning.
[0167] Existing noise removal implementations operate during both the training phase and the runtime phase. During the training phase, a network topology that receives an N×N pixel region with various per-pixel data channels such as pixel color, depth, normal, normal deviation, primitive ID, and albedo and generates a final pixel color is defined. A set of "representative" training data is generated with reference to the "desired" pixel color calculated at a very high sample count using a low sample count input for one frame. The network is trained towards these inputs, generating a set of "ideal" weights for the network. In these implementations, the reference data is used to train the weights of the network to most closely match the output of the network to the desired result.
[0168] At runtime, the pre-calculated ideal network weights are loaded and the network is initialized. For each frame, a low sample count image of the noise removal input (i.e., the same as that used for training) is generated. For each pixel, a given neighborhood of the pixel's input is run through the network to predict the "noise-removed" pixel color, generating a noise-removed frame.
[0169] FIG. 15 shows an implementation of initial training. A machine learning engine 1500 (e.g., a CNN) receives an N×N pixel region with various per-pixel data channels such as pixel color, depth, normal, normal deviation, primitive ID, and albedo as high sample count image data 1702 and generates a final pixel color. Representative training data is generated using a low sample count input 1501 for one frame. The network is trained towards these inputs, generating a set of "ideal" weights 1505, which the machine learning engine 1500 then uses to remove noise from low sample count images at runtime.
[0170] To improve the above technology, the noise removal stage for generating new training data for each frame or subset of frames (e.g., every N frames for N = 2, 3, 4, 10, 25, etc.) is enhanced. In particular, as shown in FIG. 16, this involves selecting one or more regions within each frame, herein referred to as the "new reference region" 1602, which is rendered into a separate high-sample-count buffer 1604 with a high sample count. The low-sample-count buffer 1603 stores the low-sample-count input frame 1601 (including the low-sample region 1604 corresponding to the new reference region 1602).
[0171] The position of the new reference region 1602 may be randomly selected. Alternatively, the position of the new reference region 1602 may be adjusted in a predefined manner for each new frame (e.g., using a predefined movement of regions between frames, being restricted to a specified region at the center of the frame, etc.).
[0172] Regardless of how the new reference regions are selected, this is used by the machine learning engine 1600 to continuously refine and update the trained weights 1605 used for noise removal. In particular, the reference pixel colors from each new reference region 1602 and the noisy reference pixel inputs from the corresponding low sample count regions 1607 are rendered. Then, using the high sample count reference regions 1602 and the corresponding low sample count regions 1607, supplementary training is performed on the machine learning engine 1600. In contrast to the initial training, this training is continuously performed during runtime for each new reference region 1602, thereby ensuring that the machine learning engine 1600 is accurately trained. For example, per-pixel data channels (e.g., pixel color, depth, normal, normal deviation, etc.) may be evaluated, and the machine learning engine 1600 uses the per-pixel data channels to adjust the trained weights 1605. Similar to the case of training (Figure 15), the machine learning engine 1600 is trained towards an ideal set of weights 1605 to remove noise from the low sample count input frame 1601 and generate a denoised frame 1620. However, the trained weights 1605 are continuously updated based on the new image characteristics of the new type of low sample count input frame 1601.
[0173] The retraining operation performed by the machine learning engine 1600 may be executed simultaneously as a background process on a graphics processor unit (GPU) or the host processor. The rendering loop, which may be implemented as a driver component and / or a GPU hardware component, may continuously generate new training data (e.g., in the form of new reference regions 1602) to place in the queue. The background training process executed on the GPU or the host processor may continuously read new training data from this queue, retrain the machine learning engine 1600, and update it with new weights 1605 at appropriate intervals.
[0174] FIG. 17 shows an example of such an implementation where the background training process 1700 is implemented by the host CPU 1710. In particular, the background training process 1700 uses the high sample number new reference region 1602 and the corresponding low sample region 1604 to continuously update the trained weights 1605, thereby updating the machine learning engine 1600.
[0175] As shown in FIG. 18A for a non-limiting example of a multiplayer online game, different host machines 1820-1822 individually generate reference regions that the background training processes 1700A-C send to the server 1800 (e.g., a game server, etc.). The server 1800 then uses the new reference regions received from each of the hosts 1821-1822 to perform training on the machine learning engine 1810 and update the weights 1805 as described above. The server 1800 sends these weights 1805 to the host machine 1820, and the host machine 1820 stores the weights 1605A-C, thereby updating the respective individual machine learning engines (not shown). Since the server 1800 may be provided with a large number of reference regions in a short period of time, the weights for any given application (e.g., an online game) executed by the user can be updated efficiently and accurately.
[0176] As shown in FIG. 18B, different host machines may generate new trained weights (e.g., based on the training / reference region 1602 described above) and share the new trained weights with the server 1800 (e.g., a game server or the like), or alternatively, a peer-to-peer sharing protocol may be used. The machine learning management component 1810 on the server uses the new weights received from each of the host machines to generate a set 1805 of combined weights. The combined weights 1805 may be, for example, an average value generated from the new weights and continuously updated as described herein. Once generated, copies of the combined weights 1605A - C may be sent to and stored in each of the host machines 1820 - 1821, and then the host machines 1820 - 1821 may use the combined weights as described herein to perform noise removal operations.
[0177] Semi-closed loop update mechanisms can also be used by hardware manufacturers. For example, the reference network may be included as part of a driver distributed by the hardware manufacturer. The driver generates new training data using the techniques described herein and continuously sends these back to the hardware manufacturer, who uses this information to continue to improve the implementation of machine learning for the next driver update.
[0178] In an exemplary implementation (e.g., batch movie rendering at a rendering farm), the renderer sends the newly generated training region to a dedicated server or database (within the rendering farm of its studio), and the dedicated server or database aggregates this data from multiple rendering nodes over time. Another process on a separate machine continuously improves the studio's dedicated noise removal network, and new rendering jobs always use the latest trained network.
[0179] A machine learning method is shown in FIG. 19. The method may be implemented in the architecture described herein, but is not limited to any particular system or graphics processing architecture.
[0180] In 1901, as part of the initial training phase, low-sample-number image data and high-sample-number image data are generated for a plurality of image frames. In 1902, a machine learning noise removal engine is trained using the high / low sample number image data. For example, a set of convolutional neural network weights related to pixel features may be updated according to the training. However, any machine learning architecture may be used.
[0181] In 1903, at runtime, a low-sample-number image frame is generated along with at least one reference region having a high sample number. In 1904, the high-sample-number reference region is used by the machine learning engine and / or separate training logic (e.g., the background training module 1700) to continuously refine the training of the machine learning engine. For example, the high-sample-number reference region may be used in combination with the corresponding portion of the low-sample-number image to continue to teach the machine learning engine 1904 the most effective way to perform noise removal. For example, in the implementation of a CNN, this may include updating the weights associated with the CNN.
[0182] Multiple variations such as the manner in which a feedback loop to the machine learning engine is configured, the entity that generates the training data, how the training data is fed back to the training engine, and how the improved network is provided to the rendering engine may be implemented. Further, the above example performs continuous training using a single reference region, but any number of reference regions may be used. Further, as described above, the reference regions may be of different sizes, used on different numbers of image frames, and placed at different positions within the image frames using different techniques (e.g., randomly, according to a predefined pattern, etc.).
[0183] Furthermore, although a convolutional neural network (CNN) is described as an example of the machine learning engine 1600, the underlying principle of the present invention may be implemented using any form of machine learning engine that can continuously refine the results using new training data. By way of example and not limitation, other machine learning implementations include, among others, GMDH (group method of data handling), long short-term memory, deep reservoir computing, deep belief network, tensor deep stacking network, and deep predictive coding network.
[0184] [Apparatus and Method for Efficient Distributed Noise Removal] As described above, noise removal is an important feature for real-time ray tracing using a smooth, noise-free image. Rendering can be performed among distributed systems on multiple devices, but heretofore, all existing noise removal frameworks operate on a single instance on a single machine. When rendering is performed among multiple devices, they may not have all the rendered pixels accessible for calculating the noise-removed portions of the image.
[0185] A distributed noise removal algorithm is presented that functions with both artificial intelligence (AI) and non-AI-based noise removal techniques. Regions of the image are already distributed among nodes from the distributed rendering operation or are split and distributed from a single frame buffer. If necessary, ghost regions of the neighboring area required to calculate sufficient noise removal are collected from neighboring nodes, and the resulting tiles are composited into the final image.
[0186] (Dispersion processing) Figure 20 shows a plurality of nodes 2021 to 2023 that perform rendering. Only three nodes are shown for simplicity, but the underlying principle of the present invention is not limited to any particular number of nodes. In fact, a single node may be used to implement a particular embodiment of the present invention.
[0187] Nodes 2021 to 2023 each render a part of the image, resulting in regions 2011 to 2013 in this example. Rectangular regions 2011 to 2013 are shown in Figure 20, but regions of any shape may be used, and any device can process any number of regions. The regions required for the nodes to perform a sufficiently smooth noise removal operation are called ghost regions 2011 to 2013. In other words, ghost regions 2001 to 2003 represent the entire data required to perform noise removal at a specified level of quality. Reducing the quality level reduces the size of the ghost region and thus the amount of data required, and increasing the quality level increases the ghost region and the corresponding data required.
[0188] If a node such as node 2021 has a local copy of a part of the ghost region 2001 necessary for denoising its region 2011 at a specified level of quality, the node obtains the necessary data from one or more "neighboring" nodes such as node 2022 that owns a part of the ghost region 2001 as shown in the figure. Similarly, if node 2022 has a local copy of a part of the ghost region 2002 necessary for denoising its region 2012 at a specified level of quality, node 2022 obtains the necessary ghost region data 2032 from node 2021. The obtaining may be performed over a bus, an interconnect, a high-speed memory fabric, a network (e.g., Gigabit Ethernet), or an on-chip interconnect within a multi-core chip where rendering operations can be distributed among multiple cores (e.g., used for rendering large images either at extreme resolutions or with temporal variations). Each of nodes 2021 - 2023 may include an individual execution unit or a specified set of execution units within a graphics processor.
[0189] The specific amount of data transmitted depends on the denoising technique being used. Further, the data from the ghost region may include any data necessary to improve the denoising of each region. For example, the ghost region data may include color / wavelength, intensity / alpha data, and / or normals of an image. However, the underlying principle of the present invention is not limited to any particular set of ghost region data.
[0190] (Further details) For slower networks or interconnections, compression of this data can be utilized using existing general-purpose reversible or irreversible compression. Examples include, but are not limited to, zlib, gzip, and LZMA (Lempel-Ziv-Markov chain algorithm). Further, note that the deltas in the ray hit information between frames may be very sparse, and when a node already has deltas collected from previous frames, only the samples contributing to those deltas need to be transmitted. Thus, content-specific compression may be used. These can be selectively pushed to the node i collecting these samples, or node i can request samples from other nodes. Reversible compression is used for specific types of data and program code, and irreversible data is used for other types of data.
[0191] Figure 21 shows further details of the interaction between nodes 2021 - 2022. Each of nodes 2021 - 2022 includes ray tracing rendering circuits 2081 - 2082 for rendering respective image regions 2011 - 2012 and ghost regions 2001 - 2002. Noise removers 2100 - 2111 perform noise removal operations respectively in the regions 2011 - 2012 where each of nodes 2021 - 2022 is responsible for rendering and noise removal. For example, noise removers 2021 - 2022 may each include a circuit, software, or a combination of both for generating noise removal regions 2121 - 2122. As described above, when generating the noise removal regions, noise removers 2021 - 2022 may need to depend on data within ghost regions owned by different nodes (e.g., noise remover 2100 may need data from ghost region 2002 owned by node 2022).
[0192] Therefore, the noise removers 2100 to 2111 may each generate noise removal regions 2121 to 2122 using data from regions 2011 to 2012 and ghost regions 2001 to 2002, and at least a portion of the ghost regions 2001 to 2002 may be received from other nodes. The region data managers 2101 to 2102 may manage data transfer from the ghost regions 2001 to 2002 as described herein. The compressor / decompressor units 2131 to 2132 may perform compression and decompression of ghost region data exchanged between nodes 2021 to 2022, respectively.
[0193] For example, the region data manager 2101 of node 2021, in response to a request from node 2022, transmits data from the ghost region 2001 to the compressor / decompressor 2131, and the compressor / decompressor 2131 compresses the data to generate compressed data 2106 to be transmitted to node 2022, thereby reducing the bandwidth on the interconnect, network, bus, or other data communication link. Next, the compressor / decompressor 2132 of node 2022 decompresses the compressed data 2106, and the noise remover 2111 uses the decompressed ghost data to generate a higher-quality noise removal region 2012 than would be possible with only the data from region 2012. The region data manager 2102 may store the decompressed data from the ghost region 2001 in a cache, memory, register file, or other storage for use by the noise remover 2111 when generating the noise removal region 2122. A similar set of operations may be performed to provide data from the ghost region 2002 to the noise remover 2100 on node 2021, and the noise remover 2100 uses the data in combination with the data from region 2011 to generate a higher-quality noise removal region 2121.
[0194] (Data capture or rendering) If the connection between devices such as nodes 2021 - 2022 is slow (i.e., lower than the threshold waiting time and / or threshold bandwidth), it may be faster to render the ghost region locally rather than request results from other devices. This can be determined at runtime by tracking the network transaction speed and extrapolating linearly for the rendering time for the ghost region size. If it is faster to render the entire ghost region for output, multiple devices may end up rendering the same part of the image. The resolution of the rendered part of the ghost region may be adjusted based on the variance of the base region and the determined degree of blur.
[0195] (Load balancing) Static and / or dynamic load balancing schemes may be used to distribute the processing load among various nodes 2021 - 2023. For dynamic load balancing, the distribution determined by the noise removal filter drives the amount of samples used to render a particular region of the scene such that more time is required for noise removal but fewer samples are required for less - variance blurred regions of the image. The particular region assigned to a particular node may be adjusted dynamically based on data from previous frames or may be communicated dynamically between devices while rendering so that all devices have the same amount of work.
[0196] FIG. 22 shows how to collect performance metric data including, but not limited to, the time consumed by monitors 2201-2202 executing on respective nodes 2021-2022 to send data on network interfaces 2211-2212, the time consumed when removing noise from regions (regardless of the presence or absence of ghost region data), and the time consumed to render each region / ghost region. Monitors 2201-2202 report these performance metrics to a manager or load balancing node 2201, and the manager or load balancing node 2201 analyzes the data to identify the current workload on each node 2021-2022 and potentially determines a more efficient mode for processing various noise removal regions 2121-2122. The manager node 2201 then distributes new workloads for new regions to nodes 2021-2022 according to the detected load. For example, the manager node 2201 may send more work to nodes that are not overloaded, and / or may reassign work from overloaded nodes. Additionally, the load balancing node 2201 may send reconfiguration commands to adjust the specific manner in which rendering and / or noise removal is performed by each of the nodes (some examples of which are described above).
[0197] (Determination of Ghost Region) The size and shape of the ghost regions 2001-2002 may be determined based on the noise removal algorithms implemented by the noise removers 2100-2111. Then, each of these sizes can be dynamically modified based on the detected variance of the noise-removed samples. The learning algorithm used for AI noise removal itself may be used to determine the appropriate region size, or in other cases such as bilateral blur, a predetermined filter width determines the size of the ghost regions 2001-2002. In an exemplary implementation using a learning algorithm, the machine learning engine may be executed at the manager node 2201, and / or a part of the machine learning may be executed at each of the individual nodes 2021-2023 (see, for example, FIGS. 18A-18B and the related text above).
[0198] (Collection of the final image) The final image may be generated by collecting the rendered and noise-removed regions from each of the nodes 2021-2023 without requiring a ghost region or a normal vector. In FIG. 22, for example, the noise-removed regions 2121-2122 are sent to the region processor 2280 of the manager node 2201 that combines the regions to generate the final noise-removed image 2290, and then the final noise-removed image 2290 is displayed on the display 2290. The region processor 2280 may combine the regions using various 2D compositing techniques. Although shown as separate components, the region processor 2280 and the noise-removed image 2290 may be integrated into the display 2290. The various nodes 2021-2022 may use direct transmission techniques to send the noise-removed regions 2121-2122, potentially using various non-reversible or reversible compressions of the region data.
[0199] AI noise removal remains a computationally expensive operation as games migrate to the cloud. Therefore, distributed processing of noise removal among multiple nodes between 2021 and 2022 may be necessary to achieve the real-time frame rates required for conventional games or virtual reality (VR) that require higher frame rates. Movie studios also often render on large-scale rendering farms that can be utilized for faster noise removal.
[0200] An exemplary method for performing distributed rendering and noise removal is shown in FIG. 23. The method may be implemented with respect to the system architecture described above, but is not limited to a particular system architecture.
[0201] At 2301, graphics work is dispatched to a plurality of nodes that perform ray tracing operations to render regions of an image frame. Each node may already have the data necessary to perform the operations in memory. For example, two or more nodes may share a common memory, or the local memory of a node may already store data from previous ray tracing operations. Alternatively or additionally, specific data may be sent to each node.
[0202] At 2302, a "ghost region" required for a specified level of noise removal (i.e., an acceptable level of performance) is determined. The ghost region includes any data required to perform a specified level of noise removal, including data owned by one or more other nodes.
[0203] At 2303, data associated with the ghost region (or a portion thereof) is exchanged between nodes. At 2304, each node performs noise removal on its respective region (e.g., using the exchanged data), and at 2305, the results are combined to generate a final noise-removed image frame.
[0204] A manager node or a primary node as shown in FIG. 22 may dispatch work to the nodes and then combine the work performed by the nodes to generate a final image frame. A peer-based architecture can be used, and the nodes are peers that exchange data to render and denoise the final image frame.
[0205] The nodes described herein (e.g., nodes 2021-2023) may be a graphics processing computing system interconnected via a high-speed network. Alternatively, the nodes may be individual processing elements coupled to a high-speed memory fabric. All of the nodes may share a common virtual memory space and / or a common physical memory. Alternatively, the nodes may be a combination of a CPU and a GPU. For example, the above manager node 2201 may be a CPU and / or software executed on the CPU, and nodes 2021-2022 may be a GPU and / or software executed on the GPU. Various different types of nodes may be used and still follow the underlying principles of the present invention.
[0206] (Exemplary Neural Network Implementation) There are many types of neural networks, and a simple type of neural network is a feedforward network. A feedforward network may be implemented as an acyclic graph in which nodes are arranged in layers. Typically, a feedforward network topology includes an input layer and an output layer separated by at least one hidden layer. The hidden layer transforms the input received by the input layer into a representation useful for generating an output at the output layer. Network nodes are fully connected to nodes in adjacent layers via edges, but there are no edges between nodes within each layer. Data received at nodes in the input layer of a feedforward network is propagated (i.e., "fed forward") to nodes in the output layer through an activation function that calculates the state of nodes in each successive layer within the network based on coefficients ("weights") associated with each of the edges connecting the layers. Depending on the particular model represented by the algorithm being executed, the output from a neural network algorithm can take various forms.
[0207] Before a machine learning algorithm can be used to model a particular problem, the algorithm is trained using a training data set. Training a neural network involves selecting a network topology, using a set of training data that represents the problem being modeled by the network, and adjusting the weights until the network model performs with minimal error for all instances of the training data set. For example, during the supervised learning training process of a neural network, the output generated by the network in response to an input representing an instance within the training data set is compared to the "correct" labeled output for that instance, an error signal representing the difference between the output and the labeled output is calculated, and the weights associated with the connections are adjusted to minimize that error as the error signal is propagated backward through the layers of the network. When the error for each output generated from an instance of the training data set is minimized, the network is considered to be "trained."
[0208] The accuracy of machine learning algorithms can be significantly affected by the quality of the dataset used to train the algorithms. The training process can be computationally intensive and may require a significant amount of time on conventional general-purpose processors. Therefore, parallel processing hardware is used to train many types of machine learning algorithms. This is particularly useful for optimizing the training of neural networks because the calculations performed when adjusting the coefficients in neural networks are naturally suited for parallel implementation. Specifically, many machine learning algorithms and software applications are adapted to use the parallel processing hardware within general-purpose graphics processing devices.
[0209] FIG. 24 is a generalized diagram of a machine learning software stack 2400. The machine learning application 2402 can be configured to train a neural network using a training dataset or to use a deep neural network trained to implement machine intelligence. The machine learning application 2402 can include training and inference capabilities for neural networks and / or special software that can be used to train a neural network before deployment. The machine learning application 2402 can implement any type of machine intelligence including, but not limited to, image recognition, mapping and positioning, autonomous navigation, speech synthesis, medical imaging, or language translation.
[0210] Hardware acceleration for the machine learning application 2402 can be enabled via the machine learning framework 2404. The machine learning framework 2404 may be implemented on the hardware described herein, such as the processing system 100 including the processors and components described herein. Elements described with respect to FIG. 24 having the same or similar names as elements in any of the other drawings herein describe the same elements as those in the other drawings and can operate or function in a similar manner as described elsewhere herein, can include the same components, and can be linked to other entities, but are not limited to such. The machine learning framework 2404 can provide a library of machine learning primitives. Machine learning primitives are the basic operations generally performed by machine learning algorithms. Without the machine learning framework 2404, developers of machine learning algorithms would be required to create and optimize the main computational logic associated with the machine learning algorithm and then re-optimize the computational logic as new parallel processors are developed. Instead, the machine learning application can be configured to perform the necessary computations using the primitives provided by the machine learning framework 2404. Typical primitives include tensor convolution, activation functions, and pooling, which are computational operations performed while training a convolutional neural network (CNN). The machine learning framework 2404 can also provide primitives that implement basic linear algebra subprograms performed by many machine learning algorithms, such as matrix and vector operations.
[0211] The machine learning framework 2404 can process the input data received from the machine learning application 2402 and generate appropriate input for the computing framework 2406. The computing framework 2406 can abstract the basic instructions provided to the GPGPU driver 2408 to enable the machine learning framework 2404 to utilize hardware acceleration via the GPGPU hardware 2410 without requiring the machine learning framework 2404 to have detailed knowledge of the architecture of the GPGPU hardware 2410. Further, the computing framework 2406 can enable hardware acceleration for the machine learning framework 2404 across various types and generations of GPGPU hardware 2410.
[0212] (GPGPU Machine Learning Acceleration) FIG. 25 shows a multi-GPU computing system 2500, which may be a variation of the processing system 100. Thus, the disclosure of any feature desired to be combined with the processing system 100 here also discloses, but is not limited to, corresponding combinations with the multi-GPU computing system 2500. Elements of FIG. 25 having the same or similar names as elements in any of the other drawings here describe the same elements as those in the other drawings and can operate or function in a similar manner as described elsewhere here, can include the same components, and can be linked to other entities, but are not limited to such. The multi-GPU computing system 2500 can include a processor 2502 coupled to a plurality of GPGPUs 2506A - D via a host interface switch 2504. The host interface switch 2504 can be, for example, a PCI Express switch device that couples the processor 2502 to a PCI Express bus through which the processor 2502 can communicate with a set of GPGPUs 2506A - D. Each of the plurality of GPGPUs 2506A - D can be an instance of the GPGPU described above. The GPGPUs 2506A - D can be interconnected via a set of high-speed point-to-point GPU - to - GPU links 2516. The high-speed GPU - to - GPU links can be connected to each of the GPGPUs 2506A - D via dedicated GPU links. The P2P GPU link 2516 enables direct communication between each of the GPGPUs 2506A - D without the need for communication on the host interface bus to which the processor 2502 is connected. Due to the GPU - to - GPU traffic directed to the P2P GPU link, the host interface bus remains available for system memory access or for communication with other instances of the multi-GPU computing system 2500, for example, via one or more network devices. Instead of connecting the GPGPUs 2506A - D to the processor 2502 via the host interface switch 2504, the processor 2502 can include direct support for the P2P GPU link 2516 and thus can be directly connected to the GPGPUs 2506A - D.
[0213] (Implementation of Machine Learning Neural Network) The computational architecture described herein can be configured to perform a type of parallel processing particularly suitable for training and deploying neural networks for machine learning. A neural network can be generalized as a network of functions having a graph relationship. As is well known in the art, there are various types of implementations of neural networks used in machine learning. One exemplary type of neural network is a feedforward network as described above.
[0214] A second exemplary type of neural network is a Convolutional Neural Network (CNN). A CNN is a special feedforward neural network for processing data having a known grid-like topology such as image data. Thus, CNNs are commonly used in computer vision and image recognition applications, but may also be used in other types of pattern recognition such as speech and language processing. The nodes in the CNN input layer are organized into a set of "filters" (feature detectors excited by receptive fields in the retina), and the output of each set of filters is propagated to nodes in successive layers of the network. The computation of a CNN involves applying a convolutional numerical operation to each filter to generate the output of the filter. Convolution is a special type of numerical operation performed by two functions to generate a third function that is a modified version of one of the two original functions. In the terminology of convolutional networks, the first function for convolution may be called the input, and the second function may be called the convolutional kernel. The output may be called the feature map. For example, the input to a convolutional layer can be a multi-dimensional array of data defining various color components of an input image. The convolutional kernel can be a multi-dimensional array of parameters that are adapted by a training process for the neural network.
[0215] A recurrent neural network (RNN) is a family of feedforward neural networks that includes feedback connections between layers. By sharing parameter data between different parts of the neural network, RNN enables the modeling of sequential data. The architecture of the RNN contains cycles. Since at least a portion of the output data from the RNN is used as feedback for processing subsequent inputs in the sequence, the cycles represent the influence of the current value of a variable on its own value at a future point in time. This feature makes the RNN particularly useful for language processing due to the variability that language data can exhibit.
[0216] The drawings described below illustrate exemplary feedforward, CNN, and RNN networks and explain general processes for training and deploying each of these types of networks. These explanations are exemplary and non-limiting, and it is understood that the concepts illustrated are generally applicable to deep neural networks and machine learning techniques.
[0217] The above-exemplified neural networks can be used to perform deep learning. Deep learning is machine learning using deep neural networks. The deep neural networks used in deep learning are artificial neural networks composed of multiple hidden layers, as opposed to shallow neural networks that include only a single hidden layer. Deeper neural networks are generally more computationally intensive to train. However, additional hidden layers in the network enable multi-step pattern recognition that results in reduced output error compared to shallow machine learning techniques.
[0218] Deep neural networks used in deep learning typically include a front-end network for performing feature recognition coupled to a back-end network, where the back-end network represents a numerical computation model that can perform computations (e.g., object classification, speech recognition, etc.) based on the feature representation provided to the model. Deep learning enables machine learning to be performed without the need for handcrafted feature engineering to be performed on the model. Instead, deep neural networks can learn features based on the statistical structure or correlations within the input data. The learned features can be provided to a numerical computation model that can map the detected features to the output. The numerical computation model used by the network is generally specialized for the particular task being performed, and different models are used to perform different tasks.
[0219] Once the neural network is constructed, a learning model can be applied to the network to train the network for performing a particular task. The learning model describes how to adjust the weights within the model to reduce the output error of the network. Backpropagation of error is a common method used to train neural networks. An input vector is presented to the network for processing. The output of the network is compared to the desired output using a loss function, and an error value is calculated for each neuron in the output layer. The error values are then propagated backwards until each neuron has an associated error value that roughly represents its contribution to the original output. The network can then learn from these errors using an algorithm such as the stochastic gradient descent algorithm to update the weights of the neural network.
[0220] Figures 26-27 illustrate an exemplary convolutional neural network. Figure 26 shows the various layers within a CNN. As shown in Figure 26, an exemplary CNN used to model image processing can receive an input 2602 that describes the red, green, and blue (RGB) components of an input image. The input 2602 can be processed by a plurality of convolutional layers (e.g., convolutional layer 2604, convolutional layer 2606). The output from the plurality of convolutional layers can optionally be processed by a set 2608 of fully connected layers. As described above for feed-forward networks, the neurons in a fully connected layer have a complete connection to all of the activations in the previous layer. The output from the fully connected layer 2608 can be used to generate the output result from the network. The activations within the fully connected layer 2608 can be computed using matrix multiplication instead of convolution. Not all CNN implementations utilize fully connected layers. For example, in some implementations, the convolutional layer 2606 can generate the output for the CNN.
[0221] Convolutional layers are sparsely connected and different from the traditional neural network configuration found in the fully connected layer 2608. Traditional neural network layers are fully connected such that all output units interact with all input units. However, convolutional layers are sparsely connected as the output of the convolution of a field is input into the nodes of the subsequent layer (instead of the respective state values of each node within the field) as shown. The kernel associated with the convolutional layer performs a convolution operation and its output is sent to the next layer. The dimensionality reduction performed within the convolutional layer is one aspect that enables the CNN to scale to process large images.
[0222] Figure 27 shows an exemplary computational stage within the convolutional layer of a CNN. The input to the convolutional layer 2712 of the CNN can be processed in three stages in the convolutional layer 2714. The three stages can include a convolution stage 2716, a detection stage 2718, and a pooling stage 2720. The convolutional layer 2714 can then output data to a successive convolutional layer. The final convolutional layer of the network can generate output feature map data or provide an input to a fully connected layer, for example, to generate classification values for the input to the CNN.
[0223] The convolution stage 2716 performs several convolutions in parallel to generate a set of linear activations. The convolution stage 2716 may include an affine transformation, which can be any transformation that can be specified as a linear transformation with a translation added. Affine transformations include rotation, translation, scaling (enlargement and reduction), and combinations of these transformations. The convolution stage calculates the output of a function (such as a neuron) connected to a specific region within the input. This can be determined as the local region associated with the neuron. The neuron calculates the dot product between the weights of the neuron and the region within the local input to which the neuron is connected. The output from the convolution stage 2716 defines a set of linear activations that are processed by successive stages of the convolutional layer 2714.
[0224] The linear activations can be processed by the detection stage 2718. In the detection stage 2718, each linear activation is processed by a non-linear activation function. The non-linear activation function increases the non-linear characteristics of the overall network without affecting the receptive field of the convolutional layer. Several types of non-linear activation functions may be used. One specific type is the rectified linear unit (ReLU), which uses an activation function defined as f(x) = max(0, x) such that the activation is thresholded at zero.
[0225] The pooling stage 2720 uses a pooling function that replaces the output of the convolutional layer 2706 with a summary statistic of the nearby outputs. The pooling function can be used to introduce invariance to translations in the neural network so that small translations to the input do not change the pooled output. Invariance to local translations can be useful in scenarios where the presence of features in the input data is more important than the exact location of the features. During the pooling stage 2720, various types of pooling functions can be used, including max pooling, average pooling, and l2-norm pooling. Additionally, some CNN implementations do not include a pooling stage. Instead, such implementations substitute an additional convolutional stage with an increased stride for the previous convolutional stage.
[0226] Next, the output from the convolutional layer 2714 can be processed by the next layer 2722. The next layer 2722 can be either an additional convolutional layer or a fully connected layer 2708. For example, the first convolutional layer 2704 in FIG. 27 can output to the second convolutional layer 2706, and the second convolutional layer can output to the first layer of the fully connected layer 2808.
[0227] FIG. 28 shows an exemplary recurrent neural network 2800. In a recurrent neural network (RNN), the previous state of the network affects the output of the current state of the network. RNNs can be constructed in various ways using various functions. The use of RNNs generally centers around using numerical computational models to predict the future based on a sequence of previous inputs. For example, an RNN may be used to perform statistical language modeling to predict the next word when given a sequence of previous words. The illustrated RNN 2800 can be described as having an input layer 2802 that receives an input vector, a hidden layer 2804 that implements a recurrence function, a feedback mechanism 2805 that enables "memory" of previous states, and an output layer 2806 that outputs a result. The RNN 2800 operates based on time steps. The state of the RNN at a given time step is affected based on the previous time step via the feedback mechanism 2805. For a given time step, the state of the hidden layer 2804 is defined by the previous state and the input at the current time step. The initial input (x1) at the first time step can be processed by the hidden layer 2804. The second input (x2) can be processed by the hidden layer 2804 using the state information determined during the processing of the initial input (x1). A given state can be calculated as s_t = f(Ux_t + Ws_(t - 1)), where U and W are parameter matrices. The function f is generally non-linear, such as a hyperbolic tangent function (Tanh) or a variant of the rectifier function f(x) = max(0, x). However, the specific numerical computational function used in the hidden layer 2804 can vary depending on the details of the specific implementation of the RNN 2800.
[0228] In addition to the basic CNN and RNN networks described above, variations of these networks may be possible. An example of a variation of the RNN is the long short term memory (LSTM) RNN. The LSTM RNN can learn long-term dependencies that may be required to process longer sequences of language. A variation of the CNN is the convolutional deep belief network, which has a similar structure to the CNN and is trained in a similar manner to the deep belief network. A deep belief network (DBN) is a generative neural network composed of multiple layers of probabilistic (random) variables. The DBN can be trained layer by layer using greedy unsupervised learning. The learned weights of the DBN can then be used to provide a pre-trained neural network by determining an optimal initial set of the neural network's weights.
[0229] Figure 29 shows the training and deployment of a deep neural network. When a given network is structured for a task, the neural network is trained using a training data set 2902. Various training frameworks 2904 have been developed to enable hardware acceleration of the training process. For example, the machine learning framework described above may be configured as a training framework. The training framework 2904 can draw in an untrained neural network 2906 and enable the untrained neural net to be trained using the parallel processing resources described herein to generate a trained neural net 2908.
[0230] To start the training process, the initial weights may be selected randomly or pre-trained using a deep belief network. The training cycle is then executed either in a supervised or unsupervised manner.
[0231] Supervised learning is a learning method in which training is performed as an intermediary operation when the training dataset 2902 contains inputs paired with the desired output for the inputs, or when the training dataset contains inputs with known outputs and the output of the neural network is manually scored. The network processes the inputs and compares the resulting output against the expected or desired output. The error is then backpropagated through the system. The training framework 2904 can be adjusted to adjust the weights that control the untrained neural network 2906. The training framework 2904 can provide tools to monitor how the untrained neural network 2906 is converging towards a model suitable for generating the correct answers based on the known input data. The training process occurs repeatedly as the weights of the network are adjusted to refine the output generated by the neural network. The training process can continue until the neural network reaches a statistically desired accuracy associated with the trained neural net 2908. The trained neural network 2908 can then be deployed to implement any number of machine learning operations.
[0232] Unsupervised learning is a learning method in which the network attempts to train itself using unlabeled data. Thus, for unsupervised learning, the training dataset 2902 contains input data without associated output data. The untrained neural network 2906 can learn groupings within the unlabeled inputs and determine how individual inputs relate to the overall dataset. Unsupervised training can be used to generate self-organizing maps, which are a type of trained neural network 2907 that can perform operations useful for reducing the dimensionality of the data. Unsupervised training can also be used to perform anomaly detection, which enables the identification of data points within an input dataset that deviate from the normal patterns of the data.
[0233] Variations for supervised and unsupervised training may also be used. Semi-supervised learning is a technique that includes a mixture of labeled and unlabeled data from the same distribution in the training dataset 2902. Incremental learning is a variation of supervised learning where the input data is continuously used to further train the model. Incremental learning enables the trained neural network 2908 to adapt to new data 2912 without forgetting the knowledge that has permeated into the network during initial training.
[0234] Whether supervised or unsupervised, the training process, especially for deep neural networks, can be too computationally intensive for a single computing node. Instead of using a single computing node, a distributed network of computing nodes can be used to accelerate the training process.
[0235] FIG. 30A is a block diagram showing distributed learning. Distributed learning is a training model that uses a plurality of distributed computing nodes, such as the nodes described above, to perform supervised or unsupervised training of a neural network. The distributed computing nodes can each include one or more host processors and one or more of the general-purpose processing nodes, such as a massively parallel general-purpose graphics processing unit. As shown, distributed learning can be performed with model parallelization 3002, data parallelization 3004, or a combination of model and data parallelization.
[0236] In model parallelization 3002, different computing nodes within a distributed system can execute training computations for different parts of a single network. For example, each layer of a neural network can be trained by different processing nodes of the distributed system. The advantages of model parallelization include, in particular, the ability to scale to particularly large models. Dividing the computations associated with different layers of a neural network enables the training of very large neural networks where the weights of all layers do not fit into the memory of a single computing node. In some examples, model parallelization can be particularly useful when performing unsupervised training of large neural networks.
[0237] In data parallelization 3004, different nodes of a distributed network have a complete instance of the model, and each node receives a different part of the data. Then, the results from different nodes are combined. Different techniques for data parallelization are possible, but all data parallel training techniques require a technique for combining the results and synchronizing the model parameters among the nodes. Exemplary techniques for combining data include parameter averaging and update-based data parallelization. Parameter averaging trains each node on a subset of the training data and sets the global parameters (e.g., weights, biases) to the average of the parameters from each node. Parameter averaging uses a central parameter server to maintain the parameter data. Update-based data parallelization is similar to parameter averaging except that instead of transferring the parameter data from the nodes to the parameter server, updates to the model are transferred. Additionally, update-based data parallelization can be executed in a distributed manner, and the updates are compressed and transferred among the nodes.
[0238] Combined model and data parallelization 3006 can be implemented, for example, in a distributed system where each computing node includes multiple GPUs. Each node can have a complete instance of the model, and separate GPUs within each node are used to train different parts of the model.
[0239] Distributed training has increased overhead compared to training on a single machine. However, the parallel processors and GPGPUs described herein can each implement various techniques for reducing the overhead of distributed training, including techniques for enabling wide-bandwidth GPU-to-GPU data transfer and accelerated remote data synchronization.
[0240] (Exemplary Machine Learning Applications) Machine learning can be applied to solve various technical problems, including but not limited to computer vision, autonomous driving and navigation, speech recognition, and language processing. Conventionally, computer vision has been one of the most active research areas for machine learning applications. Applications of computer vision range from the reproduction of human visual capabilities such as face recognition to the creation of new categories of visual capabilities. For example, a computer vision application can be configured to recognize sound waves from vibrations induced in objects visible within a video. Parallel processor accelerated machine learning enables training using much larger training data sets than previously achievable by computer vision applications and enables deployment of the inference system using low-power parallel processors.
[0241] Parallel processor accelerated machine learning has applications in autonomous driving, including lane and road sign recognition, obstacle avoidance, navigation, and driving control. The accelerated machine learning techniques can be used to train a driving model based on a data set that defines appropriate responses to specific training inputs. The parallel processors described herein enable rapid training of increasingly complex neural networks used in autonomous driving solutions and enable deployment of low-power inference processors in a mobile platform suitable for integration into autonomous vehicles.
[0242] Parallel processor accelerated deep neural networks enable machine learning techniques for automatic speech recognition (ASR). ASR involves the generation of a function that computes the most likely language sequence given an input acoustic sequence. Accelerated machine learning using deep neural networks has enabled the replacement of previously used hidden Markov models (HMMs) and Gaussian mixture models (GMMs) in ASR.
[0243] Parallel processor accelerated machine learning can also be used to accelerate natural language processing. The automated learning procedure can utilize statistical inference algorithms to generate models that are robust to incorrect or rare inputs. Exemplary natural language processor applications include automatic machine translation between human languages.
[0244] The parallel processing platforms used in machine learning can be divided into a training platform and a deployment platform. The training platform is generally super-parallel and includes optimizations for accelerating multi-GPU single-node training and multi-node multi-GPU training. Exemplary parallel processors suitable for training include the super-parallel general-purpose graphics processing units and / or multi-GPU computing systems described herein. In contrast, the deployed machine learning platform generally includes low-power parallel processors suitable for use in products such as cameras, autonomous robots, and self-driving vehicles.
[0245] FIG. 30B shows an exemplary inference system on a chip (SOC) 3100 suitable for performing inference using a trained model. Elements of FIG. 30B having the same or similar names as elements in any of the other figures herein describe the same elements as those in the other figures and can operate or function in a similar manner as described elsewhere herein, can include the same components, and can be linked to other entities, but are not limited to such. The SOC 3100 can integrate processing components including a media processor 3102, a vision processor 3104, a GPGPU 3106, and a multi-core processor 3108. The SOC 3100 can further include on-chip memory 3105 that can enable a shared on-chip data pool accessible by each of the processing components. The processing components can be optimized for low-power operation to enable deployment to various machine learning platforms including autonomous vehicles and autonomous robots. For example, one implementation of the SOC 3100 can be used as part of a main control system for an autonomous vehicle. When the SOC 3100 is configured for use in an autonomous vehicle, the SOC is designed and configured to conform to relevant functional safety references for the deployment jurisdiction.
[0246] During operation, the media processor 3102 and the vision processor 3104 can operate in cooperation to accelerate computer vision operations. The media processor 3102 can enable low-latency decoding of multiple high-resolution (e.g., 4K, 8K) video streams. The decoded video streams can be written to a buffer within the on-chip memory 3105. The vision processor 3104 can then perform preprocessing operations on the frames of the decoded video in preparation for analyzing the decoded video and processing the frames using a trained image recognition model. For example, the vision processor 3104 can accelerate the convolution operations of a CNN used to perform image recognition on high-resolution video data, while the back-end model calculations are performed by the GPGPU 3106.
[0247] The multi-core processor 3108 can include control logic for assisting in the ordering and synchronization of data transfers and the shared memory operations performed by the media processor 3102 and the vision processor 3104. The multi-core processor 3108 can also function as an application processor for executing software applications that can utilize the inference computing capabilities of the GPGPU 3106. For example, at least a portion of the navigation and driving logic can be implemented in software executed on the multi-core processor 3108. Such software can directly issue a compute workload to the GPGPU 3106, or alternatively, the compute workload can be issued to the multi-core processor 3108, which can offload at least a portion of these operations to the GPGPU 3106.
[0248] The GPGPU 3106 can include processing clusters, such as a low-power configuration of processing clusters DPLAB06A - DPLAB06H, within the ultra-parallel general-purpose graphics processing unit DPLAB00. The processing clusters within the GPGPU 3106 can support instructions that are particularly optimized for performing inference calculations on trained neural networks. For example, the GPGPU 3106 can support instructions for performing low-precision calculations, such as 8-bit and 4-bit integer vector operations.
[0249] [Ray Tracing Architecture] In one implementation, the graphics processor includes circuitry and / or program code for performing real-time ray tracing. A dedicated set of ray tracing cores may be included in the graphics processor for performing various ray tracing operations described herein, including ray traversal and / or ray intersection operations. In addition to the ray tracing cores, multiple sets of graphics processing cores for performing programmable shading operations, and multiple sets of tensor cores for performing matrix operations on tensor data may also be included.
[0250] Figure 31 shows an exemplary portion of one such graphics processing unit (GPU) 3105 that includes a dedicated set of graphics processing resources disposed in multi-core groups 3100A - 3100N. The graphics processing unit (GPU) 3105 may be a graphics processor 300, GPGPU 1340, and / or a variation of any other graphics processor described herein. Thus, the disclosure of any feature of a graphics processor also discloses, but is not limited to, the corresponding combination with GPU 3105. Further, elements of FIG. 31 having the same or similar names as elements of any other drawing herein describe the same elements as those of the other drawings and can operate or function in a similar manner as described elsewhere herein, can include the same components, and can be linked to other entities, but are not limited to such. Although details are provided for only a single multi-core group 3100A, it is recognized that the other multi-core groups 3100B - N may comprise the same or a similar set of graphics processing resources.
[0251] As shown in the figure, the multi-core group 2100A may include a set 3130 of graphics cores, a set 3140 of tensor cores, and a set 3150 of ray tracing cores. The scheduler / dispatcher 241 schedules and dispatches graphics threads for execution on various cores 3130, 3140, 3150. A set 3120 of register files stores operand values used by the cores 3130, 3140, 3150 when executing graphics threads. These may include, for example, integer registers for storing integer values, floating-point registers for storing floating-point values, vector registers for storing packed data elements (integer and / or floating-point data elements), and tile registers for storing tensor / matrix values. The tile registers may be implemented as a combined set of vector registers.
[0252] One or more level 1 (L1) caches and texture units 3160 locally store graphics data such as texture data, vertex data, pixel data, ray data, bounding volume data, etc. within each multi-core group 3100A. The level 2 (L2) cache 3180 is shared by all or a subset of the multi-core groups 3100A - 3100N and stores graphics data and / or instructions for multiple concurrent graphics threads. As shown in the figure, the L2 cache 3180 may be shared among multiple multi-core groups 3100A - 3100N. One or more memory controllers 3170 couple the GPU 3105 to a memory 3198, which may be system memory (e.g., DRAM) and / or dedicated graphics memory (e.g., GDDR6 memory).
[0253] The input / output (IO) circuit 3195 couples the GPU 3105 to one or more I / O devices such as a digital signal processor (DSP), a network controller, or a user input device to 3195. The on-chip interconnect may be used to couple the I / O device 3190 to the GPU 3105 and the memory 3198. One or more IO memory management units (IOMMUs) 3170 of the IO circuit 3195 directly couple the IO device 3190 to the system memory 3198. The IOMMU 3170 manages multiple sets of page tables to map virtual addresses to physical addresses within the system memory 3198. Further, the IO device 3190, the CPU 3199, and the GPU 3105 may share the same virtual address space.
[0254] The IOMMU 3170 also supports virtualization. In this case, the IOMMU 3170 may manage a first set of page tables for mapping guest / graphics virtual addresses to guest / graphics physical addresses (e.g., within the system memory 3198), and a second set of page tables for mapping guest / graphics physical addresses to system / host physical addresses. The base address of each of the first and second sets of page tables may be stored in a control register and swapped during a context switch (e.g., thereby providing access to the set of page tables associated with the new context). Although not shown in FIG. 31, each of the cores 3130, 3140, 3145, and / or the multicore groups 3100A - 3100N may include a translation lookaside buffer (TLB) for caching guest virtual to guest physical, guest physical to host physical, and guest virtual to host physical translations.
[0255] The CPU 3199, GPU 3105, and IO device 3190 can be integrated on a single semiconductor chip and / or chip package. The illustrated memory 3198 may be integrated on the same chip or coupled to the memory controller 3170 via an off-chip interface. In one implementation, the memory 3198 includes GDDR6 memory that shares the same virtual address space as other physical system-level memories, but the underlying principles of the present invention are not limited to this particular implementation.
[0256] The tensor core 3140 includes a plurality of execution units specifically designed to perform matrix operations, which are the basic computational operations used to execute deep learning operations. For example, simultaneous matrix multiplication operations may be used for neural network training and inference. The tensor core 3140 may perform matrix processing using various operand precisions including single-precision floating point (e.g., 32 bits), half-precision floating point (e.g., 16 bits), integer word (16 bits), byte (8 bits), and nibble (4 bits). Implementations of neural networks may also potentially combine details from multiple frames to build a high-quality final image and extract features of each rendered scene.
[0257] In deep learning implementations, parallel matrix multiplication operations may be scheduled to execute on the tensor core 3140. Neural network training, in particular, requires a significant number of matrix dot product operations. To process the inner product formula of an N×N×N matrix multiplication, the tensor core 3140 may include at least N dot product processing elements. Before matrix multiplication begins, one entire matrix is loaded into the tile register, and at least one column of the second matrix is loaded every N cycles. There are N dot products to be processed every cycle.
[0258] The matrix elements may be stored with different precisions depending on the specific implementation, including 16-bit words, 8-bit bytes (e.g., INT8), and 4-bit nibbles (e.g., INT4). Different precision modes may be specified for Tensor Core 3140 to ensure that the most efficient precision is used for different workloads (e.g., inference workloads that can tolerate quantization to bytes and nibbles, etc.).
[0259] Ray Tracing Core 3150 may be used to accelerate ray tracing operations for both real-time and non-real-time ray tracing implementations. In particular, Ray Tracing Core 3150 may include a ray traversal / crossing circuit that uses a bounding volume hierarchy (BVH) to perform ray traversal and identify intersections between rays and primitives enclosed within the BVH volume. Ray Tracing Core 3150 may also include circuitry for performing depth testing and culling (e.g., using a Z-buffer or similar configuration). In one embodiment, Ray Tracing Core 3150 performs traversal and crossing operations in cooperation with the image noise removal techniques described herein, at least some of which may be executed on Tensor Core 3140. For example, Tensor Core 3140 may implement a deep learning neural network to perform noise removal on frames generated by Ray Tracing Core 3150. However, CPU 3199, Graphics Core 3130, and / or Ray Tracing Core 3150 may also implement all or part of the noise removal and / or deep learning algorithms.
[0260] Furthermore, as described above, a variance technique for noise removal may be used, and the GPU 3150 is within a computing device coupled to other computing devices on a network or high-speed interconnect. The interconnected computing devices may further share neural network learning / training data to improve the rate at which the overall system learns to perform noise removal for different types of image frames and / or different graphics applications.
[0261] The ray tracing core 3150 may handle all BVH traversals and ray-primitive intersections, preventing the graphics core 3130 from being overloaded with thousands of instructions per ray. Each ray tracing core 3150 may include a first set of special circuits for performing a bounding box test (e.g., for traversal operations) and a second set of special circuits for performing a ray-triangle intersection test (e.g., for the traversed intersecting ray). Thus, the multi-core group 3100A can simply initiate a ray probe, and the ray tracing core 3150 independently performs ray traversal and intersection and returns hit data (e.g., hit, no hit, multiple hits, etc.) to the thread context. Other cores 3130, 3140 may be freed up to perform other graphics or perform work calculations while the ray tracing core 3150 performs traversal and intersection operations.
[0262] Each ray tracing core 3150 may include a traversal unit for performing BVH test operations and an intersection unit for performing ray-primitive intersection tests. The intersection unit may then generate a "hit", "no hit", or "multi-hit" response and provide it to the appropriate thread. During traversal and intersection operations, the execution resources of other cores (e.g., the graphics core 3130 and the tensor core 3140) may be freed up to perform other forms of graphics work.
[0263] A hybrid rasterization / ray tracing technique in which operations are distributed between a graphics core 3130 and a ray tracing core 3150 may also be used.
[0264] The ray tracing core 3150 (and / or other cores 3130, 3140) may include a ray tracing instruction set such as Microsoft's DXR (DirectX Ray Tracing) including the DispatchRays command, as well as hardware support for ray generation shaders, closest-hit shaders, any-hit shaders, and miss shaders, enabling the assignment of a unique set of shaders and textures for each object. Another ray tracing platform that may be supported by the ray tracing core 3150, the graphics core 3130, and the tensor core 3140 is Vulkan 1.1.85. However, it should be noted that the underlying principles of the present invention are not limited to a particular ray tracing ISA.
[0265] Generally, the various cores 3150, 3140, 3130 may support a ray tracing instruction set including instructions / functions for ray generation, closest-hit, any-hit, ray-primitive intersection, per-primitive and hierarchical bounding box construction, miss, visit, and exception. More specifically, ray tracing instructions for performing the following functions may be included.
[0266] Ray Generation - Ray generation instructions may be executed for each pixel, sample, or other user-defined work assignment.
[0267] Closest-Hit - Closest-hit instructions may be executed to find the closest intersection of a ray with a primitive in the scene.
[0268] Any-Hit - The Any-Hit instruction identifies multiple intersections between a ray and primitives within a scene in order to potentially identify the new nearest intersection.
[0269] Intersection - The Intersection instruction performs a ray-primitive intersection test and outputs the result.
[0270] Per-primitive Bounding box Construction - This instruction constructs a bounding box around a given primitive or group of primitives (e.g., when constructing a new BVH or other acceleration data structure).
[0271] Miss - Indicates that a ray misses all geometry within a scene or a specified region of the scene.
[0272] Visit - Indicates a child volume through which a ray passes.
[0273] Exceptions - Includes various types of exception handlers (e.g., called under various error conditions).
[0274] [Hierarchical Beam Tracing] Bounding volume hierarchies are commonly used to improve the efficiency with which operations are performed on graphics primitives and other graphics objects. A BVH is a hierarchical tree structure built based on a set of geometric objects. The top of the tree structure is a root node that encloses all of the geometric objects within a given scene. Individual geometric objects are enclosed within bounding volumes that form the leaf nodes of the tree. These nodes are then grouped as a small set and enclosed within a larger bounding volume. Next, these are also grouped recursively and enclosed within other larger bounding volumes, ultimately resulting in a tree structure with a single bounding volume represented by the root node at the top of the tree. Bounding volume hierarchies are used to efficiently support various operations on a set of geometric objects, such as collision detection, primitive culling, and ray traversal / intersection operations used in ray tracing.
[0275] In a ray tracing architecture, rays are traversed through the BVH to determine ray-primitive intersections. For example, if a ray does not pass through the root node of the BVH, the ray will not intersect any of the primitives enclosed by the BVH and no further processing for the ray is required with respect to this set of primitives. If a ray passes through the first child node of the BVH but not the second child node, the ray need not be tested against the primitives enclosed by the second child node. Thus, the BVH provides an efficient mechanism for testing ray-primitive intersections.
[0276] A group of continuous light rays called a "beam" may be tested against a BVH, rather than individual light rays. FIG. 32 shows an exemplary beam 3201 schematically shown by four different light rays. Any light ray that intersects a patch 3200 defined by four light rays is considered to be within the same beam. The beam 3201 in FIG. 32 is defined by a rectangular arrangement of light rays, but the beam may be defined in various other ways while still following the underlying principles of the present invention (e.g., circles, ellipses, etc.).
[0277] FIG. 33 shows how the ray tracing engine 3310 of the GPU 3320 implements the beam tracing technology described herein. In particular, the ray generation circuit 3304 generates a plurality of light rays on which traversal and intersection operations are performed. However, instead of performing traversal and intersection operations on individual light rays, the traversal and intersection operations are performed using a hierarchy of beams 3307 generated by the beam hierarchy construction circuit 3305. The beam hierarchy is similar to a bounding volume hierarchy (BVH). For example, FIG. 34 provides an example of a primary beam 3400 that may be subdivided into a plurality of different components. In particular, the primary beam 3400 may be divided into quadrants 3401 to 3404, and each quadrant itself may be divided into sub-quadrants such as sub-quadrants A to D within quadrant 3404. The primary beam may be subdivided in various ways. For example, the primary beam may be divided (not into quadrants) into halves, and each half may be divided into halves, and so on. Regardless of how the subdivision is performed, the hierarchical structure is generated in the same way as the BVH, for example, a root node representing the primary beam 3400, child nodes at the first level (each represented by quadrants 3401 to 3404), child nodes at the second level for each of the sub-quadrants A to D, etc.
[0278] Once the beam hierarchy 3307 is constructed, the traversal / intersection circuit 3306 may perform traversal / intersection operations using the beam hierarchy 3307 and the BVH 3308. In particular, beams may be tested against the BVH, and portions of the beams that do not intersect any part of the BVH may be culled. Using the data shown in FIG. 34, for example, if sub-beams associated with sub-regions 3402 and 3403 do not intersect the BVH or a particular branch of the BVH, these may be culled with respect to the BVH or the branch. The remaining portions 3401, 3404 may be tested against the BVH by performing a depth-first search or other search algorithm.
[0279] A method for ray tracing is shown in FIG. 35. The method may be implemented within the context of the graphics processing architecture described above, but is not limited to any particular architecture.
[0280] At 3500, a primary beam including a plurality of rays is constructed, and at 3501, the beam is subdivided and a hierarchical data structure is generated to create a beam hierarchy. Operations 3500 - 3501 may be performed as a single integrated operation to construct a beam hierarchy from a plurality of rays. At 3502, the beam hierarchy is used with the BVH to cull rays and / or nodes / primitives from the BVH. At 3503, ray-primitive intersections are determined for the remaining rays and primitives.
[0281] [Irreversible and Reversible Packet Compression in a Distributed Ray Tracing System] The ray tracing operation may be distributed among a plurality of computing nodes coupled together over a network. For example, FIG. 36 shows a ray tracing cluster 3600 including a plurality of ray tracing nodes 3610-3613 executing ray tracing operations in parallel and potentially combining the results at one of the nodes. In the illustrated architecture, the ray tracing nodes 3610-3613 are communicatively coupled to a client-side ray tracing application 3630 via a gateway.
[0282] One of the challenges of the distributed architecture is the large amount of packetized data that must be transmitted between each of the ray tracing nodes 3610-3613. Both reversible compression techniques and irreversible compression techniques may be used to reduce the data transmitted between the ray tracing nodes 3610-3613.
[0283] To implement reversible compression, instead of sending packets filled with the results of a particular type of operation, data or commands are sent that enable the receiving node to reconstruct the results. For example, probabilistically sampled area lights and ambient occlusion (AO) operations do not necessarily require a direction. Thus, the sending node can simply send a random seed, which is then used by the receiving node to perform random sampling. For example, if the scene is distributed among nodes 3610 - 3612, only the light ID and the origin need to be sent to nodes 3610 - 3612 to sample light 1 at points p1 - p3. Then each node may sample the light independently and probabilistically. The random seed may be generated by the receiving node. Similarly, for the primary ray hit point, ambient occlusion (AO) and soft shadow sampling can be computed on nodes 3610 - 3612 without waiting for the original points for consecutive frames. Further, if it is known that a set of rays goes to the same point light source, a command to identify the light source may be sent to the receiving node, and the receiving node applies it to the set of rays. As another example, if there are N ambient occlusion rays passing through a single point, a command to generate N samples from this point may be sent.
[0284] Various further techniques may be applied for irreversible compression. For example, quantization coefficients may be used to quantize all coordinate values related to BVHs, primitives, and rays. Further, 32-bit floating-point values used for data such as BVH nodes and primitives may be converted to 8-bit integer values. In an exemplary implementation, the bounds of a ray packet are stored with full precision, but the individual ray points P1 - P3 are transmitted as an offset with an index to the bounds. Similarly, multiple local coordinate systems may be generated that use 8-bit integer values as local coordinates. The position of the starting point of each of these local coordinate systems may be encoded using full precision (e.g., 32-bit floating-point) values to effectively connect the global and local coordinate systems.
[0285] The following are examples of reversible compression. An example of a ray data format used internally in a ray tracing program is as follows. struct Ray { uint32 pixId; uint32 materialID; uint32 instanceID; uint64 primitiveID; uint32 geometryID; uint32 lightID; float origin[3]; float direction[3]; float t0; float t; float time; float normal[3]; / / used for geometry intersections float u; float v; float wavelength; float phase; / / Interferometry float refractedOffset; / / Schlieren-esque float amplitude; float weight; }; Instead of transmitting the raw data for each generated node, this data can be compressed by grouping values and creating implicit rays using applicable metadata where possible.
[0286] (Ray bundling and grouping of ray data) The flag may be used for common data or masks with modifiers. struct RayPacket { uint32 size; uint32 flags; list <ray>rays; } For example, RayPacket.rays = ray_1 to ray_256 (The starting points are all shared) All ray data is packed except that only a single starting point is stored among all rays. RayPacket.flags is set for RAYPACKET_COMMON_ORIGIN. When RayPacket is unpacked upon reception, the starting point is embedded from a single starting point value.
[0287] (The starting points are shared among only some rays) All ray data is packed except for the rays that share a starting point. For each group's unique shared starting point, an operator that identifies the operation (shared starting point), stores the starting point, and masks which rays share the information is packed. Such an operation can be performed on any shared value among nodes such as material ID, primitive ID, starting point, direction, normal, etc. struct RayOperation { uint8 operationID; void* value; uint64 mask; } (Implicit ray transmission) Often, ray data can be derived at the receiving end by the minimum meta - information used to generate it. A very common example is generating multiple secondary rays to probabilistically sample a region. Instead of the transmitting side generating the secondary rays and sending them for the receiving side to act on, the transmitting side can send a command that the rays need to be generated with any dependent information, and the rays are generated at the receiving end. If the rays are first generated by the transmitting side and it is necessary to determine which receiving side to send them to, the rays are generated and a random seed can be sent so that the exact same rays can be regenerated.
[0288] For example, to sample hit points with 64 shadow rays that sample an area light source, all 64 rays intersect the same region from calculation N4. A RayPacket with a common starting point and normal is created. If the receiving side desires to shade the contribution of the resulting pixel, more data may be sent, but in this example, it is assumed that only whether the ray hits other node data is desired to be returned. A RayOperation is created for the generate shadow ray operation, and a value of the sampled light ID and a random number seed are assigned. When N4 receives a ray packet, to generate the same rays as originally generated by the sending side, shared starting point data is embedded in all rays and the direction is set based on the ray ID sampled probabilistically with the random number seed, generating fully embedded ray data. When the result is returned, only the binary result for all rays needs to be returned, which can be handled by a mask for the rays.
[0289] In this example, transmitting the original 64 rays would have used 104 bytes × 64 rays = 6656 bytes. If the returning rays were also transmitted in raw format, this would also be 2 × 13312 bytes. Using reversible compression by transmitting only the common ray origin, normal, and the ray generation operations at the seed and ID results in only 29 bytes being transmitted, along with 8 bytes for the crossed mask. This results in a data compression rate of ~360:1 that needs to be transmitted over the network. This does not include the overhead for processing the message itself, and the message needs to be identified somehow, which depends on the implementation. Other operations may be done to recompute the ray origin and direction from the pixel ID for primary rays, to recompute the pixel ID based on the range within the ray packet, and for many other possible implementations for value recomputation. Similar operations can be used for any single ray or group of rays being transmitted, including shadow, reflection, refraction, ambient occlusion, intersection, volume intersection, shadowing, bounced reflection, etc. in a path trace.
[0290] FIG. 37 shows further details regarding two ray tracing nodes 3710-3711 that perform compression and decompression of ray tracing packets. In particular, when the first ray tracing engine 3730 is ready to send data to the second ray tracing engine 3731, the ray compression circuit 3720 performs irreversible and / or reversible compression of the ray tracing data as described herein (e.g., converting a 32-bit value to an 8-bit value, using raw data instead of commands to reconstruct the data, etc.). The compressed ray packet 3701 is sent from network interface 3725 to network interface 3726 over a local network (e.g., a 10 Gb / s, 100 Gb / s Ethernet network). Then, if necessary, the ray decompression circuit decompresses the ray packet. For example, this may execute commands to reconstruct the ray tracing data (e.g., using a random seed to perform random sampling for lighting calculations). The ray tracing engine 3731 then uses the received data to perform ray tracing operations.
[0291] In the reverse direction, the ray compression circuit 3741 compresses the ray data, the network interface 3726 sends the compressed ray data over the network (e.g., using the techniques described herein), the ray decompression circuit 3740 decompresses the ray data if necessary, and the ray tracing engine 3730 uses the data in ray tracing operations. Although shown as separate units in FIG. 37, the ray decompression circuits 3740-3741 may be integrated within the ray tracing engines 3730-3731, respectively. For example, as long as the compressed ray data includes commands to reconstruct the ray data, these commands may be executed by the respective ray tracing engines 3730-3731.
[0292] As shown in FIG. 38, the ray compression circuit 3720 may include an irreversible compression circuit 3801 for performing the irreversible compression techniques described herein (e.g., converting 32-bit floating-point coordinates to 8-bit integer coordinates), and a reversible compression circuit 3803 for performing reversible compression techniques (e.g., transmitting commands and data that enable the ray recompression circuit 3821 to reconstruct the data). The ray decompression circuit 3721 includes an irreversible decompression circuit 3802 and a reversible decompression circuit 3804 for performing reversible decompression.
[0293] Other exemplary methods are shown in FIG. 39. The method may be implemented on a ray tracing architecture or other architectures described herein, but is not limited to any particular architecture.
[0294] At 3900, ray data transmitted from a first ray tracing node to a second ray tracing node is received. At 3901, an irreversible compression circuit performs irreversible compression on the first ray tracing data, and at 3902, a reversible compression circuit performs reversible compression on the second ray tracing data. At 3903, the compressed ray tracing data is transmitted to the second ray tracing node. At 3904, an irreversible / reversible decompression circuit performs irreversible / reversible decompression of the ray tracing data, and at 3905, the second ray tracing node performs ray tracing operations using the decompressed data.
[0295] [Graphics Processor Using Hardware-Accelerated Hybrid Ray Tracing] Next, a hybrid rendering pipeline is presented that performs rasterization on the graphics core 3130 and ray tracing operations on the ray tracing core 3150, the graphics core 3130, and / or the CPU 3199 cores. For example, the rasterizer and depth test may be performed on the graphics core 3130 instead of the primary ray casting stage. Next, the ray tracing core 3150 may generate secondary rays for ray reflection, refraction, and shadows. Further, a specific region of the scene where the ray tracing core 3150 performs ray tracing operations (e.g., based on a material property threshold such as a high reflectivity level) is selected, and other regions of the scene are rendered with rasterization on the graphics core 3130. This hybrid implementation may be used for real-time ray tracing applications where latency is a critical issue.
[0296] The ray traversal architecture described below may use an existing single instruction multiple data (SIMD) and / or single instruction multiple thread (SIMT) graphics processor to perform programmable shading and control of ray traversal while accelerating important functions such as BVH traversal and / or intersection using, for example, dedicated hardware. The SIMD occupancy of the incoherent path can be improved by regrouping shaders generated at a specific point during traversal and before shading. This is achieved using dedicated hardware that dynamically sorts shaders on-chip. Recursion is managed by splitting the function into continuations that are executed when returning and regrouping the continuations before execution for improved SIMD occupancy.
[0297] Ray traversal / crossing programmable control is achieved by decomposing the traversal function into an internal traversal that can be implemented as fixed-function hardware and an external traversal that runs on the GPU processor and enables programmable control through user-defined traversal shaders. The cost of transferring the traversal context between hardware and software is reduced by conservatively discarding the internal traversal state during the transition between the internal and external traversals.
[0298] Ray tracing programmable control can be expressed through different shader types listed in Table A below. Multiple shaders can exist for each type. For example, each material can have a different hit shader.
[0299] [Table 1] Recursive ray tracing may be initiated by an API function that instructs the graphics processor to start a set of primary shaders or intersection circuits that can generate ray-scene intersections for primary rays. This then generates other shaders such as traversal shaders, hit shaders, or miss shaders. A shader that generates a child shader can also receive a return value from that child shader. A callable shader is a general-purpose function that can be directly generated by other shaders and can also return a value to the calling shader.
[0300] FIG. 40 shows a graphics processing architecture including a shader execution circuit 4000 and a fixed function circuit 4010. The general-purpose execution hardware subsystem includes a plurality of single instruction multiple data (SIMD) and / or single instruction multiple threads (SIMT) cores / execution units (EUs) 4001 (i.e., each core may include a plurality of execution units), one or more samplers 4002, and a level 1 (L1) cache 4003 or other form of local memory. The fixed function hardware subsystem 4010 includes a message unit 4004, a scheduler 4007, a ray BVH traversal / intersection circuit 4005, a sort circuit 4008, and a local L1 cache 4006.
[0301] During operation, the primary dispatcher 4009 dispatches a set of primary rays to the scheduler 4007, and the scheduler 4007 schedules the work to be executed on the SIMD / SIMT cores / EUs 4001 by the shaders. The SIMD cores / EUs 4001 may be the ray tracing cores 3150 and / or the graphics cores 3130 described above. The execution of the primary shader generates further work to be executed (e.g., further work to be executed by one or more child shaders and / or fixed function hardware). The message unit 4004 distributes the work generated by the SIMD cores / EUs 4001 to the scheduler 4007 and accesses the free stack pool, the sorting circuit 4008, or the ray-BVH intersection circuit 4005 as needed. If further work is sent to the scheduler 4007, this is scheduled for processing on the SIMD / SIMT cores / EUs 4001. Before scheduling, the sorting circuit 4008 may sort the rays into groups or bins as described herein (e.g., grouping rays with similar characteristics). The ray-BVH intersection circuit 4005 performs ray intersection tests using the BVH volume. For example, the ray-BVH intersection circuit 4005 may compare the ray coordinates with each level of the BVH to identify the volume with which the ray intersects.
[0302] The shader can be referenced using a shader record, a user-assigned structure containing a pointer to the entry function, vendor-specific metadata, and global arguments to the shader to be executed by the SIMD cores / EUs 4001. Each execution instance of the shader is associated with a call stack that may be used to store arguments passed between the parent shader and the child shaders. The call stack may also store references to the continuation functions to be executed upon call return.
[0303] FIG. 41 shows an exemplary set of allocated stacks 4101, including a primary shader stack, a hit shader stack, a traversal shader stack, a continuation function stack, and a ray - BVH intersection stack (which may be executed by fixed - function hardware 4010 as described). A new shader call may implement a new stack from the free stack pool 4102. A call stack, e.g., a stack composed of a set of allocated stacks, may be cached in local L1 caches 4003, 4006 to reduce access latency.
[0304] There may be a limited number of call stacks, each having a fixed maximum size of "Sstack" allocated to a contiguous region of memory. Thus, the base address of the stack can be directly calculated as base address = SID * Sstack from the stack index (SID). The stack ID may be assigned and de - assigned by the scheduler 4007 when scheduling work to the SIMD core / EU 4001.
[0305] The primary dispatcher 4009 may include a graphics processor command processor that dispatches primary shaders in response to dispatch commands from a host (e.g., a CPU). The scheduler 4007 may receive these dispatch requests and start the primary shader on the SIMD processor thread if it can assign a stack ID for each SIMD lane. The stack ID may be assigned from the free stack pool 4102 initialized at the start of the dispatch command.
[0306] A running shader can generate a child shader by sending a generation message to message unit 4004. This command includes the stack ID associated with the shader and also a pointer to the child shader record for each active SIMD lane. The parent shader can issue this message only once for each active lane. After sending the generation message for all relevant lanes, the parent shader may terminate.
[0307] Shaders running on SIMD core / EU 4001 can also use the generation message with a shader record pointer reserved for fixed function hardware to generate fixed function tasks such as ray-BVH intersections. As described above, message unit 4004 sends the generated ray-BVH intersection work to fixed function ray-BVH intersection circuit 4005 and sends callable shaders directly to sort circuit 4008. The sort circuit may group shaders by shader record pointer to derive SIMD batches with similar characteristics. Thus, stack IDs from different parent shaders can be grouped within the same batch by sort circuit 4008. Sort circuit 4008 sends the grouped batch to scheduler 4007, and scheduler 4007 accesses the shader records from graphics memory 2511 or last level cache (LLC) 4020 and starts the shaders on the processor thread.
[0308] Continuations may be treated as callable shaders and may also be referenced through shader records. When a child shader is generated and returns a value to the parent shader, a pointer to the continuation shader record may be pushed onto call stack 4101. When the child shader returns, the continuation shader record is popped from call stack 4101 and a continuation shader may be generated. Optionally, the generated continuation may pass through the sort unit in the same manner as callable shaders and may be started on the processor thread.
[0309] As shown in FIG. 42, the sorting circuit 4008 groups tasks generated by shader record pointers 4201A, 4201B, 4201n to create SIMD batches for shading. The stack ID or context ID within the sorted batch can be grouped from different dispatches and different input SIMD lanes. The grouping circuit 4210 may perform sorting using a content addressable memory (CAM) structure 4201 that includes a plurality of entries each having an entry identified by a tag 4201. As described above, the tag 4201 may be the corresponding shader record pointer 4201A, 4201B, 4201n. The CAM structure 4201 may store a limited number of tags (e.g., 32, 64, 128, etc.) respectively associated with incomplete SIMD batches corresponding to the shader record pointers.
[0310] For incoming generation commands, each SIMD lane has a corresponding stack ID (shown as 16 context IDs 0 - 15 in each CAM entry) and shader record pointers 4201A - B,...n (functioning as tag values). The grouping circuit 4210 may compare the shader record pointer of each lane with the tag 4201 in the CAM structure 4201 to find a matching batch. If a matching batch is found, the stack ID / context ID may be added to the batch. Otherwise, a new entry with a new shader record pointer tag may be created, and in some cases, the old entry of the incomplete batch may be removed.
[0311] The running shader can deallocate the call stack when empty by sending a deallocation message to the message unit. The deallocation message is relayed to the scheduler, and the scheduler returns the stack ID / context ID of the active SIMD lanes to the free pool.
[0312] A hybrid approach for ray traversal operations using a combination of fixed-function ray traversal and software ray traversal is presented. Thus, this maintains the efficiency of fixed-function traversal while providing the flexibility of software traversal. FIG. 43 shows an acceleration structure that can be used for hybrid traversal, which is a two-level tree having a single top-level BVH 4300 and several lower-level BVHs 4301 and 4302. Graphical elements for showing an internal traversal path 4303, an external traversal path 4304, a traversal node 4305, a leaf node 4306 having triangles, and a leaf node 4307 having custom primitives are shown on the right side.
[0313] Leaf nodes 4306 having triangles in the top-level BVH 4300 can reference triangles, intersection shader records for custom primitives, or traversal shader records. Leaf nodes 4306 having triangles in the lower-level BVHs 4301 - 4302 can reference only triangles and intersection shader records for custom primitives. The type of reference is encoded within the leaf node 4306. The internal traversal 4303 indicates traversal within each BVH 4300 - 4302. The internal traversal operations include calculation of ray-BVH intersections, and the traversal between the BVH structures 4300 - 4302 is known as the external traversal. The internal traversal operations can be efficiently implemented in fixed-function hardware, while the external traversal operations can be executed with acceptable performance by a programmable shader. Thus, the internal traversal operations may be executed using a fixed-function circuit 4010, and the external traversal operations may be executed using a shader execution circuit 4000 including SIMD / SIMT cores / EUs 4001 for executing programmable shaders.
[0314] The SIMD / SIMT core / EU 4001 should be noted that, for the sake of simplicity in some cases, it is simply referred to here as "core", "SIMD core", "EU" or "SIMD processor". Similarly, the ray - BVH traversal / intersection circuit 4005 is simply referred to as "traversal unit", "traversal / intersection unit" or "traversal / intersection circuit" in some cases. When alternative terms are used, the specific names used to designate each circuit / logic do not change the underlying function that the circuit / logic performs, as described here.
[0315] Furthermore, for the purpose of explanation, although shown as a single component in FIG. 40, the traversal / intersection unit 4005 may include a separate traversal unit and a separate intersection unit, each of which may be implemented in circuitry and / or logic as described here.
[0316] When a ray intersects a traversal node during an internal traversal, a traversal shader may be generated. The sort circuit 4008 may group these shaders into shader record pointers 4201A - 4201B, n to create SIMD batches, and the SIMD batches are initiated by the scheduler 4007 for SIMD execution on the graphics SIMD core / EU 4001. The traversal shader can change the traversal in several ways, enabling a wide range of applications. For example, the traversal shader can select a BVH at a coarser level of detail (LOD), or transform the ray to enable rigid body transformation. Then, the traversal shader may generate an internal traversal for the selected BVH.
[0317] The internal traversal calculates ray-BVH intersections by traversing the BVH to compute ray-box and ray-triangle intersections. The internal traversal is generated in the same way as the shader by sending messages to the message circuit 4004, and the message circuit 4004 relays the corresponding generated messages to the ray-BVH intersection circuit 4005 that computes ray-BVH intersections.
[0318] The stack for the internal traversal may be stored locally within the fixed function circuit 4010 (e.g., within the L1 cache 4006). When the ray intersects a leaf node corresponding to a traversal shader or intersection shader, the internal traversal ends and the internal stack may be discarded. The discarded stack may be written to memory at a location specified by the calling shader, along with pointers to the ray and BVH, and then the corresponding traversal shader or intersection shader may be generated. If the ray intersects any triangle during the internal traversal, the corresponding hit information may be provided as input arguments to these shaders, as shown in the following code. These generated shaders may be grouped by the sort circuit 4008 to create a SIMD batch for execution. struct HitInfo { float barycentrics[2]; float tmax; bool innerTravComplete; uint primID; uint geomID; ShaderRecord* leafShaderRecord; } Discarding the internal traversal stack reduces the cost of spilling it to memory. This technique is described in Restart Trail for Stackless BVH Traversal, High Performance Graphics (2010), pp. 107-111, where a 42-bit restart trail and a 6-bit depth value may be applied to discard the stack to a few entries at the top of the stack. The restart trail indicates the branches already incorporated within the BVH, and the depth value indicates the depth of the traversal corresponding to the last stack entry. This is sufficient information to later resume the internal traversal.
[0319] The internal traversal completes when the internal stack is empty and there are no BVH nodes to test. In this case, if the external stack is not empty, an external stack handler is generated that pops the top of the external stack and resumes the traversal.
[0320] The external traversal may execute the main traversal state machine or may be implemented in program code executed by the shader execution circuit 4000. Under the following conditions, namely, (1) when a new ray is generated by the hit shader or the primary shader, (2) when the traversal shader selects a BVH for traversal, and (3) when the external stack handler resumes an internal traversal for a BVH, this may generate an internal traversal query.
[0321] As shown in FIG. 44, space is allocated on the call stack 4405 for the fixed function circuit 4010 to store the discarded internal stack 4410 before the internal traversal is generated. The offsets 4403-4404 to the top of the call stack and the internal stack are maintained in the traversal state 4400, which is also stored in the memory 2511. The traversal state 4400 also includes the ray in the world space 4401 and the object space 4402 and the hit information for the closest intersection primitive.
[0322] The traversal shader, intersection shader, and external stack handler are all generated by the ray-BVH intersection circuit 4005. The traversal shader is allocated on the call stack 4405 before starting a new internal traversal for the second level BVH. The external stack handler is the shader responsible for updating hit information and restarting any pending internal traversal tasks. The external stack handler is also responsible for generating a hit or miss shader when the traversal is complete. When there are no pending internal traversal queries to be generated, the traversal is complete. When the traversal is complete and an intersection is found, a hit shader is generated. Otherwise, a miss shader is generated.
[0323] The hybrid traversal method described above uses a two-level BVH hierarchy, but any number of BVH levels may also be implemented with corresponding changes in the implementation of the outer traversal.
[0324] Furthermore, while the fixed function circuit 4010 for performing ray-BVH intersection is described above, other system components may also be implemented within the fixed function circuit. For example, the outer stack handler described above may be an internal (invisible to the user) shader that can potentially be implemented within the fixed function BVH traversal / intersection circuit 4005. This implementation may be used to reduce the number of dispatched shader stages and round trips between the fixed function intersection hardware 4005 and the processor.
[0325] The examples described herein enable programmable shading and ray traversal control using user-defined functions that can execute with greater SIMD efficiency on existing and future GPU processors. Programmable control of ray traversal enables several important features such as procedural instancing, probabilistic level-of-detail selection, custom primitive intersection, and lazy BVH update.
[0326] A programmable multiple instruction multiple data (MIMD) ray tracing architecture that supports speculative execution of hit and intersection shaders is also provided. In particular, the architecture focuses on reducing the scheduling and communication overhead between the programmable SIMD / SIMT core / execution unit 4001 and the fixed function MIMD traversal / intersection unit 4005 as described above with respect to FIG. 40 in a hybrid ray tracing architecture. Multiple speculative execution modes of the hit and intersection shaders that can be dispatched in a single batch from the traversal hardware, avoiding some traversal and shading round trips, are described below. Dedicated circuitry for implementing these techniques may be used.
[0327] Embodiments of the present invention are particularly beneficial in use cases where execution of multiple hits or intersection shaders is desired from a ray traversal query, which imposes a significant overhead when implemented without dedicated hardware support. These include, but are not limited to, nearest k-hit queries (starting hit shaders for the k nearest intersections) and multiple programmable intersection shaders.
[0328] The techniques described herein may be implemented as an extension to the architecture shown in FIG. 40 (and described with respect to FIGS. 40-44). In particular, the present embodiments of the invention are built on top of this architecture by extensions to improve the performance of the above use cases.
[0329] The performance limitations of the ray traversal architecture for hybrid ray tracing are the overhead of starting a traversal query from an execution unit and the overhead of calling a programmable shader from ray tracing hardware. When multiple hit or intersection shaders are called during the traversal of the same ray, this overhead creates an "execution roundtrip" between the programmable core 4001 and the traversal / intersection unit 4005. This also places additional pressure on the sort unit 4008 that needs to extract SIMD / SIMT coherence from individual shader calls.
[0330] Some aspects of ray tracing require programmable control that can be expressed through the different shader types (i.e., primary, hit, any hit, miss, intersection, traversal, callable) listed in Table A above. There can be multiple shaders for each type. For example, each material can have a different hit shader. Some of these shader types are defined in the current Microsoft Ray Tracing API.
[0331] As a simple consideration, recursive ray tracing is initiated by an API function that instructs the GPU to start a set of primary shaders that can generate ray-scene intersections (implemented in hardware and / or software) for primary rays. This then generates other shaders such as traversal shaders, hit shaders, or miss shaders. A shader that generates a child shader can also receive a return value from that child shader. A callable shader is a general-purpose function that can be directly generated by other shaders and can also return a value to the calling shader.
[0332] A ray traversal calculates ray-scene intersections by traversing and intersecting nodes within a bounding volume hierarchy (BVH). Recent research has shown that the efficiency of ray-scene intersection calculations can be improved by more than an order of magnitude by using techniques more suitable for fixed-function hardware such as reduced-precision arithmetic, BVH compression, per-ray state machines, dedicated intersection pipelines, and custom caches.
[0333] The architecture shown in FIG. 40 includes a system in which an array of SIMD / SIMT cores / execution units 4001 interacts with a fixed-function ray tracing / intersection unit 4005 to perform programmable ray tracing. The programmable shaders are mapped to SIMD / SIMT threads on the execution units / cores 4001, and SIMD / SIMT utilization, execution, and data coherence are important for optimal performance. Ray queries often break coherence for a variety of reasons, as follows. · Traversal differences: The duration of BVH traversal varies widely between rays, which is advantageous for asynchronous ray processing. · Execution differences: Rays generated from different lanes of the same SIMD / SIMT thread can result in different shader calls. · Data access differences: For example, rays hitting different surfaces sample different BVH nodes, and primitives and shaders access different textures. Various other scenarios can cause data access differences.
[0334] The SIMD / SIMT core / execution unit 4001 may be a variation of the core / execution units described herein, including graphics cores 415A - 415B, shader cores 1355A - N, graphics core 3130, graphics execution unit 608, execution units 852A - 852B, or any other core / execution unit described herein. The SIMD / SIMT core / execution unit 4001 may be used in place of graphics cores 415A - 415B, shader cores 1355A - 1355N, graphics core 3130, graphics execution unit 608, execution units 852A - 852B, or any other core / execution unit described herein. Accordingly, the disclosure of any feature in combination with graphics cores 415A - 415B, shader cores 1355A - 1355N, graphics core 3130, graphics execution unit 608, execution units 852A - 852B, or any other core / execution unit described herein also discloses the corresponding combination with the SIMD / SIMT core / execution unit 4001 of FIG. 40, but is not limited thereto.
[0335] The fixed - function ray tracing / intersection unit 4005 can overcome the first two problems by processing each ray individually and out - of - order. However, this splits the SIMD / SIMT group. Accordingly, the sort unit 4008 is responsible for forming new coherent SIMD / SIMT groups of shader calls that are redispensed to the execution units.
[0336] It is easy to recognize the advantages of such an architecture compared to a direct pure - software - based ray tracing implementation on a SIMD / SIMT processor. However, there is overhead associated with message passing between the SIMD / SIMT core / execution unit 4001 (herein, sometimes simply referred to as the SIMD / SIMT processor or core / EU) and the MIMD traversal / intersection unit 4005. Further, the sort unit 4008 may not extract full SIMD / SIMT utilization from non - coherent shader calls.
[0337] Use cases in which shader calls can occur particularly frequently during traversal can be identified. An extension for a hybrid MIMD ray tracing processor is described to significantly reduce the overhead of communication between the core / EU 4001 and the traversal / intersection unit 4005. This can be particularly advantageous when finding the k closest intersections and the implementation of a programmable intersection shader. However, it should be noted that the techniques described herein are not limited to specific processing scenarios.
[0338] A summary of the high-level cost of rate-tracing context switching between the core / EU 4001 and the fixed-function traversal / intersection unit 4005 is provided below. Most of the performance overhead is caused by these two context switches each time a shader call is required during a single rate traversal.
[0339] Each SIMD / SIMT lane that starts a ray generates a spawn message to the traversal / intersection unit 4005 associated with the traversed BVH. Data (rate traversal context) is relayed to the traversal / intersection unit 4005 via the spawn message and (cached) memory. When the traversal / intersection unit 4005 is ready to assign a new hardware thread to the spawn message, it loads the traversal state and performs a traversal on the BVH. There is also a setup cost that needs to be incurred before the first traversal step of the BVH.
[0340] Figure 45 shows the operation flow of a programmable ray tracing pipeline. The shaded elements including traversal 4502 and intersection 4503 may be implemented within a fixed-function circuit, and the remaining elements may be implemented in a programmable core / execution unit.
[0341] In 4502, the primary ray shader 4501 sends work to a traversal circuit that traverses the current ray through a BVH (or other acceleration structure). When reaching a leaf node, in 4503, the traversal circuit calls an intersection circuit, and when the intersection circuit identifies a ray-triangle intersection, in 4504, it calls an any-hit shader (which may return the result to the traversal circuit as shown in the figure).
[0342] Alternatively, the traversal may end before reaching a leaf node, and in 407, a closest-hit shader is called (if a hit is recorded), or in 4506, a miss shader is called (in the case of a miss).
[0343] As shown in 4505, when the traversal circuit reaches a custom primitive leaf node, an intersection shader may be called. The custom primitive may be any non-triangle primitive such as a polygon or polyhedron (e.g., tetrahedron, voxel, hexahedron, wedge, pyramid, or other "unstructured" volume). The intersection shader 4505 identifies any intersection between the ray and the custom primitive relative to the any-hit shader 4504 that implements any-hit processing.
[0344] When the hardware traversal 4502 reaches a programmable stage, the traversal / intersection unit 4005 may generate a shader dispatch message to the relevant shaders 4505 - 4507 corresponding to a single SIMD lane of the execution units used to execute the shaders. The dispatches occur for rays in any order and diverge in the called programs, so the sort unit 4008 may accumulate multiple dispatch calls to extract coherent SIMD batches. The updated traversal state and optional shader arguments may be written to the memory 2511 by the traversal / intersection unit 4005.
[0345] In the k-nearest intersection problem, the closest hit shader 4507 is executed for the first k intersections. In the conventional method, this means that when the closest intersection is found, the ray traversal is terminated, the hit shader is called, and a new ray is generated from the hit shader to find the next closest intersection (the same intersection does not occur again due to the offset of the ray's starting point). It is easy to recognize that this implementation requires the generation of k rays for a single ray. Other implementations operate with an any hit shader 4504 that uses an insertion sort operation to maintain a global list of the closest intersections called for all intersections. The main problem with this approach is that there is no upper limit on the any hit shader calls.
[0346] As described above, the intersection shader 4505 may be called for non-triangle (custom) primitives. Depending on the result of the intersection test and the traversal state (pending nodes and primitive intersections), the traversal of the same ray may continue after the execution of the intersection shader 4505. Therefore, finding the closest hit may require several round trips to the execution unit.
[0347] Also, through changes to the traversal hardware and shader scheduling model, it is also possible to focus on reducing SIMD-MIMD context switching for the intersection shader 4505 and the hit shaders 4504, 4507. First, the ray traversal circuit 4005 delays shader calls by accumulating multiple potential calls and dispatching them in larger batches. Additionally, certain calls that are determined to be unnecessary may be deleted at this stage. Further, the shader scheduler 4007 may aggregate multiple shader calls from the same traversal context into a single SIMD batch, which results in a single ray generation message. In one exemplary implementation, the traversal hardware 4005 interrupts traversal threads and waits for the results of multiple shader calls. This mode of operation enables the dispatch of multiple shaders and is herein referred to as "speculative" shader execution, some of which may not need to be called when using sequential calls.
[0348] FIG. 46A shows an example where traversal operations face multiple custom primitives 4650 within a subtree, and FIG. 46B shows how this can be resolved in three intersection dispatch cycles C1 - C3. In particular, the scheduler 4007 may require three cycles to submit work to the SIMD processor 4001, and the traversal circuit 4005 may require three cycles to provide the results to the sort unit 4008. The traversal state 4601 required by the traversal circuit 4005 may be stored in memory such as a local cache (e.g., L1 cache and / or L2 cache).
[0349] A. Delaying Ray Tracing Shader Calls The way in which the hardware traversal state 4601 is managed to enable the accumulation of multiple potential intersection or hit calls within the list can also be changed. At a given time during traversal, each entry within the list may be used to generate a shader call. For example, the k-nearest intersection points can be accumulated in the traversal hardware 4005 and / or the traversal state 4601 within memory, and when the traversal is complete, a hit shader can be called for each element. For the hit shader, multiple potential intersections may be accumulated for a subtree within the BVH.
[0350] In the case of the k-nearest use case, the advantage of this approach is that instead of k - 1 round trips to the SIMD core / EU 4001 and k - 1 new ray generation messages, all hit shaders are called from the same traversal thread during a single traversal operation on the traversal circuit 4005. A potential implementation challenge is that it is not straightforward to guarantee the execution order of the hit shaders (the standard "round trip" approach guarantees that the hit shader for the closest intersection is executed first, etc.). This may be addressed either by synchronization of the hit shaders or relaxation of the order.
[0351] In the case of the intersection shader use case, the traversal circuit 4005 does not pre-recognize whether a given shader returns a positive intersection test. However, it is possible to execute multiple intersection shaders speculatively, and if at least one returns a positive hit result, this is merged into the global closest hit. A particular implementation finds the optimal number of deferred intersection tests to reduce the number of dispatch calls, but needs to avoid calling too many redundant intersection shaders.
[0352] B. Aggregation of Shader Calls from the Traversal Circuit When dispatching multiple shaders from the same ray generated by traversal circuit 4005, branches within the flow of the ray traversal algorithm may be created. This can be a problem for intersection shaders. The reason for this is that the remaining BVH traversal depends on the results of all dispatched intersection tests. This means that a synchronization operation is necessary to wait for the results of the shader calls, which can be an issue on asynchronous hardware.
[0353] The two points where the results of shader calls are merged are the SIMD processor 4001 and the traversal circuit 4005. Regarding the SIMD processor 4001, multiple shaders can synchronously aggregate these results using the standard programming model. One relatively simple way to do this is to use global atomics to aggregate the results of a shared data structure in memory where the intersection results of multiple shaders can be stored. Then, the last shader can resolve the data structure and callback the traversal circuit 4005 to continue the traversal.
[0354] A more efficient approach could also be implemented that restricts the execution of multiple shader calls to the lanes of the same SIMD thread on the SIMD processor 4001. Then, the intersection tests are reduced locally using SIMD / SIMT reduction operations (without relying on global atomics). This implementation may depend on new circuitry within the sorting unit 4008 to keep small batches of shader calls within the same SIMD batch.
[0355] The execution of the traversal thread may be further paused on the traversal circuit 4005. Using the conventional execution model, when a shader is dispatched during traversal, the execution unit 4001 enables the execution of other ray generation commands while processing the shader, so the shader thread ends and the ray traversal state is saved in memory. If the traversal thread is simply paused, the traversal state need not be stored and the results of each shader can be waited for separately. This implementation may include circuitry to avoid deadlocks and provide sufficient hardware utilization.
[0356] Figures 47-48 show an example of a latency model that calls a single shader invocation on a SIMD core / execution unit 4001 having three shaders 4701. When saved, all intersection tests are evaluated within the same SIMD / SIMT group. Thus, the closest intersection can also be computed on the programmable core / execution unit 4001.
[0357] As described above, all or part of the shader aggregation and / or latency may be performed by the traversal / intersection circuit 4005 and / or the core / EU scheduler 4007. Figure 47 shows how the shader latency / aggregation circuit 4706 in the scheduler 4007 delays the scheduling of shaders associated with a particular SIMD / SIMT thread / lane until a specified trigger event occurs. When the trigger event is detected, the scheduler 4007 dispatches multiple aggregated shaders within a single SIMD / SIMT batch to the core / EU 4001.
[0358] Figure 48 shows how the shader latency / aggregation circuit 4805 in the traversal / intersection circuit 4005 can delay the scheduling of shaders associated with a particular SIMD thread / lane until a specified trigger event occurs. When the trigger event is detected, the traversal / intersection circuit 4005 presents the aggregated shaders within a single SIMD / SIMT batch to the sorting unit 4008.
[0359] However, it should be noted that the shader latency and aggregation techniques may be implemented within various other components such as the sort unit 4008, or may be distributed among multiple components. For example, the traversal / intersection circuit 4005 may perform a first set of shader aggregation operations, and the scheduler 4007 may perform a second set of shader aggregation operations to ensure that the shaders of the SIMD threads are efficiently scheduled on the cores / EUs 4001.
[0360] The "triggering event" for dispatching the aggregated shaders to the cores / EUs may be a processing event such as a specific number of accumulated shaders or a minimum latency associated with a specific thread. Alternatively or additionally, the triggering event may be a time event such as a specific duration from the latency of the first shader or a specific number of processor cycles. Other variables such as the current workload on the cores / EUs 4001 and the traversal / intersection unit 4005 may also be evaluated by the scheduler 4007 to determine when to dispatch the SIMD / SIMT batches of shaders.
[0361] Different embodiments of the present invention may be implemented using different combinations of the above techniques, based on the specific system architecture and application requirements being used.
[0362] [Raytracing Instruction] The ray tracing instructions described below are included in an instruction set architecture (ISA) that supports the CPU 3199 and / or the GPU 3105. When executed by the CPU, single instruction multiple data (SIMD) instructions may utilize vector / packed source and destination registers to execute the described operations and may be decoded and executed by the CPU core. When executed by the GPU 3105, the instructions may be executed by the graphics core 3130. For example, any of the above execution units (EUs) 4001 may execute the instructions. Alternatively or additionally, the instructions may be executed by execution circuitry on the ray tracing core 3150 and / or the tensor core 3140.
[0363] FIG. 49 shows an architecture for executing the ray tracing instructions described below. The illustrated architecture may be integrated within one or more of the above cores 3130, 3140, 3150 (see, e.g., FIG. 31 and related text) or may be included in different processor architectures.
[0364] During operation, the instruction fetch unit 4903 fetches the ray tracing instruction 4900 from the memory 3198 and the decoder 4995 decodes the instruction. In one implementation, the decoder 4995 decodes the instruction to generate executable operations (e.g., uops or micro-operations within a microcoding core). Alternatively, some or all of the ray tracing instruction 4900 may be executed without being decoded and, thus, such a decoder 4904 is not required.
[0365] In any implementation, the scheduler / dispatcher 4905 schedules and dispatches instructions (or operations) among the functional units (FUs) 4910 - 4912. The illustrated implementation includes a vector FU 4910 for executing single instruction multiple data (SIMD) instructions that operate simultaneously on multiple packed data elements stored in the vector register 4915, and a scalar FU 4911 for operating on scalar values stored in one or more scalar registers 4916. An optional ray tracing FU 4912 may operate based on the packed data values stored in the vector register 4915 and / or the scalar values stored in the scalar register 4916. In an implementation without the dedicated FU 4912, the vector FU 4910 and possibly the scalar FU 4911 may execute the ray tracing instructions described below.
[0366] The various FUs 4910 - 4912 access the ray tracing data 4902 (e.g., traversal / intersection data) necessary to execute the ray tracing instructions 4900 from the vector register 4915, the scalar register 4916, and / or the local cache subsystem 4908 (e.g., L1 cache). The FUs 4910 - 4912 may also perform accesses to the memory 3198 via load and store operations, and the cache subsystem 4908 may operate independently to cache the data locally.
[0367] The ray tracing instructions may be used to increase performance for ray traversal / intersection and BVH construction, but they may also be applicable in other fields such as high performance computing (HPC) and general purpose GPU (GPGPU) implementations.
[0368] In the following description, a double word may be abbreviated as dw in some cases, and an unsigned byte may be abbreviated as ub. Further, the source registers and destination registers shown below (e.g., src0, src1, dest, etc.) may represent vector register 4915, or in some cases, may represent a combination of vector register 4915 and scalar register 4916. Typically, when the source value or destination value used by an instruction includes packed data elements (e.g., when the source or destination stores N data elements), vector register 4915 is used. Other values may use scalar register 4916 or vector register 4915.
[0369] (Inverse quantization) An example of the Inverse quantization (Dequantize) instruction "inverse quantizes" a previously quantized value. As an example, in a ray tracing implementation, a particular BVH subtree may be quantized to reduce storage and bandwidth requirements. The Inverse quantization instruction may be in the form of inverse_quantize dest src0 src1 src2, where source register src0 stores N unsigned bytes, source register src1 stores 1 unsigned byte, source register src2 stores 1 floating-point value, and destination register dest stores N floating-point values. All of these registers may be vector register 4915. Alternatively, src0 and dest may be vector register 4915, and src1 and src2 may be scalar register 4916.
[0370] The following code sequence defines a particular implementation of the Inverse quantization instruction. for (int i = 0; i < SIMD_WIDTH; i++) { if (execMask[i]) { dst[i] = src2[i] + ldexp(convert_to_float(src0[i]), src1); } } In this example, ldexp multiplies a double-precision floating-point value by a specified power of two (i.e., ldexp(x, exp) = x * 2 exp ). In the above code, if the execution mask value associated with the current SIMD data element (execMask[i]) is set to 1, the SIMD data element at position i in src0 is converted to a floating-point value and multiplied by the integer power of two (2 src1 value) of the value in src1, and this value is added to the corresponding SIMD data element in src2.
[0371] (Optional min or max) The optional min or max instruction may perform either a min or max operation for each lane, as indicated by the bits of a bitmask (i.e., return the minimum or maximum value of a set of values). The bitmask may utilize vector register 4915, scalar register 4916, or a separate set of mask registers (not shown). The following code sequence, i.e., sel_min_max dest src0 src1 src2, defines one specific implementation of the min / max instruction, where src0 stores N doublewords, src1 stores N doublewords, src2 stores one doubleword, and the destination register stores N doublewords.
[0372] The following code sequence defines one specific implementation of the optional min / max instruction. for (int i = 0; i < SIMD_WIDTH) { if (execMask[i]) { dst[i] = (1 << i) & src2? min(src0[i], src1[i]) : max(src0[i], src1[i]); } } In this example, the value of (1<<i)&src2 (1 left-shifted by i and ANDed with src2) is used to select the minimum value of the i-th data element of src0 and src1, or the maximum value of the i-th data element of src0 and src1. The operation is performed on the i-th data element only if the execution mask value (execMask[i]) associated with the current SIMD data element is set to 1.
[0373] (Shuffle index instruction) The shuffle index instruction can copy any set of input lanes to the output lanes. In the case of a 32-bit SIMD width, this instruction can be executed with a lower throughput. This instruction has the form shuffle_index dest src0 src1 <optional flag>, where src0 stores N doublewords, src1 stores N unsigned bytes (i.e., index values), and dest stores N doublewords.
[0374] The following code sequence defines one specific implementation of the shuffle index instruction. for (int i = 0; i < SIMD_WIDTH; i++) { uint8_t srcLane = src1.index[i]; if (execMask[i]) { bool invalidLane = srcLane < 0 || srcLane >= SIMD_WIDTH ||!execMask[srcLaneMod]; if (FLAG) { invalidLane |= flag[srcLaneMod]; } if (invalidLane) { dst[i] = src0[i]; } else { dst[i] = src0[srcLane]; } } } In the above code, the index of src1 identifies the current lane. If the i-th value in the execution mask is set to 1, a check is performed to ensure that the source lane is within the range from 0 to the SIMD width. If so, a flag is set (srcLaneMod), and the i-th data element of the destination is set equal to the i-th data element of src0. If the lane is within the range (i.e., valid), the index value of src1 (srcLane0) is used as the index to src0 (dst[i]=src0[srcLane]).
[0375] (Immediate Shuffle Up / Dn / XOR Instruction) The immediate shuffle instruction may shuffle input data elements / lanes based on the immediate value of the instruction. The immediate value may specify shifting the input lanes by 1, 2, 4, 8, or 16 positions based on the value of the immediate. Optionally, a further scalar input register can be specified as a fill value. When the source lane index is invalid, the fill value (if provided) is stored in the destination data element position. If no fill value is provided, all data element positions are set to 0.
[0376] The flag register may be used as a source mask. If the flag bit of the source lane is set to 1, the source lane is marked as invalid and the instruction may proceed.
[0377] The following are examples of different implementations of the immediate shuffle instruction. shuffle_ <up dn xor>_<1 / 2 / 4 / 8 / 16> dest src0 <optional src1> <optional flag> shuffle_ <up dn xor>_<1 / 2 / 4 / 8 / 16> dest src0 <optional src1> <optional flag> In this implementation, src0 stores N doublewords, src1 stores one doubleword for the fill value (if present), and dest stores N doublewords containing the result.
[0378] The following code sequence defines one specific implementation of the immediate shuffle instruction. for (int i = 0; i < SIMD_WIDTH; i++) { int8_t srcLane; switch (SHUFFLE_TYPE) { case UP: srcLane = i - SHIFT; case DN: srcLane = i + SHIFT; case XOR: srcLane = i ^ SHIFT; } if (execMask[i]) { bool invalidLane = srcLane < 0 || srcLane >= SIMD_WIDTH ||!execMask[srcLane]; if (FLAG) { invalidLane |= flag[srcLane]; } if (invalidLane) { if (SRC1) dst[i] = src1; else dst[i] = 0; } else { dst[i] = src0[srcLane]; } } } Here, the input data element / lane is shifted by only 1, 2, 4, 8, or 16 positions based on the value of an immediate. Register src1 is a further scalar source register and is used as a fill value to be stored at the destination data element position when the source lane index is invalid. If no fill value is provided and the source lane index is invalid, the destination data element position is set to 0. The flag register (FLAG) is used as a source mask. If the flag bit of a source lane is set to 1, the source lane is marked as invalid and the instruction proceeds as described above.
[0379] (Indirect Shuffle Up / Dn / XOR Instruction) The indirect shuffle instruction has a source operand (src1) that controls the mapping from source lanes to destination lanes. The indirect shuffle instruction may be in the following form: shuffle_ <up dn xor>dest src0 src1 <optional flag> src0 stores N doublewords, src1 stores 1 doubleword, and dest stores N doublewords.
[0380] The following code sequence defines one specific implementation of the immediate shuffle instruction. for (int i = 0; i < SIMD_WIDTH; i++) { int8_t srcLane; switch (SHUFFLE_TYPE) { case UP: srcLane = i - src1; case DN: srcLane = i + src1; case XOR: srcLane = i ^ src1; } if (execMask[i]) { bool invalidLane = srcLane < 0 || srcLane >= SIMD_WIDTH ||!execMask[srcLane]; if (FLAG) { invalidLane |= flag[srcLane]; } if (invalidLane) { dst[i] = 0; } else { dst[i] = src0[srcLane]; } } } Therefore, the indirect shuffle instruction operates in a similar manner to the immediate shuffle instruction above, but the mapping of source lanes to the destination lane is controlled by the source register src1 rather than an immediate value.
[0381] (Cross-lane min / max instruction) The cross-lane minimum / maximum instructions may be supported for floating-point and integer data types. The cross-lane minimum instruction may be in the form of lane_min dest src0, and the cross-lane maximum instruction may be in the form of lane_max dest src0, where src0 stores N doublewords and dest stores 1 doubleword.
[0382] As an example, the following code sequence defines one specific implementation of the cross-lane minimum. dst=src[0]; for (int i=1;i<SIMD_WIDTH) { if (execMask[i]) { dst=min(dst,src[i]); } } In this example, the doubleword value at the data element position i of the source register is compared with the data element of the destination register, and the minimum of the two values is copied to the destination register. The cross-lane maximum instruction operates in substantially the same way, with the only difference being that the maximum of the data element and the destination value at position i is selected.
[0383] (Cross-lane min / max index instructions) The cross-lane minimum index instruction may be in the form of lane_min_index dest src0, and the cross-lane maximum index instruction may be in the form of lane_max_index dest src0, where src0 stores N doublewords and dest stores 1 doubleword.
[0384] As an example, the following code sequence defines one specific implementation of the cross-lane minimum index instruction. dst_index=0; tmp=src[0] for (int i=1;i<SIMD_WIDTH) { if (src[i]<tmp&&execMask[i]) { tmp=src[i]; dst_index=i; } } In this example, the destination index spans the destination register and is incremented from 0 to the SIMD width. If the execution mask bit is set, the data element at position i in the source register is copied to the temporary storage location (tmp), and the destination index is set to the data element position i.
[0385] (Cross-lane sorting network instruction) The cross-lane sorting network instruction may sort all N input elements using an N-wide (stable) sorting network either in ascending order (sortnet_min) or descending order (sortnet_max). The min / max versions of the instruction may be in the form of sortnet_min dest src0 and sortnet_max dest src0 respectively. In one implementation, src0 and dest store N doublewords. The min / max sort is performed on the N doublewords of src0, and the ascending elements (in the case of min) or descending elements (in the case of max) are stored in dest in their respective sorted order. An example of the code sequence defining the instruction is dst = apply_N_wide_sorting_network_min / max(src0).
[0386] (Cross-lane sorting network index instruction) The crosslane sorting network index instruction may sort all N input elements using an N-wide (stable) sorting network, but returns a replacement index either in ascending (sortnet_min) or descending (sortnet_max) order. The min / max versions of the instruction may be in the form of sortnet_min_index dest src0 and sortnet_max_index dest src0, where src0 and dest each store N doublewords. An example of the code sequence defining the instruction is dst = apply_N_wide_sorting_network_min / max_index(src0).
[0387] A method for executing any of the above instructions is shown in FIG. 50. The method may be implemented in the above specific processor architecture, but is not limited to any specific processor or system architecture.
[0388] At 5001, the instructions of the primary graphics thread are executed on the processor core. This may include, for example, any of the above cores (e.g., graphics core 3130). When reaching the ray tracing operation within the primary graphics thread as determined at 5002, the ray tracing instructions are offloaded to the ray tracing execution circuit, which may be in the form of a functional unit (FU) as described above with respect to FIG. 49, or a dedicated ray tracing core 3150 as described with respect to FIG. 31.
[0389] In 5003, a ray tracing instruction is decoded and fetched from memory. In 5005, the instruction is decoded into an executable operation (e.g., in an embodiment that requires a decoder). In 5004, the ray tracing instruction is scheduled and dispatched for execution by a ray tracing circuit. In 5005, the ray tracing instruction is executed by the ray tracing circuit. For example, the instruction may be dispatched and executed on the above-mentioned FUs (e.g., vector FU4910, ray tracing FU4912, etc.) and / or the graphics core 3130 or the ray tracing core 3150.
[0390] When the execution of the ray tracing instruction is completed, in 5006, the result is stored (e.g., stored in memory 3198), and in 5007, the primary graphics thread is notified. In 5008, the ray tracing result is processed within the context of the primary thread (e.g., read from memory and integrated into the graphics rendering result).
[0391] In an embodiment, the terms "engine" or "module" or "logic" may refer to an application specific integrated circuit (ASIC), an electronic circuit, a processor (shared, dedicated, or grouped), and / or a memory (shared, dedicated, or grouped) that executes one or more software or firmware programs, combinational logic circuits, and / or other suitable components that provide the above functionality, may be a part of these, or may include these. In an embodiment, the engine, module, or logic may be implemented in firmware, hardware, software, or any combination of firmware, hardware, and software.
[0392] [Bounding Volume and Ray-Box Intersection Test] Figure 51 is a diagram of a bounding volume 5102 according to an embodiment. The illustrated bounding volume 5102 is an axis aligned with a three-dimensional axis 5100. However, embodiments are applicable to different bounding representations (e.g., oriented bounding box, discrete oriented polytope, sphere, etc.) and any number of dimensions. The bounding volume 5102 defines the minimum and maximum extents of a three-dimensional object 5104 along each dimension of the axis 5100. To generate a BVH of a scene, a bounding box is constructed for each object in a set of objects within the scene. Then, a set of parent bounding boxes can be constructed around a group of the bounding boxes constructed for each object.
[0393] Figures 52A-52B illustrate a representation of a bounding volume hierarchy for two-dimensional objects. Figure 52A shows a set of bounding volumes 5200 around a set of geometric objects. Figure 52B shows an ordered tree 5202 of the bounding volumes 5200 of Figure 52A.
[0394] As shown in Figure 52A, the set of bounding volumes 5200 includes a root bounding volume N1, which is the parent bounding volume of all the other bounding volumes N2-N7. The bounding volumes N2 and N3 are internal bounding volumes between the root volume N1 and the leaf volumes N4-N7. The leaf volumes N4-N7 contain the geometric objects O1-O8 of the scene.
[0395] FIG. 52B shows an order tree 5202 of bounding volumes N1 to N7 and geometric objects O1 to O8. The illustrated order tree 5202 is a binary tree in which each node of the tree has two child nodes. A data structure configured to include information for each node can include bounding information for the bounding volume (e.g., bounding box) of the node and at least a reference to each child node of the node.
[0396] The order tree 5202 representation of the bounding volume defines a hierarchy that can be used to perform various operations of a hierarchical version, including but not limited to collision detection and ray-box intersection. In an example of ray-box intersection, the nodes can start from the root node N1, which is the parent node to all other bounding volume nodes in the hierarchy, and can be tested hierarchically. If the ray-box intersection test for the root node N1 fails, all other nodes of the tree may be bypassed. If the ray-box intersection test for the root node N1 passes, the subtrees of the tree can be tested until at least a set of intersecting leaf nodes N4 to N7 is determined and can be traversed or bypassed in order. The exact test and traversal algorithms used can vary according to the embodiment.
[0397] FIG. 53 is a diagram of a ray-box intersection test according to an embodiment. During the ray-box intersection test, a ray 5302 is emitted, and the equation defining the ray can be used to determine whether the ray intersects the plane defining the bounding box 5300 being tested. The ray 5302 can be expressed as O + D·t, where O corresponds to the viewpoint of the ray, D is the direction of the ray, and t is a real value. The change in t can be used to define any point along the ray. The ray 5302 is said to intersect the bounding box 5300 when the maximum incident plane intersection distance is less than or equal to the minimum exit plane distance. For the ray 5302 in FIG. 53, the y-plane incident intersection distance is t min-y shown as 5304. The y-plane exit intersection distance is t max-y It is shown as 5308. The incident intersection distance on the x plane is t min-x can be calculated at 5306, and the exit intersection distance on the x plane is t t max-x It is shown as 5310. Therefore, for a given ray 5302, t min-x 5306 is t max-y Since it is less than 5308, it can be numerically shown that it intersects the bounding box at least along the x and y planes. To perform a ray-box intersection test using a graphics processor, the graphics processor is configured to store at least an acceleration data structure that defines each bounding box to be tested. For acceleration using a bounding volume hierarchy, at least a reference to a child node to the bounding box is stored.
[0398] [Bounding Volume Node Compression] For an axis-aligned bounding box in 3D space, the acceleration data structure can store the lower and upper limits of the bounding box in three dimensions. A software implementation can use 32-bit floating-point numbers to store these limits (bounds), which amounts to a total of 2×3×4 = 24 bytes per bounding box. For an N-wide BVH node, N boxes and N child references must be stored. In total, the storage for a 4-wide BVH node is N*24 bytes + N*4 bytes per reference, which is a total of (24 + 4)*N bytes, amounting to 112 bytes for a 4-wide BVH node and 224 bytes for an 8-wide BVH node in total.
[0399] In one embodiment, the size of the BVH node is reduced by storing a single, more accurate parent bounding box that encloses all child bounding boxes and storing each child bounding box with lower precision relative to that parent box. Depending on the usage scenario, different numbers of representations may be used to store the high-precision parent bounding box and the low-precision relative child bounds. FIG. 54 is a block diagram showing an exemplary quantized BVH node 5410 according to an embodiment. The quantized BVH node 5410 can include higher-precision values to define the parent bounding box of the BVH node. For example, parent_lower_x5412, parent_lower_y5414, parent_lower_z5416, parent_upper_x5422, parent_upper_y5424, and parent_upper_z5426 can be stored using single-precision or double-precision floating-point values. The values of the child bounding boxes for each child bounding box stored in the node can be quantized and stored as lower-precision values, such as a fixed-point representation relative to the bounding box values defined for the parent bounding box. For example, child_lower_x5432, child_lower_y5434, and child_lower_z5436, as well as child_upper_x5442, child_upper_y5444, and child_upper_z5446, can be stored as lower-precision fixed-point values. Additionally, child references 5452 can be stored for each child. The child reference 5452 can be an index into a table storing the positions of each child node, or can be a pointer to the child node.
[0400] As shown in FIG. 54, single-precision or double-precision floating-point values may be used to store the parent bounding box, and M-bit fixed-point values may be used to encode the relative child bounding boxes. The data structure for the quantized BVH node 5410 of FIG. 54 can be defined by the quantized N-wide BVH node shown in Table 1 below.
[0401] [Table 2] The quantized nodes in Table 1 achieve a reduced data structure size by storing higher-precision values within the range of the parent bounding box and quantizing the child values while maintaining the baseline accuracy level. In Table 1, Real indicates a higher-precision numerical representation (e.g., 32-bit or 64-bit floating-point value), UintM indicates a lower-precision unsigned integer using M bits of precision used to represent a fixed-point number, and Reference indicates the type used to represent a reference to a child node (e.g., a 4-byte index of an 8-byte pointer).
[0402] A typical instantiation of this approach can use 32-bit child references, single-precision floating-point values for the parent bounding, and M = 8 bits (1 byte) for the relative child bounds. This compressed node requires 6*4 + 6*N + 4*N bytes. For a 4-wide BVH, this amounts to a total of 64 bytes (versus 112 bytes for the uncompressed version), and for an 8-wide BVH, this amounts to a total of 104 bytes (versus 224 bytes for the uncompressed version).
[0403] To traverse such a compressed BVH node, the graphics processing logic can expand the relative child bounding box and then intersect the expanded node using standard techniques. The uncompressed lower bounds can then be obtained for each dimension x, y, and z. The following Equation 1 shows the equation for obtaining the lower_x value of a child.
[0404] [Equation] In Equation 1 above, M represents the number of bits of precision for the fixed-point representation of the child bounds. The logic for expanding child data for each dimension of the BVH node can be implemented as shown in Table 2 below.
[0405]
Table 3
[0406] In one embodiment, the performance of the expansion can be improved by storing, instead of the parent_upper_x / y / z values, the size of the scaled parent bounding box, e.g., (parent_upper_x - parent_lower_x) / (2 M-1 ). In such an embodiment, the range of the child bounding box can be calculated according to the exemplary logic shown in Table 3.
[0407]
Table 4
[0408] The above approach works well in shader or CPU-based implementations, but one embodiment provides special hardware configured to perform ray tracing operations, including ray-box intersection tests, using a bounding volume hierarchy. In such an embodiment, the special hardware can store a further quantized representation of the BVH node data and be configured to automatically dequantize such data when performing ray-box intersection tests.
[0409] FIG. 55 is a block diagram of a composite floating point data block 5500 for use by a quantized BVH node 5510, according to a further embodiment. In one embodiment, logic for supporting the composite floating point data block 5500 can be defined by special logic within the graphics processor, as opposed to a 32-bit single precision floating point representation or a 64-bit double precision floating point representation of the range of the parent bounding box. The composite floating point (CFP) data block 5500 can include a 1-bit signed bit 5502, a variable size (E-bit) signed integer exponent 5504, and a variable size (K-bit) mantissa 5506. Multiple values for E and K may be configurable by adjusting values stored in the configuration registers of the graphics processor. In one embodiment, the values of E and K may be configured independently within a range of values. In one embodiment, a fixed set of mutually related values for E and K may be selected via the configuration registers. In one embodiment, a single value for each of E and K is hard-coded into the BVH logic of the graphics processor. The values E and K enable the CFP data block 5500 to be used as a customized (e.g., special-purpose) floating point data type that can be adjusted to the dataset.
[0410] Using the CFP data block 5500, the graphics processor can be configured to store the bounding box data within the quantized BVH node 5510. In one embodiment, the lower bounds of the parent bounding box (parent_lower_x5512, parent_lower_y5514, parent_lower_z5516) are stored at a precision level determined by the E and K values selected for the CFP data block 5500. The precision level of the stored values for the lower bounds of the parent bounding box is generally set to a higher precision than the values of the child bounding boxes (child_lower_x5524, child_upper_x5526, child_lower_y5534, child_upper_y5536, child_lower_z5544, and child_upper_z5546) stored as fixed-point values. The scaled parent bounding box size is stored as a power of two exponent (e.g., exp_x5522, exp_y5532, exp_z5542). Further, a reference for each child (e.g., child reference 5552) can be stored. The size of the quantized BVH node 5510 can be scaled based on the width (e.g., number of children) stored at each node, and the amount of memory used to store the child references and the bounding box values for the child nodes increases with each additional node.
[0411] The logic for the implementation of the quantized BVH node of FIG. 55 is shown in Table 4 below.
[0412]
Table 5
[0413] In the example of E = 8, K = 16, and M = 8, using 32 bits for child references, the QuantizedNodeHW structure in Table 4 has a size of 52 bytes for a 4-wide BVH and a size of 92 bytes for an 8-wide BVH. This represents a reduction in the structure size compared to the quantized nodes in Table 1 and a significant reduction in the structure size compared to existing implementations. Note that for the mantissa value (K = 16), 1 bit of the mantissa may be implied, reducing the storage requirement to 15 bits.
[0414] The layout of the BVH node structure in Table 4 enables reduced hardware to perform ray-box intersection tests for child bounding boxes. The hardware complexity is reduced based on several factors. Since the relative child bounds add an additional M-bit precision, a smaller number of bits can be selected for K. The scaled parent bounding box size is stored as a power of 2 (exp_x / y / z fields), which simplifies the calculations. Further, the calculations are refactored to reduce the multiplier size.
[0415] In one embodiment, the ray intersection logic of the graphics processor calculates the hit distance of the ray to the axis-aligned planes in order to perform a ray-box test. The ray intersection logic can use BVH node logic that includes support for the quantized node structure of Table 4. The logic can calculate the distance to the lower bound of the parent bounding box using the higher-precision lower bound of the parent and the quantized relative range of the child box. Exemplary logic for x-plane calculations is shown in Table 5 below.
[0416]
Table 6
[0417] Using the lower bound of the parent, the intersection distances to the relative child bounding boxes can be calculated for each child bounding box as exemplified by the calculations of dist_child_lower_x and dist_child_upper_x as in Table 5. The calculations of the dist_child_lower / upper_x / y / z values can be performed using a 23-bit × 8-bit multiplier.
[0418] FIG. 56 shows a ray-box intersection that uses quantization values to define a child bounding box 5610 relative to a parent bounding box 5600, according to an embodiment. Applying the ray-box intersection distance determination formula for the x-plane shown in Table 5, the distance along ray 5602 where the ray intersects the bounds of the parent bounding box 5600 along the x-plane can be determined. The position dist_parent_lower_x5603 where ray 5602 intersects the lower bounding surface 5604 of the parent bounding box 5600 can be determined. Based on dist_parent_lower_x5603, dist_child_lower_x5605 where the ray intersects the minimum bounding surface 5606 of the child bounding box 5610 can be determined. Further, based on dist_parent_lower_x5603, for the position where the ray intersects the maximum bounding surface 5608 of the child bounding box 5610, dist_child_upper_x5607 can be determined. Similar determinations can be performed for each dimension in which the parent bounding box 5600 and the child bounding box 5610 are defined (e.g., along the y-axis and z-axis). The surface intersection distance can then be used to determine whether the ray intersects the child bounding box. In one embodiment, the graphics processing logic can use SIMD and / or vector logic to determine the intersection distances for multiple dimensions and multiple bounding boxes in parallel. Further, at least a first portion of the calculations described herein may be executed on a graphics processor, while a second portion of the calculations may be executed on one or more application processors coupled to the graphics processor.
[0419] FIG. 57 is a flowchart of BVH extension and traversal logic 5700 according to an embodiment. In one embodiment, the BVH extension and traversal logic may be present within the special purpose hardware logic of a graphics processor or may be executed by shader logic executed on the execution resources of the graphics processor. As shown in block 5702, the BVH extension and traversal logic 5700 can cause the graphics processor to perform operations for calculating the distance along a ray to the lower bounding surface of a parent bounding volume. In block 5704, the logic can calculate the distance to the lower bounding surface of a child bounding volume based in part on the calculated distance to the lower bounding surface of the parent bounding volume. In block 5706, the logic can calculate the distance to the upper bounding surface of a child bounding volume based in part on the calculated distance to the lower bounding surface of the parent bounding volume.
[0420] In block 5708, the BVH extension and traversal logic 5700 can determine ray intersection for a child bounding volume based in part on the distances to the upper and lower bounding surfaces of the child bounding volume, but the intersection distances for each dimension of the bounding box are used to determine the intersection. In one embodiment, the BVH extension and traversal logic 5700 determines ray intersection for the child bounding volume by determining whether the maximum incident surface intersection distance of the ray is less than or equal to the minimum exit surface distance. In other words, the ray intersects the child bounding volume when the ray enters the bounding volume along all of the defined surfaces before exiting the bounding volume along any of the defined surfaces. At 5710, if the BVH extension and traversal logic 5700 determines that the ray intersects the child bounding volume, then as shown in block 5712, the logic can traverse the child node for the bounding volume to test the child bounding volume within the child node. At block 5712, node traversal can be performed and access can be made to a reference to the node associated with the intersecting bounding box. The child bounding volume becomes the parent bounding volume and the children of the intersecting bounding volume can be evaluated. At 5710, if the BVH extension and traversal logic 5700 determines that the ray does not intersect the child bounding volume, then as shown in block 5714, the branch of the bounding hierarchy associated with the child bounding volume is skipped. This is because the ray does not intersect the bounding volume further down the subtree branch associated with the non-intersecting child bounding volume.
[0421] [Further Compression via Shared Plane Bounding Boxes] For any N-wide BVH using a bounding box, the bounding volume hierarchy can be constructed such that each of the six sides of the 3D bounding box is shared by at least one child bounding box. In a 3D shared-face bounding box, 6×log2 N bits can be used to indicate whether a given face of the parent bounding box is shared with a child bounding box. For N = 4 in a 3D shared-face bounding box, 12 bits are used to indicate the shared faces, and 2 bits each are used to identify which of the four children potentially reuse the face of the parent being shared. Each bit can be used to indicate whether the parent face is reused by a particular child. For a 2-wide BVH, for each face of the parent bounding box, 6 additional bits can be added to indicate whether the face (e.g., side face) of the bounding box is shared by a child. The concept of SPBB can be applied to any number of dimensions, but in one embodiment, the advantages of SPBB generally become greatest for 2-wide (e.g., binary) SPBBs.
[0422] The use of shared-face bounding boxes can further reduce the amount of data stored when using BVH node quantization, as described herein. In an example of a 3D 2-wide BVH, the 6 bits for the shared faces can reference min_x, max_x, min_y, max_y, min_z, and max_z for the parent bounding box. If the min_x bit is zero, the first child inherits the shared face from the parent bounding box. For a child that shares a face with the parent bounding box, it is not necessary to store the quantized value of that face, which reduces the storage cost and expansion cost of the node. Further, a high-precision value for the face can be used in the child bounding box.
[0423] FIG. 58 is a diagram of an exemplary two-dimensional shared plane bounding box 5800. The two-dimensional (2D) shared plane bounding box (SPBB) 5800 ...
Claims
1. A displacement mapping circuit / logic that generates an original displacement-mapped mesh by performing displacement mapping on a plurality of vertices of a base subdivision mesh, and a mesh compression circuit / logic that compresses the original displacement-mapped mesh comprising: The mesh compression circuit / logic includes a quantizer that determines a difference vector from each vertex of a coarse base mesh to a corresponding displaced vertex of the original displacement-mapped mesh and quantizes the original displacement-mapped mesh with respect to the coarse base mesh by combining the difference vectors in a displacement array. An apparatus.
2. The apparatus according to claim 1, wherein the mesh compression circuit / logic further stores a compressed displacement mesh including the difference vectors of the displacement array and original quarter base coordinates of the base subdivision mesh.
3. The apparatus according to claim 2, wherein the coarse base mesh includes the base subdivision mesh.
4. The apparatus according to claim 2 or 3, further comprising an interpolator that performs bilinear interpolation on the base subdivision mesh to generate the coarse base mesh.
5. The apparatus according to any one of claims 2 to 4, further comprising a decompression circuit / logic that decompresses the compressed displacement mesh according to a request.
6. The apparatus according to claim 5, wherein the decompression circuit / logic attempts to reconstruct the original displacement-mapped mesh by combining coordinates of the base mesh with elements of the displacement array to generate a decompressed displacement-mapped mesh.
7. The apparatus according to claim 6, wherein the decompressed displacement-mapped mesh includes an approximation of the original displacement-mapped mesh.
8. The apparatus according to claim 7, further comprising a BVH generation circuit that generates a bounding volume hierarchy (BVH) based on a plurality of primitives including a first primitive associated with the decompressed displacement-mapped mesh.
9. The apparatus according to claim 8, further comprising a ray traversal / intersection circuit that traverses one or more rays through the BVH to identify intersections with the decompressed displacement-mapped mesh.
10. Generating an original displacement-mapped mesh by performing displacement mapping on a plurality of vertices of a base subdivision mesh; Determining a difference vector from each vertex of a coarse base mesh to a corresponding displaced vertex of the original displacement-mapped mesh and combining the difference vectors in a displacement array to compress the original displacement-mapped mesh by quantizing the original displacement-mapped mesh with respect to the coarse base mesh; A method comprising the steps of: The method according to claim 10, further comprising storing a compressed displacement mesh including the difference vectors of the displacement array and original quarter base coordinates of the base subdivision mesh.
12. The method according to claim 11, wherein the coarse base mesh includes the base subdivision mesh.
13. The method according to claim 11 or 12, further comprising performing bilinear interpolation on the base subdivision mesh to generate the coarse base mesh.
14. The method according to any one of claims 11 to 13, further comprising expanding the compressed displacement mesh in response to a request.
15. The method according to claim 14, wherein the expanding step further comprises combining the coordinates of the base mesh with elements of the displacement array to generate an expanded displacement-mapped mesh.
16. The method according to claim 15, wherein the expanded displacement-mapped mesh includes an approximation of the original displacement-mapped mesh.
17. The method according to claim 16, further comprising generating a bounding volume hierarchy (BVH) based on a plurality of primitives including a first primitive associated with the expanded displacement-mapped mesh.
18. The method according to claim 17, further comprising traversing one or more rays through the BVH to identify intersections with the expanded displacement-mapped mesh.
19. When executed by a machine, for the machine, Performing displacement mapping on a plurality of vertices of a base subdivision mesh to generate an original displacement-mapped mesh; Determine a difference vector from each vertex of the coarse base mesh to the corresponding displaced vertex of the original displacement-mapped mesh, and by combining the difference vectors in the displacement array, quantize the original displacement-mapped mesh with respect to the coarse base mesh, thereby compressing the original displacement-mapped mesh and A program for causing execution. The program according to claim 19, further causing the machine to perform an operation of storing a compressed displacement mesh including the difference vectors of the displacement array and the original quarter base coordinates of the base subdivision mesh.
21. The program according to claim 20, wherein the coarse base mesh includes the base subdivision mesh.
22. The program according to claim 20 or 21, further causing the machine to perform an operation of performing bilinear interpolation on the base subdivision mesh to generate the coarse base mesh.
23. The program according to any one of claims 20 to 22, further causing the machine to perform an operation of expanding the compressed displacement mesh in response to a request.
24. The program according to claim 23, wherein the expansion further includes combining the coordinates of the base mesh with the elements of the displacement array to generate an expanded displacement-mapped mesh.
25. The program according to claim 24, wherein the expanded displacement-mapped mesh includes an approximation of the original displacement-mapped mesh.
26. A machine-readable storage medium storing the program according to any one of claims 19 to 25.
Citation Information
Patent Citations
Graphics model conversion device, graphics model processing program making computer function as graphics modeling conversion device
JP2009151754A
System, method, and program for photorealistic imaging using ambient occlusion
JP2010134919A
Encoding method, encoding device, decoding method, and decoding device
JP2013539125A
3-dimensional image data compression device, method, program, and recording medium
WO2006062199A1