Partially resident boundary volume hierarchy
The partially resident BVH technique addresses memory access issues in ray tracing by treating non-resident BVH pages as misses, improving traversal efficiency and reducing rendering delays.
Patent Information
- Application Number
- JP2022552425
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2020-03-13
- Filing Date
- 2021-03-03
- Publication Date
- 2025-07-30
- Estimated Expiration
- 2041-03-03
AI Technical Summary
Ray tracing operations are computationally expensive and can result in significant memory access stalls due to the large size of bounding volume hierarchies (BVH) that may not fit entirely in readily accessible memory, leading to inefficient traversal and rendering delays.
Implementing a partially resident bounding volume hierarchy (BVH) where memory pages classified as non-resident are treated as misses during traversal, allowing operations to proceed without waiting for data to be loaded into immediate memory, and using memory management units (MMUs) to manage residency designations.
This approach reduces memory access stalls and enhances rendering efficiency by allowing traversal to continue without full residency, thus optimizing performance in ray tracing applications.
Smart Images

Figure 0007715724000001 
Figure 0007715724000002 
Figure 0007715724000003
Abstract
Description
Technical Field
[0001] (Cross - Reference to Related Applications) This application claims the benefit of U.S. Patent Application No. 16 / 819,014, filed on March 13, 2020, the content of which is incorporated herein by reference.
Background Art
[0002] Ray tracing is a type of graphics rendering technique that casts simulated rays (rays) to test for intersections with objects and color pixels based on the results of ray casting. Ray tracing is computationally more expensive than rasterization - based techniques but yields more physically accurate results. Improvements to ray - tracing operations are constantly being made.
[0003] A more detailed understanding will become possible from the following description given by way of example in conjunction with the accompanying drawings.
Brief Description of the Drawings
[0004]
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6
Figure 7
DETAILED DESCRIPTION OF THE INVENTION
[0005] Techniques are provided for performing ray tracing of a ray. The techniques include identifying a first memory page classified as resident based on a first traversal of a boundary volume hierarchy, obtaining a first portion of the boundary volume hierarchy associated with the first memory page, traversing the first portion of the boundary volume hierarchy according to a ray intersection test, identifying a second memory page classified as valid and non-resident based on a second traversal of the boundary volume hierarchy, and determining that an error has occurred for each node of the boundary volume hierarchy within the second memory page in response to the second memory page being classified as valid and non-resident.
[0006] FIG. 1 is a block diagram of an exemplary device 100 that can implement one or more features of the present disclosure. Device 100 can be, for example, a computer, a gaming device, a handheld device, a set-top box, a television, a mobile phone, a tablet computer, or any other computing device, but is not limited thereto. Device 100 includes a processor 102, a memory 104, a storage device 106, one or more input devices 108, and one or more output devices 110. Also, device 100 includes one or more input drivers 112 and one or more output drivers 114. Any input driver 112 is embodied as hardware, a combination of hardware and software, or software, and serves to control the input device 108 (e.g., control operations, receive inputs from the input driver 112, and provide data to the input driver 112). Similarly, any output driver 114 is embodied as hardware, a combination of hardware and software, or software, and serves to control the output device 110 (e.g., control operations, receive inputs from the output driver 114, and provide data to the output driver 114). It should be understood that device 100 can include additional components not shown in FIG. 1. <http: / / www.google.com / patents / US20170076776A1?cl=en&q=FIG.1&f=0> <http: / / www.google.com / patents / US20170076776A1?cl=en&q=FIG.2&f=0>
[0007] <http: / / www.google.com / patents / US20170076776A1?cl=en&q=FIG.3&f=0> In various alternatives, processor 102 includes a central processing unit (CPU), a graphics processing unit (GPU), a CPU and GPU located on the same die, or one or more processor cores, and each processor core can be a CPU or a GPU. In various alternatives, memory 104 can be located on the same die as processor 102 or separately from processor 102. Memory 104 includes volatile or non-volatile memory (e.g., random access memory (RAM), dynamic RAM, cache). <http: / / www.google.com / patents / US20170076776A1?cl=en&q=FIG.4&f=0> <http: / / www.google.com / patents / US20170076776A1?cl=en&q=FIG.5&f=0>
[0008] <http: / / www.google.com / patents / US20170076776A1?cl=en&q=FIG.6&f=0> The memory device 106 includes a fixed or removable memory device (e.g., but not limited to, a hard disk drive, a solid state drive, an optical disk, a flash drive). The input device 108 includes, but is not limited to, a keyboard, a keypad, a touch screen, a touch pad, a detector, a microphone, an accelerometer, a gyroscope, a biometric scanner, or a network connection (e.g., a wireless local area network card for transmitting and / or receiving wireless IEEE802 signals). The output device 110 includes, but is not limited to, a display, a speaker, a printer, a tactile feedback device, one or more lights, an antenna, or a network connection (e.g., a wireless local area network card for transmitting and / or receiving wireless IEEE802 signals).
[0009] Input driver 112 and output driver 114 each include one or more hardware, software, and / or firmware components configured to interface with and drive input device 108 and output device 110, respectively. Input driver 112 communicates with processor 102 and input device 108 to enable processor 102 to receive inputs from input device 108. Output driver 114 communicates with processor 102 and output device 110 to enable processor 102 to send outputs to output device 110. Output driver 114 includes an accelerated processing device (APD) 116 coupled to display device 118, which in some examples is a physical display device or a simulated device that uses a remote display protocol to present outputs. APD 116 is configured to receive compute commands and graphics rendering commands from processor 102, process those compute and graphics rendering commands, and provide pixel outputs to display device 118 for display. As will be explained in more detail below, APD 116 includes one or more parallel processing units configured to perform computations according to a single-instruction-multiple-data (SIMD) paradigm. Thus, while various functions are described herein as being performed by or in conjunction with APD 116, in various alternatives, the functions described as being performed by APD 116 are not driven by a host processor (e.g., processor 102) and are performed additionally or alternatively by other computing devices having similar capabilities configured to provide graphics output to display device 118. For example, any processing system configured to perform processing tasks according to the SIMD paradigm is contemplated to perform the functions described herein. Alternatively, computing systems not configured to perform processing tasks according to the SIMD paradigm are contemplated to perform the functions described herein.
[0010] Figure 2 is a diagram showing details of device 100 and APD 116 according to an example. Processor 102 (FIG. 1) executes operating system 120, driver 122, and application 126, and in some situations, alternatively or additionally executes other software. The operating system 120 controls various aspects of device 100, such as managing hardware resources, processing service requests, scheduling and controlling process execution, and performing other operations. The APD driver 122 controls the operation of APD 116 and sends tasks such as graphics rendering tasks or other work to APD 116 for processing. The APD driver 122 also includes a just-in-time compiler that compiles programs for execution by processing components of APD 116 (such as SIMD unit 138 described in more detail below).
[0011] APD 116 executes commands and programs for selected functions such as graphic operations and non-graphic operations suitable for parallel processing. Based on the commands received from processor 102, APD 116 can be used to perform graphics pipeline operations such as pixel operations, geometric calculations, and rendering of images to display device 118. Also, based on the commands received from processor 102, APD 116 performs computational processing operations not directly related to graphic operations, such as operations related to video, physical simulation, computational fluid dynamics, or other tasks. In some examples, these computational processing operations are performed by executing compute shaders on SIMD unit 138.
[0012] APD116 includes a computing unit 132 that includes one or more SIMD units 138 configured to perform operations in parallel according to the SIMD paradigm at the request of the processor 102 (or another unit). The SIMD paradigm is one in which multiple processing elements share a single program control flow unit and program counter and thus execute the same program, but can execute that program with different data. In one example, each SIMD unit 138 includes 16 lanes, and each lane can execute the same instruction simultaneously with other lanes within the SIMD unit 138, but can execute that instruction with different data. The lanes can be switched off predictively if not all lanes need to execute a given instruction. Also, prediction can be used to execute programs having branch control flow. More specifically, for programs having conditional branches or other instructions where the control flow is based on calculations performed by individual lanes, prediction of lanes corresponding to currently unexecuted control flow paths, and serial execution of different control flow paths, enables any control flow.
[0013] The basic unit of execution within the computing unit 132 is the work item. Each work item represents a single instantiation of a program that is executed in parallel within a particular lane. Work items can be executed simultaneously (or partially simultaneously and partially sequentially) as a "wavefront" on a single SIMD processing unit 138. One or more wavefronts are included in a "work group", which includes a collection of work items that are specified to execute the same program. A work group can be executed by executing each of the wavefronts that make up the work group. In an alternative example, wavefronts are executed on a single SIMD unit 138 or on different SIMD units 138. A wavefront can be thought of as the largest collection of work items that can be executed simultaneously (or pseudo-simultaneously) on a single SIMD unit 138. "Pseudo-simultaneous" execution occurs in the case of a wavefront that is larger than the number of lanes within the SIMD unit 138. In such a situation, the wavefront is executed over multiple cycles, with different collections of work items being executed in different cycles. The APD scheduler 136 is configured to perform operations related to the scheduling of various work groups and wavefronts on the computing unit 132 and the SIMD unit 138.
[0014] The parallel processing provided by the computing unit 132 is suitable for graphics-related operations such as pixel value calculation, vertex transformation, and other graphics operations. Thus, in some cases, a graphics pipeline 134 that receives graphics processing commands from the processor 102 provides compute tasks to the computing unit 132 for parallel execution.
[0015] Further, the computing unit 132 is used to perform computational tasks that are not related to graphics or are not part of the "normal" operation of the graphics pipeline 134 (e.g., custom operations performed to supplement the processing done on the operation of the graphics pipeline 134). An application 126 or other software executed on the processor 102 transmits a program that defines such a computational task to the APD 116 for execution.
[0016] The APD 116 includes one or more memory management units (MMUs) 150. The MMU processes memory access requests such as requests for translation of virtual addresses to physical addresses. In various embodiments, the MMU 150 includes one or more translation lookaside buffers (TLBs) or an interface to one or more TLBs. The TLB caches virtual - physical address translations for rapid reference.
[0017] The computing unit 132 performs ray tracing, which is a technique for rendering a 3D scene by testing for intersections between simulated rays and objects in the scene. Much of the work involved in ray tracing is done by programmable shader programs executed on the SIMD unit 138 within the computing unit 132, as will be described in more detail below.
[0018] FIG. 3 is a diagram showing a ray tracing pipeline 300 for rendering graphics using ray tracing technology according to an example. The ray tracing pipeline 300 provides an overview of the operations and entities involved in rendering a scene using ray tracing. In some embodiments, the ray generation shader 302, any hit shader 306, intersection shader 307, closest hit shader 310, and miss shader 312 are shader implementation stages that represent stages of the ray tracing pipeline whose functions are performed by shader programs executed within the SIMD unit 138. Any of the specific shader programs at each specific shader implementation stage are defined by application-provided code (i.e., code provided by the application developer, pre-compiled by the application compiler and / or compiled by the driver 122). In other embodiments, any of the ray generation shader 302, any hit shader 306, closest hit shader 310, and miss shader 312 are implemented as software that runs on any type of processor, circuitry that performs the operations described herein, or a combination of hardware circuitry and software running on a processor. The acceleration structure traversal stage 304 performs a ray intersection test to determine whether the ray hits a triangle.
[0019] The ray tracing pipeline 300 means the path through which the ray tracing operation flows. To render a scene using ray tracing, a rendering orchestrator such as a program executed on the processor 102 specifies an aggregate of geometries as a "scene". Various objects in the scene are represented as an aggregate of geometric primitives, which are often triangles but can be any geometric shape. As used herein, the term "triangle" refers to these geometric primitives that make up the scene. The rendering orchestrator renders the scene by specifying the camera position and the image, and by requiring that rays be traced from the camera through the image. The ray tracing pipeline 300 performs various operations described herein to determine the color of the rays. The color is often derived from the triangles that the rays intersect. As described elsewhere in this specification, rays that do not hit a triangle call the miss shader 312. One possible operation of the miss shader 312 is to color the ray with the color from the "skybox", which is an image specified as representing the surrounding scene where there is no geometry (e.g., a scene without geometry renders only the skybox). The color of a pixel in the image is determined based on the intersection point between the ray and the image position. In some examples, after a sufficient number of rays have been traced and the pixels of the image have been assigned colors, the image is displayed on the screen or used in some other manner.
[0020] In some embodiments where the shader stages of the ray tracing pipeline 300 are implemented in software, other programmable shader stages (ray generation shader 302, any hit shader 306, closest hit shader 310, miss shader 312) are implemented as shader programs executed on the SIMD unit 138. The acceleration structure traversal stage is implemented in software (e.g., as a shader program executed on the SIMD unit 138), in hardware, or as a combination of hardware and software. The ray tracing pipeline 300 is, in various embodiments, configured partially or fully in software or partially or fully in hardware and, in various embodiments, is configured by the processor 102, the scheduler 136, by a combination thereof, or by any other hardware and / or software unit partially or fully. In an example, traversal through the ray tracing pipeline
[0021] The ray tracing pipeline 300 operates in the following manner. Ray generation shader 302 is performed. Ray generation shader 302 sets the data of the rays to be tested for a triangle and requests an acceleration structure traversal stage 304 to test the rays for intersection with the triangle.
[0022] Acceleration structure traversal stage 304 traverses an acceleration structure, which is a data structure that describes the scene volume and the objects in the scene, and tests the rays for the triangles in the scene. During this traversal, for the triangles that intersect with the ray, ray tracing pipeline 300 triggers the execution of any hit shader 306 and / or intersection shader 307 if specified by the material of the triangle that the shader intersects. Note that multiple triangles can be intersected by a single ray. It is not guaranteed that the acceleration structure traversal stage traverses the acceleration structure in the order from the one closest to the ray origin to the one farthest from the ray origin. Acceleration structure traversal stage 304 triggers the execution of the closest hit shader 310 for the triangle closest to the origin of the ray that the ray hits, or triggers a miss shader if the triangle is not hit.
[0023] Any hit shader 306 or intersection shader 307 can "reject" an intersection from the acceleration structure traversal stage 304. Thus, it should be noted that the acceleration structure traversal stage 304 triggers the execution of the miss shader 312 when no intersection with a ray is found, or when one or more intersections are found but all are rejected by any hit shader 306 and / or intersection shader 307. An exemplary situation where any hit shader 306 "rejects" a hit is when at least a part of the triangle reported when the acceleration structure traversal stage 304 hits is completely transparent. Since the acceleration structure traversal stage 304 only tests the geometry and not the transparency, any hit shader 306 called due to an intersection with a triangle having at least some transparency may determine that the reported intersection should not be counted as a hit because it "intersects" the transparent part of the triangle. A typical use of the closest hit shader 310 is to color a ray based on the texture of the material. A typical use of the miss shader 312 is to color a ray with the color set by the skybox. In various embodiments, it should be understood that the closest hit shader 310 and the miss shader 312 implement a wide variety of techniques for coloring rays and / or performing other operations. When these shaders are implemented as programmable shader stages that execute shader programs, different shader programs used in the same application can color pixels in different ways. The term "hit shader" is sometimes used herein to refer to one or more of any hit shader 306, intersection shader 307, and closest hit shader 310.
[0024] A typical way for the ray generation shader 302 to generate rays is to use a technique called backwards ray tracing. In backwards ray tracing, the ray generation shader 302 generates a ray with a starting point at the point of the camera. The point where the ray intersects the plane defined to correspond to the screen defines the pixel on the screen that the ray uses to determine its color. If the ray hits an object, the pixel is shaded based on the closest hit shader 310. If the ray does not hit an object, the pixel is shaded based on the miss shader 312. Multiple rays can be cast for each pixel, and the final color of the pixel is determined by some combination of the colors determined for each of the rays of the pixel.
[0025] Any one of the optional hit shader 306, intersection shader 307, closest hit shader 310, and miss shader 312 can cause a unique ray to enter the ray tracing pipeline 300 at the ray test point. These rays can be used for any purpose. One common use is to implement environmental lighting or reflections. In one example, when the closest hit shader 310 is called, the closest hit shader 310 causes rays in various directions. For each object or light hit by the generated rays, the closest hit shader 310 adds lighting intensity and color to the pixel corresponding to the closest hit shader 310. Although some examples of ways to render a scene using various components of the ray tracing pipeline 300 are described, it should be understood that any of a wide variety of techniques can be used alternatively.
[0026] As described above, the determination of whether a ray intersects an object is referred to in this specification as a "ray intersection test". The ray intersection test involves emitting a ray from a starting point, determining whether the ray intersects a triangle, and if so, determining how far from the starting point of the triangle intersection it is. To increase efficiency, a ray tracing test uses a representation of space called a bounding volume hierarchy. This bounding volume hierarchy is the "acceleration structure" referred to elsewhere in this specification. In the bounding volume hierarchy, each non-leaf node represents an axis-aligned bounding box that bounds the geometry of all of its children. In one example, the base node represents the maximum extent of the entire region where the ray intersection test is being performed. In this example, the base node has two children, each of which represents an axis-aligned bounding box that mutually exclusively sub-divides the entire region. Each of those two children has two child nodes that represent axis-aligned bounding boxes that sub-divide the space of their parent, and so on. A leaf node represents a triangle for which a ray intersection test can be performed. A non-leaf node may be referred to in this specification as a "box node", and a leaf node may be referred to in this specification as a "triangle node".
[0027] The bounding volume hierarchy data structure enables reducing the number of ray-triangle intersections (which are complex and thus expensive in terms of processing resources) compared to a scenario where all triangles in a scene need to be tested against a ray because such a data structure is not used. Specifically, if a ray does not intersect a particular bounding box and that bounding box bounds a number of triangles, all triangles within that box can be excluded from the test. Thus, the ray intersection test is performed as a series of tests of the ray against axis-aligned bounding boxes, followed by tests against triangles.
[0028] FIG. 4 is a diagram showing a bounding volume hierarchy by way of an example. For simplicity, the hierarchy is shown in 2D. However, it should be understood that the extension to 3D is straightforward and the tests described in this specification are generally performed in three dimensions.
[0029] The spatial representation 402 of the boundary volume hierarchy is shown on the left side of FIG. 4, and the tree representation 404 of the boundary volume hierarchy is shown on the right side of FIG. 4. In both the spatial representation 402 and the tree representation 404, non-leaf nodes are represented by the letter "N", and leaf nodes are represented by the letter "O". The ray intersection test is performed by traversing through the tree 404, and for each leaf node tested, if the test for its non-leaf node fails, the branch below that node is excluded. In one example, the ray intersects O5 but does not intersect the other triangles. The test is performed on N1 and it is determined that the test is successful. The test is performed on N2 and it is determined that the test fails (since O5 is not within N1). Note that this test excludes all sub-nodes of N2, performs a test on N3, and it is noted that the test is successful. Note that this test tests N6 and N7, N6 is successful but N7 fails. Note that this test tests O5 and O6, O5 is successful but O6 fails. Instead of testing eight triangle tests, two triangle tests (O5 and O6) and five box tests (N1, N2, N3, N6, N7) are performed.
[0030] The ray tracing pipeline 300 emits rays and detects whether the rays hit a triangle and how such a hit is shaded. Each triangle has a material assigned, and the material specifies which closest hit shader is executed for that triangle at the closest hit shader stage 310, and whether any hit shader is executed at any hit shader stage 306, whether the intersection shader is executed at the intersection shader stage 307, and the specific any hit shader and intersection shader executed at those stages if those shaders are executed.
[0031] Thus, when emitting a ray, the ray tracing pipeline 300 evaluates the intersections detected at the acceleration structure traversal stage 304 as follows. If it is determined that the ray intersects a triangle and the material of that triangle has at least any hit shader or intersection shader, the ray tracing pipeline 300 executes the intersection shader and / or any hit shader to determine whether the intersection should be considered a hit or a miss. If neither any hit shader nor intersection shader is specified for a particular material, the intersection with the triangle having that material reported by the acceleration structure traversal 304 is considered a hit.
[0032] Some examples of situations where any hit shader or intersection shader does not count an intersection as a hit are provided here. In one example, if the alpha is 0, it means that the point where the ray intersects the triangle is completely transparent, and any hit shader will consider such an intersection not a hit. In another example, any hit shader determines that the point where the ray intersects the triangle is considered to be in the "cutout" portion of the triangle (the cutout "cuts out" portions of the triangle by designating those portions as parts where the ray cannot hit), and thus considers that intersection not a hit.
[0033] When the acceleration structure has been completely traversed, the ray tracing pipeline 300 executes the closest hit shader 310 on the closest triangle determined to be hit by the ray. Similar to any hit shader 306 and intersection shader 307, the closest hit shader 310 to be executed for a particular triangle depends on the material assigned to that triangle.
[0034] In short, the ray tracing pipeline 300 traverses the acceleration structure 304 to determine which triangle is the closest hit to a given ray. Any hit shader and intersection shader evaluate the intersection (potential hit) to determine whether those intersections are counted as actual hits. Then, for the closest triangle whose intersection is counted as an actual hit, the ray tracing pipeline 300 executes the closest hit shader of that triangle. If there is no triangle to be counted as a hit, the ray tracing pipeline 300 executes the miss shader for the ray.
[0035] Here, the operation of the ray tracing pipeline 300 is considered with respect to the exemplary rays 1 to 4 shown in FIG. 4. For each of the exemplary rays 1 to 4, the ray tracing pipeline 300 determines which triangles those rays intersect. The ray tracing pipeline 300 executes any appropriate hit shader 306 and intersection shader 307 to determine the closest non-missing hit (and thus the closest hit triangle) as specified by the material of the intersecting triangle. The ray tracing pipeline 300 executes the closest hit shader for the closest hit triangle.
[0036] In one example, for ray 1, the ray tracing pipeline 300, when executed, executes the closest hit shader to O4 as long as the triangle does not have any hit shader or intersection shader that indicates that ray 1 did not hit the triangle. In that situation, the ray tracing pipeline 300 executes the closest hit shader to O1 as long as the triangle does not have any hit shader or intersection shader that indicates that the triangle was not hit by ray 1, and in that situation, the ray tracing pipeline 300 executes the miss shader 312 for ray 1. Similar operations occur for rays 2, 3, and 4. For ray 2, the ray tracing pipeline 300 determines that intersections occur with O2 and O4 and, if specified by the material, executes any hit and / or intersection shaders for those triangles and executes the appropriate closest hit or miss shader. For rays 3 and 4, the ray tracing pipeline 300 determines the intersections as shown (ray 3 intersects O3 and O7 and ray 4 intersects O5 and O6), executes the appropriate any hit and / or intersection shaders, and executes the appropriate closest hit or miss shader based on the results of the any hit and / or intersection shaders.
[0037] Bounding volume hierarchies such as BVH404 include data defining various nodes including leaf nodes and non-leaf nodes, and associated information such as the geometry of the boxes associated with non-leaf nodes, the geometry of the triangles associated with leaf nodes, and other information. The amount of data within BVH404 can span multiple memory pages, such as in the case of a very large BVH that holds the geometry of a very large scene. In one example, a video game application includes one or more "levels" that include geometry such as terrain, props, and other geometry. In such an example, the BVH404 for the entire level is calculated "offline," meaning at application development time and not during runtime. This action removes the need to recalculate the BVH404 when the player's character traverses the level. However, the amount of data in BVH404 is very large.
[0038] Because a large BVH404 is used, it is possible that not all of BVH404 is stored in memory that is readily accessible, such as cache, APD memory, or other memory at any given time. Thus, accessing a particular portion of BVH404 can sometimes result in an unacceptable stall during execution, for example, when the application waits for the accessed portion of BVH404 to become available before proceeding with other work. For the reasons above, techniques are provided herein to facilitate the processing of BVH404 that has memory pages that are not readily available upon access.
[0039] FIG. 5 is a diagram showing a BVH 500 having different BVH memory pages 502 according to an example. The BVH memory pages 502 are shown as being resident, valid and non-resident, or invalid. A valid BVH memory page 502 is a BVH memory page 502 whose virtual address of the memory page has a valid physical address translation. Also, a valid BVH memory page 502 is regarded as a resident BVH memory page 502 (and a resident BVH memory page 502 is regarded as a valid memory page). A valid, non-resident memory page is a BVH memory page 502 whose virtual memory address has a valid physical address translation but is not regarded as resident. An invalid BVH memory page 502 is a BVH memory page having a virtual memory page address that does not have a valid address translation. Nodes within the BVH 500 point to other nodes using virtual memory addresses (thus, a node includes a pointer to another node). It is possible for a pointer of a node to have an invalid address, and the translation to a physical address does not exist in the page table. This pointer is for an invalid BVH memory page 502. In FIG. 5, such an invalid memory page 502 is shown as empty because no data actually exists for that memory page.
[0040] A resident memory page is a memory page in which data is stored in memory that is considered to be immediately accessible. The specific memory considered to be immediately accessible varies in different implementations. In one example, a specific cache memory such as a level 0 cache memory is considered to be "immediately accessible", and thus, the BVH memory page 502 stored in the level 0 cache memory is considered to be resident. In another example, the APD system memory is considered to be "immediately accessible", and thus, the BVH memory page 502 stored in the APD system memory (and all memory "closer" to the computing unit 132) is considered to be resident. A BVH memory page 502 not stored in any immediately accessible memory is considered to be non-resident (or not resident). The APD system memory is within the APD 116 and is memory available for use by any of the computing units 132. An application can specify which memory pages are considered to be resident and which are considered to be valid and non-resident. The data of a non-resident memory page can be in a format incompatible with the boundary volume hierarchy. In an example, an application that reads such data generates a BVH portion from that data and loads that BVH portion into memory considered to be resident. The application then marks the page containing that data as resident.
[0041] Different designations of the BVH memory page 502 as resident, valid and non-resident, or invalid (referred to as "residency designations") enable traversal of the BVH without waiting to load the BVH into memory that is considered to be immediately accessible even if the BVH 500 is not completely resident. More specifically, when a BVH traversal entity such as the computing unit 132 executes a shader program to traverse the BVH 500, the BVH traversal entity traverses using the virtual address of the BVH memory page 502. The traversal entity provides such a virtual address to the MMU 150 for translation. The MMU 150 examines the stored translations (such as within the TLB and / or within one or more page tables), determines the physical address of the page and the residency designation of the BVH memory page 502, and returns these values to the traversal entity.
[0042] For a BVH memory page 502 that is resident, the traversal entity processes the contents of such a memory page 502 as normal (i.e., traverses the box nodes that the ray intersects until one or more triangles are found and performs an intersection test on such one or more triangles as described with respect to FIG. 4).
[0043] For a BVH memory page 502 that is valid and non-resident, the traversal entity processes the contents of such a BVH memory page 502 as if an error had occurred for all such contents. In one example, the BVH memory page 502 includes box nodes but no triangle nodes. In such an example, the traversal entity processes all such box nodes as if the ray had missed them. Thus, the traversal entity does not traverse any children of such box nodes and does not record a hit to any of the triangles that are the ultimate children of such box nodes even if the ray actually hits a node and data is resident such that an intersection test for such a triangle could be performed.
[0044] For an invalid BVH memory page 502, the MMU 150 generates a fault to be processed by a fault handler (such as an operating system running within the processor 102). Such a fault indicates that a virtual address referencing a particular BVH memory page 502 does not reference a valid memory page, and thus that the BVH 500 contains an invalid memory address.
[0045] Instead of waiting for the content of such a memory page to be loaded into immediately accessible memory, the operation involving ray tracing can proceed by treating the content of a valid non-resident BVH memory page 502 as a miss. The triangles represented by the non-resident portion of the BVH 500 are simply not displayed.
[0046] FIG. 6 is a diagram showing an exemplary system in which the technology of the present disclosure is implemented. FIG. 7 is a flowchart of a method 700 for performing ray tracing using a partially resident bounding volume hierarchy according to an example. Although described with respect to the systems of FIGS. 1 - 6, those skilled in the art will understand that any system configured to perform the steps of method 700 in any technically feasible order is within the scope of the present disclosure. Here, FIGS. 6 and 7 will be described together.
[0047] FIG. 6 is a block diagram of a system 600 for performing ray tracing using a partially resident BVH according to an example. The system 600 includes a BVH traversal unit 602, a memory management unit (MMU) 604, a memory 606 regarded as "immediately accessible" including resident BVH pages 608, a memory 610 regarded as "not immediately accessible" including valid non-resident BVH pages 612, a translation lookaside buffer 614, and one or more page tables 616.
[0048] The BVH traversal unit 602 is an entity that performs a ray intersection test. In various examples, the BVH traversal unit 602 is a shader program executed on the computing unit 132, an application executed on the processor 102, or any other entity such as a program executed on a processor, a hardware circuit or a combination of software and hardware configured to execute a ray intersection test.
[0049] In some examples, the MMU 604 is the MMU 150 of the APD 116. In other examples, the MMU 604 is a different MMU. The MMU 604 provides an address translation service and translates virtual addresses to physical addresses. Also, the MMU 604 indicates to the BVH traversal unit 602 whether the BVH memory page 502 is resident or non-resident but valid. In some configurations, the MMU 604 also indicates whether the BVH memory page 502 is invalid.
[0050] The BVH traversal unit 602 is communicatively coupled to the readily accessible memory 606. As described elsewhere in this specification, the readily accessible memory 606 is a memory that is considered to store the resident BVH pages 608. In contrast, the non-readily accessible memory 610 is a memory that is considered to store the valid non-resident memory pages 612. The data of the valid non-resident memory pages 612 can be in a form that is not immediately suitable for use as part of the BVH. The data corresponding to the valid non-resident memory pages 612 can be in the same memory as the resident memory pages 608, but the application can nevertheless indicate that the data corresponding to the non-resident memory pages 612 is non-resident. In one example, the application stores raw geometries (e.g., triangles) in the system memory along with the resident memory pages 612 of the BVH. At this point, the memory pages corresponding to the raw geometries are indicated as valid but non-resident in the page table. The application processes the raw geometries to generate a portion of the BVH and indicates to the operating system that the portion of the BVH is currently within the resident memory pages. The operating system then changes the page table to indicate that those memory pages are resident instead of being valid and non-resident.
[0051] In various embodiments, information regarding the classification of BVH memory page 502, shown as page status 618, is stored in page table 616 that is read into TLB 614 for use by MMU 604. An entity that writes page table 616, such as an operating system running on processor 102, writes this page status 618 into page table 616. In various embodiments, information indicating which memory is considered immediately accessible and which memory is not considered immediately accessible is stored for reference by an entity that writes page table 616 (such as an operating system running on processor 102). The entity references the location when the memory page migrates between memories and updates the page status 618. In some examples, the application indicates to the operating system which memory pages are resident and which memory pages are valid and non-resident.
[0052] Now refer to FIGS. 6 and 7 together. Method 700 begins at step 702, and BVH traversal unit 602 traverses BVH 500. While traversing BVH 500, BVH traversal unit 602 encounters a first BVH memory page 502 classified as resident. At step 704, BVH traversal unit 602 obtains the portion of BVH 500 within the first BVH memory page 502 and traverses the portion of BVH 500 represented in the first BVH memory page.
[0053] In step 706, while traversing the BVH 500, the traversal unit 602 encounters a second memory page that is valid but classified as non-resident. In step 708, the BVH traversal unit 602 processes the geometries within the second memory page as if the ray had missed its geometry. For a box node, the BVH traversal unit 602 handles each triangle that is a child of such a box node as if a miss had occurred. Specifically, the BVH traversal unit 602 does not execute a hit shader on that triangle. For a triangle, the BVH traversal unit 602 processes such a triangle as if the ray had missed it.
[0054] An exemplary traversal through the BVH 500 of FIG. 5 is now described using the method 700 of FIG. 7. In this example, a ray is tested against the BVH 500 to intersect a triangle. The ray intersects triangle O3 but not the other triangles. The BVH traversal unit 602 starts at the root node N1. Since N1 is in the resident page, the BVH traversal unit 602 fetches the data for N1, performs an intersection test, determines that the ray intersects the space associated with N1, and proceeds to the intersection tests for the children N2 and N3 of N1, where N2 and N3 are in the resident memory page. The BVH traversal unit 602 determines that the ray intersects the box node N2 but not the box node N3. Since a miss occurs for N3, the BVH traversal unit 602 does not proceed to the children of N3. However, since a hit occurs for N2, the BVH traversal unit 602 proceeds to nodes N4 and N5.
[0055] Nodes N4 and N5 are in a different BVH memory page 502(2) from nodes N1, N2, and N3. However, BVH memory page 502(2) is resident. Therefore, the BVH traversal unit 602 can access the data of N4 and N5 normally and proceed through the BVH 500 from that point. More specifically, node N2 stores the locations of nodes N4 and N5 using a pointer (a memory address within the virtual address space). The BVH traversal unit 602 provides this memory address to the MMU 604 for translation. The MMU 604 returns the physical addresses of nodes N4 and N5, as well as an indication that these nodes are in the resident memory page 502(2). Since these nodes are in the resident memory page, the BVH traversal unit 602 evaluates the rays against these nodes instead of treating them as misses.
[0056] The BVH traversal unit 602 evaluates the ray against node N4 and determines that there is no intersection. The BVH traversal unit 602 evaluates the ray against node N5 and determines that there is an intersection. Therefore, the BVH traversal unit 602 attempts to access the children of N5, which are the triangle nodes O3 and O4 in a different BVH memory page 502(5) from the BVH memory page 502(2) of N4 and N5. The BVH traversal unit 602 provides the memory addresses of O3 and O4 to the MMU 604, which returns an indication that the address pointing to memory page 502(5) is a valid address but that memory page 502(5) is non-resident. In response to the indication that memory page 502(5) is non-resident, the BVH traversal unit 602 treats both triangles O3 and O4 as misses. Since no other triangles have been hit in the BVH, this results in the miss shader being executed as described with respect to Figure 4.
[0057] Examples have been described where the miss shader is executed because no triangle has been hit, but other triangles can be hit by the ray. More specifically, even if some of the triangles intersected by the ray are in non-resident memory pages, other triangles are sometimes in resident memory pages. If those triangles are intersected by the ray, the hit shader is executed for at least one of those triangles, and in some cases, the miss shader is not executed.
[0058] Each of the units shown in the figures represents a hardware circuit configured to perform the operations described herein, software configured to perform the operations described herein, or a combination of software and hardware configured to perform the steps described herein. For example, the acceleration structure traversal stage 304 is implemented entirely in hardware, entirely in software executed on a processing unit (such as the computing unit 132), or as a combination thereof. In some examples, the acceleration structure traversal stage 304 is implemented partially in hardware and partially in software. In some examples, the portion of the acceleration structure traversal stage 304 that traverses the bounding volume hierarchy is software executed on a processor, and the portion of the acceleration structure traversal stage 304 that performs the ray-box intersection test and the ray-triangle intersection test is implemented in hardware. When a particular stage of the ray tracing pipeline 300 is said to be "called", this call involves performing the function of the hardware if the stage is implemented as a hardware circuit, or executing a shader program (or other software) if the stage is implemented as a shader program executed on a processor.
[0059] It should be understood that many variations are possible based on the disclosure herein. Although features and elements have been described above in specific combinations, each feature or element can be used alone without the other features and elements, or in various combinations with or without the other features and elements.
[0060] The provided method can be implemented on a general-purpose computer, processor, or processor core. Suitable processors include, by way of example, general-purpose processors, dedicated processors, conventional processors, digital signal processors (DSPs), multiple microprocessors, one or more microprocessors associated with a DSP core, controllers, microcontrollers, application specific integrated circuits (ASICs), field programmable gate array (FPGA) circuits, any other type of integrated circuit (IC), and / or state machines. Such processors can be manufactured by configuring the manufacturing process using the results of processed hardware description language (HDL) instructions and other intermediate data including netlists (such instructions can be stored on a computer-readable medium). The result of such processing may be a mask work, which is then used in a subsequent semiconductor manufacturing process to manufacture a processor implementing aspects of the embodiments.
[0061] The methods or flowcharts provided herein may be implemented in a computer program, software, or firmware incorporated in a non-transitory computer-readable storage medium for execution by a general-purpose computer or processor. Examples of non-transitory computer-readable storage media include read only memory (ROM), random access memory (RAM), registers, cache memory, semiconductor memory devices, magnetic media (e.g., internal hard disks and removable disks), magneto-optical media, and optical media (e.g., CD-ROM disks and digital versatile disks (DVDs)).
Claims
Claim 1 A method for performing ray tracing of light rays, comprising: identifying a first memory page classified as resident according to a page table based on a first traversal of a boundary volume hierarchy; obtaining a first portion of the boundary volume hierarchy associated with the first memory page; traversing the first portion of the boundary volume hierarchy according to a ray intersection test; identifying a second memory page classified as valid and non-resident according to the page table based on a second traversal of the boundary volume hierarchy; processing each node of the boundary volume hierarchy in the second memory page as if an error had occurred in response to the second memory page being classified as valid and non-resident. A method. Claim 2 The first traversal of the boundary volume hierarchy comprises: determining that the ray intersects a first parent node that is a parent of one or more nodes of the first memory page; obtaining a page address for the one or more nodes from data of the first parent node. The method of Claim 1. Claim 3 Identifying the first memory page classified as resident comprises: determining that the page address is indicated as resident according to the page table. The method of Claim 2. Claim 4 The second traversal of the boundary volume hierarchy comprises: determining that the ray intersects a second parent node that is a parent of one or more nodes of the second memory page; obtaining a page address for the one or more nodes from data of the second parent node. The method of Claim 1. Claim 5 Identifying the second memory page classified as valid and non-resident comprises: determining that the page address is indicated as valid and non-resident according to the page table. The method of Claim 4. Claim 6 further comprising: identifying a third memory page classified as invalid based on a third traversal of the boundary volume hierarchy; generating a fault for the third memory page. The method of Claim 1. Claim 7 In response to determining that an error has occurred for each node of the boundary volume hierarchy in the second memory page, further comprising not executing a hit shader for any triangle node that is a child of any node in the second memory page The method of claim 1
8. In response to determining that processing has been performed as if an error has occurred for each node of the boundary volume hierarchy in the second memory page, further comprising executing a miss shader for the ray The method of claim 1
9. Further comprising updating the status of the memory page of the boundary volume hierarchy based on migration of the memory page The method of claim 1
10. A system for performing ray tracing of a ray, comprising A memory for storing a memory page of a boundary volume hierarchy A processor, and comprising The processor Identifying a first memory page classified as resident according to a page table based on a first traversal of a boundary volume hierarchy Obtaining a first portion of the boundary volume hierarchy associated with the first memory page Traversing the first portion of the boundary volume hierarchy according to a ray intersection test Identifying a second memory page classified as valid and non-resident according to the page table based on a second traversal of the boundary volume hierarchy In response to the second memory page being classified as valid and non-resident, processing as if an error has occurred for each node of the boundary volume hierarchy in the second memory page Is configured to perform System
11. The first traversal of the boundary volume hierarchy Determining that the ray intersects a first parent node that is a parent of one or more nodes of the first memory page Obtaining a page address for the one or more nodes from the data of the first parent node, and comprising The system of claim 10
12. Identifying the first memory page classified as resident Including determining that the page address is indicated as resident according to the page table The system of claim 11
13. The second traversal of the boundary volume hierarchy determining that the ray intersects a second parent node that is a parent of one or more nodes of the second memory page; obtaining page addresses for the one or more nodes from data of the second parent node; The system of claim 10. **Claim 14** Identifying the second memory page classified as valid and non-resident includes determining that the page address is indicated as valid and non-resident according to the page table. The system of claim 13. **Claim 15** The processor is identifying a third memory page classified as invalid based on a third traversal of the boundary volume hierarchy; generating a fault for the third memory page; configured to perform The system of claim 10. **Claim 16** The processor is configured not to execute a hit shader on any triangle node that is a child of any node in the second memory page in response to determining that a miss has occurred for each node in the boundary volume hierarchy within the second memory page. The system of claim 10. **Claim 17** The processor is configured to execute a miss shader on the ray in response to determining that processing is performed as if a miss has occurred for each node in the boundary volume hierarchy within the second memory page. The system of claim 10. **Claim 18** The processor is configured to update the status of the memory pages of the boundary volume hierarchy based on migration of the memory pages. The system of claim 10. **Claim 19** A computer-readable storage medium storing instructions that, when executed by a processor, identify a first memory page classified as resident according to a page table based on a first traversal of a boundary volume hierarchy; obtain a first portion of the boundary volume hierarchy associated with the first memory page; traverse the first portion of the boundary volume hierarchy according to a ray intersection test; identify a second memory page classified as valid and non-resident according to the page table based on a second traversal of the boundary volume hierarchy; In response to the second memory page being classified as valid and non-resident, process each node of the boundary volume hierarchy within the second memory page as if an error had occurred; Cause the processor to perform ray tracing of light rays thereby; A computer-readable storage medium. **Claim 20** The first traverse of the boundary volume hierarchy Determine that the ray intersects a first parent node that is a parent of one or more nodes of the first memory page; Obtain page addresses for the one or more nodes from data of the first parent node; The computer-readable storage medium of claim 19.
Citation Information
Patent Citations
Conditional page fault control for page residency
JP2016534486A
Method for efficient grouping of cache requests for datapath scheduling
US20200050550A1