Techniques for generating bounding volume hierarchy

By predefined triangle set sets and centroid boxes, the time-consuming problem of building BVH from top to bottom is solved, the construction efficiency of the enclosing body hierarchy is improved, and the performance of ray tracing rendering is improved.

CN120345002APending Publication Date: 2025-07-18ADVANCED MICRO DEVICES INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202380084990.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2022-12-16
Filing Date
2023-11-27
Publication Date
2025-07-18

AI Technical Summary

Technical Problem

In the prior art, when building a hierarchy of enclosures, the top-down method takes a long time, resulting in inefficient ray tracing rendering process.

Method used

By predefined sets of triangles and centroid boxes at different levels of detail, the amount of calculations when building the surrounding body hierarchy is reduced, the resolution subdivision method is used to generate candidate splits, and the distribution of triangles is determined through the centroid box, which reduces the time complexity of building BVH.

Benefits of technology

It improves the construction efficiency of the enclosing body hierarchy, reduces the calculation time during the ray tracing rendering process, and improves rendering performance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120345002A_ABST
    Figure CN120345002A_ABST
Patent Text Reader

Abstract

A technique for constructing a bounding volume hierarchy is disclosed. The technique subdivides the candidate box node based on resolution to generate a plurality of cells of the candidate box node; identifying a plurality of nodes of the set of triangles that fit within the unit; generating a plurality of candidate splits based on the plurality of nodes; selecting the candidate split based on a selection criterion to obtain a selected candidate split; and generating a sub-box node for a box node of the bounding volume hierarchy being constructed based on the selected candidate split.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Cross - Reference to Related Applications

[0002] This application claims priority to U.S. Non - Provisional Patent Application No. 18 / 083,298, filed on December 16, 2022, the entire content of which is incorporated herein by reference. Background Art

[0003] In image synthesis, ray tracing is used to find the closest intersection of a given ray with a scene that models the propagation of light. Brief Description of the Drawings

[0004] A more detailed understanding can be obtained from the following description, given by way of example in conjunction with the accompanying drawings, in which:

[0005] Figure 1 is a block diagram of an example device that can implement one or more features of the present disclosure;

[0006] Figure 2 is Figure 1 a block diagram of the device in, illustrating additional details according to one example;

[0007] Figure 3 illustrates a ray - tracing pipeline for rendering graphics using ray - tracing techniques according to one example;

[0008] Figure 4 illustrates a bounding volume hierarchy (″BVH″) according to one example;

[0009] Figure 5 illustrates the generation of a BVH from scene geometry by a BVH constructor according to one example; Figure 6A illustrates aspects of generating a BVH using a top - down technique according to one example; Figure 6B illustrates a set of candidate splits for a BVH node according to one example;

[0010] Figure 7 illustrates a set of example primitive sets;

[0011] Figure 8 illustrates different partition resolutions according to one example;

[0012] Figure 9 illustrates an example of a set of box nodes that find a set of triangle sets of all primitives corresponding to descendants of a candidate box node of a BVH being constructed;

[0013] Figures 9A to 9C illustrates different candidate splits according to the example;

[0014] Figure 10Illustrates additional operations related to constructing a BVH according to an example; and

[0015] Figure 11 is a flowchart of a method for constructing a BVH according to an example. Detailed implementation

[0016] Disclosed is a technique for constructing a bounding volume hierarchy. The technique subdivides candidate box nodes based on a resolution to generate a plurality of cells of the candidate box nodes; identifies a plurality of nodes of a primitive set suitable within the cell; generates a plurality of candidate splits based on the plurality of nodes; selects a candidate split based on a selection criterion to obtain a selected candidate split; and generates child box nodes for a box node of the bounding volume hierarchy being constructed based on the selected candidate split.

[0017] Figure 1 is a block diagram of an example device 100 that can implement one or more features of the present disclosure. The device 100 can include, for example, a computer, a gaming device, a handheld device, a set-top box, a television, a mobile phone, or a tablet computer. The device 100 includes a processor 102, a memory 104, a storage device 106, one or more input devices 108, and one or more output devices 110. The device 100 can also optionally include an input driver 112 and an output driver 114. It should be understood that the device 100 can include Figure 1 additional components not shown in

[0018] In various alternatives, the processor 102 includes a central processing unit (CPU), a graphics processing unit (GPU), a CPU and a GPU on the same die, or one or more processor cores, where each processor core can be a CPU or a GPU. In various alternatives, the memory 104 is located on the same die as the processor 102 or is located separately from the processor 102. The memory 104 includes volatile or non-volatile memory, such as random access memory (RAM), dynamic RAM, or a cache.

[0019] The storage device 106 includes a fixed or removable storage device, such as a hard disk drive, a solid state drive, an optical disc, or a flash drive. The input device 108 includes, but is not limited to, a keyboard, a keypad, a touch screen, a touchpad, a detector, a microphone, an accelerometer, a gyroscope, a biometric scanner, or a network connection (e.g., a wireless local area network card for transmitting and / or receiving wireless IEEE 802 signals). The output device 110 includes, but is not limited to, a display, a speaker, a printer, a haptic feedback device, one or more lights, an antenna, or a network connection (e.g., a wireless local area network card for transmitting and / or receiving wireless IEEE 802 signals).

[0020] The input driver 112 communicates with the processor 102 and the input device 108 and allows the processor 102 to receive input from the input device 108. The output driver 114 communicates with the processor 102 and the output device 110 and allows the processor 102 to send output to the output device 110. It should be noted that the input driver 112 and the output driver 114 are optional components, and if the input driver 112 and the output driver 114 are absent, the device 100 will operate in the same manner. The output driver 114 includes an acceleration processing device (″APD") 116 coupled to the display device 118. The APD receives compute commands and graphics rendering commands from the processor 102, processes these compute commands and graphics rendering commands, and provides a pixel output to the display device 118 for display. As detailed below, the APD 116 includes one or more parallel processing units that perform computations according to the single instruction multiple data (″SIMD″) paradigm. Thus, although various functions are described herein as being performed by or in conjunction with the APD 116, in various alternative embodiments, the functions described as being performed by the APD 116 are additionally or alternatively performed by other computing devices having similar capabilities that are not driven by the host processor (e.g., the processor 102) and that provide graphics output to the display device 118. For example, any processing system that performs processing tasks according to the SIMD paradigm is contemplated to be capable of performing the functions described herein. Alternatively, a computing system that does not perform processing tasks according to the SIMD paradigm is contemplated to perform the functions described herein.

[0021] Figure 2 is a block diagram of the device 100, illustrating additional details regarding the performance of processing tasks on the APD 116 according to one example. The processor 102 maintains one or more control logic modules in the system memory 104 for execution by the processor 102. The control logic modules include an operating system 120, a driver 122, and an application 126. These control logic modules control various aspects of the operation of the processor 102 and the APD 116. For example, the operating system 120 communicates directly with the hardware and provides an interface to the hardware for other software executing on the processor 102. The driver 122 controls the operation of the APD 116 by, for example, providing an application programming interface (″API″) to software executing on the processor 102 (e.g., the application 126) to access the various functions of the APD 116. The driver 122 also includes a just-in-time compiler that compiles programs for execution by processing components of the APD 116, such as the SIMD unit 138 detailed below.

[0022] The APD 116 executes commands and programs for selected functions, such as graphics operations and non-graphics operations that are suitable for parallel processing. The APD 116 can be used to perform graphics pipeline operations, such as pixel operations, geometric calculations, and presenting an image to the display device 118 based on commands received from the processor 102. The APD 116 also performs computational processing operations not directly related to graphics operations based on commands received from the processor 102, such as operations related to video, physical simulation, computational fluid dynamics, or other tasks.

[0023] The APD 116 includes a compute unit 132 that includes one or more SIMD units 138 that perform operations in parallel according to the SIMD paradigm at the request of the processor 102. The compute unit 132 is sometimes referred to herein as the "parallel processing unit 202". Each compute unit 132 includes a local data share ("LDS") 137 that is accessible to the wavefronts executing in the compute unit 132 but not to the wavefronts executing in other compute units 132. The global memory 139 stores data that is accessible to the wavefronts executing on all compute units 132. In some examples, the local data share 137 has faster access characteristics (e.g., lower latency time and / or higher bandwidth) than the global memory 139. Although shown in the APD 116, the global memory 139 can be partially or fully located in other elements, such as in the system memory 104 or in another memory not shown or described. The SIMD paradigm is a paradigm in which multiple processing elements share a single program control flow unit and program counter and thus execute the same program but are able to execute that program with different data. In one example, each SIMD unit 138 includes sixteen lanes, where each lane executes the same instruction simultaneously with the other lanes in the SIMD unit 138 but can execute that instruction with different data. If not all lanes are required to execute a given instruction, the lanes can be turned off through prediction. Prediction can also be used to execute programs with divergent control flow. More specifically, for programs with conditional branches or other instructions where the control flow is based on calculations performed by a single lane, the lanes corresponding to the control flow paths that are not currently being executed are predicted, and the serial execution of different control flow paths can achieve any control flow.

[0024] The basic execution unit in the compute unit 132 is a work item. Each work item represents a single instantiation of a program to be executed in parallel in a particular lane. Work items can be executed simultaneously as a "wavefront" on a single SIMD processing unit 138. One or more wavefronts are included in a "workgroup", which includes a set of work items designated to execute the same program. A workgroup can be executed by executing each of the wavefronts that make up the workgroup. In an alternative, wavefronts are executed sequentially on a single SIMD unit 138, or partially or fully in parallel on different SIMD units 138. A wavefront can be considered the maximum set of work items that can be executed simultaneously on a single SIMD unit 138. Thus, if a command received from the processor 102 indicates that a particular program is to be parallelized to an extent that the program cannot be executed simultaneously on a single SIMD unit 138, the program is divided into wavefronts that are parallelized on two or more SIMD units 138 or serialized (or parallelized and serialized as needed) on the same SIMD unit 138. The scheduler 136 performs operations involving scheduling various wavefronts on different compute units 132 and SIMD units 138.

[0025] The parallelism provided by the compute unit 132 is suitable for graphics-related operations such as pixel value calculation, vertex transformation, and other graphics operations. Thus, in some instances, a graphics pipeline that receives graphics processing commands from the processor 102 provides compute tasks to the compute unit 132 for parallel execution.

[0026] The compute unit 132 is also used to execute compute tasks that do not involve graphics or are not part of a "normal" operation (e.g., custom operations performed to supplement the processing performed for operations in the graphics pipeline). An application 126 or other software executing on the processor 102 sends a program that defines such compute tasks to the APD 116 for execution.

[0027] The APD 116 is configured to implement the features of the present disclosure by performing a plurality of functions described in more detail below. For example, the APD 116 is configured to receive an image including one or more three-dimensional (3D) objects, divide the image into a plurality of tiles, perform a visibility pass for the primitives of the image, divide the image into tiles, perform a coarse-level tiling for the tiles of the image, divide the tiles into fine tiles, and perform a fine-level tiling of the image. Optionally, front-end geometry processing of the primitives determined to be in the first tile of the tiles can be performed simultaneously with the visibility pass.

[0028] Figure 3Illustrated is a ray tracing pipeline 300 for rendering graphics using ray tracing techniques. The ray tracing pipeline 300 provides an overview of the operations and entities involved in rendering a scene using ray tracing. The ray generation shader 302, the any-hit shader 306, the closest-hit shader 310, and the miss shader 312 are shader implementation levels representing ray tracing pipeline levels, and the functions of the ray tracing pipeline levels are executed by shader programs executed in the SIMD unit 138. Any one of the specific shader programs at each specific shader implementation level is defined by code provided by the application (i.e., code provided by the application developer and pre-compiled by the application compiler and / or compiled by the driver 122). The acceleration structure traversal level 304 performs ray intersection tests to determine whether a ray hits a triangle.

[0029] The various programmable shader levels (ray generation shader 302, any-hit shader 306, closest-hit shader 310, miss shader 312) are implemented as shader programs executed on the SIMD unit 138. The acceleration structure traversal level 304 is implemented in software (e.g., as a shader program executed on the SIMD unit 138), hardware, or a combination of hardware and software. The hit or miss unit 308 is implemented in any technically feasible manner, such as being part of any other unit, being implemented as a hardware acceleration structure, or being implemented as a shader program executed on the SIMD unit 138. The ray tracing pipeline 300 can be coordinated partially or fully in software or partially or fully in hardware and can be coordinated by the processor 102, the scheduler 136, by their combination, or partially or fully by any other hardware and / or software unit. The term "ray tracing pipeline processor" as used herein refers to a processor that executes software to perform the operations of the ray tracing pipeline 300, a hardware circuit hardwired to perform the operations of the ray tracing pipeline 300, or a combination of hardware and software that together perform the operations of the ray tracing pipeline 300.

[0030] The ray tracing pipeline 300 operates as follows. The ray generation shader 302 is executed. The ray generation shader 302 establishes the data of the ray to be tested against the triangle and requests the acceleration structure traversal level 304 to test the intersection of the ray with the triangle.

[0031] Acceleration Structure Traversal Stage 304 traverses an acceleration structure and tests rays against triangles in the scene. The acceleration structure is a data structure that describes the volume of the scene and objects within the scene (such as triangles). In various examples, the acceleration structure is a bounding volume hierarchy. In some specific implementations, a hit or miss unit 308, which is part of the acceleration structure traversal stage 304, determines whether the result of the acceleration structure traversal stage 304 (which may include raw data such as barycentric coordinates and a possible hit time) actually indicates a hit. For a hit triangle, the ray tracing pipeline 300 triggers the execution of an any-hit shader 306. Note that a single ray may hit multiple triangles. There is no guarantee that the acceleration structure traversal stage will traverse the acceleration structure in order from closest to the ray source to furthest from the ray source. The hit or miss unit 308 triggers the execution of the closest-hit shader 310 for the triangle that the ray hits and that is closest to the ray origin, or triggers the miss shader if no triangles are hit.

[0032] Note that an any-hit shader 306 may "reject" a hit from the ray intersection test unit 304, and thus if the ray intersection test unit 304 does not find or accept a hit, the hit or miss unit 308 triggers the execution of the miss shader 312. An example scenario where an any-hit shader 306 may "reject" a hit is when at least a portion of the triangle reported by the ray intersection test unit 304 as being hit is completely transparent. Since the ray intersection test unit 304 only tests geometry and not transparency, an any-hit shader 306 called due to hitting a triangle with at least some transparency may determine that the reported hit is not actually a hit due to the transparent portion of the "hit" triangle. A typical use of the closest-hit shader 310 is to shade a material based on the texture of the material. A typical use of the miss shader 312 is to shade pixels with a color set by a skybox. It should be understood that the shader programs defined for the closest-hit shader 310 and the miss shader 312 may implement a variety of techniques for shading pixels and / or performing other operations.

[0033] Typically, the ray generation shader 302 generates rays using a technique known as backward ray tracing. In backward ray tracing, the ray generation shader 302 generates rays with an origin at a point on the camera. The point at which the ray intersects a plane defined to correspond to the screen defines a pixel on the screen, and the ray is used to determine the color of that pixel. If the ray hits an object, the pixel is shaded based on the closest hit shader 310. If the ray does not hit an object, the pixel is shaded based on the miss shader 312. Multiple rays can be cast for each pixel, and the final color of the pixel is determined by some combination of the colors determined for each ray for the pixel. As described elsewhere herein, a single ray can produce multiple samples, where each sample indicates whether the ray hit or missed a triangle. In one example, a ray is cast with four samples. Two such samples hit a triangle and two miss. Thus, the triangle color contributes only partially (e.g., 50%) to the final color of the pixel, where the other part of the color is determined based on triangles hit by other samples, or if no triangles are hit, by the miss shader. In some examples, rendering a scene involves casting at least one ray for each pixel of an image to obtain the color of each pixel. In some examples, multiple rays are cast for each pixel to obtain multiple colors for each pixel, thus achieving a multi-sampled rendering goal. In some such examples, at a later time, the multi-sampled rendering goal is compressed through color blending to obtain a single-sampled image for display or further processing. Although multiple samples for each pixel can be obtained by casting multiple rays for each pixel, techniques are provided herein for obtaining multiple samples for each ray, such that multiple samples can be obtained for each pixel by casting only one ray per pixel. Such a task can be performed multiple times to obtain additional samples for each pixel. More specifically, multiple rays can be cast for each pixel and multiple samples can be obtained for each ray, such that the total number of samples obtained for each pixel is the number of samples per ray multiplied by the number of rays per pixel.

[0034] Any one of the any hit shader 306, the closest hit shader 310, and the miss shader 312 can generate its own rays that enter the ray tracing pipeline 300 at the ray test point. These rays can be used for any purpose. A common use is to implement ambient lighting or reflection. In one example, when the closest hit shader 310 is called, the closest hit shader 310 generates rays in various directions. For each object or light hit by the generated rays, the closest hit shader 310 adds the lighting intensity and color to the pixel corresponding to the closest hit shader 310. It should be understood that although some examples of ways in which the various components of the ray tracing pipeline 300 can be used to render a scene have been described, any one of a variety of techniques can alternatively be used.

[0035] As described above, the determination of whether a ray hits an object is referred to herein as a "ray intersection test". The ray intersection test involves casting a ray from an origin and determining whether the ray hits a triangle, and if so, determining how far from the origin the triangle is hit. For efficiency, the ray tracing test uses a spatial representation called a bounding volume hierarchy. The bounding volume hierarchy is the "acceleration structure" described above. In the bounding volume hierarchy, each non-leaf node represents an axis-aligned bounding box that bounds the geometry of all of the node's children. In one example, the root node represents the maximum extent of the entire region in which the ray intersection test is performed. In this example, the root node has two children, each of which represents a mutually exclusive axis-aligned bounding box that subdivides the entire region. Each of these two children has two child nodes that represent axis-aligned bounding boxes that subdivide the space of their parent, and so on. The leaf nodes represent triangles on which the ray test can be performed. It should be understood that if a first node points to a second node, the first node is considered the parent of the second node.

[0036] The bounding volume hierarchy data structure allows a reduction in the number of ray-triangle intersections (ray-triangle intersections are complex and thus expensive in terms of processing resources) compared to a situation where such a data structure is not used and thus all triangles in the scene must be tested against the ray. Specifically, if a ray does not intersect a particular bounding box and that bounding box bounds a large number of triangles, all of the triangles in that box can be eliminated from the test. Thus, the ray intersection test is performed as a series of tests of the ray against axis-aligned bounding boxes followed by tests against triangles.

[0037] Figure 4 is an illustration of a bounding volume hierarchy according to one example. For simplicity, the hierarchy is shown in 2D. However, the extension to 3D is straightforward, and it should be understood that the tests described herein are generally performed in three dimensions.

[0038] A spatial representation 402 of the bounding volume hierarchy is illustrated in Figure 4 the left side, and a tree representation 404 of the bounding volume hierarchy is illustrated in Figure 4 the right side. In both the spatial representation 402 and the tree representation 404, non-leaf nodes are represented by the letter "N", while leaf nodes are represented by the letter "O". The ray intersection test will be performed by traversing the tree 404, and for each non-leaf node that is tested, if the box test for that non-leaf node fails, the branches below that node are eliminated. For the leaf nodes that are not eliminated, a ray-triangle intersection test is performed to determine whether the ray intersects the triangle at that leaf node.

[0039] In one example, the ray intersects O5 but does not intersect the other triangles. This test will be performed on N1 to determine that the test is successful. This test will be performed on N2 to determine that the test fails (because O5 is not within N1). This test will eliminate all children of N2 and will be performed on N3, noting that the test is successful. This test will test N6 and N7, noting that N6 is successful but N7 fails. This test will test O5 and O6, noting that O5 is successful but O6 fails. Instead of testing 8 triangles, two triangle tests (O5 and O6) and five box tests (N1, N2, N3, N6, and N7) are performed.

[0040] The above Figures 1 to 4 describes a specific implementation in which a top-down construction of a bounding volume hierarchy can be performed. The top-down construction of the bounding volume hierarchy generates a bounding volume hierarchy for a scene, accepting the geometry of the scene (e.g., a set of triangles) as input and generating a BVH as output. Generally speaking, the top-down construction involves iteratively generating nodes for the BVH. In each node, a candidate split of the triangles in the node is determined, and the children for the node are determined based on an evaluation of the candidate split. Additional details are now provided.

[0041] Figure 5 illustrates the generation of a BVH 505 from scene geometry by a BVH constructor 501 according to one example. The BVH constructor 501 accepts the scene geometry 503 and uses a top-down technique to generate a bounding volume hierarchy 505. The scene geometry 503 includes geometric objects corresponding to the objects of the scene to be rendered. The BVH 505 is a bounding volume hierarchy that allows for a quick determination of whether a ray intersects the scene geometry, as described with reference to Figures 1 to 4 In various examples, the BVH constructor 501 is embodied entirely in software, entirely in hardware (e.g., as a circuit), or as a combination thereof. In different examples, the BVH constructor 501 is within a device 100 in which ray tracing is performed, or within a different system. In one example, an application developer creates a scene with geometry and uses the BVH constructor 501 to generate a BVH corresponding to the scene, and then delivers the application to a user for execution. In another example, an application developer uses the BVH constructor 501 to generate a BVH corresponding to a scene and also uses the BVH to execute a ray tracing-enabled application. In another example, the BVH constructor 501 present in the device 100 (e.g., within the APD 116) generates a BVH from the scene geometry for an application, and then the APD 116 uses the generated BVH to render the geometry of the scene. Although some examples describe using a scene, these examples should not be considered restrictive.

[0042] Figure 6A Illustrates aspects of generating a BVH using a top - down technique according to an example. The top - down technique iteratively constructs a BVH. The top - down technique begins with a box node having a bounding volume. In the Figure 6A example geometry 600 of, the bounding volume 601 encloses the triangle 603. The centroid 605 is illustrated for each triangle. The centroid represents the vertex positions that characterize the position of the triangle. In some examples, the centroid of a triangle is the intersection between lines that bisect each edge and terminate at the opposite vertex of that edge.

[0043] For a given box node (e.g., root node 602 or box node 604), the BVH constructor 501 identifies a set of candidate splits for the bounding volume of that box node, evaluates all candidate splits to determine a cost metric for each candidate split, and selects one of the candidate splits based on a comparison of the cost metrics. Any technically feasible cost metric can be used. In some examples, the cost metric is the sum of the areas of the faces of the tightly - fitting bounding boxes of the triangles for each part of the candidate split, and in other examples, other cost metrics are used, such as a cost metric based on the surface area of the bounding volume of the bounding boxes for the parts of the candidate split. In one example, if a candidate split defines a first set of triangles and a second set of triangles, bounding boxes that tightly bound each set are formed. Then, the area of each face of each bounding box is determined, and thus the sum of those areas is determined for each bounding box and added together. This metric is the cost metric and is generated for each candidate split. Then, the smallest such cost metric indicates which split should be selected. Similarly, although specific cost metrics are described, any technically feasible cost metric for selecting a candidate split is possible.

[0044] The selected candidate split indicates which of the triangles within the bounding volume 601 of the box node are to be included in each of the sub - box nodes 604. More specifically, the candidate split defines which geometric parts of the bounding volume 601 are associated with which sub - box nodes 604. Each sub - box node 604 is assigned a different geometric part.

[0045] In the Figure 6A example, the selected candidate split 600 indicates that one side of the split includes triangles 603 - 1 and 603 - 2, and the other side of the split includes triangles 603 - 3, 603 - 4, and 603 - 5. Thus, in this example, one sub - box node 604 of the root node 602 is generated with a bounding box that encloses the triangles (603 - 1, 603 - 2) on one side, and another sub - box node 604 of the root node 602 is generated with a bounding box that encloses the triangles (603 - 3, 603 - 4, 603 - 5) on the other side.

[0046] Figure 6B Illustrates a set of candidate splits for a box node according to an example. Nine candidate splits are illustrated - 650-1 to 650-9. In each candidate split, different boundaries 652 between different sides of the split are illustrated. It can be seen that different sets of triangles fit within different sides of each candidate split. Thus, for each candidate split, different sets of triangles will be included in different bounding boxes. Thus, each different candidate split represents a different way of subdividing triangles among the children of the box node. Additionally, each set of triangles in each side of each candidate split is associated with a different bounding volume. Thus, each candidate split has children with different bounding boxes. Although techniques for grouping primitives based on locations in three-dimensional space are described in Figure 6B , other techniques can be used to group primitives together. Such other techniques may or may not consider the location and / or extent of such primitives and may additionally or alternatively consider other aspects of such primitives, such as primitive size.

[0047] Figure 6B Illustrates nine different candidate splits. This number is small for clarity. However, there may be a very large number of candidate splits. For example, three candidate splits 650-1, 650-2, 650-3 are shown, which divide the geometry horizontally at three different points. A typical bounding volume hierarchy can include a huge number of triangles. For such a bounding volume hierarchy, the number of possible candidate splits can be very high. In addition to the above, in a naive approach to constructing a top-down bounding volume hierarchy, for each candidate split, the BVH constructor 501 must visit each triangle to identify which side of the split the triangle falls into. In this naive approach, this determination must be made for each node and for each candidate split, resulting in a high number of computations having to be performed. This means that constructing a bounding volume hierarchy in a top-down manner can be very time-consuming.

[0048] For at least these reasons, this document provides a technique that helps reduce the amount of time required to build a BVH in a top-down manner. Generally speaking, the technique includes: before building the BVH, predefined sets of triangles at different levels of detail, and centroid boxes for each set of triangles. Generally speaking, the size of the centroid box for a set of triangles at a particular level of detail is different from the size of the centroid box for sets of triangles at different levels of detail, although the actual size of these centroid boxes changes based on the actual geometry within each set of triangles. As described above, the technique includes: generating a centroid box for each set of triangles. A centroid box is a box that bounds the centroids of the triangles within a set of triangles. Note that the centroid box of a set of triangles is typically smaller than the bounding volume of the set of triangles, because the bounding volume bounds the complete geometry of the triangles, while the centroid box of the triangles bounds the centroids of the triangles. The result of the above is a data structure that includes multiple levels of detail (sometimes referred to herein as a "collection of triangle sets"), where each level of detail has a set of triangles. Each set of triangles specifies a bounding box that bounds the triangles within that set of triangles and a centroid box that bounds the centroids of those triangles. Although this document sometimes describes a "collection of triangle sets", in some examples, a "collection of primitive sets" may alternatively be used. In this document, any instance of the term "collection of triangle sets" can be replaced with the term "collection of primitive sets". A collection of primitive sets is similar to a collection of triangle sets, except that a collection of primitive sets has primitives instead of triangles. Primitives are more general than triangles and include triangles or other geometric structures that can be found at the leaf nodes of a BVH. Such other geometric structures include procedurally defined geometric structures, which are geometric structures for which the intersection between the geometric structure and a ray is determined based on the execution of a shader program or by some other technique. Other geometric structures may also include primitives that are not procedurally defined but are not triangular in shape. The primitives of a collection of primitive sets do not include the bounding boxes found in the box nodes of a BVH. In addition to the above, in the case of the triangles described herein, such descriptions also apply to non-triangular primitives. In other words, in this document, the word "primitive" can be used to replace the word "triangle".

[0049] Triangle sets allow for certain acceleration in top-down BVH creation. More specifically, as described above, in a naive implementation of the top-down approach for each node, for each candidate split, the BVH constructor must determine where each triangle is placed within that candidate split (i.e., which side of the split the triangle falls on). Thus, in such implementations, the BVH constructor must iterate over each triangle for each candidate split. Building triangle sets using the information above allows for determining which side of a candidate split each centroid box lies on for each BVH node of the BVH being constructed rather than for each triangle. By having a constant number of such centroid boxes, the time complexity of the BVH is reduced because the BVH constructor iterates over a constant number of centroid boxes rather than over the number of triangles. In other words, by pre-building triangle sets with centroid boxes and then determining which side each centroid box falls within, rather than which side each individual triangle falls on, the amount of time required for BVH construction is reduced. Indeed, the triangles initially need to be placed into centroid boxes, however this processing occurs at the start of BVH construction rather than for each node of the BVH being constructed. Thus, rather than determining which triangles fall within each side of a split for each node of the BVH being constructed, it is determined which centroid boxes fall within each side for each node of the BVH being constructed. The improvement in time complexity comes from the fact that evaluating all triangles for each node scales more significantly in time than evaluating a candidate set for each node. Specifically, once the triangle sets are built, there is a fixed number of centroid boxes within such triangle sets. Thus, the number of items being evaluated for placement in the sides of candidate splits for the BVH being constructed is constant. On the other hand, in a naive implementation, the number of items being evaluated is not constant - the number is proportional to the number of triangles to be represented by the BVH being constructed. Since time complexity is an expression of how much time an algorithm consumes as a function of the number of objects being processed by that algorithm, an algorithm that replaces a variable number of items (triangles) with a fixed number of items (centroid boxes) has a lower time complexity. It should be understood that a centroid box can include multiple triangles and that there can be a deterministic number of centroid boxes such that the number of centroid boxes can be fixed rather than changing based on the number of primitives. Additional details are now provided.

[0050] Figure 7Illustrates an example triangle set collection 700. The triangle set collection 700 includes a plurality of box nodes 704. The root node 702 is also a box node 704. Each box node is a set of triangles. Thus, each box node has an associated set of triangles, an associated bounding volume, and an associated centroid box. For a given box node, the associated triangles are shown below the reference numeral 704. For example, box node 704-1 is associated with triangles 7101 to 10. Thus, box node 704-1 has a bounding volume that bounds all of the triangles in triangles 1 to 10, and box node 704-1 has a centroid box that bounds all of the centroids of triangles 1 to 10. In some examples, as used herein, the phrase "bounds" means tightly bounds the object being referred to, which means that the box is large enough to enclose all of the items being referred to, but no larger than the item.

[0051] It should be understood that the illustrated triangle set collection 700 itself is a BVH. This BVH is used to construct a different, higher-performance BVH. In other words, the techniques described herein generate a first BVH for spatially classifying the triangles of a scene, and the spatial classification from this first BVH is used to assist in generating a second, higher-performance BVH that is the end result of the technique. In some examples, the first BVH 700 is constructed using a relatively simple BVH construction algorithm such as a parallel linear BVH ("LBVH"). Any BVH construction algorithm can be used to generate a BVH that is used to generate a set of triangles for the BVH construction techniques of the present disclosure, as long as there is no overlap in each split centroid box. In other words, the BVH construction techniques of the present disclosure can be considered a means for refining different BVHs into higher-performance BVHs. It should be understood that the topology of a BVH can greatly affect the performance of BVH traversal. The parallel linear BVH is particularly suitable for generating the first BVH because the parallel linear BVH generates the BVH in a manner of a virtual grid based on the centroid using Morton codes. Thus, the centroid box for a primitive set is an integer-based bounding box made of Morton codes that form a virtual grid. This aspect allows for easy calculation of the centroid box range corresponding to each bounding box of the LBVH and the primitives that fall within such centroid boxes. In other words, the LBVH defines a grid in which each cell corresponds to a different integer Morton code value. Additionally, the centroid of any particular primitive has one of such integer Morton code values. Thus, it is easy to determine which centroid box a primitive falls into, and thus it is easy to generate a centroid box for each bounding box of the LBVH that indicates the range of centroids within such centroid boxes. Although examples of generating a set of triangles by constructing a BVH are described, the techniques presented herein are not limited to using a set of triangles generated from a BVH.

[0052] Level 706 of the BVH represents different levels of detail of the triangle set. For example, level 706-1 represents a higher level of detail than level 706-2. Similarly, generally speaking, levels with higher levels of detail include more triangles and are typically larger than levels with lower levels of detail.

[0053] Now referring to the generation of the second, higher-performance BVH, as described with reference to Figure 6A As described, constructing such a BVH in a top-down manner involves iteratively generating children for the BVH nodes of the BVH by evaluating candidate splits for the triangles bounded by the box node. In one example, the BVH constructor 501 generates children for a box node (e.g., box node 602) of the BVH being constructed. To perform this operation, the BVH constructor 501 identifies the triangles within the bounding volume of the box node 602 and generates a plurality of candidate splits for those triangles. Each candidate split indicates a certain number of "sides" and the triangles belonging to each side. The BVH constructor 501 selects a candidate split as the accepted split and generates children for each side of the accepted split. The BVH constructor 501 then continues to perform these operations to generate a complete BVH. For example, the BVH constructor 501 generates children for newly generated box nodes in a manner similar to generating children for the root node 602, and so on. In some examples, when the BVH is complete, the BVH constructor 501 stops generating.

[0054] As described above, generating a BVH includes determining which triangles fit within each candidate split. In the techniques described herein, these steps are performed by determining which centroid boxes of the set of triangle sets are within each candidate split. Since each centroid box includes one or more triangles, determining which centroid boxes are within each candidate split necessarily results in determining which triangles are within each candidate split. More specifically, for any given BVH node of the BVH being constructed, the technique includes: determining which centroid boxes fall within each side of the candidate split. The technique then includes: selecting a candidate split based on a cost metric and generating child nodes for the BVH node of the BVH being constructed based on the candidate split, as described elsewhere herein.

[0055] For a given BVH node of the BVH being constructed, the BVH constructor 501 performs the following operations. The BVH constructor 501 determines an appropriate partition resolution. The partition resolution identifies the size of the cells of the centroid box that bounds all the centroids of the primitives descending from the BVH node. The cells define the manner in which the set of triangle sets is evaluated to determine which triangle sets fall within which partition of a non-leaf node of the BVH being constructed, as detailed below. Figure 8Illustrates different partition resolutions according to an example. In the first partition resolution 802-1, the size of the cell 804-1 is larger than the size of the cell 804-2 in the second partition resolution 802-2.

[0056] In some examples, the appropriate partition resolution is provided as a tunable parameter to the BVH constructor 501 (e.g., through application execution on the CPU 102 or by another software entity such as a shader program or a hardware entity such as the hardware within the APD 116). In some examples, the resolution specifies a number of cells by which the centroid box corresponding to a non-leaf node of the BVH being built is partitioned. Thus, in these examples, the partition resolution identifies a number of cells to divide the centroid box corresponding to the box nodes of the BVH being built into, but does not necessarily specify the absolute size of those boxes. For a given box node of the BVH being constructed, the partition resolution determines the possible number of candidate splits. More specifically, the boundaries of the cells 804 of the partition resolution indicate the boundaries of the candidate splits. In one example, the BVH constructor 501 determines a number of candidate splits, where each such candidate split has at least one side different from all the sides of the remaining candidate splits in that candidate split, and where the boundaries of each candidate split are aligned with the boundaries of the cells 804 of the partition resolution. In one example, the candidate splits for the partition resolution 802-1 may include a bottom side and a top side, the bottom side including the bottom four cells 804-1 and the top side including the top four cells 804-1. Different candidate splits for this resolution may include a left side and a right side, the left side including the left four cells 804-1 and the right side including the right four cells 804-1. For the partition resolution 802-2, many more candidate splits can occur. For example, the bottom plane of the cell 804-2 can form one side of the split, and the top three planes of the cell 804-2 can form a different side. Alternatively, the bottom two planes of the cell 804-2 can form one side of the split, and the top two planes can form the other side. "Plane" means a set of cells 804-2 having the same vertical position (but changing depth and horizontal position). As shown, it can be seen that the partition resolution of the centroid box for which candidate splits are being determined determines the number of possible candidate splits to be evaluated. A finer resolution (e.g., the partition resolution 802-2) results in a greater number of candidate splits, and a coarser resolution (e.g., the partition resolution 802-1) results in a smaller number of candidate splits.

[0057] To generate candidate splits for candidate BVH nodes of a BVH being constructed, the BVH constructor 501 uses the set of triangle sets 700. Specifically, the BVH constructor 501 traverses down the set of triangle sets 700 from a candidate BVH node to find a set of BVH nodes representing all the triangles that are descendants of the candidate BVH node. Each BVH node in this set of BVH nodes has a centroid box that fits within a cell 802 of the selected partition resolution 802.

[0058] It should be understood that within the set of triangle sets 700, box nodes have pointers to child box nodes. Traversing down the set of triangle sets 700 means following these pointers. Traversing down the set of triangle sets 700 to find the set of BVH nodes means finding the highest BVH node 704 whose centroid box fits within the cell 802, and also means finding the BVH nodes 704 that together "cover" all the triangles bounded by the candidate BVH node of the BVH being constructed.

[0059] Figure 9 An example of finding a set of box nodes representing all the triangles that are descendants of a candidate box node as described above is illustrated. In this example, the candidate box node is box node 702. The BVH constructor 501 searches down the tree for the highest box node 704 whose centroid box fits within the cell 902 (as shown, for the selected partition resolution). Additionally, the BVH constructor 501 identifies such box nodes 704 until the set of identified box nodes together bounds all the triangles bounded by the candidate box node 702.

[0060] In Figure 9In it, the BVH constructor 501 examines the box node 704-1 and determines that the centroid box (i.e., the box tightly enclosing all the centroids of all the triangles bounded by the bounding volume of the box) does not fit within a single cell 902. The BVH constructor 501 examines the two children of the box node 704-1, which are the box node 704-3 and the box node 704-4. The BVH constructor 501 determines that the centroid box of the box node 704-3 fits within the cell 902, but the centroid box of the box node 704-4 does not fit within the cell 902. The BVH constructor 501 identifies the box node 704-3 as a box node representing all the triangles that are descendants of the candidate box node 702 among this set of box nodes. The BVH constructor 501 examines the children of the box node 704-4 and determines that the centroid boxes of the box nodes 704-9 and 704-10 each fit within the corresponding cell 902. Similarly, the BVH constructor 501 determines that the centroid box of the node 704-2 does not fit within the cell 902, but determines that the centroid boxes of the nodes 704-5 and 705-6 fit within the corresponding cell 902. The identified box nodes are 704-3, 704-9, 704-10, 704-5, and 704-6. Additionally, all the triangles among the set of bounding triangles 1 to 16 are those that are all the triangles bounded by the candidate box node 702. The result is a set of identified box nodes of the triangle set collection 700. This set can be used to determine, in a more performant way than having to individually examine all the triangles, which triangles fit within each candidate split.

[0061] Figure 9 Illustrates the placement of the centroid box 904 within the cell 902 of the centroid box of the candidate box node 702. The centroid box 904-1 is associated with the node 704-3 and fits within the cell 902-1. The centroid box 904-2 is associated with the node 704-9 and fits within the cell 902-2. The centroid box 904-3 is associated with the node 704-10 and fits within the cell 902-3. The centroid box 904-4 is associated with the node 704-5 and fits within the cell 902-6. The centroid box 904-6 is associated with the node 704-6 and fits within the cell 902-7. As can be seen, a set of nodes 704 has been found that span all the triangles that are descendants of the candidate nodes of the BVH being constructed ( Figure 9 (not shown in the figure) and fit within the cell 902.

[0062] Once the nodes 704 of the triangle set that fit within the cell for the node of the BVH being constructed are found, it is relatively straightforward to determine which triangles fit within which side of the candidate split. More specifically, since the extent of the cell 902 is known and since the candidate split is defined relative to the cell boundaries, it is straightforward to determine which side of the set of triangles associated with a particular node 704 of the triangle set that fits within the cell, because each node 704 has a corresponding centroid box. For example, to determine which side of the candidate split the centroid box fits within, the bounding volume hierarchy generator 501 compares the boundaries of the centroid box with the boundaries of the side of the candidate split and identifies the side of the centroid box as the side into which the centroid box fits. Thus, for any particular candidate split, it is relatively straightforward to determine the side associated with each of the nodes 704 that are determined to fit within the cell. The relatively small number of simpler comparisons involved in this technique is much less than comparing the actual geometry of each triangle against the boundaries of the sides. This technique can thus result in a similar output to a top-down "bin-based" BVH constructor (e.g., one where the constructor evaluates for each box node which side of the candidate split each triangle falls within), but at a much lower computational cost.

[0063] Figures 9A to 9C Illustrates different candidate splits 900 according to an example. In Figure 9A , the boundary 908-1 splits the geometry into a top side (associated with the centroid box 910-1) and a bottom side (associated with the centroid box 910-2). The centroid box 910 bounds the centroid of the centroid box 904. The top side includes the centroid boxes 904-1, 904-2, and 904-3, and the bottom side includes the centroid boxes 904-4 and 904-5. The bounding volume hierarchy generator 501 determines which side of the boundary 908-1 the centroid box 904 falls within by comparing the extent of the centroid box 904 with the boundary 908-1. This determination results in determining which side of the boundary 908-1 each of the triangles corresponding to the centroid box 904 falls within. Thus, it is not necessary to test each such triangle against the boundary 908-1.

[0064] In Figure 9B , the boundary 908-2 splits the centroid box 904 as shown, resulting in the centroid box 910-3 and the centroid box 910-4. Similarly, in Figure 9C , the boundary 908-3 splits the centroid box 904 as shown, resulting in the centroid box 910-5 and the centroid box 910-6. To determine the children for the BVH node corresponding to the geometry of Figures 9A to 9C , the bounding volume hierarchy constructor 501 evaluates these candidate splits 900, selects one candidate split 900 based on a selection criterion, and generates child BVH nodes from the sides of the split, as described elsewhere herein.

[0065] Figure 10 Illustrates additional operations related to constructing a BVH according to an example. More specifically, Figure 10 Illustrates the generation of a BVH 1001 under construction based on a set of triangle sets 700 (which, in some examples, is constructed using an algorithm such as LBVH as described elsewhere herein, and which, in some examples, includes a centroid box at each BVH node 704). Phase 11000-1 results in the generation of BVH nodes 1000-2 and 1000-3 from BVH node 1002-1. Specifically, the BVH constructor 501 starts at the root node 702 of the set of triangle sets 700 and forms node 1002-1 in the BVH 1001 under construction based on that root node 702. Node 1002-1 has a centroid box that bounds all the centroids of the root node 702 and has a bounding box that bounds all the triangles of the root node 702. At this point, node 1002-1 is a candidate box node for which a split is being generated.

[0066] The BVH constructor 501 traverses down the set of triangle sets 700 to identify the highest node 704 whose centroid box fits entirely within a single cell 902( Figure 9 ) and encloses all the triangles in the triangles enclosed by the root node 702. In phase 1000-1, the BVH constructor 501 has determined that the centroid boxes for each of nodes 704-3, 704-9, 704-10, 704-5, and 704-6 fit within a single cell. As described elsewhere herein, the size of the cell can be a tunable parameter. In some examples, the size of the cell is determined by dividing the centroid box of a candidate BVH node by a resolution parameter, which is a tunable parameter or derived from a tunable parameter. In one example, the resolution parameter specifies that the centroid box should be divided into 64 cells. Thus, different BVH nodes will have different cell sizes. In examples where the resolution parameter is a tunable parameter, the cell size is determined indirectly based on the resolution parameter. The tunable parameter can be the same for different levels of the BVH under construction or can be different for different levels. In one example, the tunable parameter remains constant until a certain level is reached, and then the cell size remains constant. Any technically feasible means for setting the tunable parameter to specify the resolution and thus the cell size for any particular box node 704 in the BVH 1001 under construction is possible.

[0067] The BVH constructor 501 generates candidate splits using BVH nodes identified as suitable within a single unit, evaluates the candidate splits, and selects a candidate split based on any technically feasible criterion (e.g., the lowest sum of bounding volume surface areas). The BVH constructor 501 generates children based on the candidate split, where each side corresponds to a new node in the BVH. Each such node has a bounding volume that bounds all the primitives on the corresponding side and a centroid box that bounds the centroid of all the primitives that are descendants of the node. In phase 1000-1, the generated BVH nodes include BVH node 1002-2 and BVH node 1002-3.

[0068] In phase 1000-2, the BVH constructor 501 determines children for BVH node 1002-2, which is now a candidate node. The BVH constructor 501 starts with the BVH nodes of the triangle set collection 700 that together bound all the triangles that are descendants of the candidate node 1002-2. In this case, BVH node 704-1 bounds all such triangles. The BVH constructor 501 traverses the triangle set collection 700 downward to find the highest BVH node that fits within a unit for a certain resolution and together encloses all the triangles of the candidate node 1002-2. In the example shown, such BVH nodes 704 include BVH nodes 704-7, 704-8, 704-15, and 704-16. The BVH constructor 501 generates candidate splits from these BVH nodes, selects one of the candidate splits, and generates children for BVH node 1002-2 of the BVH 1001 being constructed according to the selected candidate split. The BVH constructor 501 repeats these steps until the complete BVH is built.

[0069] Figure 11 is a flowchart of a method 1100 for constructing a BVH according to an example. Although described with reference to Figures 1 to 10 a system, any system configured to perform the steps of method 1100 in any technically feasible order falls within the scope of the present disclosure.

[0070] At step 1102, the BVH constructor 501 determines the unit size for the centroid box of the triangle set collection 700 based on the resolution to generate units. As described above, the resolution can be associated with a tunable parameter indicating the number of units into which a box node is divided.

[0071] At step 1104, the BVH constructor 501 identifies nodes of the triangle set collection that fit within the subdivided units. More specifically, the BVH constructor 501 finds the highest box node of the triangle set collection 700 whose centroid box fits within the subdivided units.

[0072] At step 1106, the BVH constructor 501 generates candidate splits based on the identified nodes. Specifically, the BVH constructor 501 selects a plurality of boundaries for different candidate splits, where each boundary lies on a face of the cell. The BVH constructor then places the centroid box of each box node among the box nodes identified in step 1104 into one side of the candidate split by comparing the extent of the centroid box with the boundaries. The result for any particular candidate split is an indication of which centroid boxes (and thus which nodes identified in step 1104) are on each side of the candidate split.

[0073] At step 1108, the BVH constructor 501 selects one of the candidate splits in the candidate split based on a selection criterion. In various examples, the selection criterion specifies a way to evaluate different candidate splits to select the one candidate split that is considered to be "best". In one example, the selection criterion is the surface area of the bounding volume of the triangles times the number of primitives for each side.

[0074] At step 1110, the BVH constructor 501 generates children for the box nodes of the BVH being constructed based on the selected candidate split. Specifically, the BVH constructor 501 generates one child for each side of the selected candidate split, where each child has a bounding volume that bounds all of the geometry associated with the associated side.

[0075] The BVH constructor 501 repeats method 1100 any number of times to build the BVH. After step 1110, the BVH constructor 501 selects a node of the BVH being constructed to generate children for. Step 1102 for that node will calculate the cell size for the partitioning of the centroid box. Step 1104 will identify the nodes of the set of triangle sets that fall within the cell of the subdivision node and together bound all of the triangles bounded by the subdivision node. The BVH constructor 501 will continue with steps 1106, 1108, and 1110 and continue for additional nodes of the BVH being constructed.

[0076] It should be understood that there may be many variations based on the disclosure herein. Although the above features and elements are described in particular combinations, each feature or element can be used alone without the other features and elements or in various combinations with or without other features or elements.

[0077] The various functional units illustrated in the figures and / or herein (including, but not limited to, processor 102, input driver 112, input device 108, output driver 114, output device 110, acceleration processing device 116, scheduler 136, computing unit 132, SIMD unit 138) may be implemented as a general-purpose computer, processor, or processor core, or as a program, software, or firmware stored in a non-transitory computer-readable medium or another medium and executable by a general-purpose computer, processor, or processor core. The provided methods may be implemented in a general-purpose computer, processor, or processor core. By way of example, suitable processors include general-purpose processors, dedicated processors, conventional processors, digital signal processors (DSPs), multiple microprocessors, one or more microprocessors associated with a DSP core, controllers, microcontrollers, application specific integrated circuits (ASICs), field programmable gate array (FPGA) circuits, any other type of integrated circuit (IC), and / or state machines. Such processors may be manufactured by configuring a manufacturing process using hardware description language (HDL) instructions for the processing and the results of other intermediate data including netlists (such instructions capable of being stored on a computer-readable medium). The results of such processing may be a mask, which may then be used in a semiconductor manufacturing process to fabricate a processor implementing the features of the present disclosure.

[0078] The methods or flowcharts provided herein may be implemented in a computer program, software, or firmware incorporated in a non-transitory computer-readable storage medium for execution by a general-purpose computer or processor. Examples of non-transitory computer-readable storage media include read-only memory (ROM), random access memory (RAM), registers, cache memory, semiconductor memory devices, magnetic media (such as internal hard disks and removable disks), magneto-optical media, and optical media (such as CD-ROM disks) and digital versatile disks (DVDs).

Claims

1. A method for constructing a bounding volume hierarchy, the method comprising: Subdividing a candidate box node based on a resolution to generate a plurality of cells of the candidate box node; Identifying a plurality of nodes of a primitive set suitable within the cells; Generating a plurality of candidate splits based on the plurality of nodes; Selecting a candidate split based on a selection criterion to obtain a selected candidate split; And Based on the selected candidate split, generating child box nodes for a box node of the bounding volume hierarchy being constructed.

2. The method according to claim 1, wherein the resolution indicates the number of cells into which the candidate box node is subdivided.

3. The method according to claim 1, wherein identifying the plurality of nodes suitable within the unit comprises: Identifying a node whose centroid box fits within a single cell of the cells of the candidate box node.

4. The method according to claim 3, wherein the centroid box includes a box defining the centroid of the primitives of the node.

5. The method according to claim 1, wherein generating the plurality of candidate splits comprises: Splitting the bounding volume of the candidate box node based on a boundary aligned with the plurality of cells.

6. The method according to claim 1, wherein selecting the candidate split comprises: Evaluating the plurality of candidate splits according to the selection criterion and selecting one of the candidate splits.

7. The method according to claim 1, wherein the selection criterion is the lowest total bounding box surface area criterion.

8. The method according to claim 1, wherein generating the sub-box node comprises: Generating child box nodes for each side of the selected candidate split, wherein each child box node has a bounding box defining each primitive corresponding to a side of the child box node.

9. The method according to claim 1, the method further comprising: For a plurality of box nodes of the bounding volume hierarchy being constructed, repeating the subdivision, identification, generation of the plurality of candidate splits, selection, and generation of the child box nodes.

10. A system for constructing a bounding volume hierarchy, the system comprising: A memory configured to store the bounding volume hierarchy; And A bounding volume hierarchy constructor configured to construct the bounding volume hierarchy by performing operations including: Subdividing a candidate box node of the bounding volume hierarchy based on a resolution to generate a plurality of cells of the candidate box node; Identifying a plurality of nodes of a primitive set suitable within the cells; Generating a plurality of candidate splits based on the plurality of nodes; Selecting a candidate split based on a selection criterion to obtain a selected candidate split; and Based on the selected candidate split, generating child box nodes for a box node of the bounding volume hierarchy being constructed.

11. The system according to claim 10, wherein the resolution indicates the number of cells into which the candidate box node is subdivided.

12. The system according to claim 10, wherein identifying the plurality of nodes suitable within the unit comprises: Identifying a node whose centroid box fits within a single cell of the cells of the candidate box node.

13. The system according to claim 12, wherein the centroid box includes a box defining the centroid of the primitives of the node.

14. The system according to claim 10, wherein generating the plurality of candidate splits comprises: Splitting the bounding volume of the candidate box node based on a boundary aligned with the plurality of cells.

15. The system according to claim 10, wherein said selecting the candidate split comprises: Evaluating the plurality of candidate splits according to the selection criterion and selecting one of the candidate splits.

16. The system according to claim 10, wherein the selection criterion is the lowest total bounding box surface area criterion.

17. The system according to claim 10, wherein generating the sub-box node comprises: Generating child box nodes for each side of the selected candidate split, wherein each child box node has a bounding box defining each primitive corresponding to a side of the child box node.

18. The system according to claim 10, wherein the bounding volume hierarchy constructor is further configured to: for a plurality of box nodes of the bounding volume hierarchy being constructed, repeat the subdivision, identification, generation of the plurality of candidate splits, selection, and generation of the child box nodes.

19. A non-transitory computer-readable medium storing instructions that, when executed by a processor, cause the processor to perform operations, the operations including: subdividing candidate box nodes based on a resolution to generate a plurality of cells of the candidate box nodes; identifying a plurality of nodes of a primitive set collection that fit within the cells; generating a plurality of candidate splits based on the plurality of nodes; selecting a candidate split based on a selection criterion to obtain a selected candidate split; and generating child box nodes for a box node of the bounding volume hierarchy being constructed based on the selected candidate split.

20. The non-transitory computer-readable medium according to claim 19, wherein the resolution indicates the number of cells into which the candidate box node is subdivided.