Systems and methods for accelerated ray tracing using asynchronous operations and light transformations

By executing shader programs on the GPU and using RTU asynchronous traversal to accelerate structures, the problem of high computational cost in real-time applications is solved, and more efficient rendering performance is achieved.

CN114170365BActive Publication Date: 2025-07-01SONY INTERACTIVE ENTERTAINMENT LLC
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202110959031.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2020-08-20
Filing Date
2021-08-20
Publication Date
2025-07-01
Estimated Expiration
2041-08-20

AI Technical Summary

Technical Problem

Ray tracing faces high computational cost in real-time applications, resulting in limited rendering speed.

Method used

A shader program is executed on a graphics processing unit (GPU), and a ray tracing unit (RTU) implemented by the hardware is asynchronously traverses the acceleration structure, detects the intersection points of light rays and objects, and transmits the results to the shader program.

Benefits of technology

By separating calculations and improving parallel processing capabilities, the speed of ray tracing is significantly improved, the load on the shader program is reduced, and the rendering performance is improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114170365B_ABST
    Figure CN114170365B_ABST
Patent Text Reader

Abstract

A graphics processing unit (GPU) includes one or more processor cores adapted to execute software-implemented shader programs, and one or more hardware-implemented ray tracing units (RTUs) adapted to traverse an acceleration structure to compute intersections of rays with bounding volumes and graphical primitives asynchronously with shader operations. The RTU implements traversal logic to traverse the acceleration structure, including performing ray transformations as needed to account for changes in coordinate space between levels, performing stack management and other tasks to relieve the burden on the shader, delivering intersection points to the shader, and then the shader computes whether the intersection point hits a transparent or opaque portion of the intersected object. Thus, one or more processing cores within the GPU perform accelerated ray tracing by offloading aspects of the processing to the RTU, which traverses the acceleration structure representing the 3D environment therein.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application generally relates to systems and methods for accelerating ray tracing. Background Art

[0002] Ray tracing is used to simulate optical effects in computer-generated 3D graphics by tracing the optical path from the eye of a hypothetical observer (usually the camera position) to virtual objects in the graphics. Compared to other techniques such as rasterization, ray tracing produces optical effects with a higher degree of realism but at a greater computational cost. This means that in real-time applications such as video games, ray tracing poses challenges because the rendering speed is critical. Summary of the Invention

[0003] Accordingly, a method for graphics processing includes executing a shader program on a graphics processing unit (GPU), the shader program performing ray tracing of a 3D environment represented by an acceleration structure. The method includes using a hardware-implemented ray tracing unit (RTU) to traverse the acceleration structure at the request of the GPU intrinsic shader program and using the result of the acceleration structure traversal at the shader program.

[0004] In an exemplary embodiment, the acceleration structure traversal performed by the RTU can be asynchronous with respect to the shader program. In some implementations, the result of the acceleration structure traversal performed by the RTU includes detection of an intersection point between a ray and a bounding volume contained within the acceleration structure. In some examples, the RTU processing includes maintaining a stack used in the acceleration structure traversal.

[0005] The acceleration structure can be a hierarchy having multiple levels. In such embodiments, the result of the acceleration structure traversal performed by the RTU can include detection of a transition from a higher level to a lower level within the multiple levels of the acceleration structure. The result of the acceleration structure traversal performed by the RTU can also include detection of a transition from a lower level to a higher level within the multiple levels of the acceleration structure. The acceleration structure traversal performed by the RTU can include handling transitions between the multiple levels of the acceleration structure.

[0006] In non - limiting implementations, the results of the acceleration structure traversal performed by the RTU can include the detection of intersections between rays and primitives contained within the acceleration structure. In such implementations, the results of the acceleration structure traversal performed by the RTU can include the detection of the earliest intersection between a ray and a primitive contained within the acceleration structure. Additionally, the results of the acceleration structure traversal performed by the RTU can include the sorting of intersections by the RTU according to the distance of the detected intersections from the ray origin such that the RTU detects a first intersection between the ray and a primitive as it traverses the acceleration structure, the RTU detects a second intersection between the ray and a primitive as it traverses the acceleration structure, and when transmitting the results from the RTU to the shader program, the second intersection result is transmitted before the first intersection result. If needed, when the RTU detects an intersection between a ray and a primitive contained within the acceleration structure and transmits the result to the shader program, the shader program and the RTU then communicate regarding the result of the hit test performed by the shader program between the ray and the primitive.

[0007] When the RTU detects an intersection between a ray and a bounding volume contained within the acceleration structure and transmits the result to the shader program, the shader program and the RTU can then communicate regarding the shader program's determination of whether to ignore the intersection and / or the shader program's determination of the position of the intersection along the ray.

[0008] In another aspect, a graphics processing unit (GPU) includes: at least one processor core adapted to execute a software - implemented shader; and at least one hardware - implemented ray tracing unit (RTU), the RTU being separate from the processor core and adapted to traverse an acceleration structure to identify intersections of rays with objects represented in the acceleration structure, to generate results, and to return the results to the shader for the shader to identify a hit associated with the intersection.

[0009] In an exemplary implementation of this second aspect, the RTU can include hardware circuitry for identifying intersections, and the shader can be adapted to identify hits using software. The shader can be configured with instructions executable by the processor core to shade pixels in 3D computer graphics.

[0010] In some embodiments of this second aspect, the RTU can include hardware circuitry for implementing traversal logic to traverse the acceleration structure. The RTU can include hardware circuitry for implementing stack management of a stack used in the traversal of the acceleration structure. Additionally, the RTU can include hardware circuitry for sorting intersections according to the distance from the origin.

[0011] In some implementations, the RTU is adapted to identify intersection points asynchronously with shader identity hits. The shader can include instructions executable by a processor core to read the state of the RTU.

[0012] In a non - limiting implementation, the RTU can include hardware circuitry for transforming the coordinate space used by a higher level of an acceleration structure having multiple levels to the coordinate space used by a lower level of the acceleration structure. The RTU can also include hardware circuitry for transforming the coordinate space used by a lower level of an acceleration structure having multiple levels to the coordinate space used by a higher level of the acceleration structure and / or for restoring light attributes to the light attributes used when traversing the higher level of the acceleration structure.

[0013] In some examples, the RTU can include hardware circuitry for identifying a first intersection point between a first ray and a first bounding volume contained within an acceleration structure, and the shader can include instructions executable to determine whether to ignore the first intersection point and, in response to determining not to ignore the first intersection point, identify the position of the first intersection point along the first ray.

[0014] The processor core and the RTU can be supported on a common semiconductor die. Multiple processor cores and multiple RTUs can be on a common semiconductor die.

[0015] In another aspect, a component includes at least one processor core adapted to execute at least one shader to color pixels in a graphics. The component also includes at least one ray - tracing unit (RTU) separate from the processor core. The RTU includes hardware circuitry for identifying intersection points of rays and objects represented in an acceleration structure for the processor core to identify hits associated with the intersection points, implementing logic for traversing the acceleration structure, and implementing management of a data stack used when traversing the acceleration structure.

[0016] In another aspect, a method for graphics processing includes executing a shader program on a graphics processing unit (GPU), the shader program performing ray - tracing of a 3D environment represented by an acceleration structure. The method further includes: asynchronously with the shader program, using a hardware - implemented ray - tracing unit (RTU) of the GPU to traverse the acceleration structure upon request within the shader program, and using the results of the acceleration - structure traversal at the shader program.

[0017] In another aspect, a method for graphics processing includes executing a shader program on a graphics processing unit (GPU), the shader program performing ray tracing of a 3D environment represented by an acceleration structure. The method further includes: traversing, using a hardware-implemented ray tracing unit (RTU) of the GPU at the request of the GPU intrinsic shader program, the acceleration structure, and using the result of the acceleration structure traversal at the shader program. The acceleration structure is a hierarchical structure having multiple levels, and the acceleration structure traversal performed by the RTU includes handling coordinate transformations between the multiple levels of the acceleration structure.

[0018] In another aspect, a graphics processing unit (GPU) includes: at least one processor core adapted to execute a software-implemented shader; and at least one hardware-implemented ray tracing unit (RTU) separate from the processor core and adapted to traverse an acceleration structure asynchronously with respect to shader operations to identify intersection points of rays and objects represented in the acceleration structure, to generate a result, and to return the result to the shader for the shader to identify a hit associated with the intersection point.

[0019] In another aspect, a graphics processing unit (GPU) includes at least one processor core adapted to execute a software-implemented shader; and at least one hardware-implemented ray tracing unit (RTU) separate from the processor core and adapted to traverse an acceleration structure to identify intersection points of rays and objects represented in the acceleration structure, to generate a result, and to return the result to the shader for the shader to identify a hit associated with the intersection point. The RTU includes hardware circuitry for implementing traversal logic to traverse the acceleration structure, the traversal including modifying at least one ray to account for changes in the coordinate system during traversal.

[0020] Details of both the structure and operation of the present application can be best understood with reference to the accompanying drawings, in which like reference numerals refer to like parts, and in which: BRIEF DESCRIPTION OF THE DRAWINGS

[0021] Figure 1 An acceleration structure is shown;

[0022] Figure 2 Other details of a multi-level acceleration structure are shown;

[0023] Figure 3 A simplified graphics processing unit (GPU) is shown;

[0024] Figure 4 An exemplary GPU with a ray tracing unit (RTU) and a traversal diagram are shown, the ray tracing unit having hardware circuitry for identifying ray intersection points, traversal logic for traversing the acceleration structure, and stack management circuitry;

[0025] Figure 4A , Figure 4B and Figure 4C illustrate, in an exemplary flowchart format, exemplary logic consistent with Figure 4 ;

[0026] Figure 5 illustrate two exemplary GPUs that perform asynchronous processing between a shader and an RTU;

[0027] Figure 5A illustrate, in an exemplary flowchart format, exemplary logic consistent with Figure 5 ;

[0028] Figure 6 illustrate a GPU where a traversal diagram illustrates cooperative processing when a hit test is not required;

[0029] Figure 7 illustrate other examples involving multi-level acceleration structures;

[0030] Figure 7A illustrate, in an exemplary flowchart format, exemplary logic consistent with Figure 7 ; and

[0031] Figure 8 illustrate a GPU where a traversal diagram illustrates intersection points determined by a shader. DETAILED DESCRIPTION

[0032] The present disclosure generally relates to computer ecosystems, which include various aspects of a network of consumer electronics (CE) devices, such as, but not limited to, computer game networks. Systems herein may include server and client components that can be connected via a network such that data can be exchanged between the client and server components. The client components may include one or more computing devices, the computing devices including, such as Sony Game consoles such as those made by Microsoft, Nintendo, or other manufacturers, virtual reality (VR) headsets, augmented reality (AR) headsets, portable televisions (such as smart TVs, Internet-enabled TVs), portable computers (such as laptop computers and tablet computers), and other mobile devices (including smartphones and additional examples discussed below). These client devices can operate in a variety of operating environments. For example, some client computers can employ, for example, the Linux operating system, an operating system from Microsoft, or the Unix operating system, or an operating system produced by Apple, Inc. or Google. These operating environments can be used to execute one or more browsing programs, such as browsers made by Microsoft, Google, or Mozilla, or other browser programs that can access websites hosted by the Internet servers discussed below. Moreover, one or more computer game programs can be executed using an operating environment in accordance with the principles of the present invention.

[0033] The server and / or gateway can include one or more processors that execute instructions configuring the server to receive and transmit data over a network such as the Internet. Alternatively, the client and server can be connected via a local intranet or a virtual private network. The server or controller can be instantiated by a game console (such as Sony ), a personal computer, and the like.

[0034] Information can be exchanged between the client and the server over the network. To this end, and for security purposes, the server and / or client can include a firewall, a load balancer, temporary storage devices, and a proxy, as well as other network infrastructure for reliability and security. One or more servers can form a device that implements a method for providing a secure community (such as an online social networking site) to network members.

[0035] The processor or processor core can be a single-chip or multi-chip processor that can perform logic by means of various lines (such as address lines, data lines, and control lines) as well as registers and shift registers. The logic can include various forms of flowcharts herein to represent, but not limit the principles of the present invention. For example, state logic can be used when appropriate.

[0036] The components included in one embodiment can be used in other embodiments in any suitable combination. For example, any of the various components described herein and / or depicted in the figures can be combined, interchanged, or excluded from other embodiments.

[0037] "A system having at least one of A, B, and C" (similarly, "a system having at least one of A, B, or C" and "a system having at least one of A, B, and C") includes the following systems: having only A; having only B; having only C; having both A and B; having both A and C; having both B and C; and / or having A, B, and C simultaneously, etc.

[0038] Now specifically referring to Figure 1 , an acceleration structure 10 is shown, which is a data structure representing a three-dimensional (3D) environment of computer-generated objects, such as those that can be used in movies, computer simulations (such as computer games), etc. Figure 1 The architecture of the acceleration structure 10 shown in the example of is a bounding volume hierarchy (BVH), which is a tree structure having a root node 12, internal nodes 14, and leaves 16. The root node 12 and internal nodes 14 contain a plurality of bounding volumes, each of which corresponds to a child node of the node. The leaves 16 contain one or more primitives or other types of geometries.

[0039] For simplicity, Figure 2 an exemplary multi-level acceleration structure 200 having two levels is shown. The upper layer ("top-level acceleration structure") 202 does not contain primitives in its leaves 204; instead, its leaves 204 are references to the lower layer ("bottom-level acceleration structure") 206. The bottom-level acceleration structure 206 differs from the top level 202 in that the leaves 208 of the bottom level 206 contain primitives.

[0040] The top-level acceleration structure 202 may have bounding volumes (and in some applications, primitives) in world space coordinates. The bottom-level acceleration structures 206 each have their own coordinate spaces. This allows, for example, a 3D object represented in the bottom-level acceleration structure "Y" to appear twice, each time at a different position and orientation.

[0041] Figure 1 and Figure 2 shows non-limiting examples of acceleration structures that can be used in accordance with the principles of the present invention.

[0042] Figure 3 is a simplified diagram of a graphics processing unit (GPU) 300, which includes a semiconductor die 302 supporting one or more processor cores 304 (a plurality of processor cores 304 in the example) and one or more (a plurality in the example) intersection engines 306. The intersection engines assist in ray tracing. The GPU 300 may also include many other functional units, such as a cache 308, a memory interface 310, a rasterizer 312, and a rendering backend 314.

[0043] The processor core 304 executes a software-implemented shader program (also referred to herein as a "shader") to generate a light beam by initializing the light beam and then traversing the light beam. Figure 2 The ray is directed through the 3D environment represented by, for example, the acceleration structure 200 by successively colliding the ray with the bounding volumes and primitives contained in the acceleration structure. In a simplified embodiment, the intersection engine 306, which may be implemented by dedicated hardware, may calculate the intersection points of the ray with the bounding volumes and primitives. Thus, the identification of ray-bounding volume and ray-primitive intersections may be offloaded to the intersection engine 306. Figure 3 In a simplified concept of , the processor core 304 transmits the node or leaf address along with the description of the ray to the intersection engine 306, and after the intersection engine calculates the intersection points between the ray and the bounding volume or primitive, the intersection engine returns the results to the processor core.

[0044] However, as referenced in this article Figure 3 As will be appreciated, much of the processing is handled by the shader program, which must perform stack management, keep track of ray states, keep track of intermediate results, etc. As will be further appreciated herein, in practice, since much of the logic and calculations are handled directly by the shader program, performance can be lower than required. It is in this context that the Figure 4 And the following and so on.

[0045] Figure 4 A GPU 400 on a common semiconductor die 402 is shown that can be implemented in a rendering device 404, which can be (but not limited to) a computer simulation or game console, a computer server that streams games to end users, a device associated with computer enhanced graphics in movies, etc. The GPU 400 includes one or more (in the example shown, multiple) processor cores 406 for executing one or more shader programs and for communicating with one or more (in the example shown, multiple) ray tracing units (RTUs) 408. The RTU is implemented by hardware with hardware circuits for performing the tasks disclosed herein. The RTU 408 can use its own hardware-implemented traversal logic to traverse the acceleration structure. The RTU 408 may include: one or more intersection engine circuits 410, which can calculate the intersection of rays with bounding volumes or primitives (such as triangles); traversal logic circuits 412, which are used to traverse the acceleration structure; and stack management circuits 414, which are used to maintain stacks used in the traversal of the acceleration structure. RTU 408 may also include other sub-units.

[0046] Figure 4A shows the high-level logic, and Figure 4BIllustrates a coarser-grained logic that can be implemented by the hardware circuitry in the RTU 410. Starting from block 414, the shader executed on the processor core passes the root node of the acceleration structure and ray information to the RTU 410. Using this information, at block 416, the RTU traverses the acceleration structure using its traversal logic circuitry 412 and stack management circuitry 414 to identify the intersection points of the ray with the bounding volumes and primitives. At block 418, one or more intersection points are passed to the shader. At block 420, the shader identifies the hits of any intersection points and at block 422 uses the overall information obtained thereby to color the pixels for rendering.

[0047] In some embodiments, the RTU may include circuitry for performing Figure 4B the stack management and traversal of the acceleration structure shown. If at state 424 it is determined that the current node is a root node or an interior node, the RTU moves to block 426 to check for intersection points of the ray with each of the bounding volumes contained in the current node. If at state 428 it is determined that there are multiple intersection points, then at block 430 the multiple intersection points are sorted such that the earliest (i.e., closest to the ray origin) intersection point is the first intersection point and the remaining intersection points are in order of distance from the ray origin. Basically, the intersection points (starting from the origin of the ray) are sorted from shortest to longest. If there is more than one intersection point, then at block 432 the child nodes corresponding to the second and subsequent intersection points are pushed onto the stack. Depending on block 432 or depending on state 428, if the test is negative, the circuitry moves to block 434 to continue processing the child node corresponding to the first intersection point, or if there are no intersection points, to continue processing the node popped from the top of the stack. If there are no nodes to pop because the stack is empty, the acceleration structure traversal is complete. If at state 424 it is determined that the current node is a leaf, then at block 436, the RTU checks for intersection points of the ray with one or more primitives contained in the leaf.

[0048] In other embodiments, the traversal can be stackless.

[0049] Return Figure 4 The lower half of, illustrates an exemplary traversal using the strategy described above. The thick arrows illustrate the traversal of the acceleration structure. The thick solid boxes depict the nodes where the RTU has found an intersection point between the ray and the bounding volume of the node, and the thick dashed boxes depict the nodes where the RTU has found an intersection point between the ray and the primitives contained in the leaf. The leaves not marked with thick dashed lines in the traversal are the leaves where the RTU has not found an intersection point between the ray and the primitives contained in the leaf. The specific steps of the exemplary traversal are as follows.

[0050] As shown at 438, traversal begins by processing the root node A to identify the intersection point of the ray with the bounding volume contained by the root node. As shown at 440 and 442, the bounding volumes corresponding to child node E and child node J of A are identified as intersecting through the processing of the root node A, and thus E and J are pushed onto the stack. Additionally, through the processing of the root node A, the bounding volume corresponding to child node B of A is also identified as intersecting, and at 444, B is processed to identify the intersection point of the ray with the bounding volume of B. This in turn identifies primitive D, which is pushed onto the stack at 446, and primitive C, which is processed at 448. In the example shown, it is determined through this processing that there is no intersection point between the ray and primitive C. Primitive D is popped from the stack, and at 450 it is processed to identify the intersection point with the ray; in this exemplary case, the intersection point between the ray and primitive D is identified, and the primitive is passed to the shader program to identify whether there is a "hit" between the ray and the primitive (as described below and Figure 4C as described in).

[0051] As shown at 452, the bounding volume E is then popped from the stack and processed to identify the intersection point. As shown at 454, primitives H and I identified from processing the bounding volume E are pushed onto the stack, and at 456 the bounding volume F is processed to identify the intersection point. Then at 458 primitive G is processed, the intersection point is identified, and primitive G is passed to the shader program for a hit test. At 460 the next primitive H is popped from the stack and processed, and no intersection point is found. Next, at 462 primitive I is popped from the stack and processed (the intersection point is identified and primitive I is passed to the shader program for a hit test), and then at 464 the bounding volume J is popped from the stack and processed. This results in the identification of primitive K, which is processed at 466; no intersection point is identified.

[0052] The shader program running on the processor core collaborates with the RTU to perform ray tracing. The above example shows the processing when the primitives are partially transparent (e.g., they are triangles representing leaves). The purpose of the shader program is to identify the earliest (i.e., closest to the origin of the ray) intersection point ("hit") with the opaque parts of the 3D environment represented by the acceleration structure.

[0053] In Figure 4C the example of shows the communication between the shader program and the RTU to implement the above. Initially, the shader program passes the root node A and ray information to the RTU, which traverses the acceleration structure until it reaches the intersection point between the ray and the primitive (in the example of Figure 4 is primitive D) (in this example, there is no intersection point between the ray and primitive C).

[0054] In Figure 4CStarting from box 470, the RTU passes primitive D to the shader program. At state 472, the shader identifies whether the ray hits the opaque part of primitive D, which means the shader tests primitive D to determine whether the ray passes through the transparent part of the primitive or hits the solid part of the primitive. In this example, the ray hits the solid part of the primitive. The shader program records primitive D as the earliest intersection point encountered so far and passes the result of the intersection point ("hit") to the RTU at box 474.

[0055] Moving to box 476, the RTU shortens the ray because no point tests beyond the intersection point of the ray and primitive D. Box 478 indicates that the RTU continues to traverse the acceleration structure, reaches primitive G, and determines that the ray intersects it. The RTU passes G to the shader program for a hit test.

[0056] The shader program performs a hit test consistent with Figure 4 and Figure 4C to determine that the ray passes through the transparent part of primitive G and thus there is no hit. In some embodiments, the shader program passes the information that G is not hit to the RTU. In other embodiments, due to the miss, the shader program does not pass the information that primitive G is not hit to the RTU.

[0057] As discussed with respect to Figure 4 the RTU continues to traverse and reaches primitive H. If the ray was not shortened due to the "hit" on primitive D, an intersection point would be detected, but when the ray has been shortened, there is no intersection point between the ray and primitive H. The RTU continues to traverse and reaches primitive I, where an intersection point between the ray and the primitive is detected. It passes primitive I to the shader program that performs the hit test. In this example, the shader program determines that there is a hit and the hit on primitive I is earlier than the hit on primitive D (i.e., closer to the origin of the ray); thus, the shader program updates the earliest intersection point encountered to primitive I. The shader program also notifies the RTU that there is a hit on primitive I. Again, the RTU shortens the ray based on the hit on primitive I and continues the traversal of the acceleration structure, reaching primitive K. There is no intersection point between the ray and primitive K, so the RTU attempts to pop the next node to be processed from the stack, but the stack is empty. Therefore, the RTU terminates the processing (it has reached the end of the acceleration structure traversal) and notifies the shader program that it has finished. The shader program now knows that the earliest hit is on primitive I and continues processing accordingly.

[0058] Each step of this communication between the processor core and the RTU can be in Figure 4In the upper part of [the figure]. The processor core passes the root node A to the RTU; the RTU passes primitive D to the shader program for a hit test, and the shader program reports a hit; the RTU passes primitive G to the shader program for a hit test, and the shader program reports a miss; the RTU passes primitive I to the shader program for a hit test, and the shader program reports a hit; the RTU notifies the shader program that the acceleration structure traversal is complete.

[0059] The above processing strategy can lead to a significant increase in ray tracing speed because the shader program only performs hit tests. It does not perform acceleration structure traversal or manage the corresponding stack.

[0060] In the above, phrases such as "pass node A" and "pass primitive G" describe any strategy for communication, including but not limited to passing pointers to nodes or primitives, or passing the IDs of nodes or primitives. In the above, phrases such as "the RTU notifies the shader program" or "the shader program notifies the RTU" also refer to any strategy for communication, including but not limited to "push" strategies such as setting registers and ringing doorbells, or interrupt-driven communication, and "pull" strategies such as reading or polling the status of other units.

[0061] Now turning to Figure 5 and Figure 5A , it shows the asynchronous operation of the RTU relative to the shader, and the strategy for all aspects of the shader program running on the processor core to initiate communication between the shader program and the RTU. See Figure 5 In the upper part of [the figure], the GPU 500 includes one or more processor cores 502 and one or more RTUs 504 that operate asynchronously with the shader program. The GPU 500 can be substantially the same as the GPU 400 shown in Figure 4 , except as described below.

[0062] Figure 5A At box 506, it shows the shader program sending the root node A and ray information to the RTU to initiate processing. At box 508, the RTU starts traversing the acceleration structure asynchronously with the shader operation. As the RTU continues to traverse, at box 510, the shader periodically reads the status of the RTU, and at box 512, the RTU reports the status. At box 514, the shader passes the hit identification to the RTU to enable the RTU to shorten its ray.

[0063] The above operations are reflected in Figure 5In it, the arrow pointing to the left is a request for information by the shader program, and the arrow pointing to the right is the transmission of information from the shader program to the RTU. More specifically, as shown at 516, the shader program sends the root node A to the RTU and reads the status from the RTU, and receives the status "WIP" at 518, which means that the RTU has not found any intersection points yet. The shader program reads the status again and receives the status "D" at 520, which means that the RTU has found an intersection point with primitive D.

[0064] The shader program performs a hit test, finds that primitive D is hit by the ray, and notifies the RTU of this hit at 522, so that the RTU can shorten the ray. The read / report process continues as the RTU and the shader read asynchronously traverse the acceleration structure.

[0065] In Figure 5 The lower half is an example of another implementation, where the RTU sorts the intersection points by distance. This can lead to an improvement in ray tracing speed. In this example, as mentioned before, there are three intersection points between the ray and the primitive. The intersection point closest to the ray origin is "I". It is the hit shown at 524. The next intersection point closest to the ray origin is "G". This is a miss. The intersection point farthest from the ray origin is "D". This is the hit as mentioned before.

[0066] But now look at Figure 5 the lower half of Figure 5 The first step is the same as the figure at the top of

[0067] When the shader program is performing a hit test, the RTU traversal determines that there are intersection points with primitive G and primitive I, and primitive I is closer to the ray origin than primitive G (i.e., the RTU sorts the intersection points by distance from the ray origin). When the shader program next reads the state, at 532 it is notified that the hit test should be performed on primitive I (instead of G) because primitive I is the known intersection point closest to the ray origin. As shown at 534, the shader program performs the hit test, finds that primitive I is hit by the ray, and notifies the RTU of the hit. When the shader program performs the hit test, the RTU completes the traversal of the acceleration structure without finding any more intersection points. Since the intersection point of primitive G is farther from the ray origin than that of primitive I, primitive G is discarded. When the shader program reads the state next time, at 536 it is notified that the traversal of the acceleration structure is complete and there are no more primitives for which a hit test is required. Such sorting of intersection points is also valuable when performing ray tracing of an environment with transparency; in this case, it is beneficial to sort the intersection points so that the shader program is first notified of the intersection point farthest from the ray origin.

[0068] Figure 6 By showing the acceleration traversal performed by the RTU in FIG. 600 and the communication diagrams 602, 604 between one or more processor cores 606 of a GPU (such as any of the GPUs described in Figure 4 or Figure 5 or elsewhere in this document) and one or more RTUs 608, an example of cooperative processing when no hit test is required is shown. As understood herein, if no hit test is required, such as when the primitives in the acceleration structure are opaque, then the processing performed by the shader program is further reduced, resulting in performance improvement. Primitives can be indicated as opaque to the RTU by a flag or other means.

[0069] In Figure 6 where the primitives are opaque, the RTU can track the earliest intersection points between the rays and the primitives without requiring the shader program to perform a hit test. The acceleration traversal diagram 600 is substantially the same as the example shown in Figure 4 but in Figure 6 the leaves that intersect the rays are shown with thick solid boundaries instead of thick dashed boundaries. Thus, as shown in Figure 4 there are three intersection points with the primitives, in order from earliest (closest to the ray origin) to farthest as I - G - D.

[0070] Figure 6The middle communication diagram 602 shows the collaborative ray tracing between the shader program and the RTU when locating the earliest hit. Since the RTU can track the earliest intersection point, it reports the "Work In Progress" ("WIP") status until the traversal of the acceleration structure is completed, at which point it reports the "Completed" status as shown at 612 and provides I as the earliest primitive intersected by the ray.

[0071] Figure 6 The bottom communication diagram 604 shows the collaborative ray tracing between the shader program and the RTU when the shader program wants to know if there is an intersection point but does not need to know the details of the intersection point; this is typical for ray tracing of shadows and surrounding occlusions. The shader processing ends once a primitive intersected by the ray is located, and in this case the primitive is D, as shown at 614.

[0072] Figure 7 shows an acceleration structure 700 having a top-level acceleration structure (TLAS) 702 and a set of bottom-level acceleration structures (BLAS) 704 similar to that shown in Figure 2 as described below. Figure 7 Also shown is a communication diagram between one or more processor cores 706 of the GPU and one or more RTUs 708 to show an example of collaborative processing between the shader program and the RTU when the acceleration structure is a multi-level hierarchy.

[0073] The TLAS 702 has leaves (such as X 710), each leaf giving a link to a corresponding BLAS (such as BLAS X, one of the set of BLASs 704). As previously mentioned, the TLAS 702 uses world space coordinates. Each BLAS in the set of BLASs 704 has a leaf (such as D 712) containing a primitive, and also as previously discussed, each BLAS in the set of BLASs 704 has its own coordinate space. In this example, the goal of the shader program is to determine the earliest intersection point, and all primitives in the acceleration structure are opaque.

[0074] The desired traversal of the acceleration structure 700 implemented by the RTU 708 is as follows. Processing the root node A of the TLAS 720 identifies the intersection points of the ray with the bounding volumes representing the child nodes B and E; E in the TLAS 702 is pushed onto the stack maintained by the RTU in this case, and the bounding volume contained within B in the TLAS 702 is processed to identify the ray intersection point. The intersection point of the ray with the bounding volume corresponding to the leaf X within B is found; this in turn leads to the processing of the leaf X, which represents BLASX of the set of BLASs 704. The coordinate space of X is different from the world space coordinates (whose primitives have their own coordinate spaces), so ray attributes such as the ray origin must be transformed.

[0075] Process the root node of BLAS X to identify the ray intersection point, which results in the processing of C; when processing C, identify the intersection point with the bounding volume of leaf node D 712. In the example shown, the processing of leaf node D 712 identifies the intersection point with primitive D, and recall that since the primitives in Figure 7 are assumed to be opaque, the intersection point is automatically considered a hit. Since D is considered a hit, the RTU shortens the ray to the length from the ray origin to D.

[0076] Next, E in TLAS 702 is popped from the stack. The coordinate space of the BLAS part X will no longer be used, so the ray properties in world space must be restored. The ray length must be preserved during this process. Process E to identify the ray intersection point, and identify the intersection point for the bounding volume corresponding to leaf Z; this in turn results in the processing of leaf Z, which represents BLAS Z of the set of BLASs 704. Again, since the coordinate space for Z is different from the world space coordinates, the ray properties (such as the ray origin) must be transformed to the coordinate space of Z.

[0077] Process the root node of BLAS Z to identify the ray intersection point, which results in the processing of F, and then the processing of leaf G, which is identified as the intersection point in the example shown (and thus, is a hit in the "opaque" example shown). The ray is shortened to the length between the origin and primitive G.

[0078] In one embodiment, the RTU detects transitions from TLAS 702 to BLASs within the set of BLASs 704 and from BLASs within the set of BLASs 704 to TLAS 702, and the shader program performs ray transformation updates and passes the results to the RTU. In this example, the communication steps are as Figure 7 shown, starting at 714, where the shader program sends the root node A and ray information to the RTU to initiate processing.

[0079] As shown at 716, the shader program reads the status from the RTU and receives the "work in progress" ("WIP") status, which means the RTU has not found any intersection points yet. The shader program reads the status again, and at 718 receives the status "entering BLAS X", which indicates that the RTU has detected a transition to BLAS X in its traversal of the acceleration structure 700. The shader program transforms the ray properties (such as the origin) to the BLAS X coordinate space, and sends (720) the ray properties and the BLAS root node X to the RTU.

[0080] The RTU traverses BLAS X and finds the intersection point between the ray and primitive D and shortens the ray accordingly. Recall that this example assumes the primitive is opaque, so no shader unit is required to perform a hit test. The shader program reads the status and receives the status "exit BLAS" at 722, which indicates that the RTU processing of the BLAS is complete. As shown at 724, the shader program sends the world space ray attributes to the RTU.

[0081] The shader program reads the status again and, as shown at 726, receives the status "enter BLAS Z". The shader program transforms the ray attributes (such as the origin) into the BLAS Z coordinate space and sends the ray attributes and the BLAS root node Z to the RTU at 728. The RTU traverses BLAS Z and finds the intersection point between the ray and primitive G and shortens the ray accordingly. When the shader program next reads the status, it is notified at 730 that the traversal of the acceleration structure is complete and that the earliest intersection point is the intersection with primitive G.

[0082] In another embodiment, the RTU can handle the detected transitions from the TLAS to the BLAS and from the BLAS to the TLAS, updating the ray attributes as needed. In this case, in the Figure 7 example, all processing is performed by the RTU, so after the shader program sends the root node A to the RTU, the shader program will read the status "WIP" until the RTU traversal of the acceleration structure is complete, at which point the shader program will read the status "complete" and the earliest intersection point is the intersection with primitive "G".

[0083] Figure 7A The above coordinate transformation logic according to the logic that can be performed entirely by the RTU or by the cooperation between the shader and the RTU is shown. This example shows an acceleration structure with two levels. Starting from box 732, the determination of the ray-bounding volume intersection is initially performed in world space. When the RTU traverses the acceleration structure, if it is determined at status 734 that the RTU has traversed to a BLAS in another coordinate space, the logic moves to box 736 to transform the ray into the coordinate space dedicated to the BLAS and processes the ray at box 738 for intersection point identification. Similarly, when the RTU continues to traverse the acceleration structure, if it is determined at status 740 that the RTU has traversed to the TLAS (which is described in world space), the logic moves to box 742 to transform the ray into world space (or restore the ray attributes if the ray attributes are in world space) and processes the ray at box 744 for intersection point identification. In the above variant, the TLAS may not be in world coordinate space but in its own specific coordinate space. In other variants, the acceleration structure has multiple levels; the processing is similar to Figure 7Aperformed, converting the ray to a new coordinate space when traversing to lower levels of the acceleration structure and restoring or transforming it when traversing to higher levels of the acceleration structure.

[0084] Now turning to Figure 8 , an example of an intersection point determined by a shader executed on one or more processor cores 800 communicating with one or more RTUs 802 of a GPU is shown. In the previous example, the RTU determines whether the ray under test intersects a primitive, e.g., the primitive is a triangle, and the intersection point engine of the RTU can determine the intersection point between the ray and the triangle. In contrast, in Figure 8 , the geometry associated with leaf N 804 in the acceleration structure 806 has a geometry that makes it impossible for the intersection point engine of the RTU to calculate the intersection point, e.g., a sphere.

[0085] Thus, in some embodiments, the shader program and the RTU cooperate to perform ray tracing as described below. As described in other parts of this document, the RTU traverses the acceleration structure 806. The RTU tests the ray against the bounding volume in node M 808 and determines that the ray intersects the bounding volume corresponding to leaf N 804. When the shader program reads the state, it receives at 810 the state that the bounding volume of N has been intersected. The shader program performs a hit test between the ray and the sphere contained in leaf N and determines that a hit has occurred. At 812, it notifies the RTU that a hit has occurred and also notifies the RTU of the location of the hit so that the ray can be shortened accordingly by the RTU. The fact that the primitive in leaf N is a sphere can be identified by the RTU and / or the shader based on a flag or other indicator associated with the primitive being a sphere (or other geometry beyond the processing capabilities of the RTU).

[0086] It will be appreciated that while the principles of the present invention have been described with reference to some exemplary embodiments, these embodiments are not intended to be limiting, and various alternative arrangements can be used to implement the subject matter claimed herein.

Claims

1. A method for graphics processing, comprising: Executing a shader program on a graphics processing unit (GPU), the shader program performing ray tracing of a 3D environment represented by an acceleration structure; Asynchronously with the shader program, using the GPU to traverse, at the request of the shader program, a hardware-implemented ray tracing unit (RTU) of the acceleration structure; And Using the result of the acceleration structure traversal at the shader program, wherein the shader program is configured to send a root node and ray-related information to the RTU to initiate processing, in which the RTU is configured to start traversing the acceleration structure asynchronously with the operation of the shader program, wherein the shader program periodically reads at least one state of the RTU, and the RTU reports the at least one state as it continues traversing, the shader program passes a hit identifier to the RTU to enable the RTU to shorten the ray, at least one state from the RTU to the shader program indicates that the RTU has not found any intersection points, at least a second state from the RTU indicates that the RTU has found an intersection point with a first primitive, the shader program performs a hit test and in response to finding that the first primitive is hit by the ray, notifies the RTU that the hit enables the RTU to shorten the ray.

2. The method according to claim 1, wherein using the result includes coloring pixels in a computer-generated graphic.

3. The method according to claim 1, wherein the result of the acceleration structure traversal by the RTU includes detection of an intersection point between a first ray and a bounding volume contained in the acceleration structure, and / or detection of an intersection point between a second ray and a primitive contained in the acceleration structure.

4. The method according to claim 1, wherein the RTU processing includes maintaining a stack used in the acceleration structure traversal.

5. The method according to claim 1, wherein the result of the acceleration structure traversal by the RTU includes sorting of the intersection points by the RTU according to the distance of the intersection points detected by the RTU from the ray origin, such that: The RTU detects a first intersection point between a first ray and a primitive as it traverses the acceleration structure; And The RTU detects a second intersection point between the first ray and a primitive as it traverses the acceleration structure; And When transmitting the result from the RTU to the shader program, the second intersection point result is transmitted before the first intersection point result that is also transmitted to the shader program.

6. The method according to claim 3, wherein the result of the acceleration structure traversal by the RTU includes detection of the earliest intersection point between a first ray and a primitive contained in the acceleration structure.

7. A method for graphics processing, comprising: Executing a shader program on a graphics processing unit (GPU), the shader program performing ray tracing of a 3D environment represented by an acceleration structure; Using the GPU to traverse, at the request of the shader program, a hardware-implemented ray tracing unit (RTU) of the acceleration structure; And Use the result of traversing an acceleration structure at the shader program, where the acceleration structure is a hierarchical structure with multiple levels, where the RTU identifies the intersection points of rays with elements in the acceleration structure, indicates the intersection points to the shader program, and the shader program performs a hit test to determine whether the ray passes through the transparent part of the element or hits the opaque part of the element. The RTU sorts the intersection points by distance. The shader program receives the status that the RTU has found an intersection point with a first primitive. The shader program performs a hit test on the first primitive and, in response to determining that the first primitive is hit by the ray, notifies the RTU so that the RTU can shorten the ray. The RTU determines the intersection points with a second primitive and a third primitive, where the third primitive is closer to the ray origin than the second primitive. The shader program accesses the RTU information to perform a hit test on the third primitive but not on the second primitive.

8. The method according to claim 7, wherein using the result includes coloring pixels in a computer-generated graphic.

9. The method according to claim 7, wherein the result of the acceleration structure traversal by the RTU includes detection of transitions from higher levels to lower levels within the multiple levels of the acceleration structure.

10. The method according to claim 7, wherein the result of the acceleration structure traversal by the RTU includes detection of transitions from lower levels to higher levels within the multiple levels of the acceleration structure.

11. A graphics processing unit (GPU) comprising: At least one processor core adapted to execute a software-implemented shader; And At least one hardware-implemented ray tracing unit (RTU) separate from the processor core and adapted to traverse an acceleration structure asynchronously relative to shader operations to identify intersection points of rays with objects represented in the acceleration structure to generate a result and return the result to the shader for the shader to identify a hit associated with the intersection point, Wherein the shader is configured to send root node and ray information to the RTU to initiate processing, in which the RTU is configured to start traversing the acceleration structure asynchronously with respect to shader operations. The shader is configured to periodically read at least one status of the RTU. The shader is configured to pass a hit identification to the RTU to enable the RTU to shorten the ray. At least a first status from the RTU indicates that the RTU has not found any intersection points. At least a second status from the RTU indicates that the RTU has found an intersection point with a first primitive. The shader is configured to perform a hit test and, in response to finding that the first primitive is hit by the ray, the shader is configured to notify the RTU of the hit so that the RTU can shorten the ray.

12. The GPU as claimed in claim 11, wherein the RTU includes hardware circuitry for identifying the intersection points, and the shader is adapted to identify the hit using software.

13. The GPU as claimed in claim 11, wherein the shader is configured with instructions executable by the processor core to shade pixels in 3D computer graphics.

14. The GPU as claimed in claim 11, wherein the RTU includes hardware circuitry for implementing traversal logic to traverse the acceleration structure.

15. The GPU as claimed in claim 11, wherein the RTU includes hardware circuitry for implementing stack management of a stack used in traversal of the acceleration structure.

16. A graphics processing unit (GPU) comprising: at least one processor core adapted to execute a software-implemented shader; and at least one hardware-implemented ray tracing unit (RTU) separate from the processor core and adapted to traverse an acceleration structure to identify intersection points of a ray with objects represented in the acceleration structure to generate results, and return the results to the shader, the shader being configured to receive a first status that the RTU has found an intersection point with a first primitive, the shader being configured to perform a hit test on the first primitive, and in response to determining that the first primitive is hit by the ray, the shader being configured to notify the RTU such that the RTU can shorten the ray, the RTU being configured to determine intersection points with a second primitive and a third primitive, the third primitive being closer to the ray origin than the second primitive, the shader being configured to access RTU information to perform a hit test on the third primitive but not on the second primitive.

17. The GPU as claimed in claim 16, wherein the RTU includes hardware circuitry for identifying the intersection points, and the shader is adapted to identify the hit using software.

18. The GPU as claimed in claim 16, wherein the shader is configured with instructions executable by the processor core to shade pixels in 3D computer graphics.

19. The GPU as claimed in claim 16, wherein the RTU includes hardware circuitry for implementing stack management of a stack used in traversal of the acceleration structure.

20. The GPU as claimed in claim 16, wherein the RTU includes hardware circuitry for transforming at least a first ray from world space to a coordinate space corresponding to a lower level of a multi-level acceleration structure.

21. The GPU as claimed in claim 16, wherein the RTU includes hardware circuitry for transforming at least a first ray from a coordinate space corresponding to a lower level of a multi-level acceleration structure to world space.

Citation Information

Patent Citations

  • Hierarchy Merging

    US20170270146A1

  • Robust, efficient multiprocessor-coprocessor interface

    US20200050451A1