Processing Device and Processing Method for Ray Tracing Acceleration Structure

By using the ray tracing acceleration structure of TLAS and BLAS, traversing TLAS and BLAS through descriptors and pointers, the problem of low ray tracing efficiency in the prior art is solved, and efficient ray intersection testing is achieved, suitable for real-time image rendering.

CN115222868BActive Publication Date: 2025-07-04SHANGHAI BIREN TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210837376.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-07-15
Publication Date
2025-07-04
Estimated Expiration
2042-07-15

AI Technical Summary

Technical Problem

Existing ray tracing techniques are inefficient when dealing with complex scenarios, and existing acceleration structures perform intersecting test rates in real-time image rendering are not suitable.

Method used

Using a ray tracing acceleration structure including a top-level acceleration structure (TLAS) and a bottom-level acceleration structure (BLAS), points to TLAS and instance cache through descriptors, and traverses TLAS and BLAS with instance identifiers and pointers to find intersecting nodes, achieving efficient ray tracing.

Benefits of technology

It improves the efficiency of ray tracing, can effectively handle ray intersection tests in complex scenes, and is suitable for real-time image rendering.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115222868B_ABST
    Figure CN115222868B_ABST
Patent Text Reader

Abstract

The present invention provides a processing device and a processing method for a ray tracing acceleration structure. The processing device includes a machine-readable storage medium and a processor. The processor executes a descriptor to simulate the interaction between a ray and a scene, where the descriptor includes a first pointer and a second pointer. The processor obtains a TLAS (top-level acceleration structure) by using the first pointer. The processor traverses the TLAS to find leaf nodes in the TLAS that intersect with the ray, where the intersecting leaf nodes include instance identifiers. The processor obtains an intersecting instance record from an instance cache pointed to by the second pointer by using the instance identifier, where the intersecting instance record includes a third pointer. The processor obtains a BLAS (bottom-level acceleration structure) by using the third pointer. The processor traverses the BLAS to find primitive nodes in the BLAS that intersect with the ray.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to an image rendering technology, and particularly to a processing device, a processing method, and a machine-readable storage medium for a ray tracing acceleration structure. Background Art

[0002] Ray tracing technology can simulate the interaction between light and a scene. For example, ray tracing can be used in a graphics rendering system to generate three-dimensional images. Three-dimensional images typically include a large number of primitives. Primitives are usually triangular primitives, but sometimes can also be other shapes, such as other polygons, lines, or points. Ray tracing can identify the primitives in the scene that intersect with the light rays, and process the identified intersecting primitives (e.g., execute a shader program to process the primitives) to mimic the natural interaction between light and the scene. The intersection test between the light rays and the primitives in the scene involves a lot of processing. Simple ray tracing technology can test each ray against each primitive in the scene. For scenes with millions or even tens of billions of primitives, and applications that require tracing millions of light rays, this simple ray tracing technology is inefficient.

[0003] Therefore, ray tracing technology usually uses an acceleration structure. The acceleration structure can reduce the intersection test. However, even with the existing ray tracing acceleration structure, the rate of performing the intersection test may not be suitable for real-time image rendering. How to process the ray tracing acceleration structure is one of the many topics in this technical field. Summary of the Invention

[0004] The present invention provides a processing device, a processing method, and a machine-readable storage medium for a ray tracing acceleration structure to efficiently process the ray tracing acceleration structure.

[0005] In an embodiment according to the present invention, the processing device includes a machine-readable storage medium and a processor. The machine-readable storage medium includes at least one thread group (thread group or warp), at least one instance buffer, at least one top-level acceleration structure (TLAS), and at least one bottom-level acceleration structure (BLAS). The processor is coupled to the machine-readable storage medium for obtaining a thread group from the machine-readable storage medium and executing the thread group, where the thread group includes at least one descriptor. The processor executes the descriptor to simulate the interaction between light and a scene, where the descriptor includes a first pointer for pointing to the TLAS of the scene and a second pointer for pointing to the instance buffer. The processor obtains the TLAS from the machine-readable storage medium by using the first pointer. The processor traverses the TLAS based on the light to find an intersected leaf node in the TLAS that intersects with the light, where the intersected leaf node includes an instance identifier for pointing to a corresponding instance. The processor obtains the intersected instance records corresponding to the intersected leaf node from the instance buffer pointed to by the second pointer by using the instance identifier, where the intersected instance records include a third pointer for pointing to the BLAS of the scene. The processor obtains the BLAS from the machine-readable storage medium by using the third pointer. The processor traverses the BLAS based on the light to find an intersected primitive node in the BLAS that intersects with the light.

[0006] In an embodiment according to the present invention, the processing method includes: executing a descriptor to simulate the interaction between light and a scene, where the descriptor includes a first pointer for pointing to the TLAS of the scene and a second pointer for pointing to the instance buffer; obtaining the TLAS by using the first pointer; traversing the TLAS based on the light to find an intersected leaf node in the TLAS that intersects with the light, where the intersected leaf node includes an instance identifier for pointing to a corresponding instance; obtaining the intersected instance records corresponding to the intersected leaf node from the instance buffer pointed to by the second pointer by using the instance identifier, where the intersected instance records include a third pointer for pointing to the BLAS of the scene; obtaining the BLAS by using the third pointer; and traversing the BLAS based on the light to find an intersected primitive node in the BLAS that intersects with the light.

[0007] In an embodiment according to the present invention, the machine-readable storage medium is used to store at least one thread group, at least one instance cache, at least one TLAS, and at least one BLAS, wherein the thread group includes at least one descriptor. When the descriptor is executed by a processor, the processing method of the ray tracing acceleration structure can be implemented.

[0008] Based on the above, the descriptor can point to the TLAS and the instance cache through a first pointer and a second pointer, so that the processor can efficiently obtain the content of the TLAS and the instance cache corresponding to the descriptor. After traversing the TLAS, when a ray intersects a leaf node (instance) of the TLAS, the processor can find the intersecting leaf node in the TLAS. Each leaf node includes an instance identifier used to point to the corresponding instance. By using the second pointer and the instance identifier of the intersecting leaf node, the processor can efficiently obtain the intersecting instance record corresponding to the intersecting leaf node from the instance cache. Each instance record in the instance cache includes a third pointer used to point to the corresponding BLAS. By using the third pointer, the processor can efficiently obtain the corresponding BLAS. After traversing the corresponding BLAS, when a ray intersects a leaf node (primitive node) of the BLAS, the processor can find the intersecting primitive node in the BLAS. Therefore, the processor can efficiently process the ray tracing acceleration structure. Description of the Drawings

[0009] Figure 1 is a schematic diagram of a circuit block of a processing device for a ray tracing acceleration structure according to an embodiment of the present invention.

[0010] Figure 2 is a schematic flowchart of a processing method for a ray tracing acceleration structure according to an embodiment of the present invention.

[0011] Figure 3 is a schematic diagram of a top-level data structure according to an embodiment of the present invention.

[0012] Description of the Reference Numerals

[0013] 100: Processing device

[0014] 110: Processor

[0015] 120: Machine-readable storage medium

[0016] 121: Thread group

[0017] 122: Instance cache

[0018] 123: TLAS (Top-Level Acceleration Structure)

[0019] 124: BLAS (Bottom-Level Acceleration Structure)

[0020] 125: Index Cache

[0021] 126: Vertex Cache

[0022] 321: Descriptor

[0023] IID: Instance Identifier

[0024] IR: Intersection Instance Record

[0025] PTR1: First Pointer

[0026] PTR2: Second Pointer

[0027] PTR3: Third Pointer

[0028] PTR4: Fourth Pointer

[0029] PTR5: Fifth Pointer

[0030] S210 - S260: Steps

[0031] TR: Root Node

[0032] TB1, TB2: Branch Nodes

[0033] TL1, TL2, TL3, TL4: Leaf Nodes Detailed Embodiment

[0034] Reference will now be made in detail to the exemplary embodiments of the present invention, examples of which are illustrated in the accompanying drawings. Whenever possible, the same component symbols are used in the drawings and the description to represent the same or similar parts.

[0035] As used throughout this specification (including the claims), the term "coupled (or connected)" may refer to any direct or indirect means of connection. For example, if it is described in the text that a first device is coupled (or connected) to a second device, it should be interpreted that the first device can be directly connected to the second device, or the first device can be indirectly connected to the second device through other devices or some means of connection. The terms "first", "second", etc. mentioned throughout this specification (including the claims) are used to name the elements, and are not used to limit the upper or lower limits of the number of components, nor to limit the order of the components. Additionally, wherever possible, components / components / steps using the same reference numerals in the drawings and embodiments represent the same or similar parts. Components / components / steps using the same reference numerals or the same terms in different embodiments can be referred to each other's relevant descriptions.

[0036] Ray tracing acceleration structures can use Bounding Volume Hierarchies (BVHs) or Bounding Box Hierarchies. In a ray tracing acceleration structure, a large bounding box encloses one or more small bounding boxes, and each bounding box encloses multiple primitives. Based on such a bounding volume hierarchy, intersection tests for ray tracing become easier. If a ray misses a bounding box, no intersection tests need to be performed for any child nodes (primitives) within that bounding box. Thus, ray tracing acceleration structures can reduce intersection tests.

[0037] The ray tracing acceleration structure includes a bottom-level acceleration structure (BLAS) and a top-level acceleration structure (TLAS). The BLAS and TLAS can be BVH trees. The BLAS has leaf nodes that are object primitives. The top level of the BLAS is a single root node. For example, the BLAS can be used to describe the model of a single object in a scene or a group of objects in a scene. The TLAS describes the high-level scene, starting from the root node at the top level and terminating at the lowest-level BLAS. The TLAS can describe multiple instances of the same BLAS. For example, the BLAS can simulate a single chair, while the TLAS can simulate a concert hall that includes hundreds of chairs (instances), with each instance representing a different chair in different positions and / or orientations within the concert hall. Intersection tests are performed by traversing the BVH trees (BLAS and TLAS). If a given ray "hits" a bounding box (node), the ray needs to be tested against each child node of that bounding box (node). This continues down through the BVH tree until at least one primitive (leaf node) is hit, or the ray misses all child nodes of an intersection node.

[0038] Figure 1 It is a schematic diagram of a circuit block of a processing device 100 for a ray tracing acceleration structure according to an embodiment of the present invention. Figure 1The processing device 100 shown includes a processor 110 and a machine-readable storage medium 120. Depending on the actual application, the machine-readable storage medium 120 may include a main memory or any type of "non-transitory readable medium". For example, in some embodiments, the non-transitory readable medium includes, for example, semiconductor memory, programmable logic circuits, and / or storage devices. The machine-readable storage medium 120 stores at least one thread group (or warp) 121, at least one instance buffer 122, at least one TLAS 123, at least one BLAS 124, at least one index buffer 125, and at least one vertex buffer 126. The processor 110 is coupled to the machine-readable storage medium 120 to read data content. For simplicity of the figure, Figure 1 the data transfer interface circuit between the machine-readable storage medium 120 and the data processing integrated circuit 120 is not shown because the data transfer interface circuit of the machine-readable storage medium 120 is a well-known circuit.

[0039] It should be understood that the number of threads included in a thread group (warp) is the size of the thread group, and the size of the thread group is usually less than or equal to 128. For example, the size of the thread group can be 4, 16, 32, 64, 128, etc. In addition, a thread group can include multiple thread groups (warps).

[0040] In some other examples, a warp can also be referred to as a thread bundle. Correspondingly, a thread group can be referred to as a thread group. A thread group can include multiple thread bundles (warps).

[0041] According to the actual design, the processor 110 may include any type of integrated circuit. For example, the processor 110 may include a Central Processing Unit (CPU), a Graphics Processing Unit (GPU), a controller, a microcontroller, a microprocessor, an Application-Specific Integrated Circuit (ASIC), a Digital Signal Processor (DSP), a Field Programmable Gate Array (FPGA), and / or other data processing circuits. The processor 110 may execute the thread group 121 (program) of the machine-readable storage medium 120.

[0042] Figure 2 is a schematic flowchart of a processing method for a ray tracing acceleration structure according to an embodiment of the present invention. Please refer to Figure 1 and Figure 2 . In step S210, the processor 110 may obtain the thread group 121 from the machine-readable storage medium 120 and execute the thread group 121. Among them, the thread group 121 includes at least one descriptor. The processor 110 may execute the descriptor to simulate the interaction between the ray and the scene. Among them, the descriptor includes a first pointer (pointer) for pointing to the TLAS 123 of the scene and a second pointer for pointing to the instance cache 122.

[0043] Figure 3 is a schematic diagram of a top-level data structure according to an embodiment of the present invention. According to the actual design, Figure 3 the illustrated instance cache 122, TLAS 123, BLAS 124, index cache 125, and vertex cache 126 may be used as Figure 1 one of many embodiments of the illustrated instance cache 122, TLAS 123, BLAS 124, index cache 125, and vertex cache 126. In Figure 3In the illustrated embodiment, the thread group 121 includes a descriptor 321. The descriptor 321 has a first pointer PTR1 for pointing to the TLAS 123 of the scene and a second pointer PTR2 for pointing to the instance cache 122. The present embodiment does not limit the implementation manners of the first pointer PTR1 and the second pointer PTR2. For example, the first pointer PTR1 may include the address of the TLAS 123 in the machine-readable storage medium 120 (or main memory), and the second pointer PTR2 may include the address of the instance cache 122 in the machine-readable storage medium 120 (or main memory).

[0044] Please refer to Figure 1 、 Figure 2 and Figure 3 . The processor 110 may execute the descriptor 321 in step S210. In step S220, the processor 110 may obtain the TLAS 123 from the machine-readable storage medium 120 by using the first pointer PTR1 carried by the descriptor 321. Each leaf node in the TLAS 123 (such as leaf nodes TL1, TL2, TL3, and TL4) includes an instance identifier for pointing to the corresponding instance. In step S230, the processor 110 may traverse the TLAS 123 based on the ray to find the leaf node in the TLAS 123 that intersects the ray (hereinafter referred to as the intersecting leaf node). The specific method for the processor 110 to traverse the TLAS 123 may be any ray-tracing intersection test, such as an existing intersection test or other intersection tests. The processor 110 may start the intersection test from the root node TR. When the ray "hits" the root node TR (bounding box), the processor 110 may perform an intersection test for each child node of the root node TR (such as branch nodes TB1 and TB2). Since the ray does not hit the branch node TB1 (bounding box), any child nodes of this branch node TB1 (such as leaf nodes TL1 and TL2) do not need to be subjected to an intersection test. Here, it is assumed that the ray hits the branch node TB2 (bounding box), so each child node of this branch node TB2 (such as leaf nodes TL3 and TL4) needs to be subjected to an intersection test. Here, it is assumed that the ray hits the leaf node TL3.

[0045] Each instance record in the instance cache 122 includes a third pointer for pointing to the corresponding BLAS in the scene, a fourth pointer for pointing to the corresponding index cache, and a fifth pointer for pointing to the corresponding vertex cache. In step S240, by using the instance identifier IID carried by the intersecting leaf node TL3, the processor 110 can obtain the instance record IR corresponding to the intersecting leaf node TL3 (hereinafter referred to as the intersecting instance record) from the instance cache 122 pointed to by the second pointer PTR2. The intersecting instance record IR includes a third pointer PTR3 for pointing to the BLAS 124, a fourth pointer PTR4 for pointing to the index cache 125, and a fifth pointer PTR5 for pointing to the vertex cache 126. The embodiments of the present invention do not limit the implementation manners of the third pointer PTR3, the fourth pointer PTR4, and the fifth pointer PTR5. For example, the third pointer PTR3 may include the address of the BLAS 124 in the machine-readable storage medium 120 (or the main memory), the fourth pointer PTR4 may include the address of the index cache 125 in the machine-readable storage medium 120 (or the main memory), and the fifth pointer PTR5 may include the address of the vertex cache 126 in the machine-readable storage medium 120 (or the main memory).

[0046] In step S250, the processor 110 can obtain the BLAS 124 from the machine-readable storage medium 120 by using the third pointer PTR3 carried by the intersecting instance record IR. In step S260, the processor 110 can traverse the BLAS 124 based on the ray to find the leaf nodes in the BLAS 124 that intersect with the ray (hereinafter referred to as the intersecting primitive nodes). The specific method for the processor 110 to traverse the BLAS 124 can be any ray-tracing intersection test, such as existing intersection tests or other intersection tests. The intersecting primitive nodes include primitive identifiers. The processor 110 can obtain the primitive index number corresponding to the intersecting primitive node (hereinafter referred to as the intersecting primitive index number) from the index cache 125 pointed to by the fourth pointer PTR4 by using the primitive identifier of the intersecting primitive node. The processor 110 can obtain the primitive vertex coordinates corresponding to the intersecting primitive node from the vertex cache 126 pointed to by the fifth pointer PTR5 by using the intersecting primitive index number.

[0047] In summary, the descriptor 321 can point to the TLAS 123 and the instance cache 122 through the first pointer PTR1 and the second pointer PTR2. Therefore, the processor 110 can efficiently obtain the content of the TLAS 123 and the instance cache corresponding to the descriptor 321. After traversing the TLAS 123, when the ray intersects with an instance of the TLAS 123 (such as the intersection leaf node TL3), the processor can find this intersection leaf node TL3 in the TLAS 123, and then obtain the instance identifier IID carried by this intersection leaf node TL3 for pointing to the corresponding instance. By using the second pointer PTR2 and the instance identifier IID of the intersection leaf node TL3, the processor 110 can efficiently obtain the intersection instance record IR corresponding to the intersection leaf node TL3 from the instance cache 122, and then obtain the third pointer PTR3 carried by this intersection instance record IR for pointing to the corresponding BLAS 124, the fourth pointer PTR4 for pointing to the corresponding index cache 125, and the fifth pointer PTR5 for pointing to the corresponding vertex cache 126. By using the third pointer PTR3, the processor 110 can efficiently obtain the corresponding BLAS 124. After traversing the corresponding BLAS 124, when the ray intersects with a leaf node (primitive node) of the BLAS 124, the processor 110 can find the intersection primitive node in the BLAS 124. Therefore, the processor 110 can efficiently process the ray tracing acceleration structure.

[0048] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and are not intended to limit them. Although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements on some or all of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the scope of the technical solutions of the embodiments of the present invention.

Claims

1. A processing device for a ray tracing acceleration structure, comprising: A machine-readable storage medium including at least one thread group, at least one instance cache, at least one top-level acceleration structure, and at least one bottom-level acceleration structure; And A processor coupled to the machine-readable storage medium for obtaining the thread group from the machine-readable storage medium and executing the thread group, wherein the thread group includes at least one descriptor, the processor executes the descriptor to simulate the interaction between a ray and a scene, the descriptor includes a first pointer for pointing to the top-level acceleration structure of the scene and a second pointer for pointing to the instance cache, the processor obtains the top-level acceleration structure from the machine-readable storage medium by using the first pointer, the processor traverses the top-level acceleration structure based on the ray to find an intersection leaf node in the top-level acceleration structure that intersects the ray, the intersection leaf node includes an instance identifier for pointing to a corresponding instance, the processor obtains the intersection instance record corresponding to the intersection leaf node from the instance cache pointed to by the second pointer by using the instance identifier, the intersection instance record includes a third pointer for pointing to the bottom-level acceleration structure of the scene, the processor obtains the bottom-level acceleration structure from the machine-readable storage medium by using the third pointer, and the processor traverses the bottom-level acceleration structure based on the ray to find an intersection primitive node in the bottom-level acceleration structure that intersects the ray.

2. The processing device according to claim 1, characterized in that, The machine-readable storage medium further includes at least one index cache and at least one vertex cache, the intersection instance record further includes a fourth pointer for pointing to the index cache and a fifth pointer for pointing to the vertex cache, the intersection primitive node includes a primitive identifier, the processor obtains the intersection primitive index number corresponding to the intersection primitive node from the index cache pointed to by the fourth pointer by using the primitive identifier, and the processor obtains the primitive vertex coordinates corresponding to the intersection primitive node from the vertex cache pointed to by the fifth pointer by using the intersection primitive index number.

3. A processing method for a ray tracing acceleration structure, characterized in that, The processing method includes: Executing a descriptor to simulate the interaction between a ray and a scene, wherein the descriptor includes a first pointer for pointing to the top-level acceleration structure of the scene and a second pointer for pointing to an instance cache; Obtaining the top-level acceleration structure by using the first pointer; Traversing the top-level acceleration structure based on the ray to find an intersection leaf node in the top-level acceleration structure that intersects the ray, wherein the intersection leaf node includes an instance identifier for pointing to a corresponding instance; Obtaining the intersection instance record corresponding to the intersection leaf node from the instance cache pointed to by the second pointer by using the instance identifier, wherein the intersection instance record includes a third pointer for pointing to the bottom-level acceleration structure of the scene; Obtaining the bottom-level acceleration structure by using the third pointer; and Traverse the underlying acceleration structure based on the ray to find the intersection primitive nodes in the underlying acceleration structure that intersect with the ray.

4. The processing method according to claim 3, characterized in that, The intersection instance record further includes a fourth pointer for pointing to an index buffer and a fifth pointer for pointing to a vertex buffer. The intersection primitive node includes a primitive identifier, and the processing method further includes: Obtaining the intersection primitive index number corresponding to the intersection primitive node from the index buffer pointed to by the fourth pointer by using the primitive identifier; and Obtaining the primitive vertex coordinates corresponding to the intersection primitive node from the vertex buffer pointed to by the fifth pointer by using the intersection primitive index number.

5. A machine-readable storage medium for storing at least one thread group, at least one instance buffer, at least one top-level acceleration structure, and at least one underlying acceleration structure, wherein the thread group includes at least one descriptor, and when the descriptor is executed by a processor, it can implement the processing method of the ray tracing acceleration structure according to any one of claims 3-4.

Citation Information

Patent Citations

  • Enhanced techniques for traversing ray tracing acceleration structures

    CN113808245A

  • Enhanced Techniques for Traversing Ray Tracing Acceleration Structures

    US20210390758A1