CUDA-Based Parallel BVH Minimum Distance Query Method
By implementing the parallelized hierarchical body enclosing box minimum distance query method on the CUDA platform, using the two-level task queue and binary tree search strategy, the efficiency problem of the three-dimensional model close-distance query algorithm in the existing technology under high real-time requirements is solved, and efficient and real-time three-dimensional model distance query is achieved.
Patent Information
- Application Number
- CN202510402459.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-01
- Publication Date
- 2025-06-24
- Estimated Expiration
- 2045-04-01
AI Technical Summary
The existing three-dimensional model close-distance query algorithm cannot meet the high real-time requirements of industrial automation scenarios, especially when processing three-dimensional models with tens of millions of primitives, the calculation efficiency is inefficient and lack of real-time.
The parallelized hierarchical body bounding box (BVH) minimum distance query method is adopted based on CUDA, and a highly parallelized query process is realized by constructing a binary tree search strategy that matches the two-level task queue and design.
It significantly improves the minimum distance query efficiency between high-complexity and high-precision three-dimensional models, solves the problems of low computing efficiency and insufficient real-time performance, improves query speed, and improves the execution efficiency of related algorithms.
Smart Images

Figure CN119917683B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of computer graphics, and more specifically, it is a method for parallel hierarchical bounding volume minimum distance query based on CUDA. Background Art
[0002] In computer graphics, for the problem of querying the closest distance between two three-dimensional objects, currently, the hierarchical bounding volume (BVH) algorithm is mainly used for acceleration. As computer graphics is gradually applied to industrial automation scenarios such as assembly management, digital twin, and robot simulation, the requirements for the accuracy and real-time performance of related algorithms are also getting higher and higher. In some high-precision calculation scenarios, it may be necessary to perform real-time distance queries on three-dimensional models containing tens of millions or even hundreds of millions of primitives at a frequency of approximately 30 - 50 Hz. This includes the following problems:
[0003] Huge data volume; a model with tens of millions of primitives may require constructing a BVH tree with approximately 40 levels, so the number of leaf nodes contained in the hierarchical bounding volume test tree (BVTT) exceeds 10 24 ones. Strict alignment in three-dimensional space; due to the coarseness of the bounding box, it is usually impossible to exclude it by pruning the BVTT. Severe branch divergence; these vertices with close distances may ultimately be distributed in different subtrees of the BVTT, which significantly increases the workload of the query.
[0004] Currently, the existing three-dimensional model closest distance query algorithms cannot meet the high real-time requirements of industrial automation scenarios. In the prior art, the following several methods are usually adopted to partially solve the above problems:
[0005] Surface Area Heuristic (SAH) algorithm. SAH is a commonly used and effective method for constructing high-quality BVHs. However, for geometric bodies with relatively regular shapes in industrial scenarios, the segmentation mode of applying the SAH algorithm has little difference from the traditional equidistant segmentation mode, and the performance improvement is relatively limited.
[0006] Selecting denser bounding boxes. However, these denser bounding boxes usually have much more complex geometric distance calculation methods than traditional rectangular bounding boxes, resulting in a significant increase in the computational cost of each search, which is not worth the effort.
[0007] Using GPU parallel search. However, the existing parallel recursive methods are only limited to collision detection. Although the functions are similar, the parallelization difficulty of collision detection is much less than that of distance query, and due to the significant difference in algorithm logic, the existing algorithms cannot be directly transplanted. Summary of the Invention
[0008] In view of the characteristics of NVIDIA GPUs and the high-precision and high-real-time requirements of industrial application scenarios, the present invention provides a parallel BVH minimum distance query method to overcome the defects of low computational efficiency and insufficient real-time performance of existing distance query algorithms when dealing with tens of millions of 3D models.
[0009] The technical solution adopted by the present invention to achieve the above object is: a parallel BVH minimum distance query method based on CUDA, which constructs two-level task queues, uses CUDA thread blocks as batch processing units, and uses a hierarchical bounding volume test tree to implement parallel search to query the closest distance between two 3D objects, including the following steps:
[0010] Construct two-level task queues, including a global queue and a local queue;
[0011] Before starting the search, initialize all threads of the GPU in units of thread blocks; the 0th thread of each thread block first takes a node from the global queue and places it at the top of the local queue, and initializes it as the root node of the current thread block;
[0012] During the search, all threads in the same thread block batch-read child nodes from the local queue for concurrent search; after the search is completed, batch-add the child nodes to be searched next time to the local queue;
[0013] After the thread block completes the search, obtain the minimum distance of the hierarchical bounding volume.
[0014] The construction of the two-level task queues is specifically as follows:
[0015] The first-level queue is the global queue, which is visible to all threads participating in the search, and this queue is a first-in-first-out queue;
[0016] The second-level queue is the local queue, that is, the exclusive queue of the thread block. Each thread block instantiates a queue, which is visible to all threads in the current thread block; this queue adopts a double-ended design, where the front end is a buffer with a fixed length, and the length is the same as the number of threads in the thread block; the back end is a last-in-first-out queue.
[0017] All threads in the same thread block batch-read child nodes from the local queue for concurrent search, including the following steps:
[0018] In each batch search, for child nodes containing two hierarchical bounding volumes, calculate the distances of the bounding volumes in the left child node and the right child node respectively;
[0019] When the distance calculated for a certain child node is greater than or equal to the currently remembered minimum distance, directly skip this child node; otherwise, store the child node with the closer distance in the front end of the local queue, and store the child node with the farther distance in the back end of the local queue;
[0020] In addition, if the child node is a leaf node, directly calculate the distance between the primitives contained within the child node; if the calculated distance is less than the currently memorized minimum distance, replace the currently memorized minimum distance with the calculated distance to update the memory.
[0021] When storing the child node, use the parallel prefix compression algorithm to format and store the data without locks, and use the thread block shared memory.
[0022] When retrieving the child node, the thread block first batch retrieves all nodes from the front end of the local queue; when the number of child nodes in the front end of the local queue is less than the total number of threads in the thread block, the threads exceeding the length part retrieve data from the back end.
[0023] Filter the child nodes with priority for parallelization according to the determination conditions and add them to the global queue; the determination conditions are any one of the following cases:
[0024] a. The layer where the node is located is less than 1 / M of the total number of layers of BVTT, 2 ≤ M ≤ 10;
[0025] b. The distance difference between the closer child node and the farther child node is less than (1 / N) of the total volume of the bounding box -1 / 3 , 5 ≤ N ≤ 20;
[0026] c. The number of child nodes in the current global queue is not greater than the total number of thread blocks.
[0027] After the thread block completes the search, obtaining the minimum distance of the hierarchical bounding volume includes the following steps:
[0028] All thread blocks first search the nodes in their own local queues; when a thread block completes the search of all child nodes in its own local queue and no longer adds new child nodes to the local queue, this thread block retrieves a new child node from the global queue as the root node for the next round of search, and repeats the step of all threads in the same thread block batch reading child nodes from the local queue for concurrent search;
[0029] When all thread blocks have completed the search and there are no nodes in the global queue either, the algorithm exits, and the currently memorized minimum distance is the global minimum distance, which is used as the minimum distance of the hierarchical bounding volume.
[0030] The CUDA-based parallel BVH minimum distance query system includes:
[0031] A task queue construction module for constructing a two-level task queue, including a global queue and a local queue;
[0032] The initialization thread module is used to initialize all threads of the GPU in units of thread blocks before starting the search; the 0th thread of each thread block first takes a node from the global queue and places it at the top of the local queue, initializing it as the root node of the current thread block;
[0033] The batch search module is used to perform concurrent searches by all threads in the same thread block reading child nodes in batches from the local queue during the search; after the search is completed, the child nodes to be searched next are added to the local queue in batches;
[0034] The BVH minimum distance acquisition module is used to obtain the minimum distance of the hierarchical bounding volume after the thread block completes the search.
[0035] The CUDA-based parallel BVH minimum distance query device includes a memory and a processor; the memory is used to store computer programs; the processor is used to implement the CUDA-based parallel BVH minimum distance query method when executing the computer programs.
[0036] A computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the CUDA-based parallel BVH minimum distance query method is implemented.
[0037] The present invention has the following beneficial effects and advantages:
[0038] 1. The present invention designs an innovative two-layer task queue structure for the parallel distance query task of 3D models and formulates matching enqueue and dequeue rules. This design not only allows the algorithm to preferentially process BVTT nodes with closer distances like traditional single-threaded recursion in a highly parallel situation, thus effectively performing pruning operations, but also takes into account the hardware characteristics of NVIDIA GPUs, avoiding the common thread divergence phenomenon in CUDA programming;
[0039] 2. For the proposed two-layer queue structure, the present invention designs a matching binary tree search strategy to meet the requirements of processing highly complex 3D models in industrial production scenarios. This search strategy realizes a highly parallel query process, effectively avoids the problems discussed in the previous sections, and significantly improves the minimum distance query efficiency between high-complexity and high-precision 3D models. The design of this strategy not only optimizes the allocation of computing resources but also ensures high-efficiency and high-precision data processing. BRIEF DESCRIPTION OF THE DRAWINGS
[0040] Figure 1 The overall architecture diagram of the algorithm of the present invention.
[0041] Figure 2 The algorithm thread block search flow chart of the present invention.
[0042] Figure 3 Schematic diagram of the local queue design of the present invention.
[0043] Figure 4 Schematic diagram of the method for storing / retrieving local queue data of the present invention.
[0044] Figure 5 Schematic diagram of the global queue design of the present invention. Detailed implementation manners
[0045] The present invention will be further described in detail below with reference to the accompanying drawings and embodiments.
[0046] The present invention uses the hierarchical bounding volume (BVH) algorithm to accelerate the query process. This algorithm recursively groups one or a set of primitive elements into bounding volumes, and each bounding volume can be enclosed in a larger box, forming a binary tree or multi - ary tree structure that is convenient for indexing. In the collision / distance query task of two geometric bodies, by aggregating two BVH binary trees into a hierarchical volume test tree (BVTT), the query task can be converted into a classic binary tree search problem, simplifying the algorithm logic and facilitating the implementation of the program.
[0047] For example, in the industrial production and manufacturing scenario, in order to monitor the status of the production line in real time, a machine vision system uses sensors (such as depth cameras or lidar) to obtain the three - dimensional point cloud in the environment or the model of the workpiece, thereby constructing a digital twin to simulate the current real operation situation. The data volume of these point clouds or three - dimensional models can reach tens of millions or even billions. To improve the computing efficiency, the object models are stored in the computer in the form of BVH trees (such as AABB trees or OBB trees), representing the hierarchical bounding volumes of each object. When the robot performs grasping, obstacle avoidance, or detection tasks, it is necessary to quickly calculate the minimum distance between the robot and the target object. At this time, it is necessary to aggregate the BVHs of the two query objects to form a BVTT, and traverse this binary tree to achieve the minimum distance query.
[0048] The classic distance query algorithm uses a depth - first method to recursively traverse and search the BVTT. Before recursively entering a subtree, the algorithm first calculates the distance between the bounding volumes of the two child nodes, so as to preferentially access the subtree with a closer distance. In this way, the algorithm can quickly locate the pair of vertices with a closer distance. In the later stage of the recursion, according to the currently found closest distance, the algorithm can quickly exclude those child nodes with a farther distance, perform pruning on the BVTT, thereby greatly accelerating the search process. This method significantly reduces unnecessary calculations, improves the search efficiency, and optimizes the performance of the entire distance query.
[0049] Such as Figure 1As shown in the figure, in view of the hardware characteristics of NVIDIA GPUs and the characteristics of distance queries in industrial production scenarios, the present invention designs a parallel recursive search algorithm with two-level task queues based on the classic recursive algorithm according to the classification of the CUDA framework itself. Among them, the first-level queue is the global queue, which is visible to all threads participating in the search; the second-level queue is the thread block exclusive queue, and each thread block instantiates a queue, which is visible to all threads in the thread block.
[0050] As Figure 2 shown in the figure, when performing distance queries, the CUDA thread block is used as the batch processing unit, and the operation speed is improved through the batch search optimization process. Before starting the search, the 0th thread of each thread block first takes a node from the global queue and puts it into the head of the front end of the local queue, and initializes it as the root node of the thread block. Among them, the front end is a fixed-length buffer of the local queue, and the head is the front part in the front end. During the search, all threads in the same thread block batch-read nodes from the local queue and perform concurrent searches. After the search is completed, according to the recursive termination condition, the nodes that need to be searched next are batch-added to the local queue. Specifically, the main process includes the following steps:
[0051] 1. Global queue management
[0052] 1.1. Start:
[0053] Start the minimum distance query process. Initialize the root node of the BVTT to be processed in the global queue.
[0054] 1.2. Take out a node from the global queue:
[0055] The 0th thread of each thread block independently takes out a node from the queue, indicating that the thread block needs to calculate the minimum distance of the primitives contained in the subtree expanded by this node.
[0056] 1.3. Initialize the local queue:
[0057] Generate a local queue for the current thread block to manage the batch search of its child nodes. Put the current node into the head of the front end of the local queue as the initial node.
[0058] 1.4. Thread block batch search subtree:
[0059] Call the sub-process (local search process) to search the subtree expanded by the current node.
[0060] 1.5. Judge that the global queue is empty and all thread blocks are processed
[0061] Check whether the global queue is empty and whether all subtrees in the local queue have completed the search. If so, end the process and return the final minimum distance; if not, return to step 1.2 to continue.
[0062] 2. Local Search Process
[0063] 2.1. Remove Nodes from the Local Queue:
[0064] Remove nodes from the local queue in batches and assign them to each thread within the thread block for search.
[0065] 2.2. Determine Whether it is a Leaf Node
[0066] Determine whether the current node is a leaf node (i.e., the bottommost node of BVTT). If so, directly calculate the minimum distance between the primitives inside the node, update the optimal solution, and then jump to step 2.9.
[0067] 2.3. Obtain the Left and Right Child Nodes:
[0068] If the node is not a leaf node, obtain its left and right child nodes as the targets for further processing.
[0069] 2.4. Calculate the Distances between the Left and Right Child Nodes and the Target Object:
[0070] Calculate the distances between the bounding boxes inside the left and right child nodes respectively (geometric distance or other measurement methods).
[0071] 2.5. Sort in Ascending Order of Distance:
[0072] Sort the left and right child nodes from near to far according to the calculated distances to give priority to processing the nearer child nodes.
[0073] 2.6. Recursion Termination Condition Judgment: Whether the nearer child node is greater than the current global minimum value
[0074] If it is greater, put the child node at the front of the local queue, indicating that further search is needed. If it is less than or equal to, directly skip the node because there cannot be a closer distance inside the node.
[0075] After the judgment is completed, use parallel prefix scan to compress the front part of the local queue to exclude the empty spaces generated by skipping nodes.
[0076] 2.7. Recursion Termination Condition Judgment: Whether the farther child node is greater than the current global minimum value
[0077] If it is greater, enter the judgment described in step 2.8; if it is less than or equal to, directly skip the node because there cannot be a closer distance inside the node.
[0078] 2.8. Whether to add child nodes to the global queue
[0079] If the current node meets the conditions for priority search, add it to the global queue for batch processing by other thread blocks; otherwise, add the child node to the back end of the local queue for the current thread block to search later. The judgment conditions are as follows:
[0080] The layer where the node is located is less than 1 / 4 of the total number of layers of BVTT (the threshold is adjustable). For nodes at too low levels, the corresponding subtrees are not very large in scale, and the time difference between splitting for other thread blocks to search and searching within the current thread block itself is not significant, but it increases the memory read and write overhead instead.
[0081] The distance difference between the closer child nodes and the farther child nodes is less than (1 / 10) of the total volume of the bounding box -1 / 3 (the threshold is adjustable). For bounding boxes with very similar or even the same distances, similar distances are more likely to appear in the expanded subtrees, so it is worth parallelizing the processing.
[0082] The amount of data in the current global queue is not greater than the total number of thread blocks. If the number of nodes to be processed in the global queue exceeds the total number of thread blocks supported by the GPU, it may cause some nodes that are expected to be searched with priority parallelization to not be searched for a long time, but instead take longer.
[0083] 2.9. Whether the local queue is empty
[0084] If there are still unprocessed nodes in the local queue, return to step 2.1 to continue; if the local queue is empty, the local processing process ends.
[0085] Through the above process, the algorithm can achieve step-by-step processing of the nodes in the global queue, complete the block processing of nodes with complex structures based on the optimized search of the local queue, and control the search path according to the node distance and priority.
[0086] As Figure 3 、 Figure 4 shown, the local queue is a composite queue, divided into two parts. The front end is a buffer with a fixed length (the size is equal to the number of threads in the thread block), and the back end is a last-in-first-out queue. Such a design is to match the priority search mode of the binary tree. When storing data, the closer child nodes in each batch search are stored in the front end, and the farther child nodes are stored in the back end. Since the recursive termination conditions (i.e., the above steps 2.6 and 2.7) will exclude some bounding boxes with farther distances, the length relationship of the data stored in each batch must satisfy: the length of the data stored in the back end Len _ back <= the length of the data stored in the front end Len _The front <= the thread block size block_size. Therefore, such a design will not cause queue overflow. When retrieving data, the thread block preferentially retrieves all data in batches from the front end. When the amount of data in the front end is less than the total number of thread blocks, the remaining threads retrieve data from the back end. This queue also adopts a lock-free design, uses a parallel prefix compression algorithm to format the stored data, and uses thread block shared memory to accelerate the memory read and write speed.
[0087] As Figure 5 shown, the design of the global queue is relatively simple. It is a first-in-first-out queue used to store the nodes screened out in step 2.8 above. The buffer of this queue is implemented in the global video memory and is visible to all thread blocks. Since the number of times the thread block accesses this queue in one search is not large, the access and storage of data are both implemented using atomic memory operations, which is more easily implemented in programming.
[0088] By adopting the algorithm of the present invention, the speed of performing the three-dimensional model nearest distance query task in the graphics algorithm is increased by 5 to 30 times, greatly improving the execution efficiency of related algorithms. In the digital twin system of industrial production, if related algorithms are applied to perform graphics operations, the production efficiency of high-precision scenarios can be greatly improved. In the robot simulation environment (such as reinforcement learning training, etc.), if related algorithms are applied to perform physical simulations, the convergence speed of the model can be greatly increased, reducing the training time required. Or, without changing the training time, the accuracy of the simulation environment can be greatly improved, making the model obtained by training more accurate and reliable.
Claims
1. A parallel BVH minimum distance query method based on CUDA, characterized in that: A two-level task queue is constructed, CUDA thread blocks are used as batch processing units, and a hierarchical bounding box test tree is used to implement parallel search to query the shortest distance between two three-dimensional objects, including the following steps: Build a two-level task queue, including a global queue and a local queue; Before starting the search, all threads of the GPU are initialized in thread blocks. Thread 0 of each thread block first takes a node from the global queue and puts it at the top of the local queue, which is initialized as the root node of the current thread block. When searching, all threads in the same thread block read child nodes in batches from the local queue and perform concurrent searches. After the search is completed, the child nodes that need to be searched next time are added to the local queue in batches. After the thread block completes the search, the minimum distance of the bounding box of the hierarchy is obtained; All threads in the same thread block read child nodes in batches from the local queue and perform concurrent search, including the following steps: In each batch search, for the child nodes containing two hierarchical bounding boxes, the distance of the bounding box in the left child node and the distance of the bounding box in the right child node are calculated respectively; When the calculated distance of a child node is greater than or equal to the currently memorized minimum distance, the child node is directly skipped; otherwise, the child node with a closer distance is stored in the front of the local queue, and the child node with a farther distance is stored in the back of the local queue; In addition, if the child node is a leaf node, the distance between the graphics elements contained in the child node is directly calculated; if the calculated distance is less than the currently remembered minimum distance, the calculated distance replaces the currently remembered minimum distance to update the memory.
2. The CUDA-based parallel BVH minimum distance query method according to claim 1, characterized in that: The construction of the two-level task queue is as follows: The first-level queue is a global queue, which is visible to all threads participating in the search. This queue is a first-in-first-out queue; The second-level queue is a local queue, that is, a thread block-specific queue. Each thread block instantiates a queue, which is visible to all threads in the current thread block. The queue adopts a double-ended design, in which the front end is a fixed-length buffer with the same length as the number of threads in the thread block; the back end is a first-in, last-out queue.
3. The CUDA-based parallel BVH minimum distance query method according to claim 1, characterized in that: When storing in child nodes, a parallel prefix compression algorithm is used to format and store data without locks, and thread blocks are used to share memory.
4. The CUDA-based parallel BVH minimum distance query method according to claim 1, characterized in that: When taking out child nodes, the thread block takes out all nodes in batches from the front end of the local queue first; when the number of child nodes in the front end of the local queue is less than the total number of threads in the thread block, the threads exceeding the length take out data from the back end.
5. The CUDA-based parallel BVH minimum distance query method according to claim 1, characterized in that: The child nodes to be prioritized for parallelization are selected according to the judgment condition and added to the global queue; the judgment condition is any one of the following conditions: a. The number of layers the node is located in is less than 1 / M of the total number of layers of BVTT, 2≦M≦10; b. The distance between the closer child node and the farther child node is less than (1 / N) of the total volume of the bounding box. -1 / 3 , 5≦N≦20; c. The number of child nodes in the current global queue is not greater than the total number of thread blocks.
6. The CUDA-based parallel BVH minimum distance query method according to claim 1, characterized in that: After the thread block completes the search, the minimum distance of the bounding box of the hierarchical volume is obtained, which includes the following steps: All thread blocks preferentially search the nodes in their own local queues; when a thread block completes the search of all child nodes in its own local queue and no longer adds new child nodes to the local queue, the thread block takes out a new child node from the global queue as the root node for the next round of search, and repeats the steps of batch reading child nodes from the local queue by all threads in the same thread block to perform concurrent search; When all thread blocks have completed the search and there are no nodes in the global queue, the algorithm exits and the currently memorized minimum distance is the global minimum distance, which is used as the minimum distance of the bounding box of the hierarchy.
7. A CUDA-based parallel BVH minimum distance query system, applied to the CUDA-based parallel BVH minimum distance query method according to any one of claims 1 to 6, characterized in that: include: Task queue construction module, used to build two-level task queues, including global queues and local queues; The thread initialization module is used to initialize all threads of the GPU in thread blocks before starting the search. Thread No. 0 of each thread block first takes a node from the global queue and puts it at the top of the local queue, and initializes it as the root node of the current thread block. Batch search module, when used for searching, all threads in the same thread block read sub-nodes in batches from the local queue for concurrent search; after the search is completed, the sub-nodes to be searched next time are added to the local queue in batches; The BVH minimum distance acquisition module is used to obtain the minimum distance of the bounding box of the hierarchy after the thread block completes the search.
8. A parallel BVH minimum distance query device based on CUDA, characterized in that: It comprises a memory and a processor; the memory is used to store a computer program; the processor is used to implement the CUDA-based parallelized BVH minimum distance query method as described in any one of claims 1 to 6 when executing the computer program.
9. A computer-readable storage medium, characterized in that: The storage medium stores a computer program, and when the computer program is executed by the processor, the CUDA-based parallelized BVH minimum distance query method according to any one of claims 1 to 6 is implemented.
Citation Information
Patent Citations
Method for quickly constructing bounding volume hierarchy (BVH) based on GPU
CN101819675A
Rigid body obstacle avoidance method based on three-dimensional discrete model distance calculation
CN119129376A