Fast and energy-efficient k-nearest neighbor search accelerator for large-scale point cloud

By building an NSVS framework to realize k-nearest neighbor search of large-scale point cloud maps on FPGAs, the problems of redundancy in large-scale point cloud map search area and slow data transmission are solved, and fast and efficient search is achieved to meet the real-time requirements of unmanned vehicles.

WO2025148389A1PCT designated stage expired Publication Date: 2025-07-17SHANGHAI TECH UNIV
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
PCT/CN2024/118960
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-01-11
Filing Date
2024-09-14
Publication Date
2025-07-17

AI Technical Summary

Technical Problem

Large-scale point cloud map search area is redundant and data transmission is slow. The existing technology searches in large-scale point cloud maps are inefficient and cannot meet the real-time requirements of unmanned vehicles.

Method used

NSVS framework based on DSVS search structure is built, and a large-scale point cloud map k-nearest neighbor search algorithm is implemented on FPGA. By reducing the search area and adaptive data transmission technology, it includes dividing the point cloud space into voxels and subvoxels, optimizing data access using data reuse buffers, and reducing redundant search areas, and processing candidate neighboring points selection in parallel.

Benefits of technology

It realizes fast and efficient k-nearest neighbor search for large-scale point cloud maps, with a search speed of 9.1 times faster than the most advanced FPGA and 11.5 times higher than the most advanced FPGA and GPU, meeting the real-time needs of unmanned vehicles.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2024118960_17072025_PF_FP_ABST
    Figure CN2024118960_17072025_PF_FP_ABST
Patent Text Reader

Abstract

The technical solution of the present invention provides a fast and energy-efficient k-nearest neighbor search accelerator for a large-scale point cloud, and is characterized in that an NSVS framework that performs a search on the basis of a DSVS search structure is constructed, and a k-nearest neighbor search algorithm for a large-scale point cloud map is implemented on an FPGA. The technical solution comprises: constructing a DSVS search structure; and searching for k-nearest neighbors on the basis of a DSVS. An experimental result on a KITTI data set shows that the k-nearest neighbor search accelerator provided in the present invention has a search speed 9.1 times faster than an existing state-of-the-art FPGA implementation. In addition, the solution of the present invention also achieves optimal energy efficiency, and the energy efficiency of the accelerator provided in the present invention is 11.5 times and 13.5 times higher than state-of-the-art FPGA and GPU implementations, respectively.
Need to check novelty before this filing date? Find Prior Art

Description

A fast and energy-efficient large-scale point cloud K-nearest neighbor search accelerator Technical Field

[0001] The present invention relates to a large-scale point cloud K-nearest neighbor search accelerator. Background Art

[0002] k-nearest neighbor search is an important step in many lidar algorithms. It is widely used in simultaneous localization and mapping algorithms to correct positioning drift errors, and in related algorithms such as re-localizing the user terminal by matching a frame of lidar point cloud with the entire point cloud map. Although the k-nearest neighbor search algorithm has a relatively simple structure, it takes up about 80% of the matching time due to the large number of query operations in large-scale point clouds [1].

[0003] Considering the real-time requirements of simultaneous localization and mapping algorithms in complex outdoor scenes and the strict battery constraints of unmanned vehicles, how to develop an energy-efficient algorithm for fast k-nearest neighbor search is a huge challenge.

[0004] In order to improve the performance of k-nearest neighbor and reduce power consumption, relevant experts have explored different aspects. A series of works such as [2] and [3] proposed a parallel pipeline k-nearest neighbor search algorithm based on tree data structure. However, it is relatively fast only on small-scale point cloud maps. The KD-tree construction time of large-scale point cloud maps is too long, which is unacceptable for unmanned vehicles. [4] proposed a k-nearest neighbor search hardware accelerator based on DSVS (double-segmentation-voxel-structure) that can quickly build large-scale point cloud maps. It divides the space where the point cloud is located into voxels by adaptively setting the edge length, and further divides the dense voxels into sub-voxels. When processing large and uneven point clouds, DSVS can narrow the search area to nearby (sub-) voxels. However, due to slow data transmission and search area redundancy, DSVS performs slowly and inefficiently in large-scale point cloud maps. In addition, most k-nearest neighbor implementations such as [2][4] can only process medium-sized point clouds, with a maximum of about 400,000 points, and are therefore only suitable for inter-frame matching or local map matching. Searching on large-scale point clouds will lead to problems such as long search algorithm construction time, long point cloud transmission time, complex search operations, and large search areas.

[0005] References

[0006] [1] J.Zhang and S.Singh, "Low-drift and real-time lidar odometry and mapping," Autonomous Robots, vol.41, pp.401–416,02 2017.

[0007] [2] Y.Li, K.Zheng, and H.Xiao, “A knn accelerator based on approximate kd tree for icp,” in 2022 International Conference on Image Processing and Media Computing (ICIPMC), 2022, pp.124–128.

[0008] [3] F. Chen, R. Ying, J.

[0009] [4]H.Sun,

[0010] Summary of the Invention

[0011] The technical problems to be solved by the present invention are: regional redundancy and slow data transmission in large-scale point cloud search.

[0012] To solve the above technical problems, the technical solution of the present invention is to provide a fast and energy-efficient large-scale point cloud K-nearest neighbor search accelerator. The NSVS framework based on the DSVS search structure is constructed to implement the large-scale point cloud K-nearest neighbor search algorithm on the FPGA, including:

[0013] Constructing the DSVS search structure:

[0014] Divide the reference set space into different voxels, and the side length of each voxel is equal to the desired search range R in If the k-nearest neighbor result exceeds the expected search range Rin, it is considered an outlier; based on the number of points in each voxel, the voxel containing more points is further divided into sub-voxels; the reference set is sorted to ensure that the points in adjacent voxels or adjacent sub-voxels are stored in continuous memory;

[0015] Searching for k-nearest neighbors based on DSVS further includes:

[0016] Function module 1: locate the query point and use different strategies to narrow the search range according to the location of the query point:

[0017] If the neighboring voxel of the voxel containing the query point is not further divided into sub-voxels, then the neighboring voxel is used as the search voxel;

[0018] If the neighboring voxels of the voxel containing the query point are further divided into sub-voxels, then only the sub-voxel closest to the voxel containing the query point is selected as the search sub-voxel among the neighboring voxels;

[0019] The search voxel and the search subvoxels constitute a reduced search area;

[0020] Function module 2: Extract candidate neighboring points in the reduced search area, where:

[0021] Define the data reuse ratio, which is the ratio of the number of query points to the number of reference points in the search area:

[0022] If the data reuse rate exceeds the threshold, the reference set is accessed sequentially using the data reuse buffer, which is continuously updated according to the changes of the query point, thereby ensuring that the points around the current query point are always in the data reuse buffer;

[0023] If the data reuse rate is lower than the threshold, the candidate neighboring points are obtained from the external memory in a random access mode;

[0024] Function Module 3: Select k-nearest neighbors from the candidate neighboring points obtained in Function Module 2 through a highly parallel k-nearest neighbor selection accelerator;

[0025] The task of building the DSVS search structure is performed on the PS side, and the task of searching for k-nearest neighbors based on DSVS is accelerated on the PL side.

[0026] Preferably, the voxel and the sub-voxel are indexed by a hash value and a sub_hash value respectively.

[0027] Preferably, the number of the search voxels and the search sub-voxels is at most 17.

[0028] Preferably, in the second function module, if the candidate neighboring points of the query point have been stored in a data reuse buffer in a previous query point search process, the candidate neighboring points are directly obtained from the data reuse buffer.

[0029] Preferably, the throughput of the function module 1, the function module 2 and the function module 3 is optimized to 1, and the data channels between the function modules are implemented as stream types to enable task-level pipelines.

[0030] This paper proposes an FPGA implementation of a high-efficiency and fast large-scale point cloud k-nearest neighbor search algorithm based on DSVS. Compared with existing technical solutions, the innovations are:

[0031] 1) A new search technique (NSVS, nearest-sub-voxel-selection) that reduces redundant search areas based on the proximity and density of query points in the search structure. When a query point is located in a dense area with many points, the present invention significantly reduces the search area.

[0032] 2) Adaptive data transfer technology is used to efficiently transfer point clouds with different data reuse rates from external memory to the accelerator. Point clouds with low data reuse rates are transferred in a random access manner through multiple large-bitwidth ports, while point clouds with high data reuse rates are transferred in a sequential access manner through on-chip cache and FIFO.

[0033] Experimental results on the KITTI dataset show that our proposed k-nearest neighbor search accelerator is 9.1 times faster than the state-of-the-art FPGA implementation. Furthermore, our solution achieves optimal energy efficiency, with our proposed accelerator outperforming the state-of-the-art FPGA and GPU implementations by 11.5 and 13.5 times, respectively. BRIEF DESCRIPTION OF THE DRAWINGS

[0034] Figure 1 is a flowchart of the NSVS search algorithm, which is also a hardware architecture diagram;

[0035] Figures 2A to 2C illustrate the search areas of DSVS and NSVS in a two-dimensional example, wherein Figure 2A illustrates the search area of ​​DSVS, Figure 2B illustrates the search area of ​​NSVS technology in scenario A, and Figure 2C illustrates the search area of ​​NSVS technology in scenario B;

[0036] FIG3 illustrates the 17 voxels that are added to the search region when the query point is located at the upper left corner of the center voxel. DETAILED DESCRIPTION

[0037] Below in conjunction with specific embodiment, further set forth the present invention.Should be understood that these embodiments are only used to illustrate the present invention and are not used in limiting the scope of the present invention.In addition, should be understood that after reading the content taught by the present invention, those skilled in the art can make various changes or modifications to the present invention, and these equivalent forms fall equally within the scope limited by the appended claims of the application.

[0038] In order to quickly search for k-nearest neighbors in large-scale point cloud images, the NSVS framework proposed in this paper includes two parts: building a DSVS search structure and searching for k-nearest neighbors based on DSVS.

[0039] Part 1: Building the DSVS search structure:

[0040] As shown in FIG2A , the embodiment of the present invention first divides the space where the reference set is located into a large number of voxels (as shown by the solid lines in FIG2A to FIG2C ), and the side length of each voxel is equal to the desired search range R in If the k-nearest neighbor result exceeds the search range R in , it is considered an outlier. Then, the embodiment of the present invention calculates the number of points in each voxel and further divides the dense voxels containing more points into sub-voxels (as shown by the dotted lines in Figures 2A to 2C). Finally, the embodiment of the present invention sorts the reference set to ensure that the points in adjacent voxels or adjacent sub-voxels are stored in a continuous memory for fast search. Voxels and sub-voxels are indexed by hash and sub_hash values.

[0041] Part 2: Searching for k-nearest neighbors:

[0042] This embodiment of the present invention first finds the query point q by calculating the values ​​of hash and sub_hash. Next, this embodiment of the present invention uses the proposed NSVS technique to narrow the search area to a few voxels and subvoxels, as shown in the dark area in Figure 2B. Then, the adaptive data transmission technology proposed in this invention is used to extract candidate neighboring points within the narrowed search area. This embodiment of the present invention uses three 512-bit high-performance interfaces, each of which is responsible for transmitting the position coordinates of the point in the X, Y, or Z direction to maximize the transmission rate. Finally, the k-nearest neighbors are selected from the candidate neighboring points.

[0043] This embodiment of the present invention implements the NSVS framework on a heterogeneous system with powerful processors (CPUs) and programmable logic (PLs), as shown in Figure 1. Using the HLS analysis tool, we found that building a DSVS is not a computationally intensive task. Therefore, this embodiment of the present invention chooses to perform the construction task on the PS side and accelerate the complex k-nearest neighbor search task on the PL side.

[0044] The k-nearest neighbor search accelerator consists of the following four parts:

[0045] 1. Data reuse buffer: On-chip cache for large reference sets.

[0046] The data reuse rate is defined as the ratio of the number of query points to the number of reference points. When the data reuse rate exceeds a threshold, the reference set is accessed sequentially using the data reuse buffer. When the data reuse rate falls below the threshold, the reference set is accessed directly and randomly via multiple large-bitwidth ports. For small reference sets with point clouds under 400,000 points, query points are densely distributed within the reference set, resulting in a high data reuse rate. In this case, embodiments of the present invention utilize a customized data reuse buffer. Since points within the same voxel and subvoxel are sequentially placed in physically contiguous memory after point cloud sorting, circular partitioning of the block RAM on the FPGA allows for simultaneous access to points within a voxel or subvoxel. The data reuse buffer is continuously updated based on query point changes, ensuring that points surrounding the current query point are always available in the data reuse buffer. Candidate neighbors of a query point, if already stored in the data reuse buffer during a previous query point search, can be directly retrieved from the data reuse buffer. This reduces the total amount of data transferred from external memory. However, when the data reuse rate is low, query points may be so sparsely distributed that adjacent query points rarely share common candidate neighbors. Therefore, for very large reference sets, up to 6 million points, this embodiment of the present invention uses random access mode to retrieve candidate neighbors from external memory. The candidate neighbors for a query point only comprise a small portion of the reference set, eliminating the need to transfer other reference points from external memory to the accelerator. Consequently, the disclosed solution automatically selects the faster transfer mode based on data reuse, increasing the robustness of the present invention to diverse datasets.

[0047] 2. Function module 1: locate the query point and narrow the search range.

[0048] The embodiments of the present invention use different strategies to narrow the search range according to the location of the query point. When the query point is located in a subvoxel, that is, a dense area with a large number of reference points, the present invention significantly narrows the search area.

[0049] 3. Function module 2: Extract candidate neighboring points based on the search area obtained by function module 1.

[0050] Candidate neighbors may come from the data reuse buffer or directly from the external memory, which is determined by the data reuse rate.

[0051] 4. Function module 3: Select k-nearest neighbors from the candidate neighboring points obtained in function module 2 through a highly parallel k-nearest neighbor selection accelerator.

[0052] The embodiment of the present invention optimizes the throughput of these three hardware functions to 1, and implements the data channels between the functions as stream types to enable task-level pipelining.

[0053] FIG2A to FIG2C and FIG3 describe how to reduce the number of candidate neighbors based on NSVS, which further includes two steps:

[0054] Step 1: Select only the nearest subvoxel in each neighboring voxel of the voxel containing the query point.

[0055] Step 2: Select the 17 closest search voxels and search subvoxels from these voxels and subvoxels in step 1.

[0056] For example, in the embodiment of the present invention, the search area is limited to 27 voxels in a 3×3×3 area around the query point. In the DSVS data structure, the number of sub-voxels is difficult to control because it selects a radius of R. in The uncertainty of the number of subvoxels increases the difficulty of hardware implementation. Therefore, in NSVS, if adjacent voxels are further divided into subvoxels, the embodiment of the present invention only selects the subvoxel closest to the voxel containing the query point for the following reasons:

[0057] 1) The fact that a voxel is split into subvoxels means that the voxel is very dense. We can ignore the other subvoxels that are farther away as they do not have much impact on the accuracy.

[0058] 2) Regardless of whether a voxel is divided into sub-voxels, the number of search voxels or search sub-voxels is at most 27, which reduces logic complexity and is conducive to hardware implementation.

[0059] When selecting the nearest subvoxel, there are two cases:

[0060] Case A: the voxel containing the query point is not further split into subvoxels;

[0061] Case B: The voxel containing the query point is further segmented.

[0062] Figures 2A to 2C show a simplified two-dimensional example. As shown in Figure 2A, the DSVS proposed in the present invention selects all voxels and sub-voxels within the search sphere as the search area. In Figures 2B and 2C, the voxels and sub-voxels in the dark area are the search areas in Case A and Case B, respectively. In Case B, the search area can be further narrowed. Since the points in the central voxel are very dense, it is almost impossible for the five voxels in the lower right corner to contain the k-nearest neighbors of point q. Therefore, due to the NSVS technology, the number of search areas is reduced from 19 to 8 in Case A, and the number of search areas is reduced to 7 in Case B. In the three-dimensional case, the number of search voxels and search voxels in Case A and Case B are 27 and 17, respectively.

[0063] To further balance the hardware implementation load, as shown in Figure 3, this embodiment of the present invention further reduces the number of search voxels and sub-voxels in Case A to 17. First, this embodiment of the present invention determines the position of the query point relative to the voxel center, within a total of eight regions. Then, based on the query point's position, the 17 nearest voxels or sub-voxels are added to the search region. For example, if the query point is located at the upper left corner of the center voxel, the 17 nearest voxels are shown in Figure 3.

[0064] The above solution can be applied to the k-nearest neighbor search step in the simultaneous localization and mapping process of autonomous vehicles or robots. The point cloud can be composed of lidar data. The proposed solution can accurately and quickly perform the k-nearest neighbor search task for different data sets. FPGA acceleration enables the entire algorithm to achieve better real-time performance and consume less energy.

Claims

1. A fast and energy-efficient large-scale point cloud K-nearest neighbor search accelerator, characterized in that, An NSVS framework for searching based on the DSVS search structure is constructed, and a large-scale point cloud graph k-nearest neighbor search algorithm is implemented on the FPGA, including: Construct the DSVS search structure: The space where the reference set is located is divided into different voxels, and the side length of each voxel is equal to the desired search range R in , if the k-nearest neighbor results exceed the desired search range Rin, they are regarded as outliers; based on the number of points at the midpoint of each voxel, the voxels containing more points are further divided into sub-voxels; the reference set is sorted to ensure that the points in adjacent voxels or adjacent sub-voxels are stored in consecutive memory locations; Search for k-nearest neighbors based on DSVS, which further includes: Function module 1: Locate the query point and adopt different strategies according to the position of the query point to narrow the search range: If the adjacent voxels of the voxel containing the query point are not further divided into sub-voxels, then these adjacent voxels are used as search voxels; If the adjacent voxels of the voxel containing the query point are further divided into sub-voxels, then only the sub-voxel closest to the voxel containing the query point is selected as the search sub-voxel in these adjacent voxels; The search voxels and the search sub-voxels form a narrowed search area; Function module 2: Extract candidate neighboring points within the narrowed search area, where: Define the data reuse rate, which is the ratio of the number of query points to the number of reference points in the search area: If the data reuse rate exceeds the threshold, access the reference set sequentially using the data reuse buffer, and this data reuse buffer is continuously updated according to the change of the query point, so as to ensure that the points around the current query point are always in the data reuse buffer; If the data reuse rate is lower than the threshold, obtain candidate neighboring points from the external memory in a random access mode; Function module 3: Select k-nearest neighbors from the candidate neighboring points obtained in function module 2 through a highly parallel k-nearest neighbor selection accelerator; Execute the task of constructing the DSVS search structure on the PS side, and accelerate the task of searching for k-nearest neighbors based on DSVS on the PL side.

2. The fast high-energy-efficient large-scale point cloud K-nearest neighbor search accelerator according to claim 1, characterized in that The voxel and the sub-voxel are indexed by hash value and sub_hash value respectively.

3. A fast and high-energy-efficient large-scale point cloud K-nearest neighbor search accelerator as claimed in claim 1, wherein, The number of the search voxels and the search sub-voxels is at most 17.

4. A fast and high-energy-efficient large-scale point cloud K-nearest neighbor search accelerator according to claim 1, characterized in that In function module 2, if the candidate neighboring points of the query point have been stored in the data reuse buffer during the search process of the previous query point, then directly obtain the candidate neighboring points from the data reuse buffer.

5. A fast and high-energy-efficient large-scale point cloud K-nearest neighbor search accelerator as claimed in claim 1, wherein Optimize the throughput of function module 1, function module 2, and function module 3 to 1, and implement the data channels between function modules as stream type to enable task-level pipelining.

Citation Information

Patent Citations

  • Three-dimensional laser radar point cloud efficient K-nearest neighbor search algorithm for unmanned driving

    CN111860340A

  • Approximate nearest neighbor search method combining VP tree and guide nearest neighbor graph

    CN112287185A

  • Graph structure-based point cloud clustering GPU (Graphics Processing Unit) optimization method and device

    CN114240729A

  • Fast high-energy-efficiency large-scale point cloud K-nearest neighbor search accelerator

    CN117788591A

  • Method for performing efficient similarity search

    US20100106713A1