Stream Processing and Non-Blocking ORB Feature Extraction Accelerator

By designing a stream processing and non-blocking ORB feature extraction accelerator on the FPGA platform, the problem of slow ORB feature extraction speed on low-power platforms is solved, efficient feature extraction is achieved, and the operation efficiency of algorithms such as ORB_SLAM is improved.

CN117726921BActive Publication Date: 2025-06-27SHANGHAI TECH UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202311468288.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-11-06
Publication Date
2025-06-27
Estimated Expiration
2043-11-06

AI Technical Summary

Technical Problem

When the prior art runs on low-power platforms, the ORB feature extraction speed is slow, resulting in low operating efficiency of algorithms such as ORB_SLAM, especially in small robots or drones.

Method used

A low-power ORB feature extraction accelerator based on FPGA is designed, using stream processing and non-blocking hardware architecture, combined with cache management algorithm and hardware sorting module, to realize parallel and non-blocking of rBRIEF descriptor computing.

Benefits of technology

It significantly improves the speed of ORB feature extraction, reduces latency and resource occupation, ensures the quality of feature points, and does not significantly reduce the accuracy of algorithms such as ORB_SLAM.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117726921B_ABST
    Figure CN117726921B_ABST
Patent Text Reader

Abstract

The present invention discloses an FPGA-implemented ORB feature extraction accelerator based on stream processing and non-blocking, which mainly includes the following two aspects of innovation: proposing a non-blocking hardware architecture and buffer management algorithm based on stream processing. This technology precisely controls and buffers each column of the rBRIEF descriptor calculation window through an algorithm, enabling the descriptor to receive new pixel stream inputs while calculating, so as to achieve non-blocking processing. Proposing an efficient hardware sorting design embedded inside the accelerator. This technology is based on the counting sort algorithm, uses extremely few resources to implement the sorting of rBRIEF on hardware, and embeds it inside the accelerator. The present invention ensures the quality of feature points while achieving high-speed feature point extraction, and will not significantly reduce the accuracy of algorithms such as ORB_SLAM.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a low-power ORB feature extraction accelerator based on FPGA. Background Art

[0002] ORB feature extraction [1] is a common algorithm in computer vision. Its main function is to extract the pixel points (feature points) that remain unchanged in consecutive images, and calculate the descriptors of the feature points (for feature point matching). Subsequently, various SLAM algorithms (such as ORB_SLAM) can be applied to estimate the motion trajectory of the camera by extracting and comparing the features of the front and rear frames of pictures, and at the same time, a point cloud map composed of feature points can be obtained. This algorithm is widely used in the fields of visual positioning and 3D reconstruction.

[0003] Currently, with the continuous improvement of the resolution and frame rate of cameras, the speed of ORB feature extraction is greatly challenged, especially when running on low-power platforms such as small robots or drones. Taking ARM Cortex-A53 as an example, the running speed of ORB_SLAM is only 6 frames per second, far from meeting the frame rate of the camera, and the ORB feature extraction accounts for more than 60% of the processing time per frame. By accelerating the speed of ORB feature extraction, the running efficiency of algorithms such as ORB_SLAM can be significantly improved.

[0004] To solve this problem, relevant scholars have proposed a series of hardware designs to accelerate the operation of ORB feature extraction. One of the difficulties in the hardware design of ORB feature extraction is that the calculation of the rBRIEF descriptor [5] is complex and difficult to parallelize. eSLAM [2] proposed the RS-BRIEF descriptor, which can significantly reduce the computational complexity, but leads to a significant reduction in accuracy. FSLAM [3] uses methods such as quantization and look-up tables to accelerate the calculation of the feature point angle, but the overall calculation speed of rBRIEF is limited. On the other hand, due to the difficulty of applying pipeline processing to the calculation of the rBRIEF descriptor, when calculating a descriptor, the input pixel stream will be blocked until the descriptor calculation is completed. The designs in [2][3][4] are all based on blocking calculations, and the delay caused by blocking accounts for 30%. In addition, the output descriptors need to be sorted according to their response values. Since the data bit width of the descriptors is large and the number is large, the usual hardware sorting will occupy a large amount of on-chip memory and logic resources. eSLAM [3] uses hardware heap sorting, and ac2SLAM [4] adds ping-pong buffers on the basis of heap sorting to reduce resource occupancy.

[0005] References

[0006] [1]E. Rublee, V. Rabaud, K. Konolige, and G. Bradski, “Orb: An efficient alternative to sift or surf,” in 2011 International Conference on Computer Vision, 2011, pp. 2564–2571.

[0007] [2]R. Liu, J. Yang, Y. Chen, and W. Zhao, “eslam: An energy-efficient accelerator for real-time orb-slam on fpga platform,” in Proceedings of the 56th Annual Design Automation Conference 2019, 2019, pp. 1–6.

[0008] [3]V. Vemulapati and D. Chen, “Fslam: an efficient and accurate slam accelerator on soc fpgas,” in 2022 International Conference on Field-Programmable Technology(ICFPT), 2022, pp. 1–9.

[0009] [4]C. Wang, Y. Liu, K. Zuo, J. Tong, Y. Ding, and P. Ren, “ac 2slam: Fpga accelerated high-accuracy slam with heapsort and parallel keypoint extractor,” in 2021 International Conference on Field-Programmable Technology(ICFPT), 2021, pp. 1–9.

[0010] [5]M. Calonder, V. Lepetit, C. Strecha, and P. Fua, “Brief: Binary robust independent elementary features,” in Proc. 11th European Conference on Computer Vision, 2010, pp. 778–792. Summary of the Invention

[0011] The objective of the present invention is to accelerate the speed of ORB feature extraction to significantly improve the operation efficiency of algorithms such as ORB_SLAM.

[0012] To achieve the above objective, the technical solution of the present invention is to provide an ORB feature extraction accelerator based on stream processing and non-blocking implemented by FPGA, including

[0013] A downsampling module, which is used to adjust the input image to the required size to obtain the original image;

[0014] A Gaussian filtering module, which is used to blur the original image to obtain a blurred image;

[0015] A corner detection module, which is used to determine the feature point positions in the blurred image and output a feature point mask;

[0016] A non-maximum suppression module, which is used to perform non-maximum suppression on the feature point mask. It is characterized in that it further includes:

[0017] An rBRIEF calculation module, which enables parallel calculation of window column selection, rBRIEF calculation, and window column FIFO through a cache management algorithm, including the following steps:

[0018] The working area of the rBRIEF descriptor is a 37x37 sliding window, which is updated in each clock cycle. Whenever the sliding window is updated: the middle pixel left pixel of the leftmost column of pixels of the feature point mask and the middle pixel right pixel of the rightmost column of pixels are detected; if the middle pixel right pixel is a feature point, the middle column of the blurred image is regarded as the first column of a certain window, the middle column of the blurred image is written into the window column FIFO, and at the same time the counter is cleared, and the next 36 middle columns are also written into the window column FIFO; if the middle pixel left pixel is a feature point, the middle column of the blurred image is regarded as the last column of a certain window, and a label is attached when each middle column is written into the window column FIFO to indicate whether it is the starting column or the ending column, and the rBRIEF calculation module decides whether to perform rBRIEF calculation according to the label;

[0019] After the rBRIEF calculation module reads data from the window column FIFO, it concatenates the data into another window matrix. A new column is inserted at the end of the window matrix in each clock cycle, and the other columns of the window matrix are shifted left in sequence. The window concatenation runs at a throughput of 1 column per cycle until a certain column is the end column; then the reading of the window column FIFO is stopped, and then the rBRIEF calculation is performed. When calculating rBRIEF, the original image is used to calculate the centroid direction, and then the BRIEF coordinates are rotated according to the direction angle to calculate the rBRIEF descriptor.

[0020] The sorting module is a counting sort implemented in hardware.

[0021] Preferably, 256 containers are used inside the sorting module to cache the indices of the feature points, and the indices corresponding to the feature points with the same response value are cached in the same container.

[0022] Preferably, the sorting module is integrated with the rBRIEF calculation module. When an rBRIEF descriptor is calculated, its index is immediately stored in the container; when all rBRIEF descriptors are calculated, the indices are output in sequence.

[0023] The above technical solutions mainly include the following two aspects of innovation:

[0024] 1) A non-blocking hardware architecture and buffer management algorithm based on stream processing are proposed. This technology precisely controls and buffers each column of the rBRIEF descriptor calculation window through an algorithm, enabling the descriptor to receive new pixel stream inputs while calculating to achieve non-blocking processing.

[0025] 2) An efficient hardware sorting design embedded inside the accelerator is proposed. This technology is based on the counting sort algorithm, uses extremely few resources to implement the sorting of rBRIEF in hardware, and embeds it inside the accelerator.

[0026] The technical solutions disclosed in the present invention are mainly applied to visual SLAM on low-power platforms, and achieve high-speed and low-power ORB feature extraction through a unique architecture. The non-blocking rBRIEF descriptor calculation in the present invention significantly improves the data throughput, and the integrated hardware sorting module further reduces the overall latency and resource occupancy. The present invention ensures the quality of feature points while achieving high-speed feature point extraction, and does not significantly reduce the accuracy of algorithms such as ORB_SLAM. BRIEF DESCRIPTION OF THE DRAWINGS

[0027] Figure 1 Schematically shows the system architecture of the accelerator;

[0028] Figure 2Schematically shows the hardware architecture for rBRIEF descriptor calculation;

[0029] Figure 3 Schematically shows the hardware architecture of the sorting module;

[0030] Figure 4 Schematically shows the comparison of experimental data;

[0031] Figure 5 Schematically shows the verification of non-blocking calculation. Detailed implementation manners

[0032] The following further elaborates the present invention in conjunction with specific embodiments. It should be understood that these embodiments are only used to illustrate the present invention and not to limit the scope of the present invention. In addition, it should be understood that after reading the content taught by the present invention, those skilled in the art can make various changes or modifications to the present invention, and these equivalent forms also fall within the scope defined by the appended claims of this application.

[0033] The accelerator disclosed by the present invention has been fully implemented to run on the FPGA. The system framework of the entire accelerator is as Figure 1 shown, and the specific working process of this accelerator is as follows:

[0034] Image data is input from the ARM processor ( Figure 1 PS part in it) to the FPGA ( Figure 1 PL part in it) in the form of a pixel stream through AXIDMA. The FPGA accelerator includes six sub-modules: downsampling, Gaussian filtering, corner detection (FAST), non-maximum suppression (NMS), descriptor calculation (rBRIEF), and sorting (Sort), where:

[0035] The downsampling module (Resizer) adjusts the input image to the required size using the linear interpolation method.

[0036] The image with adjusted size is input to the Gaussian filtering module (Gaussian Filter), and the image is blurred using a fixed Gaussian convolution kernel.

[0037] The blurred image and the original image are synchronously input to the corner detection module (FAST). This module determines the positions of feature points in the blurred image through the FAST algorithm. The mask of the feature points (KP mask), the blurred image, and the original image are synchronously output. Among them, the data value on the mask of the feature points is the harris response value of this point

[0038] The non-maximum suppression module (NMS) performs 3×3 non-maximum suppression on the mask of the feature points, and synchronously outputs the blurred image, the original image, and the mask of the feature points after non-maximum suppression.

[0039] Calculation of rBRIEF descriptor ( Figure 1 in rBRIEF) is divided into three parts: window selection, rBRIEF calculation, and window column FIFO. The present invention enables parallel calculation of these three parts by proposing a new cache management algorithm. The overall algorithm structure is as Figure 2 shown. The following is the process description of the algorithm:

[0040] 1) Window column selection

[0041] The working area for rBRIEF descriptor calculation is a 37x37 sliding window, which is updated every clock cycle. Whenever the sliding window is updated, we detect the two middle pixels of the leftmost and rightmost columns of the feature point mask ( Figure 2 left pixel and right pixel in Figure 2 ), and then decide whether to write the middle column of the blurred image into the FIFO. If the right pixel is a feature point, then the middle column (the middle column) of the blurred image can be regarded as the first column (from left to right) of a certain window, as shown by the dashed box 3 in Figure 2 . This middle column will be written into the window column FIFO, and the counter is cleared. Moreover, the next 36 middle columns will also be written into the window column FIFO. If the left pixel is a feature point, then the middle column can be regarded as the last column of a certain window (shown by the dashed box 1 in

[0042] b) rBRIEF calculation.

[0043] As Figure 2 shown in region 5 in

[0044] c) Window column FIFO.

[0045] Window column buffering ( Figure 2The existence of the middle region 4 allows for different production and consumption rates of the window columns. By controlling the depth of the window column FIFO, the rBRIEF calculation can be achieved while new pixel inputs are not interrupted.

[0046] The last step is sorting. This design uses hardware-implemented counting sort, as Figure 3 shown. The hardware sort internally uses 256 containers to cache the indices of the feature points. The indices corresponding to the feature points with the same response value are cached in the same container. The sorting module is actually integrated with the rBRIEF module. When a descriptor calculation is completed, its index is immediately stored in the container. When all descriptor calculations are completed, the indices are output sequentially in order. Since only the indices of the feature points are cached, the additional resource occupancy only requires one BRAM and a few hundred LUTs.

[0047] Figure 4 Shows the comparison between the present invention and the original ORB_SLAM3 and existing accelerator designs (eSLAM, ac2SLAM, and FSLAM). The experimental data shows that the present invention has the lowest latency, resource occupancy, and the highest accuracy, as well as the lowest energy consumption. Figure 5 is the curve of the accelerator's latency varying with the feature point density, indicating that under actual working conditions, the present invention realizes the non-blocking calculation of the rBRIEF descriptor.

Claims

1. A stream processing and non-blocking ORB feature extraction accelerator, comprising a downsampling module for resizing an input image to a required size to obtain an original image; a Gaussian filtering module for blurring the original image to obtain a blurred image; a corner detection module for determining feature points in the blurred image and outputting a feature point mask; a non-maximum suppression module for performing non-maximum suppression on the feature point mask, characterized in that it further comprises: an rBRIEF calculation module, which takes the original image, the blurred image and the feature point mask output by the non-maximum suppression module as inputs, calculates the rBRIEF descriptor and outputs it to the sorting module, and enables parallel calculation of window column selection, rBRIEF calculation and window column FIFO through a cache management algorithm, including the following steps: The working area of the rBRIEF descriptor is a 37x37 sliding window, which is updated every clock cycle. Whenever the sliding window is updated: detect the middle pixel left pixel of the leftmost column of pixels of the feature point mask and the middle pixel right pixel of the rightmost column of pixels; if the middle pixel right pixel is a feature point, the middle column of the blurred image is regarded as the first column of a certain window, write the middle column of the blurred image into the window column FIFO, and at the same time clear the counter, and the next 36 middle columns are also written into the window column FIFO; if the middle pixel left pixel is a feature point, the middle column of the blurred image is regarded as the last column of a certain window, and a label is attached when each middle column is written into the window column FIFO to indicate whether it is the starting column or the ending column, and the rBRIEF calculation module decides whether to perform rBRIEF calculation according to the label; After the rBRIEF calculation module reads data from the window column FIFO, it splices them into another window matrix. A new column is inserted into the end of the window matrix every clock cycle, and the other columns of the window matrix are shifted left in turn. The window splicing runs at a throughput of 1 column per cycle until a certain column is the ending column; then stop reading the window column FIFO, and then perform rBRIEF calculation. When calculating rBRIEF, the original image is used to calculate the centroid direction, and then the BRIEF coordinates are rotated according to the direction angle to calculate the rBRIEF descriptor; a sorting module, which is a counting sort implemented in hardware; The sorting module is integrated with the rBRIEF calculation module. When an rBRIEF descriptor is calculated, its index is immediately stored in a container; when all rBRIEF descriptors are calculated, the indexes are output in sequence.

2. The ORB feature extraction accelerator based on stream processing and non-blocking as claimed in claim 1, wherein The sorting module internally uses 256 containers to cache the indexes of feature points, and the indexes corresponding to feature points with the same response value are cached in the same container.

Citation Information

Patent Citations

  • An ORB-SLAM hardware accelerator

    CN109919825A

  • ORB_SLAM relocation feature point retrieval acceleration method based on FPGA

    CN113536024A