Acceleration method and system of laser radar odometer based on FPGA (Field Programmable Gate Array)

Through point cloud acquisition, floating point quantization and hash mapping optimization on the FPGA platform, the compatibility problem of lidar odometer operation speed and hardware resource consumption in miniaturized equipment is solved, and efficient and accurate lidar point cloud data processing and pose calculation are achieved.

CN120468872APending Publication Date: 2025-08-12XI AN JIAOTONG UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510589407.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-08
Publication Date
2025-08-12

AI Technical Summary

Technical Problem

In the prior art, the lidar odometer based on the FPGA platform is difficult to compatible with the operating speed and hardware resource consumption, making it difficult to achieve a balance between real-time and computing efficiency in miniaturized smart terminal devices.

Method used

A FPGA-based lidar odometer acceleration system is adopted, through the collaborative work of the PS and PL ends, including point cloud acquisition, floating point quantization, local map management and point cloud sorting modules, the data interaction is used by AXI4 and AXI Lite buses, and the point cloud processing is optimized using hash mapping and pipeline structure to achieve efficient storage and calculation of point cloud data.

Benefits of technology

It significantly improves processing speed, reduces computing and storage costs, ensures the accuracy of adjacent point cloud search and the accuracy of pose calculation, and realizes efficient and accurate lidar point cloud data processing and pose calculation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120468872A_ABST
    Figure CN120468872A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of FPGA platform algorithm hardware acceleration, in particular to an FPGA-based laser radar odometer acceleration method and system, and the system is composed of a PS end and a PL end cooperatively. The PS end comprises a point cloud acquisition module (for acquiring a plurality of frames of 32-bit floating point laser radar point clouds and transmitting the point clouds to the PL end) and a pose calculation module (for calculating a current pose based on a search result of the adjacent point clouds of the PL end). The PL end comprises a floating point number quantization module (converting two frames of 32-bit floating point cloud into 16-bit fixed points), a local map management module (building a local map for the first frame of 16-bit point cloud in a Hash mapping manner and searching a secondary frame of neighborhood voxel to obtain coordinates) and a point cloud sorting module (calculating the distance between the neighborhood point cloud and the first frame, sorting the distance and matching the nearest neighbor point cloud to form a search result). The system realizes point cloud efficient processing and pose estimation through software and hardware cooperation.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of FPGA platform algorithm hardware acceleration, and in particular to an acceleration method and system for a laser radar odometer based on FPGA. Background Art

[0002] LiDAR odometry is a technology that uses only three-dimensional point cloud observations to perform high-precision pose estimation without requiring external signals. Its stable and accurate positioning capabilities in complex scenarios, including strong lighting, dark environments, and environments with irregular geometries, have garnered widespread attention and application in areas such as autonomous drones and intelligent mobile robots.

[0003] The positioning accuracy of lidar odometry is closely linked to the performance of the sensor. In recent years, with the rapid development of solid-state laser technology, the density of point clouds captured by lidar has increased significantly. This has enabled the odometry to obtain richer geometric feature observations, but it has also led to a significant increase in the computational complexity of point cloud processing. When deployed on general-purpose processor platforms, this cannot meet real-time requirements, making it difficult to deploy in miniaturized drones or other small intelligent robots, which are limited in computing resources, battery power, capacity, and payload. FPGAs, with their flexible programmability, low cost, and low power consumption, are well-suited for deploying lidar odometry systems in some small intelligent terminal devices. The high parallelism of FPGAs ensures real-time operation.

[0004] In the paper "An SoC-FPGA-based iterative-closest-point accelerator enabling faster picking robots[J]," the authors proposed a point cloud registration algorithm accelerator based on a CPU-FPGA computing platform. Leveraging the heterogeneous FPGA platform, the authors designed a dynamically reconfigurable KD tree construction and search unit. They also designed a parallel distance calculation unit and sorting circuit using the FPGA. They also reduced memory resource consumption by reusing factor graphs. However, implementing the iterative KD tree search process on an FPGA platform did not achieve ideal acceleration, was difficult to implement, and consumed a lot of hardware resources.

[0005] In the paper "An optimized FPGA-based real-time NDT for 3D-LiDAR localization in smart vehicles[J]," the authors proposed an FPGA-based NDT lidar odometry accelerator design. This work uses 3D voxel and subvoxel partitioning to perform nearest neighbor indexing of point clouds. Compared to global search algorithms such as KD trees, this method is faster. However, further partitioning of 3D voxels consumes significant on-chip memory resources.

[0006] In summary, although there are multiple solutions for the hardware acceleration design of lidar odometry based on FPGA platforms, there is still a lack of a method that takes into account both the operating speed and hardware resource consumption of the lidar odometry. Summary of the Invention

[0007] The technical problem to be solved by the present invention is to address the deficiencies in the above-mentioned prior art and provide an acceleration method and system for a lidar odometer based on FPGA, so as to solve the technical problem that the operating speed of the lidar odometer is incompatible with the hardware resource consumption.

[0008] The purpose of the present invention is achieved by the following technical solutions: In a first aspect, the present invention provides an acceleration system for a laser radar odometer based on FPGA, the system including a PS end and a PL end for data interaction; The PS side includes a point cloud acquisition module and a pose calculation module. The point cloud acquisition module is used to acquire multi-frame lidar point clouds and send the acquired point clouds to the PL side. The lidar point cloud is a 32-bit floating point point cloud. The pose calculation module is used to obtain the neighboring point cloud search results from the PL side and calculate the current pose of the lidar based on the neighboring point cloud search results. The PL side includes a floating-point quantization module, a local map management module, and a point cloud sorting module; The floating-point quantization module is used to receive two consecutive frames of 32-bit floating-point point clouds and quantize them into 16-bit fixed-point point clouds; The local map management module is used to obtain the 16-bit fixed-point point cloud of the first frame, perform hash mapping based on the voxel coordinates of the first frame point cloud to form a local map of the point cloud, and search the neighborhood voxel range of the second frame point cloud based on the local map of the point cloud to obtain the neighborhood point cloud coordinates; The point cloud sorting module is used to calculate the distance between the neighborhood point cloud coordinates and the first frame point cloud, sort the distance results, and match the closest point cloud with the first frame point cloud according to the sorting results to form a neighborhood point cloud search result.

[0009] As a further improvement of the present invention, the point cloud acquisition module exchanges data with the PL end through the AXI4 bus; the posture calculation module exchanges data with the PL end through the AXI Lite bus.

[0010] As a further improvement of the present invention, the floating-point quantization module includes an exponent adjustment submodule, a shift submodule and a complement calculation submodule; The exponent adjustment submodule is used to receive the exponent part of the 32-bit floating point cloud and calculate the exponent part according to the set rules to obtain the exponent value; The shift submodule is used to receive the decimal part of the 32-bit floating point cloud, shift the decimal part to the right according to the exponent value obtained by the exponent adjustment submodule, and truncate the high bits of the shifted data according to the set bit width to obtain tmp_dat; The complement calculation submodule is used to calculate the complement of tmp_dat to obtain a quantized 16-bit fixed-point point cloud.

[0011] As a further improvement of the present invention, the local map management module includes a hash controller and a physical memory; The hash controller is used to obtain two consecutive frames of 16-bit fixed-point point clouds, calculate voxel coordinates based on the first frame of fixed-point point cloud, perform hash mapping based on the voxel coordinates, store the first frame of point cloud in the voxels to form a point cloud local map, search the neighborhood voxel range of the second frame of point cloud based on the point cloud local map to obtain neighborhood point cloud coordinates, generate control signals and physical storage addresses based on the neighborhood point cloud coordinates, and perform data interaction with the physical memory based on the physical storage addresses; The physical memory is used to store and / or read the point cloud local map and the corresponding point cloud under the control of the hash controller.

[0012] As a further improvement of the present invention, the hash controller includes a point cloud coordinate conversion unit, a hash state machine, a hash mapping unit, and a distance comparison unit; The point cloud coordinate conversion unit is used to calculate the spatial voxels corresponding to the point cloud according to the point cloud coordinates, and convert the point cloud coordinates into the index of the spatial voxels; The hash state machine is used to read the point cloud voxel index after coordinate transformation, and then control the hash mapping unit to complete the hash function calculation. It also controls the physical memory module through the external RAM bus to complete the storage and reading of point cloud and voxel information in the local map. The hash mapping unit includes a two-stage pipeline structure. The first stage pipeline performs an XOR calculation on the hash function through a parallel DSP to obtain an XOR operation result. The second stage pipeline includes a modulus operation unit, which is used to perform a modulus calculation on the XOR operation result to obtain a reduction result. The distance comparison unit is used to perform distance comparison between the neighboring point cloud and the first frame point cloud according to the reduction result to obtain a comparison result.

[0013] As a further improvement of the present invention, the working state of the hash state machine includes an insert mode and a search mode; When the working state is in insert mode: the hash state machine is in the initial state, waiting for a valid start signal to be input, and then jumps to the data handshake state; In the data handshake state, the hash state machine pulls the ready signal in the current module high and waits for the point cloud coordinate conversion unit to output a valid point cloud voxel index. When the ready signal and the valid signal are detected to be high at the same time in this cycle, the hash state machine splices the point cloud index data corresponding to the three direction axes into a hash key, and in the next clock cycle, the hash state machine jumps to the hash mapping state. In the hash mapping state, the hash state machine starts hash function calculation; after the calculation is completed, it jumps to the hash search state; In the hash search state, if the same hash key as the point cloud to be inserted is found in all slots of the current hash bucket, a synchronous downsampling operation is performed on the voxels, and the hash state machine jumps to the distance comparison state. In the distance comparison state, the queried point cloud and the point cloud to be inserted are passed to the distance comparison unit. The current state of the hash state machine is switched according to the result obtained by the distance comparison unit. If the same hash key as the point cloud to be inserted is not found in all slots of the current hash bucket, the hash state machine jumps to the hash insertion state. When the working state is in search mode: when the hash state machine is in the data handshake state, it waits for the next frame of valid point cloud handshake signal, and jumps to the voxel index calculation state after the handshake is successful; In the voxel index calculation state, the hash state machine groups the neighboring voxels and calculates the index of the neighboring voxels in each group; After the two neighbor voxel indices in the current group are calculated, the state machine jumps to the hash mapping state and passes the calculation result of the hash key to the hash mapping unit; In the hash mapping state, the hash state machine waits for the hash mapping calculation to be completed, and after the completion flag is pulled high, it jumps to the hash search state; In the hash search state, the hash state machine reads the corresponding hash key value in the physical memory; by comparing the hash key value of each slot in the current hash bucket with the hash key to be queried one by one, it determines whether there is a point cloud insertion in the current voxel to be queried, and locates the hash path, slot and point cloud frame index of the inserted point cloud; the hash state machine generates the corresponding selection control signal based on the queried hash path, slot and other information, and sends it to the physical memory to control the multi-level MUX network, select the queried point cloud frame index information to the read address bus of the physical memory, and the hash state machine jumps to the result output state; In the result output state, the hash state machine will pull up the point cloud output valid flag corresponding to the queried point cloud coordinates, and perform data handshake with the point cloud sorting module; after the handshake is successful, if the current voxel to be queried has not completed all neighbor searches, the hash state machine will jump to the voxel index calculation state and perform the next set of voxel index calculations; if the current voxel has completed all neighbor searches, the state machine will detect whether the end flag of the current frame is pulled high. If so, the state machine will jump to the initial state.

[0014] As a further improvement of the present invention, the hash mapping unit chooses to use the Montgomery algorithm for optimization.

[0015] As a further improvement of the present invention, the hash bucket in the hash state machine includes a plurality of storage slots; When all storage slots in the hash bucket are occupied, a random slot under all insertion paths is selected, the hash key-value pair stored in the random slot is kicked out, and replaced with the hash key-value pair currently to be inserted; the hash key-value pair randomly kicked out is re-inserted.

[0016] As a further improvement of the present invention, the point cloud sorting module includes a distance calculation unit based on a parallel pipeline and a sorting unit based on a linear feedback shift register; The distance calculation unit consists of two pipelines. The first pipeline is used to store the input reference frame target point cloud and neighboring point cloud coordinates. The second pipeline is used to parallelly calculate the squared distances in different directions and accumulate the squared distances through an addition tree before storing the output. The sorting unit includes several registers, several multiplexers and several comparators; the several registers respectively store the distance calculation results output by the distance calculation unit, the comparators respectively compare the distance calculation results, and after obtaining the comparison results, they are transmitted to the multiplexers as control signals; and are used to transmit to the linear feedback shift register through the multiplexers.

[0017] In a second aspect, the present invention provides an acceleration method for a laser radar odometer based on an FPGA, which is implemented based on the above-mentioned acceleration system of the laser radar odometer based on an FPGA, including: S1. Use the PS end to obtain multi-frame lidar point clouds and read two consecutive frames of point clouds to the PL end. The PL end receives 32-bit floating-point point cloud data and obtains 16-bit fixed-point point clouds after hardware quantization. After receiving the first frame of fixed-point point cloud, S2 and PL calculate the voxel coordinates and perform hash mapping to store the first frame of point cloud in the corresponding voxels to obtain a local point cloud map. After the local map is constructed, the second frame of point cloud is received and the neighborhood voxel range of the second frame of point cloud is searched in the local point cloud map to obtain the neighborhood point cloud coordinates. S3. Calculate the distance between the coordinates of the neighborhood point cloud and the reference point cloud of the first frame, sort the points according to the distance results, match the point cloud with the closest distance to the reference point cloud of the first frame, obtain the neighboring point cloud search results and store them; S4. The PS side sends the neighboring point cloud search results to the PL side; the PL side constructs the point cloud matching residual based on the neighboring point cloud search results, minimizes the matching residual through the Gauss-Newton method, and finally calculates the current position of the lidar.

[0018] The beneficial effects of the present invention are as follows: the present invention provides an acceleration system for a lidar odometer based on FPGA, which quantizes 32-bit floating-point data into 16-bit fixed-point data through a floating-point quantization module, significantly reducing the amount of data, improving the processing speed, and reducing the calculation and storage costs. Based on the collaborative work of the local map management module and the point cloud sorting module, the accuracy and efficiency of the neighboring point cloud search are ensured, thereby improving the accuracy and reliability of the pose calculation. Through the collaborative work of the PS end and the PL end, efficient and accurate lidar point cloud data processing and pose calculation are achieved. The entire system can effectively capture and process lidar point cloud data, accurately calculate the pose of the lidar at different times, and achieve precise positioning and navigation of the environment.

[0019] Furthermore, the main technical effects of this floating-point quantization module can be described in a simple and understandable way: First, the exponent adjustment submodule receives and processes the exponent portion of the 32-bit floating-point point cloud, calculating the exponent value according to the set rules; then, the shift submodule uses this exponent value to right-shift the decimal portion of the floating-point point cloud and truncate the high bits to obtain a 16-bit tmp_dat; finally, the complement calculation submodule calculates the complement of tmp_dat, thereby achieving the quantization conversion from 32-bit floating-point point cloud to 16-bit fixed-point point cloud. The entire working principle is to ensure that the quantized point cloud data not only maintains the main characteristics of the original data but also effectively reduces the consumption of storage and computing resources by accurately adjusting the exponent and reasonably truncating the shifted data.

[0020] Furthermore, the function of the hash controller is to obtain two consecutive frames of 16-bit fixed-point point cloud data, calculate the voxel coordinates based on the first frame of point cloud data, and then store the first frame of point cloud information in the corresponding voxels through hash mapping, thereby constructing a local point cloud map. The physical memory is responsible for storing and reading the local point cloud map and its related information under the guidance and control of the hash controller. The working principle is as follows: First, the hash controller receives two consecutive frames of 16-bit fixed-point point cloud data, calculates the voxel coordinates based on the first frame of data, and stores the corresponding first frame of point cloud information in the corresponding voxels through hash mapping, thereby constructing a local point cloud map. Then, based on the constructed local point cloud map, the hash controller searches the neighborhood voxel range of the second frame of point cloud, determines the neighborhood point cloud coordinates, and generates corresponding control signals and address information of the physical memory. Finally, the physical memory exchanges data with the hash controller based on the received address information, thereby realizing efficient storage and reading of point cloud data.

[0021] Furthermore, the point cloud coordinate conversion unit is responsible for converting the point cloud coordinates into spatial voxel indices for subsequent processing. The hash state machine is responsible for reading these indices and completing the calculation of the hash function by controlling the hash mapping unit, ultimately storing the point cloud information in an external physical memory module. The hash mapping unit adopts a two-stage pipeline structure. First, the hash function calculation is performed through the parallel DSP to obtain a preliminary result. Then, the modulo operation unit performs a modulo operation to obtain the position reduction result of the voxel information. Based on these reduction results, the distance comparison unit compares the distance between the neighboring point cloud and the first frame point cloud to generate a distance comparison result. Overall, the system achieves efficient management and real-time processing of point cloud data by converting point cloud coordinates into spatial voxel indices, using the hash mapping unit and pipeline structure for efficient data processing and storage, and using the distance comparison unit for distance judgment. BRIEF DESCRIPTION OF THE DRAWINGS

[0022] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0023] Figure 1 This is a framework diagram of the FPGA lidar odometer acceleration system of the present invention.

[0024] Figure 2 This is a schematic diagram of the single-precision floating-point number format of the present invention.

[0025] Figure 3This is a schematic diagram of the 16-bit fixed-point number format customized by the present invention.

[0026] Figure 4 This is a hardware circuit diagram of the floating-point quantization unit of the present invention.

[0027] Figure 5 This is the architecture diagram of the local map management module of the present invention.

[0028] Figure 6 This is the address and control bus structure diagram of the primary memory of the present invention.

[0029] Figure 7 This is a data bus structure diagram of the primary RAM and secondary RAM of the present invention.

[0030] Figure 8 This is a hardware circuit diagram of the point cloud coordinate conversion unit of the present invention.

[0031] Figure 9 This is a state transition diagram of the state machine in the insertion mode of the present invention.

[0032] Figure 10 This is a hash insertion state flow chart of the present invention.

[0033] Figure 11 This is the state transition diagram of the state machine in search mode.

[0034] Figure 12 This is a hardware circuit diagram of the hash mapping unit of the present invention.

[0035] Figure 13 This is a diagram of the Montgomery reduction unit pipeline division structure of the present invention.

[0036] Figure 14 This is a hardware circuit diagram of the modulo operation unit of the present invention.

[0037] Figure 15 This is a diagram of three hash conflict avoidance strategies designed for the present invention.

[0038] Figure 16 This is a hardware circuit diagram of the distance calculation unit of the present invention.

[0039] Figure 17 This is a hardware circuit diagram of the sorting unit of the present invention.

[0040] Figure 18 This is a flowchart of the data set testing of the present invention.

[0041] Figure 19 Schematic diagram of the data set test results of the present invention. DETAILED DESCRIPTION

[0042] In order to make the purpose and technical solution of the present invention clearer and easier to understand, the present invention is further described in detail below with reference to the accompanying drawings and embodiments. The specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.

[0043] The technical solutions of the present invention will be described clearly and completely below with reference to the accompanying drawings and specific embodiments. The described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments.

[0044] Example 1 like Figures 1 to 19 As shown, this embodiment provides an acceleration system for a laser radar odometer based on FPGA. The following is a specific implementation method.

[0045] This system includes the PS side and PL side for data interaction. Figure 1 As shown in the figure, data exchange is achieved between the PS and PL ends through the AXI4 data bus and the AXI Lite bus.

[0046] The PS side includes a point cloud acquisition module and a pose calculation module. The point cloud acquisition module is used to obtain multi-frame LiDAR point clouds and send them to the PLA side. The LiDAR point clouds are 32-bit floating point point clouds. The pose calculation module is used to obtain the neighboring point cloud search results from the PLA side and calculate the LiDAR's current pose based on these neighboring point cloud search results.

[0047] Specifically, the point cloud acquisition module on the PS side includes DDR4 memory and a DDR controller. Multiple frames of LiDAR point clouds (in this embodiment, point clouds refer to point cloud data) are sequentially written to the DDR4 memory via the DDR controller. Two consecutive frames of point clouds are then read from the DDR4 memory via the DMA controller and transmitted to the PLA side via the AXI4 data bus.

[0048] The pose calculation module includes ARM. The ARM processor constructs the point cloud matching residual with the neighboring point cloud search results, minimizes the matching residual through the Gauss-Newton method, and finally calculates the current pose of the lidar.

[0049] The PL side includes a floating-point quantization module, a local map management module, and a point cloud sorting module.

[0050] The floating-point quantization module is used to receive two consecutive frames of 32-bit floating-point point clouds and quantize them into 16-bit fixed-point point clouds.

[0051] The floating point quantization module includes the exponent adjustment submodule, the shift submodule and the complement calculation submodule. Specifically, when the LiDAR sends point cloud data through the Ethernet port, it is packaged and sent in the floating point format, such as Figure 2The storage format of a 32-bit floating point number is shown in FIG. In this embodiment, a 32-bit floating point number is quantized into a 16-bit fixed point number by a floating point quantization unit.

[0052] like Figure 4 The figure shows the hardware circuit diagram of the floating-point quantization unit. The exponent adjustment submodule receives the exponent portion of a 32-bit floating-point point cloud and calculates it according to a set rule to obtain an exponent value. The shift submodule receives the decimal portion of a 32-bit floating-point point cloud and right-shifts it according to the exponent value obtained by the exponent adjustment submodule. The high-order bits of the shifted data are truncated according to the set bit width to obtain tmp_dat. The two's complement calculation submodule calculates the two's complement of tmp_dat to obtain a quantized 16-bit fixed-point point cloud. The set rule in this embodiment is based on the IEEE 754 standard.

[0053] The shift submodule specifically includes: the shift submodule shifts the decimal part to the right according to the exponent value calculated by the exponent adjustment module, discards the bits shifted to the right, extends the high bits with the sign bit, and truncates the high bits of the shifted data according to the set bit width to obtain tmp_dat (temporary data).

[0054] The two's complement calculation module is used because fixed-point numbers need to be stored in two's complement form in subsequent hardware circuits. This module consists of a two's complement calculation circuit and a two-to-one multiplexer. It calculates the two's complement of tmp_dat by checking whether the sign bit of the floating-point number is "1".

[0055] Finally, the register outputs the quantized 16-bit fixed-point number. The entire floating-point quantization module is implemented using a single-stage pipeline structure. For a 32-bit floating-point number, the quantization process can be completed within a single clock cycle.

[0056] The local map management module is used to obtain the 16-bit fixed-point point cloud of the first frame, perform hash mapping according to the voxel coordinates of the first frame point cloud to form a local map of the point cloud, and search the neighborhood voxel range of the second frame point cloud based on the point cloud local map to obtain the neighborhood point cloud coordinates.

[0057] Among them, such as Figure 5As shown, the local map management module includes a hash controller and a physical memory; the hash controller is used to receive two consecutive frames of quantized fixed-point point cloud data, and complete the hash mapping process of local map construction, update, and maintenance, generate control signals and physical storage addresses, and interact with the physical memory for data. Specifically, two consecutive frames of 16-bit fixed-point point clouds are obtained, and the voxel coordinates are calculated based on the first frame of fixed-point point cloud. Hash mapping is performed based on the voxel coordinates, and the first frame of point cloud is stored in the voxels to form a point cloud local map. Based on the point cloud local map, the neighborhood voxel range of the second frame of point cloud is searched to obtain the neighborhood point cloud coordinates, and the control signal and physical storage address are generated based on the neighborhood point cloud coordinates, and data interaction is performed with the physical memory based on the physical storage address. The physical memory is used to store and / or read the point cloud local map and the corresponding point cloud under the control of the hash controller.

[0058] The physical memory chooses to store hash key values and point cloud coordinates through two levels of memory. Considering that the depth of the hash table is generally greater than the amount of data to be stored, the bottom layer needs to open up additional storage space when storing hash key values. If the point cloud fixed-point coordinates are directly stored as hash key values, the storage bit width is 16bit×3. The larger storage bit width will result in a large consumption of storage resources when the hash table depth is deep. Therefore, when storing hash key values, the present invention does not directly store the point cloud coordinates, but stores the frame index of the current point cloud in the primary memory according to the result of the hash mapping, and then stores the point cloud coordinates in the secondary memory according to the frame index order. The frame index of the point cloud only requires 8bit data to represent, and the bit width is reduced by 6 times compared to directly storing the point cloud coordinates, which greatly alleviates the storage pressure of the primary memory. In the secondary memory, because the point cloud coordinates are stored in index order, no additional storage space is required, and its storage resource consumption is also within an acceptable range. During the search process, the index of the current point cloud must first be searched in the primary memory according to the result of the hash function mapping. Then, the point cloud coordinates are searched in the secondary memory based on the index. This may cause an additional clock cycle delay when searching the secondary memory, but this storage structure splits the original large-bit-width, large-depth memory into a small-bit-width, large-depth memory and a large-bit-width, small-depth memory, to a certain extent balancing the storage bit width and depth and reducing storage resource consumption.

[0059] like Figure 6 The address and control bus circuit structure of the primary memory are shown in Figure 1. In order to improve the parallelism of the hash search process, both the primary memory and the secondary memory use dual-port RAM (i.e. Figure 1 Each port has the ability to read and write independently and in parallel.

[0060] like Figure 7 Figure 2 shows the data bus structure of the primary RAM and secondary RAM designed in this invention. The hash controller reads the hash key values of all slots at the current physical address from the primary RAM, selects the corresponding point cloud index to the secondary RAM's read address port via slot selection signals and path selection signals, and finally reads the corresponding point cloud coordinates from the secondary RAM's read data port, completing the hash search process.

[0061] The hash controller mainly provides a control interface and a data interface for interacting with the physical memory module, and completes the hash function mapping and the storage and reading process of the hash key value through internal state machine control.

[0062] Specifically, the hash controller includes a point cloud coordinate conversion unit, a hash state machine, a hash mapping unit, and a distance comparison unit; the point cloud coordinate conversion unit is used to calculate the spatial voxels corresponding to the point cloud according to the point cloud coordinates, and convert the point cloud coordinates into the index of the spatial voxel; the hash state machine is used to read the point cloud voxel index that has undergone coordinate conversion, and then control the hash mapping unit to complete the hash function calculation, and control the physical memory module through the external RAM bus to complete the storage and reading of the point cloud and voxel information in the local map; the hash mapping unit includes a two-stage pipeline structure, the first stage pipeline calculates the hash function through parallel DSP to obtain the operation result; the second stage pipeline includes a modulo operation unit, and the modulo operation unit is used to perform modulo calculation on the operation result to obtain a reduction result; the distance comparison unit is used to compare the distance between the neighboring point cloud and the first frame point cloud according to the reduction result to obtain a comparison result.

[0063] The main function of the point cloud coordinate conversion unit is to calculate the spatial voxel in which the point cloud is located based on the current point cloud coordinates and convert the point cloud coordinates into the index of the spatial voxel. Assume that there is a point cloud whose coordinates in the laser radar coordinate system are , and the side length of the specified space voxel is , then divide the point cloud coordinates of each axis by the side length of the spatial voxel and round down to get the voxel index of the point cloud in space .

[0064]

[0065] The data interface at the input of the point cloud coordinate conversion unit receives 16-bit fixed-point cloud coordinates in the order of the x-axis, y-axis, and z-axis, and can only accept coordinates in one direction per clock cycle. After coordinate conversion, the output interface outputs an 8-bit voxel index in each direction. The control signals at both the input and output terminals consist of a set of handshake signals, each of which includes a data valid signal (tvalid) and a ready handshake signal (tready). The signal handshake at the current data interface is successful only when both tvalid and tready are high. The internal circuitry of this unit consists of two main parts. The upper part consists of a three-stage pipeline shift register, which primarily calculates the voxel coordinates of the point cloud. The lower part, consisting of a timing counter and comparator, primarily controls the timing of the output data.

[0066] The hash state machine is the core control unit of the hash controller module. Its main function is to control the timing of the construction, maintenance, update, and search processes of the local map, and to generate corresponding control signals and flags. The hash state machine reads the coordinate-converted point cloud voxel index through the internal bus, then controls the hash mapping unit to complete the hash function calculation. It also controls the physical memory module through the external RAM bus to complete the storage and reading of the point cloud and voxel information in the local map.

[0067] Furthermore, the hash state machine has two main working modes, namely insert mode and search mode. The selection and switching of these two modes are indicated by a mode_select signal. When mode_select is 0x01, it indicates that the current state machine is working in insert mode, and when mode_select is 0x02, it indicates that the current state machine is working in search mode. In the default initial state, the state machine works in insert mode and waits for the point cloud data of the previous frame to arrive. After the local map management module completes the map construction, the top-level module will switch the state machine mode to search mode and wait for the arrival of the point cloud data of the current frame. After completing the KNN search task of all point clouds in a frame, the top level of the system will reset the state machine to insert mode again and wait for the input of the next frame of point cloud.

[0068] like Figure 9 As shown in the figure, the state transition diagram of the hash state machine in insert mode. In this mode, the state machine mainly jumps between seven states, including: The state machine initially works in S0, that is, the initial state, and after waiting for a valid start signal to be input, it jumps to S1, that is, the data handshake state.

[0069] In the S1 state, the state machine actively pulls the ready signal of the current module high and waits for the previous module to output a valid point cloud voxel index. When both the ready and valid signals are pulled high in this cycle, the state machine concatenates the point cloud index data of the three axes into a hash key and, in the next clock cycle, jumps to S2, the hash mapping state.

[0070] In the hash map state, the state machine starts the hash function calculation. When the calculation is completed, it jumps to S3, the hash search state, in the next clock cycle.

[0071] In the hash lookup state, if the same hash key as the point cloud to be inserted is found in all slots of the current hash bucket, this means that the voxel where the point cloud to be inserted is already occupied by another point cloud. In this case, the local map management module performs synchronous downsampling on that voxel. The specific process is as follows: First, the state machine jumps to S5, the distance comparison state. It simultaneously passes the point cloud queried in S3 and the point cloud to be inserted to the distance comparison module. It waits for the distance comparison module's calculation completion flag to go high. If the comparison result indicates that the point cloud to be inserted is closer to the voxel center, the state machine jumps to S6, the hash replacement state. Otherwise, the state machine abandons inserting the current point cloud and jumps to S1 to perform the handshake for the next point cloud. In the hash replacement state, the state machine marks the slot number of the already inserted point cloud and inserts the hash key value corresponding to the point cloud to be inserted into the marked slot, replacing the previously inserted point cloud.

[0072] If the same hash key is not found in the hash slot, it means that the voxel where the point cloud to be inserted is located is blank. At this time, the state machine directly jumps to S4, the hash insertion state. The processing logic of the state machine in the hash insertion state is as follows: Figure 10 shown.

[0073] When the hash state machine completes the insertion of the previous frame of point cloud in the insertion mode, the top module will send the mode handshake signal again and switch the state machine to the search mode.

[0074] like Figure 11 The state transition diagram of the hash state machine in search mode is shown. The state machine is initially in state S0, waiting for the mode switch signal to arrive. The next cycle jumps to state S1. In state S1, the state machine waits for the next frame of valid point cloud data handshake signal. After the handshake is successful, it jumps to state S7, which is the voxel index calculation state.

[0075] In the S7 state, the state machine divides the 18 neighbor voxels into 9 groups and calculates the index of the neighbor voxels in each group. After the index calculation of the two neighbor voxels in the current group is completed, the state machine jumps to the S2 state and passes the calculated hash key to the hash map unit.

[0076] The state machine waits for the hash map calculation to be completed in the S2 state, and after the completion flag is pulled high, it jumps to S3, the hash lookup state.

[0077] In the S3 state, the state machine reads the corresponding hash key value in the primary memory. By comparing the hash key value of each slot in the current hash bucket with the hash key to be queried one by one, it can be determined whether there is a point cloud insertion in the current voxel to be queried, and the hash path, slot and point cloud frame index of the inserted point cloud can be located. Then, the state machine will generate the corresponding selection control signal based on the queried hash path, slot and other information, and send it to the physical memory module to control the multi-level MUX network to select the queried point cloud frame index information to the read address bus of the secondary memory. Subsequently, the state machine jumps to S8, that is, the result output state.

[0078] In the S8 state, the state machine places the queried point cloud coordinates on the external data bus, then pulls the point cloud output valid flag high and performs a data handshake with the next level module. After the handshake is successful, if the current voxel to be queried has not completed all neighbor searches, the state machine jumps to S7 and calculates the next set of voxel indices. If the current voxel has completed all neighbor searches, the state machine checks whether the end flag of the current frame is pulled high. If so, the state machine jumps to S0 and waits for the working mode to switch or the arrival of a new frame of point cloud data. Otherwise, the state machine jumps to S1 and waits for the next point cloud data handshake in the current frame.

[0079] In this embodiment, hash buckets are independent storage units used to store data in a hash table. Each bucket corresponds to a range of hash values. A hash table is a key-value storage structure that uses a hash function to map keys to fixed-size storage areas (hash buckets), enabling efficient data insertion, querying, and deletion.

[0080] When the hash mapping unit selects a spatial hash function, it multiplies the spatial coordinates by a very large prime number to make them more evenly distributed, thus better preventing hash collisions. The hash function then performs a bitwise exclusive-or operation on the three multiplication results. Finally, a modulo-N operation is performed on the result of the exclusive-or operation, where N is the depth of the hash table. This embodiment also constructs different hash mapping paths based on the two sets of parameters to further reduce the probability of hash collisions.

[0081] The two sets of parameters are: (1) =73856093、 =19349663、 =83492791、(2) =92837111、 =689287499, =283923481.

[0082] The corresponding hash function is:

[0083] Where, Bit-space hash functions, To take the operation, It is an exclusive OR operation.

[0084] The hash map unit utilizes an overall pipeline architecture. In the first pipeline stage, DSP resources are instantiated in parallel to perform multiplication operations on the three spatial coordinates x, y, and z with their corresponding parameters, completing the calculation within a single clock cycle. The second pipeline stage instantiates a two-stage XOR unit to process the calculation results output by the DSP unit and outputs the bitwise XOR result to the modulo unit. The modulo unit receives the XOR result and the modulus N and outputs the modulo result of the two to the final pipeline register.

[0085] In order to improve the computational efficiency of the modulo operation unit, the present invention uses the Montgomery algorithm for hardware optimization. The product term in the Montgomery algorithm is considered as a whole, and the result of the hash function XOR operation is passed to the REDC function as a whole.

[0086] The Montgomery algorithm creates a mapping relationship to convert the multiplication factors under the general domain , Mapped to the Montgomery domain. It is defined as a power of 2 and is greater than the modulus N. It is usually taken as the value closest to the modulus N. In addition, the parameter And the modulus N must also be mutually exclusive. According to the extended Euclidean algorithm, it can be obtained that R and N satisfy the following relationship, and can be further Expressed as the opposite of the multiplicative inverse of N.

[0087]

[0088]

[0089]

[0090] Where, is the parameter, , is the multiplication factor in the general field, For the model, is the greatest common divisor.

[0091] In order to quickly solve the multiplication modular operation, the Montgomery modular reduction function (REDC) is defined. The specific expression is:

[0092]

[0093]

[0094] Where, is the Montgomery modular reduction function, with X as the input variable. The REDC function multiplies the input variable X by the inverse of R and then modulo N. In modular multiplication problems, X is the product of two multiplication factors in the Montgomery domain. The result of one REDC operation remains in the Montgomery domain. To convert X from the Montgomery domain back to the conventional domain, the output of the first REDC operation is used as an argument and input into the REDC function again. After a second REDC operation, the product ab modulo N in the conventional domain is obtained.

[0095] The definition of the REDC formula reveals that the calculation involves dividing X by R. Because R can be expressed as a power of 2, this calculation can be performed using shift operations. However, this calculation method erases the low-order bits of X, resulting in errors. To address the problem of X not being divisible by R, the Montgomery algorithm adds m times the modulus N to X, making it divisible by R. The following formula further simplifies this expression using the extended Euclidean algorithm to obtain the expression for the parameter m.

[0096]

[0097]

[0098]

[0099]

[0100]

[0101] After clarifying the calculation method of parameter m, the calculation process of the REDC function can be summarized into the following four steps: (1) Calculate parameters according to formula 4-9 and formula 4-17 and parameter m; (2) Calculate the value of X + mN, shift the result right by k bits, and assign it to the variable y; (3) Determine whether y is greater than modulo N. If it is greater than N, update y = yN; (4) Return the variable y as the output result of the REDC function.

[0102] In order to apply the Montgomery modular multiplication algorithm in the design of the modular unit, this embodiment regards the product term in the Montgomery algorithm as a whole, and passes the result of the XOR operation of the hash function as a whole to the REDC function.

[0103] like Figure 13 As shown, this embodiment divides the Montgomery reduction unit into a four-stage pipeline. In the pre-calculation stage, the parameters After the calculation is completed, the square of the parameter R is pre-calculated modulo N, and the calculated parameter is directly reused in the subsequent Montgomery domain conversion and reduction. The first stage of the reduction unit pipeline primarily processes the sign of the input data, converting it from two's complement format to its corresponding absolute value. The second stage uses the pre-calculated parameter Rs to convert the input data to the Montgomery domain. The third stage performs the main steps of the Montgomery reduction, including the calculation of the parameter m and the calculation of the preliminary reduction result. To improve computational efficiency and reduce hardware resource overhead, the design of this stage implements the modulo operation through bit width truncation and the division operation through shift operations. Finally, the fourth stage is responsible for adjusting the preliminary reduction result. To reduce the consumption of resources such as lookup tables, all large-bitwidth multiplication and addition operations in the pipeline are implemented using instantiated DSP units.

[0104] From the above formula, we can see that the input variable X of the REDC function must be smaller than the square of the modulus N. The storage bit width of the modulus N is 16 bits, while the storage bit width of the input variable X is 48 bits, which is much larger than the modulus N and the parameter R. If the large bit width X is directly passed to the REDC unit, it will violate the basic constraints of the Montgomery algorithm. Therefore, when taking the modulus N, multiple iterative operations are still required. Figure 14 To address this issue, the present invention proposes a block-based optimized modulo operation unit. This unit first divides the 48-bit width of X into three 16-bit blocks, then computes the REDC function for each block in parallel. After the computation is complete, the output of each REDC unit is passed to a second-level REDC unit for further computation. The results are then accumulated and combined, and then adjusted by an adjustment unit to a value within the range of 0 to N for output.

[0105]

[0106]

[0107]

[0108]

[0109] Where, , , After the blocks are split, the large bit width X can be divided into small blocks. , , , and then perform the modular expansion to its outermost layer. After expansion, we can see that the original product term is greater than the square of the modulus N. However, after splitting it again and taking the modulus N, each part is constrained to fall within the range of the modulus N, so its product term is now less than the square of the modulus N. Then, performing the modular expansion on the product term again will not violate the basic constraints of the Montgomery algorithm.

[0110] like Figure 14 As shown in the figure, the block-optimized modulo operation unit consists of two REDC units and one DSP unit. Each REDC unit level is internally composed of a four-stage pipeline. Therefore, the computational latency for performing a modulo operation on a 48-bit input is 9 clock cycles. Furthermore, this module is a fully pipelined architecture. Timing analysis shows that the circuit has a maximum operating clock frequency of 550MHz. Compared to third-party divider IP cores such as Vivado, this achieves a 3.3x increase in maximum clock frequency at similar single-time computation latency.

[0111] In addition, in order to reduce the probability of hash conflicts, the present invention also designs the following in the hash state machine: Figure 15 Three strategies are shown: One is the dual hash function: by instantiating two different hash functions, multiple optional insertion paths are provided for the same hash key.

[0112] The second is multiple slots: Instantiate multiple storage slots under the same hash bucket.

[0113] The third strategy is random ejection: If the first two strategies fail to avoid hash collisions, meaning all slots under multiple insertion paths are occupied, the hash table selects a random slot under all insertion paths, ejects the key-value pair stored there, and replaces it with the key-value pair currently being inserted. The randomly ejected hash key is returned to the top level of the hash table and inserted again.

[0114] The point cloud sorting module is used to calculate the distance between the neighborhood point cloud coordinates and the first frame point cloud, sort the distance results, and match the closest point cloud with the first frame point cloud according to the sorting results to form a neighborhood point cloud search result.

[0115] Among them, the point cloud sorting module includes a distance calculation unit based on a parallel pipeline and a sorting unit based on a linear feedback shift register.

[0116] Specifically, if Figure 16 Figure 2 shows the hardware circuit structure of the distance calculation unit. This unit utilizes a parallel pipeline architecture and can be divided into two stages. The first stage is responsible for registering the coordinates of the input reference frame target point cloud and neighboring point clouds. The second stage instantiates three DSP units to parallelly calculate the squared distances in the x, y, and z directions. Finally, an adder tree accumulates these squared distances and stores them as output.

[0117] The pipelined architecture ensures that the distance calculation unit can process a pair of point cloud coordinates in every clock cycle, increasing circuit throughput, optimizing circuit timing, and boosting the maximum clock frequency. Within a single-stage pipeline, a resource replication strategy is employed to instantiate multiple DSP units in parallel to simultaneously calculate distances in multiple dimensions. This reduces processing latency within a single-stage pipeline and further optimizes pipeline timing.

[0118] like Figure 17 The figure shows the hardware circuit structure of the sorting unit designed in this embodiment. To optimize circuit timing and maintain a fully pipelined architecture, the sorting unit hardware circuit is designed based on a linear feedback shift register. The entire sorting unit consists of k registers, k multiplexers (MUXs), and k comparators. The distance information stored in the k registers increases in order from left to right. The input distance calculation result is simultaneously fed to the k comparators and compared with the distance information in the k registers. The comparator output is transmitted as a control signal to the MUX. If the current distance calculation result is exactly between register m and register m+1, the MUX controls the first m registers to retain the original data. The m+1 register selects the current calculation result as its input. The outputs of the subsequent k registers are then connected to the inputs of the next register, forming a k-level shift register. This is equivalent to inserting the current calculation result into the queue at position m+1. From this position, the data in the queue are pushed one storage location to the right, and the last data in the queue is discarded. The sorting unit can update the sorting queue once every clock cycle. The point cloud sorting process is complete when the last distance calculation result is passed to the linear feedback shift register. The data stored in the k registers at this point is the sorted result. The point cloud sorting module ultimately matches the point cloud coordinates corresponding to these k distances with the target point cloud in the reference frame and outputs them to the next level module.

[0119] A fast point cloud sorting module was designed based on a parallel pipeline architecture and linear feedback shift registers. This pipeline architecture improves circuit throughput and maximum operating clock frequency, reducing overall point cloud sorting latency. The linear feedback shift register ensures that the sorting queue is updated once per clock cycle, accelerating the sorting process.

[0120] This embodiment designs an efficient point cloud storage and indexing architecture based on hash tables and spatial hash functions. This improves the parallelism of hash lookups, accelerating search speed. Furthermore, the Montgomery algorithm, based on block optimization, accelerates modulo operations. Furthermore, a two-level physical storage architecture reduces the consumption of redundant storage resources and improves on-chip storage efficiency.

[0121] Example 2 This embodiment provides an FPGA-based acceleration method for a lidar odometer. Based on the system implementation in Example 1, the method includes the following steps: S1. Use the PS end to obtain multi-frame lidar point clouds and read two consecutive frames of point clouds to the PL end; the PL end receives 32-bit floating-point point cloud data and obtains 16-bit fixed-point point clouds after hardware quantization.

[0122] Furthermore, the 32-bit floating point cloud data is quantized and converted into a fixed-point format for on-chip storage and calculation. Considering the scanning resolution and maximum detection distance of the laser radar, the quantized fixed-point format selected by the present invention is as follows: Figure 3 As shown, the fixed-point number 16-bit complement Indicates that the highest bit is the sign bit, the next highest 8 bits are the integer bits, and the lowest 7 bits are the decimal bits. The conversion relationship between it and the decimal system is shown in the following formula:

[0123]

[0124] Where, is a fixed-point number, is a decimal fixed-point number, The original 32-bit floating point number is converted to a decimal number. Indicates rounding down and converting to binary.

[0125] After receiving the first frame of fixed-point point cloud, S2 and PL calculate the voxel coordinates and perform hash mapping to store the first frame of point cloud in the corresponding voxels to obtain a local map of the point cloud. After the local map is constructed, the second frame of point cloud is received and the neighborhood voxel range of the second frame of point cloud is searched in the local map of the point cloud to obtain the neighborhood point cloud coordinates.

[0126] S3. Calculate the distance between the coordinates of the neighborhood point cloud and the reference point cloud of the first frame, sort the points according to the distance results, match the point cloud with the closest distance to the reference point cloud of the first frame, obtain the neighboring point cloud search results and store them; S4. The PS side sends the neighboring point cloud search results to the PL side; the PL side constructs the point cloud matching residual based on the neighboring point cloud search results, minimizes the matching residual through the Gauss-Newton method, and finally calculates the current position of the lidar.

[0127] Example 3 As a preferred embodiment of Example 1, this embodiment is based on the Xilinx Zynq UltraScale+MPSoCs platform with an operating frequency of 150 MHz, and experimental tests are performed on the system in Example 1. Figure 18 This article demonstrates the experimental process for testing a lidar odometry acceleration system using the NTU-VIRAL dataset. First, the dataset is converted from the bag format under the ROS operating system to a TXT text format. The TXT dataset is then written to an SD card, and the PS reads the data from the SD card and transfers it to the DDR4 memory mounted on the PS. Next, the DMA controller on the PL reads two consecutive frames of point cloud data and transfers them to the lidar odometry hardware acceleration system. Finally, after the lidar odometry acceleration system completes pose calculations, the PS saves the system's output pose, generates a trajectory file in TUM format, and updates the current DDR memory read pointer position. Figure 19 The results of the lidar odometry acceleration system tested on the three sequences eee, nya, and sbs in the NTU-VIRAL dataset are shown.

[0128] As shown in Table 1, in terms of operating speed, the LiDAR odometer acceleration system designed in the present invention can achieve an average processing speed of 9.6ms / frame when deployed on the Xilinx Zynq UltraScale+ MPSoCs platform, which is 2.1 times faster than the odometer system deployed on the Intel i5 processor. In terms of energy consumption, the power deployed on the AXU15EG platform is 4.525W, which is 3.4 times lower than that of the Intel i5 processor. In addition, the energy efficiency of the LiDAR odometer system running on the AXU15EG platform is significantly better than that running on the Intel i5 processor. The test results show that the LiDAR odometer acceleration system designed in the present invention can meet the low power consumption and high real-time performance requirements of small autonomous navigation drones.

[0129] Table 1

[0130] Table 2

[0131] As shown in Table 2, in terms of positioning accuracy, the analysis results show that the RMSE of the absolute error of the output pose trajectory of the acceleration system designed in the present invention is 2.87m. Its positioning accuracy is close to that of lidar odometry systems such as LOAM, LeGO-LOAM, and F-LOAM, and it fully meets the positioning accuracy requirements of smart terminal devices in simple scenarios.

[0132] Table 3

[0133] Table 3 shows the resource consumption of the implemented hardware system. The lookup table consumed 10,506 memory blocks, the flip-flops 8,278, the DSP 78, and the BRAM 203.5 (a total of 7.1 MB). This was achieved by increasing the parallelism of the hash lookup process. Modulo operations were also accelerated using the block-based optimized Montgomery algorithm. Furthermore, a two-level physical memory architecture reduced redundant memory resource consumption, improving on-chip storage efficiency.

Claims

1. An acceleration system for a laser radar odometer based on FPGA, characterized in that: The system includes the PS side and PL side for data interaction; The PS side includes a point cloud acquisition module and a pose calculation module. The point cloud acquisition module is used to acquire multi-frame lidar point clouds and send the acquired point clouds to the PL side. The lidar point cloud is a 32-bit floating point point cloud. The pose calculation module is used to obtain the neighboring point cloud search results from the PL side and calculate the current pose of the lidar based on the neighboring point cloud search results. The PL side includes a floating-point quantization module, a local map management module, and a point cloud sorting module; The floating-point quantization module is used to receive two consecutive frames of 32-bit floating-point point clouds and quantize them into 16-bit fixed-point point clouds; The local map management module is used to obtain the 16-bit fixed-point point cloud of the first frame, perform hash mapping based on the voxel coordinates of the first frame point cloud to form a local map of the point cloud, and search the neighborhood voxel range of the second frame point cloud based on the local map of the point cloud to obtain the neighborhood point cloud coordinates; The point cloud sorting module is used to calculate the distance between the neighborhood point cloud coordinates and the first frame point cloud, sort the distance results, and match the closest point cloud with the first frame point cloud according to the sorting results to form a neighborhood point cloud search result.

2. The FPGA-based laser radar odometer acceleration system according to claim 1, characterized in that: The point cloud acquisition module exchanges data with the PL end through the AXI4 bus; the posture calculation module exchanges data with the PL end through the AXI Lite bus.

3. The FPGA-based laser radar odometer acceleration system according to claim 1, characterized in that: The floating point quantization module includes an exponent adjustment submodule, a shift submodule and a complement calculation submodule; The exponent adjustment submodule is used to receive the exponent part of the 32-bit floating point cloud and calculate the exponent part according to the set rules to obtain the exponent value; The shift submodule is used to receive the decimal part of the 32-bit floating point cloud, shift the decimal part to the right according to the exponent value obtained by the exponent adjustment submodule, and truncate the high bits of the shifted data according to the set bit width to obtain tmp_dat; The complement calculation submodule is used to calculate the complement of tmp_dat to obtain a quantized 16-bit fixed-point point cloud.

4. The FPGA-based laser radar odometer acceleration system according to claim 1, characterized in that: The local map management module includes a hash controller and a physical memory; The hash controller is used to obtain two consecutive frames of 16-bit fixed-point point clouds, calculate voxel coordinates based on the first frame of fixed-point point cloud, perform hash mapping based on the voxel coordinates, store the first frame of point cloud in the voxels to form a point cloud local map, search the neighborhood voxel range of the second frame of point cloud based on the point cloud local map to obtain neighborhood point cloud coordinates, generate control signals and physical storage addresses based on the neighborhood point cloud coordinates, and perform data interaction with the physical memory based on the physical storage addresses; The physical memory is used to store and / or read the point cloud local map and the corresponding point cloud under the control of the hash controller.

5. The FPGA-based laser radar odometer acceleration system according to claim 4, characterized in that: The hash controller includes a point cloud coordinate conversion unit, a hash state machine, a hash mapping unit, and a distance comparison unit; The point cloud coordinate conversion unit is used to calculate the spatial voxels corresponding to the point cloud according to the point cloud coordinates, and convert the point cloud coordinates into the index of the spatial voxels; The hash state machine is used to read the point cloud voxel index after coordinate transformation, and then control the hash mapping unit to complete the hash function calculation. It also controls the physical memory module through the external RAM bus to complete the storage and reading of point cloud and voxel information in the local map. The hash mapping unit includes a two-stage pipeline structure. The first stage pipeline performs XOR calculation on the hash function through parallel DSP to obtain the XOR operation result; The second stage pipeline includes a modulo operation unit, which is used to perform a modulo calculation on the XOR operation result to obtain a reduction result; The distance comparison unit is used to perform distance comparison between the neighboring point cloud and the first frame point cloud according to the reduction result to obtain a comparison result.

6. The FPGA-based laser radar odometer acceleration system according to claim 5, characterized in that: The working states of the hash state machine include insert mode and search mode; When the working state is in insert mode: the hash state machine is in the initial state, waiting for a valid start signal to be input, and then jumps to the data handshake state; In the data handshake state, the hash state machine pulls the ready signal in the current module high and waits for the point cloud coordinate conversion unit to output a valid point cloud voxel index. When the ready signal and the valid signal are detected to be high at the same time in this cycle, the hash state machine splices the point cloud index data corresponding to the three direction axes into a hash key, and in the next clock cycle, the hash state machine jumps to the hash mapping state. In the hash mapping state, the hash state machine starts hash function calculation; after the calculation is completed, it jumps to the hash search state; In the hash search state, if the same hash key as the point cloud to be inserted is found in all slots of the current hash bucket, a synchronous downsampling operation is performed on the voxels, and the hash state machine jumps to the distance comparison state. In the distance comparison state, the queried point cloud and the point cloud to be inserted are passed to the distance comparison unit. The current state of the hash state machine is switched according to the result obtained by the distance comparison unit. If the same hash key as the point cloud to be inserted is not found in all slots of the current hash bucket, the hash state machine jumps to the hash insertion state. When the working state is in search mode: when the hash state machine is in the data handshake state, it waits for the next frame of valid point cloud handshake signal, and jumps to the voxel index calculation state after the handshake is successful; In the voxel index calculation state, the hash state machine groups the neighboring voxels and calculates the index of the neighboring voxels in each group; After the two neighbor voxel indices in the current group are calculated, the state machine jumps to the hash mapping state and passes the calculation result of the hash key to the hash mapping unit; In the hash mapping state, the hash state machine waits for the hash mapping calculation to be completed, and after the completion flag is pulled high, it jumps to the hash search state; In the hash search state, the hash state machine reads the corresponding hash key value in the physical memory; by comparing the hash key value of each slot in the current hash bucket with the hash key to be queried one by one, it determines whether there is a point cloud inserted in the current voxel to be queried, and locates the hash path, slot and point cloud frame index of the inserted point cloud; The hash state machine generates the corresponding strobe control signal based on the queried hash path, slot and other information, and sends it to the physical memory to control the multi-level MUX network, strobe the queried point cloud frame index information to the read address bus of the physical memory, and the hash state machine jumps to the result output state; In the result output state, the hash state machine will pull up the point cloud output valid flag corresponding to the queried point cloud coordinates and perform data handshake with the point cloud sorting module; After the handshake is successful, if the current voxel to be queried has not completed all neighbor searches, the hash state machine jumps to the voxel index calculation state and performs the next set of voxel index calculations; if the current voxel has completed all neighbor searches, the state machine will detect whether the end flag of the current frame is pulled high. If it is pulled high, the state machine jumps to the initial state.

7. The FPGA-based laser radar odometer acceleration system according to claim 5, characterized in that: The hash map unit is optimized using the Montgomery algorithm.

8. The FPGA-based laser radar odometer acceleration system according to claim 6, characterized in that: The hash bucket in the hash state machine includes multiple storage slots; When all storage slots in the hash bucket are occupied, a random slot under all insertion paths is selected, the hash key-value pair stored in the random slot is kicked out, and replaced with the hash key-value pair currently to be inserted; the hash key-value pair randomly kicked out is re-inserted.

9. The FPGA-based laser radar odometer acceleration system according to claim 1, characterized in that: The point cloud sorting module includes a distance calculation unit based on a parallel pipeline and a sorting unit based on a linear feedback shift register; The distance calculation unit consists of two pipelines. The first pipeline is used to store the input reference frame target point cloud and neighboring point cloud coordinates. The second pipeline is used to parallelly calculate the squared distances in different directions and accumulate the squared distances through an addition tree before storing the output. The sorting unit includes several registers, several multiplexers and several comparators; the several registers respectively store the distance calculation results output by the distance calculation unit, the comparators respectively compare the distance calculation results, and after obtaining the comparison results, they are transmitted to the multiplexers as control signals; and are used to transmit to the linear feedback shift register through the multiplexers.

10. An acceleration method for a laser radar odometer based on FPGA, implemented based on the acceleration system for the laser radar odometer based on FPGA according to any one of claims 1 to 9, characterized in that: include: S1. Use the PS end to obtain multi-frame lidar point cloud and read two consecutive frames of point cloud to the PL end; The PL side receives 32-bit floating-point point cloud data and obtains 16-bit fixed-point point cloud after hardware quantization; After receiving the first frame of fixed-point point cloud, S2 and PL calculate the voxel coordinates and perform hash mapping to store the first frame of point cloud in the corresponding voxels to obtain a local point cloud map. After the local map is constructed, the second frame of point cloud is received and the neighborhood voxel range of the second frame of point cloud is searched in the local point cloud map to obtain the neighborhood point cloud coordinates. S3. Calculate the distance between the coordinates of the neighborhood point cloud and the reference point cloud of the first frame, sort the points according to the distance results, match the point cloud with the closest distance to the reference point cloud of the first frame, obtain the neighboring point cloud search results and store them; S4. The PS sends the neighboring point cloud search results to the PL. The PL side constructs the point cloud matching residual based on the point neighboring point cloud search results, minimizes the matching residual through the Gauss-Newton method, and finally calculates the current position of the lidar.