Hardware Acceleration Positioning System and Robot under the Collaboration of Hardware and Software

Through a hardware accelerated positioning system that combines software and hardware, the correlation scanning matching hardware circuits are used to process low-resolution maps and CPUs in parallel to process high-resolution maps, which solves the real-time positioning problem of robot SLAM systems in complex environments, and improves the computing efficiency and frame rate of loopback detection.

CN115984347BActive Publication Date: 2025-07-29AMICRO SEMICONDUCTOR CO LTD +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202111202214.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-10-15
Publication Date
2025-07-29
Estimated Expiration
2041-10-15

AI Technical Summary

Technical Problem

In the prior art, the real-time nature of the mobile robot SLAM system based on lidar is difficult to meet the needs of loop detection in complex environments, especially the requirements of robots for instant positioning speed in public places, and the calculation speed of traditional branch bounding algorithms is difficult to meet the needs of complex map scenarios.

Method used

Using a hardware accelerated positioning system with hardware and software collaboration, the low-resolution multi-layer map is processed in parallel through correlation scanning and matching hardware circuits, and combined with CPU processing of high-resolution multi-layer maps, a pipeline parallel architecture and interrupt control of the branch bounding algorithm are designed to reasonably allocate computing resources and improve the frame rate of the branch bounding algorithm.

Benefits of technology

It improves the real-time and computing efficiency of loop detection, meets the positioning needs of robots in complex environments, simplifies hardware system design, and saves hardware resources.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115984347B_ABST
    Figure CN115984347B_ABST
Patent Text Reader

Abstract

The present invention discloses a hardware-accelerated positioning system and a robot under the cooperation of software and hardware. The hardware-accelerated positioning system includes a CPU and a correlation scanning and matching hardware circuit; the correlation scanning and matching hardware circuit is used to search each layer of the first map to be searched in a hardware parallel processing manner and update the first upper bound value of the currently searched node; the CPU is used to, whenever the correlation scanning and matching hardware circuit searches for a node corresponding to a grid point of a map at a preset depth, search each layer of the second map to be searched by executing the branch and bound algorithm, and perform scanning and matching of the point cloud collected by the lidar on the corresponding layer of the second map to be searched by executing the correlation scanning and matching algorithm, and then calculate and update the second upper bound value of the currently searched node by using the matched result; the CPU is further used to configure the pose corresponding to the latest obtained second upper bound value or the pose corresponding to the latest obtained first upper bound value as the pose of the robot.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of SLAM mapping and positioning, and relates to a hardware-accelerated positioning system and a robot under the cooperation of software and hardware. Background Art

[0002] The core problem of the SLAM solution for mobile robots based on lidar is to use the correlation scan matching algorithm to achieve accurate positioning of the robot. The correlation scan matching (CSM) algorithm is one of the most widely used scan matching algorithms. This method often optimizes based on a prior estimate. Within a search window with limited size, the lidar point cloud at the current moment is aligned with the map through rotation and translation transformations.

[0003] The positioning part of the classic lidar SLAM framework often consists of three parts: front-end real-time scan matching, back-end non-linear optimization, and loop detection. Among them, the front-end first obtains the initial pose from the odometer, and then obtains the pose after laser constraint through continuous frame matching of lidar data. In the case of a small window, the calculation amount is small and the algorithm frame rate is high. The back-end uses loop detection to construct a closed-loop constraint, and obtains the optimized pose by minimizing the observation and estimation residuals. Errors introduced by the estimation will be generated in the front-end. As the movement range expands, this error will gradually accumulate, resulting in incorrect results in the end. In order to reduce the cumulative error of robot positioning, loop detection often performs scan matching based on a search window with a large size, adding a closed-loop constraint with strong constraints. The overall computational complexity of the algorithm is related to the square of the search window size, and the computational amount of loop detection is often thousands of times that of the front-end real-time scan matching.

[0004] Therefore, the real-time performance of loop detection is extremely poor. The branch-and-bound algorithm adopted by the prior art can only narrow the search space in a limited map scenario, such as in a grid map with a specific resolution, and improve the search speed of the algorithm to a certain extent. However, when facing a more complex map scenario, the calculation speed is difficult to meet the navigation and positioning requirements of the robot, especially the demand for instant positioning speed of service robots performing delivery and other services in public places. Summary of the Invention

[0005] To solve the above technical deficiencies, the present invention discloses a hardware-accelerated positioning system and a robot under software-hardware cooperation. From the perspective of hardware acceleration, a branch and bound algorithm under software-hardware cooperation is designed. Specifically, a dedicated acceleration circuit with parallel pipelining is used to process multi-layer maps with lower resolution, and the CPU is called to process multi-layer maps with higher resolution, improving the calculation efficiency of branch and bound and meeting the real-time requirements. The main contribution lies in designing a pipelined parallel architecture and interrupt control for calculating score values, coordinating the resources of the CPU and the dedicated acceleration circuit with parallel pipelining, so as to share the calculation amount of the traditional branch and bound algorithm and improve the frame rate of the branch and bound algorithm. The specific technical solutions are as follows:

[0006] A hardware-accelerated positioning system under software-hardware cooperation, where the hardware-accelerated positioning system has an electrical connection relationship with a lidar; among them, the lidar is installed on the body of the robot; the hardware-accelerated positioning system includes a CPU and a correlation scanning and matching hardware circuit; the correlation scanning and matching hardware circuit is used to search each layer of the first map to be searched in a hardware parallel processing manner and update the first upper bound value of the currently searched node; the CPU is used to, whenever the correlation scanning and matching hardware circuit searches for a node corresponding to a grid point of a map layer at a preset depth, search each layer of the second map to be searched by executing the branch and bound algorithm, and perform scanning and matching of the point cloud collected by the lidar on the corresponding layer of the second map to be searched by executing the correlation scanning and matching algorithm, and then calculate and update the second upper bound value of the currently searched node using the matching result; the CPU is further used to configure the pose corresponding to the latest obtained second upper bound value or the pose corresponding to the latest obtained first upper bound value as the pose of the robot; among them, more than one branch and bound search tree is constructed for all layers of the first map to be searched and all layers of the second map to be searched, so that the grid points corresponding to the root nodes of all branch and bound search trees are in the first map to be searched with the lowest resolution; the resolution of each layer of the first map to be searched is lower than the resolution of any layer of the second map to be searched; the first map to be searched and the second map to be searched are set in layers.

[0007] The hardware-accelerated positioning system disclosed in this technical solution performs branch and bound processing on multi-layer maps through the method of software and hardware module scheduling, which is specifically used to handle the real-time problem of loop detection. This hardware-accelerated positioning system hands over the search process of multi-layer maps with relatively low resolution to the correlation scanning and matching hardware circuit to be realized in a hardware parallel processing manner, and hands over the search process of multi-layer maps with relatively high resolution to the CPU to be realized, thereby reasonably allocating computing resources within the same system, maximizing the computing efficiency of executing the algorithm in a pure software manner, and saving some hardware resources.

[0008] Further, the hardware acceleration positioning system stores a pre-constructed grid map with a first resolution. The grid map with the first resolution is constructed by the CPU into a multi-level grid map, which includes a first map to be searched with a first preset number of layers and a second map to be searched with a second preset number of layers. The maximum resolution in all layers of the second map to be searched is equal to the first resolution. In this multi-level grid map, the first map to be searched and the second map to be searched are arranged in layers; the first map to be searched with the first preset number of layers and the second map to be searched with the second preset number of layers are vertically arranged in ascending order of resolution. Among them, the sequence numbers of the first map to be searched with the first preset number of layers and the second map to be searched with the second preset number of layers both represent the depth of the corresponding layer of the map.

[0009] This technical solution generates a low-resolution map based on the highest-resolution map to form a multi-level grid map arranged from top to bottom, so as to facilitate the construction of a tree structure and reduce the probability of iteration falling into local extrema. Among them, the higher the layer of the grid map, the lower the resolution and the earlier the grid map of the higher layer is transmitted to the hardware acceleration positioning system. The lower the layer of the grid map, the higher the resolution, realizing the hierarchical fuzzy processing of the map, making the map at low resolution store the maximum occupancy probability value of adjacent grids, ensuring that the lower the resolution during the regional search of the point cloud, the higher the accumulated score value. Therefore, during the subsequent branch and bound process, for the pose with the highest score value, it is necessary to continue the branch search in a map with a higher resolution layer.

[0010] Furthermore, the hardware acceleration positioning system further includes a bus interface, and both the CPU and the correlation scanning and matching hardware circuit are electrically connected to the bus interface; the CPU is configured to execute the following steps: Step A, the CPU configures, through the bus interface, parameters required for executing the branch and bound algorithm and parameters required for executing the correlation scanning and matching algorithm for the correlation scanning and matching hardware circuit; among them, the parameters required for executing the branch and bound algorithm include the first preset number of layers, the resolution of the first search map to be searched for each layer, the second preset number of layers, the resolution of the second search map to be searched for each layer, and the number of root nodes; Step B, the CPU transmits, through the bus interface, the first search map to be searched for with the lowest resolution for one layer to the correlation scanning and matching hardware circuit; Step C, the CPU starts the correlation scanning and matching hardware circuit to search for all root nodes, and triggers the correlation scanning and matching hardware circuit to calculate the score values of the corresponding root nodes; at the same time, the CPU discretizes the point cloud according to the parameters required for executing the correlation scanning and matching algorithm; Step D, the CPU transmits, through the bus interface, the remaining first search maps to be searched for to the correlation scanning and matching hardware circuit; at the same time, the CPU sorts all the root nodes currently searched by the correlation scanning and matching hardware circuit in ascending order of the score values, and then stores the root nodes in the stack according to the current sorting; wherein, the stack is a cache space opened inside the hardware acceleration positioning system that supports first in last out; the depth of the first search map to be searched for each layer and the depth of the second search map to be searched for each layer are respectively equal to the depths of the nodes at the same vertical height of the branch and bound search tree.

[0011] In this technical solution, the parameters of the loop detection algorithm (parameters required for executing the branch and bound algorithm and parameters required for executing the correlation scanning and matching algorithm) are configured into the parameter registers of the correlation scanning and matching hardware circuit to maintain the normal execution of the branch operation and the correlation scanning and matching by the correlation scanning and matching hardware circuit; the correlation scanning and matching hardware circuit searches for the root node corresponding to the pose of the first search map to be searched for with the lowest resolution to start searching for child nodes in the branch of the first search map to be searched for with the highest resolution, and at the same time schedules the CPU to discretize the point cloud as the standby coordinate points for aligning and matching to the discrete grid points of the second search map; moreover, when scheduling the CPU to receive the first search map with a higher resolution by the correlation scanning and matching hardware circuit, the score values of the root nodes already calculated by the correlation scanning and matching hardware circuit are sorted to obtain the maximum score value; improving the search efficiency of the hardware acceleration positioning system.

[0012] On the other hand, the CPU's point cloud discretization operation is parallel to the CPU's starting the correlation scanning and matching hardware circuit to search for root nodes, realizing software and hardware cooperation.

[0013] Further, the CPU is further configured to perform the following steps: Step E, the CPU controls the node at the top of the stack to pop out of the stack, and then the CPU determines whether the currently popped node is a leaf node. If so, it proceeds to Step F; otherwise, it proceeds to Step G. Step F, the CPU determines whether the score value of the currently popped node is greater than the current optimal upper bound value. If so, it updates the score value of the node to the current optimal upper bound value, and determines the updated current optimal upper bound value as the first upper bound value on the corresponding branch of the currently searched node or the second upper bound value on the corresponding branch of the currently searched node, and then proceeds to Step H; otherwise, it directly proceeds to Step H, where the initial value of the current optimal upper bound value is a preset initial threshold. Step G, when the CPU determines that the score value of the currently popped node is greater than the current optimal upper bound value, the CPU determines whether the absolute value of the difference between the depth of the currently popped node and the depth of the first map to be searched with the lowest resolution in a layer is greater than or equal to the reference depth difference. If so, it proceeds to Step I, and determines that the correlation scan matching hardware circuit searches for a node located on the first map to be searched with the highest resolution; otherwise, it proceeds to Step J. The reference depth difference is equal to the difference between the first preset number of layers and the value 1. When the CPU determines that the score value of the currently popped node is less than or equal to the current optimal upper bound value, no branch operation is performed, and it is determined that a pruning operation is completed, and then Step H is executed. Step I, the CPU performs a branch operation on the node described in Step G, searches for all child nodes of the node described in Step G, and then calculates the score values of the currently searched child nodes, which are used to be updated to the second upper bound value when they are greater than the current optimal upper bound value and become the popped node in Step E. Then it proceeds to Step K. The branch operation in Step I is to search for all child nodes of the node described in Step G in the adjacent second map to be searched with a higher resolution than the layer where the node described in Step G is located. Step J, the correlation scan matching hardware circuit performs a branch operation in a hardware parallel processing manner, searches for all child nodes of the node described in Step G, and calculates the score values of the currently searched child nodes, which are used to be updated to the first upper bound value when they are greater than the current optimal upper bound value and become the popped node in Step E. Then it proceeds to Step K. The branch operation in Step J is to search for all child nodes of the node described in Step G in the adjacent first map to be searched with a higher resolution than the layer where the node described in Step G is located. Step K, the CPU sorts the child nodes searched in Step I or the child nodes searched in Step J in the order described in Step D, and then controls the currently sorted child nodes to be pushed onto the stack in the order described in Step D. Then it returns to Step E. Step H, the CPU determines whether the stack is empty. If so, it configures the node corresponding to the current optimal upper bound value described in Step F as the position of the robot; otherwise, it returns to Step E to follow the principle of depth-first search;Among them, each of the root nodes corresponds to a branch and bound search tree. When searching a node of each branch and bound search tree, a branching operation included in the branch and bound algorithm is adopted; the depth of the first map to be searched in each layer and the depth of the second map to be searched in each layer are respectively equivalent to the depth of the corresponding node on the branch and bound search tree.

[0014] Based on the above technical solution, in a software and hardware co - operating system, according to the depth where the current branch is located, this technical solution allocates the boundary value calculation to the correlation scanning and matching hardware circuit or the CPU. Specifically, first, the first map to be searched with the largest amount of computation and lower resolution in the first preset number of layers is transmitted to the correlation scanning and matching hardware circuit for score value calculation, and then the second map to be searched with a smaller amount of computation in the second preset number of layers is left in the CPU for score value calculation. Thus, the computation amount of the root node and most of the upper bound calculations is allocated to the hardware for execution, and the CPU is allowed to search the tree according to a specific rule to reach the optimal score value as early as possible without traversing the entire tree, realizing partial branching, pruning branches, sorting, and partial upper bound calculation by vertically searching maps of each layer, thereby narrowing the search scope. Therefore, this technical solution reasonably schedules the computing resources of the CPU and the hardware resources of the correlation scanning and matching hardware circuit, greatly simplifying the design of the hardware system.

[0015] Further, in step E, the top - of - stack node of the stack is the first node or the second node. The first node is configured with the coordinates of the grid point of the first map to be searched and the depth information of the first map to which the grid point belongs, and the second node is configured with the coordinates of the grid point of the second map to be searched and the depth information of the second map to which the grid point belongs. Whenever the CPU determines in step F that the score value of the currently popped - out first node is greater than the current optimal upper bound value, the score value of this first node is updated to the current optimal upper bound value, and it is determined that the score value of this first node is the first upper bound value of the currently searched node. Whenever the CPU determines in step F that the score value of the currently popped - out second node is greater than the current optimal upper bound value, the score value of this second node is updated to the current optimal upper bound value, and it is determined that the score value of this second node is the second upper bound value of the currently searched node. This technical solution classifies and sorts the candidate score values obtained from each branching operation, and the last - in - first - out storage mechanism of the stack selects the child node with the largest score value of the branch on the current - layer map, accelerating the search for the optimal node representing the robot's position.

[0016] Furthermore, during the process of the correlation scanning and matching hardware circuit calculating the score value of the root node in step C, if the correlation scanning and matching hardware circuit triggers an interruption, the correlation scanning and matching hardware circuit will instead enter step D; during the process of the CPU discretizing the point cloud in the second map to be searched in step C, if it detects that an interruption is triggered, the CPU will instead execute step D. The interruption mechanism is adopted to improve the processing efficiency of the branch and bound algorithm, avoiding missing hundreds of thousands of instruction cycles due to the CPU pausing and waiting for the correlation scanning and matching hardware circuit to calculate the score value, because the time for searching the root node is relatively long, longer than the interruption time.

[0017] During the process of the correlation scanning and matching hardware circuit performing a branch operation on the remaining first map to be searched, the CPU adopts a polling mechanism to wait for the correlation scanning and matching hardware circuit to calculate and output the score value of the corresponding node, so as to avoid delays caused by adopting the interruption mechanism, where the currently searched node is not the root node. Frequent interruptions of the CPU with a low frequency will cause a large amount of delay problems, because the scanning and matching time of the correlation scanning and matching hardware circuit for non-root nodes is relatively short.

[0018] Furthermore, the correlation scanning and matching hardware circuit includes a memory module, a point cloud processing module, a node search module, a state machine control module, and an interconnect bus. Among them, the memory module, the point cloud processing module, the node search module, and the state machine control module establish data transmission relationships through the interconnect bus; the point cloud processing module is used to read the currently stored point cloud from the memory module under the control of the state machine control module, and then control the read point cloud to perform a rotation transformation; the point cloud processing module is used to read the result of the previous rotation transformation under the control of the state machine control module, then control the result of the previous rotation transformation to perform a coordinate system transformation, and then set the result of the coordinate system transformation as a discrete point cloud, so as to align the currently stored point cloud into the coordinate system of the pre-expanded grid map, and at the same time determine the discretization of the currently stored point cloud; then store the discrete point cloud into the memory module through the interconnect bus; where the pre-expanded grid map is a first layer of the map to be searched required for searching the root node in step C, or a first layer of the map to be searched required for searching the child node in step J; the node search module is used to obtain the discrete point cloud from the memory module under the control of the state machine control module, and then calculate the index value of the discrete point cloud mapped to the map storage space according to the preset coordinate offset value and search step, and set the index value as the read address of the occupancy probability of the point cloud; the node search module is also used to read the occupancy probability of the corresponding point cloud from the map storage space according to the currently set read address under the control of the state machine control module, and call the accumulator built in the node search module to accumulate the occupancy probability of the currently read point cloud, and then output the accumulated result and set it as the score value of the corresponding node in the pre-expanded grid map; the memory module is provided with a map storage space for storing the first map to be searched transmitted by the bus interface; the memory module is used to store the preset trigonometric function values; where the number of rotation transformations is equal to the number of the aforementioned root nodes, and one or more coordinate system transformations are performed corresponding to one rotation transformation; the currently searched node in the search window is determined by a specific rotation parameter and a specific translation parameter; where the occupancy probability of the point cloud is the probability value of the grid point matched by the point cloud in the pre-expanded grid map after being processed by the point cloud processing module; there is a corresponding index value for the occupancy probability of the point cloud in each layer of the first map to be searched.

[0019] The calculation of the score value of the node in this technical solution is split into: performing coordinate system transformation on the read point cloud according to the rotation angle, initial position, and coordinate offset, converting the discrete point cloud coordinates into index values in the map storage space, reading the occupancy probability of the point cloud in the map storage space by the index value, and cyclically accumulating the occupancy probability to obtain the score value of the node searched under the corresponding rotation transformation and coordinate system transformation, so that the calculation of the score value is obtained by relying on the parameters required by the correlation scanning matching algorithm.

[0020] On the other hand, the point cloud processing module is preferably designed as a pipeline structure for coordinate system transformation, and the node search module is preferably designed as a pipeline structure for indexing and calculating score values, realizing the parallel execution of the process of scheduling the point cloud processing module to read the latest point cloud and the process of the point cloud processing module performing coordinate system transformation on the original point cloud at a specific clock cycle, or scheduling the coordinate system transformation of the point cloud processing module to be parallel with the index value calculation of the node search module, or scheduling the index value calculation of the node search module to be parallel with the cumulative calculation of the occupancy probability of the node search module. It also ensures that the score value result is shared on the interconnection bus for the CPU to read and use for sorting.

[0021] Further, the map storage space includes a first map storage space and a second map storage space; the first map storage space is used to store the first to-be-searched map with the first prefabricated resolution, and the second map storage space is used to store the first to-be-searched map with the second prefabricated resolution. The depth level of the first to-be-searched map with the first prefabricated resolution is one level different from the depth level of the first to-be-searched map with the second prefabricated resolution, so that the two layers of the first to-be-searched maps are adjacent; the first prefabricated resolution is less than the second prefabricated resolution; the correlation scanning matching hardware circuit is used to schedule the point cloud processing module to align and match the currently stored point cloud to the first to-be-searched map with the second prefabricated resolution while performing a branch operation once, and also schedule the node search module to calculate the index value of the occupancy probability of the currently stored point cloud in the first map storage space. Thus, by utilizing the parallel characteristics of the point cloud processing module and the node search module, searching for one layer of the first to-be-searched map with the first resolution and converting the point cloud to one layer of the first to-be-searched map with the second resolution in the same clock cycle, the advantage of the parallel pipeline of the correlation scanning matching hardware circuit is exerted, and the calculation process of the score value of the currently searched node is accelerated.

[0022] Further, the first preset layer number is 3, and the second preset layer number is 4. During the process of the CPU executing step B, the first first search map of the first layer transmitted by the CPU to the correlation scanning and matching hardware circuit through the bus interface is then written into the first map storage space by the bus interface. When the CPU executes step D, the CPU first transmits the second first search map of the second layer to the correlation scanning and matching hardware circuit, and then the CPU transmits the third first search map of the third layer to the correlation scanning and matching hardware circuit. Then, the second first search map of the second layer is written into the first map storage space by the bus interface, overwriting the first first search map of the first layer; and then the third first search map of the third layer is written into the second map storage space by the bus interface. Among them, before completely overwriting the first first search map of the first layer, the score values of all root nodes have been obtained.

[0023] This technical solution constructs three first search maps of the first layer for the correlation scanning and matching hardware circuit to search, and leaves four second search maps with higher resolution for the CPU to process. However, this technical solution does not require three map storage spaces to be specifically designed in the memory module of the correlation scanning and matching hardware circuit, but only two memories are designed, and the three first search maps of the first layer are alternately stored under the scheduling of the state machine control module, so that after the calculation of the score values of all root nodes is completed, the second first search map of the second layer overwrites the first first search map of the first layer, and then the third first search map of the third layer overwrites the second first search map of the second layer. This saves the consumption of memory resources.

[0024] Further, the node search module includes a selector and an accumulator; the accumulator is configured to accumulate the occupancy probabilities of the corresponding point clouds from the map storage space under the control of the state machine control module, wherein the occupancy probabilities of the corresponding point clouds in the map storage space are first transmitted to the interconnection bus according to the address corresponding to the index value under the control of the state machine control module, and then sequentially transmitted to the input end of the accumulator by the interconnection bus; the input end of the selector is connected to the output end of the accumulator; the selector is configured to select and output the accumulation result of the accumulator every other preset counting period, and set the currently output accumulation result as the score value of a corresponding node in the first map to be searched for a corresponding layer; wherein, every time the point cloud processing module completes a coordinate system transformation in the correlation scanning and matching algorithm of the map, the selector is triggered to output the currently accumulated accumulation result of the accumulator to the interconnection bus, and configure the currently obtained accumulation result as the score value of a new node corresponding to the search in the search window. This technical solution uses the accumulator to accumulate the occupancy probabilities corresponding to the index values obtained by hardware scanning and matching, so that the corresponding accumulation result is output as the score value at the pose of the first map to be searched for a corresponding layer every other specific counting period, that is, the score value of the currently searched node. Among them, the result accumulated under each coordinate system transformation is used as the score value of a node, which is equivalent to the accumulated result of the occupancy probability indexed after the point cloud is matched to a corresponding layer of the map, that is, the sum value of the occupancy probabilities, which is equivalent to the sum value of the positioning probabilities of a node corresponding to the search in the search window.

[0025] Further, the node search module further includes a grid index sub-module, and the grid index sub-module includes a first index subtractor, a second index subtractor, a first index adder, and a second index adder, which are configured to be in a pipeline structure; wherein, the first index subtractor and the second index subtractor both belong to subtractors, and the first index adder and the second index adder both belong to adders; the first input end of the first index subtractor is used to receive the abscissa of the discrete point cloud transmitted by the interconnection bus, wherein the abscissa of the discrete point cloud transmitted by the interconnection bus to the first index subtractor is sourced from the memory module; the second input end of the first index subtractor is used to receive the horizontal axis coordinate offset value transmitted by the interconnection bus, wherein the preset coordinate offset value includes the horizontal axis coordinate offset value, and the horizontal axis coordinate offset value transmitted by the interconnection bus is sourced from the preset coordinate offset value stored in the map offset value register; the map offset value register is a parameter register set inside the correlation scan matching hardware circuit, and is used to store the coordinate offset value associated with the pre-expanded grid map; the first input end of the first index adder is connected to the output end of the first index subtractor; the second input end of the first index adder is used to receive the horizontal axis coordinate search step transmitted by the interconnection bus; wherein, the search step includes the horizontal axis coordinate search step, and the horizontal axis coordinate search step transmitted by the interconnection bus is sourced from the search step stored in the search window parameter register; the search window parameter register is a parameter register set inside the correlation scan matching hardware circuit, and is used to store the positions of the nodes existing in the search window and the associated sub-node search information, including the search step; the sum value output by the first index adder is configured as the horizontal axis direction index value of the discrete point cloud mapped to the map storage space; the first input end of the second index subtractor is used to receive the ordinate of the discrete point cloud transmitted by the interconnection bus, wherein the ordinate of the discrete point cloud transmitted by the interconnection bus to the second index subtractor is sourced from the memory module; the second input end of the second index subtractor is used to receive the vertical axis coordinate offset value transmitted by the interconnection bus, wherein the preset coordinate offset value further includes the vertical axis coordinate offset value, and the vertical axis coordinate offset value transmitted by the interconnection bus is also sourced from the preset coordinate offset value stored in the map offset value register; the first input end of the second index adder is connected to the output end of the second index subtractor; the second input end of the second index adder is used to receive the vertical axis coordinate search step transmitted by the interconnection bus; wherein, the search step further includes the vertical axis coordinate search step, and the vertical axis coordinate search step transmitted by the interconnection bus is also sourced from the search step stored in the search window parameter register; the sum value output by the second index adder is configured as the vertical axis direction index value of the discrete point cloud mapped to the map storage space.

[0026] This technical solution uses two parallel addition and subtraction operation combination structures to calculate the horizontal-axis direction index value and the vertical-axis direction index value mapped into the map storage space respectively, so as to realize the parallel conversion of the abscissa and ordinate of the discrete point cloud into the discrete index information of a layer of the map to be searched currently participating in the search.

[0027] Further, the grid index sub-module further includes a third index adder and an index multiplier. The third index adder belongs to the adder, and the index multiplier belongs to the multiplier. The first input end of the index multiplier is connected to the output end of the second index adder. The second input end of the index multiplier is used to receive the number of row grids transmitted by the interconnection bus, where the number of row grids is the number of grids that each row of the map storage space can occupy, stored in the map size register and transmitted from the map size register to the interconnection bus. The map size register is a parameter register set inside the correlation scan matching hardware circuit, used to store the size range of the first map to be searched and the associated extended information. The first input end of the third index adder is connected to the output end of the first index adder, and the second input end of the third index adder is connected to the output end of the index multiplier. The third index adder is used to control the sum of the product of the number of row grids and the sum value output by the second index adder and the horizontal-axis direction index value to be added, and then set the sum value obtained by the addition as the index value of the discrete point cloud mapped into the map storage space, and send it to the interconnection bus, so as to read out the occupancy probability matching the index value from the memory module.

[0028] In this technical solution, the index multiplier controls the multiplication of the number of row grids and the vertical-axis direction index value, and then the third index adder controls the addition of the horizontal-axis direction index value and the product output by the index multiplier. The index value of the discrete point cloud mapped to the occupancy probability is calculated in the form of row scanning query, which is also used as the index value mapped into the map storage space, equivalent to the storage address of the occupancy probability of the point cloud in the map storage space, equivalent to the read address for externally reading the occupancy probability stored in the map storage space, so as to facilitate the interconnection bus to read out the occupancy probability from the memory module.

[0029] Furthermore, the grid index sub-module further includes a third index adder and an index multiplier. The third index adder belongs to the adder, and the index multiplier belongs to the multiplier. The first input terminal of the index multiplier is connected to the output terminal of the first index adder. The second input terminal of the index multiplier is used to receive the number of column grids transmitted by the interconnection bus. The number of column grids is the number of grids that each column of the map storage space can occupy, which is stored in the map size register and transmitted from the map size register to the interconnection bus. The map size register is a parameter register set inside the correlation scan matching hardware circuit and is used to store the size range of the first map to be searched and associated extended information. The first input terminal of the third index adder is connected to the output terminal of the second index adder, and the second input terminal of the third index adder is connected to the output terminal of the index multiplier. The third index adder is used to control the sum of the product of the number of column grids and the output of the first index adder and the index value in the vertical axis direction to be added, and then set the sum obtained by the addition as the index value of the discrete point cloud mapped into the map storage space, and send it to the interconnection bus, so as to facilitate reading the occupancy probability matching the index value from the memory module.

[0030] In this technical solution, the index multiplier controls the multiplication of the number of column grids and the index value in the horizontal axis direction, and then the third index adder controls the addition of the index value in the vertical axis direction and the product output by the index multiplier. The index value of the discrete point cloud mapped to the occupancy probability is calculated by column scanning, which is also used as the index value mapped to the map storage space, equivalent to the storage address of the occupancy probability of the point cloud in the map storage space, and equivalent to the read address for externally reading the occupancy probability stored in the map storage space, so as to facilitate the interconnection bus to read the occupancy probability from the memory module.

[0031] Further, the multi-level grid map is obtained through expansion processing, so that the first map to be searched is the pre-expanded grid map and adapts to the bus bit width of the bus interface; the expansion processing is specifically as follows: based on an original map of a specific resolution at one layer, successively perform expansion transformations according to the expansion parameters required for the maximum detection radius reached by the currently stored point cloud, the expansion parameters required for performing the coordinate transformation, the expansion parameters required for performing the rotation transformation, and the expansion parameters required for aligning the bus bit width, to obtain the corresponding layer of map in the multi-level grid map; among them, the parameters required for performing the correlation scan matching algorithm include the expansion parameters required for performing the coordinate transformation and the expansion parameters required for performing the rotation transformation; among them, the map delineated by the currently stored point cloud is a grid map delineated with the lidar as the center and the maximum detection diameter reached in the currently stored point cloud as the side length of the map, and the size of this grid map includes the maximum height value and the minimum height value restricted by the occlusion of obstacles in the height direction of the map, and the maximum width value and the minimum width value restricted by the occlusion of obstacles in the width direction of the map; among them, aligning the bus bit width means that the memory occupied by the boundary grid of the pre-expanded grid map is equal to the bit width of the bus for transmitting this grid map; among them, the foregoing expansion parameters include coordinate translation parameters; the sum value of the coordinate translation parameters required for the maximum detection radius reached by the currently stored point cloud, the coordinate translation parameters required for performing the coordinate transformation, the coordinate translation parameters required for performing the rotation transformation, and the coordinate translation parameters required for aligning the bus bit width determines the preset coordinate offset value; the offset value register is used to store the foregoing expansion parameters, and the offset value register is a parameter register set inside the correlation scan matching hardware circuit. The pre-expanded grid map is a grid map expanded successively according to the coordinate translation parameters, the coordinate rotation parameters, and the memory parameters occupied by the side length of the map based on the maximum coverage radius of the point cloud, which can not only constrain the actual use size of the map, but also adapt to the bus bit width used during map transmission, ensure data alignment, and reduce the storage space occupied by the map.

[0032] Further, the construction steps of one layer of the original map at a specific resolution include: using a preset sliding window to sample the grid map at the first resolution with a sampling interval of the side length of the preset grid matching at the specific resolution; each time the preset sliding window samples, merging all the grids covered by the preset sliding window in the grid map at the first resolution into a preset grid, and then forming a grid of one layer of the original map at the specific resolution with this preset grid, where each preset grid stores a matching index value; when the preset sliding window has covered all the grid areas of the grid map at the first resolution row by row or column by column, all the merged preset grids form one layer of the original map at the specific resolution; among them, the larger the size of the preset sliding window, the smaller the specific resolution; where the aforementioned grid point is the central position of the grid where it is located.

[0033] This technical solution designs a sliding window to perform grid merging processing on the original grid map at the first resolution. The number of grids representing the same regional position in each layer of the merged grid map is different, realizing hierarchical fuzzy processing, that is, realizing different resolutions for each layer of the map; furthermore, ensuring that the occupancy probability of the map storage at low resolution is relatively large, and ensuring that the lower the resolution of the map matched by the point cloud, the higher the cumulative score calculated. Therefore, when it is necessary to search for a pose or a matching node with a higher score value, this technical solution starts searching layer by layer from the map with the highest resolution.

[0034] Further, the point cloud processing module includes a point cloud rotation sub-module and a point cloud discretization sub-module; the point cloud rotation sub-module includes a first register, a second register, a first point cloud multiplier, a second point cloud multiplier, a third point cloud multiplier, a fourth point cloud multiplier, a point cloud adder, and a point cloud subtractor, and is configured as a pipeline structure; the first point cloud multiplier, the second point cloud multiplier, the third point cloud multiplier, and the fourth point cloud multiplier all belong to multipliers, the point cloud adder belongs to an adder, and the point cloud subtractor belongs to a subtractor; the first input end of the first point cloud multiplier is connected to the output end of the first register; wherein, the interconnection bus transmits the abscissa of the point cloud to the input end of the first register; the abscissa of the point cloud transmitted by the interconnection bus to the first register is derived from the abscissa of the currently stored point cloud; the second input end of the first point cloud multiplier is used to receive the cosine function value corresponding to the rotation angle reached by the current rotation transformation; wherein, the rotation angle reached by the current rotation transformation is the sum of the angle search step size and the rotation angle reached by the previous rotation transformation; the angle search step size is stored in the search window parameter register; the search window parameter register is a parameter register set inside the correlation scan matching hardware circuit, and is used to store the poses and associated step search information included in the search window; the first input end of the second point cloud multiplier is connected to the output end of the second register; wherein, the interconnection bus caches the ordinate of the point cloud to the second register; the ordinate of the point cloud transmitted by the interconnection bus to the second register is derived from the ordinate of the currently stored point cloud; the second input end of the second point cloud multiplier is used to receive the sine function value corresponding to the rotation angle reached by the current rotation transformation; the first input end of the third point cloud multiplier is connected to the output end of the first register; the second input end of the third point cloud multiplier is used to receive the sine function value corresponding to the rotation angle reached by the current rotation transformation; the first input end of the fourth point cloud multiplier is connected to the output end of the second register; the second input end of the fourth point cloud multiplier is used to receive the cosine function value corresponding to the rotation angle reached by the current rotation transformation; the first input end of the point cloud subtractor is connected to the output end of the first point cloud multiplier, and the second input end of the point cloud subtractor is connected to the output end of the second point cloud multiplier; the point cloud subtractor is used to transmit the output difference to the point cloud discretization sub-module and set the output difference as the abscissa result of the current rotation transformation; the first input end of the point cloud adder is connected to the output end of the third point cloud multiplier, and the second input end of the point cloud adder is connected to the output end of the fourth point cloud multiplier; the point cloud adder is used to transmit the output sum to the point cloud discretization sub-module and set the output sum as the ordinate result of the current rotation transformation.

[0035] This technical solution uses a rotation matrix as the basic operation architecture, and constructs a point cloud rotation sub-module only using adders, subtracters and multipliers, which is used to receive and process the trigonometric function values corresponding to the step rotation angles pre-calculated by software. It not only eliminates the trigonometric function hardware calculation module, but also controls the iterative execution of the rotation transformation of the point cloud rotation sub-module on the basis of following the step rotation, obtains the point cloud data after the rotation transformation, and provides it to the point cloud discretization sub-module to continue the coordinate system transformation, determines the traversal of the rotation angle currently searched in the search window, so as to speed up the cumulative calculation operation of the score value.

[0036] Compared with the rotation based on the initial position, the method of this module in this paper avoids the hardware trigonometric function calculation without increasing the additional storage space of the hardware and without losing the calculation accuracy, and makes full use of the update transformation operation of the point cloud when scanning and matching to a new layer of map to obtain the rotation angle obtained under the step-by-step rotation.

[0037] Further, the point cloud discretization sub-module includes a first discrete adder, a first discrete subtractor, a first discrete multiplier, a second discrete adder, a second discrete subtractor, and a second discrete multiplier, which are configured to form a pipeline structure; both the first discrete adder and the second discrete adder belong to adders, both the first discrete subtractor and the second discrete subtractor belong to subtractors, and both the first discrete multiplier and the second discrete multiplier belong to multipliers; the first input terminal of the first discrete adder is connected to the output terminal of the point cloud subtractor, and the first input terminal of the first discrete adder is used to receive the abscissa result obtained from the current rotation transformation; the second input terminal of the first discrete adder is used to receive the abscissa of the robot position transmitted by the interconnection bus, where the abscissa of the robot position is the abscissa value of the robot in the world coordinate system calculated in advance; the first discrete adder is used to control the addition of the abscissa result obtained from the current rotation transformation and the abscissa of the robot position, and then output the sum value obtained by the addition, so that the sum value becomes the abscissa value of the point cloud in the world coordinate system; the first input terminal of the first discrete subtractor is connected to the output terminal of the first discrete adder; the second input terminal of the first discrete subtractor is used to receive the maximum abscissa value of the map transmitted by the interconnection bus, and the maximum abscissa value of the map is derived from the map size register and transmitted to the interconnection bus by the map size register; where the map size register is a parameter register set inside the correlation scan matching hardware circuit and is used to store the size range of the grid map that meets the transmission requirements of the bus bit width; the first input terminal of the first discrete multiplier is connected to the output terminal of the first discrete subtractor; the second input terminal of the first discrete multiplier is used to receive the reciprocal of the map resolution transmitted by the interconnection bus, and the reciprocal of the map resolution is derived from the map resolution register and transmitted to the interconnection bus by the map resolution register; where the map resolution register is a parameter register set inside the correlation scan matching hardware circuit and is used to store the resolution information of the first map to be searched for a corresponding layer; the product value output by the first discrete multiplier is configured as the abscissa value of the discrete point cloud to complete the coordinate system transformation, realizing the alignment of the abscissa value of the currently stored point cloud to the coordinate system of the first map to be searched for a corresponding layer, and at the same time determining the completion of the discretization of the abscissa of the currently stored point cloud; the first discrete multiplier is used to output the product value obtained by multiplication to the memory module through the interconnection bus and determine that the product value is the abscissa that matches the currently aligned and matched node on the first map to be searched for a corresponding layer; the first input terminal of the second discrete adder is connected to the output terminal of the point cloud adder, and the first input terminal of the second discrete adder is used to receive the ordinate result obtained from the current rotation transformation;The second input terminal of the second discrete adder is used to receive the vertical coordinate of the robot position transmitted by the interconnection bus, where the vertical coordinate of the robot position is the pre-calculated vertical axis coordinate value of the robot in the world coordinate system; the robot position register is used to store the vertical coordinate of the robot position and the horizontal coordinate of the robot position, and the robot position register is a parameter register set inside the correlation scanning and matching hardware circuit; the second discrete adder is used to control the addition of the vertical coordinate result of the current rotation transformation and the vertical coordinate of the robot position, and then output the sum value obtained by the addition, so that the sum value becomes the vertical axis coordinate value of the point cloud transformed into the world coordinate system; the first input terminal of the second discrete subtractor is connected to the output terminal of the second discrete adder; the second input terminal of the second discrete subtractor is used to receive the maximum vertical coordinate value of the map transmitted by the interconnection bus, and the maximum vertical coordinate value of the map is derived from the map size register and transmitted from the map size register to the interconnection bus; the first input terminal of the second discrete multiplier is connected to the output terminal of the second discrete subtractor; the second input terminal of the second discrete multiplier is used to receive the reciprocal of the map resolution transmitted by the interconnection bus; the product value output by the second discrete multiplier is configured as the vertical coordinate value of the discrete point cloud to complete the coordinate system transformation, so as to align the vertical coordinate value of the currently stored point cloud into the coordinate system of the corresponding first search map layer, and at the same time determine the discretization of the horizontal coordinate of the currently stored point cloud; the second discrete multiplier is used to output the product value obtained by multiplication to the memory module through the interconnection bus, and determine that the product value is the vertical coordinate that matches the currently aligned and matched node on the corresponding first search map layer.

[0038] For each point cloud, this technical solution reads the rotated horizontal coordinate and the rotated vertical coordinate from the point cloud rotation sub-module, and then through the adder and subtractor, the horizontal and vertical coordinates of the above point cloud are transformed from the point cloud coordinate system to the coordinate system of the corresponding first search map layer along the respective adapted coordinate axis directions in the same offset manner, including swapping the coordinate axes where the point cloud is located to achieve the coordinate system transformation; then multiply by the reciprocal of the grid map resolution to complete the discretization process of the point cloud.

[0039] Furthermore, the state machine control module belongs to a finite state machine; the state machine control module is used to schedule the working states of the memory module, the point cloud processing module, and the node search module; wherein, the working states include the multi-resolution map reading state, the search state, and the loop state; the state machine control module is used to execute: in the multi-resolution map reading state, control the point cloud processing module to receive a first to-be-searched map of one layer; then switch to the search state, control the point cloud processing module to perform the aforementioned coordinate system transformation on the currently stored point cloud to obtain a discrete point cloud, and then control the node search module to calculate the index value of the discrete point cloud mapped to the map storage space where the current first to-be-searched map of one layer is located, and use the currently calculated index value to search for the occupancy probability in the corresponding map storage space, control the aforementioned accumulator to perform an accumulative calculation on the searched occupancy probability to obtain the score value of the corresponding node; after saving the currently calculated score value, switch to the loop state; in the loop state, if it is determined that the currently processed node is the root node and the occupancy probabilities corresponding to all root nodes have not been indexed yet, then return to the search state, continue to use the newly calculated index value to search for the occupancy probability in the corresponding map storage space until the occupancy probabilities corresponding to all root nodes supported by the search window have been completed, and then allow to choose whether to trigger an interruption; in the loop state, if it is determined that there are still nodes whose score values have not been calculated and the currently processed node is not the root node, then return to the search state, continue to control the aforementioned accumulator to perform an accumulative calculation on the searched occupancy probability until the score values of all nodes supported by the search window have been calculated, and then allow to choose whether to trigger an interruption; after triggering an interruption, the working states supported for the correlation scan matching hardware circuit to jump include but are not limited to the multi-resolution map reading state or the search state; wherein, the search window is a set of pose parameters and depth information of nodes, including the pose parameters and depth information of one or more nodes; the parameters required to execute the correlation scan matching algorithm include the pose parameters of nodes, and the pose parameters of nodes include a matching rotation angle and horizontal and vertical coordinates.

[0040] In this technical solution, the state machine control module completes the correlation scan matching of the poses corresponding to the search in the search window on each layer of the first to-be-searched map in the way of hardware-scheduling the working state and the interruption signal. Among them, in each correlation scan matching, every time the node search module calculates the score value of a pose in the search window, it controls the point cloud processing module to perform a rotation and a translation process on each point cloud in sequence to read out the occupancy probability matching the index value from the memory module, realizes the operation of the loop state, ensures the real-time performance of the hardware-correlation scan matching algorithm to search the branch and bound search tree, and improves the convergence speed of the algorithm.

[0041] Further, in the step I, the CPU searches for 4 child nodes in the second map to be searched at the corresponding layer for each branch of the non-leaf node. Among them, these 4 child nodes belong to the same branch and bound search tree, so that the rotation parameters required by these 4 child nodes are the same, the coordinate translation parameters required by these 4 child nodes are different from each other, and the rotation parameters and coordinate translation parameters required by these 4 child nodes are included within the same search window; both the rotation parameters and the coordinate translation parameters are the parameters required for performing the correlation scan matching algorithm. Thereby, the search range is reduced, and the leaf nodes on different branch and bound search trees can be searched more quickly or pruning can be accelerated to enter another branch for continued search by scanning the first map to be searched and the second map to be searched at each layer through the correlation scan matching algorithm, and based on the leaf node scores, the current optimal upper bound value is updated, thereby improving the execution efficiency of the branch and bound algorithm.

[0042] Further, a bus interface module is externally provided for the hardware acceleration positioning system. The bus interface module includes a DMA controller module and a transmission bus; the DMA controller module is used for continuously transmitting the data stored in the physical storage space with discontinuous addresses in batches, reducing the triggering times of the software interrupt of the CPU; the transmission bus includes a first bus and a second bus. The first bus has signal transceiver connections with the memory module, the point cloud processing module, the node search module, the state machine control module, the interconnection bus, and the DMA controller module respectively. The first bus is used for configuring the data transmission parameters for the DMA controller module, and the first bus is also used for configuring the parameters stored inside the map size register, the parameters stored inside the map resolution register, the extended parameters required for the maximum detection radius reached by the currently stored point cloud, the extended parameters required for performing the coordinate system transformation, the extended parameters required for performing the rotation transformation, and the extended parameters required to meet the transmission requirements of the bus bit width, so as to realize the memory mapping communication of the correlation scan matching hardware circuit; the second bus is connected to the DMA controller module and is used for transmitting the multi-level grid map pre-constructed by the CPU, the sine function values under the pre-configured corresponding rotation parameters, the cosine function values under the pre-configured corresponding rotation parameters, and the point cloud currently collected by the lidar to the memory module; wherein, the transmission bus follows the AMBA protocol.

[0043] This technical solution provides a bus interface architecture module for the correlation scan matching hardware circuit and the CPU, and designs a first bus suitable for memory - mapped communication with simple and low throughput (transmitting extended parameters, map feature parameters, and associated basic control signals to the registers inside the correlation scan matching hardware circuit) and a second bus for high - speed data streams (point clouds collected in real - time by lidar, transmitting each layer of the first map to be searched and trigonometric function values that urgently need to be calculated to the memory module) according to data transmission performance; it improves the real - time performance of the arithmetic operations of the correlation scan matching hardware circuit, and the real - time performance of the CPU running the branch - and - bound algorithm and the correlation scan matching algorithm.

[0044] A robot is equipped with a lidar, and the hardware acceleration positioning system described above is arranged inside the robot to solve the re - positioning problem on the mobile robot embedded platform. Compared with the existing - technology robot that executes the branch - and - bound algorithm in a pure - software platform manner, even when the frequency of the correlation scan matching hardware circuit is not high, this technical solution still has an obvious advantage in running speed and can better meet the re - positioning requirements of the mobile robot platform. Brief Description of the Drawings

[0045] Figure 1 It is a schematic diagram of the circuit module of the hardware acceleration positioning system under the cooperation of software and hardware disclosed in an embodiment of the present invention.

[0046] Figure 2 It is a flowchart of the method steps executed by the CPU in the hardware acceleration positioning system disclosed in an embodiment of the present invention. Detailed Embodiments

[0047] The following further describes the detailed embodiments of the present invention with reference to the drawings. It should be noted that if there is no conflict, the various features in the embodiments of the present invention can be combined with each other, and all are within the protection scope of the present invention. In addition, although functional module division is carried out in the system schematic diagram and the logical sequence is shown in the flowchart, in some cases, the steps shown or described can be executed in a different order from the module division in the system or the sequence in the flowchart. Furthermore, the terms "first", "second", "third", etc. adopted by the present invention do not limit the data and execution order, but only distinguish the same items or similar items with basically the same functions and effects.

[0048] The robot according to the embodiment of the present invention can be configured into any suitable shape to implement specific business function operations. For example, the robot according to the embodiment of the present invention can be a delivery robot, a handling robot, a nursing robot, a cleaning robot, etc. The robot generally includes a housing, a sensor unit, a driving wheel component, a storage component, and a controller. The outer shape of the housing is generally circular. In some embodiments, the outer shape of the housing can be generally elliptical, triangular, D-shaped, cylindrical, or other shapes. The sensor unit is used to collect some motion parameters of the robot and various data in the environmental space. In some embodiments, the sensor unit includes a lidar mounted above the housing, and its mounting height is higher than the height of the top surface shell of the housing. The lidar is used to detect the obstacle distance between the robot and the obstacle.

[0049] In the branch and bound algorithm, pruning is performed by bounding, trimming the branches that are not likely to exceed the current best value, thereby screening the branches, and seeking the optimal solution only among a part of the feasible solutions instead of enumerating all of them and then searching, which improves the search efficiency. Also, since the branch and bound algorithm is an iterative and search method, when the branch and bound algorithm is involved in the loop detection of the laser SLAM backend, a strongly constrained loop detection is formed, but the overall computational amount is huge, the computational complexity is extremely high, and the real-time performance of the algorithm is poor. Especially for the optimization of the grid map pre-constructed by the robot in a large-scale scene, the time spent on eliminating the pose error is relatively long, resulting in poor real-time performance of the robot's pose estimation and affecting the real-time effect of the robot's navigation and positioning.

[0050] To improve the problem of poor real-time performance of the aforementioned loop detection, the present invention discloses a hardware-accelerated positioning system under software-hardware collaboration. The hardware-accelerated positioning system is electrically connected to a lidar. The lidar is installed on the body of a robot. The hardware-accelerated positioning system includes a CPU and a correlation scanning and matching hardware circuit. An improvement in hardware scheduling is added based on the branch-and-bound algorithm. It should be noted that the main idea of the branch-and-bound algorithm is to construct the possible poses of the robot into the nodes of a branch-and-bound search tree, and the root node represents all possible solutions. The child nodes are the branch nodes of the grids of the parent node, and the score of each node is the upper bound of all child nodes and leaf nodes under that node. Therefore, in the same branch-and-bound search tree (tree-shaped search structure or skeleton search structure), corresponding nodes are selected for branching operations from the latest searched child nodes in the order of the score values. For the score value of the currently searched child node not being greater than the current optimal upper bound value, the currently searched child node is pruned, and the CPU and the correlation scanning and matching hardware circuit are no longer allowed to expand the search, so as to accelerate the algorithm search process. Among them, the branch-and-bound algorithm adopted in this embodiment generates the nodes of the branch-and-bound search tree based on a depth-first search strategy, that is, along a branch, search to the bottom from top to bottom, then return to the node at the previous depth level, and continue to search along another branch, so as to search for the optimal solution in a multi-resolution multi-layer grid map constructed in a large-scale scenario. Especially when selecting the optimal solution from the candidate solutions, the pose included in the node corresponding to the optimal solution is set as the pose of the robot. Searching a node of each branch-and-bound search tree adopts the branching operation included in the branch-and-bound algorithm.

[0051] As an embodiment, the correlation scan matching hardware circuit is used to search each layer of the first map to be searched in a hardware parallel processing manner, and update the first upper bound value corresponding to the currently searched node. The search method of the correlation scan matching hardware circuit conforms to the search concept of the branch and bound algorithm; the process of searching each layer of the first map to be searched is to search for nodes on the first map to be searched in the corresponding layer along the branches according to the branch and bound algorithm; the correlation scan matching hardware circuit starts from the topmost first map to be searched, calculates the score value of the root node, which is equivalent to designing a dedicated pipeline structure on the hardware circuit of this embodiment to calculate all candidate solutions at the lowest resolution. Then, search for child nodes in each remaining layer of the first map to be searched along the branches, and calculate the upper bound score value of the pose of the first map to be searched in the layer where the child node is located, that is, the first upper bound value corresponding to the currently searched node, that is, the first upper bound value corresponding to the branch where the currently searched node is located. It should be noted that if the score value of the currently searched child node is not greater than the current optimal upper bound value, the upper bound score value of the pose of the first map to be searched in the layer where the child node is located is the current optimal upper bound value, and it is determined as the updated first upper bound value; if the score value of the currently searched child node is greater than the current optimal upper bound value, the upper bound score value of the pose of the first map to be searched in the layer where the child node is located is the score value of this child node, and the score value of this child node is updated to the current optimal upper bound value, and it is determined as the updated first upper bound value. Among them, the first map to be searched is a grid map with a relatively low resolution and belongs to the few most blurred layers in a multi-layer resolution map. However, it occupies the vast majority of the computational workload in the process of using the branch and bound operation to search for nodes and calculate score values. Therefore, in this embodiment, a dedicated pipeline structure of the correlation scan matching hardware circuit is designed to calculate the score value of the node corresponding to the pose of each layer of the first map to be searched, and update the first upper bound value. Since the resolution of the first map to be searched is relatively low, the score value corresponding to the pose in the second map to be searched is relatively large, and it is easily updated to the current optimal upper bound value during the branch and bound process, so that the optimal solution of the branch and bound algorithm appears in the updated first upper bound value.

[0052] The CPU is used to, whenever the correlation scan matching hardware circuit searches for a node corresponding to a grid point of a map layer at a preset depth, preferably when the correlation scan matching hardware circuit searches for a node corresponding to a grid point of the first map to be searched with the highest resolution, search each layer of the second map to be searched by executing the branch and bound algorithm, perform scan matching on the corresponding layer of the second map to be searched by executing the correlation scan matching algorithm, and then calculate the second upper bound value corresponding to the currently searched node using the matched result. When the correlation scan matching hardware circuit searches all the first maps to be searched along the branches, it starts to be handed over to the CPU to search the second map to be searched by executing the branch and bound algorithm, where the second map to be searched is located below the bottom layer of the first map to be searched, and the resolution of each layer of the first map to be searched is lower than that of any layer of the second map to be searched, so that the CPU is used to search each layer of the second map to be searched with a higher resolution. Since the resolution of the second map to be searched is higher, the score value corresponding to the pose in the second map to be searched is smaller, and it is easy to be pruned during the branch and bound process, so that most of the second upper bound values can only be maintained at the current optimal upper bound value, that is, not greater than the first upper bound value, and the computational amount of the CPU is relatively low compared with the hardware parallel processing method of the correlation scan matching hardware circuit, but the real-time performance of the hardware acceleration positioning system running the branch and bound algorithm is improved. During the process of the CPU searching each layer of the second map to be searched by executing the branch and bound algorithm, the point cloud collected by the lidar is scanned and matched on the corresponding layer of the second map to be searched by executing the correlation scan matching algorithm to calculate the occupancy probability of the node searched by the branch and bound algorithm in the matched pose, specifically the occupancy probability on the corresponding layer of the second map to be searched, and the score value corresponding to the matched pose (the score value of the node) is obtained by accumulation, and then the second upper bound value corresponding to the currently searched node is updated by comparing with the current optimal upper bound value.

[0053] The CPU is also used to configure the pose corresponding to the latest obtained second upper bound value or the pose corresponding to the latest obtained first upper bound value as the pose of the robot; specifically, for each layer of the first map to be searched, the relevance scan matching hardware circuit calculates score values for the currently searched nodes and sorts them to obtain the maximum score value in the currently traversed layer of the first map to be searched, and uses it to update the current optimal upper bound value; when traversing each layer of the second map to be searched, the CPU calculates score values for the currently searched nodes and sorts them to obtain the maximum score value in the currently traversed layer of the second map to be searched, and uses it to update the current optimal upper bound value, but does not sort by combining the first layer of the map to be searched and the previously traversed second map to be searched, but only sorts the nodes corresponding to the current search on the current layer of the map; after traversing to the bottom layer of the second map to be searched, based on the depth-first search strategy, return to the map with a higher depth level and repeat the above search and calculation steps until all the sorted score values in batches are used to complete the update operation of the current optimal upper bound value, so as to configure the grid point corresponding to the maximum score value as the position of the robot, that is, during the traversal process, the pose corresponding to the latest obtained second upper bound value or the pose corresponding to the latest obtained first upper bound value. In this embodiment, generally, the latest obtained first upper bound value is set as the maximum score value.

[0054] Therefore, when the hardware acceleration positioning system runs the branch and bound algorithm, according to the depth of the current branch, it allocates to the CPU to calculate the second upper bound value and the relevance scan matching hardware circuit to calculate the first upper bound value, that is, the boundary value of the nodes searched on the current branch. This enables the relevance scan matching hardware circuit not to transmit multiple layers of the second map to be searched into the relevance scan matching hardware circuit, saving the RAM resources inside the relevance scan matching hardware circuit. Specifically, when the operating frequency of the CPU is 599.99 MHz and the operating frequency of the relevance scan matching hardware circuit is 133.33 MHz, the relevance scan matching hardware circuit increases the frame rate of the traditional branch and bound algorithm from 3 FPS to 19.4 FPS. The hardware acceleration positioning system disclosed in this embodiment performs branch and bound processing on multiple layers of maps through software and hardware module scheduling, and is specifically used to handle the real-time problem of loop detection. This hardware acceleration positioning system hands over the search process of multiple layers of maps with relatively low resolution to the relevance scan matching hardware circuit to be implemented in a hardware parallel processing manner, and hands over the search process of multiple layers of maps with relatively high resolution to the CPU to be implemented, thus reasonably allocating computing resources within the same system, maximizing the computing efficiency of the algorithm executed in a pure software manner, and saving some hardware resources.

[0055] In some embodiments, the CPU can be a general-purpose processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), a single-chip microcomputer, an AR (Acorn RISC Machine), or other programmable logic devices, discrete gate or transistor logic, discrete hardware components, or any combination of these components. Additionally, the controller can also be any conventional processor, controller, microcontroller, or state machine. The controller can also be implemented as a combination of computing devices, for example, a combination of a DSP and a microprocessor, multiple microprocessors, one or more microprocessors in combination with a DSP and / or any other such configuration.

[0056] It should be noted that the aforementioned first map to be searched and the second map to be searched are hierarchically arranged. Among them, all the first maps to be searched are processed by the correlation scanning and matching hardware circuit, and all the second maps to be searched are processed by the CPU. The first map to be searched and the second map to be searched belong to different layers of maps and are used to form a multi-layer grid map.

[0057] More than one branch-and-bound search tree is constructed for all layers of the first map to be searched and all layers of the second map to be searched. In fact, the branch-and-bound search tree is constructed based on the resolution of the first map to be searched and the resolution of the second map to be searched. The pose of each layer of the map corresponds to the node at the same depth of the branch-and-bound search tree, so that the grid points corresponding to the root nodes of all the branch-and-bound search trees are in the first map to be searched with the lowest resolution. Specifically, the hardware-accelerated positioning system stores a pre-constructed grid map with the first resolution, which is constructed from the point cloud collected by the lidar;

[0058] The grid map with the first resolution is constructed by the CPU into a multi-level grid map, which includes the first map to be searched with the first preset number of layers and the second map to be searched with the second preset number of layers. The first map to be searched with the first preset number of layers is hierarchically arranged; the second map to be searched with the second preset number of layers is also hierarchically arranged; preferably, in this multi-level grid map, the maps with the number of layers arranged from 1 to the first preset number of layers are the first maps to be searched; the maps with the number of layers arranged from (the first preset number of layers + 1) to the second preset number of layers are the second maps to be searched.

[0059] Among them, the CPU is pre-configured with the first preset number of layers and the second preset number of layers, that is, the depth of the branch and bound search tree is configured; the CPU is pre-configured with the map resolution of each layer, including the highest resolution and the lowest resolution. For example, assuming that the grid map to be layered is h layers, the initial map resolution is r, that is, the first resolution is r, and the area represented by each grid of the grid map with the first resolution is the square of r. Then, in the most blurred layer of the map (the layer of the map with the lowest resolution), the area represented by each grid is the square of the product of 2 to the power of h and r. When the multi-level grid map is configured to be 7 layers, the grid points under each layer of the map respectively store the maximum probability values within the adjacent 1, 4, 16, 64 up to the square of 64 grids. Preferably, r is equal to 5 cm.

[0060] Preferably, the grid map with the first resolution is composed of a certain number of laser 360-degree scan points, specifically constructed by probability grids of 5 cm * 5 cm {Pmin, Pmax}. When the map is created, if the grid probability is less than Pmin, it means the point is unobstructed, between Pmin and Pmax means unknown, and greater than Pmax means the point is obstructed. Each frame of laser scan generates a set of grid points, and each set of grid points is assigned an occupancy probability value. If the grid point already has a probability value previously, the probability value of the grid point needs to be updated. In this embodiment, the probability of each grid point in the map is replaced by the maximum probability value in its nearby area and configured as the occupancy probability. Thus, by directly querying the map with the corresponding resolution, the index can be obtained, and then the score of the grid point can be obtained through accumulation, that is, the aforementioned score value, or it can be considered that the score of the corresponding grid point is directly indexed, and this score is regarded as the maximum probability value that has been previously configured in the nearby area.

[0061] It should be noted that for maps with lower resolutions (that is, higher-level maps), the maximum probability value should be obtained in a larger area, and its size is consistent with the size of the grid at the current resolution. The grid point is used to represent the center of the grid and stores the abscissa and ordinate of the grid. The occupancy probability of the grid points in the map with a lower resolution is closer to 1.

[0062] In this embodiment, the maximum resolution in the second map to be searched for all layers is equal to the first resolution; the first map to be searched for the first preset number of layers and the second map to be searched for the second preset number of layers are arranged vertically in ascending order of resolution to construct a branch and bound search tree. The depth of each layer of the first map to be searched and the depth of each layer of the second map to be searched are respectively equal to the depth of the nodes at the same vertical height of the branch and bound search tree. Of course, the first map to be searched for the first preset number of layers and the second map to be searched for the second preset number of layers can also be arranged vertically and horizontally in ascending order of resolution. Among them, the sequence numbers of the first map to be searched for the first preset number of layers and the second map to be searched for the second preset number of layers are used to represent the depth of the corresponding layer of the map. Preferably, the first map to be searched for the first preset number of layers and the second map to be searched for the second preset number of layers are arranged from high to low. The higher the layer of the map, the lower the resolution and the larger the sequence number. Then, the depth corresponding to the map arranged from high to low is decreasing. Preferably, the first map to be searched for the first preset number of layers and the second map to be searched for the second preset number of layers are arranged from high to low. The higher the layer of the map, the lower the resolution and the smaller the sequence number. Then, the depth corresponding to the map arranged from high to low is increasing.

[0063] In the hardware-accelerated positioning system, in order to reduce memory occupancy and data transmission time, the original grid map and the constructed multi-level grid map are represented with different bit widths. Since the data stored in the grid map is fixed-point numbers, when blurring in the X direction (separating a low-resolution map along the X axis), the bit width is first extended to a floating-point number value. Then, when blurring in the Y direction (changing lines along the Y axis to continue separating a low-resolution map along the X axis in the next line), it is then converted to be stored with a lower data bit width.

[0064] It should be noted that each node in each branch and bound search tree contains the index of the pose component and the depth information of the branch and bound search tree. At a specific rotation angle and specific depth, each root node contains multiple searchable child nodes. The leaf node is the end node on a branch of the branch and bound search tree, and the branch and bound search tree is a node search structure formed by performing branch operations on the multi-level grid map.

[0065] Therefore, a low-resolution map is generated based on the highest-resolution map to form a multi-layer grid map arranged from top to bottom, so as to facilitate the construction of a tree structure. Among them, the higher the layer of the grid map, the lower the resolution, and the higher-layer grid map is transmitted to the hardware-accelerated positioning system first. The lower the layer of the grid map, the higher the resolution, realizing hierarchical blurring processing of the map. Let the map at low resolution store the maximum occupancy probability value of adjacent grids, ensuring that the lower the resolution during regional search of the point cloud, the higher the accumulated score value. Thus, during the subsequent branch and bound process, for the pose with the highest score value, it is necessary to continue branch search in a map with a higher resolution layer.

[0066] As an embodiment, the hardware-accelerated positioning system further includes a bus interface, and both the CPU and the correlation scanning and matching hardware circuit are electrically connected to the bus interface. Among them, the bus interface belongs to a software and hardware interface; as Figure 2 shown, the CPU is configured to execute the following steps:

[0067] Step A: The CPU configures, through the bus interface, the parameters required for executing the branch and bound algorithm and the parameters required for executing the correlation scanning and matching algorithm for the correlation scanning and matching hardware circuit, that is, configures the parameters of the loop detection algorithm; then enters Step B. Among them, the parameters required for executing the branch and bound algorithm include the first preset number of layers, the resolution of the first map to be searched for each layer, the second preset number of layers, the resolution of the second map to be searched for each layer, and the number of root nodes, all of which are pre-configured by the CPU and stored in relevant registers.

[0068] Step B: The CPU transmits, through the bus interface, the first map to be searched for the layer with the lowest resolution to the correlation scanning and matching hardware circuit and stores it in the internal storage space of the correlation scanning and matching hardware circuit; then the CPU sends a "start" instruction to the correlation scanning and matching hardware circuit and enters Step C. Among them, the map at the depth where the root node is located is the first map to be searched, and it is the first layer of the map searched by the correlation scanning and matching hardware circuit, and it is the first map to be searched for the layer with the lowest resolution. Thus, the correlation scanning and matching hardware circuit undertakes the search work of the layer of the map with the largest computational load.

[0069] It should be added that before executing Step B, the CPU configures the address and the size of the data to be transmitted (including pre-computed trigonometric function values, point clouds, and the map of the layer with the lowest resolution among the currently constructed layers of the map) to the bus interface; then the CPU starts the bus interface for data transmission.

[0070] Step C: The CPU activates the correlation scan matching hardware circuit to search for all root nodes and triggers the correlation scan matching hardware circuit to calculate the score values of the corresponding root nodes. Specifically, while searching for the next root node, the score value of the currently searched root node is calculated. Meanwhile, the CPU discretizes the point cloud according to the parameters required for executing the correlation scan matching algorithm. Specifically, the point cloud at all rotation angles is discretized for further calculation to obtain the index value of the occupancy probability of each point cloud in the corresponding layer of the second map to be searched. Then, it proceeds to Step D. All the rotation angles are all the rotation angles matched by a search window.

[0071] Step D: The CPU transmits the remaining first map to be searched to the correlation scan matching hardware circuit through the bus interface. Meanwhile, the CPU sorts all the root nodes currently searched by the correlation scan matching hardware circuit in ascending order of the score values, and then stores the root nodes in the stack according to the current sorting. Then, it proceeds to Step E. Among them, the stack is a cache space internally opened by the hardware acceleration positioning system that supports last-in, first-out, and is used to store the currently searched nodes in a batch and ensure that the node that exits the stack first is the node with the largest score value among the currently batch of nodes pushed onto the stack. The nodes pushed onto the stack in the previous batch are pushed to the bottom of the stack, so that the score value of the currently exiting node is not necessarily the largest, but after multiple exits, the node with the largest or optimal score value is obtained.

[0072] It should be added that between Step C and Step D, the bus interface returns the currently searched root nodes and the calculated score values to the DRAM for the CPU to read, but not necessarily all the searched root nodes. Among them, the DRAM can be internal to the hardware acceleration positioning system. Meanwhile, the CPU configures the address and the size of the data to be transmitted to the bus interface, which can be the address and the size of the data to be transmitted of two adjacent layers of the first map to be searched, and then activates the CPU to activate the bus interface for data transmission, enters Step D, and transmits the data to the map storage space in the correlation scan matching hardware circuit.

[0073] It should be emphasized that in the hardware acceleration positioning system, the resolution of the first map to be searched transmitted layer by layer to the correlation scan matching hardware circuit increases layer by layer, and the resolution of the second map to be searched processed layer by layer by the CPU also increases, so as to ensure that the map with a lower resolution is searched first to obtain the node with a relatively large or the largest score value as early as possible, and reserve it for updating the upper bound value of the node at the current depth subsequently.

[0074] Preferably, in the process that the correlation scan matching hardware circuit calculates the score value of the root node in step C (where the bus interface may simultaneously configure the relevant access information of a new layer of the first map to be searched), or in the process that the correlation scan matching hardware circuit searches for the root node in step C, if the correlation scan matching hardware circuit triggers an interruption, the correlation scan matching hardware circuit responds to the interruption signal and then proceeds to step D to receive the remaining layers of the first map to be searched transmitted by the bus interface. During the process that the CPU discretizes the point cloud in the second map to be searched in step C, if it detects that an interruption is triggered, the CPU responds to the interruption and then proceeds to step D to sort the currently searched root nodes; among them, the CPU and the correlation scan matching hardware circuit mostly work simultaneously. In the above process, the CPU does not adopt a polling mechanism. If the CPU pauses and waits for the correlation scan matching hardware circuit to calculate the score value, tens of thousands of instruction cycles will be missed, which is a huge waste of resources for the CPU because the time for searching the root node is relatively long and longer than the interruption duration. Adopting the interruption mechanism will greatly improve the processing efficiency of the program. It should be noted that the CPU is used to respond to the interruption triggered by the correlation scan matching hardware circuit and execute the interruption handling program.

[0075] During the process that the correlation scan matching hardware circuit performs branch operations on the remaining first map to be searched, for example, during the process that the correlation scan matching hardware circuit searches for the map with a resolution of 32r and the map with a resolution of 64r successively, the CPU adopts a polling mechanism to wait for the correlation scan matching hardware circuit to calculate and output the score values of the corresponding nodes, and then sorts the nodes searched in the same layer of the first map to be searched to avoid delays caused by adopting the interruption mechanism, where the currently searched node is not the root node. This embodiment addresses the problem that frequent interruptions of a CPU with a low frequency will cause a large amount of delays because the scanning and matching time of the correlation scan matching hardware circuit for non-root nodes is relatively short. Among them, the map with a resolution of 32r and the map with a resolution of 64r are both first maps to be searched with a resolution larger than that of the first map to be searched at the depth where the root node is located.

[0076] Step E: The CPU controls the node at the top of the stack to pop out of the stack, so that the node with the largest score value in the latest batch of nodes pushed onto the stack pops out first; then the CPU determines whether the currently popped node is a leaf node. If so, it proceeds to step F; otherwise, the CPU determines whether the score value of the currently popped node is greater than the current optimal upper bound value.

[0077] Step F: The CPU determines whether the score value of the currently popped node is greater than the current optimal upper bound value. If so, it updates the score value of this node to the current optimal upper bound value, and determines the updated current optimal upper bound value as the first upper bound value corresponding to the currently searched node or the second upper bound value corresponding to the currently searched node, so as not to discard any node corresponding to a grid point where an optimal solution may exist, and then proceeds to Step H; otherwise, it directly proceeds to Step H, where the initial value of the current optimal upper bound value is a preset initial threshold.

[0078] Preferably, in Step E, the top node of the stack is the first node or the second node, where the first node is configured with the coordinates of the grid point of the first map to be searched and the depth information of the first map to be searched to which this grid point belongs, and the second node is configured with the coordinates of the grid point of the second map to be searched and the depth information of the second map to be searched to which this grid point belongs. Whenever the CPU determines in Step F that the score value of the currently popped first node is greater than the current optimal upper bound value, it updates the score value of this first node to the current optimal upper bound value, and determines that the score value of this first node is the first upper bound value of the currently searched node; whenever the CPU determines in Step F that the score value of the currently popped second node is greater than the current optimal upper bound value, it updates the score value of this second node to the current optimal upper bound value, and determines that the score value of this second node is the second upper bound value of the currently searched node. In this embodiment, the candidate score values obtained from each branch operation are classified and sorted, and the types of nodes processed by different modules are distinguished, making the traversal priority of the nodes more obvious; and the last-in, first-out storage mechanism of the stack is used to select the child node with the largest score value on the current layer of the map for the branch, accelerating the search for the optimal node representing the robot's position.

[0079] Step G: When the CPU determines that the score value of the currently popped node is greater than the current optimal upper bound value, the CPU determines whether the absolute value of the difference between the depth of the currently popped node and the depth of the first map to be searched with the lowest resolution is greater than or equal to the reference depth difference. If so, it proceeds to Step I, and determines that the node corresponding to the grid point of the map at the preset depth searched by the correlation scanning and matching hardware circuit, and according to the rule of the vertical arrangement of the first map to be searched (the map with lower resolution belongs to a higher layer), the first map to be searched at the depth where the currently searched node is located is the first map to be searched with the highest resolution. Otherwise, it proceeds to Step J; where the reference depth difference is equal to the difference between the first preset number of layers and the value 1.

[0080] When the CPU determines that the score value of the currently popped node is less than or equal to the current optimal upper bound value, no branch operation is performed, then it is determined that a pruning operation is completed, and then Step H is executed.

[0081] Step I: The CPU performs a branching operation on the node described in Step G to search for all child nodes of the node described in Step G, and then calculates the score values of the currently searched child nodes for subsequent update to the second upper bound value when it is greater than the current optimal upper bound value and becomes the node popped out in Step E; then proceed to Step K. Among them, the branching operation in Step I is in the adjacent first search map of the next layer with a resolution higher than that of the layer map to which the node described in Step G belongs. Specifically, the pose parameters and depth information included in the search window are used to search for all child nodes of the node described in Step G among the grid points of this layer of the first search map. Among them, the number of child nodes is related to the pose parameters included in the search window, and preferably there are 4 child nodes, excluding the root node.

[0082] Step J: The correlation scan matching hardware circuit performs a branching operation in a hardware parallel processing manner to search for all child nodes of the node described in Step G, and calculates the score values of the currently searched child nodes for subsequent update to the first upper bound value when it is greater than the current optimal upper bound value and becomes the node popped out in Step E; then proceed to Step K. Among them, the branching operation in Step J is in the adjacent second search map of the next layer with a resolution higher than that of the layer map to which the node described in Step G belongs. Specifically, the pose parameters and depth information included in the search window are used to search for all child nodes of the node described in Step G among the grid points of this layer of the second search map. Among them, the number of child nodes is related to the pose parameters included in the search window, and preferably there are 4 child nodes, excluding the root node.

[0083] Step K: The CPU sorts the child nodes searched in Step I or the child nodes searched in Step J in the order described in Step D, and then controls the currently sorted child nodes to be pushed onto the stack in the order described in Step D; that is, sort the child nodes whose score values are currently calculated by the correlation scan matching hardware circuit in ascending order, and then control these child nodes to be pushed onto the stack in ascending order of the score values, and then return to Step E to ensure that the node popped out in Step E is the node with the largest score value among these child nodes. Since the nodes pushed down towards the bottom of the stack are prone to be pruned after subsequent popping out, compared with the node caching method of pushing onto the stack in descending order, it can reduce the search area and search time.

[0084] Step H: The CPU determines whether the stack is empty. If it is, the pose corresponding to the node with the current optimal upper bound value in Step F is configured as the pose of the robot, which is also the operation result of the loop detection algorithm based on the software and hardware co - design; otherwise, return to Step E to follow the principle of depth - first search. Before determining the pose of the robot in this embodiment, it is necessary to traverse all the nodes that may have the optimal solution (optimal upper bound value) searched during the operation of the branch - and - bound algorithm (i.e., all the nodes cached in the stack), and only discard the child nodes that definitely do not have the optimal solution (the nodes that have been pruned). Therefore, when the stack is empty, it means that all the root nodes (nodes with relatively higher score values) pushed into the bottom of the stack earlier have been traversed. After all the root nodes in the stack are popped out, the stack becomes empty. Among them, one root node corresponds to a branch - and - bound search tree; the depth of the first map to be searched in each layer and the depth of the second map to be searched in each layer are respectively equivalent to the depth of the corresponding nodes on the branch - and - bound search tree. The depth and the number of nodes of the branch - and - bound tree constructed by the search map area limit are limited, and only the nodes belonging to this interval are searched, excluding the rest of the nodes.

[0085] In the embodiments described in the foregoing Steps A to K, in this embodiment in the software - hardware collaborative system, according to the depth of the current branch, the correlation scanning and matching hardware circuit or the CPU is assigned to calculate the boundary value. Specifically, first, the first map to be searched with the first preset number of layers, which has the largest amount of computation and a lower resolution, is transmitted to the correlation scanning and matching hardware circuit for score value calculation, and then the second map to be searched with the second preset number of layers, which has a smaller amount of computation, is left in the CPU for score value calculation. Thus, the computation amount of the root nodes and most of the upper - bound calculations is allocated to the hardware for execution, and the CPU is allowed to search the tree according to a specific rule to reach the optimal score value as early as possible without traversing the entire tree, realizing the execution of expanding some branches, pruning some branches, sorting, and partial upper - bound calculation by vertically searching each layer of the map, thereby narrowing the search range. Therefore, this embodiment reasonably schedules the computing resources of the CPU and the hardware resources of the correlation scanning and matching hardware circuit, greatly simplifying the design of the hardware system.

[0086] As an embodiment, such as Figure 2As shown in the figure, the correlation scan matching hardware circuit includes a memory module, a point cloud processing module, a node search module, a state machine control module, and an interconnect bus. Among them, the memory module, the point cloud processing module, the node search module, and the state machine control module all establish data transmission relationships through the interconnect bus; the interconnect bus serves as the interconnect structure of the correlation scan matching hardware circuit, playing the role of data transmission and data sharing among various circuit modules inside the correlation scan matching hardware circuit, enabling the point cloud processing module, the node search module, and the state machine control module to be connected to the memory space of the memory module through the interconnect structure. The interconnect bus communicates and transmits data with the CPU or an external bus through the bus interface set in the correlation scan matching hardware circuit; among them, the state machine control module, as a finite state machine, is configured as a control unit in the correlation scan matching hardware circuit and is used to schedule the data transmission of the interconnect bus.

[0087] The point cloud processing module is used to read the currently stored point cloud from the memory module under the control of the state machine control module. Among them, the currently stored point cloud is a batch of point clouds that perform correlation scan matching operations in the correlation scan matching hardware circuit, simply referred to as a batch of currently processed point clouds. The number of cycles of the working state scheduled by the state machine control module is the number of this batch of point clouds. Specifically, it reads the existing point cloud from the square box marked with Point Cloud as shown in the figure; then controls the read point cloud to perform a rotation transformation to achieve the coverage of the angular search range of the search window by the point cloud according to the preset angular search step size, and then updates the result of the rotation transformation to a batch of currently processed point clouds, and then controls the updated point cloud to perform the next rotation transformation. Therefore, in this embodiment, the rotation transformation of the currently read point cloud is optimized to a step-by-step rotation, ensuring that the angle of each rotation is obtained by rotating another preset angular search step size on the basis of the angle obtained by the previous rotation, so that the rotation transformation does not start from the initial position every time; it should be emphasized that each matching process of the correlation scan matching of the map includes a rotation transformation and a coordinate system transformation, and the foregoing transformations are performed based on the pre-expanded grid map. Figure 2 As shown in the figure, the existing point cloud is read from the square box marked with Point Cloud; then it controls the read point cloud to perform a rotation transformation to achieve the coverage of the angular search range of the search window by the point cloud according to the preset angular search step size, and then updates the result of the rotation transformation to a batch of currently processed point clouds, and then controls the updated point cloud to perform the next rotation transformation. Therefore, in this embodiment, the rotation transformation of the currently read point cloud is optimized to a step-by-step rotation, ensuring that the angle of each rotation is obtained by rotating another preset angular search step size on the basis of the angle obtained by the previous rotation, so that the rotation transformation does not start from the initial position every time; it should be emphasized that each matching process of the correlation scan matching of the map includes a rotation transformation and a coordinate system transformation, and the foregoing transformations are performed based on the pre-expanded grid map.

[0088] The point cloud processing module is used to read the result of the aforementioned rotation transformation (the latest batch of point clouds being currently processed) under the control of the state machine control module, and control the read point clouds to perform coordinate system transformation, including the offset transformation of the horizontal and vertical coordinate values and the scale adjustment of the map resolution, so as to achieve coordinate transformation first and then discretization processing, and finally obtain the discretization result as the coordinate system transformation result defined in this embodiment. Then, the result of the coordinate system transformation is set as the discrete point cloud, which falls into the coordinate system of the pre-expanded grid map, so as to align the currently stored point clouds into the coordinate system of the pre-expanded grid map; then, the discrete point clouds are stored in the memory module through the interconnection bus; wherein, the pre-expanded grid map is the first map to be searched for one layer required for searching the root node in step C, or the first map to be searched for one layer required for searching the sub-node in step J; wherein, the currently stored point clouds, that is, the point clouds stored in the box marked with Point Cloud, are sourced from the lidar collection. It should be noted that, in the aforementioned grid map, the value in each grid represents the occupancy probability of that grid; the alignment of the point cloud with the pre-expanded grid map is understood as: the process of aligning the point cloud used to represent obstacles scanned by the lidar with the obstacle grids in the pre-expanded grid map. Preferably, the point cloud processing unit calls the associated computing unit to perform rotation and translation to achieve the coincidence of the point cloud and the obstacles.

[0089] The node search module is used to obtain the discrete point cloud from the memory module under the control of the state machine control module, and then calculate the index value of the discrete point cloud mapped to the map storage space according to the preset coordinate offset value and search step size, match the rotation and translation operations passed by the pre-expanded grid map (the first map to be searched for the current branch operation), and set the index value as the read address of the occupancy probability of the point cloud; wherein, the search step size is the number of grid searches that need to be traversed for each pose search within the same search window.

[0090] The node search module is also used to, under the control of the state machine control module, read out the occupancy probability of the corresponding point cloud from the map storage space according to the currently set reading address, call the accumulator built in the node search module to accumulate the occupancy probability of the currently read point cloud, and then output the accumulated result and set it as the score value of the corresponding node in the pre-expanded grid map, which is used to represent the probability sum corresponding to a node currently searched by the search window; where a node currently searched by the search window corresponds to: the grid point obtained by transforming the currently read point cloud based on the pose parameters of the search window. Each time a coordinate system transformation is performed, the node search module performs an accumulation operation once to obtain the probability sum corresponding to a node currently searched by the search window, which is used as the score value of this node, that is, the score value of the corresponding grid point (the grid point determined by the pose parameters included in this node) in the first map to be searched.

[0091] For each pose in the search window within the pre-expanded grid map, the correlation scan matching hardware circuit uses a pipeline parallel architecture to calculate the score value. Specifically, the calculation of the score value is split into: performing a coordinate system transformation on the read point cloud according to the rotation angle, initial position, and coordinate offset, discretizing the point cloud into the grid coordinate system by the resolution of the grid map and the maximum size boundary of the grid map, converting the discrete point cloud coordinates into index values in the map storage space, reading the occupancy probability of the point cloud in the map storage space by the index values, and cyclically accumulating the occupancy probability to obtain the score value of the node searched under the corresponding rotation transformation and coordinate system transformation, so that the calculation of the score value is obtained based on the parameters required by the correlation scan matching algorithm.

[0092] On the other hand, the point cloud processing module is designed as a pipeline structure for coordinate system transformation, and the node search module is designed as a pipeline structure for indexing and calculating score values, realizing the parallel execution of the process of scheduling the point cloud processing module to read the latest point cloud and the process of the point cloud processing module performing coordinate system transformation on the original point cloud at a specific clock cycle, or scheduling the coordinate system transformation of the point cloud processing module to be parallel with the index value calculation of the node search module, or scheduling the index value calculation of the node search module to be parallel with the accumulation calculation of the occupancy probability of the node search module. It also ensures that the score value result is shared on the interconnection bus for the CPU to read for sorting.

[0093] A memory module is provided with a map storage space for storing a first map to be searched transmitted by the bus interface; there is a corresponding index value for the occupancy probability of the point cloud in each layer of the first map to be searched; those skilled in the art can easily understand, based on the raster map constructed by the lidar point cloud, that the occupancy probability of the point cloud is an index value corresponding to the occupancy probability of the point cloud in a corresponding layer of the first map to be searched, and the occupancy probability of the point cloud is also the probability value that the corresponding grid point is occupied after the point cloud is converted into a grid in the pre-expanded raster map through the point cloud processing module.

[0094] The map storage space disclosed in this embodiment stores the position information of each grid point in each layer of the first map to be searched and the probability that the grid point is occupied by the point cloud, that is, the occupancy probability; therefore, the index value is set as the read address of the occupancy probability of the point cloud in a corresponding layer of the first map to be searched; the memory module is further configured to store the currently stored point cloud and the preset trigonometric function values, and the preset trigonometric function values are stored in Figure 2 the box marked with Trigonometric as shown; wherein, the currently stored point cloud is a circle of point clouds collected by the lidar; each type of data is stored in different block storage spaces of the same memory module; wherein, the number of rotation transformations is equal to the number of the foregoing root nodes, and one or more coordinate system transformations are performed corresponding to one rotation transformation; a node currently searched by the search window is determined by a specific rotation parameter and a specific translation parameter, and both the specific rotation parameter and the specific translation parameter belong to the parameters required for performing the correlation scan matching algorithm; the specific rotation parameter determines one rotation transformation, and the specific translation parameter determines one coordinate system transformation performed corresponding to one rotation transformation to obtain the occupancy probability of the pose matching in a corresponding layer of the first map to be searched where the node is located.

[0095] It should be noted that the lidar can be installed on a mobile robot. The laser probe of the lidar keeps collecting laser point cloud frame data during rotation, which is recorded as point cloud. The scanning range of the lidar can be centered on the body center of the robot or the center of the lidar. In this embodiment, every time the lidar rotates one circle, the point cloud collected in this circle is sent to the correlation scanning matching hardware circuit to cooperate with each branch operation executed by the CPU, and a node corresponding to a pose parameter that meets the search window is searched in each branch to calculate the score value corresponding to the grid point of a layer of the first map to be searched (corresponding to a layer of the first map to be searched) where the node is located. It should be noted that there is a specific coordinate transformation relationship between the world coordinate system and the two-dimensional grid map coordinate system, and based on this specific coordinate transformation relationship, the possible position of the body center of the robot in the two-dimensional grid map coordinate system is calculated, corresponding to a node in a specific pose.

[0096] In the correlation scanning matching hardware circuit, each computing unit and the aforementioned state machine control module can be, but are not limited to, a digital circuit module compiled by the designer using the hardware description language Verilog HDL, or a digital circuit module drawn or compiled by the designer on software with circuit drawing or compilation functions. In addition, in each embodiment of the present invention, each functional unit can be integrated in a processing module, or each unit can exist physically alone, or two or more units can be integrated in one module.

[0097] It should be noted that the memory module in the above embodiments is essentially a storage medium, and the storage medium can be, but is not limited to, various storage media that can store program codes such as read-only memory (ROM, Read-Only Memory), random access memory (RAM, RandomAccess Memory), etc.

[0098] As an embodiment, such as Figure 2As shown, the map storage space includes a first map storage space and a second map storage space; the first map storage space is used to store the first searchable map with the first pre - set resolution, and the second map storage space is used to store the first searchable map with the second pre - set resolution. The depth level of the first searchable map with the first pre - set resolution is one level different from the depth level of the first searchable map with the second pre - set resolution, so that the two layers of the first searchable maps are adjacent, including but not limited to: the depth of the first searchable map with the first pre - set resolution is greater than the depth of the first searchable map with the second pre - set resolution; among them, the first pre - set resolution is less than the second pre - set resolution. The correlation scan - matching hardware circuit is used to, during the process of executing a branch operation once, schedule the point - cloud processing module to align and match the currently stored point cloud to a layer of the first searchable map with the second pre - set resolution, and at the same time, schedule the node search module to calculate the index value of the occupancy probability of the currently stored point cloud in the first map storage space, which is equivalent to obtaining the index value of the occupancy probability of the currently stored point cloud in the first map storage space while keeping the same batch of point clouds aligned and matched to the first searchable map with the second pre - set resolution. Thus, by using the parallel pipeline structure of the point - cloud processing module and the node search module, search for a layer of the first searchable map with the first resolution and convert the point cloud to a layer of the first searchable map with the second resolution in the same clock cycle, giving play to the advantages of the parallel pipeline of the correlation scan - matching hardware circuit and accelerating the calculation of the score value of the currently searched node.

[0099] Preferably, the first pre - set number of layers is 3, and the second pre - set number of layers is 4; during the process of the CPU executing step B, the CPU transmits the first layer of the first searchable map to the correlation scan - matching hardware circuit through the bus interface, and then the bus interface writes the first layer of the first searchable map into the first map storage space. Among them, all root nodes only have corresponding grid points in the first layer of the first searchable map, and the depth of all root nodes is equal to the depth of the first layer of the first searchable map; from the first layer of the first searchable map to the third layer of the first searchable map, the resolution (abbreviation of map resolution) increases layer by layer; from the third layer of the first searchable map to the first layer of the second searchable map, the resolution (abbreviation of map resolution) increases layer by layer; from the first layer of the second searchable map to the fourth layer of the second searchable map, the resolution (abbreviation of map resolution) increases layer by layer; from the first layer of the first searchable map to the fourth layer of the second searchable map, they are arranged vertically from top to bottom to construct a branch - and - bound search tree with consistent depth values.

[0100] When the CPU executes step D, the CPU first transmits the first search map of the second layer to the correlation scanning and matching hardware circuit, and then transmits the first search map of the third layer to the correlation scanning and matching hardware circuit. Then, the bus interface writes the first search map of the second layer into the first map storage space and overwrites the first search map of the first layer; then the bus interface writes the first search map of the third layer into the second map storage space; among them, before completely overwriting the first search map of the first layer, the score values of all root nodes have been obtained. It can be understood that the process of overwriting the first search map of the first layer is also the process of continuously searching for the remaining root nodes and calculating their score values. In this embodiment, the most blurred first search map of three layers is constructed for the correlation scanning and matching hardware circuit to search, while leaving the second search map with higher resolution of four layers for the CPU to process. However, in this technical solution, it is not necessary to specifically design three map storage spaces in the memory module of the correlation scanning and matching hardware circuit, but only two memories are designed, and the first search map of three layers is alternately stored under the scheduling of the state machine control module, so that after calculating the score values of all root nodes, the first search map of the second layer overwrites the first search map of the first layer, and then the first search map of the third layer overwrites the first search map of the second layer. This saves the consumption of memory resources.

[0101] As an embodiment, as Figure 2 shown, the node search module further includes a grid indexing sub-module. The grid indexing sub-module includes a first index subtractor, a second index subtractor, a first index adder, and a second index adder, which are used to construct a pipeline structure; among them, both the first index subtractor and the second index subtractor belong to subtractors, and in Figure 2 the grid indexing sub-module are circles marked with "-"; both the first index adder and the second index adder belong to adders, and in Figure 2 the grid indexing sub-module are circles marked with "+".

[0102] The first input end of the first index subtractor is used to receive the abscissa of the discrete point cloud transmitted by the interconnection bus. Among them, the abscissa of the discrete point cloud transmitted by the interconnection bus to the first index subtractor comes from the memory module; preferably, the abscissa of the discrete point cloud transmitted by the interconnection bus supports being stored in the register inside the grid indexing sub-module, and the first input end of the first index subtractor receives the abscissa of the discrete point cloud transmitted by the interconnection bus through the connected register. In Figure 2 the shown embodiment, the first input end of the first index subtractor (in Figure 2 the grid indexing sub-module, the first subtractor arranged from top to bottom) is connected to the register DP x (marked with DPx is connected to the square frame), where the register DP x is used to cache the abscissa of the discrete point cloud transmitted by the interconnection bus, and then transmit the abscissa of the currently cached discrete point cloud to the first input end of the first index subtractor, so that the first input end of the first index subtractor completes receiving the abscissa of the discrete point cloud transmitted by the interconnection bus; among them, the discrete point cloud transmitted by the interconnection bus to the first index subtractor is from the memory module.

[0103] The second input end of the first index subtractor is used to receive the horizontal axis coordinate offset value transmitted by the interconnection bus. Preferably, the horizontal axis coordinate offset value transmitted by the interconnection bus supports being stored in the register inside the grid index sub-module. The second input end of the first index subtractor receives the horizontal axis coordinate offset value transmitted by the interconnection bus through the connected register. In Figure 2 In the shown embodiment, the second input end of the first index subtractor is connected to the register Offset x (marked with Offset x is connected to the square frame), where the register Offset x is used to cache the horizontal axis coordinate offset value transmitted by the interconnection bus, and then transmit the currently cached horizontal axis coordinate offset value to the second input end of the first index subtractor, so that the second input end of the first index subtractor completes receiving the horizontal axis coordinate offset value transmitted by the interconnection bus. Among them, the preset coordinate offset value includes the horizontal axis coordinate offset value, and the horizontal axis coordinate offset value transmitted by the interconnection bus is from the preset coordinate offset value stored in the map offset value register; the map offset value register is a parameter register set inside the correlation scan matching hardware circuit, and is used to store the coordinate offset value associated with the pre-expanded grid map; Figure 1 The shown parameter register set includes a map offset value register. The first index subtractor is used to control the subtraction of the abscissa of the discrete point cloud received at the first input end from the horizontal axis coordinate offset value received at the second input end, and then output the difference value.

[0104] The first input end of the first index adder is connected to the output end of the first index subtractor to construct a combinational logic circuit; the second input end of the first index adder is used to receive the horizontal axis coordinate search step transmitted by the interconnection bus; preferably, the horizontal axis coordinate search step transmitted by the interconnection bus supports being stored in the register inside the grid index sub-module. The second input end of the first index adder receives the horizontal axis coordinate search step transmitted by the interconnection bus through the connected register. In Figure 2 In the shown embodiment, the second input end of the first index adder is connected to the register Step x (marked with Stepx is connected to the box), where the register Step x is used to cache the horizontal axis coordinate search step transmitted by the interconnection bus, and then transmit the currently cached horizontal axis coordinate search step to the second input end of the first index adder. Among them, the search step includes the horizontal axis coordinate search step, and the horizontal axis coordinate search step transmitted by the interconnection bus is derived from the search step stored in the search window parameter register; the search window parameter register is a parameter register set inside the correlation scan matching hardware circuit, and is used to store the positions of the nodes existing in the search window and the search information of the associated child nodes; Figure 1 The parameter register group shown includes a search window parameter register. The first index adder is used to control the addition of the difference between the horizontal axis coordinate search step and the output of the first index subtractor, and then output the sum obtained by the addition. The sum value output by the first index adder is configured as the horizontal axis direction index value of the discrete point cloud mapped to the map storage space.

[0105] The first input end of the second index subtractor is used to receive the ordinate of the discrete point cloud transmitted by the interconnection bus. Preferably, the ordinate of the discrete point cloud transmitted by the interconnection bus supports being stored in the register inside the grid index sub-module. The first input end of the second index subtractor receives the ordinate of the discrete point cloud transmitted by the interconnection bus through the connected register. In Figure 2 In the shown embodiment, the first input end (in Figure 2 inside the grid index sub-module of, the second subtractor arranged from top to bottom) is connected to the output end of the register DP y (marked with DP y of the box), where the register DP y is used to cache the ordinate of the discrete point cloud transmitted by the interconnection bus, and then transmit the currently cached ordinate of the discrete point cloud to the first input end of the second index subtractor, so that the first input end of the second index subtractor completes receiving the ordinate of the discrete point cloud transmitted by the interconnection bus; among them, the ordinate of the discrete point cloud transmitted by the interconnection bus to the second index subtractor is derived from the memory module.

[0106] The second input end of the second index subtractor is used to receive the vertical axis coordinate offset value transmitted by the interconnection bus. Preferably, the vertical axis coordinate offset value transmitted by the interconnection bus supports being stored in the register inside the grid index sub-module. The second input end of the second index subtractor receives the vertical axis coordinate offset value transmitted by the interconnection bus through the connected register. In Figure 2 In the shown embodiment, the second input end of the second index subtractor is connected to the register Offset y(The box marked with Offset y ) is connected, where the register Offset y is used to cache the vertical axis coordinate offset value transmitted by the interconnection bus, and then transmit the currently cached vertical axis coordinate offset value to the second input terminal of the second index subtractor; where the preset coordinate offset value further includes a vertical axis coordinate offset value, and the vertical axis coordinate offset value transmitted by the interconnection bus also comes from the preset coordinate offset value stored in the map offset value register. The second index subtractor is used to control the subtraction of the ordinate of the discrete point cloud received at the first input terminal from the vertical axis coordinate offset value received at the second input terminal, and then output the difference value.

[0107] The first input terminal of the second index adder is connected to the output terminal of the second index subtractor to construct a combinational logic circuit; the second input terminal of the second index adder is used to receive the vertical axis coordinate search step length transmitted by the interconnection bus; preferably, the vertical axis coordinate search step length transmitted by the interconnection bus supports being stored in a register inside the grid index sub-module, and the second input terminal of the second index adder receives the vertical axis coordinate search step length transmitted by the interconnection bus through the connected register. In Figure 2 the illustrated embodiment, the second input terminal of the second index adder is connected to the register Step y (the box marked with Step y ), where the register Step y is used to cache the vertical axis coordinate search step length transmitted by the interconnection bus, and then transmit the currently cached vertical axis coordinate search step length to the second input terminal of the second index adder. The search step length further includes a vertical axis coordinate search step length, and the vertical axis coordinate search step length transmitted by the interconnection bus also comes from the search step length stored in the search window parameter register; where the search step length further includes a vertical axis coordinate search step length, and the vertical axis coordinate search step length transmitted by the interconnection bus also comes from the search step length stored in the search window parameter register; the second index adder is used to control the addition of the vertical axis coordinate search step length and the difference value output by the second index subtractor, and then output the sum value obtained by the addition. The sum value output by the second index adder is configured as the vertical axis direction index value for mapping the discrete point cloud to the map storage space.

[0108] The node search module uses two parallel addition and subtraction operation combination structures to respectively calculate the horizontal axis direction index value and the vertical axis direction index value mapped into the map storage space, and realizes the parallel conversion of the abscissa and ordinate of the discrete point cloud into the discrete index information of a layer of the map to be searched currently participating in the search.

[0109] As a preferred example, the grid index sub-module further includes a third index adder and an index multiplier. The third index adder belongs to the adder and is the first adder arranged from left to right within the grid index sub-module corresponding to Figure 2 ; the index multiplier belongs to the multiplier and is the only circle marked with "×" in the grid index sub-module of Figure 2.

[0110] The first input end of the index multiplier is connected to the output end of the second index adder to construct a combinational logic circuit. The first input end of the index multiplier is used to receive the sum value output by the second index adder, and the second input end of the index multiplier is used to receive the number of row grids transmitted by the interconnection bus. Preferably, the number of row grids transmitted by the interconnection bus supports being stored in the register inside the grid index sub-module, and the second input end of the index multiplier receives the number of row grids transmitted by the interconnection bus through the connected register. In Figure 2 the illustrated embodiment, the second input end of the index multiplier is connected to the register NUM x (the box marked with NUM x ), where the register NUM x is used to cache the number of row grids transmitted by the interconnection bus and then transmit the currently cached number of row grids to the second input end of the index multiplier, so that the second input end of the index multiplier completes receiving the number of row grids transmitted by the interconnection bus; wherein, the number of row grids is the number of grids that each row of the map storage space can occupy, is stored in the map size register, and is transmitted from the map size register to the interconnection bus; Figure 1 the illustrated parameter register group includes the map size register. The map size register is a parameter register set inside the correlation scan matching hardware circuit and is used to store the size range of the pre-expanded grid map and the associated expansion information; the index multiplier is used to control the multiplication of the number of row grids and the sum value output by the second index adder, and then output the product.

[0111] The first input end of the third index adder is connected to the output end of the first index adder. The first input end of the third index adder is used to receive the horizontal axis direction index value output by the first index adder, and the second input end of the third index adder is used to receive the product output by the index multiplier to construct a combinational logic circuit; the third index adder is used to control the addition of the product of the number of row grids and the sum value output by the second index adder and the horizontal axis direction index value, and then set the obtained sum value as the index value of the discrete point cloud mapped into the map storage space, and send it to the interconnection bus, so as to facilitate reading out the occupancy probability matching the index value from the memory module.

[0112] In this preferred example, the index multiplier controls the multiplication of the number of row grids and the index value in the vertical axis direction, and then the third index adder controls the addition of the index value in the horizontal axis direction and the product output by the index multiplier. The index value of the discrete point cloud mapped to the occupancy probability is calculated in the form of row scanning query, which is also used as the index value mapped to the map storage space, equivalent to the storage address of the occupancy probability of the point cloud in the map storage space, and equivalent to the read address for externally reading the occupancy probability stored in the map storage space, so as to facilitate the interconnection bus to read out the occupancy probability from the memory module.

[0113] As another preferred example, the grid index sub-module further includes a third index adder and an index multiplier. The third index adder belongs to the adder, and the index multiplier belongs to the multiplier. The first input end of the index multiplier is connected to the output end of the first index adder. The second input end of the index multiplier is used to receive the number of column grids transmitted by the interconnection bus. The number of column grids is the number of grids that each column of the map storage space can occupy, stored in the map size register and transmitted to the interconnection bus by the map size register. The map size register is a parameter register set inside the correlation scanning and matching hardware circuit, used to store the size range of the first map to be searched and related extended information. The first input end of the third index adder is connected to the output end of the second index adder, and the second input end of the third index adder is connected to the output end of the index multiplier. The third index adder is used to control the addition of the product of the sum of the number of column grids and the output of the first index adder and the index value in the vertical axis direction, and then set the obtained sum value as the index value of the discrete point cloud mapped to the map storage space, and send it to the interconnection bus, so as to facilitate reading out the occupancy probability matching the index value from the memory module. In this preferred example, the index multiplier controls the multiplication of the number of column grids and the index value in the horizontal axis direction, and then the third index adder controls the addition of the index value in the vertical axis direction and the product output by the index multiplier. The index value of the discrete point cloud mapped to the occupancy probability is calculated in the form of column scanning query, which is also used as the index value mapped to the map storage space, equivalent to the storage address of the occupancy probability of the point cloud in the map storage space, and equivalent to the read address for externally reading the occupancy probability stored in the map storage space, so as to facilitate the interconnection bus to read out the occupancy probability from the memory module.

[0114] As an embodiment, such as Figure 2As shown, the point cloud processing module includes a point cloud rotation sub-module and a point cloud discretization sub-module; the point cloud rotation sub-module includes a first register, a second register, a first point cloud multiplier, a second point cloud multiplier, a third point cloud multiplier, a fourth point cloud multiplier, a point cloud adder, and a point cloud subtractor, which are configured as a pipeline structure; the first point cloud multiplier, the second point cloud multiplier, the third point cloud multiplier, and the fourth point cloud multiplier all belong to multipliers, and are represented as Figure 2 in the point cloud rotation sub-module, the circles marked with "×"; the point cloud adder belongs to an adder, and is represented as Figure 2 in the point cloud rotation sub-module, the circles marked with "+"; the point cloud subtractor belongs to a subtractor, and is represented as Figure 2 in the point cloud rotation sub-module, the circles marked with "-".

[0115] The first input end of the first point cloud multiplier is connected to the output end of the first register; among them, the interconnection bus transmits the abscissa of the point cloud to the input end of the first register; the abscissa of the point cloud transmitted by the interconnection bus to the first register is derived from the abscissa of the currently stored point cloud; in Figure 2 the shown embodiment, the first input end of the first point cloud multiplier (in Figure 2 the point cloud rotation sub-module, the first multiplier arranged from top to bottom) is connected to the output end of register P x (the box marked with P x ), where register P x is the first register.

[0116] The second input end of the first point cloud multiplier is used to receive the cosine function value corresponding to the rotation angle reached by the current rotation transformation. Among them, the rotation angle reached by the current rotation transformation is the sum value of the angle search step and the rotation angle reached by the previous rotation transformation, and is derived from the rotation angle and its corresponding cosine function value obtained by the step rotation operation executed at the CPU or software level, so as to avoid increasing the complexity of the hardware unit due to trigonometric function operations and prevent hardware transmission delays. Preferably, the cosine function value transmitted by the interconnection bus supports being stored in the register inside the point cloud rotation sub-module, and the second input end of the first point cloud multiplier receives the cosine function value transmitted by the interconnection bus through the connected register. In Figure 2In the illustrated embodiment, the second input terminal of the first point cloud multiplier is connected to the output terminal of the register COS (the box marked with COS), where the register COS is used to cache the cosine function value transmitted by the interconnection bus, and then transmit the currently cached cosine function value to the second input terminal of the first point cloud multiplier. The first point cloud multiplier is used to control the multiplication of the abscissa of the point cloud received by the first input terminal of the first point cloud multiplier and the cosine function value corresponding to the rotation angle reached by the current rotation transformation, and then output the product value obtained by the multiplication. The angle search step size is stored in the search window parameter register; the search window parameter register is a parameter register set inside the correlation scan matching hardware circuit, and is used to store the poses and associated step search information included in the search window.

[0117] The first input terminal of the second point cloud multiplier is connected to the output terminal of the second register; wherein, the interconnection bus caches the ordinate of the point cloud into the second register; the ordinate of the point cloud transmitted by the interconnection bus to the second register is derived from the ordinate of the currently stored point cloud; Figure 2 In the illustrated embodiment, the first input terminal of the second point cloud multiplier (the second multiplier arranged from top to bottom in the point cloud rotation sub-module) Figure 2 is connected to the output terminal of the register P y (the box marked with P y ), where the register P y is the second register.

[0118] The second input terminal of the second point cloud multiplier is used to receive the sine function value corresponding to the rotation angle reached by the current rotation transformation; wherein, the rotation angle reached by the current rotation transformation is the sum value of the angle search step size and the rotation angle reached by the previous rotation transformation, and is the rotation angle and its corresponding sine function value obtained from the step rotation operation performed at the CPU or software level, so as to avoid increasing the complexity of the hardware unit due to trigonometric function operations and prevent hardware transmission delays. Preferably, the sine function value transmitted by the interconnection bus supports being registered into the register inside the point cloud rotation sub-module, and the second input terminal of the second point cloud multiplier receives the sine function value transmitted by the interconnection bus through the connected register. Figure 2In the illustrated embodiment, the second input terminal of the second point cloud multiplier is connected to the output terminal of the register SIN (the box marked with SIN), where the register SIN is used to cache the sine function value transmitted by the interconnect bus, and then transmit the currently cached sine function value to the second input terminal of the second point cloud multiplier. The second point cloud multiplier is used to control the multiplication of the ordinate of the point cloud received by the first input terminal of the second point cloud multiplier and the sine function value corresponding to the rotation angle achieved by the current rotation transformation, and then output the product value obtained by the multiplication.

[0119] The first input terminal of the third point cloud multiplier is connected to the output terminal of the first register and is used to receive the abscissa of the point cloud output by the first register. Figure 2 In the illustrated embodiment, the first input terminal of the third point cloud multiplier (the third multiplier arranged from top to bottom in the point cloud rotation sub-module) is connected to an output terminal of the register P Figure 2 (the box marked with P x ( x )

[0120] The second input terminal of the third point cloud multiplier is used to receive the sine function value corresponding to the rotation angle achieved by the current rotation transformation; preferably, the sine function value transmitted by the interconnect bus supports being stored in the register inside the point cloud rotation sub-module, and the second input terminal of the second point cloud multiplier receives the sine function value transmitted by the interconnect bus through the connected register. In Figure 2 the illustrated embodiment, the second input terminal of the second point cloud multiplier is connected to the output terminal of the register SIN (the box marked with SIN), where the register SIN is used to cache the sine function value transmitted by the interconnect bus, and then transmit the currently cached sine function value to the second input terminal of the third point cloud multiplier. The third point cloud multiplier is used to control the multiplication of the ordinate of the point cloud currently transmitted by the interconnect bus and the sine function value corresponding to the rotation angle achieved by the current rotation transformation, and then output the product value obtained by the multiplication. The third point cloud multiplier is used to control the multiplication of the abscissa of the point cloud received by the first input terminal of the third point cloud multiplier and the sine function value corresponding to the rotation angle achieved by the current rotation transformation, and then output the product value obtained by the multiplication.

[0121] The first input terminal of the fourth point cloud multiplier is connected to the output terminal of the second register and is used to receive the ordinate of the point cloud output by the second register. Figure 2 In the illustrated embodiment, the first input terminal of the fourth point cloud multiplier (in Figure 2in the point cloud rotation sub-module, the fourth multiplier arranged from top to bottom) and register P y (the box marked with P y is connected to an output terminal of).

[0122] The second input terminal of the fourth point cloud multiplier is used to receive the cosine function value corresponding to the rotation angle reached by the current rotation transformation; in Figure 2 the illustrated embodiment, the second input terminal of the fourth point cloud multiplier is connected to the output terminal of register COS (the box marked with COS), where register COS is used to cache the cosine function value transmitted by the interconnection bus and then transmit the currently cached cosine function value to the second input terminal of the fourth point cloud multiplier. The fourth point cloud multiplier is used to control the multiplication of the ordinate of the point cloud received by the first input terminal of the fourth point cloud multiplier and the cosine function value corresponding to the rotation angle reached by the current rotation transformation received by the second input terminal of the fourth point cloud multiplier, and then output the product value obtained by the multiplication.

[0123] The first input terminal of the point cloud subtractor is connected to the output terminal of the first point cloud multiplier, and the second input terminal of the point cloud subtractor is connected to the output terminal of the second point cloud multiplier to construct a combinational logic circuit; the point cloud subtractor is used to control the subtraction of the product value output by the first point cloud multiplier and the product value output by the second point cloud multiplier, and then output the difference obtained by the subtraction, and set it as the abscissa value of the rotated point cloud obtained after the point cloud undergoes the current rotation transformation; the point cloud subtractor is also used to transmit the output difference to the point cloud discretization sub-module and determine to update the abscissa result of the current rotation transformation to the abscissa of the currently stored point cloud for the next execution of the rotation transformation.

[0124] The first input terminal of the point cloud adder is connected to the output terminal of the third point cloud multiplier, and the second input terminal of the point cloud adder is connected to the output terminal of the fourth point cloud multiplier; the first input terminal of the point cloud adder is used to receive the product value output by the third point cloud multiplier; the second input terminal of the point cloud adder is used to receive the product value output by the fourth point cloud multiplier; the point cloud adder is used to control the addition of the product value output by the third point cloud multiplier and the product value output by the fourth point cloud multiplier, and then output the sum value obtained by the addition, and set it as the ordinate value of the rotated point cloud obtained after the point cloud undergoes the current rotation transformation; the point cloud adder is also used to transmit the output difference to the point cloud discretization sub-module and determine to update the ordinate result of the current rotation transformation to the ordinate of the currently stored point cloud for the next execution of the rotation transformation.

[0125] It should be noted that the angular search step size is stored in the search window parameter register. The search window parameter register is a parameter register set inside the correlation scan matching hardware circuit, and is used to store the poses covered by the search window and the associated step-by-step search information; Figure 2 The parameter register set shown includes the search window parameter register. The number of rotations is determined by the angular search range of the pre-configured search window and the pre-configured angular search step size, and both are related to the number of the aforementioned root nodes, so that all poses within the search window can be traversed (queried) within the corresponding number of rotations. Each rotation transformation corresponds to a coordinate system transformation executed by the point cloud processing module, which belongs to a matching process of the correlation scan matching of the map, and a score value is obtained by accumulation.

[0126] In this embodiment, with the rotation matrix as the basic operation architecture, only adders, subtracters and multipliers are used to construct the point cloud rotation sub-module, which is used to receive and process the trigonometric function values corresponding to the step-by-step rotation angles pre-calculated by software. It not only eliminates the trigonometric function hardware calculation module, but also controls the point cloud rotation sub-module to iteratively execute the rotation transformation on the basis of following the step-by-step rotation, obtains the point cloud data after the rotation transformation, and provides it to the point cloud discretization sub-module to continue the coordinate system transformation, determines that the traversal of the rotation angle currently searched by the search window is completed, so as to accelerate the accumulation calculation operation of the score value.

[0127] Compared with the rotation based on the initial position, the method of this module in this paper avoids the hardware trigonometric function calculation without increasing the extra storage space of the hardware and without losing the calculation accuracy, and makes full use of the update transformation operation of the point cloud when scanning and matching to a new layer of the map to obtain the rotation angle obtained under the step-by-step rotation.

[0128] As an embodiment, as Figure 2 shown, the point cloud processing module includes a point cloud discretization sub-module; the point cloud discretization sub-module includes a first discrete adder, a first discrete subtracter, a first discrete multiplier, a second discrete adder, a second discrete subtracter and a second discrete multiplier, and is used to be constructed into a pipeline structure; both the first discrete adder and the second discrete adder belong to adders, and are marked as circles with a "+" in the Figure 2 point cloud discretization sub-module; both the first discrete subtracter and the second discrete subtracter belong to subtracters, and are shown as circles with a "-" in the Figure 2 point cloud discretization sub-module; both the first discrete multiplier and the second discrete multiplier belong to multipliers, and are shown as Figure 2 circles with a "×" in the point cloud discretization sub-module.

[0129] The first input terminal of the first discrete adder is connected to the output terminal of the point cloud subtractor. The first input terminal of the first discrete adder is used to receive the abscissa result obtained from the current rotation transformation.

[0130] The second input terminal of the first discrete adder is used to receive the abscissa of the robot position transmitted by the interconnection bus. The abscissa of the robot position is the abscissa value of the horizontal axis of the robot in the world coordinate system calculated in advance. Preferably, the abscissa of the robot position transmitted by the interconnection bus supports being stored in the register inside the point cloud discretization sub-module. The second input terminal of the first discrete adder receives the abscissa of the robot position transmitted by the interconnection bus through the connected register. In Figure 2 the illustrated embodiment, the second input terminal of the first discrete adder is connected to the register Pose x (the box marked with Pose x ). The register Pose x is used to cache the abscissa of the robot position transmitted by the interconnection bus, and then transmit the currently cached abscissa of the robot position to the second input terminal of the first discrete adder. The abscissa of the robot position is the abscissa value of the horizontal axis of the robot in the world coordinate system calculated in advance (by software or the CPU unit). The lidar is mounted on the robot and serves as a sensor device for robot positioning. The first discrete adder is used to control the addition of the abscissa result obtained from the current rotation transformation (the output result of the point cloud subtractor) and the abscissa of the robot position, and then output the sum obtained by the addition, so that the sum value becomes the abscissa value of the horizontal axis of the point cloud in the world coordinate system.

[0131] The first input terminal of the first discrete subtractor is connected to the output terminal of the first discrete adder. The second input terminal of the first discrete subtractor is used to receive the maximum abscissa value of the map transmitted by the interconnection bus. Preferably, the maximum abscissa value of the map transmitted by the interconnection bus supports being stored in the register inside the point cloud discretization sub-module. The second input terminal of the first discrete subtractor receives the maximum abscissa value of the map transmitted by the interconnection bus through the connected register. In Figure 2 the illustrated embodiment, the second input terminal of the first discrete subtractor is connected to the register MAX x (the box marked with MAX x ). The register MAX xUsed to cache the maximum abscissa value of the map transmitted by the interconnect bus, and then transmit the currently cached maximum abscissa value of the map to the second input terminal of the first discrete subtractor. The maximum abscissa value of the map is derived from the map size register and is transmitted by the map size register to the interconnect bus; wherein, the map size register is a parameter register set inside the correlation scan matching hardware circuit, and is used to store the size range of the grid map that meets the transmission requirements of the bus width; Figure 2 The parameter register group shown includes a map size register. The first discrete subtractor is used to control the subtraction of the maximum abscissa value of the map from the sum value output by the first discrete adder, and then output the obtained difference value. The difference value output by the first discrete subtractor becomes the abscissa value of the grid map that meets the transmission requirements of the bus width.

[0132] The first input terminal of the first discrete multiplier is connected to the output terminal of the first discrete subtractor and is used to receive the difference value output by the first discrete subtractor. In this embodiment, the first input terminal of the first discrete multiplier is connected to the output terminal of the first discrete subtractor to construct a combinational logic circuit; the second input terminal of the first discrete multiplier is used to receive the reciprocal of the map resolution transmitted by the interconnect bus. Preferably, the reciprocal of the map resolution transmitted by the interconnect bus supports being stored in the register inside the point cloud discretization sub-module. The second input terminal of the first discrete multiplier receives the reciprocal of the map resolution transmitted by the interconnect bus through the connected register. In Figure 2 In the shown embodiment, the second input terminal of the first discrete multiplier is connected to the register 1 / R (the box marked with 1 / R). Among them, the register 1 / R is used to cache the reciprocal of the map resolution transmitted by the interconnect bus, and then transmit the currently cached reciprocal of the map resolution to the second input terminal of the first discrete multiplier, so that the second input terminal of the first discrete multiplier completes receiving the reciprocal of the map resolution transmitted by the interconnect bus; wherein, the reciprocal of the map resolution is derived from the map resolution register and is transmitted by the map resolution register to the interconnect bus; wherein, the map resolution register is a parameter register set inside the correlation scan matching hardware circuit and is used to store the resolution information of the first map to be searched for a corresponding layer; Figure 2The parameter register set shown includes a map resolution register. A first discrete multiplier is used to control the multiplication of the reciprocal of the map resolution and the difference output by the first discrete subtractor, and then output the product value obtained by the multiplication, and configure the product value as the abscissa value of the discrete point cloud to complete the coordinate system transformation, so as to align the abscissa value of the currently stored point cloud into the coordinate system of the corresponding first map to be searched; the first discrete multiplier is also used to output the product value obtained by the multiplication to the memory module through the interconnection bus, and determine that the product value is the abscissa that matches the currently aligned and matched node on the corresponding first map to be searched.

[0133] The first input end of the second discrete adder is connected to the output end of the point cloud adder, and the first input end of the second discrete adder is used to receive the ordinate result obtained by the current rotation transformation.

[0134] The second input end of the second discrete adder is used to receive the ordinate of the robot position transmitted by the interconnection bus, where the ordinate of the robot position is the longitudinal axis coordinate value of the robot in the world coordinate system calculated in advance; preferably, the ordinate of the robot position transmitted by the interconnection bus supports being stored in the register inside the point cloud discretization sub-module, and the second input end of the second discrete adder receives the ordinate of the robot position transmitted by the interconnection bus through the connected register. In Figure 2 the embodiment shown, the second input end of the second discrete adder is connected to the register Pose y (the box marked with Pose y ), where the register Pose y is used to cache the ordinate of the robot position transmitted by the interconnection bus, and then transmit the currently cached ordinate of the robot position to the second input end of the second discrete adder. The robot position register is used to store the ordinate and abscissa of the robot position, and the robot position register is a parameter register set inside the correlation scan matching hardware circuit; the second discrete adder is used to control the addition of the ordinate result obtained by the current rotation transformation (the output result of the point cloud adder) and the ordinate of the robot position, and then output the sum value obtained by the addition, so that the sum value becomes the longitudinal axis coordinate value of the point cloud transformed into the world coordinate system.

[0135] The first input of the second discrete subtractor is connected to the output of the second discrete adder; the second input of the second discrete subtractor is used to receive the maximum vertical coordinate value of the map transmitted from the interconnection bus; preferably, the maximum vertical coordinate value of the map transmitted from the interconnection bus supports storage in a register inside the point cloud discrete submodule, and the second input of the second discrete subtractor receives the maximum vertical coordinate value of the map transmitted from the interconnection bus through the connected register. Figure 2 In the embodiment shown, the second input terminal of the second discrete subtractor is connected to the register MAX y (Marked with MAX y The block) is connected, where register MAX y The second discrete subtractor is configured to cache the maximum vertical coordinate value of the map transmitted from the interconnect bus and then transmit the currently cached maximum vertical coordinate value of the map to the second input of the second discrete subtractor. The maximum vertical coordinate value of the map is derived from the map size register and transmitted to the interconnect bus from the map size register. The second discrete subtractor is configured to control the subtraction of the maximum vertical coordinate value of the map from the sum output by the second discrete adder and then output the difference. The difference output by the second discrete subtractor becomes the vertical coordinate value of the grid map that meets the bus bit width transmission requirements.

[0136] The first input of the second discrete multiplier is connected to the output of the second discrete subtractor, and is used to receive the difference output by the second discrete subtractor; the second input of the second discrete multiplier is used to receive the inverse of the map resolution transmitted from the interconnection bus; preferably, the inverse of the map resolution transmitted from the interconnection bus supports storage in a register inside the point cloud discrete submodule, and the second input of the second discrete multiplier receives the inverse of the map resolution transmitted from the interconnection bus through the connected register. The second input of the second discrete multiplier is connected to register 1 / R (the box marked with 1 / R); the second discrete multiplier is used to control the inverse of the map resolution and the second discrete

[0137] In summary, for each point cloud, this embodiment reads the horizontal coordinate and the vertical coordinate after the rotation transformation from the point cloud rotation submodule, and then uses an adder and a subtractor to transform the horizontal and vertical coordinates of the above point cloud along each adaptive coordinate axis direction in the same offset manner from the point cloud coordinate system to the coordinate system of the first layer of the map to be searched, including swapping the coordinate axes of the point cloud to achieve the coordinate system transformation; and then multiplying by the inverse of the grid map resolution to complete the discretization processing of the point cloud.

[0138] As an embodiment, the state machine control module belongs to a finite state machine; the state machine control module is used to schedule the working states of the memory module, the point cloud processing module, and the node search module, so that each time the point cloud processing module performs the coordinate system transformation and the rotation transformation once, the node search module calculates the score value of a pose within the search window each time to obtain the score value of the currently searched node. Among them, the working states specifically include the idle state, the point cloud reading state, the trigonometric function reading state, the multi-resolution map reading state, the search state, the loop state, the interrupt state, and the clear waiting state; the initial state of the correlation scan matching hardware circuit is the idle state.

[0139] The state machine control module is used to execute:

[0140] When the correlation scan matching hardware circuit receives the algorithm start signal, the working state of the state machine control module jumps to the point cloud reading state. At this time, the point cloud collected by the lidar is transmitted from the external read FIFO to the block storage space corresponding to the box marked with Point Cloud.

[0141] When all the lidar point coordinates are read, the working state jumps to the trigonometric function reading state. At this time, the pre-calculated trigonometric function values are transmitted from the external read FIFO to the block storage space corresponding to the box marked with Trigonometric.

[0142] When all the pre-calculated trigonometric function values are read, the working state jumps to the multi-resolution map reading state. At this time, the first map to be searched is transmitted from the external read FIFO to the first map storage space or the second map storage space.

[0143] In the multi-resolution map reading state, control the point cloud processing module to receive a layer of the first map to be searched; then switch to the search state, control the point cloud processing module to perform the aforementioned coordinate system transformation (including point cloud rotation and translation) on the currently stored point cloud to obtain the discrete point cloud, and then control the node search module to calculate the index value of the discrete point cloud mapped to the map storage space where the current layer of the first map to be searched is located, and use the currently calculated index value to search for the occupancy probability in the corresponding map storage space, and then control the aforementioned accumulator to perform an accumulative calculation on the searched occupancy probability to obtain the score value of the corresponding node, that is, the score value of the currently searched node.

[0144] After saving the currently calculated score value, it goes to the loop state. In this embodiment, when searching for the root node (specifically the first map to be searched in the first layer transmitted to the map storage space), the score value of the currently searched root node is written into the write data FIFO. After the writing to the register is completed, it indicates that the currently calculated score value has been saved. When searching for a map with a resolution higher than that of the first map to be searched in the first layer, the score value of the currently searched node is written into the register. After the writing to the register is completed, it indicates that the currently calculated score value has been saved.

[0145] In the loop state, if it is determined that the currently processed node is the root node and not all occupancy probabilities corresponding to the root nodes have been indexed, then it returns to the search state and continues to search for the occupancy probability in the corresponding map storage space using the newly calculated index value until all occupancy probabilities corresponding to the root nodes supported by the search window have been completed, and then it jumps to the interrupt state. At this time, it is determined whether an interrupt needs to be triggered according to the interrupt enable control signal. If the interrupt clear control signal is true, it enters the clear waiting state.

[0146] In the loop state, if it is determined that there are still nodes whose score values have not been calculated and the currently processed node is not the root node, then it returns to the search state and continues to control the aforementioned accumulator to perform cumulative calculation on the searched occupancy probability until the score values of all nodes supported by the search window have been calculated, and then it jumps to the interrupt state. At this time, it is determined whether an interrupt needs to be triggered according to the interrupt enable control signal. If the interrupt clear control signal is true, it enters the clear waiting state.

[0147] After an interrupt is triggered, if the read map signal is true, the next working state jumps to the state of reading multi-resolution maps to control the point cloud processing module to start receiving the remaining first maps to be searched in other layers. After an interrupt is triggered, if the branch signal (enabling the point cloud processing module to perform the aforementioned coordinate system transformation) is true, the next working state jumps to the search state. Thus, after an interrupt is triggered, it supports the correlation scan matching hardware circuit to jump to the state of reading multi-resolution maps or the search state. Preferably, after an interrupt is triggered, if the task completion signal is true, the next working state jumps to the idle state, waiting for the start of the next correlation scan matching algorithm.

[0148] It should be noted that the search window is a set of pose parameters and depth information of nodes, including the pose parameters and depth information of one or more nodes. The parameters required to perform the correlation scan matching algorithm include the pose parameters of the nodes, and the pose parameters of the nodes include a matching rotation angle and horizontal and vertical coordinates.

[0149] In an embodiment, the state machine control module, by means of hardware-scheduled working states and interrupt signals, successively completes the correlation scan matching of the poses corresponding to the search window on the first map to be searched in each layer. In each correlation scan matching, every time the node search module calculates the score value of a pose within the search window, it controls the point cloud processing module to perform a rotation, a translation, and a multiplication by the reciprocal of the resolution on each point cloud in sequence, so as to read out the occupancy probability with a matching index value from the memory module, realize the operation of the loop state, ensure the real-time performance of the hardware-based correlation scan matching algorithm for searching the branch-and-bound search tree, and improve the convergence speed of the algorithm.

[0150] In the foregoing embodiment, the multi-level grid map is obtained through expansion processing, so that both the first map to be searched and the second map to be searched are the pre-expanded grid maps, and it adapts to the bus width of the bus interface to ensure the alignment of the data transmitted on the bus.

[0151] The expansion processing is specifically as follows: Based on a layer of the original map at a specific resolution, the expansion transformation is successively performed according to the expansion parameters required for the maximum detection radius reached by the currently stored point cloud, the expansion parameters required for performing the coordinate system transformation, the expansion parameters required for performing the rotation transformation, and the expansion parameters required for aligning the bus width, to obtain the corresponding layer of the multi-level grid map; among them, the parameters required for performing the correlation scan matching algorithm include the expansion parameters required for performing the coordinate system transformation and the expansion parameters required for performing the rotation transformation; the foregoing expansion parameters are also stored in the parameter register group and are used as the parameters for the hardware to perform the rotation and translation of the point cloud. The foregoing expansion parameters are all related to the maximum range reached by the point cloud. The specific expansion parameters and the calculation method of the expansion transformation of the grid map are prior arts and will not be elaborated here.

[0152] Among them, the map delineated by the currently stored point cloud is a grid map centered on the lidar and delineated with the maximum detection diameter reached by the currently stored point cloud as the side length of the map. The size of this grid map includes the maximum height value and the minimum height value restricted by the occlusion of obstacles in the height direction of the map, and the maximum width value and the minimum width value restricted by the occlusion of obstacles in the width direction of the map, which is equivalent to cropping the grid map default constructed by the system based on the occlusion effect of obstacles; or centered on the robot, cropping the grid map default constructed by the system with the maximum range reached by the point cloud as the radius. Among them, aligning the bus width means that the memory occupied by the boundary grid of the pre-expanded grid map is equal to the bit width of the bus for transmitting this grid map.

[0153] Among them, the aforementioned expansion parameters include coordinate translation parameters; the sum value of the coordinate translation parameters required for the maximum detection radius reached by the currently stored point cloud, the coordinate translation parameters required for performing the coordinate transformation, the coordinate translation parameters required for performing the rotation transformation, and the coordinate translation parameters required for aligning the bus width determines the preset coordinate offset value; the offset value register is used to store the aforementioned expansion parameters, and the offset value register is a parameter register set inside the correlation scanning and matching hardware circuit. The pre-expanded grid map is based on the maximum coverage radius of the point cloud and is expanded in sequence according to the coordinate translation parameters, coordinate rotation parameters, and memory parameters occupied by the map side length, which can not only constrain the actual usage size of the map, but also adapt to the bus width used during map transmission, ensure data alignment, and reduce the storage space occupied by the map.

[0154] Specifically, the construction steps of one layer of the original map with a specific resolution include:

[0155] Taking the preset grid side length matched at a specific resolution as the sampling interval, use a preset sliding window to sample the grid map with the first resolution; among them, the larger the preset grid side length, the smaller the specific resolution.

[0156] During each sampling of the preset sliding window, all the grids covered by the preset sliding window in the grid map with the first resolution are merged into a preset grid, and then this preset grid forms a grid of one layer of the original map with a specific resolution. Thus, compared with the grid map with the first resolution, the number of grids used to describe the same environmental area in one layer of the original map with a specific resolution becomes smaller, and the side length of the corresponding formed grid becomes larger. Among them, each preset grid stores a matched index value. When one layer of the original map with a specific resolution is transmitted to the map storage space, it is configured as the index value of the map storage space and corresponds to the occupancy probability of a point cloud on the grid point of one layer of the original map with a specific resolution.

[0157] After the preset sliding window completes sampling of all grid regions of the grid map with the first resolution row by row or column by column, all merged preset grids form an original map of a specific resolution at one layer; preferably, the preset sliding window first performs interval sampling along the X-axis direction, slides and traverses from one boundary to another boundary, then changes rows along the Y-axis direction, and repeats the foregoing sampling steps until sampling of all grid regions of the grid map with the first resolution is completed; or, the preset sliding window first performs interval sampling along the Y-axis direction, slides and traverses from one boundary to another boundary, then changes columns along the X-axis direction, and repeats the foregoing sampling steps until coverage of all grid regions of the grid map with the first resolution is completed. Among them, the size of the preset sliding window is negatively correlated with the specific resolution, so as to sequentially construct a first map to be searched with the first preset number of layers and a second map to be searched with the second preset number of layers longitudinally, construct the multi-level grid map including multiple layers of original maps with specific resolutions. In the multi-level grid map, from top to bottom, the resolution increases layer by layer, the resolution of the topmost first map to be searched is the lowest, and the resolution of the bottommost second map to be searched is the highest.

[0158] It should be added that the foregoing grid point is the central position of the grid where it is located.

[0159] In this embodiment, a sliding window is designed to perform grid merging processing on the original grid map with the first resolution. The number of grids representing the same regional position in each merged grid map is different, realizing hierarchical fuzzy processing, that is, realizing different resolutions for each layer of the map; furthermore, ensuring that the occupancy probability of map storage at low resolution is relatively large, and ensuring that the lower the resolution of the map matched by the point cloud, the higher the accumulated score value calculated. Therefore, when it is necessary to search for a pose or a matching node with a higher score value, this embodiment starts searching layer by layer from the map with the highest resolution.

[0160] As an embodiment, a bus interface module is externally provided for the correlation scan matching hardware circuit. The bus interface module includes a DMA controller module and a transmission bus; preferably, the bus interface module includes an AXI DMA module and an AXI bus.

[0161] The DMA controller module is used to continuously transfer data stored in the physical storage space with discontinuous addresses in batches, reducing the number of trigger times of the software interrupt of the CPU; specifically, the DMA controller module requires the device driver of the CPU to generate a linked list of storage data addresses, describe the discontinuous physical space with the linked list, and send the initial address of the linked list to the DMA controller module. After the DMA controller module transfers a section of data to the correlation scan matching hardware circuit, it sequentially reads the address of the next linked list until all the data is transferred.

[0162] The transmission bus includes a first bus and a second bus. Preferably, the transmission bus is an AXI bus, which includes two types of AXI buses, namely the AXI-Lite bus and the AXI-Stream bus. The former is suitable for simple and low-throughput memory-mapped communication, while the latter is for high-speed stream data. The first bus is preferably the AXI-Lite bus, which establishes connection relationships between the CPU and the correlation scan matching hardware circuit and the DMA controller module respectively. It can send control signals to the memory module, the point cloud processing module, the node search module, the state machine control module, and the interconnection bus, and monitor the working state of the read state machine control module. The first bus has signal transceiver connections with the memory module, the point cloud processing module, the node search module, the state machine control module, the interconnection bus, and the DMA controller module. The first bus is used to configure data transmission parameters for the DMA controller module. The first bus is also used to configure the parameters stored in the map size register, the parameters stored in the map resolution register, the expansion parameters required for the maximum detection radius reached by the currently stored point cloud, the expansion parameters required for performing the coordinate transformation, the expansion parameters required for performing the rotation transformation, the expansion parameters required to meet the transmission requirements of the bus bit width, and the search step, so as to achieve memory-mapped communication for the memory of the correlation scan matching hardware circuit.

[0163] The second bus is preferably the AXI-Stream bus. The second bus is connected to the DMA controller module and is used to transmit the multi-level grid map pre-constructed by the CPU, the sine function values under the pre-configured corresponding rotation parameters, the cosine function values under the pre-configured corresponding rotation parameters, and the point cloud currently collected by the lidar to the memory module. Among them, the transmission bus follows the AMBA protocol. This embodiment provides a bus interface architecture module for the correlation scan matching hardware circuit and an external data source (controller). According to the data transmission performance, a first bus suitable for simple and low-throughput memory-mapped communication (transmitting expansion parameters, map feature parameters, and associated basic control signals to the registers inside the correlation scan matching hardware circuit) and a second bus for high-speed data streams (transmitting truncated signed distance function values and weight values to the memory module) are designed respectively, improving the real-time performance of the arithmetic operations of the correlation scan matching hardware circuit.

[0164] This embodiment provides a bus interface architecture module for the correlation scanning matching hardware circuit and the CPU, and designs a first bus suitable for simple and low-throughput memory-mapped communication (transmitting extended parameters, map feature parameters, and associated basic control signals to the registers inside the correlation scanning matching hardware circuit) and a second bus for high-speed data streams (point clouds collected in real time by lidar, transmitting each layer of the first map to be searched and trigonometric function values that need to be calculated urgently to the memory module) according to data transmission performance, improving the real-time performance of the arithmetic operations of the correlation scanning matching hardware circuit, the real-time performance of the CPU running the branch and bound algorithm and the correlation scanning matching algorithm.

[0165] As an embodiment, in the process of the CPU executing the branch and bound algorithm in step I, for each branch of a non-leaf node, 4 child nodes are searched in the second map to be searched at the corresponding layer. Among them, these 4 child nodes belong to the same branch and bound search tree, so that the rotation parameters required by these 4 child nodes are the same, the coordinate translation parameters required by these 4 child nodes are different from each other, and the rotation parameters and coordinate translation parameters required by these 4 child nodes are included within the same search window. The rotation parameters and the coordinate translation parameters are both parameters required for executing the correlation scanning matching algorithm. Thereby, the search range is narrowed, and the leaf nodes on different branch and bound search trees can be searched faster or pruning can be accelerated to enter another branch for continued search (equivalent to depth-first search can accelerate the speed of branch and bound) by scanning the first map to be searched and the second map to be searched at each layer through the correlation scanning matching algorithm, and based on the leaf node scores, the current optimal upper bound value is updated, thereby improving the execution efficiency of the branch and bound algorithm.

[0166] It should be noted that the ways for the CPU to execute the correlation scan matching algorithm include: according to the rotation parameters and the coordinate translation parameters, converting and matching the point cloud to a second to-be-searched map of the layer where the currently searched node is located, obtaining a discrete point cloud, and determining the completion of the discretization of the point cloud; specifically including rotation transformation, coordinate translation, and multiplying by the reciprocal of the resolution; for a point cloud (including the abscissa and the ordinate), each time a rotation transformation and a corresponding coordinate translation are executed, and then multiplying by the reciprocal of the resolution (the resolution of the currently searched second to-be-searched map of a layer), a discrete point cloud can be obtained, and the discretization of the point cloud is completed. The specific software execution steps of the CPU can refer to the function description part of the aforementioned point cloud processing module, which will not be elaborated here because the execution method processes of the two are the same, except that in this embodiment, the corresponding steps are not hardwareized with logic devices but executed in the form of a computer program. Then, according to the preset coordinate offset value and the search step included in the search window, calculate the index value of the occupancy probability at the corresponding grid point of the discrete point cloud on the second to-be-searched map, and then use the currently obtained index value to perform an accumulative calculation on the corresponding occupancy probability, and set the accumulative result as the score value of the currently searched node; the specific software execution steps can refer to the function description of the aforementioned node search module, which will not be elaborated here because the execution method processes of the two are the same, except that in this embodiment, the corresponding steps are not hardwareized with logic devices but executed in the form of a computer program. Among them, the search window is a set of pose parameters required for the correlation scan matching algorithm to perform scan matching on the second to-be-searched map, and is used to complete the search for all child nodes from one branch to the corresponding second to-be-searched map of a layer. In this embodiment, the CPU scans and matches the second to-be-searched maps of the second preset number of layers. On the one hand, it adapts to the fact that the score values (occupancy probabilities) assigned to the grid points of the map with a higher resolution are smaller, so that it is easy to be pruned during the search process on the second to-be-searched map, simplifies the algorithm, and reduces the computationally intensive processes.

[0167] Based on the foregoing embodiments, the present invention also discloses a robot, which is equipped with a lidar, and the hardware acceleration positioning system as described above is provided inside the robot to solve the relocalization problem on the embedded platform of the mobile robot. Compared with the existing technology robots that execute the branch and bound algorithm in a pure software platform manner, even when the frequency of the correlation scan matching hardware circuit in this technical solution is not high, it still has an obvious advantage in running speed and can better meet the relocalization requirements of the mobile robot platform.

[0168] As described above, it is only the preferred embodiment of the present invention, and it is not intended to limit the present invention in other forms. Any person skilled in the art may use the technical content disclosed above to make changes or modifications into equivalent embodiments with equivalent changes. However, any simple modification, equivalent change, and modification made to the above embodiments based on the technical essence of the present invention without departing from the technical solution content of the present invention still fall within the protection scope of the technical solution of the present invention.

Claims

1. A hardware acceleration positioning system under the cooperation of software and hardware, where the hardware acceleration positioning system is electrically connected to a lidar; among them, The lidar is installed on the body of the robot; characterized in that the hardware acceleration positioning system includes a CPU and a correlation scan matching hardware circuit; The correlation scan matching hardware circuit is used to search each layer of the first map to be searched in a hardware parallel processing manner and update the first upper bound value of the currently searched node; The CPU is used to, whenever the correlation scan matching hardware circuit searches for the node corresponding to the grid point of a layer of the map with a preset depth, search each layer of the second map to be searched by executing the branch and bound algorithm, and perform scan matching of the point cloud collected by the lidar on the corresponding layer of the second map to be searched by executing the correlation scan matching algorithm, and then calculate and update the second upper bound value of the currently searched node by using the matched result; The CPU is further used to configure the pose corresponding to the latest obtained second upper bound value or the pose corresponding to the latest obtained first upper bound value as the pose of the robot; Wherein, more than one branch and bound search tree is constructed for all layers of the first map to be searched and all layers of the second map to be searched, so that the grid points corresponding to the root nodes of all the branch and bound search trees are in the first map to be searched of the lowest resolution layer; the resolution of each layer of the first map to be searched is lower than the resolution of any layer of the second map to be searched; the first map to be searched and the second map to be searched are set in layers.

2. The hardware acceleration positioning system according to claim 1, wherein The hardware acceleration positioning system stores a pre-constructed grid map with a first resolution, and the grid map with the first resolution is constructed by the CPU into a multi-level grid map, and the multi-level grid map includes the first maps to be searched with a first preset number of layers and the second maps to be searched with a second preset number of layers, and the maximum resolution in all layers of the second maps to be searched is equal to the first resolution; The first maps to be searched with the first preset number of layers and the second maps to be searched with the second preset number of layers are arranged vertically in ascending order of resolution. Among them, the serial numbers of the first maps to be searched with the first preset number of layers and the serial numbers of the second maps to be searched with the second preset number of layers both represent the depth of the corresponding layer of the map.

3. The hardware acceleration positioning system according to claim 2, wherein The hardware acceleration positioning system further includes a bus interface, and both the CPU and the correlation scan matching hardware circuit are electrically connected to the bus interface; the CPU is configured to execute the following steps: Step A, the CPU configures the parameters required for executing the branch and bound algorithm and the parameters required for executing the correlation scan matching algorithm for the correlation scan matching hardware circuit through the bus interface; wherein, the parameters required for executing the branch and bound algorithm include the first preset number of layers, the resolution of each layer of the first map to be searched, the second preset number of layers, the resolution of each layer of the second map to be searched, and the number of root nodes; Step B, the CPU transmits the first map to be searched of the lowest resolution layer to the correlation scan matching hardware circuit through the bus interface; Step C: The CPU activates the correlation scan matching hardware circuit to search for all root nodes, and triggers the correlation scan matching hardware circuit to calculate the score values of the corresponding root nodes; meanwhile, the CPU discretizes the point cloud according to the parameters required to execute the correlation scan matching algorithm. Step D: The CPU transmits the remaining first map to be searched to the correlation scan matching hardware circuit through the bus interface; meanwhile, the CPU sorts all the root nodes currently searched by the correlation scan matching hardware circuit in ascending order of the score values, and then stores the root nodes in the stack according to the current sorting. Among them, the stack is a cache space internally opened by the hardware acceleration positioning system that supports last-in, first-out; the depth of each layer of the first map to be searched and the depth of each layer of the second map to be searched are respectively equal to the depths of the nodes at the same vertical height of the branch and bound search tree.

4. The hardware acceleration positioning system according to claim 3, wherein The CPU is further configured to execute the following steps: Step E: The CPU controls the node at the top of the stack to pop out of the stack, and then the CPU determines whether the currently popped node is a leaf node. If so, it proceeds to Step F; otherwise, it proceeds to Step G. Step F: The CPU determines whether the score value of the currently popped node is greater than the current optimal upper bound value. If so, it updates the score value of the node to the current optimal upper bound value, and determines the updated current optimal upper bound value as the first upper bound value or the second upper bound value of the currently searched node, and then proceeds to Step H; otherwise, it directly proceeds to Step H, where the initial value of the current optimal upper bound value is a preset initial threshold. Step G: When the CPU determines that the score value of the currently popped node is greater than the current optimal upper bound value, the CPU determines whether the absolute value of the difference between the depth of the currently popped node and the depth of the first map to be searched at the lowest resolution is greater than or equal to the reference depth difference. If so, it proceeds to Step I, and determines the node corresponding to the grid point of the map at a preset depth searched by the correlation scan matching hardware circuit; otherwise, it proceeds to Step J; where the reference depth difference is equal to the difference between the first preset number of layers and the value 1; when the CPU determines that the score value of the currently popped node is less than or equal to the current optimal upper bound value, no branching operation is performed, and then Step H is executed. Step I: The CPU performs a branching operation on the node described in Step G to search for all child nodes of the node described in Step G, and then calculates the score values of the currently searched child nodes; then it proceeds to Step K; where the branching operation in Step I is to search for all child nodes of the node described in Step G in the adjacent second map to be searched at a resolution higher than that of the layer of the map to which the node described in Step G belongs. Step J: The correlation scanning and matching hardware circuit performs a branch operation in a hardware parallel processing manner, searches for all child nodes of the node described in Step G, and calculates the score value of the currently searched child node; then proceeds to Step K. Among them, the branch operation in Step J is to search for all child nodes of the node described in Step G in the adjacent first map to be searched at a resolution higher than that of the layer of the map to which the node described in Step G belongs. Step K: The CPU sorts the child nodes searched in Step I or the child nodes searched in Step J in the order described in Step D, and then controls the currently sorted child nodes to be pushed onto the stack in the order described in Step D; then returns to Step E. Step H: The CPU determines whether the stack is empty. If it is, it configures the pose corresponding to the node with the current optimal upper bound value described in Step F as the pose of the robot; otherwise, it returns to Step E. Among them, one root node corresponds to a branch and bound search tree, and a branch operation included in the branch and bound algorithm is adopted to search for a node of each branch and bound search tree; the depth of each layer of the first map to be searched and the depth of each layer of the second map to be searched are respectively equivalent to the depths of the corresponding nodes on the branch and bound search tree. Among them, in the hardware-accelerated positioning system, the resolution of the first map to be searched transmitted layer by layer to the correlation scanning and matching hardware circuit increases, and the resolution of the second map to be searched processed layer by layer by the CPU also increases.

5. The hardware acceleration positioning system according to claim 4, wherein In Step E, the top node of the stack is the first node or the second node. Among them, the first node is configured with the coordinates of the grid point of the first map to be searched and the depth information of the first map to which the grid point belongs, and the second node is configured with the coordinates of the grid point of the second map to be searched and the depth information of the second map to which the grid point belongs. Whenever the CPU determines in Step F that the score value of the currently popped first node is greater than the current optimal upper bound value, it updates the score value of this first node to the current optimal upper bound value and determines that the score value of this first node is the first upper bound value of the currently searched node. Whenever the CPU determines in Step F that the score value of the currently popped second node is greater than the current optimal upper bound value, it updates the score value of this second node to the current optimal upper bound value and determines that the score value of this second node is the second upper bound value of the currently searched node.

6. The hardware acceleration positioning system according to claim 4, wherein During the process of the correlation scanning and matching hardware circuit calculating the score value of the root node in Step C, if the correlation scanning and matching hardware circuit triggers an interruption, the correlation scanning and matching hardware circuit will instead enter Step D; during the process of the CPU discretizing the point cloud in the second map to be searched in Step C, if it detects that an interruption is triggered, the CPU will instead execute Step D. During the process that the remaining first map to be searched is branched by the correlation scan and match hardware circuit, the CPU adopts a polling mechanism to wait for the correlation scan and match hardware circuit to calculate and output the score value of the corresponding node, so as to avoid delay caused by adopting the interrupt mechanism, where the currently searched node is not the root node.

7. The hardware-accelerated positioning system according to any one of claims 4 to 6, characterized in that The correlation scan and match hardware circuit includes a memory module, a point cloud processing module, a node search module, a state machine control module, and an interconnection bus. The memory module, the point cloud processing module, the node search module, and the state machine control module establish data transmission relationships through the interconnection bus. The point cloud processing module is configured to read the currently stored point cloud from the memory module under the control of the state machine control module, and then control the read point cloud to perform a rotation transformation. The point cloud processing module is configured to read the result of the current rotation transformation under the control of the state machine control module, then control the result of the current rotation transformation to perform a coordinate system transformation, and then set the result of the coordinate system transformation as a discrete point cloud, so as to align the currently stored point cloud into the coordinate system of the pre-expanded grid map, and at the same time determine the discretization of the currently stored point cloud; then store the discrete point cloud into the memory module through the interconnection bus; where the pre-expanded grid map is a layer of the first map to be searched required for searching the root node in step C, or a layer of the first map to be searched required for searching the child node in step J. The node search module is configured to obtain the discrete point cloud from the memory module under the control of the state machine control module, and then calculate the index value of the discrete point cloud mapped to the map storage space according to the preset coordinate offset value and search step, and set the index value as the read address of the occupancy probability of the point cloud. The node search module is further configured to read the occupancy probability of the corresponding point cloud from the map storage space according to the currently set read address under the control of the state machine control module, call the accumulator built in the node search module to accumulate the occupancy probability of the currently read point cloud, and then output the accumulation result and set it as the score value of the corresponding node in the pre-expanded grid map. The memory module is provided with a map storage space for storing the first map to be searched transmitted by the bus interface. The memory module is used to store the preset trigonometric function values. Wherein, the number of rotation transformations is equal to the number of the aforementioned root nodes, and one or more coordinate system transformations are performed corresponding to one rotation transformation; a node currently searched by the search window is determined by a specific rotation parameter and a specific translation parameter. Wherein, the occupancy probability of the point cloud is the probability value of the grid point matched by the point cloud in the pre-expanded grid map after being processed by the point cloud processing module; there is a corresponding index value for the occupancy probability of the point cloud in each layer of the first map to be searched.

8. The hardware-accelerated positioning system according to claim 7, wherein The map storage space includes a first map storage space and a second map storage space. The first map storage space is used to store the first map to be searched with the first prefabricated resolution, and the second map storage space is used to store the first map to be searched with the second prefabricated resolution. The depth levels of the first map to be searched with the first prefabricated resolution and the first map to be searched with the second prefabricated resolution differ by one level, so that the two layers of the first map to be searched are adjacent; the first prefabricated resolution is less than the second prefabricated resolution; The correlation scanning and matching hardware circuit is used to schedule the point cloud processing module to align and match the currently stored point cloud to the first map to be searched with the second prefabricated resolution during one execution of the branch operation, and also schedule the node search module to calculate the index value of the occupancy probability of the currently stored point cloud in the first map storage space.

9. The hardware-accelerated positioning system according to claim 8, wherein The first preset number of layers is 3, and the second preset number of layers is 4; During the CPU's execution of step B, the CPU transmits the first layer of the first map to be searched to the correlation scanning and matching hardware circuit through the bus interface, and then the bus interface writes the first layer of the first map to be searched into the first map storage space; When the CPU executes step D, the CPU first transmits the second layer of the first map to be searched to the correlation scanning and matching hardware circuit, and then the CPU transmits the third layer of the first map to be searched to the correlation scanning and matching hardware circuit. Then, the bus interface writes the second layer of the first map to be searched into the first map storage space and overwrites the first layer of the first map to be searched; then the bus interface writes the third layer of the first map to be searched into the second map storage space; Among them, before completely overwriting the first layer of the first map to be searched, the score values of all root nodes have been obtained.

10. The hardware acceleration positioning system according to claim 7, wherein, The node search module includes a selector and an accumulator; the accumulator is used to accumulate the occupancy probability of the corresponding point cloud in the map storage space under the control of the state machine control module. Among them, the occupancy probability of the corresponding point cloud in the map storage space is first transmitted to the interconnect bus according to the address corresponding to the index value under the control of the state machine control module, and then sequentially transmitted to the input end of the accumulator by the interconnect bus; The input end of the selector is connected to the output end of the accumulator; the selector is used to select and output the accumulation result of the accumulator every other preset counting cycle, and set the currently output accumulation result as the score value of a corresponding node in the first map to be searched for a corresponding layer; Among them, every time the point cloud processing module completes a coordinate system transformation in the map correlation scanning and matching algorithm, it triggers the selector to output the currently accumulated accumulation result of the accumulator to the interconnect bus, and configures the currently obtained accumulation result as the score value of a new node corresponding to the search of the search window.

11. The hardware acceleration positioning system according to claim 10, wherein The node search module further includes a grid index sub-module, which includes a first index subtractor, a second index subtractor, a first index adder, and a second index adder, and is configured as a pipeline structure; wherein, the first index subtractor and the second index subtractor both belong to subtractors, and the first index adder and the second index adder both belong to adders; The first input end of the first index subtractor is used to receive the abscissa of the discrete point cloud transmitted by the interconnection bus. Among them, the abscissa of the discrete point cloud transmitted by the interconnection bus to the first index subtractor is sourced from the memory module; The second input end of the first index subtractor is used to receive the abscissa axis coordinate offset value transmitted by the interconnection bus. Among them, the preset coordinate offset value includes the abscissa axis coordinate offset value, and the abscissa axis coordinate offset value transmitted by the interconnection bus is sourced from the preset coordinate offset value stored in the map offset value register; the map offset value register is a parameter register set inside the correlation scan matching hardware circuit and is used to store the coordinate offset value associated with the pre-expanded grid map; The first input end of the first index adder is connected to the output end of the first index subtractor; the second input end of the first index adder is used to receive the abscissa axis coordinate search step transmitted by the interconnection bus. Among them, the search step includes the abscissa axis coordinate search step, and the abscissa axis coordinate search step transmitted by the interconnection bus is sourced from the search step stored in the search window parameter register; the search window parameter register is a parameter register set inside the correlation scan matching hardware circuit and is used to store the positions of the nodes existing in the search window and the associated sub-node search information, including the search step; the sum value output by the first index adder is configured as the abscissa axis direction index value of the discrete point cloud mapped to the map storage space; The first input end of the second index subtractor is used to receive the ordinate of the discrete point cloud transmitted by the interconnection bus. Among them, the ordinate of the discrete point cloud transmitted by the interconnection bus to the second index subtractor is sourced from the memory module; The second input end of the second index subtractor is used to receive the ordinate axis coordinate offset value transmitted by the interconnection bus. Among them, the preset coordinate offset value further includes the ordinate axis coordinate offset value, and the ordinate axis coordinate offset value transmitted by the interconnection bus is also sourced from the preset coordinate offset value stored in the map offset value register; The first input end of the second index adder is connected to the output end of the second index subtractor; the second input end of the second index adder is used to receive the ordinate axis coordinate search step transmitted by the interconnection bus. Among them, the search step further includes the ordinate axis coordinate search step, and the ordinate axis coordinate search step transmitted by the interconnection bus is also sourced from the search step stored in the search window parameter register; the sum value output by the second index adder is configured as the ordinate axis direction index value of the discrete point cloud mapped to the map storage space.

12. The hardware acceleration positioning system according to claim 11, wherein The grid index sub-module further includes a third index adder and an index multiplier. The third index adder belongs to the adder, and the index multiplier belongs to the multiplier; The first input terminal of the index multiplier is connected to the output terminal of the second index adder; the second input terminal of the index multiplier is used to receive the number of row grids transmitted by the interconnection bus. The number of row grids is the number of grids that each row of the map storage space can occupy, stored in the map size register and transmitted from the map size register to the interconnection bus. The map size register is a parameter register set inside the correlation scan matching hardware circuit, used to store the size range of the first map to be searched and associated extended information; The first input terminal of the third index adder is connected to the output terminal of the first index adder, and the second input terminal of the third index adder is connected to the output terminal of the index multiplier. The third index adder is used to control the sum of the product of the number of row grids and the sum value output by the second index adder and the horizontal axis direction index value, and then set the sum value obtained by the addition as the index value of the discrete point cloud mapped into the map storage space, and send it to the interconnection bus, so as to read out the occupancy probability matching the index value from the memory module.

13. The hardware acceleration positioning system according to claim 11, characterized in that, The grid index sub-module further includes a third index adder and an index multiplier. The third index adder belongs to the adder, and the index multiplier belongs to the multiplier; The first input terminal of the index multiplier is connected to the output terminal of the first index adder; the second input terminal of the index multiplier is used to receive the number of column grids transmitted by the interconnection bus. The number of column grids is the number of grids that each column of the map storage space can occupy, stored in the map size register and transmitted from the map size register to the interconnection bus. The map size register is a parameter register set inside the correlation scan matching hardware circuit, used to store the size range of the first map to be searched and associated extended information; The first input terminal of the third index adder is connected to the output terminal of the second index adder, and the second input terminal of the third index adder is connected to the output terminal of the index multiplier. The third index adder is used to control the sum of the product of the number of column grids and the sum value output by the first index adder and the vertical axis direction index value, and then set the sum value obtained by the addition as the index value of the discrete point cloud mapped into the map storage space, and send it to the interconnection bus, so as to read out the occupancy probability matching the index value from the memory module.

14. The hardware-accelerated positioning system according to claim 7, wherein The multi-level grid map is obtained through expansion processing, so that the first map to be searched is the pre-expanded grid map and adapts to the bus bit width of the bus interface; The specific expansion process is as follows: Based on a layer of original map at a specific resolution, expansion transformations are successively performed according to the expansion parameters required for the maximum detection radius reached by the currently stored point cloud, the expansion parameters required for performing the coordinate transformation, the expansion parameters required for performing the rotation transformation, and the expansion parameters required for aligning the bus bit width, to obtain the corresponding layer of map in the multi-level grid map; among them, the parameters required for performing the correlation scan matching algorithm include the expansion parameters required for performing the coordinate transformation and the expansion parameters required for performing the rotation transformation. Among them, the map delineated by the currently stored point cloud is a grid map delineated with the lidar as the center and the maximum detection diameter reached in the currently stored point cloud as the side length of the map. The size of this grid map includes the maximum height value and the minimum height value restricted by the occlusion of obstacles in the height direction of the map, and the maximum width value and the minimum width value restricted by the occlusion of obstacles in the width direction of the map. Among them, the aligned bus bit width means that the memory occupied by the boundary grid of the pre-expanded grid map is equal to the bit width of the bus for transmitting this grid map. Among them, the aforementioned expansion parameters include coordinate translation parameters; the sum value of the coordinate translation parameters required for the maximum detection radius reached by the currently stored point cloud, the coordinate translation parameters required for performing the coordinate transformation, the coordinate translation parameters required for performing the rotation transformation, and the coordinate translation parameters required for aligning the bus bit width determines the preset coordinate offset value; the offset value register is used to store the aforementioned expansion parameters, and the offset value register is a parameter register set inside the correlation scan matching hardware circuit.

15. The hardware-accelerated positioning system according to claim 14, wherein The construction steps of a layer of original map at a specific resolution include: Taking the preset grid side length matched at a specific resolution as the sampling interval, and sampling the grid map at the first resolution using a preset sliding window. During each sampling of the preset sliding window, all the grids covered by the preset sliding window in the grid map at the first resolution are merged into a preset grid, and then this preset grid is used as a grid of a layer of original map at a specific resolution. Among them, each preset grid stores a matched index value. When the preset sliding window completes the coverage of all grid areas of the grid map at the first resolution row by row or column by column, all the merged preset grids form a layer of original map at a specific resolution. Among them, the larger the size of the preset sliding window, the smaller the specific resolution. Among them, the aforementioned grid point is the central position of the grid where it is located.

16. The hardware-accelerated positioning system according to claim 14, wherein The point cloud processing module includes a point cloud rotation sub-module and a point cloud discretization sub-module. The point cloud rotation sub-module includes a first register, a second register, a first point cloud multiplier, a second point cloud multiplier, a third point cloud multiplier, a fourth point cloud multiplier, a point cloud adder, and a point cloud subtractor, and is configured as a pipeline structure; the first point cloud multiplier, the second point cloud multiplier, the third point cloud multiplier, and the fourth point cloud multiplier all belong to multipliers, the point cloud adder belongs to an adder, and the point cloud subtractor belongs to a subtractor. The first input terminal of the first point cloud multiplier is connected to the output terminal of the first register; wherein, the interconnection bus transmits the abscissa of the point cloud to the input terminal of the first register; the abscissa of the point cloud transmitted by the interconnection bus to the first register is derived from the abscissa of the currently stored point cloud; The second input terminal of the first point cloud multiplier is used to receive the cosine function value corresponding to the rotation angle reached by the current rotation transformation; Wherein, the rotation angle reached by the current rotation transformation is the sum value of the angle search step size and the rotation angle reached by the previous rotation transformation; the angle search step size is stored in the search window parameter register; the search window parameter register is a parameter register set inside the correlation scan matching hardware circuit, and is used to store the poses and associated step search information included in the search window; The first input terminal of the second point cloud multiplier is connected to the output terminal of the second register; wherein, the interconnection bus caches the ordinate of the point cloud to the second register; the ordinate of the point cloud transmitted by the interconnection bus to the second register is derived from the ordinate of the currently stored point cloud; The second input terminal of the second point cloud multiplier is used to receive the sine function value corresponding to the rotation angle reached by the current rotation transformation; The first input terminal of the third point cloud multiplier is connected to the output terminal of the first register; The second input terminal of the third point cloud multiplier is used to receive the sine function value corresponding to the rotation angle reached by the current rotation transformation; The first input terminal of the fourth point cloud multiplier is connected to the output terminal of the second register; The second input terminal of the fourth point cloud multiplier is used to receive the cosine function value corresponding to the rotation angle reached by the current rotation transformation; The first input terminal of the point cloud subtractor is connected to the output terminal of the first point cloud multiplier, and the second input terminal of the point cloud subtractor is connected to the output terminal of the second point cloud multiplier; The point cloud subtractor is used to transmit the output difference to the point cloud discretization sub-module, and set the output difference as the abscissa result of the current rotation transformation; The first input terminal of the point cloud adder is connected to the output terminal of the third point cloud multiplier, and the second input terminal of the point cloud adder is connected to the output terminal of the fourth point cloud multiplier; The point cloud adder is used to transmit the output sum value to the point cloud discretization sub-module, and set the output sum value as the ordinate result of the current rotation transformation.

17. The hardware acceleration positioning system according to claim 16, wherein The point cloud discretization sub-module includes a first discrete adder, a first discrete subtractor, a first discrete multiplier, a second discrete adder, a second discrete subtractor and a second discrete multiplier, and is configured as a pipeline structure; the first discrete adder and the second discrete adder both belong to adders, the first discrete subtractor and the second discrete subtractor both belong to subtractors, and the first discrete multiplier and the second discrete multiplier both belong to multipliers; The first input terminal of the first discrete adder is connected to the output terminal of the point cloud subtractor, and the first input terminal of the first discrete adder is used to receive the abscissa result of the current rotation transformation; The second input terminal of the first discrete adder is used to receive the abscissa of the robot position transmitted by the interconnection bus, where the abscissa of the robot position is the pre-calculated value of the horizontal axis coordinate of the robot in the world coordinate system; The first discrete adder is used to control the addition of the abscissa result obtained from the current rotation transformation and the abscissa of the robot position, and then output the sum value obtained by the addition, so that the sum value becomes the horizontal axis coordinate value of the point cloud transformed into the world coordinate system; The first input terminal of the first discrete subtractor is connected to the output terminal of the first discrete adder; the second input terminal of the first discrete subtractor is used to receive the maximum abscissa value of the map transmitted by the interconnection bus, and the maximum abscissa value of the map is derived from the map size register and transmitted from the map size register to the interconnection bus; where the map size register is a parameter register set inside the correlation scan matching hardware circuit and is used to store the size range of the grid map that meets the transmission requirements of the bus bit width; The first input terminal of the first discrete multiplier is connected to the output terminal of the first discrete subtractor; the second input terminal of the first discrete multiplier is used to receive the reciprocal of the map resolution transmitted by the interconnection bus, and the reciprocal of the map resolution is derived from the map resolution register and transmitted from the map resolution register to the interconnection bus; where the map resolution register is a parameter register set inside the correlation scan matching hardware circuit and is used to store the resolution information of the first map to be searched for a corresponding layer; the product value output by the first discrete multiplier is configured as the abscissa value of the discrete point cloud to complete the coordinate system transformation, so as to align the abscissa value of the currently stored point cloud into the coordinate system of the first map to be searched for a corresponding layer, and at the same time determine the discretization of the abscissa of the currently stored point cloud; The first discrete multiplier is used to output the product value obtained by multiplication to the memory module through the interconnection bus, and determine that the product value is the abscissa that matches the currently aligned and matched node on the first map to be searched for a corresponding layer; The first input terminal of the second discrete adder is connected to the output terminal of the point cloud adder, and the first input terminal of the second discrete adder is used to receive the ordinate result obtained from the current rotation transformation; The second input terminal of the second discrete adder is used to receive the ordinate of the robot position transmitted by the interconnection bus, where the ordinate of the robot position is the pre-calculated value of the vertical axis coordinate of the robot in the world coordinate system; the robot position register is used to store the ordinate and abscissa of the robot position, and the robot position register is a parameter register set inside the correlation scan matching hardware circuit; The second discrete adder is used to control the addition of the ordinate result obtained from the current rotation transformation and the ordinate of the robot position, and then output the sum value obtained by the addition, so that the sum value becomes the vertical axis coordinate value of the point cloud transformed into the world coordinate system; The first input terminal of the second discrete subtractor is connected to the output terminal of the second discrete adder; the second input terminal of the second discrete subtractor is used to receive the maximum ordinate value of the map transmitted by the interconnection bus, and the maximum ordinate value of the map is derived from the map size register and transmitted by the map size register to the interconnection bus; The first input terminal of the second discrete multiplier is connected to the output terminal of the second discrete subtractor; the second input terminal of the second discrete multiplier is used to receive the reciprocal of the map resolution transmitted by the interconnection bus; the product value output by the second discrete multiplier is configured as the ordinate value of the discrete point cloud to complete the coordinate system transformation, so as to align the ordinate value of the currently stored point cloud into the coordinate system of the corresponding first map to be searched in a layer, and at the same time determine to complete the discretization of the abscissa of the currently stored point cloud; The second discrete multiplier is used to output the product value obtained by multiplication to the memory module through the interconnection bus, and determine that the product value is the ordinate corresponding to the currently aligned and matched node on the corresponding first map to be searched in a layer.

18. The hardware-accelerated positioning system according to claim 17, wherein The state machine control module belongs to a finite state machine; the state machine control module is used to schedule the working states of the memory module, the point cloud processing module and the node search module; wherein, the working states include the state of reading the multi-resolution map, the search state and the loop state; The state machine control module is used to execute: In the state of reading the multi-resolution map, control the point cloud processing module to receive the first map to be searched in a layer; then in the search state, control the point cloud processing module to perform the aforementioned coordinate system transformation on the currently stored point cloud to obtain a discrete point cloud, and then control the node search module to calculate the index value of the discrete point cloud mapped to the map storage space where the current first map to be searched in a layer is located, and use the currently calculated index value to search the occupancy probability in the corresponding map storage space, and control the aforementioned accumulator to perform an accumulative calculation on the searched occupancy probability to obtain the score value of the corresponding node; After saving the currently calculated score value, go to the loop state; In the loop state, if it is judged that the currently processed node is the root node and the occupancy probabilities corresponding to all root nodes have not been indexed, then return to the search state, and continue to use the newly calculated index value to search the occupancy probability in the corresponding map storage space until the occupancy probabilities corresponding to all root nodes supported by the search window have been completed, and then the option to trigger an interrupt is allowed; In the loop state, if it is judged that there are still nodes whose score values have not been calculated and the currently processed node is not the root node, then return to the search state, and continue to control the aforementioned accumulator to perform an accumulative calculation on the searched occupancy probability until the score values of all nodes supported by the search window have been calculated, and then the option to trigger an interrupt is allowed; After triggering an interrupt, the working states supported for the correlation scan matching hardware circuit to jump include but are not limited to the state of reading the multi-resolution map or the search state; Among them, the search window is a set of pose parameters and depth information of nodes, including the pose parameters and depth information of one or more nodes; the parameters required for performing the correlation scan matching algorithm include the pose parameters of nodes, and the pose parameters of nodes include a matching rotation angle, abscissa, and ordinate.

19. The hardware-accelerated positioning system according to claim 4, wherein In the step I, the CPU searches for 4 child nodes in the second map to be searched at the corresponding layer for each branch of the non-leaf node. Among them, these 4 child nodes belong to the same branch and bound search tree, so that the rotation parameters required by these 4 child nodes are the same, the coordinate translation parameters required by these 4 child nodes are different from each other, and the rotation parameters and coordinate translation parameters required by these 4 child nodes are included in the same search window; the rotation parameters and the coordinate translation parameters are both parameters required for performing the correlation scan matching algorithm.

20. The hardware-accelerated positioning system according to claim 19, wherein A bus interface module is externally provided for the hardware acceleration positioning system, and the bus interface module includes a DMA controller module and a transmission bus; The DMA controller module is used to continuously transfer the data stored in the physical storage space with discontinuous addresses in batches, reducing the triggering times of software interrupts of the CPU; The transmission bus includes a first bus and a second bus. The first bus has signal transceiver connections with the memory module, the point cloud processing module, the node search module, the state machine control module, the interconnection bus, and the DMA controller module respectively. The first bus is used to configure data transmission parameters for the DMA controller module, and the first bus is also used to configure the parameters stored inside the map size register, the parameters stored inside the map resolution register, the expansion parameters required for the maximum detection radius reached by the currently stored point cloud, the expansion parameters required for performing the coordinate system transformation, the expansion parameters required for performing the rotation transformation, and the expansion parameters required to meet the transmission requirements of the bus width, so as to realize the mapping communication of the memory of the correlation scan matching hardware circuit; The second bus is connected to the DMA controller module and is used to transmit the multi-level grid map pre-constructed by the CPU, the sine function values under the pre-configured corresponding rotation parameters, the cosine function values under the pre-configured corresponding rotation parameters, and the point cloud currently collected by the lidar to the memory module; Among them, the transmission bus follows the AMBA protocol.

21. A robot, the robot is equipped with a lidar, characterized in that, The robot internally is provided with the hardware acceleration positioning system according to any one of claims 1 to 20.

Citation Information

Patent Citations

  • Robot environment map real-time updating method

    CN113375683A

  • Method and device for real-time mapping and localization

    EP3078935A1