Method and device for generating voxel grid
By separating, superimposing, registering and densely optimizing multi-frame point clouds, the existing OCC representation method has solved the problem of large calculation amount and inaccurate voxel representation, and a voxel grid with small calculation amount and high accuracy is generated.
Patent Information
- Application Number
- CN202510113517.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-23
- Publication Date
- 2025-05-27
AI Technical Summary
The existing OCC representation method has a large amount of calculation and cannot be deployed on more chips. Moreover, the three-dimensional voxel-level representation method cannot accurately reflect the height in reality due to the preset size of the voxels, and the data generated is rough.
By separating the dynamic object point clouds and static scene point clouds in the multi-frame original point clouds, and performing point cloud superposition and registration respectively, combining them for dense optimization, and giving newly generated point semantic categories, the target point clouds are voxelized to generate an accurate voxel grid.
The calculation amount is reduced, the generated voxel grid is more accurate and smooth, supports the deployment of more chips, and can more accurately reflect the height in reality.
Smart Images

Figure CN120047645A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of three-dimensional image processing, and particularly to a method and device for generating a voxel grid. Background Art
[0002] The OCCupancy Networks (OCC) technology can convert point cloud data into a voxel grid, determine the occupancy state of each voxel, and then reconstruct the three-dimensional shapes of various objects in the scene. Generally speaking, it divides the environment into small cubes through an occupancy grid mapping and determines which small cubes are occupied and which are free. Due to its ability to accurately locate objects and recognize object shapes in three-dimensional space, this technology has been applied in the field of autonomous driving. In terms of object detection, this technology can be used to avoid collisions, and in terms of environmental understanding, it can be used to understand complex scenes.
[0003] The existing OCC representation methods are often three-dimensional voxel-level representations. The OCC data includes five dimensions: batch, number of channels, length, width, and height. This method requires a large amount of computation. Moreover, the inference algorithms of some chips in actual deployment only support the first four dimensions, so this method cannot be actually deployed. Secondly, due to the preset size of voxels, the three-dimensional voxel-level representation method cannot accurately reflect the height in reality, and phenomena such as uneven ground will occur during the generation process. In addition, the existing technology does not consider the error interference during the coordinate transformation of point clouds in different frames, which will also lead to rough generated data. Therefore, how to generate an accurate voxel grid while reducing the computation amount has become a problem to be solved. Summary of the Invention
[0004] Based on the above problems, the present application provides a method and device for generating a voxel grid to generate an accurate voxel grid while reducing the computation amount.
[0005] The present application discloses a method for generating a voxel grid, and the method includes:
[0006] Separating the original dynamic object point cloud and the original static scene point cloud in multiple frames of original point clouds;
[0007] Performing point cloud superposition and registration on the original dynamic object point cloud according to the front and back frame point clouds of the original dynamic object point cloud to obtain a dynamic object point cloud;
[0008] Performing point cloud superposition on the original static scene point cloud according to the front and back frame point clouds of the original static scene point cloud to obtain a static scene point cloud;
[0009] Combine the dynamic object point cloud and the static scene point cloud, perform dense optimization on the combined point cloud, and assign semantic categories to the newly generated points in the dense optimization operation to obtain the target point cloud;
[0010] Voxelize the target point cloud to obtain the original voxel grid, and obtain the lowest-height voxel and the highest-height voxel of each pillar in the original voxel grid;
[0011] In the pillar, use the highest height of the point cloud in the lowest-height voxel as the starting height value of the object corresponding to the pillar, use the lowest height of the point cloud in the highest-height voxel as the top height value of the object corresponding to the pillar, and use the semantic category of the lowest-height voxel as the semantic category of the object corresponding to the pillar to generate the voxel grid.
[0012] Optionally, the separating the original dynamic object point cloud and the original static scene point cloud from the multi-frame original point clouds includes:
[0013] Wrap the original dynamic object point cloud with a bounding box and separate it from the original static scene point cloud.
[0014] Optionally, the performing point cloud superposition and registration on the original dynamic object point cloud according to the front and back frame point clouds of the original dynamic object point cloud to obtain the dynamic object point cloud includes:
[0015] Unify multiple frames of original dynamic object point clouds belonging to the same object into the bounding box coordinate system to obtain the superimposed original dynamic object point cloud;
[0016] Map the superimposed original dynamic object point cloud to the camera coordinate system to obtain the object tile corresponding to the superimposed original dynamic object point cloud;
[0017] Reconstruct the object tile into a new point cloud, and register the new point cloud with the superimposed original dynamic object point cloud to obtain the dynamic object point cloud.
[0018] Optionally, the reconstructing the object tile into a new point cloud and registering the new point cloud with the superimposed original dynamic object point cloud to obtain the dynamic object point cloud includes:
[0019] The registration operation is implemented by the following formula:
[0020]
[0021] In the formula, O d represents the dynamic object point cloud, ICP represents the registration operation, p iThe object tile representing the i-th frame, with superscripts f and r representing object tiles from the front-view and rear-view orientations respectively, Merge representing the point cloud fusion operation, and Recon representing the reconstruction operation.
[0022] Optionally, the step of performing point cloud superposition on the original static scene point cloud based on the front and rear frame point clouds of the original static scene point cloud to obtain the static scene point cloud includes:
[0023] Predicting the previous frame predicted point cloud and the next frame predicted point cloud based on the original static scene point cloud of the current frame;
[0024] Transforming the previous frame predicted point cloud and the original static scene point cloud of the previous frame into the same coordinate system, and transforming the next frame predicted point cloud and the original static scene point cloud of the next frame into the same coordinate system to obtain the superimposed static scene point cloud.
[0025] Optionally, the step of predicting the previous frame predicted point cloud and the next frame predicted point cloud based on the original static scene point cloud of the current frame includes:
[0026] Based on the original static scene point cloud of the current frame, predicting the previous frame predicted point cloud and the next frame predicted point cloud in the coordinate system of the point cloud acquisition device of the current frame according to the known speed of the point cloud acquisition device and the time difference between adjacent frames.
[0027] Optionally, the step of transforming the previous frame predicted point cloud and the original static scene point cloud of the previous frame into the same coordinate system, and transforming the next frame predicted point cloud and the original static scene point cloud of the next frame into the same coordinate system to obtain the superimposed static scene point cloud includes:
[0028] The superposition operation is implemented by the following formula:
[0029] O j = Merge(T i (o i , o i + v i * Δt), T i+1 (o i+1 , o i+1 + v i+1 * Δt),...)
[0030] In the formula, O j represents the static scene point cloud, o i represents the i-th frame original static scene point cloud, T i represents the coordinate transformation operation of the i-th frame original static scene point cloud, v i represents the speed of the point cloud acquisition device of the i-th frame, Δt represents the adjacent frame time difference, and Merge represents the point cloud fusion operation.
[0031] Optionally, combining the dynamic object point cloud and the static scene point cloud and performing dense optimization on the combined point cloud includes:
[0032] Rigidly transforming the dynamic object point clouds belonging to the same object through rotation and translation parameters, and restoring them to the static scene point cloud at the position and orientation of the original dynamic object point cloud, and combining the dynamic object point cloud and the static scene point cloud;
[0033] Applying the Poisson surface reconstruction algorithm to perform dense optimization on the combined point cloud.
[0034] Optionally, endowing the newly generated points in the dense optimization operation with semantic categories to obtain a target point cloud includes:
[0035] Applying the nearest neighbor algorithm to endow the newly generated points with semantic categories to obtain a target point cloud.
[0036] Based on the above method for generating a voxel grid, the present application also discloses a device for generating a voxel grid, including: a point cloud separation unit, a point cloud registration unit, a point cloud superposition unit, a point cloud combination unit, a voxelization unit, and a voxel grid generation unit;
[0037] The point cloud separation unit is used to separate the original dynamic object point cloud and the original static scene point cloud from multiple frames of original point clouds;
[0038] The point cloud registration unit is used to perform point cloud superposition and registration on the original dynamic object point cloud according to the front and back frame point clouds of the original dynamic object point cloud to obtain a dynamic object point cloud;
[0039] The point cloud superposition unit is used to perform point cloud superposition on the original static scene point cloud according to the front and back frame point clouds of the original static scene point cloud to obtain a static scene point cloud;
[0040] The point cloud combination unit is used to combine the dynamic object point cloud and the static scene point cloud, perform dense optimization on the combined point cloud, and endow the newly generated points in the dense optimization operation with semantic categories to obtain a target point cloud;
[0041] The voxelization unit is used to voxelize the target point cloud to obtain an original voxel grid, and obtain the lowest height voxel and the highest height voxel of each pillar in the original voxel grid;
[0042] The voxel grid generation unit is configured to generate a voxel grid in the strut. The starting height value of the object corresponding to the strut is the highest height of the point cloud in the lowest height voxel, the top height value of the object corresponding to the strut is the lowest height of the point cloud in the highest height voxel, and the semantic category of the lowest height voxel is the semantic category of the object corresponding to the strut.
[0043] Optionally, the point cloud separation unit is configured to:
[0044] Wrap the original dynamic object point cloud with a bounding box and separate it from the original static scene point cloud.
[0045] Optionally, the point cloud registration unit includes:
[0046] A dynamic transformation sub-unit for uniformly transforming multiple frames of the original dynamic object point cloud belonging to the same object into the bounding box coordinate system to obtain the superimposed original dynamic object point cloud;
[0047] A mapping sub-unit for mapping the superimposed original dynamic object point cloud to the camera coordinate system to obtain the object tile corresponding to the superimposed original dynamic object point cloud;
[0048] A registration sub-unit for reconstructing the object tile into a new point cloud and registering the new point cloud with the superimposed original dynamic object point cloud to obtain the dynamic object point cloud.
[0049] Optionally, the registration sub-unit is configured to:
[0050] The registration operation is implemented by the following formula:
[0051]
[0052] In the formula, O d represents the dynamic object point cloud, ICP represents the registration operation, p i represents the object tile of the i-th frame, and the superscripts f and r respectively represent the object tiles from the front view and rear view orientations. Merge represents the point cloud fusion operation, and Recon represents the reconstruction operation.
[0053] Optionally, the point cloud superposition unit includes:
[0054] A prediction sub-unit for predicting the previous frame prediction point cloud and the next frame prediction point cloud based on the original static scene point cloud of the current frame;
[0055] A static superposition subunit, configured to transform the previous-frame predicted point cloud and the original static scene point cloud of the previous frame into the same coordinate system, and transform the next-frame predicted point cloud and the original static scene point cloud of the next frame into the same coordinate system, so as to obtain the superimposed static scene point cloud.
[0056] Optionally, the prediction subunit is configured to:
[0057] Based on the original static scene point cloud of the current frame, according to the known speed of the point cloud acquisition device and the time difference between adjacent frames, predict the previous-frame predicted point cloud and the next-frame predicted point cloud in the coordinate system of the point cloud acquisition device in the current frame.
[0058] Optionally, the static superposition subunit is configured to:
[0059] The superposition operation is implemented by the following formula:
[0060] O j = Merge(T i (o i , o i + v i * Δt), T i+1 (o i+1 , o i+1 + v i+1 * Δt),...)
[0061] In the formula, O j represents the static scene point cloud, o i represents the i-th frame of the original static scene point cloud, T i represents the coordinate transformation operation of the i-th frame of the original static scene point cloud, v i represents the speed of the point cloud acquisition device in the i-th frame, Δt represents the adjacent frame time difference, and Merge represents the point cloud fusion operation.
[0062] Optionally, the point cloud combination unit includes:
[0063] A combination subunit, configured to perform a rigid transformation on the dynamic object point clouds belonging to the same object through rotation and translation parameters, and restore them to the static scene point cloud in the position and orientation of the original dynamic object point cloud, and combine the dynamic object point cloud and the static scene point cloud;
[0064] A densification subunit, configured to perform densification optimization on the combined point cloud by applying the Poisson surface reconstruction algorithm.
[0065] Optionally, the point cloud combination unit is configured to:
[0066] Apply the nearest neighbor algorithm to assign semantic categories to the newly generated points to obtain the target point cloud.
[0067] The present application discloses a method and apparatus for generating a voxel grid. The raw dynamic object point cloud and the raw static scene point cloud in multiple frames of raw point clouds are separated and respectively superposed and optimized, taking into account the error interference during coordinate transformation of point clouds in different frames. The superposed point clouds are combined to generate more delicate point cloud data. The combined point cloud is densely optimized, and semantic categories are assigned to the newly generated points to obtain the target point cloud. The target point cloud is voxelized to obtain a three-dimensional original voxel grid, and the lowest height voxel and the highest height voxel of each pillar in the original voxel grid are obtained. Taking the highest height of the point cloud in the lowest height voxel as the starting height value of the object corresponding to the pillar, and taking the lowest height of the point cloud in the highest height voxel as the top height value of the object, it is possible to be not limited to the preset size of the voxel, more accurately and smoothly reflect the height in reality, and reduce phenomena such as uneven ground during the generation process. Finally, defining the semantic category of the lowest height voxel as the semantic category of the object, a two-dimensional voxel grid is generated. The method described in the present application requires less computational effort, can be deployed on more chips, and at the same time, the generated voxel grid is more accurate and smooth. BRIEF DESCRIPTION OF THE DRAWINGS
[0068] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are only the embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained according to the provided drawings.
[0069] Figure 1a It is a schematic flowchart of a method for generating a voxel grid disclosed in an embodiment of the present application;
[0070] Figure 1b It is a schematic diagram of the voxel grid disclosed in an embodiment of the present application;
[0071] Figure 1c It is a schematic diagram of generating a ground voxel grid disclosed in an embodiment of the present application;
[0072] Figure 1d It is a schematic diagram of a multi-category voxel grid disclosed in an embodiment of the present application;
[0073] Figure 2 It is a schematic flowchart of another method for generating a voxel grid disclosed in an embodiment of the present application;
[0074] Figure 3 It is a schematic structural diagram of an apparatus for generating a voxel grid disclosed in an embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0075] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in the present application without creative efforts shall fall within the protection scope of the present application.
[0076] Embodiment 1: The present application discloses a method for generating a voxel grid.
[0077] Specifically, please refer to Figure 1a , a method for generating a voxel grid disclosed in this embodiment includes the following steps:
[0078] Step 101: Separate the original dynamic object point cloud and the original static scene point cloud in multiple frames of the original point cloud.
[0079] In the method described in this embodiment, the original point cloud can be obtained by a lidar installed on the vehicle itself. The obtained point cloud is a dense three-dimensional image composed of individual points. The bounding box algorithm is a method for solving the optimal bounding space of a discrete point set. Its basic idea is to use a geometric body with a slightly larger volume and simple characteristics, that is, a bounding box, to approximately replace the point set. Generally speaking, it is to wrap the point cloud with a bounding box. Among them, the shape of the bounding box can be a cube, a sphere, etc. The bounding box can quickly obtain the overlapping relationship between objects and improve the efficiency of geometric operations. It should be noted that the bounding box has an identifier for identifying the object corresponding to the original dynamic object point cloud it wraps. For example, in multiple frames of point clouds, there is a bounding box marked as "A" in each frame, then the object wrapped by the "A" bounding box in each frame is the same object.
[0080] In the method described in this embodiment, the original dynamic object point cloud wrapped by the bounding box can be separated from the original static scene point cloud by using the position, length, width, and height in the bounding box parameters. For example, in a road surface scene, there are an original dynamic object point cloud corresponding to a pedestrian and an original static scene point cloud corresponding to the road. After the "pedestrian point cloud" is wrapped by the bounding box, the "pedestrian point cloud" can be taken out from the "road surface point cloud", leaving only the "road surface point cloud".
[0081] Step 102: Perform point cloud superposition and registration on the original dynamic object point cloud according to the front and rear frame point clouds of the original dynamic object point cloud to obtain a dynamic object point cloud.
[0082] In the method described in this embodiment, the identification of the bounding box can be used to uniformly transform multiple frames (front and rear frames) of the original dynamic object point cloud belonging to the same object into the bounding box coordinate system, perform point cloud superposition, and obtain the superposed original dynamic object point cloud. Then, using the center position of the bounding box (the geometric center of the cubic bounding box, the center of the sphere of the spherical bounding box, etc.), the superposed original dynamic object point cloud is mapped to the camera coordinate system to obtain a camera image, that is, the object patch corresponding to the superposed original dynamic object point cloud. At this time, since the original dynamic object point cloud is obtained by superposing multiple frames of point clouds, the obtained object patch also comes from multiple frames and multiple perspectives.
[0083] In the method described in this embodiment, the object patch can be input into a 3D reconstruction model and reconstructed into a new point cloud. Then, the new point cloud is registered with the superposed original dynamic object point cloud to obtain the dynamic object point cloud. Among them, the registration operation can be implemented by the Iterative Closest Point (ICP) algorithm. The ICP algorithm estimates the corresponding points by the nearest neighbor method. Simply put, for each point in the point cloud to be registered, the corresponding point with the closest distance in the registered point cloud is solved, and then the corresponding points are iteratively optimized by the least squares method. In the method described in this embodiment, the ICP algorithm can be implemented by the following formula:
[0084]
[0085] In the formula, O d represents the dynamic object point cloud, ICP represents the registration operation, p i represents the object patch of the i-th frame, and the superscripts f and r respectively represent the object patches from the front view and rear view orientations. Merge represents the point cloud fusion operation, and Recon represents the reconstruction operation.
[0086] Step 103: Perform point cloud superposition on the original static scene point cloud according to the front and rear frame point clouds of the original static scene point cloud to obtain the static scene point cloud.
[0087] In the method described in this embodiment, when collecting the original point cloud, the lidar may be shaken and jittered, affecting the sensor parameters, resulting in errors when the point clouds of different frames are transformed into the same coordinate system, and the point clouds cannot be aligned, thereby causing the original point cloud to be inaccurate and delicate. For example, there may be problems with uneven ground when superposing the static scene point cloud. Therefore, based on the original static scene point cloud of the current frame, the known speed of the point cloud acquisition device, and the time difference between adjacent frames, the predicted point cloud of the previous frame and the predicted point cloud of the next frame can be predicted. Among them, the point cloud acquisition device in the method described in this embodiment can be the lidar on the vehicle itself, so the speed of the point cloud acquisition device is the vehicle speed itself.
[0088] As an alternative method, the predicted point cloud of the previous frame and the predicted point cloud of the next frame can be predicted in the lidar coordinate system of the current frame. Then, the predicted point cloud of the previous frame and the original static scene point cloud of the previous frame are transformed into the same coordinate system, and the predicted point cloud of the next frame and the original static scene point cloud of the next frame are transformed into the same coordinate system to obtain the superimposed and smoothed static scene point cloud. For example, if the point cloud of the current frame is o 1 , the point cloud of the previous frame is o 0 , and the point cloud of the next frame is o 2 . Then, based on the coordinate system of o 1 , the predicted point cloud o' 0 of the previous frame and the predicted point cloud o' 2 of the next frame that are the same as this coordinate system can be predicted. Transform o' 0 to the same coordinate system as o 0 and perform point cloud superposition to obtain a new o 0 . Transform o' 2 to the same coordinate system as o 2 and perform point cloud superposition to obtain a new o 2 .
[0089] The superposition operation is implemented by the following formula:
[0090] O j = Merge(T i (o i , o i + v i * Δt), T i+1 (o i+1 , o i+1 + v i+1 * Δt),...) In the formula, O j represents the static scene point cloud, o i represents the original static scene point cloud of the i-th frame, T i represents the coordinate transformation operation of the original static scene point cloud of the i-th frame, v i represents the speed of the point cloud acquisition device of the i-th frame, and Δt represents the time difference between adjacent frames.
[0091] Step 104: Combine the dynamic object point cloud and the static scene point cloud, perform dense optimization on the combined point cloud, and assign semantic categories to the newly generated points in the dense optimization operation to obtain the target point cloud.
[0092] In the method described in this embodiment, based on the acquisition position of the original point cloud, the dynamic object point clouds belonging to the same object are rigidly transformed by rotation and translation parameters, and restored to the static scene point cloud according to the position and orientation of the original dynamic object point cloud for combination, so as to supplement the missing part due to the perspective. Then, the Poisson surface reconstruction algorithm is applied to densely optimize the combined point cloud. Among them, the Poisson surface reconstruction algorithm can perform a smoothing operation on the combined point cloud, and new points will be generated during the process, and the finally output point cloud contains the new points. For the newly generated points in this step of operation, in the method described in this embodiment, the nearest neighbor algorithm can be applied to assign semantic categories to the newly generated points to obtain the target point cloud. Briefly speaking, the nearest neighbor algorithm is that when multiple nearest points of the newly generated point in the feature space belong to a certain category, the newly generated point is also considered to belong to this category.
[0093] Step 105: Voxelize the target point cloud to obtain an original voxel grid, and obtain the lowest height voxel and the highest height voxel of each pillar in the original voxel grid.
[0094] In the method described in this embodiment, the target point cloud can be converted into an original voxel grid through coordinate transformation and rounding, and the original voxel grid is a three-dimensional representation method. In the method described in this embodiment, a new voxel grid representation is proposed, which can reduce the three-dimensional voxel representation to a two-dimensional voxel representation, including pillars. Pillars are the output of three-dimensional voxel simplification, and the three-dimensional voxels can be simplified to two-dimensional pillars for two-dimensional convolution processing. For example, the dimension of the original voxel grid (length, width, height) is (200, 200, 16), and the method described in this embodiment reduces the three-dimensional (200, 200, 16) label to three two-dimensional (200, 200) labels, namely the starting height matrix, the top height matrix, and the semantic category matrix. In this two-dimensional label, the height label is not included, but the height is represented by the number of elements of the pillar.
[0095] Step 106: In the pillar, take the highest height of the point cloud in the lowest height voxel as the starting height value of the object corresponding to the pillar, take the lowest height of the point cloud in the highest height voxel as the top height value of the object corresponding to the pillar, and take the semantic category of the lowest height voxel as the semantic category of the object corresponding to the pillar to generate a voxel grid.
[0096] In the method described in this embodiment, based on the example in step 105, as an optional method, this operation can traverse all the pillars containing 16 elements in the two-dimensional (200, 200) label in a bird's-eye view (top-down perspective) to obtain the representation of the height.
[0097] In the method described in this embodiment, for each traversed pillar, the lowest height voxel index and the highest height voxel index of the pillar in the original voxel grid are obtained through an index acquisition algorithm. Among them, as an optional method, the index acquisition algorithm may specifically include: traversing from the bottom end to the top end of the pillar to determine whether there are objects on the ground and above the ground, that is, whether the voxels are occupied. For the found occupied objects, continue to find their lowest height voxels and highest height voxels. For example, Figure 1b is a schematic diagram of the voxel grid disclosed in the embodiment of the present application, Figure 1b The dots in it are the point clouds representing an object, and the squares are the voxels representing the object. The voxel with the highest position, that is, the voxel in the upper row, is the highest height voxel of the object, and the three voxels in the lower row are the lowest height voxels of the object.
[0098] In the method described in this embodiment, the point clouds included are found through the two obtained voxel index values. The highest height of the point cloud in the lowest height voxel is used as the starting height value of the object corresponding to the pillar, the lowest height of the point cloud in the highest height voxel is used as the top height value of the object corresponding to the pillar, and the semantic category of the lowest height voxel is used as the semantic category of the object corresponding to the pillar. Thus, the voxel grid represented by the two-dimensional voxels of the method described in this embodiment is obtained. This operation of defining the height value is to make the voxel grid generated by the method described in this embodiment more accurate and smooth.
[0099] As an optional example, Figure 1c is a schematic diagram of generating a ground voxel grid disclosed in the embodiment of the present application. As Figure 1c shown, the squares in the figure are ground voxels, the dots are ground point clouds, and the upper right corner is the coordinate system. Figure 1c The left half of is the traditional voxel grid representation method, and the right half is the method described in this embodiment. It can be seen that the traditional method uses the height of the center point of the entire ground voxel, that is, the square, to represent the ground, that is, h 1 . The method described in this embodiment will use the height of the lowest ground point cloud in the highest voxel and the height difference between the highest ground point cloud in the lowest voxel to represent the ground, that is, h 2 . Therefore, when the ground point clouds are the same, the ground obtained by the traditional method is uneven, and the ground obtained by the method described in this embodiment is smoother than the traditional method. Comparing with the road surface height difference h, it can be clearly seen that h 1 is greater than h 2 , that is, the method described in this embodiment can more truly reflect the ground situation.
[0100] Moreover, in a complex real-world environment, there may be situations where the sidewalk and the road surface have similar heights. To distinguish the categories of objects with similar heights, the method described in this embodiment determines the category of an object based on the semantic category of the lowest-height voxel. Figure 1d It is a schematic diagram of a multi-category voxel grid disclosed in an embodiment of the present application. As Figure 1d shown, the two dashed arrows respectively represent two non-occupied areas. There is a cone standing on the road surface under the left non-occupied area, and there are no other objects on the cone. Therefore, in this non-occupied area, the highest-height point cloud in the lowest-height voxel is the top height value of the cone, the lowest-height point cloud in the highest-height voxel is the height of the sky (when it is determined that the area above an object is the sky when there are no other objects above the object), and the semantic value of the lowest-height voxel is the cone category. The right non-occupied area is under a tree. Then, in this non-occupied area, the highest-height point cloud in the lowest-height voxel is the top height value of the ground, the lowest-height point cloud in the highest-height voxel is the lowest height of the leaf point cloud, and the semantic value of the lowest-height voxel is the ground category. The letters in the figure represent semantic categories, C represents a cone, G represents the ground, T represents the grassland terrain, and V represents vegetation.
[0101] The method described in this embodiment uses a two-dimensional voxel grid representation method, reducing the data dimension compared with traditional methods, and correspondingly reducing the calculation amount and prediction parameters. Representing the three-dimensional height in the form of a pillar also better supports the inference and actual deployment of the method in this embodiment on most chips. In addition, the height of an object is based on the actual height of the point cloud, which can output a voxel grid more accurately and smoothly, and is more in line with the real-world environment. At the same time, the semantic category label of an object can also assist in reflecting the road conditions, helping the ego vehicle to achieve more accurate control in vehicle obstacle avoidance. On the other hand, the method in this embodiment optimizes the dynamic object point cloud and the static scene point cloud respectively to improve the point cloud quality. For the dynamic object point cloud, the point cloud holes caused by missing perspectives are supplemented, and for the static scene point cloud, the error interference caused by inevitable situations such as vibrations during the movement process is reduced.
[0102] Embodiment 2: The present application discloses another method for generating a voxel grid. Please refer to Figure 2 This embodiment describes the entire process of generating a voxel grid.
[0103] Step 201: Use the bounding box algorithm to separate the original dynamic object point cloud and the original static scene point cloud in the original point cloud.
[0104] Step 202: Overlay the original dynamic object point clouds representing the same object in the front and back frame point clouds of the original dynamic object point cloud.
[0105] Step 203: Convert the overlaid original dynamic object point cloud into a camera image and perform three-dimensional reconstruction.
[0106] Step 204: Register the newly reconstructed point cloud with the superimposed original dynamic object point cloud to obtain the dynamic object point cloud.
[0107] Step 205: Predict the front and back frame predicted point clouds of the original static scene point cloud.
[0108] Step 206: Superimpose the front and back frame predicted point clouds with the front and back frame original static scene point clouds to obtain the static scene point cloud.
[0109] Step 207: Combine the dynamic object point cloud and the static scene point cloud, and apply the Poisson surface reconstruction algorithm to densely optimize the combined point cloud.
[0110] Step 208: Apply the nearest neighbor algorithm to assign semantic categories to the newly generated points of the Poisson surface reconstruction algorithm to obtain the target point cloud.
[0111] Step 209: Voxelize the target point cloud to obtain the original voxel grid, and obtain the lowest height voxel and the highest height voxel of each pillar in the original voxel grid.
[0112] In the method described in this embodiment, the lowest height voxel and the highest height voxel of each pillar in the original voxel grid can be obtained through the index acquisition algorithm. This algorithm traverses from the bottom end to the top end of the pillar to determine whether there are objects above the road surface, that is, whether the voxel is occupied. Among them, as an example, the height of the pillar can be set to 16. For the found occupied entity (object), continue to find the index of its lowest height voxel and the index of its highest height voxel. The point cloud included in these two indices is the point cloud representing the occupied entity. When the index of the highest height voxel cannot be found, it is possible that there are no other objects above the road surface (this object).
[0113] Step 210: Based on the lowest height voxel and the highest height voxel, convert the three-dimensional represented original voxel grid into a two-dimensional represented voxel grid.
[0114] In the method described in this embodiment, the method of defining the object height and semantic category in the two-dimensional represented voxel grid based on the lowest height voxel and the highest height voxel is the same as that in Embodiment 1, and will not be repeated here.
[0115] Based on the method of generating a voxel grid disclosed in the above embodiment, this embodiment correspondingly discloses a device for generating a voxel grid. Please refer to Figure 3 The device for generating a voxel grid includes: a point cloud separation unit 301, a point cloud registration unit 302, a point cloud superposition unit 303, a point cloud combination unit 304, a voxelization unit 305, and a voxel grid generation unit 306;
[0116] The point cloud separation unit 301 is configured to separate the original dynamic object point cloud and the original static scene point cloud from multiple frames of original point clouds;
[0117] The point cloud registration unit 302 is configured to perform point cloud superposition and registration on the original dynamic object point cloud according to the front and back frame point clouds of the original dynamic object point cloud to obtain a dynamic object point cloud;
[0118] The point cloud superposition unit 303 is configured to perform point cloud superposition on the original static scene point cloud according to the front and back frame point clouds of the original static scene point cloud to obtain a static scene point cloud;
[0119] The point cloud combination unit 304 is configured to combine the dynamic object point cloud and the static scene point cloud, perform dense optimization on the combined point cloud, and assign semantic categories to the newly generated points in the dense optimization operation to obtain a target point cloud;
[0120] The voxelization unit 305 is configured to voxelize the target point cloud to obtain an original voxel grid, and obtain the lowest height voxel and the highest height voxel of each strut in the original voxel grid;
[0121] The voxel grid generation unit 306 is configured to, in the strut, use the highest height of the point cloud in the lowest height voxel as the starting height value of the object corresponding to the strut, use the lowest height of the point cloud in the highest height voxel as the top height value of the object corresponding to the strut, and use the semantic category of the lowest height voxel as the semantic category of the object corresponding to the strut to generate a voxel grid.
[0122] Optionally, the point cloud separation unit 301 is configured to:
[0123] Wrap the original dynamic object point cloud with a bounding box and separate it from the original static scene point cloud.
[0124] Optionally, the point cloud registration unit 302 includes:
[0125] A dynamic conversion sub-unit is configured to uniformly convert multiple frames of original dynamic object point clouds belonging to the same object into the bounding box coordinate system to obtain a superimposed original dynamic object point cloud;
[0126] A mapping sub-unit is configured to map the superimposed original dynamic object point cloud to the camera coordinate system to obtain an object tile corresponding to the superimposed original dynamic object point cloud;
[0127] A registration sub-unit is configured to reconstruct the object tile into a new point cloud and register the new point cloud with the superimposed original dynamic object point cloud to obtain the dynamic object point cloud.
[0128] Optionally, the registration subunit is configured to:
[0129] The registration operation is implemented by the following formula:
[0130]
[0131] In the formula, O d represents the point cloud of the dynamic object, ICP represents the registration operation, p i represents the object tile of the i-th frame, and the superscripts f and r respectively represent the object tiles from the front view and the rear view directions. Merge represents the point cloud fusion operation, and Recon represents the reconstruction operation.
[0132] Optionally, the point cloud superposition unit 303 includes:
[0133] A prediction subunit, configured to predict a previous-frame predicted point cloud and a next-frame predicted point cloud according to the original static scene point cloud of the current frame;
[0134] A static superposition subunit, configured to transform the previous-frame predicted point cloud and the original static scene point cloud of the previous frame into the same coordinate system, and transform the next-frame predicted point cloud and the original static scene point cloud of the next frame into the same coordinate system, to obtain the superposed static scene point cloud.
[0135] Optionally, the prediction subunit is configured to:
[0136] Based on the original static scene point cloud of the current frame, according to the known speed of the point cloud acquisition device and the time difference between adjacent frames, predict the previous-frame predicted point cloud and the next-frame predicted point cloud in the coordinate system of the point cloud acquisition device in the current frame.
[0137] Optionally, the static superposition subunit is configured to:
[0138] The superposition operation is implemented by the following formula:
[0139] O j = Merge(T i (o i , o i + v i * Δt), T i+1 (o i+1 , o i+1 + v i+1 * Δt),...)
[0140] In the formula, O j represents the static scene point cloud, o i represents the original static scene point cloud of the i-th frame, T i represents the coordinate transformation operation of the original static scene point cloud of the i-th frame, v i$v_i$ represents the velocity of the point cloud acquisition device in the $i$-th frame, $\Delta t$ represents the time difference between adjacent frames, and Merge represents the point cloud fusion operation.
[0141] Optionally, the point cloud combining unit 304 includes:
[0142] A combining subunit, configured to perform a rigid transformation on the dynamic object point clouds belonging to the same object through rotation and translation parameters, and restore them to the static scene point cloud at the positions and orientations of the original dynamic object point clouds, and combine the dynamic object point clouds and the static scene point clouds;
[0143] A densification subunit, configured to perform densification optimization on the combined point cloud by applying the Poisson surface reconstruction algorithm.
[0144] Optionally, the point cloud combining unit 304 is configured to:
[0145] Apply the nearest neighbor algorithm to assign semantic categories to the newly generated points to obtain a target point cloud.
[0146] The embodiments in this specification are described in a progressive manner. For the apparatuses disclosed in the embodiments, since they correspond to the methods disclosed in the embodiments, the descriptions are relatively simple. For related parts, refer to the descriptions in the method section.
[0147] It should also be noted that in this article, relational terms such as "first" and "second" are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variation thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements, but also includes other elements not expressly listed, or elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "including a..." does not exclude the existence of additional identical elements in the process, method, article or device including the element.
[0148] The steps of the methods or algorithms described in combination with the embodiments disclosed in this article can be directly implemented by hardware, software modules executed by a processor, or a combination of both. The software modules can be placed in a random access memory (RAM), memory, read-only memory (ROM), electrically programmable ROM, electrically erasable programmable ROM, registers, hard disk, removable disk, CD-ROM, or any other form of storage medium well-known in the technical field.
[0149] The features described in the embodiments in this specification can be replaced or combined with each other, enabling those skilled in the art to implement or use this application.
[0150] The above description of the disclosed embodiments enables those skilled in the art to implement or use this application. Various modifications to these embodiments will be obvious to those skilled in the art, and the general principles defined herein can be implemented in other embodiments without departing from the spirit or scope of this application. Therefore, this application will not be limited to the embodiments shown herein, but rather to the broadest scope consistent with the principles and novel features disclosed herein.
Claims
1. A method for generating a voxel grid, characterized in that: include: Separate the original dynamic object point cloud and the original static scene point cloud in the multi-frame original point cloud; According to the point clouds of the preceding and following frames of the original dynamic object point cloud, the original dynamic object point cloud is superimposed and registered to obtain a dynamic object point cloud; According to the previous and next frame point clouds of the original static scene point cloud, the original static scene point cloud is superimposed to obtain a static scene point cloud; Combining the dynamic object point cloud with the static scene point cloud, performing dense optimization on the combined point cloud, and assigning semantic categories to the points newly generated in the dense optimization operation to obtain a target point cloud; voxelize the target point cloud to obtain an original voxel grid, and obtain the lowest height voxel and the highest height voxel of each pillar in the original voxel grid; In the pillar, the highest height of the point cloud in the lowest height voxel is used as the starting height value of the object corresponding to the pillar, the lowest height of the point cloud in the highest height voxel is used as the top height value of the object corresponding to the pillar, and the semantic category of the lowest height voxel is used as the semantic category of the object corresponding to the pillar to generate a voxel grid.
2. The method according to claim 1, characterized in that The step of separating the original dynamic object point cloud and the original static scene point cloud in the multiple frames of original point cloud comprises: The original dynamic object point cloud is wrapped by a bounding box to separate it from the original static scene point cloud.
3. The method according to claim 2, characterized in that The step of performing point cloud superposition and registration on the original dynamic object point cloud according to the previous and next frame point clouds of the original dynamic object point cloud to obtain the dynamic object point cloud comprises: The original dynamic object point clouds of multiple frames belonging to the same object are uniformly converted into the bounding box coordinate system to obtain the superimposed original dynamic object point clouds; Mapping the superimposed original dynamic object point cloud to the camera coordinate system to obtain an object block corresponding to the superimposed original dynamic object point cloud; The object image block is reconstructed into a new point cloud, and the new point cloud is registered with the superimposed original dynamic object point cloud to obtain the dynamic object point cloud.
4. The method according to claim 3, characterized in that The step of reconstructing the object block into a new point cloud and registering the new point cloud with the superimposed original dynamic object point cloud to obtain the dynamic object point cloud includes: The registration operation is implemented as follows: In the formula, O d represents the dynamic object point cloud, ICP represents the registration operation, p i represents the object block in the i-th frame, the superscripts f and r represent the object blocks from the front and rear view directions respectively, Merge represents the point cloud fusion operation, and Recon represents the reconstruction operation.
5. The method according to claim 1, characterized in that The step of performing point cloud superposition on the original static scene point cloud according to the previous and next frame point clouds of the original static scene point cloud to obtain the static scene point cloud comprises: According to the original static scene point cloud of the current frame, the predicted point cloud of the previous frame and the predicted point cloud of the next frame are predicted; The predicted point cloud of the previous frame and the original static scene point cloud of the previous frame are transformed into the same coordinate system, and the predicted point cloud of the next frame and the original static scene point cloud of the next frame are transformed into the same coordinate system to obtain the superimposed static scene point cloud.
6. The method according to claim 5, characterized in that The step of predicting a predicted point cloud of a previous frame and a predicted point cloud of a next frame according to the original static scene point cloud of the current frame includes: Based on the original static scene point cloud of the current frame, the predicted point cloud of the previous frame and the predicted point cloud of the next frame are predicted in the coordinate system of the point cloud acquisition device of the current frame according to the known speed of the point cloud acquisition device and the time difference between adjacent frames.
7. The method according to claim 6, characterized in that The step of transforming the previous frame prediction point cloud and the previous frame original static scene point cloud into the same coordinate system, and transforming the next frame prediction point cloud and the next frame original static scene point cloud into the same coordinate system to obtain the superimposed static scene point cloud includes: The superposition operation is achieved by: o j =Merge(T i (o i ,o i +v i *Δt),T i+1 (o i+1 ,o i+1 +v i+1 *Δt),…) In the formula, O j represents the static scene point cloud, o i represents the original static scene point cloud of the i-th frame, T i represents the coordinate transformation operation of the original static scene point cloud of the i-th frame, v i represents the speed of the point cloud acquisition device in the i-th frame, Δt represents the time difference between adjacent frames, and Merge represents the point cloud fusion operation.
8. The method according to claim 1, characterized in that The step of combining the dynamic object point cloud and the static scene point cloud and performing dense optimization on the combined point cloud includes: The dynamic object point cloud belonging to the same object is rigidly transformed by rotation and translation parameters, and the position and orientation of the original dynamic object point cloud are restored to the static scene point cloud, and the dynamic object point cloud and the static scene point cloud are combined; The Poisson surface reconstruction algorithm is applied to perform dense optimization on the combined point cloud.
9. The method according to claim 1, characterized in that: The step of assigning semantic categories to the points newly generated in the dense optimization operation to obtain a target point cloud includes: The nearest neighbor algorithm is applied to assign semantic categories to the newly generated points to obtain a target point cloud.
10. A device for generating a voxel grid, characterized in that: include: Point cloud separation unit, point cloud registration unit, point cloud superposition unit, point cloud combination unit, voxelization unit and voxel grid generation unit; The point cloud separation unit is used to separate the original dynamic object point cloud and the original static scene point cloud in the multiple frames of original point cloud; The point cloud registration unit is used to perform point cloud superposition and registration on the original dynamic object point cloud according to the previous and next frame point clouds of the original dynamic object point cloud to obtain the dynamic object point cloud; The point cloud overlay unit is used to perform point cloud overlay on the original static scene point cloud according to the previous and next frame point clouds of the original static scene point cloud to obtain a static scene point cloud; The point cloud combining unit is used to combine the dynamic object point cloud and the static scene point cloud, perform dense optimization on the combined point cloud, and assign semantic categories to the points newly generated in the dense optimization operation to obtain a target point cloud; The voxelization unit is used to voxelize the target point cloud to obtain an original voxel grid, and obtain the lowest height voxel and the highest height voxel of each pillar in the original voxel grid; The voxel grid generation unit is used to generate a voxel grid in the pillar by taking the highest height of the point cloud in the lowest height voxel as the starting height value of the object corresponding to the pillar, taking the lowest height of the point cloud in the highest height voxel as the top height value of the object corresponding to the pillar, and taking the semantic category of the lowest height voxel as the semantic category of the object corresponding to the pillar.