A Fast Dense Occupancy Ground Truth Generation Method and Device

By performing front and back scene segmentation and rasterization of point clouds, intensive occupation truth values are generated, which solves the object recognition problem of visual 3D perception method and the high cost of sparse radar point cloud computing, and realizes efficient reconstruction of autonomous driving scenarios.

CN119941589BActive Publication Date: 2025-07-08ADVANCED TECH RES INST OF BEIJING UNIV OF TECH +3
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510445293.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-04-10
Publication Date
2025-07-08
Estimated Expiration
2045-04-10

AI Technical Summary

Technical Problem

In the prior art, vision-based 3D perception methods are difficult to identify all categories of objects in autonomous driving, and the three-dimensional generated by sparse radar point clouds occupies sparseness, high calculation cost and high time consumption.

Method used

By obtaining the semantics of the driving scene and the point cloud marked by the three-dimensional bounding box, front and back scene segmentation and rasterization are performed, and the foreground and back scene goals are processed respectively by fine and rough rasters, combining Poisson reconstruction and general raster models to generate intensive occupancy truth values.

Benefits of technology

It reduces the calculation time cost, improves the density and accuracy of the truth value occupies, and improves the training effect of subsequent supervision models.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119941589B_ABST
    Figure CN119941589B_ABST
Patent Text Reader

Abstract

The present invention relates to a method and device for quickly generating dense occupancy ground truth. The present invention obtains point clouds annotated with semantics and 3D bounding boxes in a driving scenario, extracts consecutive frame point clouds, and transforms each consecutive frame point cloud into a unified coordinate system; divides the foreground and background frame point clouds according to semantics and 3D bounding boxes; takes a set number of consecutive frame point clouds of foreground targets for point cloud superposition to obtain foreground target superposed point clouds, and in the corresponding time period, extracts key frame point clouds from the consecutive frame point clouds of background targets at a longer step size for point cloud superposition to obtain background target superposed point clouds; determines the ground for the background part, and performs ground paving point processing on the determined ground to obtain a dense ground; rasterizes the foreground targets using a fine grid and rasterizes the background targets using a rough grid; completes the geometric appearance of the foreground targets according to the geometric appearance missing degree of the foreground targets; merges the rasterized foreground targets with the background raster to generate occupancy ground truth.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of generating dense occupancy ground truth based on point clouds, and in particular, to a fast method and device for generating dense occupancy ground truth. Background Art

[0002] Understanding the 3D geometric structure of the surrounding scene is a fundamental step in autonomous driving systems. Although LiDAR is a direct solution for capturing geometric information, its high-cost sensors and sparse scanning points limit its further application. In recent years, vision-centered autonomous driving has attracted extensive attention as a promising direction. By taking multi-camera images as input, it has shown competitiveness in various 3D perception tasks, including depth estimation, 3D object detection, and semantic map construction. Although 3D object detection using multi-cameras has achieved certain results in vision-based 3D perception, it is vulnerable to the long-tail problem and difficult to identify all categories of objects in the real world. One improvement method is to supervise 3D object detection using multi-cameras through the three-dimensional occupancy of objects in the scene. Three-dimensional occupancy describes the scene by assigning occupancy probabilities to each voxel in the three-dimensional space. Three-dimensional occupancy is a good representation for multi-camera scene reconstruction because it naturally guarantees multi-camera geometric consistency and can recover occluded parts. In the prior art, many network models use sparse radar point clouds for supervision and predict three-dimensional occupancy, but the three-dimensional occupancy generated in this way is sparse and has a poor effect as an occupancy label. Existing occupancy ground truth generation algorithms use consecutive frame semantic point clouds as input. First, dynamic objects in the semantic point clouds are filtered, and the point clouds are segmented into static point clouds and multiple dynamic point clouds. Subsequently, the dynamic and static point clouds are respectively stacked according to the world coordinates and three-dimensional ground truth box information to achieve densification. After that, Poisson reconstruction is performed on the complete scene obtained by merging the dense dynamic and static point clouds to achieve hole filling on the surface of a single point cloud object and densification inside. Finally, the scene after Poisson reconstruction is assigned semantic ground truth through the nearest neighbor algorithm to obtain the occupancy ground truth result. The advantage of this method is that the generated occupancy ground truth has a high fine-grained level and is relatively accurate for scene reconstruction. The disadvantage is that the process of Poisson reconstruction for large-scale scenes consumes a large amount of computing power and has a huge time cost. Summary of the Invention

[0003] To solve the above technical problems or at least partially solve the above technical problems, the present invention provides a fast method and device for generating dense occupancy ground truth.

[0004] In a first aspect, the present invention provides a fast method for generating dense occupancy ground truth, including:

[0005] Obtain the point cloud labeled with semantics and three-dimensional bounding boxes of the driving scene, extract consecutive frame point clouds, and convert the consecutive frame point clouds into a unified coordinate system;

[0006] Foreground and background segmentation of each frame of point cloud is performed according to semantics and 3D bounding boxes to obtain foreground targets and background targets. The continuous frame point clouds are decoupled into continuous frame point clouds of foreground targets and continuous frame point clouds of background targets. Among them, the foreground targets are movable targets in the driving scene, and the background targets are immovable targets in the driving scene.

[0007] Take a set number of continuous frame point clouds of foreground targets for point cloud superposition to obtain superposed point clouds of foreground targets. In the time period corresponding to the set number of continuous frame point clouds of foreground targets, key frame point clouds are extracted from the continuous frame point clouds of background targets at a longer step length for point cloud superposition to obtain superposed point clouds of background targets. Determine the ground for the background part, and perform dense processing on the determined ground paving points to obtain a dense ground.

[0008] Use a fine grid to rasterize the foreground targets and a rough grid to rasterize the background targets.

[0009] After rasterization is completed, determine the degree of geometric appearance loss of the foreground targets, and perform geometric appearance completion on the foreground targets according to the degree of geometric appearance loss.

[0010] Merge the rasterized foreground targets with geometric appearance completion and the background raster to generate a complete occupancy ground truth for the autonomous driving scene.

[0011] Furthermore, determining the ground for the background part and performing dense processing on the determined ground paving points to obtain a dense ground includes:

[0012] S410, divide the superposed point clouds of background targets into concentric circular ring regions around the detection vehicle:

[0013] , ;

[0014] Among them, the th concentric circular ring region in the superposed point clouds of background targets , is the inner diameter of the th concentric circular ring region, is the outer diameter of the th

[0015] concentric circular ring region, and is the radial distance of point

[0016] ,

[0017] Among them, is the m-th in the radial direction and the n-th sub-region in the circumferential direction within the concentric circular ring region. The concentric circular ring region is divided into M regions in the radial direction and N regions in the circumferential direction. , , is the point the radial distance from the concentric region to the inner diameter. is the point azimuth angle;

[0018] S430, a height threshold dynamically adjusted according to the average value and standard deviation of the ground elevation, removes the bottom points with a height lower than the height threshold and a reflection intensity lower than the set reflection intensity threshold;

[0019] S440, performs vertical plane fitting on all sub-regions to remove the vertical points on the vertical plane;

[0020] S450, performs ground plane fitting on all sub-regions to obtain all the ground planes in the sub-regions;

[0021] S460, uses the normal uprightness, elevation, and flatness of the ground plane point set, and the normal uprightness, elevation, and flatness of all the ground in the concentric circular region where the ground flat point set is located for ground matching to exclude the non-ground ground plane point sets in each ground plane point set and determine the ground

[0022] S470, after determining the ground, constructs a uniform set of points to cover the ground.

[0023] Furthermore, the sum of the average value of all the determined ground in the concentric circular region and the gain standard deviation is used as the new height threshold.

[0024] Furthermore, the process of vertical plane fitting includes:

[0025] S441, starting from any lowest point in any sub-region, selects a set number of points as seed points and calculates the position mean and unit normal vector of the seed points;

[0026] S442, for any candidate point not determined as a vertical point, when the product of the difference between its position and the position mean of the seed segment and the unit normal vector of the seed point is less than the set distance tolerance for estimating the vertical plane, this candidate point is selected as a potential vertical point. This is performed multiple times to select the potential vertical points in the sub-region;

[0027] S443, if the reciprocal of the cosine of the product between the unit normal vector of the potential vertical point and the unit normal vector of the z-axis is greater than the difference between

[0028] Furthermore, fitting the ground plane to all sub-regions to obtain all ground planes in the sub-regions includes:

[0029] S451: Select a set number of lowest points from any sub-region as seed points, calculate the average height of the seed points, and use the points with a height lower than the sum of the average height and the height threshold as the initially estimated ground plane point set;

[0030] S452: Calculate the normal vector of the ground plane point set obtained each time, calculate the average position of the ground plane point set obtained each time, and calculate the plane coefficient of the ground plane point set using the normal vector of the ground plane point set and the average position of the ground plane point set;

[0031] S453: When iteratively selecting subsequent ground plane point sets, according to the coordinates of the candidate points and the normal vector of the previous ground plane point set, evaluate according to the plane coefficient formula. When the difference between the obtained value and the plane coefficient of the previous ground plane point set is less than the set plane distance threshold, use the candidate point as an element of the next ground plane point set.

[0032] Furthermore, based on the normal vector of the ground plane point set and the angle threshold the uprightness screening function is:

[0033] ;

[0034] Based on the average height of the points in the ground plane point set and the elevation threshold, the elevation screening function is:

[0035] ;

[0036] Among them, is the elevation threshold that changes exponentially with the distance of the point to the center , is the elevation recognizable range threshold;

[0037] Based on the flatness function of the minimum eigenvalue of the ground plane point set is:

[0038] ;

[0039] Among them, is the gain amplitude, is the minimum eigenvalue of the ground plane point set, is the flatness threshold. For the part where the elevation screening function is less than 0.5, that is, the difference between the average height of the points in the ground plane point set and the elevation threshold is greater than zero. If the ground plane point set is flat, a relatively large value will be assigned through the flatness function, so as to restore the lost true examples.

[0040] Further, the mean value of the elevation thresholds and flatness thresholds of all the determined ground surfaces in the concentric annular region and the sum of the gain standard deviation, where the gain standard deviation is the product of the gain coefficient and the standard deviation, are used as the new threshold; within a recent period of time, the mean value of the elevation thresholds and flatness thresholds of all the determined ground surfaces in the concentric annular region and the sum of the gain standard deviation are used as the recovery threshold. When the threshold is less than the recovery threshold, the ground plane point set between the threshold and the recovery threshold is restored.

[0041] Further, the geometric appearance missing degree of the foreground target is determined, and the geometric appearance completion of the foreground target according to the geometric appearance missing degree includes:

[0042] S610. First, based on the three-dimensional intersection over union (IoU) between the actual three-dimensional bounding box and the three-dimensional ground truth box of the foreground target's point cloud, the three-dimensional IoU is used as the geometric appearance missing degree;

[0043] The three-dimensional IoU is compared with the set three-dimensional IoU threshold. When the three-dimensional IoU is lower than the set three-dimensional IoU threshold, S630 is executed; when the three-dimensional IoU is not lower than the set three-dimensional IoU threshold, S640 is executed;

[0044] S630. Instead, a pre-constructed dense occupancy grid model is selected according to the semantic category;

[0045] S640. Poisson reconstruction is used for geometric appearance completion.

[0046] Further, Poisson reconstruction is used for geometric appearance completion, and the detailed steps of the completion process include:

[0047] S641. One of the two symmetric planes perpendicular to the horizontal plane of the three-dimensional ground truth box is used as the mirror plane, and one mirroring is performed based on maximizing the new occupancy grid result generated after mirroring;

[0048] S642. The center point set of the foreground target after mirroring is regarded as point cloud for Poisson reconstruction to generate a continuous three-dimensional network;

[0049] S643. The center points of the grids in the space covered by the three-dimensional network of Poisson reconstruction are used as the grid center points generated by the geometric appearance completion, and the category is unified as the foreground target category.

[0050] In a second aspect, the present invention provides a fast dense occupancy ground truth generation device, including: at least one processing unit, the processing unit is connected to a storage unit through a bus unit, the storage unit stores a computer program, and when the computer program is executed by the processing unit, the fast dense occupancy ground truth generation method as described above is implemented.

[0051] The above technical solution provided by the embodiments of the present invention has the following advantages compared with the prior art:

[0052] Obtain the point cloud labeled with semantics and 3D bounding boxes in the driving scenario, extract the continuous frame point cloud, and convert the continuous frame point cloud into a unified coordinate system; perform foreground and background segmentation on each frame of the point cloud according to the semantics and 3D bounding boxes to obtain foreground targets and background targets, and the continuous frame point cloud is decoupled into the foreground target continuous frame point cloud and the background target continuous frame point cloud; among them, the foreground target is a movable target in the driving scenario, and the background target is an immovable target in the driving scenario; take a set number of foreground target continuous frame point clouds for point cloud superposition to obtain the foreground target superposed point cloud, and in the time period corresponding to the set number of foreground target continuous frame point clouds, extract key frame point clouds from the background target continuous frame point cloud at a longer step size for point cloud superposition to obtain the background target superposed point cloud; determine the ground for the background part, and perform dense processing on the determined ground paving points to obtain a dense ground; rasterize the foreground target using a fine grid and rasterize the background target using a rough grid; after rasterization, determine the geometric appearance missing degree of the foreground target, and complete the geometric appearance of the foreground target according to the geometric appearance missing degree; merge the raster of the foreground target with the geometric appearance completed with the background raster to generate a complete occupancy ground truth for the autonomous driving scenario. The fast dense occupancy ground truth generation method proposed in this patent avoids the reconstruction process of a large range of backgrounds, and the object of Poisson reconstruction is only the foreground target point cloud downsampled by rasterization, which greatly reduces the time cost of the calculation process. For the ground part, dense processing is performed by ground matching and adding points, replacing Poisson reconstruction, which greatly reduces the calculation amount. In addition, using a general grid model instead of the foreground target with serious geometric appearance loss greatly improves the number of dirty data in the occupancy ground truth, making the ground truth have a better effect on the training of the subsequent supervised occupancy prediction model. BRIEF DESCRIPTION OF THE DRAWINGS

[0053] The drawings here are incorporated into the specification and form a part of this specification, showing the embodiments consistent with the present invention, and are used together with the specification to explain the principles of the present invention.

[0054] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, for those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0055] Figure 1 It is a flowchart of a fast dense occupancy ground truth generation method provided by an embodiment of the present invention;

[0056] Figure 2Schematic diagram of the evolution from consecutive frame point clouds to the final dense occupancy ground truth during the execution of the fast dense occupancy ground truth generation method provided by the embodiments of the present invention;

[0057] Figure 3 Flowchart for determining the ground for a part of the background and performing dense processing on the determined ground points to obtain a dense ground provided by the embodiments of the present invention;

[0058] Figure 4 Flowchart for fitting the vertical plane of the point cloud provided by the embodiments of the present invention;

[0059] Figure 5 Flowchart for fitting the point cloud ground plane for all sub-regions to obtain all the ground planes in the sub-regions provided by the embodiments of the present invention;

[0060] Figure 6 Flowchart for fitting the point cloud ground plane for all sub-regions to obtain all the ground planes in the sub-regions provided by the embodiments of the present invention;

[0061] Figure 7 Schematic diagram of the fast dense occupancy ground truth generation device provided by the embodiments of the present invention. Detailed implementation manners

[0062] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0063] It should be noted that in this article, the term "comprising", "including" or any other variation thereof is intended to cover a non-exclusive inclusion, such that a process, method, article, or device comprising a series of elements not only includes those elements but also includes other elements not expressly listed, or further includes elements inherent to such process, method, article, or device. Without further limitation, an element defined by the phrase "comprising a..." does not exclude the existence of additional identical elements in the process, method, article, or device comprising the element.

[0064] Embodiment 1

[0065] The technology of the present invention realizes a fast dense occupancy ground truth generation method, aiming to use dense occupancy grids to label targets in complex autonomous driving scenarios. The dense occupancy grids of the targets can be used to train a neural network model for target detection or semantic segmentation tasks.

[0066] Combination Figure 1 and Figure 2 As shown, the fast dense occupancy truth value generation method provided by the present application includes:

[0067] S100, obtaining the point cloud of the driving scene annotated with semantics and 3D bounding boxes, and extracting continuous frame point clouds ,in, Indicates time The point cloud frame at each moment contains several points ,in, For the moment The first point cloud frame at Convert the continuous frame point clouds into a unified coordinate system.

[0068] S200: Segment the foreground and background of each frame point cloud according to semantics and 3D bounding boxes to obtain foreground targets and background targets. The continuous frame point clouds are decoupled into continuous frame point clouds of foreground targets. And background target continuous frame point cloud ,in, For the moment Point cloud frame Middle Point cloud of foreground objects, For the moment Point cloud frame Middle The semantic annotation of the point cloud defines what kind of target each point of the point cloud belongs to, and the 3D bounding box defines the spatial range of points belonging to the same semantic classification. During semantic annotation, a 3D box is drawn to encircle points in a certain range, and then all points in this range belong to the same semantic category. This 3D box is called a 3D bounding box. According to the semantics of the point cloud and the 3D bounding box annotation, the driving scene targets contained in the point cloud are divided into foreground targets and background targets. The foreground targets are movable targets in the driving scene, including motor vehicles, pedestrians and non-motor vehicles, and the background targets are immovable targets in the driving scene, including the ground, buildings, and green plants.

[0069] S300, in the autonomous driving scenario, the foreground target is more concerned than the background target. Take a set number of foreground target continuous frame point clouds and perform point cloud superposition to obtain the foreground target superimposed point cloud In addition to the ground, the constraints of the background targets on autonomous driving are limited and do not require high-precision representation. In the time period corresponding to the set number of foreground target continuous frame point clouds, the key frame point clouds are extracted from the background target continuous frame point clouds according to a longer step length to perform point cloud superposition to obtain the background target superimposed point cloud. .

[0070] In one example, foreground target point clouds of 10 consecutive frames are taken. During the same time period, starting from the first background target point cloud, one key frame is taken every 3 frames, and a total of 3 key frame point clouds of the background target are selected for point cloud superposition. The 3 selected key frame point clouds of the background target are the first frame, the fifth frame, and the ninth frame of the consecutive frames of the background target respectively.

[0071] S400. In an autonomous driving scenario, the ground provides a lot of useful information for autonomous driving. Therefore, it is necessary to focus on the information of the ground part. As Figure 3 shown, the present application determines the ground for a part of the background and performs dense processing on the determined ground paving points to obtain a dense ground, including:

[0072] S410, divide the background target superposition point cloud into concentric ring regions around the detection vehicle:

[0073] , ;

[0074] wherein, for the th concentric ring region , is the inner diameter of the th concentric ring region, is the outer diameter of the th concentric ring region, is the radial distance of point from the center of the circle O.

[0075] In one example, takes a value of 4, , , , , , wherein, , and is in the form of , and Z is a positive integer.

[0076] S420, further divide each concentric ring region radially and circumferentially to obtain a number of sub-regions:

[0077] ,

[0078] wherein, is the mth in the radial direction and the nth in the circumferential direction in the concentric ring region . The concentric ring region is divided into M regions in the radial direction and N regions in the circumferential direction, , is a point the radial distance from the concentric region to the inner diameter, is a point azimuth angle.

[0079] By means of the above splitting method, the size of the sub-region is adaptively configured to adapt to the situation that the point cloud becomes sparser as the distance increases, avoiding the problem that the ground normal vector cannot be estimated for the sub-region due to sparse point cloud in the distance, and solving the problem of incorrect estimation of the ground normal vector caused by too small sub-region in the case of short distance.

[0080] S430. Most of the points with low reflection intensity at the bottom layer of the point cloud are noises formed by multiple reflections. According to the height threshold dynamically adjusted according to the average value and standard deviation of the ground elevation, the bottom points with a height lower than the height threshold and a reflection intensity lower than the set reflection intensity threshold are removed. The sum of the average value of all determined ground in the concentric annular region and the gain standard deviation is used as the new height threshold.

[0081] S440. Perform point cloud vertical plane fitting on all sub-regions to remove the vertical points on the vertical plane. As Figure 4 shown, the process of point cloud vertical plane fitting includes:

[0082] S441. Starting from any lowest point in any sub-region, select a set number of points as seed points, and calculate the position mean value and unit normal vector of the seed points;

[0083] S442. For any candidate point not determined as a vertical point, when the product of the difference between its position and the position mean value of the seed segment and the unit normal vector of the seed point is less than the set distance tolerance for estimating the vertical plane, the candidate point is selected as a potential vertical point;

[0084] S441 and S442 are executed multiple times to select the potential vertical points in the sub-region;

[0085] S443. If the reciprocal of the cosine of the product between the unit normal vector of the potential vertical point and the unit normal vector of the z-axis is greater than the difference from the unit normal vector tolerance of the estimated plane, the potential vertical point is selected as a vertical point.

[0086] S450. Perform point cloud ground plane fitting on all sub-regions to obtain all the ground planes in the sub-regions. As Figure 5 shown, the process includes:

[0087] S451. Select a set number of lowest points from any sub-region as seed points, calculate the average height of the seed points, and use the points with a height lower than the sum of the average height and the height threshold as the initial estimated ground plane point set;

[0088] S452. Calculate the normal vector of the ground plane point set obtained each time, calculate the average position of the ground plane point set obtained each time, and calculate the plane coefficient of the ground plane point set by using the normal vector and the average position of the ground plane point set.

[0089] S453. When iteratively selecting subsequent ground plane point sets, evaluate according to the plane coefficient formula using the coordinates of the candidate points and the normal vector of the previous ground plane point set. When the difference between the obtained value and the plane coefficient of the previous ground plane point set is less than the set plane distance threshold, take the candidate point as an element of the next ground plane point set.

[0090] S460. Use the normal vector uprightness, elevation, and flatness of the ground plane point set, and the normal vector uprightness, elevation, and flatness of all the ground in the concentric annular region where the ground plane point set is located for ground matching to exclude the ground plane point sets that are not the ground in each ground plane point set and determine the ground.

[0091] Among them, based on the normal vector of the ground plane point set and the angle threshold the uprightness screening function is:

[0092] ;

[0093] Through the normal vector uprightness, the ground plane point sets that are generally horizontal are initially screened out. However, only through the normal vector uprightness, the ground plane point sets formed by the car hood or roof, etc., and the ground plane point sets generated by occlusion cannot be filtered out. When a large object approaches the radar, occlusion will occur, and some of the measurement point clouds above the occluded space will be predicted as the ground plane point set. To solve this problem, an elevation screening function is proposed. Based on the average height of the points in the ground plane point set

[0094] ;

[0095] Among them, is the elevation threshold that changes exponentially with the distance from the point to the center. is the elevation recognizable range threshold; when the average height of the points in the ground plane point set is relatively high, it can be seen from the elevation screening function that in the ground plane point set close to the radar, the value of the elevation screening function decays rapidly with the increase of the difference between the average height of the points in the ground plane point set and the elevation threshold thus screening out the false positives in all the ground plane point sets. In this process, some true positives will be lost.

[0096] To restore the lost true positives, a flatness function based on the minimum eigenvalue of the ground plane point set is proposed as:

[0097] ;

[0098] wherein, is the gain amplitude, is the minimum eigenvalue of the ground plane point set, is the flatness threshold. For the part where the elevation screening function is less than 0.5, i.e., the average height of the points in the ground plane point set and the elevation threshold

[0099] When the difference is greater than zero, if the ground plane point set is flat, a relatively large value will be assigned through the flatness function, so as to recover the lost true positive examples. The mean value of the elevation threshold and the flatness threshold of all the determined ground in the concentric annular region and the sum of the gain standard deviation are used as the new threshold. This process will apply the elevation and flatness of the ground plane point set that is more easily determined as the ground to the threshold update process; once accumulated for a long time, it will lead to a relatively small threshold, and a lot of rough ground will be additionally filtered. This application uses the mean value of the elevation threshold

[0100] S470. After determining the ground, construct a uniform set of points to cover the ground to obtain a dense ground.

[0101] S500. After rasterizing the foreground object, non-ground background object, and ground according to the set fineness. To save the calculation time of rasterization, the foreground object and the background object are rasterized with different precisions. Since the foreground object is more important and has a relatively small overall size, a fine grid is used for rasterization, and the background object is rasterized with a rough grid. The influence degrees of the foreground and background objects on the constraints of autonomous driving are different. For the foreground object, a relatively fine grid is required to construct the geometric space contour, while the grid size of the background object can be relatively large. In this method, the grid sizes of the foreground and background objects are defined as 0.2m * 0.2m * 0.2m and 0.4m * 0.4m * 0.4m respectively, and these two numbers can be modified according to the accuracy requirements of the actual occupancy truth value.

[0102] S600. After completing the rasterization, determine the degree of geometric appearance loss of the foreground object, and complement the geometric appearance of the foreground object according to the degree of geometric appearance loss, as Figure 6 shown, including:

[0103] S610. First, determine the degree of geometric appearance deficiency based on the size of the three-dimensional ground truth box of the foreground target and the actual occupied space, and perform different processing according to the degree of geometric appearance deficiency.

[0104] In the specific implementation process, the degree of geometric appearance deficiency is the three-dimensional intersection over union (IoU) of the actual three-dimensional bounding box of the point cloud and the three-dimensional ground truth box. The actual three-dimensional bounding box is the minimum bounding cube directly calculated by calling the library function for a point cloud. The three-dimensional IoU is an important indicator to measure the overlap degree between two three-dimensional bounding boxes. The degree of geometric appearance deficiency is defined by calculating the volume of their intersection divided by the volume of their union.

[0105] S620. Compare the three-dimensional IoU with the set three-dimensional IoU threshold. When the three-dimensional IoU is lower than the set three-dimensional IoU threshold, execute S630. When the three-dimensional IoU is not lower than the set three-dimensional IoU threshold, execute S640.

[0106] S630. Replace it with a pre-constructed dense occupancy grid model according to the semantic category.

[0107] S640. Use Poisson reconstruction to complete the geometric appearance. The detailed steps of the completion process include:

[0108] S641. Take one of the two symmetric planes perpendicular to the horizontal plane of the three-dimensional ground truth box as the mirror plane, and perform one mirroring based on maximizing the new occupancy grid result generated after mirroring.

[0109] S642. Consider the point set of the center points of the foreground target after mirroring as a point cloud for Poisson reconstruction to generate a continuous three-dimensional network.

[0110] S643. Take the center points of the grids in the space enclosed by the three-dimensional network of Poisson reconstruction as the grid center points generated after the geometric appearance is completed, and the category is unified as the foreground target category. For example: In the point cloud space of the foreground target, generate continuous grids in the x, y, and z directions with a side length of 0.2m. If the center point of the grid is in the three-dimensional network, keep this grid, otherwise delete it. The remaining grids are the occupancy ground truth grids after the geometric appearance is completed.

[0111] For targets with a small geometric appearance deficiency, form a basic point cloud from the center points of all grids, mirror the basic point cloud along the symmetric plane of the three-dimensional ground truth box, then perform Poisson reconstruction, and then regard the reconstructed dense point cloud as the grid center points to generate a dense grid. For targets with a large geometric appearance deficiency, such as the vehicle in front that the ego vehicle continuously follows, usually only the tail sheet-like point cloud generates a grid. According to the semantic category, import the general grid model of this category to replace it.

[0112] The S700 merges the foreground target grid that completes the geometric appearance with the background grid to generate the complete ground truth of the occupancy of the autonomous driving scene. It should be noted that to ensure the uniformity of the grid size in the generated ground truth of occupancy and facilitate subsequent model training and testing, when loading the background grid, eight small grids with a side length of 0.2 m are used to replace the large grid with a side length of 0.4 m.

[0113] The fast and dense ground truth generation method proposed in this patent avoids the reconstruction process of a large range of background, and the object of Poisson reconstruction is only the foreground target point cloud downsampled by rasterization, greatly reducing the time cost of the calculation process. In addition, using the general grid model to replace the foreground target with severely missing geometric appearance greatly improves the quantity of dirty data (such as the flake-shaped grid at the vehicle tail) in the ground truth of occupancy, making the ground truth have a better effect on the training of the subsequent supervised occupancy prediction model.

[0114] Embodiment 2

[0115] Refer to Figure 7 As shown, the embodiment of the present invention provides a fast and dense ground truth generation device, including: at least one processing unit, and the processing unit is connected to a storage unit through a bus unit. The storage unit, as a computer-readable storage medium, can be used to store software programs, computer-executable programs, and modules, such as the software programs, computer-executable programs, and modules corresponding to a fast and dense ground truth generation method in the embodiment of the present invention. The processing unit realizes the above-mentioned fast and dense ground truth generation method by running the software programs, computer-executable programs, and modules stored in the storage unit, including:

[0116] Obtain the point cloud labeled by semantics and three-dimensional bounding boxes of the driving scene, extract the continuous frame point cloud, and convert the continuous frame point cloud into a unified coordinate system;

[0117] Perform foreground and background segmentation on each frame of point cloud according to semantics and three-dimensional bounding boxes to obtain foreground targets and background targets. The continuous frame point cloud is decoupled into the foreground target continuous frame point cloud and the background target continuous frame point cloud; among them, the foreground target is the movable target in the driving scene, and the background target is the immovable target in the driving scene;

[0118] Take a set number of foreground target continuous frame point clouds for point cloud superposition to obtain the foreground target superposed point cloud. In the time period corresponding to the set number of foreground target continuous frame point clouds, extract key frame point clouds from the background target continuous frame point cloud at a longer step length for point cloud superposition to obtain the background target superposed point cloud;

[0119] Determine the ground for a part of the background, and perform dense processing on the determined ground by laying points to obtain a dense ground;

[0120] The foreground targets are rasterized using a fine grid, and the background targets are rasterized using a coarse grid;

[0121] After rasterization, determine the degree of geometric appearance loss of the foreground targets, and perform geometric appearance completion on the foreground targets according to the degree of geometric appearance loss;

[0122] Merge the raster of the foreground targets with geometric appearance completion and the background raster to generate a complete ground truth occupancy for the autonomous driving scene.

[0123] Of course, the storage unit in a fast dense ground truth generation device provided by an embodiment of the present invention stores a computer program that is not limited to the method operations described above, and can also execute relevant operations in a fast dense ground truth generation method provided by any embodiment of the present invention.

[0124] Embodiment 3

[0125] An embodiment of the present invention provides a computer-readable storage medium that stores a computer program, and when the computer program is executed, it implements the fast dense ground truth generation method, including:

[0126] Obtain the point cloud labeled by semantics and three-dimensional bounding boxes of the driving scene, extract the continuous frame point cloud, and convert the continuous frame point cloud into a unified coordinate system;

[0127] Perform foreground and background segmentation on each frame of the point cloud according to semantics and three-dimensional bounding boxes to obtain foreground targets and background targets. The continuous frame point cloud is decoupled into a continuous frame point cloud of foreground targets and a continuous frame point cloud of background targets; among them, the foreground targets are movable targets in the driving scene, and the background targets are immovable targets in the driving scene;

[0128] Take a set number of continuous frame point clouds of foreground targets for point cloud superposition to obtain a superposed point cloud of foreground targets. In the time period corresponding to the set number of continuous frame point clouds of foreground targets, extract key frame point clouds from the continuous frame point cloud of background targets at a longer step size for point cloud superposition to obtain a superposed point cloud of background targets;

[0129] Determine the ground for the background part, and perform dense processing on the determined ground by laying points to obtain a dense ground;

[0130] The foreground targets are rasterized using a fine grid, and the background targets are rasterized using a coarse grid;

[0131] After rasterization, determine the degree of geometric appearance loss of the foreground targets, and perform geometric appearance completion on the foreground targets according to the degree of geometric appearance loss;

[0132] Merge the foreground target grid with the geometric appearance completed and the background grid to generate a complete ground truth for the occupancy of the autonomous driving scene.

[0133] A computer-readable storage medium provided by an embodiment of the present invention, the computer program stored therein is not limited to the method operations as described above, and can also execute the related operations in a fast dense ground truth generation method provided by any embodiment of the present invention.

[0134] In the embodiments provided by the present invention, it should be understood that the disclosed structure and method can be implemented in other ways. For example, the structural embodiments described above are merely illustrative. For example, the division of the units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point, the displayed or discussed coupling or direct coupling or communication connection between each other can be through some interfaces, and the indirect coupling or communication connection of the structure or unit can be in an electrical, mechanical or other form.

[0135] The units described as separate components may or may not be physically separated. The components displayed as units may or may not be physical units, that is, they can be located in one place, or they can be distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0136] In addition, each functional unit in various embodiments of the present invention can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The above integrated unit can be implemented in the form of hardware or in the form of a software functional unit.

[0137] The above are only specific embodiments of the present invention, which enable those skilled in the art to understand or implement the present invention. Various modifications to these embodiments will be obvious to those skilled in the art. The general principles defined herein can be implemented in other embodiments without departing from the spirit or scope of the present invention. Therefore, the present invention will not be limited to the embodiments shown herein, but will conform to the widest scope consistent with the principles and novel features claimed herein.

Claims

1. A fast dense occupancy ground truth generation method, characterized in that, Including: Obtain the point cloud annotated with semantics and 3D bounding boxes of the driving scenario, extract the continuous frame point clouds, and transform the continuous frame point clouds into a unified coordinate system; Perform foreground and background segmentation on each frame of the point cloud according to the semantics and 3D bounding boxes to obtain foreground targets and background targets. The continuous frame point clouds are decoupled into foreground target continuous frame point clouds and background target continuous frame point clouds; among them, the foreground targets are movable targets in the driving scenario, and the background targets are immovable targets in the driving scenario; Take a set number of foreground target continuous frame point clouds for point cloud superposition to obtain foreground target superimposed point clouds. In the time period corresponding to the set number of foreground target continuous frame point clouds, extract key frame point clouds from the background target continuous frame point clouds at a longer step size for point cloud superposition to obtain background target superimposed point clouds; Determine the ground for part of the background, and perform dense processing on the determined ground by paving points to obtain a dense ground; Use a fine grid to rasterize the foreground targets and a rough grid to rasterize the background targets; After rasterization is completed, determine the degree of geometric appearance loss of the foreground targets, and complement the geometric appearance of the foreground targets according to the degree of geometric appearance loss; Merge the rasterized foreground targets with geometric appearance complementation and the background raster to generate a complete occupancy ground truth for the autonomous driving scenario.

2. The rapid dense occupancy truth value generation method according to claim 1, characterized in that Determine the ground for part of the background, and perform dense processing on the determined ground by paving points to obtain a dense ground, including: S410, superimpose the background target point cloud and divide it into concentric circular regions around the detection vehicle: , ; Among them, the background target superposed point cloud in the th concentric ring region , is the inner diameter of the th concentric ring region, is the outer diameter of the th concentric ring region, is the radial distance of point from the center of the circle O; S420, further divide each concentric circular ring region radially and circumferentially to obtain a number of sub-regions: , Among them, is the m-th in the radial direction and the n-th sub-region in the circumferential direction belonging to the concentric circular ring region in the concentric circular ring region, which is radially divided into M regions and circumferentially divided into N regions , is the point the radial distance from the concentric region to the inner diameter is the azimuth angle of the point ;​ S430, Remove the bottom points whose height is lower than the height threshold dynamically adjusted according to the average value and standard deviation of the ground elevation and whose reflection intensity is lower than the set reflection intensity threshold; S440, Perform vertical plane fitting on all sub-regions to remove the vertical points on the vertical planes; S450, Perform ground plane fitting on all sub-regions to obtain all the ground planes in the sub-regions; S460, Use the normal uprightness, elevation, and flatness of the ground plane point set, and the normal uprightness, elevation, and flatness of all the ground in the concentric annular region where the ground flat point set is located to perform ground matching to exclude the non-ground ground plane point sets in each ground plane point set and determine the ground; S470, After determining the ground, construct a uniform point to cover the ground.

3. The rapid dense occupancy truth value generation method according to claim 2, wherein The sum of the average value and the gain standard deviation of all the determined ground in the concentric annular region is used as the new height threshold.

4. The rapid dense occupancy truth value generation method according to claim 2, characterized in that The process of vertical plane fitting includes: S441, Starting from any lowest point in any sub-region, select a set number of points as seed points, and calculate the position mean and unit normal vector of the seed points; S442, For any candidate point that has not been determined as a vertical point, when the product of the difference between its position and the position mean of the seed segment and the unit normal vector of the seed points is less than the set distance tolerance for estimating the vertical plane, select this candidate point as a potential vertical point. Perform this multiple times to select the potential vertical points in the sub-region; S443, if the reciprocal of the cosine of the product between the unit normal vector of the potential vertical point and the unit normal vector of the z-axis is greater than the difference from the unit normal vector tolerance of the estimated plane, the potential vertical point is selected as the vertical point.

5. The rapid dense occupancy truth value generation method according to claim 2, characterized in that Performing ground plane fitting on all sub-regions to obtain all the ground planes in the sub-regions includes: S451, Select a set number of lowest points from any sub-region as seed points, calculate the height mean of the seed points, and use the points whose height is lower than the sum of the height mean and the height threshold as the initially estimated ground plane point set; S452. Calculate the normal vector of the ground plane point set obtained each time, calculate the average position of the ground plane point set obtained each time, and calculate the plane coefficient of the ground plane point set by using the normal vector of the ground plane point set and the average position of the ground plane point set. S453. When iteratively selecting subsequent ground plane point sets, when evaluating according to the coordinates of the candidate points and the normal vector of the previous ground plane point set according to the plane coefficient formula, if the difference between the obtained value and the plane coefficient of the previous ground plane point set is less than the set plane distance threshold, use this candidate point as an element of the next ground plane point set.

6. The rapid dense occupancy true value generation method according to claim 2, characterized in that Based on the normal vector of the ground plane point set and the angle threshold the uprightness screening function is as follows: ; Based on the average height of points in the ground plane point set The elevation screening function based on the elevation threshold is as follows: ; Among them, is the elevation threshold that varies exponentially with the distance from the point to the center, and is the threshold for the elevation recognizable range. The flatness function based on the minimum eigenvalue of the ground plane point set is: ; Among them, is the gain amplitude, is the minimum eigenvalue of the ground plane point set, is the flatness threshold. For the part where the elevation screening function is less than 0.5, that is, the average height of the points in the ground plane point set minus the elevation threshold, when the difference is greater than zero, if the ground plane point set is flatter, a larger value is assigned through the flatness function, so as to recover the lost true positive examples.

7. The rapid dense occupancy true value generation method according to claim 6, wherein Elevation thresholds of all determined ground in the concentric annular region and flatness thresholds The sum of the mean value and the gain standard deviation is used as the new threshold, where the gain standard deviation is the product of the gain coefficient and the standard deviation; using the elevation thresholds of all determined ground in the concentric annular region and flatness thresholds The sum of the mean value and the gain standard deviation is used as the recovery threshold. When the threshold is less than the recovery threshold, the ground plane point set between the threshold and the recovery threshold is restored.

8. The rapid intensive occupancy true value generation method according to claim 1, characterized in that Determine the degree of geometric appearance loss of the foreground target, and geometric appearance completion of the foreground target according to the degree of geometric appearance loss includes: S610. First, based on the three-dimensional intersection over union (IoU) of the actual three-dimensional bounding box and the three-dimensional ground truth box of the foreground target, use the three-dimensional IoU as the degree of geometric appearance loss. Compare the three-dimensional IoU with the set three-dimensional IoU threshold. When the three-dimensional IoU is lower than the set three-dimensional IoU threshold, execute S630; when the three-dimensional IoU is not lower than the set three-dimensional IoU threshold, execute S640. S630. Replace it with a pre-constructed dense occupancy grid model according to the semantic category. S640. Use Poisson reconstruction for geometric appearance completion.

9. The rapid dense occupancy truth value generation method according to claim 8, wherein Using Poisson reconstruction for geometric appearance completion, the detailed steps of the completion process include: S641. Use one of the two symmetric planes perpendicular to the horizontal plane of the three-dimensional ground truth box as the mirror plane, and perform one mirroring based on maximizing the new occupancy grid result generated after mirroring. S642. Consider the center point set of the foreground target after mirroring as a point cloud for Poisson reconstruction to generate a continuous three-dimensional network. S643. Use the center point of the grid in the space enclosed by the three-dimensional network of Poisson reconstruction as the grid center point generated by the geometric appearance completion, and the category is unified as the foreground target category.

10. A fast dense occupancy ground truth generation device, characterized in that, Including: At least one processing unit, the processing unit is connected to the storage unit through the bus unit, the storage unit stores a computer program, and when the computer program is executed by the processing unit, it implements the fast dense occupancy ground truth generation method as described in any one of claims 1-9.

Citation Information

Patent Citations

  • Dynamic and static truth value dense generation method for three-dimensional occupancy task

    CN118196762A

  • Occupied grid-based mining area data set construction method, device and system

    CN118968451A