Rapid dense occupancy true value generation method and device
By performing front and back scene segmentation, superposition, ground determination and rasterization of point clouds, combined with Poisson reconstruction and raster model, the problems of sparseness and high computational cost in the existing technology are solved, and a method of quickly generating intensive occupancy truth values is realized.
Patent Information
- Application Number
- CN202510445293.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-10
- Publication Date
- 2025-05-06
- Estimated Expiration
- 2045-04-10
AI Technical Summary
In the prior art, the generated three-dimensional occupancy is sparse, resulting in poor effect of occupying labels, and the Poisson reconstruction process for large-scale scenarios consumes a lot of computing power, which is huge in time.
By obtaining point clouds marked by semantic and three-dimensional bounding boxes, extracting continuous frame point clouds and converting them to a unified coordinate system, performing front and back scene segmentation, point cloud overlay, ground determination and rasterization, combined with Poisson reconstruction and the use of raster models, intensive occupation truth values are generated.
It realizes the rapid generation of intensive occupancy truth values, avoids the reconstruction of large-scale future scenarios, reduces the time cost of the calculation process, improves the accuracy of occupancy truth values and the effect of subsequent supervision model training.
Smart Images

Figure CN119941589A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of generating dense occupancy truth values based on point clouds, and in particular to a fast dense occupancy truth value generation method and device. Background Art
[0002] Understanding the 3D geometry of the surrounding scene is a fundamental step in autonomous driving systems. Although LiDAR is a direct solution to capture geometric information, its high-cost sensor and sparse scanning points limit its further application. In recent years, vision-centric autonomous driving has attracted widespread attention as a promising direction. By taking multi-camera images as input, it has demonstrated competitiveness in various 3D perception tasks, including depth estimation, 3D object detection, and semantic map construction. Although multi-camera 3D object detection has achieved certain results in vision-based 3D perception, it is susceptible to the long-tail problem and has difficulty in identifying all categories of objects in the real world. One way to improve it is to supervise multi-camera 3D object detection by the 3D occupancy of the objects in the scene. 3D occupancy describes the scene by assigning occupancy probabilities to each voxel in 3D space. 3D occupancy is a good representation for multi-camera scene reconstruction because it naturally guarantees multi-camera geometric consistency and can recover occluded parts. In the prior art, many network models use sparse radar point clouds for supervision to predict 3D occupancy, but the 3D occupancy generated in this way is sparse and has poor performance as occupancy labels. The existing occupancy truth generation algorithm uses continuous frame semantic point cloud as input. First, the dynamic targets in the semantic point cloud are filtered and the point cloud is divided into static point cloud and multiple dynamic point clouds. Then, the dynamic and static point clouds are overlapped according to the world coordinates and the three-dimensional truth frame information to achieve densification. Then, the complete scene obtained by merging the dense dynamic and static point clouds is Poisson reconstructed to achieve the loss of holes on the surface of a single point cloud object and the densification of the interior. Finally, the scene after Poisson reconstruction is assigned semantic truth by the nearest neighbor algorithm to obtain the occupancy truth result. The advantage of this method is that the generated occupancy truth value is highly granular and the scene reconstruction is relatively accurate. The disadvantage is that the process of Poisson reconstruction of large-scale scenes consumes a lot of computing power and has a huge time cost. Summary of the invention
[0003] In order to solve the above technical problem or at least partially solve the above technical problem, the present invention provides a fast dense occupancy true value generation method and device.
[0004] In a first aspect, the present invention provides a fast method for generating dense occupancy truth values, comprising: Obtain the point cloud of the driving scene annotated with semantics and 3D bounding boxes, extract the continuous frame point cloud, and transform the continuous frame point cloud into a unified coordinate system; The foreground and background of each frame point cloud are segmented according to semantics and 3D bounding boxes to obtain foreground targets and background targets. The continuous frame point clouds are decoupled into continuous frame point clouds of foreground targets and continuous frame point clouds of background targets. Among them, the foreground targets are movable targets in the driving scene, and the background targets are immovable targets in the driving scene. A set number of continuous frame point clouds of foreground targets are taken for point cloud superposition to obtain a foreground target superposition point cloud. In a time period corresponding to the set number of continuous frame point clouds of foreground targets, key frame point clouds are extracted from the continuous frame point clouds of background targets according to a longer step length to perform point cloud superposition to obtain a background target superposition point cloud; the background part is ground determined, and the determined ground points are densely processed to obtain a dense ground. Use fine grid to rasterize foreground objects, and use coarse grid to rasterize background objects; After rasterization is completed, the degree of geometric appearance loss of the foreground object is determined, and the geometric appearance of the foreground object is completed according to the degree of geometric appearance loss; The geometrically completed foreground object grid is merged with the background grid to generate the complete ground truth of the autonomous driving scene occupancy.
[0005] Furthermore, the ground of the background is determined, and the determined ground points are densely processed to obtain a dense ground, including: S410, superimpose point cloud on background target Divided into the area surrounding the detection vehicle Concentric ring areas: , ; Among them, the background target superimposed point cloud Middle Concentric ring areas , For the The inner diameter of the concentric ring area is For the The outer diameter of the concentric ring area, For point The radial distance from the center O; S420, each concentric ring area Further division along the radial and circumferential directions yields several sub-areas: , in, The concentric ring area The mth radial sub-area, the nth circumferential sub-area, the concentric ring area The radial direction is divided into M areas, and the circumferential direction is divided into N areas. , For point Distance concentric area Radial distance of the inner diameter, For point azimuth; S430, removing bottom points whose height is lower than the height threshold and whose reflection intensity is lower than the set reflection intensity threshold according to the height threshold dynamically adjusted by the average value and standard deviation of the ground elevation; S440, performing vertical plane fitting on all sub-regions to remove vertical points on the vertical plane; S450, performing ground plane fitting on all sub-regions to obtain all ground planes in the sub-regions; S460, using the normal verticality, elevation and flatness of the ground plane point set, and the normal verticality, elevation and flatness of all the ground in the concentric annular area where the ground plane point set is located, ground matching is performed to exclude the non-ground ground plane point sets in each ground plane point set, and determine the ground plane point set. S470, after the ground is determined, uniform points are constructed to cover the ground.
[0006] Furthermore, the sum of the average value of all determined ground values in the concentric annular areas and the gain standard deviation is taken as a new height threshold.
[0007] Furthermore, the process of vertical plane fitting includes: S441, starting from any lowest point in any sub-region, selecting a set number of points as seed points, and calculating the position mean and unit normal vector of the seed points; S442, for any candidate point that is not determined to be a vertical point, when the product of the difference between its position and the position mean of the seed segment and the unit normal vector of the seed point is less than the set distance tolerance of the estimated vertical plane, the candidate point is selected as a potential vertical point, and the process is repeated multiple times to select potential vertical points in the sub-region; S443, if the inverse of the cosine of the product of the unit normal vector of the potential vertical point and the unit normal vector of the z-axis is greater than Potential perpendicular points are chosen as perpendicular points when the difference is less than the tolerance of the unit normal vector of the estimated plane.
[0008] Furthermore, ground plane fitting is performed on all sub-areas to obtain all ground planes in the sub-areas including: S451, selecting a set number of lowest points from any sub-region as seed points, calculating the average height of the seed points, and taking points whose heights are lower than the sum of the average height and the height threshold as an initial estimated ground plane point set; S452, calculating the normal vector of the ground plane point set obtained each time, calculating the average position of the ground plane point set obtained each time, and calculating the plane coefficient of the ground plane point set using the normal vector of the ground plane point set and the average position of the ground plane point set; S453, when iteratively selecting the subsequent ground plane point set, the plane coefficient formula is used to evaluate the coordinates of the candidate point and the normal vector of the previous ground plane point set. When the difference between the evaluated value and the plane coefficient of the previous ground plane point set is less than the set plane distance threshold, the candidate point is used as an element of the next ground plane point set.
[0009] Furthermore, based on the normal vector of the ground plane point set and angle threshold The upright screening function is: ; The average height of the points in the ground plane. The elevation screening function with elevation threshold is: ; in, As the distance from the point to the center An exponentially changing elevation threshold, is the elevation recognizable range threshold; The flatness function based on the minimum eigenvalue of the ground plane point set is: ; in, is the gain amplitude, is the minimum eigenvalue of the ground plane point set, is the flatness threshold, for the part where the elevation screening function is less than 0.5, that is, the average height of the points in the ground plane point concentration With elevation threshold When the difference is greater than zero, if the ground plane point set is flat, a larger value will be assigned by the flatness function, thereby recovering the lost true positive examples.
[0010] Furthermore, the elevation threshold of the concentric ring areas is determined by and flatness threshold The sum of the mean and the gain standard deviation is taken as the new threshold, and the gain standard deviation is the product of the gain coefficient and the standard deviation; the elevation threshold of the ground is determined by using all the concentric annular areas in the recent period of time and flatness threshold The sum of the mean and the gain standard deviation is taken as the restoration threshold. When the threshold is less than the restoration threshold, the ground plane point set between the threshold and the restoration threshold is restored.
[0011] Furthermore, the degree of geometric appearance loss of the foreground object is determined, and the geometric appearance of the foreground object is completed according to the degree of geometric appearance loss, including: S610, first, according to the 3D intersection-and-union ratio of the actual 3D enclosing frame of the point cloud of the foreground object and the 3D truth frame, the 3D intersection-and-union ratio is used as the degree of geometric appearance loss; The 3D intersection-and-union ratio is compared with a set 3D intersection-and-union ratio threshold. If the 3D intersection-and-union ratio is lower than the set 3D intersection-and-union ratio threshold, S630 is executed. If the 3D intersection-and-union ratio is not lower than the set 3D intersection-and-union ratio threshold, S640 is executed. S630, selecting a pre-built dense occupancy grid model as a substitute according to the semantic category; S640, geometric appearance completion using Poisson reconstruction.
[0012] Furthermore, Poisson reconstruction is used to complete the geometric appearance. The detailed steps of the completion process include: S641, taking one of the two symmetric planes of the three-dimensional truth box that are perpendicular to the horizontal plane as a mirror plane, and performing a mirroring operation based on maximizing a new occupancy grid result generated after the mirroring; S642, treating the central point set of the mirrored foreground object as a point cloud and performing Poisson reconstruction to generate a continuous three-dimensional network; S643, the center point of the grid in the space covered by the Poisson reconstructed three-dimensional network is used as the center point of the grid generated after the geometric appearance is completed, and the category is unified as the foreground target category.
[0013] In a second aspect, the present invention provides a fast densely occupied true value generating device, comprising: at least one processing unit, the processing unit is connected to a storage unit via a bus unit, the storage unit stores a computer program, and when the computer program is executed by the processing unit, the fast densely occupied true value generating method is implemented.
[0014] The above technical solution provided by the embodiment of the present invention has the following advantages compared with the prior art: The point cloud of the driving scene annotated with semantics and three-dimensional bounding boxes is obtained, and the continuous frame point cloud is extracted, and the continuous frame point cloud is converted into a unified coordinate system; the foreground and background of each frame point cloud are segmented according to the semantics and three-dimensional bounding boxes to obtain the foreground target and the background target, and the continuous frame point cloud is decoupled into the foreground target continuous frame point cloud and the background target continuous frame point cloud; among which, the foreground target is the movable target in the driving scene, and the background target is the immovable target in the driving scene; take a set number of foreground target continuous frame point clouds to superimpose the point clouds to obtain the foreground target superimposed point cloud, and at the time corresponding to the set number of foreground target continuous frame point clouds In the segment, the key frame point cloud is extracted from the background target continuous frame point cloud according to a longer step size to perform point cloud superposition to obtain the background target superposition point cloud; the ground is determined for the background part, and the determined ground points are densely processed to obtain a dense ground; the foreground target is rasterized using a fine grid, and the background target is rasterized using a coarse grid; after rasterization, the degree of geometric appearance loss of the foreground target is determined, and the geometric appearance of the foreground target is completed according to the degree of geometric appearance loss; the foreground target grid with geometric appearance completion is merged with the background grid to generate a complete autonomous driving scene occupancy truth value. The fast dense occupancy truth value generation method proposed in this patent avoids the reconstruction of a large range of backgrounds, and the object of Poisson reconstruction is only the foreground target point cloud downsampled by rasterization, which greatly reduces the time cost of the calculation process. For the ground part, dense processing is performed by ground matching and adding points to replace Poisson reconstruction, which greatly reduces the amount of calculation. In addition, using a general grid model to replace foreground objects with severe geometric appearance loss can greatly improve the amount of dirty data in the true occupancy value, making the true value more effective for subsequent supervised occupancy prediction model training. BRIEF DESCRIPTION OF THE DRAWINGS
[0015] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with the invention and, together with the description, serve to explain the principles of the invention.
[0016] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, for ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative labor.
[0017] Figure 1 A flowchart of a method for quickly generating a dense occupancy true value provided by an embodiment of the present invention; Figure 2 A schematic diagram of the evolution from continuous frame point clouds to final dense occupancy true values during the execution of a fast dense occupancy true value generation method provided by an embodiment of the present invention; Figure 3A flowchart of determining the ground surface of a background portion provided by an embodiment of the present invention, and densely processing the determined ground points to obtain a dense ground surface; Figure 4 A flowchart of point cloud vertical plane fitting provided by an embodiment of the present invention; Figure 5 A flowchart of performing point cloud ground plane fitting on all sub-areas to obtain all ground planes in the sub-areas provided in an embodiment of the present invention; Figure 6 A flowchart of performing point cloud ground plane fitting on all sub-areas to obtain all ground planes in the sub-areas provided in an embodiment of the present invention; Figure 7 A schematic diagram of a fast dense occupancy true value generating device provided by an embodiment of the present invention. DETAILED DESCRIPTION
[0018] In order to make the purpose, technical solution and advantages of the embodiments of the present invention clearer, the technical solution in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0019] It should be noted that, in this article, the terms "include", "comprises" or any other variations thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. In the absence of further restrictions, an element defined by the sentence "comprises a ..." does not exclude the existence of other identical elements in the process, method, article or device including the element.
[0020] Example 1 The technology of the present invention realizes a fast dense occupancy truth value generation method, the purpose of which is to use dense occupancy grids to annotate targets in complex autonomous driving scenarios. The dense occupancy grids of targets can be used to train neural network models to achieve target detection or semantic segmentation tasks.
[0021] Combination Figure 1 and Figure 2 As shown, the fast dense occupancy truth value generation method provided by the present application includes: S100, obtaining the point cloud of the driving scene annotated with semantics and 3D bounding boxes, and extracting continuous frame point clouds ,in, Indicates time The point cloud frame at each moment contains several points ,in, For the moment The first point cloud frame at Convert the continuous frame point clouds into a unified coordinate system.
[0022] S200: Segment the foreground and background of each frame point cloud according to semantics and 3D bounding boxes to obtain foreground targets and background targets. The continuous frame point clouds are decoupled into continuous frame point clouds of foreground targets. And background target continuous frame point cloud ,in, For the moment Point cloud frame Middle Point cloud of foreground objects, For the moment Point cloud frame Middle The semantic annotation of the point cloud defines what kind of target each point of the point cloud belongs to, and the 3D bounding box defines the spatial range of points belonging to the same semantic classification. During semantic annotation, a 3D box is drawn to encircle points in a certain range, and then all points in this range belong to the same semantic category. This 3D box is called a 3D bounding box. According to the semantics of the point cloud and the 3D bounding box annotation, the driving scene targets contained in the point cloud are divided into foreground targets and background targets. The foreground targets are movable targets in the driving scene, including motor vehicles, pedestrians and non-motor vehicles, and the background targets are immovable targets in the driving scene, including the ground, buildings, and green plants.
[0023] S300, in the autonomous driving scenario, the foreground target is more concerned than the background target. Take a set number of foreground target continuous frame point clouds and perform point cloud superposition to obtain the foreground target superimposed point cloud In addition to the ground, the constraints of the background targets on autonomous driving are limited and do not require high-precision representation. In the time period corresponding to the set number of foreground target continuous frame point clouds, the key frame point clouds are extracted from the background target continuous frame point clouds according to a longer step length to perform point cloud superposition to obtain the background target superimposed point cloud. .
[0024] In one example, 10 consecutive frames of foreground target point clouds are taken. Within the same time period, starting from the first frame of background target point cloud, a key frame is taken every 3 frames, and a total of 3 frames of background target key frame point clouds are selected for point cloud superposition. The selected 3 frames of background target key frames are the first frame, the fifth frame and the ninth frame of the background target continuous frames respectively.
[0025] S400, in the autonomous driving scenario, the ground will provide a lot of useful information for autonomous driving, so it is necessary to focus on the information on the ground. Figure 3 As shown, the present application determines the ground surface of the background part, and performs dense processing on the determined ground points to obtain a dense ground surface, including: S410, superimpose point cloud on background target Divided into the area surrounding the detection vehicle Concentric ring areas: , ; Among them, the background target superimposed point cloud Middle Concentric ring areas , For the The inner diameter of the concentric ring area is For the The outer diameter of the concentric ring area, For point The radial distance from the center O.
[0026] An example The value is 4. , , , , ,in, ,and The form is , Z is a positive integer.
[0027] S420, each concentric ring area Further division along the radial and circumferential directions yields several sub-areas: , in, The concentric ring area The mth radial sub-area, the nth circumferential sub-area, the concentric ring area The radial direction is divided into M areas, and the circumferential direction is divided into N areas. , For point Distance concentric area Radial distance of the inner diameter, For point azimuth.
[0028] The size of the sub-region is adaptively configured through the above segmentation method to adapt to the situation where the point cloud becomes increasingly sparse with increasing distance, avoid the problem that the ground normal vector cannot be estimated in the sub-region due to the sparse point cloud at a distance, and solve the problem that the ground normal vector is estimated incorrectly due to the small size of the sub-region at a close distance.
[0029] S430: Point clouds with low reflection intensity at the bottom of the point cloud are mostly noise formed by multiple reflections. The height threshold is dynamically adjusted according to the average and standard deviation of the ground elevation, and the bottom points with a height lower than the height threshold and a reflection intensity lower than the set reflection intensity threshold are removed. The sum of the average value of all the ground determined in the concentric annular area and the gain standard deviation is used as the new height threshold.
[0030] S440, performing point cloud vertical plane fitting on all sub-areas to remove vertical points on the vertical plane, such as Figure 4 As shown in the figure, the process of point cloud vertical plane fitting includes: S441, starting from any lowest point in any sub-region, selecting a set number of points as seed points, and calculating the position mean and unit normal vector of the seed points; S442, for any candidate point that is not determined to be a vertical point, when the product of the difference between its position and the position mean of the seed segment and the unit normal vector of the seed point is less than the set distance tolerance of the estimated vertical plane, the candidate point is selected as a potential vertical point; S441 and S442 are executed multiple times to select potential vertical points in the sub-region; S443, if the inverse of the cosine of the product of the unit normal vector of the potential vertical point and the unit normal vector of the z-axis is greater than Potential perpendicular points are chosen as perpendicular points when the difference is less than the tolerance of the unit normal vector of the estimated plane.
[0031] S450, performing point cloud ground plane fitting on all sub-regions to obtain all ground planes in the sub-regions, such as Figure 5 As shown, the process includes: S451, selecting a set number of lowest points from any sub-region as seed points, calculating the average height of the seed points, and taking points whose heights are lower than the sum of the average height and the height threshold as an initial estimated ground plane point set; S452, calculating the normal vector of the ground plane point set obtained each time, calculating the average position of the ground plane point set obtained each time, and calculating the plane coefficient of the ground plane point set using the normal vector of the ground plane point set and the average position of the ground plane point set; S453, when iteratively selecting the subsequent ground plane point set, the plane coefficient formula is used to evaluate the coordinates of the candidate point and the normal vector of the previous ground plane point set. When the difference between the evaluated value and the plane coefficient of the previous ground plane point set is less than the set plane distance threshold, the candidate point is used as an element of the next ground plane point set.
[0032] S460, using the normal verticality, elevation and flatness of the ground plane point set and the normal verticality, elevation and flatness of all the ground in the concentric annular area where the ground plane point set is located, ground matching is performed to exclude the non-ground ground plane point sets in each ground plane point set to determine the ground.
[0033] Among them, based on the normal vector of the ground plane point set and angle threshold The upright screening function is: ; The normal verticality is used to preliminarily filter out the overall horizontal ground plane point set, but the normal verticality alone cannot filter out the ground plane point set formed by the car hood or roof and the ground plane point set caused by occlusion. When a large object approaches the radar, occlusion will occur, and part of the measured point cloud above the occlusion space will be predicted as the ground plane point set. In order to solve this problem, an elevation filtering function is proposed, which is based on the average height of the points in the ground plane point set. The elevation screening function with elevation threshold is: ; in, As the distance from the point to the center An exponentially changing elevation threshold, is the elevation recognizable range threshold; the average height of the local plane point concentration points When the elevation is relatively high, it can be seen from the elevation filtering function that the ground plane points close to the radar are concentrated, and the elevation filtering function value increases with the average height of the ground point concentration points. With elevation threshold The increase of the difference decays rapidly, thereby filtering out the false positives in the entire ground plane point set, and this process will lose some true positives.
[0034] In order to recover the lost true examples, the flatness function based on the minimum eigenvalue of the ground plane point set is proposed as follows: ; in, is the gain amplitude, is the minimum eigenvalue of the ground plane point set, is the flatness threshold, for the part where the elevation screening function is less than 0.5, that is, the average height of the points in the ground plane point concentration With elevation threshold When the difference is greater than zero, if the ground plane point set is flat, a larger value will be assigned by the flatness function, thereby recovering the lost true positive examples.
[0035] The elevation threshold of all concentric ring areas to determine the ground and flatness threshold The sum of the mean and the gain standard deviation is used as the new threshold. This process will apply the elevation and flatness of the ground plane point set that is easier to determine as the ground to the threshold update process; once accumulated for a long time, it will cause the threshold to be too small, and many rough ground will be filtered out additionally. This application uses the elevation threshold of all concentric annular areas in the recent period to determine the ground and flatness threshold The sum of the mean and the gain standard deviation is taken as the restoration threshold. When the threshold is less than the restoration threshold, the ground plane point set between the threshold and the restoration threshold is restored.
[0036] S470, after the ground is determined, uniform points are constructed to cover the ground to obtain a dense ground.
[0037] S500, rasterizes the foreground targets, non-ground background targets and the ground with the set degree of precision. In order to save the calculation time of rasterization, the foreground targets and background targets are rasterized with different precisions respectively. Since the foreground targets are more important and have a smaller overall size, a fine grid is used for rasterization, and a coarse grid is used for rasterization of the background targets. The degree of influence of foreground and background targets on autonomous driving constraints is different. For foreground targets, a relatively fine grid is required to construct the geometric space contour, while the grid size of background targets can be relatively large. In this method, the grid sizes of foreground and background targets are defined as 0.2m*0.2m*0.2m and 0.4m*0.4m*0.4m, respectively, and these two numbers can be modified according to the accuracy of the actual occupancy true value.
[0038] S600, after rasterization is completed, the degree of geometric appearance loss of the foreground object is determined, and the geometric appearance of the foreground object is completed according to the degree of geometric appearance loss, such as Figure 6 As shown, including: S610, first determine the degree of geometric appearance loss according to the three-dimensional true value box size and the actual occupied space size of the foreground object, and perform different processing according to the degree of geometric appearance loss.
[0039] In the specific implementation process, the degree of geometric appearance loss is the 3D intersection-and-union ratio of the actual 3D enclosing box of the point cloud and the 3D true value box. The actual 3D enclosing box is the minimum enclosing cube of a point cloud directly calculated by calling the library function. The 3D intersection-and-union ratio is an important indicator to measure the overlap between two 3D bounding boxes. The degree of geometric appearance loss is defined by calculating the intersection volume of the two divided by their union volume.
[0040] S620, comparing the 3D I / O ratio with a set 3D I / O ratio threshold. If the 3D I / O ratio is lower than the set 3D I / O ratio threshold, S630 is executed. If the 3D I / O ratio is not lower than the set 3D I / O ratio threshold, S640 is executed. S630, selecting a pre-built densely occupied grid model as a substitute according to the semantic category.
[0041] S640 uses Poisson reconstruction to complete the geometric appearance. The detailed steps of the completion process include: S641, taking one of the two symmetric planes of the three-dimensional truth box perpendicular to the horizontal plane as the mirror plane, and performing a mirroring operation based on maximizing the result of a new occupied grid generated after the mirroring.
[0042] S642, treating the central point set of the mirrored foreground object as a point cloud and performing Poisson reconstruction to generate a continuous three-dimensional network; S643, the center point of the grid in the space covered by the Poisson reconstructed three-dimensional network is used as the center point of the grid generated after geometric appearance completion, and the category is unified as the foreground target category; for example: in the point cloud space of the foreground target, a continuous grid in the xyz direction is generated with a side length of 0.2m. If the center point of the grid is in the three-dimensional network, the grid is retained, otherwise it is deleted, and the remaining grid is the true value grid of the occupancy after geometric appearance completion.
[0043] For targets with relatively small geometric appearance missing, the center points of all grids are used to form a basic point cloud, which is mirrored along the symmetry plane of the three-dimensional truth box and then Poisson reconstruction is performed. The reconstructed dense point cloud is then regarded as the grid center points to generate a dense grid. For targets with relatively large geometric appearance missing, such as the vehicle in front that the ego vehicle is continuously following, usually only the flake point cloud at the rear of the vehicle generates a grid. According to the semantic category, the general grid model of the category is imported to replace it.
[0044] S700, merges the foreground target grid with geometric appearance completion with the background grid to generate a complete autonomous driving scene occupancy truth value. It is worth noting that in order to ensure the uniformity of the grid size in the generated occupancy truth value and facilitate subsequent model training and testing, when loading the background grid, 8 small grids with a side length of 0.2m are used instead of the large grid with a side length of 0.4m.
[0045] The fast dense occupancy truth value generation method proposed in this patent avoids the reconstruction of a large range of backgrounds, and the object of Poisson reconstruction is only the foreground target point cloud that is downsampled by rasterization, which greatly reduces the time cost of the calculation process. In addition, the use of a general grid model to replace the foreground target with serious geometric appearance loss greatly improves the amount of dirty data (such as the sheet grid at the rear of the car) in the occupancy truth value, making the truth value more effective for the subsequent supervised occupancy prediction model training.
[0046] Example 2 See also Figure 7As shown, an embodiment of the present invention provides a fast dense occupancy truth value generation device, including: at least one processing unit, the processing unit is connected to a storage unit through a bus unit, the storage unit is a computer-readable storage medium, and can be used to store software programs, computer executable programs and modules, such as the software programs, computer executable programs and modules corresponding to a fast dense occupancy truth value generation method in an embodiment of the present invention. The processing unit implements the above-mentioned fast dense occupancy truth value generation method by running the software programs, computer executable programs and modules stored in the storage unit, including: Obtain the point cloud of the driving scene annotated with semantics and 3D bounding boxes, extract the continuous frame point cloud, and transform the continuous frame point cloud into a unified coordinate system; The foreground and background of each frame point cloud are segmented according to semantics and 3D bounding boxes to obtain foreground targets and background targets. The continuous frame point clouds are decoupled into continuous frame point clouds of foreground targets and continuous frame point clouds of background targets. Among them, the foreground targets are movable targets in the driving scene, and the background targets are immovable targets in the driving scene. Taking a set number of foreground target continuous frame point clouds for point cloud superposition to obtain a foreground target superposition point cloud, in a time period corresponding to the set number of foreground target continuous frame point clouds, extracting key frame point clouds from the background target continuous frame point clouds according to a longer step length to perform point cloud superposition to obtain a background target superposition point cloud; Determine the ground surface of the background part, and perform dense processing on the determined ground points to obtain a dense ground surface; Use fine grid to rasterize foreground objects, and use coarse grid to rasterize background objects; After rasterization is completed, the degree of geometric appearance loss of the foreground object is determined, and the geometric appearance of the foreground object is completed according to the degree of geometric appearance loss; The geometrically completed foreground object grid is merged with the background grid to generate the complete ground truth of the autonomous driving scene occupancy.
[0047] Of course, the computer program stored in the storage unit of a fast dense occupancy true value generating device provided by an embodiment of the present invention is not limited to the method operations described above, and can also execute related operations in a fast dense occupancy true value generating method provided by any embodiment of the present invention.
[0048] Example 3 An embodiment of the present invention provides a computer-readable storage medium, wherein the computer-readable storage medium stores a computer program. When the computer program is executed, the method for quickly generating a densely occupied true value is implemented, including: Obtain the point cloud of the driving scene annotated with semantics and 3D bounding boxes, extract the continuous frame point cloud, and transform the continuous frame point cloud into a unified coordinate system; The foreground and background of each frame point cloud are segmented according to semantics and 3D bounding boxes to obtain foreground targets and background targets. The continuous frame point clouds are decoupled into continuous frame point clouds of foreground targets and continuous frame point clouds of background targets. Among them, the foreground targets are movable targets in the driving scene, and the background targets are immovable targets in the driving scene. Taking a set number of foreground target continuous frame point clouds for point cloud superposition to obtain a foreground target superposition point cloud, in a time period corresponding to the set number of foreground target continuous frame point clouds, extracting key frame point clouds from the background target continuous frame point clouds according to a longer step length to perform point cloud superposition to obtain a background target superposition point cloud; Determine the ground surface of the background part, and perform dense processing on the determined ground points to obtain a dense ground surface; Use fine grid to rasterize foreground objects, and use coarse grid to rasterize background objects; After rasterization is completed, the degree of geometric appearance loss of the foreground object is determined, and the geometric appearance of the foreground object is completed according to the degree of geometric appearance loss; The geometrically completed foreground object grid is merged with the background grid to generate the complete ground truth of the autonomous driving scene occupancy.
[0049] A computer-readable storage medium provided in an embodiment of the present invention stores a computer program which is not limited to the method operations described above, but can also execute related operations in a fast dense occupancy true value generation method provided in any embodiment of the present invention.
[0050] In the embodiments provided by the present invention, it should be understood that the disclosed structures and methods can be implemented in other ways. For example, the structural embodiments described above are only schematic. For example, the division of the units is only a logical function division. There may be other division methods in actual implementation, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be an indirect coupling or communication connection through some interfaces, structures or units, which can be electrical, mechanical or other forms.
[0051] The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed on multiple network units. Some or all of the units may be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0052] In addition, each functional unit in each embodiment of the present invention may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit. The above-mentioned integrated unit may be implemented in the form of hardware or in the form of software functional units.
[0053] The foregoing is merely a specific embodiment of the present invention, which enables those skilled in the art to understand or implement the present invention. Various modifications to these embodiments will be apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the present invention. Therefore, the present invention will not be limited to the embodiments shown herein, but rather to the widest scope consistent with the principles and novel features claimed herein.
Claims
1. A fast method for generating dense occupancy truth values, characterized in that: include: Obtain the point cloud of the driving scene annotated with semantics and 3D bounding boxes, extract the continuous frame point cloud, and transform the continuous frame point cloud into a unified coordinate system; The foreground and background of each frame point cloud are segmented according to semantics and 3D bounding boxes to obtain foreground targets and background targets. The continuous frame point clouds are decoupled into continuous frame point clouds of foreground targets and continuous frame point clouds of background targets. Among them, the foreground targets are movable targets in the driving scene, and the background targets are immovable targets in the driving scene. Taking a set number of foreground target continuous frame point clouds for point cloud superposition to obtain a foreground target superposition point cloud, in a time period corresponding to the set number of foreground target continuous frame point clouds, extracting key frame point clouds from the background target continuous frame point clouds according to a longer step length to perform point cloud superposition to obtain a background target superposition point cloud; Determine the ground surface of the background part, and perform dense processing on the determined ground points to obtain a dense ground surface; Use fine grid to rasterize foreground objects, and use coarse grid to rasterize background objects; After rasterization is completed, the degree of geometric appearance loss of the foreground object is determined, and the geometric appearance of the foreground object is completed according to the degree of geometric appearance loss; The geometrically completed foreground object grid is merged with the background grid to generate the complete ground truth of the autonomous driving scene occupancy.
2. The fast dense occupancy truth value generation method according to claim 1, characterized in that: The ground of the background is determined, and the determined ground points are densely processed to obtain a dense ground, including: S410, superimpose point cloud on background target Divided into the area surrounding the detection vehicle Concentric ring areas: , ; Among them, the background target superimposed point cloud Middle Concentric ring areas , For the The inner diameter of the concentric ring area is For the The outer diameter of the concentric ring area, For point The radial distance from the center O; S420, each concentric ring area Further division along the radial and circumferential directions yields several sub-areas: , in, The concentric ring area The mth radial sub-area, the nth circumferential sub-area, the concentric ring area The radial direction is divided into M areas, and the circumferential direction is divided into N areas. , For point Distance concentric area Radial distance of the inner diameter, For point azimuth; S430, removing bottom points whose height is lower than the height threshold and whose reflection intensity is lower than the set reflection intensity threshold according to the height threshold dynamically adjusted by the average value and standard deviation of the ground elevation; S440, performing vertical plane fitting on all sub-regions to remove vertical points on the vertical plane; S450, performing ground plane fitting on all sub-regions to obtain all ground planes in the sub-regions; S460, using the normal verticality, elevation and flatness of the ground plane point set, and the normal verticality, elevation and flatness of all the ground in the concentric annular area where the ground plane point set is located, ground matching is performed to exclude the non-ground ground plane point sets in each ground plane point set, and determine the ground plane point set. S470, after the ground is determined, uniform points are constructed to cover the ground.
3. The fast dense occupancy truth value generation method according to claim 2, characterized in that: The sum of the average value of all the concentric annular areas and the gain standard deviation is used as the new height threshold.
4. The fast dense occupancy truth value generation method according to claim 2, characterized in that: The process of vertical plane fitting includes: S441, starting from any lowest point in any sub-region, selecting a set number of points as seed points, and calculating the position mean and unit normal vector of the seed points; S442, for any candidate point that is not determined to be a vertical point, when the product of the difference between its position and the position mean of the seed segment and the unit normal vector of the seed point is less than the set distance tolerance of the estimated vertical plane, the candidate point is selected as a potential vertical point, and the process is repeated multiple times to select potential vertical points in the sub-region; S443, if the inverse of the cosine of the product of the unit normal vector of the potential vertical point and the unit normal vector of the z-axis is greater than Potential perpendicular points are chosen as perpendicular points when the difference is less than the tolerance of the unit normal vector of the estimated plane.
5. The fast dense occupancy truth value generation method according to claim 2, characterized in that: Fitting the ground plane to all sub-areas to obtain all ground planes in the sub-areas includes: S451, selecting a set number of lowest points from any sub-region as seed points, calculating the average height of the seed points, and taking points whose heights are lower than the sum of the average height and the height threshold as an initial estimated ground plane point set; S452, calculating the normal vector of the ground plane point set obtained each time, calculating the average position of the ground plane point set obtained each time, and calculating the plane coefficient of the ground plane point set using the normal vector of the ground plane point set and the average position of the ground plane point set; S453, when iteratively selecting the subsequent ground plane point set, the plane coefficient formula is used to evaluate the coordinates of the candidate point and the normal vector of the previous ground plane point set. When the difference between the evaluated value and the plane coefficient of the previous ground plane point set is less than the set plane distance threshold, the candidate point is used as an element of the next ground plane point set.
6. The fast dense occupancy truth value generation method according to claim 2, characterized in that: Normal vector based on ground plane point set and angle threshold The upright screening function is: ; The average height of the points in the ground plane. The elevation screening function with elevation threshold is: ; in, As the distance from the point to the center An exponentially changing elevation threshold, is the elevation recognizable range threshold; The flatness function based on the minimum eigenvalue of the ground plane point set is: ; in, is the gain amplitude, is the minimum eigenvalue of the ground plane point set, is the flatness threshold, for the part where the elevation screening function is less than 0.5, that is, the average height of the points in the ground plane point concentration With elevation threshold When the difference is greater than zero, if the ground plane point set is flat, a larger value will be assigned by the flatness function, thereby recovering the lost true positive examples.
7. The fast dense occupancy truth value generation method according to claim 6, characterized in that: The elevation threshold of all concentric ring areas to determine the ground and flatness threshold The sum of the mean and the gain standard deviation is taken as the new threshold, and the gain standard deviation is the product of the gain coefficient and the standard deviation; the elevation threshold of the ground is determined by using all the concentric annular areas in the recent period of time and flatness threshold The sum of the mean and the gain standard deviation is taken as the restoration threshold. When the threshold is less than the restoration threshold, the ground plane point set between the threshold and the restoration threshold is restored.
8. The fast dense occupancy truth value generation method according to claim 1, characterized in that: Determine the degree of geometric appearance loss of the foreground target, and complete the geometric appearance of the foreground target according to the degree of geometric appearance loss, including: S610, first, according to the 3D intersection-and-union ratio of the actual 3D enclosing frame of the point cloud of the foreground object and the 3D truth frame, the 3D intersection-and-union ratio is used as the degree of geometric appearance loss; The 3D intersection-and-union ratio is compared with a set 3D intersection-and-union ratio threshold. If the 3D intersection-and-union ratio is lower than the set 3D intersection-and-union ratio threshold, S630 is executed. If the 3D intersection-and-union ratio is not lower than the set 3D intersection-and-union ratio threshold, S640 is executed. S630, selecting a pre-built dense occupancy grid model as a substitute according to the semantic category; S640, geometric appearance completion using Poisson reconstruction.
9. The fast dense occupancy truth value generation method according to claim 8, characterized in that: Use Poisson reconstruction to complete the geometric appearance. The detailed steps of the completion process include: S641, taking one of the two symmetric planes of the three-dimensional truth box that are perpendicular to the horizontal plane as a mirror plane, and performing a mirroring operation based on maximizing a new occupancy grid result generated after the mirroring; S642, treating the central point set of the mirrored foreground object as a point cloud and performing Poisson reconstruction to generate a continuous three-dimensional network; S643, the center point of the grid in the space covered by the Poisson reconstructed three-dimensional network is used as the center point of the grid generated after the geometric appearance is completed, and the category is unified as the foreground target category.
10. A fast dense occupancy truth value generation device, characterized in that: include: At least one processing unit, the processing unit is connected to a storage unit via a bus unit, the storage unit stores a computer program, and when the computer program is executed by the processing unit, the fast dense occupancy true value generation method as described in any of claims 1-9 is implemented.
Citation Information
Patent Citations
Dynamic and static truth value dense generation method for three-dimensional occupancy task
CN118196762A
Occupied grid-based mining area data set construction method, device and system
CN118968451A
Method, apparatus, and storage medium for three-dimensional reconstruction of buildings based on missing point cloud data
US20240257462A1
Free space generation method, movable platform and storage medium
WO2023000221A1
Shielding relationship determination method and apparatus, and storage medium and electronic device
WO2023065313A1