3D occupancy prediction method, device and system, and storage medium
By separating and merging point clouds and spreading semantics to voxels, the problems of data sparsity and labeling cost in 3D occupancy prediction are solved, and high-precision 3D occupancy prediction and voxel semantic consistency are achieved.
Patent Information
- Application Number
- CN202510101373.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-22
- Publication Date
- 2025-05-06
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
The existing 3D occupancy prediction methods face problems such as data sparsity, high labeling cost and difficulty in multi-frame integration, resulting in low accuracy and high calculation cost.
By acquiring point cloud and 3D object detection tags, dynamic and static point clouds are separated, multi-frame point cloud timing merging and density processing are performed, semantics are propagated to voxels, and voxels containing too few points are eliminated to generate dense 3D occupancy data.
Effectively reduce labeling costs, improve the generation quality of dense voxels and the accuracy of occupation prediction, and is suitable for complex dynamic scenarios.
Smart Images

Figure CN119942524A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of information processing technology, and in particular relates to a 3D occupancy prediction method and device, a system, and a storage medium. Background Art
[0002] 3D occupancy prediction is a key perception task in autonomous driving and robotic systems, which is used to generate occupancy information of the three-dimensional space in the scene. Currently, most 3D occupancy prediction methods rely on LiDAR data and require high-precision annotated datasets. However, traditional methods face the following problems:
[0003] 1. Data sparsity: LiDAR point clouds are inherently sparse, especially for long-distance targets, which results in low occupancy prediction accuracy.
[0004] 2. Complexity and cost: Annotating dense 3D point cloud datasets requires a lot of manpower and time, and the computational cost is high;
[0005] 3. Difficulty in multi-frame integration: In dynamic scenes, point clouds of different time frames are difficult to align effectively, which makes it difficult to accurately obtain the three-dimensional occupancy information of dynamic targets. Summary of the invention
[0006] The technical problem to be solved by the present invention is to provide a 3D occupancy prediction method and device, system, and storage medium.
[0007] To achieve the above object, the present invention adopts the following technical solution:
[0008] A 3D occupancy prediction method, comprising:
[0009] Step S1, obtaining point cloud and 3D object detection labels;
[0010] Step S2, separating the point cloud into a dynamic point cloud and a static point cloud;
[0011] Step S3, merging the dynamic point cloud and the static point cloud in a multi-frame point cloud time sequence to obtain a dense static point cloud and a dense dynamic point cloud respectively;
[0012] Step S4, merging dense dynamic point cloud and static point cloud;
[0013] Step S5, performing point cloud hole filling and densification on the merged dense dynamic point cloud and static point cloud to obtain a dense point cloud;
[0014] Step S6, voxelizing the dense point cloud;
[0015] Step S7: propagate the semantics in the 3D object detection label to the voxel;
[0016] Step S8: Check the semantic consistency of adjacent voxels and remove voxels containing too few points.
[0017] Preferably, step S3 is specifically as follows: reading dynamic targets, estimating the movement of targets through the labels of dynamic targets, restoring dynamic targets of different frames to the same frame and then merging them, and then merging multiple frames of static point clouds to obtain dense static point clouds and dense dynamic point clouds respectively.
[0018] Preferably, in step S5, the moving least square method is used to interpolate the hole area.
[0019] The present invention also provides a 3D occupancy prediction device, comprising:
[0020] The first processing module is used to obtain point clouds and 3D object detection labels;
[0021] The second processing module is used to separate the point cloud into a dynamic point cloud and a static point cloud;
[0022] The third processing module is used to perform multi-frame point cloud temporal merging on the dynamic point cloud and the static point cloud to obtain a dense static point cloud and a dense dynamic point cloud respectively;
[0023] A fourth processing module for merging dense dynamic point clouds and static point clouds;
[0024] The fifth processing module is used to fill holes and densify the merged dense dynamic point cloud and static point cloud to obtain a dense point cloud;
[0025] a sixth processing module for voxelizing the dense point cloud;
[0026] A seventh processing module for propagating semantics in 3D object detection labels to voxels;
[0027] The eighth processing module is used to check the semantic consistency of adjacent voxels and remove voxels containing too few points.
[0028] Preferably, the second processing module is used to estimate the movement of the target through the label of the dynamic target, restore the dynamic targets of different frames to the same frame and then merge them, and merge multiple frames of static point clouds to obtain dense static point clouds and dense dynamic point clouds respectively.
[0029] Preferably, the fifth processing module uses a moving least squares method to perform interpolation of the hole area.
[0030] The present invention also provides a 3D occupancy prediction system, comprising: a memory and a processor, wherein the memory stores a computer program executed by the processor, and the computer program executes a 3D occupancy prediction method when executed by the processor.
[0031] The present invention also provides a storage medium, on which a computer program is stored, and the computer program executes the 3D occupancy prediction method when running.
[0032] By utilizing the point cloud and 3D detection labels in the target detection data, point cloud separation, dynamic point cloud correction, multi-frame merging, densification processing and semantic propagation are performed to generate dense 3D occupancy data; the technical solution of the present invention is adopted to effectively reduce the annotation cost, improve the generation quality of dense voxels and the accuracy of occupancy prediction. BRIEF DESCRIPTION OF THE DRAWINGS
[0033] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on the provided drawings without paying creative work.
[0034] Figure 1 This is a flow chart of a 3D occupancy prediction method according to an embodiment of the present invention. DETAILED DESCRIPTION
[0035] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0036] In order to make the above-mentioned objects, features and advantages of the present invention more obvious and easy to understand, the present invention is further described in detail below with reference to the accompanying drawings and specific embodiments.
[0037] Embodiment 1:
[0038] like Figure 1 As shown, an embodiment of the present invention provides a 3D occupancy prediction method, including:
[0039] Step S1: Obtain target detection dataset, which should contain point cloud and 3D target detection labels
[0040] Step 1.1, point cloud data representation. Point cloud data consists of dimensional coordinate points and intensity information collected by the laser radar, which is specifically represented as follows:
[0041] P={(x i ,y i ,z i ,A i )|i=1,2,…,n}
[0042] Among them, P represents a single-frame point cloud, that is, the set of all points in a frame; i represents the ID of a single point in a single-frame point cloud; x i ,y i ,z i A represents the three-dimensional coordinates of the point with ID i in the laser radar coordinate system; i Indicates the intensity value returned by the point with ID i; n is the number of points contained in the point cloud of this frame.
[0043] Step 1.2: 3D target detection label representation. Use a bounding box to represent the annotation information of target detection, which is defined as:
[0044] B=(x j ,y j ,z j ,w j ,h j ,l j ,r j ,S j ,ID j )
[0045] Among them, B represents a label; x j ,y j ,z j represents the three-dimensional center coordinates of the jth bounding box; w j ,h j ,l j represents the width, height and length of the jth bounding box; r j Indicates the orientation angle of the bounding box; S j Indicates the semantic label of the jth bounding box, such as vehicle or pedestrian; ID j is the target ID of the jth bounding box.
[0046] Step S2: Separation of dynamic and static point clouds
[0047] Define pedestrians and vehicle traffic participants as dynamic targets. The point clouds returned by such targets generally change over time, mainly in terms of position and posture. The point clouds returned by such targets are called "dynamic point clouds". Define other targets except dynamic targets as static environments. The point clouds returned by static environments will not change over time. The point clouds returned by such environments are called "static point clouds". Use the labels defined in step 1.2 to separate the dynamic point clouds and static point clouds from the single-frame point cloud P:
[0048]
[0049] P static =P\P dynamic
[0050] Among them, P dynamic represents a dynamic point cloud; pi Represents a dynamic point; B j Represents the j-th 3D object detection label corresponding to the point cloud P; M represents the number of objects contained in the point cloud P; P static Represents a static point cloud.
[0051] Step S3, Temporal merging of multi-frame point clouds
[0052] The point clouds separated in Step 2 are sparse. Therefore, it is necessary to merge the points of multi-frame point clouds to generate a dense point cloud. First, read the dynamic objects, estimate the motion of the objects through the labels of the dynamic objects, then restore the dynamic objects in different frames to the same frame and merge them, and then merge the multi-frame static point clouds to obtain a dense static point cloud and a dense dynamic point cloud respectively.
[0053] Step 3.1, Read the motion information of dynamic objects. Using the instance ID and bounding box information provided in the label, directly associate the point clouds belonging to the same dynamic object in consecutive frames. The temporal point cloud trajectory of the dynamic object is expressed as:
[0054]
[0055] Among them, Represents the temporal point cloud trajectory belonging to the i-th dynamic object, that is, the point cloud in consecutive frames is associated according to the ID j To form a set of points; Represents a single point numbered i in the point cloud of the t-th frame, including the three-dimensional coordinates and intensity information of the point Represents the bounding box of the i-th dynamic object in the r-th frame; ID j Represents the unique object ID of the dynamic object, used to share the same ID for the bounding boxes of the same object across different frames at different times j ; T represents the total number of frames in the time series.
[0056] Step 3.2, Estimate the pose transformation of the dynamic object between different frames. Calculate the motion parameters of the dynamic object through the bounding box position and orientation information provided in the label. The position change (Δx j , Δy j , Δz j ) and the orientation change (Δr j ) are used to generate a translation vector and a rotation matrix:
[0057]
[0058] The combined motion parameters are the pose transformation matrix:
[0059]
[0060] Cumulatively calculate the pose transformation matrix:
[0061]
[0062] in, represents the change in the center position of the jth dynamic target between the tth frame and the t-1th frame, which are the displacements in the x, y, and z directions respectively; represents the change in the orientation angle of the i-th dynamic target between the t-th frame and the t-1-th frame; represents the translation vector of the i-th dynamic target between the t-th frame and the t-1-th frame; Represents the rotation matrix of the i-th dynamic target between the t-th frame and the t-1-th frame, which is used to describe the change in the orientation of the target, assuming that the orientation angle changes only in the xy plane; Represents the pose transformation matrix of the i-th dynamic target between the t-th frame and the t-1-th frame, which is used to transform the point from the coordinate system of the t-th frame to the coordinate system of the t-1 frame; Represents the cumulative pose transformation matrix of the i-th dynamic target from the initial frame to the t-th frame, which is used to restore the point cloud of the target to the initial coordinate system. The initial pose transformation matrix I is the identity matrix.
[0063] Step 3.3: Use the accumulated pose transformation matrix to restore the dynamic point cloud to a unified coordinate system:
[0064]
[0065] in, Represents the restored point cloud set of the jth dynamic target; represents the i-th point in the t-th frame; represents the cumulative pose transformation matrix of the jth dynamic target in the tth frame; Represents the bounding box of the jth dynamic target in the tth frame.
[0066] Step 3.4: Merge the static point clouds in all frames into a unified coordinate system:
[0067]
[0068] in, Represents the static point cloud set after all frames are merged; Represents the static environment point cloud set in the tth frame.
[0069] Step S4: Merge dense dynamic and static point clouds
[0070]
[0071] Step S5: Hole filling and densification of point cloud
[0072] Hole filling and densification of point clouds. In the generated dense point cloud, there may be holes caused by sensor occlusion or distance. In order to ensure the integrity of the point cloud, the present invention introduces the moving least squares method to interpolate the hole area.
[0073] Step 5.1: Determine the hole area. Calculate each point p i ∈P final The number of neighboring points |N(p i )|:
[0074] N(p i )={q∈P final |||p i -q||≤r}
[0075] Where r is the neighborhood search radius.
[0076] If |N(p i )|<τ(threshold), then point p i The area is considered a void area:
[0077] SparseRegion = {p i ∈P final ||N(p i )|<τ}
[0078] Step 5.2: Fill the holes. Define the weight function as:
[0079]
[0080] Among them, h d is the kernel bandwidth parameter of the spatial distance, h i is the kernel bandwidth parameter of the point intensity, and both parameters are specified manually. ||p i -q|| is point p i The Euclidean space distance from the neighboring point q, |A i -A q | represents the difference between point intensities.
[0081] In each neighborhood, the densification of the sparse area is completed by minimizing the fitting error of the local plane. The error function is defined as:
[0082]
[0083] Where n is the normal vector of the local plane. N(p i ) represents point p i Neighborhood, N(p i )={q∈P final |||p i -q||≤r}, r is the neighborhood search radius.
[0084] The error function is optimized through error, and a smooth local surface is fitted to fill the hole area:
[0085]
[0086] Among them, p interpolated Represents the newly generated interpolation points.
[0087] Finally, the dense point cloud P dense It is expressed as:
[0088] P dense =P final ∪P interpolated
[0089] Step S6: Point cloud voxelization
[0090] The dense point cloud P dense Convert to voxel representation for further processing and analysis.
[0091] Step 6.1: Divide the voxel grid. Divide the three-dimensional space of the point cloud into voxel grids of uniform size, and define the voxel side length as (d x ,d y ,d z ). Calculate the boundary range of the point cloud:
[0092] [x min ,x max ],[y min ,y max ],[z min ,z max ]
[0093] Calculate the number of voxel grids:
[0094]
[0095] in, Indicates rounding up.
[0096] Step 6.2: Voxelize the point cloud. Traverse the point cloud set P dense ; Calculate the voxel index (i, j, k) of each point according to its coordinates:
[0097]
[0098] in, Indicates rounding down.
[0099] Let point p∈P dense Assign to the corresponding voxel:
[0100] V(i,j,k)={p|p∈P dense , index is (i,j,k)}
[0101] Among them, V(i,j,k) represents the voxel with index (i,j,k).
[0102] Step S7: assign semantics to voxels based on 3D object detection labels
[0103] Capture the semantics in the labels and then propagate the semantics to the voxels.
[0104] Step 7.1: Get semantic labels. For each point cloud point p = (x, y, z, A), read the semantic information S(p) from the 3D object detection label.
[0105] Step 7.2: Propagate semantics to voxels. For each voxel V(i,j,k), count the frequency of each semantic label:
[0106] f s =|{p|p∈V(i,j,k),S(p)=s}|
[0107] Among them, f s Represents the frequency of occurrence of semantic tag s; s represents a semantic category.
[0108] Select the semantic label with the highest frequency as the semantic label of the voxel:
[0109]
[0110] If multiple semantic tags have the same frequency, the tag is determined based on the custom priority.
[0111] Step S8: Voxel optimization
[0112] Neighboring voxels are checked for semantic consistency and voxels containing too few points are removed.
[0113] Step 8.1. Check the semantic consistency of adjacent voxels. Define the neighborhood of voxel V(i,j,k):
[0114] N(V)={V(i',j',k')||i-i'|≤1,|j-j'|≤1,|k-k'|≤1,V(i',j',k')≠V(i,j,k)}
[0115] If the semantic labels of most neighboring voxels are the same, the semantic label of the current voxel is corrected.
[0116] Step 8.2: Remove voxels with too few points. If a voxel V(i,j,k) contains less than 2 points, the voxel is considered to be a sparse voxel. For non-sparse voxels, retain their original semantic labels obtained through semantic propagation and statistical calculation:
[0117]
[0118] The voxelized semantic point cloud can be represented as a set: P voxel-semantic ={(i,j,k,S(V(i,j,k)))|(i,j,k) is the voxel index}, where each element contains the voxel index (i,j,k) and the corresponding semantic label S(V(i,j,k)).
[0119] The present invention uses target detection tags to realize the automatic generation of 3D occupancy prediction data sets without additional annotation, which greatly reduces the cost of data preparation. It improves the densification and semantic consistency of point clouds, effectively improves the quality of dense point clouds through dynamic point cloud correction, multi-frame merging and densification processing; ensures the accuracy and consistency of voxel semantics through semantic propagation and optimization. It supports dynamic and static scene modeling to separate dynamic targets from static environments, and corrects the motion information of dynamic targets, which is suitable for complex dynamic scenes.
[0120] Embodiment 2:
[0121] An embodiment of the present invention further provides a 3D occupancy prediction device, comprising:
[0122] The first processing module is used to obtain point clouds and 3D object detection labels;
[0123] The second processing module is used to separate the point cloud into a dynamic point cloud and a static point cloud;
[0124] The third processing module is used to perform multi-frame point cloud temporal merging on the dynamic point cloud and the static point cloud to obtain a dense static point cloud and a dense dynamic point cloud respectively;
[0125] A fourth processing module for merging dense dynamic point clouds and static point clouds;
[0126] The fifth processing module is used to fill holes and densify the merged dense dynamic point cloud and static point cloud to obtain a dense point cloud;
[0127] a sixth processing module for voxelizing the dense point cloud;
[0128] A seventh processing module for propagating semantics in 3D object detection labels to voxels;
[0129] The eighth processing module is used to check the semantic consistency of adjacent voxels and remove voxels containing too few points.
[0130] As an implementation mode of an embodiment of the present invention, the second processing module is used to estimate the movement of the target through the label of the dynamic target, restore the dynamic targets of different frames to the same frame and then merge them, and merge multiple frames of static point clouds to obtain dense static point clouds and dense dynamic point clouds respectively.
[0131] As an implementation of the embodiment of the present invention, the fifth processing module uses the moving least square method to perform interpolation of the hole area.
[0132] Embodiment 3:
[0133] An embodiment of the present invention further provides a 3D occupancy prediction system, comprising: a memory and a processor, wherein a computer program executed by the processor is stored in the memory, and the computer program executes a 3D occupancy prediction method when executed by the processor.
[0134] Embodiment 4:
[0135] An embodiment of the present invention further provides a storage medium, on which a computer program is stored, and the computer program executes the 3D occupancy prediction method when running.
[0136] The embodiments described above are only descriptions of the preferred embodiments of the present invention and are not intended to limit the scope of the present invention. Without departing from the design spirit of the present invention, various modifications and improvements made to the technical solutions of the present invention by ordinary technicians in this field should all fall within the protection scope determined by the claims of the present invention.
Claims
1. A 3D occupancy prediction method, characterized in that: include: Step S1, obtaining point cloud and 3D object detection labels; Step S2, separating the point cloud into a dynamic point cloud and a static point cloud; Step S3, merging the dynamic point cloud and the static point cloud in a multi-frame point cloud time sequence to obtain a dense static point cloud and a dense dynamic point cloud respectively; Step S4, merging dense dynamic point cloud and static point cloud; Step S5, performing point cloud hole filling and densification on the merged dense dynamic point cloud and static point cloud to obtain a dense point cloud; Step S6, voxelizing the dense point cloud; Step S7: propagate the semantics in the 3D object detection label to the voxel; Step S8: Check the semantic consistency of adjacent voxels and remove voxels containing too few points.
2. The 3D occupancy prediction method according to claim 1, characterized in that: Step S3 is specifically as follows: reading dynamic targets, estimating the movement of targets through the labels of dynamic targets, restoring dynamic targets of different frames to the same frame and then merging them, and then merging multiple frames of static point clouds to obtain dense static point clouds and dense dynamic point clouds respectively.
3. The 3D occupancy prediction method according to claim 2, characterized in that: In step S5, the moving least square method is used to interpolate the hole area.
4. A 3D occupancy prediction device, characterized in that: include: The first processing module is used to obtain point clouds and 3D object detection labels; The second processing module is used to separate the point cloud into a dynamic point cloud and a static point cloud; The third processing module is used to perform multi-frame point cloud temporal merging on the dynamic point cloud and the static point cloud to obtain a dense static point cloud and a dense dynamic point cloud respectively; A fourth processing module for merging dense dynamic point clouds and static point clouds; The fifth processing module is used to fill holes and densify the merged dense dynamic point cloud and static point cloud to obtain a dense point cloud; a sixth processing module for voxelizing the dense point cloud; A seventh processing module for propagating semantics in 3D object detection labels to voxels; The eighth processing module is used to check the semantic consistency of adjacent voxels and remove voxels containing too few points.
5. The 3D occupancy prediction device according to claim 4, characterized in that: The second processing module is used to estimate the movement of the target through the label of the dynamic target, restore the dynamic targets of different frames to the same frame and then merge them, and merge the static point clouds of multiple frames to obtain dense static point clouds and dense dynamic point clouds respectively.
6. The 3D occupancy prediction device according to claim 5, characterized in that: The fifth processing module uses the moving least square method to interpolate the hole area.
7. A 3D occupancy prediction system, characterized in that: include: A memory and a processor, wherein the memory stores a computer program executed by the processor, and when the computer program is executed by the processor, the 3D occupancy prediction method according to any one of claims 1 to 3 is executed.
8. A storage medium, characterized in that: The storage medium stores a computer program, which executes the 3D occupancy prediction method according to any one of claims 1 to 3 when running.
Citation Information
Patent Citations
Earthwork volume calculation method, device, equipment and medium
CN112906124A
Three-dimensional reconstruction-based method for identifying and detecting haemaphysalis bloodleri
CN115205898A
Point cloud semantic segmentation method and system based on voxel clustering and sparse convolution
CN115984564A
Label determination method and device, equipment and medium
CN118762235A
Deep learning system for performing private inference and operating method thereof
KR1020230159204A