A method for pre-training lidar point clouds
The voxel encoder and backbone of the detection network are pre-trained by lidar point cloud pre-training method, which solves the problem of high dependence on labeled data in the prior art, and achieves a significant improvement in performance under limited labeled data.
Patent Information
- Application Number
- CN202310161269.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-02-23
- Publication Date
- 2025-05-30
- Estimated Expiration
- 2043-02-23
AI Technical Summary
The existing lidar point cloud detection network has a high dependence on labeled data, which makes it difficult to improve performance under limited labeled data.
A lidar point cloud pre-training method is proposed, which reduces the dependence on labeled data by dividing point cloud space, sampling voxels, performing position or shape masking, and input masked voxels and non-masked voxels into voxel encoder and Transformer backbone.
This significantly improves the performance of the detection network under limited annotated data, and can achieve higher detection performance when using less annotated data.
Smart Images

Figure CN116228704B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of autonomous driving, and more particularly, to a method for pre-training lidar point clouds. Background Art
[0002] Lidar point cloud detection is to detect objects from the scene point clouds obtained by lidar scanning, and output the three-dimensional bounding boxes and categories of the objects, which is often applied to the outdoor vehicle autonomous driving scenario.
[0003] Existing lidar point cloud detection networks require a large amount of labeled data for training. The detection network has a very high dependence on the labeled data, and it is difficult to improve the performance of the network under limited labeled data.
[0004] In view of this, this application is specifically proposed. Summary of the Invention
[0005] The purpose of the present invention is to provide a method for pre-training lidar point clouds, which can pre-train the voxel encoder and backbone of the detection network, thereby significantly reducing the dependence of the detection network on labeled data, and enabling the network to significantly improve its performance even when trained with limited labeled data.
[0006] The embodiments of the present invention are implemented as follows:
[0007] A method for pre-training lidar point clouds, comprising the following steps:
[0008] S1: Define the point cloud space and input the point cloud of the entire scene into the point cloud space;
[0009] S2: Divide the point cloud space into a number of voxels of equal size, and group the point clouds located within the same voxel;
[0010] S3: Determine the non-empty voxels in the point cloud space;
[0011] S4: Sample the non-empty voxels, and the sampled non-empty voxels are used as non-masked voxels, and the non-empty voxels that are not sampled are used as masked voxels;
[0012] S5: Perform position masking or shape masking on the masked voxels, and replace the information to be masked with learnable feature information; wherein, one masked voxel adopts one masking form;
[0013] S6: Input the masked voxels and non-masked voxels into the voxel encoder, and output the feature vectors representing each voxel;
[0014] S7: Input the feature vectors of each voxel into the Transformer backbone to extract features, and output the feature data of each voxel;
[0015] S8: Recover the masked information of each masked voxel according to the feature data of the masked voxel; and
[0016] S9: Retain and transfer the weights of the voxel encoder and the Transformer backbone pre-trained in steps S1 to S8 for use in downstream tasks.
[0017] Further, in step S4, farthest point sampling is used to sample non-empty voxels according to a preset sampling ratio.
[0018] Further, in step S4, the sampling process is repeatedly executed to obtain voxels for position masking and shape masking respectively.
[0019] Further, each point in a non-empty voxel has three-dimensional coordinates in the point cloud space, a first offset coordinate relative to the voxel center, and a second offset coordinate relative to the cluster center of the points within the voxel;
[0020] In step S5, the position mask includes: masking the three-dimensional coordinates of all points within the masked voxel while retaining the first offset coordinate and the second offset coordinate;
[0021] The shape mask includes: within the masked voxel, selecting a point as a reference point among the points belonging to the same group, and masking the three-dimensional coordinates, the first offset coordinate, and the second offset coordinate of all points except the reference point within the same group. The reference point is used to provide position information.
[0022] Further, in step S8, when recovering the masked information of the masked voxel, the recovered position mask includes: inputting the feature data of the position masked voxel to the classification head and outputting the position index of the voxel.
[0023] Further, the cross-entropy loss function is calculated through the output position index and the ground truth to measure the effect of recovering the masked information.
[0024] Further, in step S8, when recovering the masked information of the masked voxel, the recovered shape mask includes: inputting the feature data of the shape masked voxel to the reconstruction head and outputting the first offset coordinate of the points within the voxel.
[0025] Further, the two-norm chamfer distance is calculated using the output first offset coordinate and the true offset of all points within the voxel relative to the voxel center to measure the effect of recovering the masked information.
[0026] The beneficial effects of the technical solution of the embodiment of the present invention include:
[0027] The lidar point cloud pre-training method provided by the embodiments of the present invention jointly models the distribution of voxels and the distribution of point clouds in voxels in the scene by using the hierarchical relationship of the scene, voxels, and point clouds in the point cloud space. When modeling the voxel distribution within the scene, the method masks the position information of the voxels, retains the geometric structure information of the voxels, and constructs a jigsaw task for the neural network to restore the positions of the voxels. When modeling the distribution of point clouds within the voxels, the method masks the geometric structure information of the voxels, retains the position information of the voxels, and allows the neural network to restore the geometric structure of the point clouds within the voxels.
[0028] The lidar point cloud pre-training method can not only model the structure of points within the voxels, but also model the positional relationship between the voxels, while taking into account the distribution characteristics of the lidar point cloud.
[0029] The lidar point cloud pre-training method does not need to construct corresponding relationships like the contrastive learning method, avoiding the ambiguity and instability of training. At the same time, it can jointly model the distribution of voxels and the distribution of points within the voxels, and can be better used to model the key information of lidar point cloud detection. The two masking methods it proposed are unified on the point cloud data, making it more unified and concise, and can better adapt to the characteristics of uneven distribution and denser near and sparser far of the lidar point cloud. This method can significantly improve the performance of the detector.
[0030] Generally speaking, the lidar point cloud pre-training method provided by the embodiments of the present invention can pre-train the voxel encoder and backbone of the detection network, thus significantly reducing the dependence of the detection network on labeled data, and enabling the network to significantly improve its performance even when trained with limited labeled data. BRIEF DESCRIPTION OF THE DRAWINGS
[0031] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following will briefly introduce the drawings required in the embodiments. It should be understood that the following drawings only show some embodiments of the present invention, and therefore should not be regarded as limiting the scope. For those of ordinary skill in the art, without creative efforts, other related drawings can be obtained based on these drawings.
[0032] Figure 1 It is a schematic flowchart of the lidar point cloud pre-training method provided by the embodiments of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0033] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the drawings in the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. Usually, the components of the embodiments of the present invention described and shown in the drawings here can be arranged and designed in various different configurations.
[0034] Accordingly, the following detailed description of the embodiments of the present invention provided in the accompanying drawings is not intended to limit the scope of the claimed invention, but merely represents selected embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the scope of protection of the present invention.
[0035] It should be noted that like reference numerals and letters denote like items in the following drawings. Therefore, once an item is defined in one drawing, it does not need to be further defined and explained in subsequent drawings.
[0036] The terms "first", "second", etc. are used only for descriptive distinction and should not be construed as indicating or implying relative importance.
[0037] As shown in this specification and the claims, unless the context clearly indicates otherwise, words such as "a", "the", etc. do not specifically refer to the singular and may also include the plural. Generally speaking, the terms "comprising" and "including" only indicate the inclusion of the clearly identified steps and elements, and these steps and elements do not constitute an exclusive list. The method or device may also include other steps or elements.
[0038] The flowcharts used in this specification are used to illustrate the operations performed by the system according to the embodiments of this specification. It can be understood that the operations of each step do not necessarily need to be executed precisely in sequence. On the contrary, they can be executed in reverse order or simultaneously. At the same time, other operations can also be added to these processes, or one or several steps of operations can be removed from these processes.
[0039] Embodiment
[0040] Please refer to Figure 1 , this embodiment provides a method for pre-training lidar point clouds. The method for pre-training lidar point clouds includes the following steps:
[0041] S1: Define the point cloud space and input the point cloud of the entire scene into the point cloud space;
[0042] S2: Divide the point cloud space into a number of voxels of equal size, and group the point clouds located within the same voxel;
[0043] S3: Determine the non-empty voxels in the point cloud space;
[0044] S4: Sample the non-empty voxels. The sampled non-empty voxels are used as non-mask voxels, and the non-empty voxels that are not sampled are used as mask voxels;
[0045] S5: Perform position masking or shape masking on the masked voxels, and replace the information to be masked with learnable feature information; wherein, one masked voxel adopts one masking form;
[0046] S6: Input the masked voxels and unmasked voxels into the voxel encoder, and output the feature vectors representing each voxel;
[0047] S7: Input the feature vectors of each voxel into the Transformer backbone to extract features, and output the feature data of each voxel;
[0048] S8: Restore the masked information of the masked voxel according to the feature data of each masked voxel; and
[0049] S9: Retain and transfer the weights of the voxel encoder and Transformer backbone pre-trained through steps S1 to S8 for use in downstream tasks.
[0050] In this embodiment, in step S4, farthest point sampling is used to sample the non-empty voxels according to a preset sampling ratio. Farthest point sampling is used to sample the voxels, and the sampled voxels are retained, while the non-sampled voxels are subjected to corresponding masking and restoration, so as to better adapt to the characteristics of uneven distribution of lidar point clouds, dense near and sparse far.
[0051] In step S4, the sampling process is repeatedly executed to obtain voxels for position masking and voxels for shape masking respectively.
[0052] Wherein, each point in each non-empty voxel has three-dimensional coordinates in the point cloud space, a first offset coordinate relative to the voxel center, and a second offset coordinate relative to the clustering center of the points inside the voxel;
[0053] In step S5, the position masking includes: masking the three-dimensional coordinates of all points in the masked voxel, while retaining the first offset coordinate and the second offset coordinate. Thus, the position information is masked while retaining the geometric structure information of the voxel.
[0054] The shape masking includes: in the masked voxel, select one point as a reference point among the points belonging to the same group, and mask the three-dimensional coordinates, the first offset coordinate, and the second offset coordinate of all points except the reference point in the same group. The reference point is used to provide position information. Thus, the geometric structure information of the voxel is masked, and the remaining one point is used to provide position information.
[0055] Further, in step S8, when restoring the masked information of the masked voxel, the restored position mask includes: inputting the feature data of the position masked voxel into the classification head, outputting the position index of the voxel, and establishing the index of the position, similar to the jigsaw process. This process is also equivalent to a classification problem, and the total number of classification categories is the number of all possible positions. The effect of restoring the masked information can be measured by calculating the cross-entropy loss function between the output position index and the ground truth.
[0056] In step S8, when restoring the masked information of the masked voxel, the restored shape mask includes: inputting the feature data of the shape masked voxel into the reconstruction head, and outputting the first offset coordinates of the points within the voxel. The effect of restoring the masked information is measured by calculating the two-norm chamfer distance between the output first offset coordinates and the true offsets of all points within the voxel relative to the voxel center.
[0057] Wherein, after inputting the feature data of the shape masked voxel into the reconstruction head, the second offset coordinates of the points within the voxel can also be output.
[0058] The lidar point cloud pre-training method jointly models the distribution of voxels and the distribution of point clouds within voxels in the scene using the hierarchical relationship among the scene, voxels, and point clouds in the point cloud space. When modeling the voxel distribution within the scene, this method masks the position information of the voxels, retains the geometric structure information of the voxels, and constructs a jigsaw task for the neural network to restore the positions of the voxels. When modeling the distribution of point clouds within voxels, this method masks the geometric structure information of the voxels, retains the position information of the voxels, and allows the neural network to restore the geometric structure of the point clouds within the voxels.
[0059] The lidar point cloud pre-training method can not only model the structure of the points within the voxels but also model the position relationship between the voxels, while considering the distribution characteristics of the lidar point cloud.
[0060] The lidar point cloud pre-training method does not need to construct corresponding relationships like the contrastive learning method, avoiding the ambiguity and instability of training. At the same time, it can jointly model the distribution of voxels and the distribution of points within voxels, and can be better used to model the key information of lidar point cloud detection. The two proposed masking methods are unified on the point cloud data, making it more unified and concise, and can better adapt to the characteristics of uneven distribution and denser near and sparser far of the lidar point cloud. This method can significantly improve the performance of the detector.
[0061] Experiments on the Waymo Autonomous Driving Dataset and the KITTI Dataset show that after pre-training with the lidar point cloud pre-training method provided by the embodiments of the present invention, the lidar point cloud detector can achieve higher performance than without pre-training when using different amounts of labeled data for training in downstream tasks.
[0062] Specifically, on the Waymo dataset, when only 5% of the labeled data is used for training, the L2 mAPH metric for measuring detection performance increases from 40.34% to 46.68%; when 10% of the data is used, it increases from 50.46% to 54.06%. In contrast, the currently best published method, ProposalContrast, increases from 40.34% to 42.58% at 5%; and at 10% of the data, it harms the detector performance, decreasing from 50.46% to 50.13%.
[0063] In addition, the pre-training of the present invention can improve the detection performance of vehicles, pedestrians, and bicycles, not limited to the detection of a certain type of object, comprehensively demonstrating the effectiveness of the method.
[0064] In summary, the lidar point cloud pre-training method provided by the embodiments of the present invention can pre-train the voxel encoder and backbone of the detection network, thereby significantly reducing the dependence of the detection network on labeled data and enabling the network to significantly improve performance even when trained with limited labeled data.
[0065] The above are only the preferred embodiments of the present invention and are not intended to limit the present invention. For those skilled in the art, the present invention may have various modifications and changes. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.
Claims
1. A pre-training method for lidar point cloud, characterized in that, it includes the following steps: S1: Define the point cloud space and input the point cloud of the entire scene into the point cloud space; S2: Divide the point cloud space into a number of voxels of equal size, and group the point clouds located within the same voxel; S3: Determine the non-empty voxels in the point cloud space; S4: Sample the non-empty voxels. The sampled non-empty voxels are used as non-masked voxels, and the non-empty voxels that are not sampled are used as masked voxels; S5: Perform position masking or shape masking on the masked voxels, and replace the information to be masked with learnable feature information; wherein, one masked voxel adopts one masking form; S6: Input the masked voxels and non-masked voxels into the voxel encoder to output the feature vector representing each voxel; S7: Input the feature vector of each voxel into the Transformer backbone to extract features and output the feature data of each voxel; S8: Restore the masked information of the masked voxel according to the feature data of each masked voxel; and S9: Retain and transfer the weights of the voxel encoder and Transformer backbone pre-trained in steps S1 to S8 for use in downstream tasks; Each point in each non-empty voxel has three-dimensional coordinates in the point cloud space, a first offset coordinate relative to the voxel center, and a second offset coordinate relative to the clustering center of the points within the voxel; In step S5, the position masking includes: masking the three-dimensional coordinates of all points within the masked voxel while retaining the first offset coordinate and the second offset coordinate; The shape masking includes: within the masked voxel, select a point as a reference point among the points belonging to the same group, and mask the three-dimensional coordinates, the first offset coordinate, and the second offset coordinate of all points except the reference point within the same group. The reference point is used to provide position information.
2. The lidar point cloud pre-training method according to claim 1, characterized in that, in step S4, farthest point sampling is used to sample the non-empty voxels according to a preset sampling ratio.
3. The lidar point cloud pre-training method according to claim 1, characterized in that, in step S4, the sampling process is repeatedly executed to obtain voxels for position masking and shape masking respectively.
4. The lidar point cloud pre-training method according to claim 1, characterized in that, in step S8, when restoring the masked information of the masked voxel, the restored position masking includes: inputting the feature data of the position masked voxel into the classification head and outputting the position index of the voxel.
5. The lidar point cloud pre-training method according to claim 4, characterized in that, the cross-entropy loss function is calculated through the output position index and the ground truth to measure the effect of restoring the masked information.
6. The lidar point cloud pre-training method according to claim 1, characterized in that, In the step S8, when restoring the masked information of the restoration mask voxel, the restoration shape mask includes: reconstructing the feature data of the shape mask voxel at the input of the reconstruction head and outputting the first offset coordinates of the points within the voxel.
7. The lidar point cloud pre-training method according to claim 6, wherein, the effect of restoring the mask information is measured by calculating the two-norm chamfer distance using the output first offset coordinates and the true offsets of all points within the voxel relative to the voxel center.
Citation Information
Patent Citations
Three-dimensional point cloud object prediction method and device based on self-attention mechanism
CN114663619A
Voxel-based feature learning network
US10970518B1