Knowledge distillation-based semantic segmentation method for automatic driving four-dimensional millimeter wave radar point cloud
The knowledge distillation method enhances four-dimensional millimeter wave radar point cloud segmentation by transferring knowledge from laser radar networks, improving segmentation accuracy and addressing sparse data density and noise sensitivity issues.
Patent Information
- Application Number
- CN202510378269.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-28
- Publication Date
- 2025-07-15
AI Technical Summary
The existing four-dimensional millimeter-wave radar point cloud semantic segmentation methods lack effective research, and there are problems such as high sparsity, lack of structural information and noise sensitivity, making it difficult to effectively classify obstacles in bad weather conditions.
Using a knowledge distillation-based method, the four-dimensional millimeter-wave radar point cloud is combined with the lidar point cloud. By building a knowledge distillation network, the frozen lidar point cloud is trained to improve its semantic segmentation performance.
It significantly improves the semantic segmentation performance of the four-dimensional millimeter-wave radar point cloud, can effectively classify obstacles in severe weather conditions, and provides reliable environmental perception and path planning information for autonomous driving.
Smart Images

Figure CN120318509A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of autonomous driving, and particularly relates to a semantic segmentation method for four-dimensional millimeter-wave radar point clouds of autonomous driving based on knowledge distillation. Background Art
[0002] Point cloud semantic segmentation refers to assigning a semantic class label to each point in the input point cloud data, such as vehicle, person, traffic sign, etc., which belongs to a classification task. At present, semantic segmentation methods based on images or lidar point clouds have been widely studied and tend to be mature. However, under some harsh lighting or weather conditions, the segmentation based on images or lidar point clouds still faces challenges. On the other hand, millimeter-wave radar can work normally in harsh weather such as rain and snow, and obtain point clouds with Doppler velocity, which makes it a commonly used sensor in environmental perception.
[0003] According to the type of point cloud, millimeter-wave radar can be divided into three-dimensional and four-dimensional. Three-dimensional millimeter-wave radar can generate two-dimensional point clouds, while four-dimensional millimeter-wave radar can generate three-dimensional point clouds similar to three-dimensional lidar. However, compared with lidar, the point clouds generated by four-dimensional millimeter-wave radar are very sparse, and their density is only about one-tenth of that of lidar. In addition, due to the multipath effect of millimeter-wave radar, its data interference is greater. Therefore, in methods such as Rcfusion, Deepfusion, Cramnet, Mvfusion, etc., millimeter-wave radar is used as an auxiliary sensor and fused with other modalities for processing.
[0004] Existing millimeter-wave radar semantic segmentation methods mainly focus on three-dimensional millimeter-wave radar. According to the radar data representation form used, these methods can be divided into tensor processing-based methods and point processing-based methods. Tensor processing-based methods, such as Rss-net, Polarnet, Peakconv, etc., take range-Doppler (RD) tensors or range-azimuth (RA) tensors or range-azimuth-Doppler (RAD) tensors as inputs, and process and segment them into different semantic regions. These methods usually require specific designs and processing flows for the morphology of the tensors to achieve segmentation. Point processing-based methods are more efficient and intuitive. Such methods can draw on mature algorithms in lidar processing without having to design from scratch. For example, using Pointnet as a feature processor to directly process millimeter-wave radar point clouds, drawing on KPConv for point cloud feature extraction, and introducing LSTM to utilize the context information in the radar sequence, etc.
[0005] In the methods introduced above, the semantic segmentation method of the three-dimensional millimeter-wave radar based on tensor processing has defects such as information loss, noise sensitivity, and insufficient flexibility. The point-based method has defects such as high sparsity and lack of structural information. Moreover, the four-dimensional millimeter-wave radar point cloud dataset is lacking, and there is currently no good research method. Summary of the Invention
[0006] To solve the problems and requirements in the background technology, the present invention proposes a semantic segmentation method for the four-dimensional millimeter-wave radar point cloud of autonomous driving based on knowledge distillation, which significantly improves the performance of the four-dimensional millimeter-wave point cloud semantic segmentation network.
[0007] The technical solutions adopted by the present invention include the following steps:
[0008] S1. Use a four-dimensional millimeter-wave radar and a lidar to collect four-dimensional millimeter-wave radar point cloud data and lidar point cloud data in an autonomous driving scenario respectively, and label the four-dimensional millimeter-wave radar point cloud data and lidar point cloud data respectively to obtain a four-dimensional millimeter-wave radar point cloud dataset and a lidar point cloud dataset.
[0009] S2. Build two identical three-dimensional point cloud semantic segmentation networks and knowledge distillation networks in a computer respectively, and use the two three-dimensional point cloud semantic segmentation networks as the lidar point cloud semantic segmentation network and the millimeter-wave radar point cloud semantic segmentation network respectively.
[0010] S3. Input the lidar point cloud dataset into the lidar point cloud semantic segmentation network for training to obtain a trained lidar point cloud semantic segmentation network, and freeze the trained lidar point cloud semantic segmentation network to obtain a frozen lidar point cloud semantic segmentation network.
[0011] The freezing process means that the parameter weights in the trained lidar point cloud semantic segmentation network are no longer updated.
[0012] S4. Use the frozen lidar point cloud semantic segmentation network combined with the knowledge distillation network to train the millimeter-wave radar point cloud semantic segmentation network through the lidar point cloud dataset and the four-dimensional millimeter-wave radar point cloud dataset to obtain a trained millimeter-wave radar point cloud semantic segmentation network.
[0013] S5. Input the to-be-tested four-dimensional millimeter-wave radar point cloud data into the trained millimeter-wave radar point cloud semantic segmentation network for semantic segmentation to obtain the semantic segmentation result of the millimeter-wave radar point cloud.
[0014] The semantic segmentation result of the millimeter-wave radar point cloud obtained is the category and three-dimensional coordinates of the obstacle in autonomous driving. The semantic segmentation result of the millimeter-wave radar point cloud obtained can be further used for environmental perception and obstacle detection, path planning and object motion trajectory prediction in autonomous driving.
[0015] The step S1 is specifically as follows:
[0016] S11. Use four-dimensional millimeter-wave radar and lidar to respectively collect single-frame four-dimensional millimeter-wave radar point cloud data and corresponding single-frame lidar point cloud data in the same scene of autonomous driving.
[0017] S12. Annotate the single-frame four-dimensional millimeter-wave radar point cloud data and the corresponding single-frame laser radar point cloud data respectively to obtain the annotated single-frame four-dimensional millimeter-wave radar point cloud data and the corresponding annotated single-frame laser radar point cloud data.
[0018] S13. Use the same method as step S1-step S2 to continuously collect data in an autonomous driving scenario to obtain several frames of annotated four-dimensional millimeter-wave radar point cloud data and several frames of corresponding annotated lidar point cloud data.
[0019] S14. The annotated four-dimensional millimeter-wave radar point cloud data of all frames are summarized to obtain a four-dimensional millimeter-wave radar point cloud data set, and the annotated lidar point cloud data of all frames are summarized to obtain a corresponding lidar point cloud data set.
[0020] In step S2:
[0021] Each of the three-dimensional point cloud semantic segmentation networks adopts a Cylinder3D network, which includes a voxelization layer, eight three-dimensional sparse convolutional layers and a semantic segmentation detection head connected in series. The input end of the voxelization layer serves as the input end of the three-dimensional point cloud semantic segmentation network, and the output end of the semantic segmentation detection head serves as the output end of the three-dimensional point cloud semantic segmentation network.
[0022] The knowledge distillation network includes a laser feature aggregation module, a millimeter wave feature aggregation module, a millimeter wave feature adaptation module and a knowledge distillation loss module; the laser feature aggregation module and the millimeter wave feature aggregation module serve as the input end of the knowledge distillation network at the same time, the output end of the millimeter wave feature aggregation module is connected to the input end of the millimeter wave feature adaptation module, the output end of the laser feature aggregation module and the output end of the millimeter wave feature adaptation module are both connected to the input end of the knowledge distillation loss module, and the output end of the knowledge distillation loss module serves as the output end of the knowledge distillation network.
[0023] The Cylinder3D network is a conventional three-dimensional point cloud semantic segmentation network, and the voxelization layer, eight three-dimensional sparse convolutional layers (SPConv convolutional layers), and the semantic segmentation detection head are all conventional modules of the Cylinder3D network.
[0024] The millimeter-wave feature adaptive module uses Conv convolution.
[0025] The laser feature aggregation module and the millimeter-wave feature aggregation module are set according to the following formula:
[0026]
[0027] Among them, represents the output result of the laser feature aggregation module, represents the output result of the millimeter-wave feature aggregation module, scatter_max() represents the scatter_max() function, represents the input of the laser feature aggregation module, represents the input of the millimeter-wave feature aggregation module.
[0028] The knowledge distillation loss module is set according to the following formula:
[0029]
[0030] Among them, L Fkd represents the value of the knowledge distillation loss function output by the knowledge distillation loss module, K represents the number of layers using the knowledge distillation module, represents the value of the knowledge distillation loss function obtained in the l-th layer, ‖*‖ represents the two-norm of the vector, represents the output result of the laser feature aggregation module, represents the output result of the millimeter-wave feature adaptive module, represents the output result of the millimeter-wave feature aggregation module, λ l represents the proportionality coefficient of the l-th layer, and l represents the layer number.
[0031] The specific step S4 is as follows:
[0032] S41. Input the lidar point cloud data into the voxelization layer of the frozen lidar point cloud semantic segmentation network to obtain the voxelized lidar point cloud feature map, and then input the voxelized lidar point cloud feature map into eight cascaded three-dimensional sparse convolutional layers in the frozen lidar point cloud semantic segmentation network for processing, and each three-dimensional sparse convolutional layer outputs a lidar point cloud sparse convolutional feature map.
[0033] S42. Input the four-dimensional millimeter-wave radar point cloud data into the voxelization layer of the millimeter-wave radar point cloud semantic segmentation network to obtain a voxelized millimeter-wave point cloud feature map. The voxelized millimeter-wave point cloud feature map is then successively input into eight cascaded three-dimensional sparse convolutional layers and a semantic segmentation detection head in the millimeter-wave radar point cloud semantic segmentation network for processing to obtain a millimeter-wave point cloud semantic segmentation result and a millimeter-wave point cloud loss function value. Each three-dimensional sparse convolutional layer also outputs a millimeter-wave point cloud sparse convolutional feature map.
[0034] S43. Simultaneously input the lidar point cloud sparse convolutional feature map and the millimeter-wave point cloud sparse convolutional feature map obtained from the two three-dimensional sparse convolutional layers where each layer of the lidar point cloud semantic segmentation network and the millimeter-wave radar point cloud semantic segmentation network is located into the knowledge distillation module for processing, and finally obtain a knowledge distillation loss function value.
[0035] S44. Sum the obtained knowledge distillation loss function value and the millimeter-wave point cloud loss function value as the loss function value of the millimeter-wave radar point cloud semantic segmentation network at the current training step, and use the millimeter-wave radar point cloud semantic segmentation result obtained in step S42 as the semantic segmentation result of the millimeter-wave radar point cloud semantic segmentation network at the current training step.
[0036] S45. Perform the next training according to the same method as steps S41 - S44 based on the loss function value of the millimeter-wave radar point cloud semantic segmentation network obtained in step S44 until the training of the millimeter-wave radar point cloud semantic segmentation network is completed, and obtain a trained millimeter-wave radar point cloud semantic segmentation network.
[0037] The specific implementation of step S43 is as follows: Simultaneously input the lidar point cloud sparse convolutional feature map and the millimeter-wave point cloud sparse convolutional feature map obtained from the two three-dimensional sparse convolutional layers where each layer of the lidar point cloud semantic segmentation network and the millimeter-wave radar point cloud semantic segmentation network is located into the lidar feature aggregation module and the millimeter-wave feature aggregation module for processing respectively. The result obtained by the millimeter-wave feature aggregation module is then input into the feature adaptation module for processing. The result obtained by the lidar feature aggregation module and the result obtained by the feature adaptation module are input into the knowledge distillation loss module together for processing to obtain the knowledge distillation loss function value of the current layer. The eight processed knowledge distillation loss function values are then processed together to obtain the final knowledge distillation loss function value.
[0038] The beneficial effects of the present invention are as follows:
[0039] 1. The present invention utilizes the knowledge distillation network to absorb the relatively dense point cloud distribution characteristics and three-dimensional structure information of the three-dimensional lidar point cloud, thereby further improving the performance of the four-dimensional millimeter-wave radar point cloud semantic segmentation network.
[0040] 2. This paper explores the feasibility of semantic segmentation of four-dimensional millimeter-wave radar point cloud, laying the foundation for further expanding the semantic segmentation method of four-dimensional millimeter-wave radar. BRIEF DESCRIPTION OF THE DRAWINGS
[0041] Figure 1 It is a schematic diagram of the process of the present invention;
[0042] Figure 2 This is a comparison chart of the prediction errors of the present invention and the Cylinder3D backbone network. DETAILED DESCRIPTION
[0043] The present invention is further described in detail below in conjunction with the accompanying drawings and specific embodiments.
[0044] like Figure 1 As shown, the present invention is implemented according to the following steps:
[0045] S1. Use OCULii four-dimensional millimeter-wave radar and Avia lidar to collect four-dimensional millimeter-wave radar point cloud data and lidar point cloud data in autonomous driving scenarios respectively, and annotate the four-dimensional millimeter-wave radar point cloud data and lidar point cloud data respectively to obtain four-dimensional millimeter-wave radar point cloud dataset and lidar point cloud dataset.
[0046] S11. Use OCULii four-dimensional millimeter-wave radar and Avia lidar to collect single-frame four-dimensional millimeter-wave radar point cloud data and corresponding single-frame lidar point cloud data in the same scene of autonomous driving.
[0047] S12. Annotate the single-frame four-dimensional millimeter-wave radar point cloud data and the corresponding single-frame laser radar point cloud data respectively to obtain the annotated single-frame four-dimensional millimeter-wave radar point cloud data and the corresponding annotated single-frame laser radar point cloud data.
[0048] S13. Use the same method as step S1-step S2 to continuously collect data in an autonomous driving scenario to obtain several frames of annotated four-dimensional millimeter-wave radar point cloud data and several frames of corresponding annotated lidar point cloud data.
[0049] S14. The annotated four-dimensional millimeter-wave radar point cloud data of all frames are summarized to obtain a four-dimensional millimeter-wave radar point cloud data set, and the annotated lidar point cloud data of all frames are summarized to obtain a corresponding lidar point cloud data set.
[0050] Since there is no publicly available four-dimensional millimeter-wave radar semantic segmentation dataset at present, in this embodiment, an OCULii four-dimensional millimeter-wave radar and an Avia lidar are used to collect 19,000 frames of four-dimensional millimeter-wave radar point cloud data and the corresponding 19,000 frames of lidar point cloud data respectively in the manner of steps S11 - S14. The four-dimensional millimeter-wave radar point cloud data and the corresponding lidar point cloud data are respectively labeled to obtain a four-dimensional millimeter-wave radar point cloud dataset and a lidar point cloud dataset.
[0051] A total of 9 valid categories are labeled in the two datasets, namely buildings, fences, vegetation, cars, bicycles, pedestrians, trucks, buses, and tricycles. In the experiment, 1 NVIDIA 4090 graphics card is used, and the batch size is set to 4 during training and 1 during inference.
[0052] S2. Respectively construct two identical three-dimensional point cloud semantic segmentation networks and knowledge distillation networks in the computer, and use the two three-dimensional point cloud semantic segmentation networks as the lidar point cloud semantic segmentation network and the millimeter-wave radar point cloud semantic segmentation network respectively.
[0053] Each three-dimensional point cloud semantic segmentation network adopts the Cylinder3D network. The Cylinder3D network includes a voxelization layer, eight three-dimensional sparse convolution layers (SPConv convolution layers) connected in series, and a semantic segmentation detection head. The input end of the voxelization layer is used as the input end of the three-dimensional point cloud semantic segmentation network, and the output end of the semantic segmentation detection head is used as the output end of the three-dimensional point cloud semantic segmentation network.
[0054] The knowledge distillation network includes a lidar feature aggregation module, a millimeter-wave feature aggregation module, a millimeter-wave feature adaptation module, and a knowledge distillation loss module; the lidar feature aggregation module and the millimeter-wave feature aggregation module are simultaneously used as the input ends of the knowledge distillation network. The output end of the millimeter-wave feature aggregation module is connected to the input end of the millimeter-wave feature adaptation module. The output ends of the lidar feature aggregation module and the millimeter-wave feature adaptation module are both connected to the input end of the knowledge distillation loss module. The output end of the knowledge distillation loss module is used as the output end of the knowledge distillation network.
[0055] Both the lidar feature aggregation module and the millimeter-wave feature aggregation module adopt the scatter_max function module in Pytorch.
[0056] The millimeter-wave feature adaptation module adopts Conv convolution in Pytorch.
[0057] The lidar feature aggregation module and the millimeter-wave feature aggregation module are set according to the following formula:
[0058]
[0059] Among them, represents the output result of the laser feature aggregation module, represents the output result of the millimeter-wave feature aggregation module, and scatter_max() represents the scatter_max() function in Pytorch, represents the input of the laser feature aggregation module, represents the input of the millimeter-wave feature aggregation module.
[0060] The knowledge distillation loss module is set according to the following formula:
[0061]
[0062] Among them, L Fkd represents the value of the knowledge distillation loss function output by the knowledge distillation loss module, K represents the number of layers using the knowledge distillation module, and in this embodiment, K = 8, represents the value of the knowledge distillation loss function obtained in the l-th layer, ‖*‖ represents the two-norm of the vector, represents the output result of the laser feature aggregation module, represents the output result of the millimeter-wave feature adaptive module, represents the output result of the millimeter-wave feature aggregation module, λ l represents the proportionality coefficient of the l-th layer, and l represents the layer number.
[0063] The Cylinder3D network is a conventional three-dimensional point cloud semantic segmentation network. The voxelization layer, eight three-dimensional sparse convolution layers (SPConv convolution layers), and the semantic segmentation detection head are all conventional modules of the Cylinder3D network.
[0064] The loss function of the Cylinder3D network is composed of the Lovasz loss of semantic segmentation and the cross-entropy loss of semantic segmentation added together.
[0065] The cross-entropy loss function L ce is as follows:
[0066]
[0067] Among them, L ce represents the cross-entropy loss function, represents the probability prediction value that the i-th point belongs to the c-th class, y ic represents the probability true value that the i-th point belongs to the c-th class, c represents the category, H is the total number of categories, and N is the number of points in the current frame.
[0068] The Lovasz loss L Lovasz is as follows:
[0069]
[0070] L seg = L ce + L Lovasz
[0071] Among them, L seg represents the loss function of the lidar point cloud semantic segmentation network, TP c represents the total number of points correctly classified as the c-th class in the current frame, FP c represents the total number of points misclassified as the c-th class in the current frame, FN c represents the total number of points misclassified as non-c class in the current frame, IoU c represents the quotient of the intersection and union of the predicted value and the true value of the c-th class, IoU c The closer the IoU value is to 1, the better the prediction result of this class, m ic represents the error, L Lovasz represents the Lovasz loss, represents the predicted class of the i-th point, c represents the class, y i represents the true class of the i-th point, i represents the index, N represents the number of points in the current frame, p ic is the probability prediction value that the i-th point belongs to the c-th class, represents sorting the errors in descending order, taking the partial derivative.
[0072] S3. Input the lidar point cloud data set into the lidar point cloud semantic segmentation network for training to obtain a trained lidar point cloud semantic segmentation network, and freeze the trained lidar point cloud semantic segmentation network to obtain a frozen lidar point cloud semantic segmentation network.
[0073] The freezing process means that the parameter weights in the trained lidar point cloud semantic segmentation network are no longer updated.
[0074] S4. Use the frozen lidar point cloud semantic segmentation network combined with the knowledge distillation network to train the millimeter-wave radar point cloud semantic segmentation network through the lidar point cloud data set and the four-dimensional millimeter-wave radar point cloud data set to obtain a trained millimeter-wave radar point cloud semantic segmentation network.
[0075] S41. Input the lidar point cloud data into the voxelization layer of the frozen lidar point cloud semantic segmentation network to obtain a voxelized lidar point cloud feature map, and then input the voxelized lidar point cloud feature map into eight cascaded three-dimensional sparse convolutional layers in the frozen lidar point cloud semantic segmentation network for processing, and each three-dimensional sparse convolutional layer outputs a lidar point cloud sparse convolutional feature map.
[0076] S42. Input the four-dimensional millimeter-wave radar point cloud data into the voxelization layer of the millimeter-wave radar point cloud semantic segmentation network to obtain a voxelized millimeter-wave point cloud feature map. Then, the voxelized millimeter-wave point cloud feature map is successively input into eight cascaded three-dimensional sparse convolutional layers and a semantic segmentation detection head in the millimeter-wave radar point cloud semantic segmentation network for processing to obtain a millimeter-wave point cloud semantic segmentation result and a millimeter-wave point cloud loss function value. Each three-dimensional sparse convolutional layer also outputs a millimeter-wave point cloud sparse convolutional feature map.
[0077] S43. Simultaneously input the lidar point cloud sparse convolutional feature map and the millimeter-wave point cloud sparse convolutional feature map obtained from the two three-dimensional sparse convolutional layers where each layer of the lidar point cloud semantic segmentation network and the millimeter-wave radar point cloud semantic segmentation network is located into the knowledge distillation module for processing, and finally obtain a knowledge distillation loss function value.
[0078] Step S43 is specifically as follows: Simultaneously input the lidar point cloud sparse convolutional feature map and the millimeter-wave point cloud sparse convolutional feature map obtained from the two three-dimensional sparse convolutional layers where each layer of the lidar point cloud semantic segmentation network and the millimeter-wave radar point cloud semantic segmentation network is located into the lidar feature aggregation module and the millimeter-wave feature aggregation module respectively for processing. The result obtained by the millimeter-wave feature aggregation module is then input into the feature adaptation module for processing. The result obtained by the lidar feature aggregation module and the result obtained by the feature adaptation module are input into the knowledge distillation loss module together for processing to obtain the knowledge distillation loss function value of the current layer. The eight knowledge distillation loss function values obtained by processing are then processed together to obtain the final knowledge distillation loss function value.
[0079] In specific implementation, represents the input of the lidar feature aggregation module and also represents the l-th layer lidar point cloud sparse convolutional feature map. represents the input of the millimeter-wave feature aggregation module and also represents the l-th layer millimeter-wave point cloud sparse convolutional feature map.
[0080] S44. Sum the obtained knowledge distillation loss function value and the millimeter-wave point cloud loss function value as the loss function value of the millimeter-wave radar point cloud semantic segmentation network at the current training step. Use the millimeter-wave radar point cloud semantic segmentation result obtained in step S42 as the semantic segmentation result of the millimeter-wave radar point cloud semantic segmentation network at the current training step.
[0081] S45. According to the loss function value of the millimeter-wave radar point cloud semantic segmentation network obtained in step S44, perform the next training in the same method as steps S41 - S44 until the training of the millimeter-wave radar point cloud semantic segmentation network is completed, and obtain a trained millimeter-wave radar point cloud semantic segmentation network.
[0082] In specific implementation, the millimeter-wave point cloud loss function value obtained in step S42 is Lce +L Lovasz In step S44, the loss function value of the millimeter-wave radar point cloud semantic segmentation network obtained by summing the knowledge distillation loss function value and the millimeter-wave point cloud loss function value is
[0083] S5. Input the measured four-dimensional millimeter-wave radar point cloud data into the trained millimeter-wave radar point cloud semantic segmentation network for semantic segmentation to obtain the semantic segmentation result of the millimeter-wave radar point cloud.
[0084] The obtained semantic segmentation result of the millimeter-wave radar point cloud is the information such as the category and three-dimensional coordinates of the obstacles in autonomous driving. The obtained semantic segmentation result of the millimeter-wave radar point cloud can further be used for environmental perception, obstacle detection, path planning, and prediction of object motion trajectories in autonomous driving.
[0085] In specific implementation, the batch size of the two Cylinder3D networks is set to 4, the Adam optimizer is used, the learning rate is set to 0.001, the periodic cosine annealing method is used to control the change of the learning rate, the initial period length is 1, the change period is 2, and the minimum learning rate is 1×10 -5 , and a total of 130 epochs are trained.
[0086] On the four-dimensional millimeter-wave radar point cloud dataset, the classification effect comparison between the present invention and the Cylinder3D backbone network is shown in Table 1. The indicators in Table 1 are the intersection over union IoU of each category and the mean intersection over union mIoU.
[0087] Table 1 is the classification effect comparison table between the present invention and the Cylinder3D backbone network
[0088]
[0089] The increment in Table 1 represents the improvement of the method of the present invention compared with the Cylinder3D backbone network. The above results show that the present invention effectively improves the performance of the millimeter-wave radar point cloud semantic segmentation network, and there is an obvious improvement in each category. As Figure 2 shown, the comparison diagram of the prediction error results between the present invention and the Cylinder3D backbone network shows that the prediction error of the method of the present invention is lower.
[0090] The present invention realizes using knowledge distillation to learn the rich structural information contained in the three-dimensional lidar point cloud, thereby enhancing the performance of the four-dimensional millimeter-wave radar point cloud semantic segmentation network and being able to improve the semantic segmentation result of the millimeter-wave radar lidar point cloud. For the four-dimensional millimeter-wave radar point cloud backbone network, the feature map features are strengthened through the knowledge distillation network at each feature scale, and finally the enhanced four-dimensional millimeter-wave radar point cloud semantic segmentation output result is obtained.
[0091] The above is only a specific implementation manner of the present invention, and it cannot be used to limit the scope of the present invention. Equivalent changes made by those of ordinary skill in the art according to this creation, as well as changes well-known to those skilled in the art, should still fall within the scope covered by the present invention.
Claims
1. A semantic segmentation method for four-dimensional millimeter-wave radar point clouds of autonomous driving based on knowledge distillation, characterized in that The following steps are involved: S1. Use four-dimensional millimeter-wave radar and laser radar to collect four-dimensional millimeter-wave radar point cloud data and laser radar point cloud data in the autonomous driving scenario respectively, and annotate the four-dimensional millimeter-wave radar point cloud data and laser radar point cloud data respectively to obtain four-dimensional millimeter-wave radar point cloud data set and laser radar point cloud data set; S2. Build two identical 3D point cloud semantic segmentation networks and knowledge distillation networks respectively, and use the two 3D point cloud semantic segmentation networks as the LiDAR point cloud semantic segmentation network and the millimeter wave radar point cloud semantic segmentation network respectively; S3, inputting the laser radar point cloud data set into the laser radar point cloud semantic segmentation network for training to obtain a trained laser radar point cloud semantic segmentation network, and freezing the trained laser radar point cloud semantic segmentation network to obtain a frozen laser radar point cloud semantic segmentation network; S4, using the frozen LiDAR point cloud semantic segmentation network combined with the knowledge distillation network to train the millimeter wave radar point cloud semantic segmentation network through the LiDAR point cloud dataset and the four-dimensional millimeter wave radar point cloud dataset to obtain a trained millimeter wave radar point cloud semantic segmentation network; S5. Input the four-dimensional millimeter-wave radar point cloud data to be tested into the trained millimeter-wave radar point cloud semantic segmentation network for semantic segmentation to obtain the semantic segmentation result of the millimeter-wave radar point cloud.
2. A semantic segmentation method for four-dimensional millimeter-wave radar point cloud of autonomous driving based on knowledge distillation according to claim 1, wherein The step S1 is specifically as follows: S11. Use a four-dimensional millimeter-wave radar and a laser radar to collect single-frame four-dimensional millimeter-wave radar point cloud data and corresponding single-frame laser radar point cloud data in the same scene of autonomous driving respectively; S12, respectively annotating the single-frame four-dimensional millimeter-wave radar point cloud data and the corresponding single-frame laser radar point cloud data to obtain the annotated single-frame four-dimensional millimeter-wave radar point cloud data and the corresponding annotated single-frame laser radar point cloud data; S13, using the same method as step S1-step S2 to continuously collect data in an autonomous driving scenario, and obtain several frames of annotated four-dimensional millimeter-wave radar point cloud data and several frames of corresponding annotated lidar point cloud data; S14. The annotated four-dimensional millimeter-wave radar point cloud data of all frames are summarized to obtain a four-dimensional millimeter-wave radar point cloud data set, and the annotated lidar point cloud data of all frames are summarized to obtain a corresponding lidar point cloud data set.
3. A semantic segmentation method for four-dimensional millimeter-wave radar point clouds of autonomous driving based on knowledge distillation according to claim 1, characterized in that In step S2: Each of the three-dimensional point cloud semantic segmentation networks adopts a Cylinder3D network, which includes a voxelization layer, eight three-dimensional sparse convolutional layers and a semantic segmentation detection head connected in series, wherein the input end of the voxelization layer serves as the input end of the three-dimensional point cloud semantic segmentation network, and the output end of the semantic segmentation detection head serves as the output end of the three-dimensional point cloud semantic segmentation network; The knowledge distillation network includes a laser feature aggregation module, a millimeter wave feature aggregation module, a millimeter wave feature adaptation module and a knowledge distillation loss module; The laser feature aggregation module and the millimeter-wave feature aggregation module simultaneously serve as the input ends of the knowledge distillation network. The output end of the millimeter-wave feature aggregation module is connected to the input end of the millimeter-wave feature adaptive module. The output ends of the laser feature aggregation module and the millimeter-wave feature adaptive module are both connected to the input end of the knowledge distillation loss module. The output end of the knowledge distillation loss module serves as the output end of the knowledge distillation network.
4. A semantic segmentation method for four-dimensional millimeter-wave radar point cloud of autonomous driving based on knowledge distillation according to claim 3, characterized in that: The millimeter-wave feature adaptive module uses Conv convolution.
5. A semantic segmentation method for four-dimensional millimeter-wave radar point cloud of autonomous driving based on knowledge distillation according to claim 3, characterized in that: The laser feature aggregation module and the millimeter-wave feature aggregation module are set according to the following formula: Among them, represents the output result of the laser feature aggregation module, represents the output result of the millimeter-wave feature aggregation module, and scatter_max() represents the scatter_max() function. represents the input of the laser feature aggregation module, represents the input of the millimeter-wave feature aggregation module.
6. A semantic segmentation method for four-dimensional millimeter-wave radar point cloud of autonomous driving based on knowledge distillation according to claim 3, characterized in that: The knowledge distillation loss module is set according to the following formula: Among them, L Fkd represents the value of the knowledge distillation loss function output by the knowledge distillation loss module, K represents the number of layers using the knowledge distillation module, represents the value of the knowledge distillation loss function obtained in the l-th layer, ‖*‖ represents the two-norm of the vector, represents the output result of the laser feature aggregation module, represents the output result of the millimeter-wave feature adaptation module, represents the output result of the millimeter-wave feature aggregation module, λ l represents the proportionality coefficient of the l-th layer, and l represents the layer number where it is located.
7. A semantic segmentation method for four-dimensional millimeter-wave radar point cloud of autonomous driving based on knowledge distillation according to claim 1, characterized in that The specific steps of S4 are as follows: S41. Input the lidar point cloud data into the voxelization layer of the frozen lidar point cloud semantic segmentation network to obtain a voxelized lidar point cloud feature map. The voxelized lidar point cloud feature map is then sequentially input into eight cascaded three-dimensional sparse convolutional layers in the frozen lidar point cloud semantic segmentation network for processing, and each three-dimensional sparse convolutional layer outputs a lidar point cloud sparse convolutional feature map. S42. Input the four-dimensional millimeter-wave radar point cloud data into the voxelization layer of the millimeter-wave radar point cloud semantic segmentation network to obtain a voxelized millimeter-wave point cloud feature map. The voxelized millimeter-wave point cloud feature map is then sequentially input into eight cascaded three-dimensional sparse convolutional layers and a semantic segmentation detection head in the millimeter-wave radar point cloud semantic segmentation network for processing to obtain a millimeter-wave point cloud semantic segmentation result and a millimeter-wave point cloud loss function value. Each three-dimensional sparse convolutional layer also outputs a millimeter-wave point cloud sparse convolutional feature map. S43. Simultaneously input the lidar point cloud sparse convolutional feature map and the millimeter-wave point cloud sparse convolutional feature map obtained from the two three-dimensional sparse convolutional layers where each layer of the lidar point cloud semantic segmentation network and the millimeter-wave radar point cloud semantic segmentation network is located into the knowledge distillation module for processing, and finally obtain a knowledge distillation loss function value. S44. The sum of the obtained knowledge distillation loss function value and the millimeter-wave point cloud loss function value is used as the loss function value of the millimeter-wave radar point cloud semantic segmentation network at the current training step. The millimeter-wave radar point cloud semantic segmentation result obtained in step S42 is used as the semantic segmentation result of the millimeter-wave radar point cloud semantic segmentation network at the current training step. S45. Perform the next training according to the same method as steps S41 - S44 based on the loss function value of the millimeter-wave radar point cloud semantic segmentation network obtained in step S44 until the training of the millimeter-wave radar point cloud semantic segmentation network is completed, and obtain a trained millimeter-wave radar point cloud semantic segmentation network.
8. A semantic segmentation method for four-dimensional millimeter-wave radar point clouds of autonomous driving based on knowledge distillation according to claim 7, characterized in that, The specific steps of S43 are as follows: The sparse convolution feature maps of lidar point clouds and millimeter-wave radar point clouds obtained from the two three-dimensional sparse convolution layers in each layer of the lidar point cloud semantic segmentation network and the millimeter-wave radar point cloud semantic segmentation network are simultaneously input into the lidar feature aggregation module and the millimeter-wave feature aggregation module for processing. The result obtained by the millimeter-wave feature aggregation module is then input into the feature adaptive module for processing. The result obtained by the lidar feature aggregation module and the result obtained by the feature adaptive module are input into the knowledge distillation loss module together for processing to obtain the knowledge distillation loss function value of the current layer. The eight knowledge distillation loss function values obtained by processing are then processed together to obtain the final knowledge distillation loss function value.
Citation Information
Cited By
Semantic segmentation method and system for point cloud encoder
CN122223341A