A radar panoptic segmentation method based on decoupled dynamic convolution kernel

By introducing a decoupled dynamic convolution kernel in radar panoramic segmentation, and uniformly processing semantics and instance segmentation, the problem of inefficiency of existing methods is solved, and a high-performance panoramic segmentation effect is achieved.

CN116704504BActive Publication Date: 2025-05-13ZHEJIANG UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310496335.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-05-05
Publication Date
2025-05-13
Estimated Expiration
2043-05-05

AI Technical Summary

Technical Problem

Existing radar panoramic segmentation methods explicitly separate semantic and instance segmentation tasks, resulting in inefficiency and poor performance.

Method used

A radar panoramic segmentation method based on decoupling dynamic convolution kernel is proposed. Semantic and instance segmentation are uniformly processed through dynamic kernel generator and decoder, and the dynamic kernel and point cloud feature convolution prediction mask is used.

Benefits of technology

The most advanced panoramic segmentation performance was achieved on the SemanticKITTI dataset, with PQ reaching 59.0% and semantic index mIoU reaching 67.7%, significantly improving the efficiency and accuracy of radar point cloud panoramic segmentation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116704504B_ABST
    Figure CN116704504B_ABST
Patent Text Reader

Abstract

The present invention relates to the field of 3D vision technology, and proposes a radar panoramic segmentation method based on decoupled dynamic convolution kernels, which mainly includes five stages of step implementation: constructing a point cloud feature extractor, constructing a dynamic convolution kernel generation module, constructing a dynamic convolution kernel decoding module, model training and model inference. The present invention proposes a new Panoptic DKNet network model to realize radar point cloud panoramic segmentation in a unified workflow with decoupled dynamic convolution kernels. Panoptic DKNet decouples the dynamic kernels of instance targets and background categories to promote their respective learning processes, and implements a decoupling strategy for classification and segmentation to avoid mutual competition between different categories. The algorithm designed by the present invention shows good panoramic segmentation performance on the SemanticKITTI benchmark dataset, and has good segmentation robustness for small target point clouds and scattered large target point clouds that are close to each other in the environment.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of 3D vision technology, and in particular to a radar panoramic segmentation method based on a decoupled dynamic convolution kernel. Background Art

[0002] In recent years, with the rapid development of autonomous driving, LiDAR perception technology has been widely studied. LiDAR segmentation can achieve point-level prediction of the entire scene, which is fundamental and critical in radar perception tasks. LiDAR panoramic segmentation can not only predict point-by-point semantic labels of background categories (such as roads and vegetation), but also predict semantic labels and instance IDs of foreground categories (such as cars and people). Due to its ability to achieve both semantic and instance segmentation in one network architecture, panoramic segmentation plays an important role in LiDAR perception and has broad application prospects.

[0003] According to the implementation method of instance segmentation, existing radar panoramic segmentation methods can be divided into two categories: top-down and bottom-up algorithms. Top-down methods adopt the detection-segmentation paradigm to extract foreground points in the bounding box based on semantic prediction to achieve instance segmentation. Bottom-up methods follow the regression-clustering paradigm, first predicting the semantic labels of all points, and then performing heuristic clustering to aggregate foreground points based on point cost or feature embedding. However, both methods explicitly separate the two segmentation tasks and use two independent branches to achieve panoramic segmentation. Summary of the invention

[0004] In view of the above problems, the present invention proposes a unified radar panoramic segmentation network model Panoptic DKNet, which mainly includes a dynamic kernel generator and a dynamic kernel decoder, which are responsible for generating dynamic kernels with semantic categories and decoding predicted segmentation masks respectively. The dynamic kernel generator predicts the center point position of the instance target and the area of ​​the background category from the perspective of the bird's eye view (BEV), and then extracts the BEV features at the corresponding positions to generate dynamic kernel weights of the instance and background classes. The dynamic kernel decoder includes a position decoder for capturing the spatial position information of the instance object and outputting a position-aware instance dynamic kernel; and a mask decoder for establishing correlations between dynamic kernels and aggregating global context features. Finally, the segmentation mask corresponding to each dynamic kernel is generated by convolving the weight of the dynamic kernel with the feature embedding of each point. The instance / background dynamic convolution kernel proposed in the present invention realizes the panoramic segmentation of radar point clouds in a unified process.

[0005] In order to achieve the above object, the present invention provides a radar panoramic segmentation method based on a decoupled dynamic convolution kernel, comprising the following steps:

[0006] S1. Build a point cloud feature extractor to extract voxel features and point-level feature embedding of the input point cloud based on voxel representation and 3D sparse convolution;

[0007] S2. Construct a dynamic convolution kernel generation module to generate the initial weights of the dynamic convolution kernel based on the bird's-eye view heat map and the predicted instance and background point cloud positions from the BEV perspective;

[0008] S3. Construct a dynamic convolution kernel decoding module, use the instance dynamic kernel to predict the 3D bounding box of the instance target, use kNN-Transformer to weight the dynamic kernel to fuse the point cloud features, and finally predict the mask by convolving the dynamic kernel with the point cloud features to output the scene point cloud panoramic segmentation result;

[0009] S4. Using the server, the network parameters are optimized by reducing the network loss function until the network converges, thereby obtaining a radar panoramic segmentation method based on a decoupled dynamic convolution kernel;

[0010] S5. Perform panoramic segmentation prediction on the new point cloud using the radar panoramic segmentation method based on the decoupled dynamic convolution kernel.

[0011] Preferably, the step S1 specifically includes the following steps:

[0012] S11. Input sparse point cloud The point cloud is converted into voxels using voxelization operation, and then the initialization features of each voxel are extracted;

[0013] S12. Construct a 3D sparse feature encoder and a multi-scale global attention module to extract sparse voxel feature expressions At the same time, we can get point-level feature embedding

[0014] Preferably, the specific process of the voxelization operation is: first set the voxel resolution s, and then for a given point p in the point cloud P i =(x i ,y i , z i ), the index of the voxel to which it belongs is The initial feature extraction method for each voxel is as follows: the coordinates of all points in a voxel are sent to the multi-layer perceptron MLP to extract the features of the points, and then the features of multiple points are fused using the Max-Pooling function to obtain the feature expression f of a voxel. v ;

[0015] The extraction process of the sparse voxel feature expression and point-level feature embedding is as follows: first, a 3D sparse convolution with a Bottleneck structure is used to extract the local information of the voxel, and then a cross-scale global attention module is used to build a long-distance dependency relationship and establish the correlation between voxels. A four-layer 3D sparse feature encoder is constructed, and the feature of the last encoder is used as the sparse voxel feature F v , whose dense spatial resolution is LxHxW, and the point-level feature embedding is:

[0016] fe=Concat(f v,1 ,f v,2 ,f v,3 ,f v,4 ,MLP(f p ))

[0017] where f v,j is the i-th layer encoding feature of the voxel where the point is located, f p are the spatial coordinates of the points, and the feature embedding of each point contains the global encoding information and its unique representation.

[0018] Preferably, the step S2 specifically includes the following steps:

[0019] S21, the sparse voxel feature F v Max-Pooling is performed on the Z axis to obtain the BEV feature Then, based on the BEV feature, the instance center point heat map is predicted. and background area map

[0020] S22, based on the instance center point heat map M th , Background area map M st , and BEV characteristic F bev , extract BEV features at the corresponding positions of high response in the heat map to obtain the initial instance dynamic kernel and background class dynamic kernel

[0021] Preferably, the BEV feature-based prediction instance center point heat map in step S21 The specific method is: in F bev Multi-layer 2D convolution is used to predict the center point heat map, where N ins is the number of semantic categories of the instance, and each channel represents the score of a category;

[0022] The background area map predicted based on the BEV feature in step S21 The specific method is: in F bevA small U-Net structure is used to predict the 2D background area segmentation result under the BEV perspective, where N st is the number of semantic categories of the background class;

[0023] The heat map M based on the instance center point in step S22 th and BEV characteristic F bev Generate instance dynamic core The specific method is: select M th The highest score N th positions, which indicate potential instance objects; then extract F bev The features at the corresponding position are used to generate the dynamic kernel of each instance object;

[0024] The background area map M in step S22 st and BEV characteristic F bev Generate background class dynamic core The specific method is: multiply the background area map with the BEV feature to obtain the BEV feature of the background Then, an adaptive average pooling operation is performed in the spatial dimension to obtain the dynamic kernel for each background category.

[0025] Preferably, the step S3 specifically includes the following steps:

[0026] S31, based on the instance dynamic kernel K th , using the Transformer module and multi-layer perceptron, predict the 3D bounding box and overlap intersection over union (IoU) of each dynamic kernel corresponding to the instance target; then use the non-maximum suppression method to eliminate redundant bounding boxes and obtain a streamlined bounding box set At the same time, the corresponding redundant instance dynamic core is also changed from Eliminate them to obtain a streamlined set of instance dynamic cores Where N' th <=N th is the number of predicted instances in the scene;

[0027] S32. Build kNN-Transformer, query feature is instance dynamic kernel K' th and background dynamic kernel K st , the key / value feature is the feature of the k nearest neighbor points of each dynamic kernel in the spatial position, and the attention mechanism is used to weightedly fuse the adjacent features to enhance the feature expression of the dynamic kernel;

[0028] S33, using the instance dynamic kernel K' thAnd the corresponding 3D bounding box B predicts instance segmentation: the bounding box is expanded by a certain proportion to obtain the search area of ​​the target object, the feature embedding of the points in the search area is subjected to dot product with the dynamic kernel and the sigmoid activation function to predict the mask score, and the points in the same search area with a mask score greater than 0.5 are assigned the same instance ID; for background class points, their feature embeddings are convolved with the background class dynamic convolution kernel, and then the argmax function is used to assign the semantic category of the dynamic kernel to each background point.

[0029] Preferably, the calculation process of the features of the k neighboring points in step S32 is specifically as follows: the dynamic kernel K' of each instance th The spatial position of is set to the center of its corresponding 3D bounding box, and the k points closest to the spatial position are indexed as the focus point; for the background class dynamic kernel, calculate The cosine similarity between each pixel feature and the dynamic kernel is calculated, and then the top-k pixels are taken and projected back into the 3D space to obtain the features of the corresponding points;

[0030] The calculation process of the weighted fusion of adjacent features by the attention mechanism in step S32 is specifically as follows:

[0031]

[0032] in For location-aware instance dynamic kernels and background class dynamic kernels, key / value It is a linear mapping of the point features of the k nearest neighbors.

[0033] Preferably, the step S4 specifically includes the following steps:

[0034] S41, collect the point cloud data in the training data set, use the server to execute the point cloud feature extraction process in step S1, output voxel features and point-level feature embedding; input the voxel features into the BEV bird's-eye view heat map prediction module in step S2, and generate the instance center point heat map M th , Background area map M st , FocalLoss is used to supervise the two heat maps, which can be expressed as:

[0035] L pos =FL(M th ,Y th ) / N th +FL(M st ,Y st ) / N st

[0036] S42, using the server to execute the dynamic convolution kernel decoding module in step S3, using the instance dynamic kernel to predict the 3D bounding box of the instance target, and supervising the predicted bounding box and IoU, which can be specifically expressed as:

[0037] L det =L box +λ IoU ·L IoU

[0038] Where L box is the L1 loss for bounding box regression, L IoU is the SmoothL1 loss of the bounding box IoU;

[0039] S43, using the server to execute the kNN-Transformer in step S3 to weight the fusion point cloud features with the dynamic kernel, and finally predicting the mask by convolution of the dynamic kernel and the point cloud features, outputting the scene point cloud panoramic segmentation result, and supervising the segmentation result, which can be specifically expressed as:

[0040]

[0041] S44, use the server to train the network, the loss function L total is the weighted sum of the losses in step S41, step S42 and step S43:

[0042] L total =L pos +L det +L mask

[0043] S45. Utilize the server to optimize the objective function and obtain local optimal network parameters.

[0044] Preferably, the instance center point heat map M in step S41 th The real label is generated using a Gaussian kernel, mapping the center of an object with semantic class c to the sth channel; the background area map M st The real label It is to map each background class point to the corresponding channel of the BEV map;

[0045] The specific calculation method of the loss of the segmentation result in step S43 is as follows: in the mask decoder training stage, the Hungarian matching algorithm is used to match the predicted instance box and the real label, and the Focal Loss (L fl ) and Dice Loss(L dl ) to supervise the segmentation mask of the instance object, using the cross entropy loss (L ce ) and Lovasz softmax loss (Lls )Supervised background category segmentation mask.

[0046] Preferably, the step S5 specifically includes the following steps:

[0047] S51, obtain the laser radar point cloud in the 3D environment, input it into the trained point cloud feature extractor, extract voxel features and point-level feature embedding;

[0048] S52, inputting the voxel features into the trained dynamic convolution kernel generation module to generate the initial weights of the dynamic convolution kernel of the instance / background class;

[0049] S53, input the dynamic convolution kernel into the trained dynamic convolution kernel decoding module, predict the 3D bounding box of the instance target, decode the dynamic kernel features, and finally predict and output the scene point cloud panoramic segmentation result.

[0050] Compared with the prior art, the present invention has the following beneficial effects:

[0051] The present invention provides a radar panoramic segmentation method based on decoupled dynamic convolution kernel. Based on the instance / background class dynamic convolution kernel decoupling strategy and the way of embedding dynamic kernel and point feature into convolution prediction mask, the proposed model completes the panoramic segmentation of lidar point cloud in a unified pipeline, and achieves the most advanced panoramic segmentation performance on the outdoor autonomous driving scene dataset SemanticKITTI, with the panoramic segmentation index PQ reaching 59.0% and the semantic index mIoU reaching 67.7%. BRIEF DESCRIPTION OF THE DRAWINGS

[0052] Figure 1 The overall algorithm framework diagram of a radar panoramic segmentation method based on a decoupled dynamic convolution kernel provided by the present invention;

[0053] Figure 2 This is a visualization diagram of the algorithm proposed in the present invention. DETAILED DESCRIPTION

[0054] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0055] The present invention aims to achieve semantic and instance segmentation in a unified workflow using decoupled dynamic convolution kernels. Specifically, the present invention generates instance / semantic segmentation masks using a set of learnable convolution kernels, where each instance dynamic kernel is responsible for predicting the mask and semantic category of the instance object, while the background dynamic kernel is used to distinguish the semantic labels of background points. In addition, in order to explicitly decouple classification and segmentation, the present invention assigns semantic labels to the generated dynamic kernels in advance, and these dynamic kernels are further used to predict segmentation masks. In this way, the present invention achieves unified prediction of instance and background categories, while improving the learning process of instance / background dynamic kernels through a decoupling strategy.

[0056] In view of the problems and shortcomings existing in the prior art, the present invention proposes a radar panoramic segmentation method based on a decoupled dynamic convolution kernel, which mainly includes five stages of designing a point cloud feature extractor, designing a dynamic convolution kernel generation module, designing a dynamic convolution kernel decoding module, model training and model inference.

[0057] The present invention proposes a radar panoramic segmentation method based on a decoupled dynamic convolution kernel, comprising the following steps:

[0058] S1. Build a point cloud feature extractor to extract voxel features and point-level feature embedding of the input point cloud based on voxel representation and 3D sparse convolution;

[0059] S2. Construct a dynamic convolution kernel generation module to generate the initial weights of the dynamic convolution kernel based on the bird's-eye view heat map and the predicted instance and background point cloud positions from the BEV perspective;

[0060] S3. Construct a dynamic convolution kernel decoding module, use the instance dynamic kernel to predict the 3D bounding box of the instance target, use kNN-Transformer to weight the dynamic kernel to fuse the point cloud features, and finally predict the mask by convolving the dynamic kernel with the point cloud features to output the scene point cloud panoramic segmentation result;

[0061] S4. Using the server, the network parameters are optimized by reducing the network loss function until the network converges, thereby obtaining a radar panoramic segmentation method based on a decoupled dynamic convolution kernel;

[0062] S5. Perform panoramic segmentation prediction on the new point cloud (simultaneously performing semantic segmentation and instance segmentation) using the radar panoramic segmentation method based on the decoupled dynamic convolution kernel.

[0063] The following is a detailed description of each step.

[0064] Specifically, step S1, construct a point cloud feature extractor, extract voxel features and point-level feature embedding of the input point cloud based on voxel representation and 3D sparse convolution. Figure 1(a) shows the backbone network of the point cloud feature extractor constructed by the present invention, and the specific steps are as follows:

[0065] S11. Input sparse point cloud The point cloud is converted into voxels by voxelization, and then the initialization features of each voxel are extracted; specifically, the specific process of the voxelization operation is: first set the voxel resolution s, and then for a given point p in the point cloud P i =(x i ,y i , z i ), the index of the voxel to which it belongs is The initial feature extraction method for each voxel is as follows: the coordinates of all points in a voxel are sent to the multi-layer perceptron MLP to extract the features of the points, and then the features of multiple points are fused using the Max-Pooling function to obtain the feature expression f of a voxel. v .

[0066] S12. Construct a 3D sparse feature encoder and a multi-scale global attention module to extract sparse voxel feature expressions At the same time, we can get point-level feature embedding Specifically, the extraction process of the sparse voxel feature expression and point-level feature embedding is as follows: for voxel features, a 3D sparse convolution with a Bottleneck structure is first used to extract the local information of the voxel, and then a cross-scale global attention module is used to construct a long-distance dependency relationship and establish the correlation between voxels. The feature extraction backbone network of the present invention includes a four-layer 3D sparse feature encoder, and the feature of the last encoder is used as the sparse voxel feature expression F v , whose dense spatial resolution is LxHxW, and the point-level feature embedding is:

[0067] fe=Concat(f v,1 ,f v,2 ,f v,3 ,f v,4 ,MLP(f p ))

[0068] where f v,j is the i-th layer encoding feature of the voxel where the point is located, f p are the spatial coordinates of the point. Therefore, the feature embedding of each point contains the global encoding information and its unique representation.

[0069] Step S2: construct a dynamic convolution kernel generation module, based on the bird's-eye view heat map, and generate the dynamic convolution kernel initial weights according to the predicted instance and background point cloud positions from the BEV perspective; Figure 1 (b) shows a dynamic convolution kernel generation module constructed by the present invention, and the specific steps are as follows:

[0070] S21, the sparse voxel feature F v Max-Pooling is performed on the Z axis to obtain the BEV feature Then, based on the BEV feature, the instance center point heat map is predicted. and background area map Specifically, the BEV feature-based prediction instance center point heat map in step S21 The specific method is: in F bev Multi-layer 2D convolution is used to predict the center point heat map, where N ins is the number of semantic categories of the instance, and each channel represents the score of a category; the background area map predicted based on the BEV feature in step S21 The specific method is: in F bev A small U-Net structure is used to predict the 2D background area segmentation result under the BEV perspective, where N st is the number of semantic categories of the background class.

[0071] S22, based on the instance center point heat map M th , Background area map M st , and BEV characteristic F bev , extract BEV features at the corresponding positions of high response in the heat map to obtain the initial instance dynamic kernel and background class dynamic kernel Specifically, in step S22, the heat map M based on the instance center point th and BEV characteristic F bev Generate instance dynamic core The specific method is: select M th The highest score N th positions, which indicate potential instance objects; then extract F bev The features at the corresponding position are used to generate the dynamic kernel of each instance object;

[0072] The background area map M in step S22 st and BEV characteristic F bev Generate background class dynamic core The specific method is: multiply the background area map with the BEV feature to obtain the BEV feature of the background Then, an adaptive average pooling operation is performed in the spatial dimension to obtain the dynamic kernel for each background category.

[0073] Step S3: construct a dynamic convolution kernel decoding module, use the instance dynamic kernel to predict the 3D bounding box of the instance target, use kNN-Transformer to weight the dynamic kernel to fuse the point cloud features, and finally predict the mask by convolving the dynamic kernel with the point cloud features to output the scene point cloud panoramic segmentation result. Figure 1 (c) shows the dynamic convolution kernel decoding module constructed by the present invention, and the specific steps are as follows:

[0074] S31, based on the instance dynamic kernel K th , using the Transformer module and multi-layer perceptron, predict the 3D bounding box and overlap intersection over union (IoU) of each dynamic kernel corresponding to the instance target; then use the non-maximum suppression method to eliminate redundant bounding boxes and obtain a streamlined bounding box set At the same time, the corresponding redundant instance dynamic core is also changed from Eliminate them to obtain a streamlined set of instance dynamic cores Where N' th <=N th is the number of predicted instances in the scene.

[0075] S32. Build kNN-Transformer, query feature is instance dynamic kernel K' th and background dynamic kernel K st , the key / value feature is the feature of the k neighboring points of each dynamic kernel in the spatial position, and the adjacent features are weighted and fused by the attention mechanism to enhance the feature expression of the dynamic kernel; specifically, the calculation process of the features of the k neighboring points in step S32 is as follows: each instance dynamic kernel K' th The spatial position of is set to the center of its corresponding 3D bounding box, and the k points closest to the spatial position are indexed as the focus point; for the background class dynamic kernel, calculate The cosine similarity between each pixel feature and the dynamic kernel is calculated, and then the top-k pixels are taken and projected back into the 3D space to obtain the features of the corresponding points;

[0076] The calculation process of the weighted fusion of adjacent features by the attention mechanism in step S32 is specifically as follows:

[0077] in For location-aware instance dynamic kernels and background class dynamic kernels, key / value It is a linear mapping of the point features of the k nearest neighbors.

[0078] S33, using the instance dynamic kernel K' thAnd the corresponding 3D bounding box B predicts instance segmentation: the bounding box is expanded by a certain proportion to obtain the search area of ​​the target object, the feature embedding of the points in the search area is subjected to dot product with the dynamic kernel and the sigmoid activation function to predict the mask score, and the points in the same search area with a mask score greater than 0.5 are assigned the same instance ID; for background class points, their feature embeddings are convolved with the background class dynamic convolution kernel, and then the argmax function is used to assign the semantic category of the dynamic kernel to each background point.

[0079] Step S4: Using the server, the network parameters are optimized by reducing the network loss function until the network converges, thereby obtaining a radar panoramic segmentation method based on a decoupled dynamic convolution kernel. Figure 1 The figure shows an overall algorithm framework diagram of a radar panoramic segmentation method based on a decoupled dynamic convolution kernel provided by the present invention, and the specific steps are as follows:

[0080] S41, collect the point cloud data in the training data set, use the server to execute the point cloud feature extraction process in step S1, output voxel features and point-level feature embedding; input the voxel features into the BEV bird's-eye view heat map prediction module in step S2, and generate the instance center point heat map M th , Background area map M st , FocalLoss is used to supervise the two heat maps, which can be expressed as:

[0081] L pos =FL(M th ,Y th ) / N th +FL(M st ,Y st ) / N st

[0082] in, M th The true label of is generated by mapping the center of the object with semantic class c to the cth channel using a Gaussian kernel; M st The true label of , which maps each background class point to the corresponding channel of the BEV map.

[0083] S42, using the server to execute the dynamic convolution kernel decoding module in step S3, using the instance dynamic kernel to predict the 3D bounding box of the instance target, and supervising the predicted bounding box and IoU, which can be specifically expressed as:

[0084] L det =L box +λ IoU ·L IoU

[0085] Where Lbox is the L1 loss for bounding box regression, L IoU is the SmoothL1 loss of the bounding box IoU;

[0086] S43, using the server to execute the kNN-Transformer in step S3 to weight the fusion point cloud features with dynamic kernels, and finally predicting the mask by convolution of the dynamic kernel and the point cloud features, outputting the scene point cloud panoramic segmentation results, and supervising the segmentation results. In the mask decoder training stage, the Hungarian matching algorithm is used to match the predicted instance box and the true label, and the loss function can be specifically expressed as:

[0087]

[0088] Among them, the segmentation mask of the instance object uses FocalLoss (L fl ) and DiceLoss(L dl ) for supervision, and the background category segmentation mask uses the cross entropy loss (L ce ) and Lovasz softmax loss (L ls ) supervision.

[0089] S44, use the server to train the network, the loss function L total is the weighted sum of the losses in step S41, step S42 and step S43:

[0090] L total =L pos +L det +L mask

[0091] S45. Utilize the server to optimize the objective function and obtain local optimal network parameters.

[0092] Step S5: Use the radar panoramic segmentation method based on decoupled dynamic convolution kernel to perform panoramic segmentation prediction on the new point cloud (simultaneously perform semantic segmentation and instance segmentation). The specific steps are as follows:

[0093] S51, obtain the laser radar point cloud in the 3D environment, input it into the trained point cloud feature extractor, extract voxel features and point-level feature embedding;

[0094] S52, inputting the voxel features into the trained dynamic convolution kernel generation module to generate the initial weights of the dynamic convolution kernel of the instance / background class;

[0095] S53, input the dynamic convolution kernel into the trained dynamic convolution kernel decoding module, predict the 3D bounding box of the instance target, decode the dynamic kernel features, and finally predict and output the scene point cloud panoramic segmentation result. Figure 2The figure shows the visualization effect diagram of the algorithm proposed in the present invention. The algorithm proposed in the present invention has good segmentation results on adjacent small targets and large targets.

[0096] The main innovative features of the present invention are as follows:

[0097] (1) This paper proposes a novel Panoptic DKNet network model to achieve panoramic segmentation of radar point clouds in a unified workflow with decoupled dynamic convolution kernels.

[0098] (2) Panoptic DKNet decouples the dynamic kernels of instance target and background categories to facilitate their respective learning processes, while implementing a decoupling strategy for classification and segmentation to avoid competition between different categories.

[0099] (3) A large number of experiments on the SemanticKITTI benchmark dataset show that the PanopticDKNet proposed in this paper achieves comparable performance in the radar point cloud panoptic segmentation task. As shown in the following table:

[0100] Table 1 Comparison of the proposed method (Panoptic DKNet) and other methods on the SemanticKITTI validation set

[0101]

[0102] Although the present invention is described herein with reference to specific embodiments, it should be understood that these embodiments are merely examples of the principles and applications of the present invention. It should therefore be understood that many modifications may be made to the exemplary embodiments and that other arrangements may be devised without departing from the spirit and scope of the present invention as defined by the appended claims. It should be understood that the various dependent claims and features described herein may be combined in a manner different from that described in the original claims. It should also be understood that features described in conjunction with individual embodiments may be used in other described embodiments.

Claims

1. A radar panoramic segmentation method based on decoupled dynamic convolution kernel, characterized in that: The following steps are involved: S1. Build a point cloud feature extractor to extract voxel features and point-level feature embedding of the input point cloud based on voxel representation and 3D sparse convolution; S2. Construct a dynamic convolution kernel generation module to generate the initial weights of the dynamic convolution kernel based on the bird's-eye view heat map and the predicted instance and background point cloud positions from the BEV perspective; S3. Construct a dynamic convolution kernel decoding module, use the instance dynamic kernel to predict the 3D bounding box of the instance target, use kNN-Transformer to weight the dynamic kernel to fuse the point cloud features, and finally predict the mask by convolving the dynamic kernel with the point cloud features to output the scene point cloud panoramic segmentation result; S4. Using the server, the network parameters are optimized by reducing the network loss function until the network converges, thereby obtaining a radar panoramic segmentation method based on a decoupled dynamic convolution kernel; S5. Performing panoramic segmentation prediction on the new point cloud using the radar panoramic segmentation method based on decoupled dynamic convolution kernel; The step S4 specifically comprises the following steps: S41, collect the point cloud data in the training data set, use the server to execute the point cloud feature extraction process in step S1, output voxel features and point-level feature embedding; input the voxel features into the BEV bird's-eye view heat map prediction module in step S2, and generate the instance center point heat map M th , Background area map M st , Focal Loss is used to supervise the two heat maps, which can be expressed as: L pos =FL(M th ,Y th ) / N th +FL(M st ,Y st ) / N st S42, using the server to execute the dynamic convolution kernel decoding module in step S3, using the instance dynamic kernel to predict the 3D bounding box of the instance target, and supervising the predicted bounding box and IoU, which can be specifically expressed as: L det =L box +λ IoU ·L IoU Where L box is the L1 loss for bounding box regression, L IoU is the Smooth L1 loss of the bounding box IoU; S43, using the server to execute the kNN-Transformer in step S3 to weight the fusion point cloud features with the dynamic kernel, and finally predicting the mask by convolution of the dynamic kernel and the point cloud features, outputting the scene point cloud panoramic segmentation result, and supervising the segmentation result, which can be specifically expressed as: S44, use the server to train the network, the loss function L total is the weighted sum of the losses in step S41, step S42 and step S43: L total =L pos +L det +L mask S45, using the server to optimize the objective function and obtain local optimal network parameters; The specific calculation method of the loss of the segmentation result in step S43 is as follows: in the mask decoder training stage, the Hungarian matching algorithm is used to match the predicted instance box and the real label, and the Focal Loss (L fl ) and Dice Loss(L dl ) to supervise the segmentation mask of the instance object, using the cross entropy loss (L ce ) and Lovasz softmax loss (L ls )Supervised background category segmentation mask.

2. The radar panoramic segmentation method based on decoupled dynamic convolution kernel according to claim 1, characterized in that: The step S1 specifically includes the following steps: S11. Input sparse point cloud The point cloud is converted into voxels using voxelization operation, and then the initialization features of each voxel are extracted; S12. Construct a 3D sparse feature encoder and a multi-scale global attention module to extract sparse voxel feature expressions At the same time, we can get point-level feature embedding 3. The radar panoramic segmentation method based on decoupled dynamic convolution kernel according to claim 2, characterized in that: The specific process of the voxelization operation is as follows: first, the voxel resolution s is set, and then for a point p in a given point cloud P i =(x i ,y i , z i ), the index of the voxel to which it belongs is The initial feature extraction method for each voxel is as follows: the coordinates of all points in a voxel are sent to the multi-layer perceptron MLP to extract the features of the points, and then the features of multiple points are fused using the Max-Pooling function to obtain the feature expression f of a voxel. v ; The extraction process of the sparse voxel feature expression and point-level feature embedding is as follows: first, a 3D sparse convolution with a Bottleneck structure is used to extract the local information of the voxel, and then a cross-scale global attention module is used to build a long-distance dependency relationship and establish the correlation between voxels. A four-layer 3D sparse feature encoder is constructed, and the feature of the last encoder is used as the sparse voxel feature F v , whose dense spatial resolution is L×H×W, and the point-level feature embedding is: fe=Concat(f v,1 ,f v,2 ,f v,3 ,f v,4 ,MLP(f p )) where f v,j is the i-th layer encoding feature of the voxel where the point is located, f p are the spatial coordinates of the points, and the feature embedding of each point contains the global encoding information and its unique representation.

4. The radar panoramic segmentation method based on decoupled dynamic convolution kernel according to claim 3, characterized in that: The step S2 specifically includes the following steps: S21, the sparse voxel feature F v Max-Pooling is performed on the Z axis to obtain the BEV feature Then, based on the BEV feature, the instance center point heat map is predicted. and background area map S22, based on the instance center point heat map M th , Background area map M st , and BEV characteristic F bev , extract BEV features at the corresponding positions of high response in the heat map to obtain the initial instance dynamic kernel and background class dynamic kernel 5. The radar panoramic segmentation method based on decoupled dynamic convolution kernel according to claim 4, characterized in that: The heat map of the center point of the BEV feature prediction instance in step S21 The specific method is: in F bev Multi-layer 2D convolution is used to predict the center point heat map, where N ins is the number of semantic categories of the instance, and each channel represents the score of a category; The background area map predicted based on the BEV feature in step S21 The specific method is: in F bev A small U-Net structure is used to predict the 2D background area segmentation result under the BEV perspective, where N st is the number of semantic categories of the background class; The heat map M based on the instance center point in step S22 th and BEV characteristic F bev Generate instance dynamic core The specific method is: select M th The highest score N th positions, which indicate potential instance objects; then extract F bev The features at the corresponding position are used to generate the dynamic kernel of each instance object; The background area map M in step S22 st and BEV characteristic F bev Generate background class dynamic core The specific method is: multiply the background area map with the BEV feature to obtain the BEV feature of the background Then, an adaptive average pooling operation is performed in the spatial dimension to obtain the dynamic kernel for each background category.

6. The radar panoramic segmentation method based on decoupled dynamic convolution kernel according to claim 5, characterized in that: The step S3 specifically comprises the following steps: S31, based on the instance dynamic kernel K th , using the Transformer module and multi-layer perceptron, predict the 3D bounding box and overlap intersection over union (IoU) of each dynamic kernel corresponding to the instance target; then use the non-maximum suppression method to eliminate redundant bounding boxes and obtain a streamlined bounding box set At the same time, the corresponding redundant instance dynamic core is also changed from Eliminate them to obtain a streamlined set of instance dynamic cores Where N' th <=N th is the number of predicted instances in the scene; S32. Build kNN-Transformer, query feature is instance dynamic kernel K' th and background dynamic kernel K st , the key / value feature is the feature of the k nearest neighbor points of each dynamic kernel in the spatial position, and the attention mechanism is used to weightedly fuse the adjacent features to enhance the feature expression of the dynamic kernel; S33, using the instance dynamic kernel K' th And the corresponding 3D bounding box B predicts instance segmentation: the bounding box is expanded by a certain proportion to obtain the search area of ​​the target object, the feature embedding of the points in the search area is subjected to dot product with the dynamic kernel and the sigmoid activation function to predict the mask score, and the points in the same search area with a mask score greater than 0.5 are assigned the same instance ID; for background class points, their feature embeddings are convolved with the background class dynamic convolution kernel, and then the argmax function is used to assign the semantic category of the dynamic kernel to each background point.

7. The radar panoramic segmentation method based on decoupled dynamic convolution kernel according to claim 6, characterized in that: The calculation process of the features of the k nearest neighbor points in step S32 is specifically as follows: t ' h The spatial position of is set to the center of its corresponding 3D bounding box, and the k points closest to the spatial position are indexed as the focus point; for the background class dynamic kernel, calculate The cosine similarity between each pixel feature and the dynamic kernel is calculated, and then the top-k pixels are taken and projected back into the 3D space to obtain the features of the corresponding points; The calculation process of the attention mechanism weighted fusion of adjacent features in step S32 is specifically as follows: in For location-aware instance dynamic kernels and background class dynamic kernels, key / value It is a linear mapping of the point features of the k nearest neighbors.

8. The radar panoramic segmentation method based on decoupled dynamic convolution kernel according to claim 7, characterized in that: The instance center point heat map M in step S41 th The real label is generated using a Gaussian kernel, mapping the center of an object with semantic class c to the sth channel; the background area map M st The real label The point of each background class is mapped to the corresponding channel of the BEV map.

9. The radar panoramic segmentation method based on decoupled dynamic convolution kernel according to claim 8, characterized in that: The step S5 specifically comprises the following steps: S51, obtain the laser radar point cloud in the 3D environment, input it into the trained point cloud feature extractor, extract voxel features and point-level feature embedding; S52, inputting the voxel features into the trained dynamic convolution kernel generation module to generate the initial weights of the dynamic convolution kernel of the instance / background class; S53, input the dynamic convolution kernel into the trained dynamic convolution kernel decoding module, predict the 3D bounding box of the instance target, decode the dynamic kernel features, and finally predict and output the scene point cloud panoramic segmentation result.

Citation Information

Patent Citations

  • Urban street point cloud semantic segmentation method based on self-attention global feature enhancement

    CN115147601A

  • Three-dimensional target detection method and apparatus

    WO2022206414A1