Three-dimensional point cloud data processing method and device based on point cloud backbone network

By adopting a three-dimensional point cloud data processing method based on point cloud backbone network in the point cloud perception model, the problems of slow inference speed and difficult deployment in the existing technology are solved, efficient and accurate point cloud perception is achieved, and easy deployment is achieved on low computing power platforms.

CN120071277APending Publication Date: 2025-05-30SHENZHEN DEEPROUTE AI CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202311631003.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-11-29
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

Existing point cloud-aware models have shortcomings in inference speed and deployment ease of use, especially on low-computing platforms.

Method used

The three-dimensional point cloud data processing method based on the point cloud backbone network is adopted to obtain 3D voxel features through preprocessing, and voxel feature enhancement, conversion and dense feature enhancement processing are performed in the backbone network, and downsampling and enlarged receptive field processing are combined with the sub-network, and dense features of multi-scale receptive field sensitivity are finally obtained through upsampling and fusion.

Benefits of technology

It improves the accuracy of the point cloud perception model, reduces the computing requirements, speeds up the model inference speed, and makes it easier to deploy on low-computing platforms.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120071277A_ABST
    Figure CN120071277A_ABST
Patent Text Reader

Abstract

The invention discloses a three-dimensional point cloud data processing method and device based on a point cloud backbone network, and the method comprises the steps: obtaining point cloud data, and carrying out the preprocessing of the point cloud data; inputting the 3D voxel features into a backbone network, and performing voxel feature enhancement, conversion from the voxel features to upright post features, upright post feature enhancement, conversion from the upright post features to dense features and dense feature enhancement processing through the backbone network; the dense features with the high-order semantic information are input into a sub-network, downsampling and receptive field expansion processing are carried out through the sub-network, and low-resolution dense features of a large receptive field are obtained; and performing up-sampling on the low-resolution features of the large receptive field, and fusing the low-resolution features of the large receptive field with high-order dense features output by the backbone network to obtain multi-scale receptive field sensitive dense features for 3D point cloud detection or segmentation tasks. Under the condition of meeting a large receptive field, the precision of the point cloud sensing model is improved, the calculation requirement of the point cloud sensing model is reduced, and the model reasoning speed is increased.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of autonomous driving technology, and in particular to a three-dimensional point cloud data processing method and device based on a point cloud backbone network. Background Art

[0002] 3D detection and segmentation based on point cloud data scanned by LiDAR is an indispensable processing process in the autonomous driving perception system.

[0003] In the autonomous driving perception system, the commonly used point cloud perception models include: pillar-based detection model and voxel-based segmentation model. In these two point cloud perception models, directly transplanting the pillar-based detection model to the segmentation task will cause a serious drop in indicators due to the lack of feature granularity in height; at the same time, directly transplanting the voxel-based segmentation model to the detection task will also significantly increase the computational cost due to the finer feature granularity.

[0004] Therefore, how to design a point cloud backbone network model and three-dimensional point cloud data processing method that unifies different downstream tasks, is efficient, and easy to deploy is a difficult problem that plagues the industry. Summary of the invention

[0005] The technical problem to be solved by the present invention is that, in view of the defects of the prior art, the present invention provides a three-dimensional point cloud data processing method and device based on a point cloud backbone network to solve the problem that the existing point cloud perception model has slow reasoning speed and is difficult to be deployed on a low computing power platform.

[0006] The technical solution adopted by the present invention to solve the technical problem is as follows:

[0007] In a first aspect, the present invention provides a three-dimensional point cloud data processing method based on a point cloud backbone network, comprising:

[0008] Acquire point cloud data scanned by the laser radar, and pre-process the point cloud data to obtain 3D voxel features;

[0009] The 3D voxel features are input into a backbone network, and the backbone network is used to perform voxel feature enhancement, voxel feature conversion to column feature, column feature enhancement, column feature conversion to dense feature, and dense feature enhancement processing to obtain dense features with high-order semantic information;

[0010] Inputting the dense features with high-order semantic information into a sub-network, performing downsampling and expanding the receptive field through the sub-network, and obtaining low-resolution dense features with a large receptive field;

[0011] Upsample the low-resolution features of the large receptive field and fuse them with the high-order dense features output by the backbone network to obtain dense features sensitive to multi-scale receptive fields for 3D point cloud detection or segmentation tasks.

[0012] In one implementation, the preprocessing of the point cloud data includes:

[0013] Divide the point cloud data into voxel grids and resample the point clouds covered in each voxel grid;

[0014] Use several voxel feature encoding layers to encode the features of non-empty voxel grids to obtain the encoded features of the point clouds in each non-empty voxel grid;

[0015] Perform a pooling operation on the point cloud features in each voxel grid to obtain the 3D voxel features.

[0016] In one implementation, inputting the 3D voxel features into the backbone network and performing voxel feature enhancement, voxel feature to column feature conversion, column feature enhancement, column feature to dense feature conversion, and dense feature enhancement processing through the backbone network includes:

[0017] Input the 3D voxel features into the first stage of the backbone network and obtain the underlying voxel features through feature extraction;

[0018] Input the underlying voxel features into the second stage of the backbone network, obtain the enhanced voxel features through voxel feature enhancement processing, and fuse the enhanced voxel features in the height direction to convert them into column features;

[0019] Input the column features into the third stage of the backbone network, obtain high-order column features through column feature enhancement processing, and convert the high-order column features into dense features through sparse coordinate indexing;

[0020] Input the dense features into the fourth stage of the backbone network and obtain the dense features with high-order semantic information through dense feature enhancement processing.

[0021] In one implementation, the conversion of the high-order column features into dense features through sparse coordinate indexing includes:

[0022] Find the corresponding positions of the dense features in the top view according to the sparse coordinate indexing;

[0023] Fill the C-dimensional features into the found positions and fill the positions where the index coordinates do not appear with all-zero features of C dimensions to obtain the top view dense features with the shape of B*C*H*W.

[0024] In one implementation, inputting the dense features with high-order semantic information into the sub-network and performing downsampling and receptive field expansion processing through the sub-network includes:

[0025] Input the dense features with high-order semantic information into the sub-network, perform downsampling in the sub-network to obtain high-order dense features with a smaller resolution;

[0026] Perform dense feature enhancement processing on the high-order dense features with a smaller resolution to obtain low-resolution dense features with a large receptive field.

[0027] In one implementation, the dense feature enhancement processing on the high-order dense features with a smaller resolution includes:

[0028] Pass the high-order dense features with a smaller resolution through the first convolutional layer and perform compression processing in the channel dimension to obtain feature e;

[0029] Pass the high-order dense features with a smaller resolution through the second convolutional layer, perform compression processing in the channel dimension, and obtain intermediate features through a large receptive field extraction unit. Pass the intermediate features through the third convolutional layer to obtain feature f;

[0030] Perform channel concatenation on feature e and feature f, and pass the concatenated features through the fourth convolutional layer to obtain the low-resolution dense features with a large receptive field.

[0031] In one implementation, the upsampling of the low-resolution features with a large receptive field and the fusion with the high-order dense features output by the backbone network include:

[0032] Perform upsampling on the low-resolution features with a large receptive field, and fuse the upsampled dense features with the high-order dense features before downsampling to obtain the dense features sensitive to multi-scale receptive fields.

[0033] In a second aspect, the present invention provides a three-dimensional point cloud data processing device based on a point cloud backbone network, including:

[0034] A preprocessing module for acquiring point cloud data scanned by a lidar and preprocessing the point cloud data to obtain 3D voxel features;

[0035] A backbone network module for inputting the 3D voxel features into the backbone network, and performing voxel feature enhancement, voxel feature to column feature conversion, column feature enhancement, column feature to dense feature conversion, and dense feature enhancement processing through the backbone network to obtain dense features with high-order semantic information;

[0036] A sub-network module, configured to input the dense features with high-order semantic information into the sub-network, and perform downsampling and receptive field expansion processing through the sub-network to obtain low-resolution dense features with a large receptive field;

[0037] A fusion processing module, configured to upsample the low-resolution features with a large receptive field and fuse them with the high-order dense features output by the backbone network to obtain dense features sensitive to multi-scale receptive fields for 3D point cloud detection or segmentation tasks.

[0038] In a third aspect, the present invention provides a terminal, including: a processor and a memory, where the memory stores a three-dimensional point cloud data processing program based on a point cloud backbone network, and when the three-dimensional point cloud data processing program based on the point cloud backbone network is executed by the processor, it is used to implement the operations of the three-dimensional point cloud data processing method based on the point cloud backbone network as described in the first aspect.

[0039] In a fourth aspect, the present invention further provides a medium, where the medium is a computer-readable storage medium, and the medium stores a three-dimensional point cloud data processing program based on a point cloud backbone network. When the three-dimensional point cloud data processing program based on the point cloud backbone network is executed by a processor, it is used to implement the operations of the three-dimensional point cloud data processing method based on the point cloud backbone network as described in the first aspect.

[0040] The present invention adopts the above technical solutions and has the following effects:

[0041] By inputting the 3D voxel features after preprocessing the point cloud data into the backbone network, and using the backbone network for voxel feature enhancement, voxel feature to pillar feature conversion, pillar feature enhancement, pillar feature to dense feature conversion, and dense feature enhancement processing, the present invention can obtain dense features with high-order semantic information; and through the sub-network for downsampling and receptive field expansion processing, low-resolution dense features with a large receptive field can be obtained; and by upsampling the low-resolution features with a large receptive field and fusing them with the high-order dense features, dense features sensitive to multi-scale receptive fields for 3D point cloud detection or segmentation tasks can be obtained. The present invention improves the accuracy of the point cloud perception model, reduces the computational requirements of the point cloud perception model, and speeds up the model inference speed while meeting the large receptive field. Description of the Drawings

[0042] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the following drawings are only some embodiments of the present invention, and for those of ordinary skill in the art, other drawings can be obtained based on the structures shown in these drawings without creative efforts.

[0043] Figure 1 It is a flowchart of a 3D point cloud data processing method based on a point cloud backbone network in an implementation manner of the present invention.

[0044] Figure 2 It is a processing flowchart of a point cloud backbone network in an implementation manner of the present invention.

[0045] Figure 3 It is a processing schematic diagram of an efficient large receptive field feature extraction unit in an implementation manner of the present invention.

[0046] Figure 4 It is a functional schematic diagram of a terminal in an implementation manner of the present invention.

[0047] The realization of the purpose, functional characteristics and advantages of the present invention will be further described with reference to the embodiments and the accompanying drawings. Specific Embodiments

[0048] In order to make the purpose, technical solutions and advantages of the present invention clearer and more definite, the following further describes the present invention in detail with reference to the accompanying drawings and by way of examples. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.

[0049] Exemplary Method

[0050] In an autonomous driving perception system, common point cloud perception models include: a detection model based on columns and a segmentation model based on voxels. In these two point cloud perception models, directly transplanting the detection model based on column representation to the segmentation task will cause a serious decline in indicators due to the lack of feature granularity in the height direction; at the same time, directly transplanting the segmentation model based on voxel representation to the detection task will also significantly increase the computational cost due to the finer feature granularity. Therefore, how to design a point cloud backbone network model and a 3D point cloud data processing method that unify different downstream tasks, are efficient and easy to deploy is a difficult problem that plagues the industrial community.

[0051] In view of the above technical problems, an embodiment of the present invention provides a three-dimensional point cloud data processing method based on a point cloud backbone network. The method inputs the 3D voxel features after preprocessing the point cloud data into the backbone network, and uses the backbone network to perform voxel feature enhancement, voxel feature to pillar feature conversion, pillar feature enhancement, pillar feature to dense feature conversion, and dense feature enhancement processing, so as to obtain dense features with high-order semantic information; and performs downsampling and receptive field expansion processing through a sub-network to obtain low-resolution dense features with a large receptive field; and, by upsampling the low-resolution features with a large receptive field and fusing them with the high-order dense features, dense features sensitive to multi-scale receptive fields for 3D point cloud detection or segmentation tasks can be obtained. Therefore, the embodiment of the present invention can improve the accuracy of the point cloud perception model, reduce the computational requirements of the point cloud perception model, and accelerate the model inference speed while satisfying a large receptive field.

[0052] As Figure 1 shown, an embodiment of the present invention provides a three-dimensional point cloud data processing method based on a point cloud backbone network, including the following steps:

[0053] Step S100, obtain the point cloud data scanned by the lidar, and preprocess the point cloud data to obtain 3D voxel features.

[0054] In this embodiment, an efficient and general point cloud backbone network model for autonomous driving vehicle chips is designed. This model unifies the feature representations of different point cloud downstream tasks, has stronger versatility, is not limited to one downstream task, and can be friendly applied to various three-dimensional point cloud downstream tasks such as detection and segmentation. At the same time, the model is designed more efficiently and has stronger feature extraction ability, and can be better deployed to autonomous driving vehicle chips.

[0055] Specifically: in the macro architecture, in this embodiment, a method of mixed representation using three representation methods (i.e., voxel, pillar, and dense feature) in three-dimensional visual representation is used to extract high-order semantic information of the point cloud, which can well balance speed and accuracy, and at the same time meet the needs of feature extraction for various different types of tasks such as detection and segmentation; moreover, in the design of the neck (neck is a small network for fusing / enhancing the output features of the backbone network) sub-network, in this embodiment, a structure that meets a large receptive field and is very efficient is designed to further improve the overall accuracy of the model and accelerate the model inference speed.

[0056] Specifically, in an implementation manner of this embodiment, step S100 includes the following steps:

[0057] Step S101, divide the point cloud data into voxel grids, and resample the point cloud covered in each voxel grid;

[0058] Step S102: Use a number of voxel feature encoding layers to encode the features of non-empty voxel grids, and obtain the encoded features of the point cloud within each non-empty voxel grid.

[0059] Step S103: Perform a pooling operation on the point cloud features within each voxel grid to obtain the 3D voxel features.

[0060] In this embodiment, the point cloud data scanned by the lidar is processed to obtain dense features sensitive to multi-scale receptive fields for 3D point cloud detection or segmentation tasks. First, obtain the point cloud data from the lidar, and obtain the 3D voxel features through the preprocessing operation of the point cloud data.

[0061] It can be understood that in this embodiment, the point cloud data is obtained from the lidar, and these raw data can be directly used as the input data for downstream tasks such as detection / segmentation. In this embodiment, in order to simultaneously meet the needs of feature extraction for various different types of tasks such as detection and segmentation, these raw data are processed through the backbone network. Among them, the backbone network mainly extracts features from the input data of the downstream tasks, and then gives the extracted features to other modules of the downstream tasks to finally complete the downstream tasks.

[0062] As an example, in this embodiment, before the backbone network performs feature extraction, the point cloud data obtained from the lidar is processed through the preprocessing operation of the point cloud data, and then the 3D voxel features are obtained. Among them, the preprocessing (i.e., pretreatment) operation includes:

[0063] First, divide the point cloud data into voxel grids, and resample the point cloud covered in the voxel grids. Then, use a number of voxel feature encoding layers to encode the features of non-empty voxel grids to obtain the encoded features of the point cloud within each non-empty voxel grid. Finally, perform a pooling operation on the point cloud features within a voxel grid to obtain the final 3D voxel features.

[0064] As Figure 1 shown, in an implementation manner of the embodiment of the present invention, the method for processing three-dimensional point cloud data based on the point cloud backbone network further includes the following steps:

[0065] Step S200: Input the 3D voxel features into the backbone network, and perform voxel feature enhancement, voxel feature to column feature conversion, column feature enhancement, column feature to dense feature conversion, and dense feature enhancement processing through the backbone network to obtain dense features with high-order semantic information.

[0066] In this embodiment, for the above-obtained 3D voxel features, these 3D voxel features can be input into the backbone network, and feature extraction is performed through four feature extraction stages of the backbone network to obtain enhanced high-order dense features.

[0067] As Figure 2 shown, in this embodiment, the four feature extraction stages of the backbone network sequentially include: the first stage, the second stage, the third stage, and the fourth stage; among them, the voxel features obtained in the first stage are the underlying voxel features, the voxel features obtained in the second stage are the enhanced voxel features, the high-order column features obtained in the third stage are the high-order column features, and the enhanced dense features obtained in the fourth stage are the enhanced dense features.

[0068] Specifically, in one implementation manner of this embodiment, step S200 includes the following steps:

[0069] Step S201, input the 3D voxel features into the first stage of the backbone network, and obtain the underlying voxel features through feature extraction;

[0070] Step S202, input the underlying voxel features into the second stage of the backbone network, obtain the enhanced voxel features through voxel feature enhancement processing, and fuse the enhanced voxel features in the height direction to transform them into column features;

[0071] Step S203, input the column features into the third stage of the backbone network, obtain high-order column features through column feature enhancement processing, and transform the high-order column features into dense features through sparse coordinate indexing.

[0072] In this embodiment, it is necessary to input the 3D voxel features into the backbone network, and obtain the underlying voxel features through feature extraction in the first stage (stage1).

[0073] It is worth mentioning that the feature extraction operation in the first stage (stage1) in this embodiment is arbitrary, and it can be the operation of the existing Model A or the operation of the existing Model B, as long as the underlying voxel features can be extracted; similarly, the second stage, the third stage, and the fourth stage are not limited to the operations of specific models.

[0074] After obtaining the underlying voxel features, input the underlying voxel features into the second stage (stage2) of the backbone network for feature extraction. This process is a process of voxel feature enhancement processing to obtain the enhanced voxel features; at the same time, fuse the voxel features in the height direction to transform them into column features. This process is a process of transforming voxel features into column features.

[0075] After that, input the obtained column features into the third stage (stage3) of the backbone network for feature extraction. This process is a process of column feature enhancement processing to obtain high-order column features; finally, transform the high-order column features into dense features through the method of sparse coordinate indexing. This process is a process of transforming column features into dense features.

[0076] In one implementation of this embodiment, step S203 includes the following steps:

[0077] Step S203a, find the corresponding positions of the dense features in the top view according to the sparse coordinate index;

[0078] Step S203b, fill the C-dimensional features into the found positions, and fill the all-zero features of C dimensions at the positions where the index coordinates do not appear, to obtain the top view dense features with the shape of B*C*H*W.

[0079] In this embodiment, the high-order column features extracted through the third stage (stage3) mainly include two parts: one part is the column feature itself, with the shape of B*N*C. Wherein, B represents the batch size, N represents the number of columns, and C represents the representation dimension of each column feature; the other part is the position information of the column feature, with the shape of B*N*2. Wherein, B represents the batch size, N represents the number of columns, and 2 represents the two-dimensional coordinates in the top view.

[0080] As an example, the method of converting into dense features through the sparse coordinate index is:

[0081] First, according to the sparse coordinate index, find the corresponding positions of the dense features in the top view, then fill the C-dimensional features into this position, and fill the all-zero features of C dimensions at the positions where the index coordinates do not appear. Finally, obtain the top view dense features with the shape of B*C*H*W. Wherein, H and W represent the number of representations of the top view dense features calculated according to the coordinate range of the point cloud data and the size of a voxel grid in the top view.

[0082] After converting into dense features through the above sparse coordinate index method, the dense features can be enhanced, and thus the enhanced high-order dense features can be obtained.

[0083] Specifically, in one implementation of this embodiment, step S200 further includes the following steps:

[0084] Step S204, input the dense features into the fourth stage of the backbone network, and obtain the dense features with high-order semantic information through dense feature enhancement processing.

[0085] In this embodiment, the dense features obtained in the third stage (stage3) are sent to the fourth stage (stage4) of the backbone network for feature extraction. This process is the process of dense feature enhancement, and the enhanced high-order dense features are obtained. The enhanced high-order dense features are dense features with high-order semantic information. Thus, the process of feature extraction of the backbone network is completed.

[0086] Such asFigure 1 As shown, in one implementation of the embodiment of the present invention, the method for processing three-dimensional point cloud data based on a point cloud backbone network further includes the following steps:

[0087] Step S300: Input the dense feature with high-order semantic information into the sub-network, and perform downsampling and receptive field expansion processing through the sub-network to obtain a low-resolution dense feature with a large receptive field.

[0088] In this embodiment, for the enhanced high-order dense feature output by the backbone network, the neck sub-network performs downsampling and receptive field expansion processing on the feature output by the backbone network to obtain a low-resolution dense feature with a large receptive field.

[0089] Specifically, in one implementation of this embodiment, step S300 includes the following steps:

[0090] Step S301: Input the dense feature with high-order semantic information into the sub-network, and perform downsampling in the sub-network to obtain a high-order dense feature with a smaller resolution.

[0091] Step S302: Perform dense feature enhancement processing on the high-order dense feature with a smaller resolution to obtain the low-resolution dense feature with a large receptive field.

[0092] In this embodiment, first, the enhanced high-order dense feature output by the backbone network is fed into the neck sub-network. In the neck sub-network, a downsampling is required to obtain a high-order dense feature with a smaller resolution. As Figure 2 shown in the process of the neck sub-network, the shape of the dense feature with a smaller resolution obtained after downsampling is N, C, H, W; where N represents the number of input data, C represents the number of feature channels, and H and W respectively represent the height and width of the feature.

[0093] Then, perform feature enhancement processing on the high-order dense feature with a smaller resolution: that is, feed the feature into the Figure 3 shown efficient large receptive field block (a feature extraction unit) to obtain an enhanced dense feature.

[0094] In one implementation of this embodiment, step S302 includes the following steps:

[0095] Step S302a: Pass the high-order dense feature with a smaller resolution through the first convolutional layer to perform compression processing in the channel dimension to obtain feature e.

[0096] Step S302b: Pass the high-order dense features with smaller resolution through the second convolutional layer, perform compression processing in the channel dimension, and obtain intermediate features through a large receptive field extraction unit. Then pass the intermediate features through the third convolutional layer to obtain feature f.

[0097] Step S302c: Concatenate the features e and f in the channel dimension, and pass the concatenated features through the fourth convolutional layer to obtain the low-resolution dense features with a large receptive field.

[0098] In this embodiment, the internal operations of the efficient large receptive field block are as follows:

[0099] As Figure 3 shown, for the dense features N, C, H, W with a smaller resolution, they can be processed in two paths. One path passes through a convolution with a kernel size of 1*1 (i.e., the first convolutional layer), which compresses the original features to half of the original in the channel dimension to obtain feature e.

[0100] The other path first passes through a convolution with a kernel size of 1*1 (i.e., the second convolutional layer), which compresses the original features to half of the original in the channel dimension. Then, it passes through a large receptive field block to obtain intermediate features, and finally passes through a 1*1 convolution (i.e., the third convolutional layer) to obtain feature f.

[0101] Finally, concatenate the features e and f in the channel dimension and pass through a 1*1 convolution (i.e., the fourth convolutional layer) to obtain the final feature output (i.e., obtain the low-resolution dense features with a large receptive field).

[0102] As Figure 1 shown, in an implementation manner of this embodiment of the present invention, the 3D point cloud data processing method based on the point cloud backbone network further includes the following steps:

[0103] Step S400: Upsample the low-resolution features with a large receptive field and fuse them with the high-order dense features output by the backbone network to obtain multi-scale receptive field-sensitive dense features for 3D point cloud detection or segmentation tasks.

[0104] In this embodiment, for the obtained low-resolution dense features with a large receptive field, through corresponding upsampling and fusion processing, multi-scale receptive field-sensitive dense features for downstream detection or segmentation tasks can be obtained.

[0105] Specifically, in an implementation manner of this embodiment, step S400 includes the following steps:

[0106] Step S401: Upsample the low-resolution features of the large receptive field, and fuse the densely sampled features obtained by upsampling with the high-order dense features before downsampling to obtain the dense features sensitive to multi-scale receptive fields.

[0107] In this embodiment, the features finally output by the above neck sub-network are upsampled, and then fused with the dense features before downsampling of the neck sub-network (i.e., the enhanced high-order dense features output by the backbone network) to obtain multi-scale dense features, which can be used as the input features for different downstream task heads.

[0108] It is worth mentioning that in this embodiment, the input for different downstream tasks will go through the feature extraction process from the first stage (stage1) to the fourth stage (stage4). The feature extraction process for the neck sub-network part is mainly used for downstream tasks sensitive to multi-scale receptive fields such as detection.

[0109] The corresponding backbone network in this embodiment can be adapted to different downstream tasks (detection, segmentation, etc.), and the indicators are all very high. Moreover, in this embodiment, efficient inference can be performed, with small inference time consumption and less video memory resource occupation, and it is easy to be deployed to low-computing-power platforms.

[0110] This embodiment achieves the following technical effects through the above technical solutions:

[0111] This embodiment extracts the high-order semantic information of the point cloud through the method of hybrid representation of voxels, columns, and dense features in 3D visual representation, which can well balance speed and accuracy, and at the same time meet the needs of feature extraction for various different types of tasks such as detection and segmentation; therefore, this embodiment can improve the accuracy of the point cloud perception model, reduce the computational requirements of the point cloud perception model, and speed up the model inference speed while satisfying the large receptive field.

[0112] Exemplary Device

[0113] Based on the above embodiments, the present invention further provides a three-dimensional point cloud data processing device based on a point cloud backbone network, including:

[0114] A preprocessing module for obtaining the point cloud data scanned by the lidar and preprocessing the point cloud data to obtain 3D voxel features;

[0115] A backbone network module for inputting the 3D voxel features into the backbone network, and performing voxel feature enhancement, voxel feature to column feature conversion, column feature enhancement, column feature to dense feature conversion, and dense feature enhancement processing through the backbone network to obtain dense features with high-order semantic information;

[0116] A sub-network module for inputting the dense features with high-order semantic information into the sub-network, performing downsampling and expanding the receptive field processing through the sub-network, and obtaining low-resolution dense features with a large receptive field;

[0117] A fusion processing module for upsampling the low-resolution features with a large receptive field and fusing them with the high-order dense features output by the backbone network to obtain multi-scale receptive field-sensitive dense features for 3D point cloud detection or segmentation tasks.

[0118] This embodiment achieves the following technical effects through the above technical solutions:

[0119] This embodiment extracts the high-order semantic information of the point cloud through the method of hybrid representation of voxels, columns, and dense features in 3D visual representation, which can well balance speed and accuracy, and at the same time meet the needs of feature extraction for various different types of tasks such as detection and segmentation; therefore, this embodiment can improve the accuracy of the point cloud perception model, reduce the computational requirements of the point cloud perception model, and accelerate the model inference speed while satisfying a large receptive field.

[0120] Based on the above embodiments, the present invention also provides a terminal, and its principle block diagram can be as Figure 4 shown.

[0121] The terminal includes: a processor, a memory, an interface, a display screen, and a communication module connected through a system bus; wherein, the processor of the terminal is used to provide computing and control capabilities; the memory of the terminal includes a storage medium and an internal memory; the storage medium stores an operating system and a computer program; the internal memory provides an environment for the operation of the operating system and the computer program in the storage medium; the interface is used to connect external devices; the display screen is used to display corresponding information; the communication module is used to communicate with a cloud server or other devices.

[0122] When the computer program is executed by the processor, it is used to implement the operations of the 3D point cloud data processing method based on the point cloud backbone network.

[0123] Those skilled in the art can understand that Figure 4 the principle block diagram shown in

[0124] In one embodiment, a terminal is provided, which includes: a processor and a memory. The memory stores a three-dimensional point cloud data processing program based on a point cloud backbone network. When the three-dimensional point cloud data processing program based on the point cloud backbone network is executed by the processor, it is used to implement the operations of the above-mentioned three-dimensional point cloud data processing method based on the point cloud backbone network.

[0125] In one embodiment, a storage medium is provided, which stores a three-dimensional point cloud data processing program based on a point cloud backbone network. When the three-dimensional point cloud data processing program based on the point cloud backbone network is executed by the processor, it is used to implement the operations of the above-mentioned three-dimensional point cloud data processing method based on the point cloud backbone network.

[0126] Those of ordinary skill in the art can understand that all or part of the processes in the above-mentioned embodiment methods can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile storage medium. When the computer program is executed, it can include the processes of the above-mentioned method embodiments. Among them, any reference to a memory, storage, database, or other medium used in the embodiments provided by the present invention can include non-volatile and volatile memories.

[0127] In summary, the present invention provides a three-dimensional point cloud data processing method and device based on a point cloud backbone network. The method includes: acquiring point cloud data and preprocessing the point cloud data; inputting 3D voxel features into the backbone network, and performing voxel feature enhancement, voxel feature to column feature conversion, column feature enhancement, column feature to dense feature conversion, and dense feature enhancement processing through the backbone network; inputting the dense features with high-order semantic information into the sub-network, and performing downsampling and receptive field expansion processing through the sub-network to obtain low-resolution dense features with a large receptive field; upsampling the low-resolution features with a large receptive field and fusing them with the high-order dense features output by the backbone network to obtain multi-scale receptive field-sensitive dense features for 3D point cloud detection or segmentation tasks. The present invention improves the accuracy of the point cloud perception model, reduces the computational requirements of the point cloud perception model, and speeds up the model inference speed while satisfying a large receptive field.

[0128] It should be understood that the application of the present invention is not limited to the above examples. For those of ordinary skill in the art, improvements or transformations can be made according to the above description. All such improvements and transformations should fall within the protection scope of the appended claims of the present invention.

Claims

1. A method for processing 3D point cloud data based on a point cloud backbone network, characterized in that, it includes: Obtain the point cloud data scanned by the lidar, and preprocess the point cloud data to obtain 3D voxel features; Input the 3D voxel features into the backbone network, and perform voxel feature enhancement, voxel feature to pillar feature conversion, pillar feature enhancement, pillar feature to dense feature conversion, and dense feature enhancement processing through the backbone network to obtain dense features with high-order semantic information; Input the dense features with high-order semantic information into the sub-network, and perform downsampling and receptive field expansion processing through the sub-network to obtain low-resolution dense features with a large receptive field; Upsample the low-resolution features with a large receptive field, and fuse them with the high-order dense features output by the backbone network to obtain dense features sensitive to multi-scale receptive fields for 3D point cloud detection or segmentation tasks.

2. The method for processing 3D point cloud data based on a point cloud backbone network according to claim 1, characterized in that, The preprocessing of the point cloud data includes: Perform voxel grid division on the point cloud data, and resample the point cloud covered in each voxel grid; Use several voxel feature encoding layers to perform feature encoding on non-empty voxel grids to obtain the encoded features of the point cloud in each non-empty voxel grid; Perform a pooling operation on the point cloud features in each voxel grid to obtain the 3D voxel features.

3. The method for processing 3D point cloud data based on a point cloud backbone network according to claim 1, characterized in that, The input of the 3D voxel features into the backbone network, and the voxel feature enhancement, voxel feature to pillar feature conversion, pillar feature enhancement, pillar feature to dense feature conversion, and dense feature enhancement processing through the backbone network include: Input the 3D voxel features into the first stage of the backbone network, and obtain the underlying voxel features through feature extraction; Input the underlying voxel features into the second stage of the backbone network, perform voxel feature enhancement processing to obtain enhanced voxel features, and fuse the enhanced voxel features in the height direction to convert them into pillar features; Input the pillar features into the third stage of the backbone network, perform pillar feature enhancement processing to obtain high-order pillar features, and convert the high-order pillar features into dense features through sparse coordinate indexing; Input the dense features into the fourth stage of the backbone network, and perform dense feature enhancement processing to obtain the dense features with high-order semantic information.

4. The method for processing 3D point cloud data based on a point cloud backbone network according to claim 3, characterized in that, The conversion of the high-order pillar features into dense features through sparse coordinate indexing includes: Find the corresponding positions of the dense features in the top view according to the sparse coordinate indexing; Fill the C-dimensional features into the found positions, and fill the all-zero features of the C dimension at the positions where the index coordinates do not appear to obtain the top view dense features with the shape of B*C*H*W.

5. The method for processing 3D point cloud data based on a point cloud backbone network according to claim 1, characterized in that, Inputting the dense features with high - order semantic information into the sub - network, and performing down - sampling and receptive field expansion processing through the sub - network, includes: Inputting the dense features with high - order semantic information into the sub - network, and performing down - sampling in the sub - network to obtain high - order dense features with a smaller resolution; Performing dense feature enhancement processing on the high - order dense features with a smaller resolution to obtain low - resolution dense features with a large receptive field.

6. The 3D point cloud data processing method based on a point cloud backbone network according to claim 5, wherein, The performing dense feature enhancement processing on the high - order dense features with a smaller resolution includes: Passing the high - order dense features with a smaller resolution through a first convolutional layer to perform compression processing in the channel dimension to obtain feature e; Passing the high - order dense features with a smaller resolution through a second convolutional layer to perform compression processing in the channel dimension, and passing through a large receptive field extraction unit to obtain intermediate features, and passing the intermediate features through a third convolutional layer to obtain feature f; Performing channel concatenation on feature e and feature f, and passing the concatenated features through a fourth convolutional layer to obtain the low - resolution dense features with a large receptive field.

7. The 3D point cloud data processing method based on a point cloud backbone network according to claim 1, wherein, The performing up - sampling on the low - resolution features with a large receptive field and fusing them with the high - order dense features output by the backbone network includes: Performing up - sampling on the low - resolution features with a large receptive field, and fusing the up - sampled dense features with the high - order dense features before down - sampling to obtain the dense features sensitive to multi - scale receptive fields.

8. A 3D point cloud data processing device based on a point cloud backbone network, wherein, It includes: A pre - processing module, configured to obtain the point cloud data scanned by a lidar and perform pre - processing on the point cloud data to obtain 3D voxel features; A backbone network module, configured to input the 3D voxel features into the backbone network, and perform voxel feature enhancement, voxel feature to column feature conversion, column feature enhancement, column feature to dense feature conversion, and dense feature enhancement processing through the backbone network to obtain dense features with high - order semantic information; A sub - network module, configured to input the dense features with high - order semantic information into the sub - network, and perform down - sampling and receptive field expansion processing through the sub - network to obtain low - resolution dense features with a large receptive field; A fusion processing module, configured to perform up - sampling on the low - resolution features with a large receptive field and fuse them with the high - order dense features output by the backbone network to obtain dense features sensitive to multi - scale receptive fields for 3D point cloud detection or segmentation tasks.

9. A terminal, wherein, It includes: A processor and a memory, the memory stores a 3D point cloud data processing program based on a point cloud backbone network, and when the 3D point cloud data processing program based on a point cloud backbone network is executed by the processor, it is used to implement the operations of the 3D point cloud data processing method based on a point cloud backbone network according to any one of claims 1 - 7.

10. A computer - readable storage medium, It is characterized in that the computer-readable storage medium stores a three-dimensional point cloud data processing program based on a point cloud backbone network, and when the three-dimensional point cloud data processing program based on the point cloud backbone network is executed by a processor, it is used to implement the operations of the three-dimensional point cloud data processing method based on the point cloud backbone network described in any one of claims 1-7.