Method and apparatus for object detection in point cloud data
A hierarchical encoder-decoder network with sparse and dense blocks addresses the challenge of capturing long-range dependencies in 3D object detection, improving accuracy and efficiency in point clouds.
Patent Information
- Application Number
- PCT/CN2023/120608
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2023-09-22
- Publication Date
- 2025-07-03
AI Technical Summary
Existing methods for 3D object detection in point clouds face challenges in capturing long-range dependencies among features while maintaining low computational costs, as submanifold sparse convolutions hinder information exchange and regular sparse convolutions are computationally expensive.
A hierarchical encoder-decoder network is introduced, utilizing sparse and dense encoder-decoder blocks to extract multi-scale features, facilitating information exchange and maintaining sparsity, and incorporating a dense backbone to expand features towards object centers.
The network effectively captures long-range dependencies and maintains low computational costs, enhancing detection accuracy for large and distant objects, suitable for applications like autonomous driving.
Smart Images

Figure CN2023120608_03072025_PF_FP_ABST
Abstract
Description
METHOD AND APPARATUS FOR OBJECT DETECTION IN POINT CLOUD DATAFIELD
[0001] Aspects of the present disclosure relate generally to artificial intelligence, and more particularly, to method and apparatus provided for object detection in point cloud data.BACKGROUND
[0002] Object detection in point cloud data refers to automatically detecting and recognizing target objects in 3D point cloud data, and is important for tasks such as machine vision or autonomous driving. A primary challenge in 3D object detection stems from the sparse distribution of points within the 3D scene, thus learning effective representations from sparse input data is a key for 3D object detection in point clouds.
[0003] Existing methods typically employ 3D sparse convolutional neural networks with small kernels to extract features. To reduce computational costs, these methods resort to submanifold sparse convolutions, which prevent the information exchange among spatially disconnected features. Some approaches attempt to address this problem by introducing large-kernel convolutions or self-attention mechanisms, but they either achieve limited accuracy improvements or incur excessive computational costs.
[0004] Therefore, a method for object detection in point cloud data which is able to capture long-range dependencies among features in the spatial space, particularly for large and distant objects, and meanwhile maintain low computational cost is needed.SUMMARY
[0005] The following presents a simplified summary of one or more aspects to provide a basic understanding of such aspects. This summary is not an extensive overview of all contemplated aspects, and is intended to neither identify key or critical elements of all aspects nor delineate the scope of any or all aspects. Its sole purpose is to present some concepts of one or more aspects in a simplified form as a prelude to the more detailed description that is presented later.
[0006] Learning effective representations from sparse input data is a key challenge for 3D object detection in point clouds. Existing point-based methods and range-based methods either suffer from high computational costs or exhibit inferior detection accuracy. Currently, voxel-based methods dominate high-performance 3D object detection.
[0007] The voxel-based methods partition the unstructured point clouds into regular voxels and utilize sparse conventional neural network (CNNs) or transformers as backbones for feature extraction. Most existing sparse CNNs are primarily built by stacking submanifold sparse residual (SSR) blocks, each consisting of two submanifold sparse convolutions with small kernels. However, submanifold sparse convolutions maintain the same sparsity between input and output features, and therefore hinder the exchange of information among spatially disconnected features. Consequently, models employing SSR blocks face challenges in effectively capturing long-range dependencies among features. One potential solution is to replace the submanifold sparse convolutions in SSR block with regular sparse convolutions. However, this leads to a significant decrease in feature sparsity as the network deepens, resulting in substantial computational costs.
[0008] It is proposed herein a method and an apparatus for object detection in point cloud data which is able to capture long-range dependencies among features in the spatial space and meanwhile maintain low computational cost. Firstly, an encoder-decoder structure is disclosed herein to be used for capture long-range dependencies among features, in which an encoder is designed to extract multi-scale features through feature down-sampling, facilitating information exchange among spatially disconnected regions, meanwhile, the decoder is designed to incorporate multi-scale feature fusion to recover the lost details. Additionally, since object detectors typically rely on object centers for detection, a similar encoder-decoder structure is adopted, which expands the extracted sparse features of the previous encoder-decoder structure towards object centers.
[0009] Leveraging the encoder-decoder structures disclosed herein, it is introduced a hierarchical encoder-decoder network for object detection in point clouds. The disclosed network can learn powerful representations for the detection of objects, particularly for large and distant objects, thus making it suitable for outdoor autonomous driving among others. It is noted that although the hierarchical encoder-decoder network is described herein mainly to be used for 3D object detection, it can be applied to 2D object detection as well.
[0010] In an aspect, a computer implemented method for object detection in point cloud data is disclosed. The method comprises obtaining raw point cloud data and generating a 2D or 3D grid based on the raw point cloud data; extracting a feature map from the 2D or 3D grid through multiple convolution blocks sequentially, at least one convolution block comprises a first part and a second part followed the first part, wherein the first part comprises a number of down-sampling layers with each down-samples an input feature map of the down-sampling layer to extract valid features and obtain an output feature map with a smaller resolution relative to the input feature map, and one corresponding convolution residual block following each down-sampling layer, and wherein the second part comprises a same number of up-sampling layers as the down-sampling layers with each up-samples an input feature map of the up-sampling layer to a same resolution as its corresponding down-sampling layer; and feeding the extracted feature map into a detection head for outputting one or more objection predictions.
[0011] In a further aspect, wherein the convolution residual block is a submanifold sparse residual (SSR) block comprising two submanifold sparse convolutions.
[0012] In a further aspect, wherein up-sampling the input feature map of the up-sampling layer to the same resolution as its corresponding down-sampling layer comprises only up-sampling features in the input feature map of the up-sampling layer to the regions covered by the valid features in an input feature map of its corresponding down-sampling layer.
[0013] In a further aspect, wherein the at least one convolution block further comprises one or more SSR blocks prior to the first part.
[0014] In a further aspect, wherein regular sparse convolution or spatially sparse convolution is adopted as the down sampling layer, and / or sparse inverse convolution or plain deconvolution is adopted as the up-sampling layer.
[0015] In a further aspect, wherein an output feature map of one up-sampling layer are fused with an input feature map of its corresponding down-sampling layer with skip connections.
[0016] In a further aspect, wherein the convolution residual block is a dense residual (DR) block comprising two plain convolutions.
[0017] In a further aspect, wherein the at least one convolution block further comprises one or more DR blocks prior to the first part.
[0018] In a further aspect, wherein plain convolution or a DR block is adopted as the down sampling layer, such as a DR block having a stride of 2, and / or plain deconvolution is adopted as the up-sampling layer.
[0019] In a further aspect, wherein an output feature map of a previous convolution block is compressed into a Bird’s -Eye-View feature map and input to the at least one convolution block.
[0020] In a further aspect, the raw point cloud data are obtained from a LiDAR, a camera and / or an optical sensor.
[0021] In a further aspect, the method further comprises assisting a task of an apparatus based on the one or more object predictions, wherein the task comprises one of autonomous driving or intelligent manufacturing.
[0022] In a further aspect, the method further comprises calculating a loss between the one or more object predictions and corresponding ground-truth labels; and training a neural network model comprising the multiple convolution blocks based on minimizing the loss.
[0023] In an aspect, a vehicle is disclosed. The vehicle comprises one or more sensors, configured for obtaining raw point cloud data related to environment around the vehicle; one or more controllers; and one or more storage devices storing computer-executable instructions that, when executed, cause the one or more controllers to perform the operations of one of the methods disclosed herein.
[0024] In an aspect, a robot is disclosed. The robot comprises one or more sensors, configured for obtaining raw point cloud data related to environment around the robot; one or more processors; and one or more storage devices storing computer-executable instructions that, when executed, cause the one or more processors to perform the operations of one of the methods disclosed herein.
[0025] In an aspect, a computer system is disclosed. The computer system comprises one or more processors; and one or more storage devices storing computer-executable instructions that, when executed, cause the one or more processors to perform the operations of one of the methods disclosed herein.
[0026] In an aspect, one or more computer readable storage media storing computer-executable instructions that, when executed, cause one or more processors to perform the operations of one of the methods disclosed herein is disclosed.
[0027] In an aspect, a computer program product comprising computer-executable instructions that, when executed, cause one or more processors to perform the operations of one of the methods disclosed herein is disclosed.BRIEF DESCRIPTION OF THE DRAWINGS
[0028] The disclosed aspects will be described in connection with the appended drawings that are provided to illustrate and not to limit the disclosed aspects.
[0029] Fig. 1 illustrates an exemplary block diagram 100 of an example apparatus, in accordance with various aspects of the present disclosure.
[0030] Fig. 2a illustrates an exemplary diagram 200a of a submanifold sparse residual (SSR) block, in accordance with various aspects of the present disclosure.
[0031] Fig. 2b illustrates an exemplary diagram 200b of a regular sparse residual (RSR) block, in accordance with various aspects of the present disclosure.
[0032] Fig. 2c illustrates an exemplary diagram 200c of a sparse encoder-decoder (SED) block, in accordance with various aspects of the present disclosure.
[0033] Fig. 3a illustrates an exemplary architecture 300a of a sparse encoder-decoder (SED) block, in accordance with various aspects of the present disclosure.
[0034] Fig. 3b illustrates an exemplary architecture 300b of a dense encoder-decoder (DED) block, in accordance with various aspects of the present disclosure.
[0035] Fig. 4 illustrates an exemplary architecture 400 of a hierarchical encoder-decoder network, in accordance with various aspects of the present disclosure.
[0036] Fig. 5 illustrates an exemplary flow chart 500 for object detection in point cloud data, in accordance with various aspects of the present disclosure.
[0037] Fig. 6 illustrates an exemplary computer system 600, in accordance with various aspects of the present disclosure.DETAILED DESCRIPTION
[0038] The present disclosure will now be discussed with reference to several example implementations. It is to be understood that these implementations are discussed only for enabling those skilled in the art to better understand and thus implement the embodiments of the present disclosure, rather than suggesting any limitations on the scope of the present disclosure.
[0039] Various embodiments will be described in detail with reference to the accompanying drawings. Wherever possible, the same reference numbers will be used throughout the drawings to refer to the same or like parts. References made to examples and embodiments are for illustrative purposes, and are not intended to limit the scope of the disclosure. It is noted that “based on” used in the disclosure should be understood as “based at least on” , rather than “solely based on” or “merely based on” .
[0040] It is anticipated that applying the disclosed method herein may involve the use of user-related information, such as data, images or videos captured by sensors or cameras on a vehicle during driving. It should be noted that the use of user-related information requires user authorization and may not exceed the scope of users’a uthorization.
[0041] Object detection in point cloud data refers to automatically detecting and recognizing target objects in 3D point cloud data, and is important for tasks such as machine vision or autonomous driving. Fig. 1 illustrates an exemplary block diagram 100 of an example apparatus, in accordance with various aspects of the present disclosure.
[0042] The apparatus 100 illustrated in Fig. 1 may be a vehicle such as an autonomous vehicle, a self-controlled machine such as a robot, an agent for a certain task, or may be a part of the vehicle, the robot, the agent, and / or the like. The autonomous vehicle is taken as an example of the apparatus in Fig. 1 in the following description but not limiting.
[0043] A vehicle can be equipped with a sensor 110 for sensing information related to the road environment or surroundings in which the vehicle is travelling, and the sensed information may be obtained as point cloud data. In Fig. 1, only one sensor 110 is shown, but more than one sensor is possible. In an embodiment, sensor 110 may be one or more of camera, LiDAR, optical sensor and / or any combination thereof. It is possible that sensor 110 including other type of sensors suitable for obtaining point cloud data.
[0044] The vehicle may include a processing system 120. The processing system 120 may be implemented in various approaches, for example, the processing system 120 may include one or more processors and / or controllers as well as one or more memories, the processors and / or controllers may execute software to perform various operations or functions, such as operations or functions according to various aspect of the disclosure.
[0045] The processing system 120 may receive raw point cloud data from the sensor 110, and perform various operations by processing the raw point cloud data. In the example of Fig. 1, the processing system 120 may include a voxelization module 120-1, a feature extraction module 120-2 and an object prediction module 120-3. It is appreciated that the modules 120-1 to 120-3 may be implemented in various approaches, for example, may be implemented as software modules or functions which are executable by processors and / or controllers.
[0046] The voxelization module 120-1 may be configured to receive the point cloud data from the sensor 110 and perform voxelization to the raw point cloud data to generate a grid of voxels. In an embodiment, the voxels may be generated by uniformly partition along different axis. In an embodiment, the voxels may be generated by partitioning along different axis in different size. In an embodiment, the voxels may be generated by be generated with setting the size along one axis equal to the full size of the input point cloud to form a 2D grid.
[0047] The feature extraction module 120-2 may be configured to perform feature extraction to the generated grid from the voxelization module 120-1. Feature extraction is vital due to the extracted features may have the greatest impact on object prediction. In an embodiment, the feature extraction module 120-2 may be implemented as a part of a neural network model, which can be trained to learn effective representations from sparse input data for object detection.
[0048] The object prediction module 120-3 may be configured to receive the extracted features from the feature extraction module 120-2 and output one or more object predictions based on the extracted features. In an embodiment, the object prediction module 120-3 may be implemented as a part of a neural network model, which can be trained to effectively predict an object from known features.
[0049] In an embodiment, the feature extraction module 120-2 and the object prediction module 120-3 may consist one neural network model. In an embodiment, in case the neural network model is to be trained, the output object predictions can be used for calculating an object function for training. In another embodiment, in case the neural network model is trained, the output object predictions can be used in autonomous driving for assisting travel path planning. In an embodiment, the output object predictions may be indicated by the object prediction module 120-3 or other components of processing system 120 to the corresponding control unit of the vehicle.
[0050] Fig. 1 is shown as an example merely, and other implementations are possible.
[0051] For 3D object detection in point clouds, methods can be categorized into three groups: point-based, range-based, and voxel-based. Point-based methods directly extract geometric features from raw point clouds and make predictions, but require computationally intensive point sampling and neighbor search procedures. Range-based methods convert point clouds into pseudo images, thus benefiting from the well-established designs of 2D object detectors. While computationally efficient, these methods often exhibit lower accuracy. Voxel-based methods employ sparse CNNs that consist of submanifold and regular sparse convolutions with small kernels to extract features. Regular sparse convolutions can capture distant contextual information but are computationally expensive. On the other hand, submanifold sparse convolutions prioritize efficiency but sacrifice the model’s ability to capture long-range dependencies.
[0052] To capture long-range dependencies for 3D object detection while achieving competitive inference speed, it is disclosed herein an encoder-decoder structure to reduce the spatial distance between distant features through feature down-sampling and recover the lost details through feature fusion, and maintain the same sparsity of the output feature map as the input feature map.
[0053] Fig. 2a illustrates an exemplary diagram 200a of a submanifold sparse residual (SSR) block, in accordance with various aspects of the present disclosure.
[0054] The structure of a single SSR block is shown in Fig. 2a, an SSR block is illustrated as consisting of two submanifold sparse convolutions (represented as SS Conv in Fig. 2a) , which is not limiting and other numbers of submanifold sparse convolutions is possible. For the submanifold sparse (SS) convolution, the convolutional output is only calculated when the center of the kernel covers a valid feature.
[0055] As shown in Fig. 2a, two SS convolutions are sequentially applied to the input feature map, with skip connections (represented as SC in Fig. 2a) incorporated between the input and output feature maps of the SSR block. The point cloud data captured for object detection would generally be sparse, the sparsity of a feature map is defined as the ratio of the regions that are not occupied by valid (nonzero) features to the total area of the feature map. As illustrated in Fig. 2a, an empty feature is shown as blank, and a valid feature is shown as black. SS convolution only operates on valid features, allowing the output feature map of the SSR block to maintain the same sparsity as the input feature map.
[0056] However, this design hinders the exchange of information among spatially disconnected features. For instance, in the top feature map, the output feature marked by a star cannot receive information from the other three feature points outside the dashed square in the bottom feature map (marked by the triangles) . This poses a challenge for the model in capturing long-range dependencies.
[0057] Fig. 2a is shown as an example merely, and other implementations are possible.
[0058] Fig. 2b illustrates an exemplary diagram 200b of a regular sparse residual (RSR) block, in accordance with various aspects of the present disclosure.
[0059] The structure of a single RSR block is shown in Fig. 2b, an RSR block is illustrated as consisting of two regular sparse convolutions (represented as RS Conv in Fig. 2b) , which is not limiting and other numbers of regular sparse convolutions is possible. For the regular sparse (RS) convolution, the convolutional output is calculated on both valid features and expanded features, the expanded features correspond to the features that fall within the neighborhood of the valid features.
[0060] As shown in Fig. 2b, two RS convolutions are sequentially applied to the input feature map, with skip connections (represented as SC in Fig. 2b) incorporated between the input and output feature maps of the RSR block. As illustrated in Fig. 2b, an empty feature is shown as blank, a valid feature is shown as black, and an expanded feature is shown as grey. RS convolution operates on both valid and expanded features, that is, the convolution kernel center traverses the regions covered by these features. Taking a 2D RS convolution with a kernel size of 3×3 as an example, the neighborhood of a certain valid feature consists of the eight positions around it.
[0061] This design is able to capture long-range dependencies better compared with SSR block, the dashed square highlights the regions from which the output feature marked by a star can receive information. However, the RSR block leads to an output feature map with a lower sparsity compared with the input feature map. Stacking RS convolutions reduces the feature sparsity dramatically, which in turn leads to a notable decrease in model efficiency compared with using SS convolutions.
[0062] Fig. 2b is shown as an example merely, and other implementations are possible.
[0063] Fig. 2c illustrates an exemplary diagram 200c of a sparse encoder-decoder (SED) block, in accordance with various aspects of the present disclosure.
[0064] The SED block disclosed herein is designed to mainly overcome the limitations of SSR block described with Fig. 2a. The fundamental idea behind this design is to reduce the spatial distance between distant features through feature down-sampling and recover the lost details through multi-scale feature fusion.
[0065] The structure of a single SED block is shown in Fig. 2c, the SED block is illustrated as consisting of down-sampling (represented as D in Fig. 2c) , an SSR block and up-sampling (represented as U in Fig. 2c) , which is not limiting and other numbers of down-sampling / up-sampling layers and SSR blocks is possible. Besides, skip connections (represented as SC in Fig. 2b) is used to incorporate features between the input and output feature maps of the SED block
[0066] As shown in Fig. 2c, after feature down-sampling, the spatially disconnected valid features in the bottom feature map are integrated into the adjacent valid features in the middle feature map. An SSR block is subsequently applied to the middle feature map to promote interaction among valid features. Finally, the middle feature map is up-sampled to match the resolution of the input feature map. It is noted that the feature up-sampling layer (U) only up-samples features to the regions covered by the valid features in the input feature map in order to keep the same sparsity.
[0067] As a result, the proposed SED block can maintain the same sparsity between input and output feature maps. This characteristic prevents the introduction of excessive computational costs when stacking multiple SED blocks. Also, since the output feature marked by a star can receive information from other valid features, the proposed SED block can capture long-range dependencies.
[0068] Fig. 2c is shown as an example merely, and other implementations are possible.
[0069] Fig. 3a illustrates an exemplary architecture 300a of a sparse encoder-decoder (SED) block, in accordance with various aspects of the present disclosure. As discussed with Fig. 2c illustrating a two-scale SED block, Fig. 3a shows a three-scale SED block, although the SED block is shown as three-scale, other multi-scale is possible.
[0070] The three-scale SED block adopts an asymmetric encoder-decoder structure, with the encoder (represented as EN) responsible for extracting multi-scale features and the decoder (represented as DE) sequentially fusing the extracted multi-scale features with the help of skip connections.
[0071] As shown in Fig. 3a, the input feature map of SED block sequentially passes multiple layers or blocks, wherein SSR represents an SSR block as described with Fig. 2c, D represents a down-sampling layer, and U represents an up-sampling layer. F1 / F2 / F3 / F4 / F5 are the names of the corresponding feature maps. The number in parentheses indicates the resolution ratio of the corresponding feature map relative to the block input. Given the input feature map X, the function of the SED block can be formulated as follows: F1=SSRm (X) (1) F2=SSRm (Down1 (F1) ) (2) F3=SSRm (Down2 (F2) ) (3) F4=UP2 (F3) +F2 (4) F5=UP1 (F4) +F1 (5)
[0072] where F4 denotes the feature map has a same resolution as F2, and F5 denotes the output feature map has the same resolution as the input X. The resolution ratios of the intermediate feature maps F1, F2, F3, and F4 relative to the input X are 1, 1 / 2, 1 / 4, and 1 / 2, respectively. SSRm indicates m consecutive SSR blocks.
[0073] In an embodiment, RS convolution may be used as the feature down-sampling layer. In another embodiment, spatially sparse convolution may be used as the feature down-sampling layer, which calculate a convolution when a kernel covers a valid feature, no matter the valid feature is at the center or not. In other embodiments, any suitable convolution for dimensionality reduction may be used as the feature down-sampling layer.
[0074] In an embodiment, sparse inverse convolution may be used as the feature up-sampling layer. In another embodiment, plain deconvolution may be used as the feature up-sampling layer. In other embodiments, any suitable deconvolution for dimensionality ascent may be used as the feature up-sampling layer.
[0075] In an embodiment, the first m SSR blocks shown in Fig. 3a may be omitted.
[0076] With the disclosed encoder-decoder structure, the SED block facilitates information exchange among spatially disconnected features, thereby enabling the model to capture long-range dependencies. Additionally, the SED block can maintain the same sparsity between input and output feature maps, preventing excessive computational costs when stacking multiple SED blocks.
[0077] It is noted that the disclosed SED block is capable of processing both 2D and 3D features, depending on whether 2D or 3D sparse convolutions are used.
[0078] Fig. 3a is shown as an example merely, and other implementations are possible.
[0079] Additionally, existing high-performance 3D object detectors usually rely on object centers for detection. However, the feature maps extracted by purely sparse CNNs may have empty holes around object centers, especially for large objects. To overcome this issue, it is disclosed herein a dense encoder-decoder (DED) block that expands sparse features towards object centers.
[0080] Fig. 3b illustrates an exemplary architecture 300b of a dense encoder-decoder (DED) block, in accordance with various aspects of the present disclosure.
[0081] As shown in Fig. 3b, DED block shares a similar structure with the SED block but utilizes dense convolutions instead. The “dense” convolution refers to plain convolution, named to differentiate it from sparse convolution mentioned herein.
[0082] Specifically, the SSR blocks in the SED block may be replaced with a dense residual (DR) block, which is similar to the SSR block but consists of two plain convolutions. Modifications can be made to enable the DED block to effectively expand sparse features towards object centers.
[0083] In an embodiment, a DR block may be used as feature down-sampling layer, such as a DR block having a stride of 2. In another embodiment, a plain convolution may be used as feature down-sampling layer. In other embodiments, any suitable convolution for dimensionality reduction may be used as the feature down-sampling layer.
[0084] In an embodiment, a plain deconvolution may be used as feature up-sampling layer. In other embodiments, any suitable deconvolution for dimensionality ascent may be used as the feature up-sampling layer.
[0085] In an embodiment, the first m DR blocks in Fig. 3b may be omitted.
[0086] It is noted that the disclosed DED block is capable of processing 2D and 3D features, depending on whether 2D or 3D sparse convolutions are used.
[0087] Fig. 3b is shown as an example merely, and other implementations are possible.
[0088] Based on the disclosed SED block and DED block described with Fig. 3a and Fig. 3b, a hierarchical encoder-decoder network designed for object detection is disclosed. Fig. 4 illustrates an exemplary architecture 400 of a hierarchical encoder-decoder network, in accordance with various aspects of the present disclosure.
[0089] As shown in Fig. 4, given the raw point clouds, a voxelization module may used to perform voxelization to generate a grid of voxels denoted as F0.
[0090] In an embodiment, the point clouds may be obtained by sensor 110 described in Fig. 1. In an embodiment, the voxelization may be realized by voxelization module 120-1 described in Fig. 1. Other implementations are possible.
[0091] Subsequently, a sparse backbone (represented as B#1) including multiple convolution blocks may be used to extract sparse features, the sparse backbone is illustrated as comprising two SSR blocks and several SED blocks. The illustration is merely an example, other numbers or arrangements of SSR blocks and SED blocks are possible. F1 / F2 / F3 / F4 are the names of the corresponding feature maps. The number in parentheses indicates the resolution ratio of the corresponding feature map relative to the input F0. The down-sampling layers after F1 / F2 / F3 are omitted for simplicity.
[0092] The SSR block may perform feature extraction as described in Fig. 2a.
[0093] The SED block may perform feature extraction as described in Fig. 2c or Fig. 3a.
[0094] In an embodiment, the sparse backbone is processing 2D features, using 2D sparse convolutions through the blocks.
[0095] In another embodiment, the sparse backbone is processing 3D features, using 3D sparse convolutions through the blocks.
[0096] In an embodiment, subsequently, a dense backbone (represented as B#2) including multiple convolution blocks may be used to expanding the sparse features towards object centers, the dense backbone is illustrated as comprising n DED blocks. The illustration is merely an example, other numbers or arrangements of DED blocks in combination with other convolution blocks are possible. F5 / F6 are the names of the corresponding feature maps. The number in parentheses indicates the resolution ratio of the corresponding feature map relative to the input F0.
[0097] The DED block may perform feature extension as described in Fig. 3b.
[0098] In an embodiment, the dense backbone is processing 2D features, using 2D sparse convolutions through the blocks.
[0099] In another embodiment, the dense backbone is processing 3D features, using 3D sparse convolutions through the blocks.
[0100] In an embodiment, before being fed into the dense backbone, the sparse features are compressed into dense Bird’s -Eye-View features.
[0101] In an embodiment, the feature extraction and extension of the sparse backbone and the dense backbone may be realized by feature extraction module 120-2 described in Fig. 1. Other implementations are possible.
[0102] Finally, the output feature map is fed into the detection head (represented as DH) for object predictions.
[0103] It is noted that a stride for convolution herein can be designed according to implementations but not limiting. Additionally, the scale of each SED block or DED block can be designed according to implementations but not limiting.
[0104] It is also to be noted that although the SED block and DED block are described hierarchically stacked and used, they can be used separately in different implementations with other preamble or follow-up components.
[0105] Fig. 5 illustrates an exemplary flow chart 500 for object detection in point cloud data, in accordance with various aspects of the present disclosure. As described below, some or all illustrated features may be omitted in an implementation within the scope of the present disclosure, and some illustrated features may not be required for implementation of all embodiments. Further, some of the blocks may be performed parallel or in a different order. In some examples, the method may be carried out by any suitable apparatus or means for carrying out the functions or algorithm described below.
[0106] The method begins at block 501, with obtaining raw point cloud data and generating a 2D or 3D grid based on the raw point cloud data.
[0107] In an embodiment, the obtaining raw point cloud data may be performed by sensor 110 described in Fig. 1.
[0108] In an embodiment, the raw point cloud data may be obtained from a LiDAR, a camera and / or an optical sensor.
[0109] In an embodiment, the generating a 2D or 3D grid may be performed by voxelization module 120-1 described in Fig. 1.
[0110] In an example, the voxels may be generated by uniformly partition along different axis. In an example, the voxels may be generated by partitioning along different axis in different size. In an example, the voxels may be generated by be generated with setting the size along one axis equal to the full size of the input point cloud to form a 2D grid.
[0111] The method then proceeds to block 502, with extracting a feature map from the 2D or 3D grid through multiple convolution blocks sequentially.
[0112] In an embodiment, the extracting a feature map may be performed by feature extraction module 120-2 described in Fig. 1.
[0113] In an embodiment, the at least one convolution block performs 2D convolution for 2D input features, or the at least one convolution block perform 3D convolution for 3D input features.
[0114] In an embodiment, at least one convolution block of the multiple convolution blocks comprises a first part and a second part, wherein the second part follows the first part.
[0115] In an example, the first part comprises a number of down-sampling layers and one corresponding convolution residual block following each down-sampling layer, wherein each down-sampling layer down-samples an input feature map of the down-sampling layer to extract valid features and obtain an output feature map with a smaller resolution relative to the input feature map. And the second part comprises a same number of up-sampling layers as the down-sampling layers, with each up-sampling layer up-samples an input feature map of the up-sampling layer to a same resolution as its corresponding down-sampling layer.
[0116] In an aspect, the at least one convolution block may be an SED block described in Fig. 3a. For example, the first part may be the encoder structure and the second part may be the decoder structure described in Fig. 3a. For example, the convolution residual block may be a submanifold sparse residual (SSR) block described in Fig. 2a. For example, the SSR block may comprise two submanifold sparse convolutions.
[0117] In an aspect, the up-sampling layer may only up-samples features in the input feature map of the up-sampling layer to the regions covered by the valid features in an input feature map of its corresponding down-sampling layer, in order to maintain a same sparsity as its corresponding down-sampling layer.
[0118] In an aspect, the at least one convolution block may comprise one or more SSR blocks prior to the first part, for example, as the first m SSR blocks shown in Fig. 3a.
[0119] In an aspect, regular sparse convolution, spatially sparse convolution or any suitable convolution may be adopted as the down-sampling layer.
[0120] In an aspect, sparse inverse convolution, plain deconvolution or any suitable convolution may be adopted as the up-sampling layer.
[0121] In an aspect, an output feature map of one up-sampling layer are fused with an input feature map of its corresponding down-sampling layer with skip connections, for example, as shown in Fig, 3a and 3b.
[0122] In an aspect, the at least one convolution block may be an DED block described in Fig. 3b.For example, the first part may be the encoder structure and the second part may be the decoder structure described in Fig. 3b. For example, the convolution residual block may be a dense residual (DR) block described in Fig. 3b. For example, the DR block may comprise two plain convolutions.
[0123] In an aspect, the at least one convolution block may comprise one or more DR blocks prior to the first part, for example, as the first m DR blocks shown in Fig. 3b.
[0124] In an aspect, plain convolution, a DR block or any suitable convolution may be adopted as the down-sampling layer.
[0125] In an aspect, plain deconvolution or any suitable convolution may be adopted as the up-sampling layer.
[0126] In an aspect, an output feature map of a previous convolution block is compressed into a Bird’s-Eye-View feature map and input to the DED block, for example, as described in Fig. 4.
[0127] The method then proceeds to block 303, with feeding the extracted feature map into a detection head for outputting one or more objection predictions.
[0128] In an embodiment, the outputting one or more objection predictions may be performed by object prediction module 120-3 described in Fig. 1.
[0129] In an example, the output objection predictions can be used for assisting a downstream task, such as autonomous driving for a vehicle, intelligent manufacturing for a robot, or other suitable tasks performed by an agent. For example, the one or more objection predictions may be indicated to a controller in a vehicle, a processor in a robot for making decisions about the task.
[0130] In an example, the output objection predictions can be used for training the neural network model comprising the multiple convolution blocks. For example, a loss between the one or more object predictions and corresponding ground-truth labels may be calculated; and the neural network model can be trained based on minimizing the loss. For example, the training of the neural network model can be performed on an apparatus with adequate computation ability, and the trained neural network model may be deployed to a vehicle, a robot or an agent for a task.
[0131] Fig. 5 is shown as an example merely, and other implementations are possible.
[0132] It is noted that a stride for convolution herein can be designed according to implementations but not limiting. Additionally, the scale of each SED block or DED block can be designed according to implementations but not limiting.
[0133] In an aspect of the disclosure, a vehicle capable of autonomous driving is provided. For example, as illustrated in Fig. 1, the vehicle comprises one or more sensors configured for obtaining raw point cloud data related to its surroundings; one or more controllers; and one or more storage devices storing computer-executable instructions that, when executed, cause the one or more controllers to perform the operations of the method of the embodiments described in the disclosure.
[0134] In an aspect of the disclosure, a robot is provided. For example, as illustrated in Fig. 1 which may also represent the structure of the robot, the robot comprises one or more sensors configured for obtaining raw point cloud data related to environment around the robot; one or more processors; and one or more storage devices storing computer-executable instructions that, when executed, cause the one or more processors to perform the operations of the method of the embodiments described in the disclosure.
[0135] Fig. 6 illustrates an exemplary computer system 600, in accordance with various aspects of the present disclosure. The computer system may comprise at least one processor 610. The computer system may further comprise at least one storage device 620. It should be appreciated that the storage device 620 may store computer-executable instructions that, when executed, cause the processor 510 to perform any operations according to the embodiments of the present disclosure as described in connection with Figs 1-5.
[0136] The embodiments of the present disclosure may be embodied in one or more computer-readable medium such as non-transitory computer-readable medium. The non-transitory computer-readable medium may store computer-executable instructions that, when executed, cause one or more processors to perform any operations according to the embodiments of the present disclosure as described in connection with Figs 1-5.
[0137] The embodiments of the present disclosure may be embodied in a computer program product comprising computer-executable instructions that, when executed, cause one or more processors to perform any operations according to the embodiments of the present disclosure as described in connection with Figs 1-5.
[0138] It should be appreciated that all the operations in the methods described above are merely exemplary, and the present disclosure is not limited to any operations in the methods or sequence orders of these operations, and should cover all other equivalents under the same or similar concepts.
[0139] It should also be appreciated that all the modules in the apparatuses described above may be implemented in various approaches. These modules may be implemented as hardware, software, or a combination thereof. Moreover, any of these modules may be further functionally divided into sub-modules or combined together.
[0140] The previous description is provided to enable any person skilled in the art to practice the various aspects described herein. Various modifications to these aspects will be readily apparent to those skilled in the art, and the generic principles defined herein may be applied to other aspects. Thus, the claims are not intended to be limited to the aspects shown herein. All structural and functional equivalents to the elements of the various aspects described throughout the present disclosure that are known or later come to be known to those of ordinary skill in the art are expressly incorporated herein by reference and are intended to be encompassed by the claims.
Claims
1.A computer implemented method for object detection in point cloud data, the method comprising:obtaining raw point cloud data and generating a 2D or 3D grid based on the raw point cloud data;extracting a feature map from the 2D or 3D grid through multiple convolution blocks sequentially, at least one convolution block comprises a first part and a second part followed the first part,wherein the first part comprises a number of down-sampling layers with each down-samples an input feature map of the down-sampling layer to extract valid features and obtain an output feature map with a smaller resolution relative to the input feature map, and one corresponding convolution residual block following each down-sampling layer, andwherein the second part comprises a same number of up-sampling layers as the down-sampling layers with each up-samples an input feature map of the up-sampling layer to a same resolution as its corresponding down-sampling layer; andfeeding the extracted feature map into a detection head for outputting one or more objection predictions.2.The computer implemented method of claim 1, wherein the convolution residual block is a submanifold sparse residual (SSR) block comprising two submanifold sparse convolutions.3.The computer implemented method of claim 2, wherein up-sampling the input feature map of the up-sampling layer to the same resolution as its corresponding down-sampling layer comprising:only up-sampling features in the input feature map of the up-sampling layer to the regions covered by the valid features in an input feature map of its corresponding down-sampling layer.4.The computer implemented method of claim 2, wherein the at least one convolution block further comprising:one or more SSR blocks prior to the first part.5.The computer implemented method of claim 2, wherein regular sparse convolution or spatially sparse convolution is adopted as the down sampling layer, and / orsparse inverse convolution or plain deconvolution is adopted as the up-sampling layer.6.The computer implemented method of claim 1, wherein an output feature map of one up-sampling layer are fused with an input feature map of its corresponding down-sampling layer with skip connections.7.The computer implemented method of claim 1, wherein the convolution residual block is a dense residual (DR) block comprising two plain convolutions.8.The computer implemented method of claim 7, wherein the at least one convolution block further comprising:one or more DR blocks prior to the first part.9.The computer implemented method of claim 7, wherein plain convolution or a DR block is adopted as the down sampling layer, and / orplain deconvolution is adopted as the up-sampling layer.10.The computer implemented method of claim 7, wherein an output feature map of a previous convolution block is compressed into a Bird’s-Eye-View feature map and input to the at least one convolution block.11.The computer implemented method of claim 1, the raw point cloud data are obtained from a LiDAR, a camera and / or an optical sensor.12.The computer implemented method of claim 1, further comprising:assisting a task of an apparatus based on the one or more object predictions, wherein the task comprises one of autonomous driving or intelligent manufacturing.13.The computer implemented method of claim 1, further comprising:calculating a loss between the one or more object predictions and corresponding ground-truth labels; andtraining a neural network model comprising the multiple convolution blocks based on minimizing the loss.14.A vehicle, comprising:one or more sensors, configured for obtaining raw point cloud data related to environment around the vehicle;one or more controllers; andone or more storage devices storing computer-executable instructions that, when executed, cause the one or more controllers to perform the operations of the method of one of claims 1-13.15.A robot, comprising:one or more sensors, configured for obtaining raw point cloud data related to environment around the robot;one or more processors; andone or more storage devices storing computer-executable instructions that, when executed, cause the one or more processors to perform the operations of the method of one of claims 1-13.16.A computer system, comprising:one or more processors; andone or more storage devices storing computer-executable instructions that, when executed, cause the one or more processors to perform the operations of the method of one of claims 1-13.17.One or more computer readable storage medium storing computer-executable instructions that, when executed, cause one or more processors to perform the operations of the method of one of claims 1-13.18.A computer program product comprising computer-executable instructions that, when executed, cause one or more processors to perform the operations of the method of one of claims 1-13.