3D Object Detection Method and System Based on Sparse Attention and Dynamic Diffusion
By adopting sparse attention and dynamic diffusion technologies in three-dimensional object detection, the shortcomings of existing methods in real-time and feature capture capabilities are solved, and efficient and high-precision three-dimensional object detection is achieved.
Patent Information
- Application Number
- CN202410910590.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-07-09
- Publication Date
- 2025-05-27
- Estimated Expiration
- 2044-07-09
AI Technical Summary
The existing voxel-based three-dimensional object detection methods perform poorly in real-time and computational efficiency, and are difficult to effectively capture long-distance dependencies and object center features, resulting in limited accuracy in identifying and positioning larger targets in complex environments.
The three-dimensional object detection method based on sparse attention and dynamic diffusion is adopted to generate the initial 3D sparse feature map through voxelization processing, combining the 3D sparse convolution backbone network and the 2D sparse convolution backbone network to enhance the capture ability of long-distance dependencies, and improve the problem of missing features in the center of the object through dynamic feature diffusion.
It improves the computing efficiency and accuracy of three-dimensional object detection, enhances the understanding of multi-object spatial layout in complex scenarios, improves the recognition ability of large-scale and long-distance objects, and improves the accuracy and robustness of object detection.
Smart Images

Figure CN118823314B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of three-dimensional object detection, and relates to the object detection of three-dimensional point clouds, specifically to a three-dimensional object detection method and system based on sparse attention and dynamic diffusion. Background Art
[0002] Three-dimensional object detection is a technology aimed at accurately identifying and locating objects from rich three-dimensional data. It is a core technology for environmental perception in autonomous driving systems. By accurately identifying and locating vehicles, pedestrians, obstacles, and road signs on the road, autonomous driving vehicles can accurately understand the surrounding environment. This process not only provides necessary environmental information but also serves as the basis for safe decision-making and control. This information helps the autonomous driving system perform precise path planning, obstacle avoidance operations, speed adjustment, and traffic behavior prediction, ensuring the safety and efficiency of driving.
[0003] Three-dimensional data acquisition sensors mainly include cameras, lidars, and millimeter-wave radars, etc. Among them, the point cloud data collected by lidars is one of the most commonly used three-dimensional data. Point clouds can capture and represent the details of objects or scenes with very high precision, which is crucial for the development of three-dimensional environmental perception.
[0004] One of the main research directions of three-dimensional point cloud object detection is the voxel-based method. Voxelization is to divide sparse and irregular point cloud data into regular three-dimensional voxel grids. This process involves quantifying the point cloud data into a predefined grid space, and each voxel represents the point cloud information inside it. Through this conversion, the original point cloud data is transformed into a more standardized and structured format, and a three-dimensional sparse convolutional network can be used to convert it into dense two-dimensional data, so that mature two-dimensional object detection methods can be used for processing.
[0005] However, such voxel-based three-dimensional object detection methods perform poorly in terms of real-time performance. Due to the sparsity and large data volume of point cloud data, when the point cloud is converted into a dense two-dimensional feature map, most regions actually do not contain any effective point cloud information, which leads to a large amount of computing resources being used to process these blank regions, resulting in a serious waste of computing power.
[0006] In addition, voxel-based methods use traditional three-dimensional convolutional networks, which mainly focus on the extraction of local features. Due to the limited receptive field of convolution, they cannot effectively obtain information from regions farther away. Therefore, they perform poorly in capturing long-range dependencies in point cloud data. This limitation restricts the ability of the method to perform accurate identification and location in complex environments, especially when there are multiple interrelated objects in the scene.
[0007] In addition, lidar scanning mainly captures the point cloud on the surface of an object, resulting in relatively few or completely missing key information inside or at the center. After voxelization and 3D convolution, the problem of missing center features is further exacerbated, which has a greater impact on the detection of larger objects and makes the method have problems in the accuracy of identifying and positioning larger targets. Summary of the Invention
[0008] The object of the present invention is to address the deficiencies in the prior art and propose a 3D object detection method and system based on sparse attention and dynamic diffusion. This method improves the computational efficiency and reduces resource consumption through a 3D object detection process that is completely based on sparse data processing. It also combines sparse attention encoding and dynamic feature diffusion to enhance the ability to capture long-range dependencies, expand the receptive field, and improve the problem of missing object center features, achieving high-precision and high-efficiency 3D object detection.
[0009] The technical solution for achieving the object of the present invention is as follows:
[0010] A 3D object detection method based on sparse attention and dynamic diffusion includes the following steps:
[0011] S1. Convert the original point cloud data into a 3D voxel grid through voxelization processing, and then extract features from each non-empty voxel containing point cloud data to obtain an initial 3D sparse feature map with a spatial structure;
[0012] S2. Input the initial 3D sparse feature map into a 3D sparse convolutional backbone network. The 3D sparse convolutional backbone network uses multiple sparse attention encoding blocks composed of dilated attention, submanifold sparse convolution, and ordinary sparse convolution, and multi-scale feature fusion to perform feature learning on the input initial 3D sparse feature map to generate a 3D sparse feature map;
[0013] S3. Perform sparse height compression processing on the 3D sparse feature map to obtain an initial 2D sparse feature map;
[0014] S4. Input the initial 2D sparse feature map into a 2D sparse convolutional backbone network. The 2D sparse convolutional backbone network performs feature enhancement and extraction on the initial 2D sparse feature map through dynamic feature diffusion and 2D sparse coding blocks to generate a 2D sparse feature map;
[0015] S5. Use the 2D sparse feature map to predict the target category and generate its bounding box.
[0016] The specific operation of voxelization in step S1 is as follows: First, the depth, height, and width of the three-dimensional point cloud space of the input point cloud data are cropped along the z-axis, y-axis, and x-axis. Then, the size of the voxel is defined, and based on the cropped three-dimensional point cloud space and the defined voxel size, the size of the generated three-dimensional voxel grid is calculated. Then, according to the position of each point in the input point cloud data in the three-dimensional point cloud space, all points are divided into the voxels they belong to. For each voxel, if the number of points contained in the voxel exceeds the threshold, random sampling is performed on the voxel to ensure that the number of points in the voxel does not exceed the specified threshold, facilitating subsequent feature extraction by the backbone network.
[0017] In step S1, features are further extracted from each non-empty voxel containing point cloud data. Specifically, for each non-empty voxel containing point cloud data, that is, a voxel with at least one point, the mean value of the contained points is calculated as the feature representation of the voxel.
[0018] In step S2, feature learning is performed on the input initial 3D sparse feature map to generate a 3D sparse feature map. Specifically, the initial 3D sparse feature map is input into a 3D sparse convolutional backbone network for feature learning. The 3D sparse convolutional backbone network first uses four sparse attention encoding blocks in sequence for feature extraction and four downsamplings. Then, multi-scale feature fusion is used to perform two downsamplings on the feature map output by the sparse attention encoding blocks, and the feature maps obtained from the two downsamplings are fused with the feature map output by the sparse attention encoding blocks. Among them, the sparse attention encoding block consists of two submanifold sparse convolutional modules using dilated attention and one ordinary sparse convolutional module using dilated attention in sequence. The submanifold sparse convolutional module is mainly responsible for feature enhancement and extraction, and the ordinary sparse convolutional module is mainly responsible for downsampling. Multi-scale feature fusion performs two downsamplings on the feature map output by the fourth sparse attention encoding block using two ordinary sparse convolutions in sequence. Then, the feature maps obtained from the two downsamplings are respectively upsampled to the size of the feature map output by the fourth sparse attention encoding block using sparse transposed convolution, and finally, the feature maps obtained from the two upsamplings are concatenated with the feature map output by the fourth sparse attention encoding block to generate a 3D sparse feature map.
[0019] The specific operation of sparse height compression in step S3 is as follows: The input 3D sparse feature map is transformed to the bird's-eye view plane, and by accumulating and summing the voxel features with the same coordinates on the bird's-eye view plane, an initial 2D sparse feature map with a two-dimensional spatial representation is generated.
[0020] In the step S4, the specific operation of enhancing and extracting features by the 2D sparse convolution backbone network through dynamic feature diffusion and 2D sparse coding blocks is as follows: First, perform voxel classification on the input initial 2D sparse feature map to predict whether the center point of each voxel belongs to a certain size category group or the background; according to the result of voxel classification, use dynamic feature diffusion to spread the voxel features belonging to large objects to large areas, spread the voxel features belonging to small objects to small areas, and do not perform feature diffusion on the points belonging to the background; finally, use 2D sparse convolution blocks to extract features from the 2D sparse feature map after feature diffusion to generate the final 2D sparse feature map for object detection.
[0021] In the step S5, the specific operation of predicting the object category and generating its bounding box using the 2D sparse feature map is as follows: Predict the probability that each voxel in the 2D sparse feature map belongs to a certain type of object to generate a heat map with the same number of channels as the number of categories; use sparse max pooling to extract the voxels with the highest probability from the heat map; for each extracted voxel with the highest probability, further regress the position offset, height, 3D size, and rotation angle of the bounding box.
[0022] A three-dimensional object detection system based on sparse attention and dynamic diffusion, which is used for the above-mentioned three-dimensional object detection method based on sparse attention and dynamic diffusion. The system includes:
[0023] Point cloud voxelization module: This module is used to convert the input point cloud data into an initial 3D sparse feature map. First, divide the cropped point cloud into three-dimensional voxel grids through the defined voxel size, then divide each point in the input point cloud data into the voxel it belongs to, and finally extract features from the points contained in each voxel to obtain the 3D sparse feature map;
[0024] 3D sparse backbone network module: This module is used to extract 3D sparse features from the initial 3D sparse feature map. First, use four sparse attention coding blocks to extract features from the initial 3D sparse feature map and gradually downsample, and then use multi-scale feature fusion to perform two downsamplings on the feature map output by the fourth sparse attention coding block and then splice it with the feature map output by the fourth sparse attention coding through an upsampling operation to generate the 3D sparse feature map;
[0025] 2D sparse backbone network module: This module is used to process the initial 2D sparse feature map after compressing the 3D sparse feature map to generate a 2D sparse feature map for object detection. First, classify the voxels in the initial 2D sparse feature map to find the voxels that need to be diffused, then dynamically spread the features of the voxels that need to be diffused to adjacent regions according to the category size of the classification, and finally use 2D sparse convolution blocks to extract features from the diffused feature map;
[0026] Sparse Detection Head Module: This module uses a 2D sparse feature map for final object classification and bounding box regression. This module employs sparse max pooling to select the voxel with the highest probability from the classification heatmap, and then regresses the bounding box size, position offset, height, and orientation for the selected voxel with the highest probability.
[0027] Advantages or Beneficial Effects of the Present Application:
[0028] This technical solution is entirely based on the method of sparse data processing, reducing unnecessary computational resources and memory consumption during the three-dimensional object detection process while maintaining data sparsity and improving computational efficiency. It can meet application scenarios with high real-time requirements, such as autonomous driving and robot navigation.
[0029] This technical solution uses a 3D sparse convolutional backbone network composed of multiple sparse attention encoding blocks and multi-scale feature fusion to enhance the ability to capture long-range spatial relationships and expand the receptive field, which can help applications such as autonomous driving and robot navigation understand the spatial layout of multiple objects in complex scenarios and enhance the recognition ability of large-scale and distant objects.
[0030] This technical solution uses a 2D sparse convolutional backbone network composed of dynamic feature diffusion and 2D sparse convolutional blocks to improve the problem of missing object center features in the two-dimensional sparse feature map. While maintaining feature sparsity, it enhances the local feature expression around key voxels, helps the three-dimensional object detection system better recognize objects of different sizes, and improves the accuracy and robustness of object detection. Description of the Drawings
[0031] Figure 1 It is the overall structure diagram of the embodiment;
[0032] Figure 2 It is the module diagram of the 3D sparse convolutional backbone network of the embodiment;
[0033] Figure 3 It is the module diagram of the 2D sparse convolutional backbone network of the embodiment. Detailed Embodiment
[0034] The following further describes the present invention in detail in conjunction with the drawings and specific embodiments, but it is not a limitation of the present invention.
[0035] Embodiment:
[0036] Referring to Figure 1 , a three-dimensional object detection method based on sparse attention and dynamic diffusion includes the following steps:
[0037] S1. Convert the original point cloud data into a three-dimensional voxel grid through voxelization processing, and then extract features for each non-empty voxel containing point cloud data to obtain an initial 3D sparse feature map with spatial structure;
[0038] The voxelization processing operation is as follows: First, clip the depth, height, and width of the three-dimensional point cloud space of the input point cloud data along the z-axis, y-axis, and x-axis to D, H, and W respectively. Define the sizes of the depth, height, and width of the voxel as v D , v H and v W . The size of the generated three-dimensional voxel grid can be obtained as D′×H′×W′, where the depth, height, and width of the three-dimensional voxel grid are calculated by the following formulas respectively: Then, according to the position of each voxel in space and the area it contains with size v D ×v H ×v W , divide each point in the input point cloud data into the voxel it belongs to. It is stipulated that the maximum number of points contained in each non-empty voxel is T, and randomly sample T points from the voxels with more than T points;
[0039] For each non-empty voxel containing point cloud data (i.e., a voxel with at least one point), calculate the feature mean of these T points as the feature representation of the voxel;
[0040] S2. Input the initial 3D sparse feature map into a 3D sparse convolutional backbone network. The 3D sparse convolutional backbone network uses multiple sparse attention encoding blocks composed of dilated attention, submanifold sparse convolution, and ordinary sparse convolution, and multi-scale feature fusion to perform feature learning on the input initial 3D sparse feature map to generate a 3D sparse feature map;
[0041] As Figure 2 shown, the 3D sparse convolutional backbone network consists of four sparse attention encoding blocks and multi-scale feature fusion. Each sparse attention encoding block includes two submanifold convolution modules using the dilated attention mechanism and one ordinary sparse convolution module using the dilated attention mechanism. The multi-scale feature fusion performs two downsamplings on the features output by the sparse attention encoding blocks and concatenates them with the features output by the sparse attention encoding blocks to generate the final 3D sparse feature map;
[0042] The 3D sparse convolutional backbone network first uses four sparse attention encoding blocks to gradually perform 1×, 2×, 4×, and 8× downsamplings on the input initial 3D sparse feature map;
[0043] In the sparse attention encoding block, before each sparse convolution, dilated attention is used to enhance the perception range of features, capture important information from a relatively long distance, and dilated attention is achieved by defining multiple different search ranges on the input 3D sparse feature map. The specific implementation steps are as follows:
[0044] 2.1) Define voxels and features: Define each non-empty voxel on the input 3D sparse feature map as a query voxel v i , each query voxel v i has its own feature f i and position P i ,
[0045] 2.2) Select voxels to be attended to: Define a dilated attention range to determine the influence of other voxels (referred to as "attended voxels") on the feature update of the query voxel v i within the dilated attention range centered on the query voxel v i . Dilated attention is achieved by defining multiple different dilated attention ranges for the query voxel v i , and each range is called a "layer". These layers are defined by three parameters:
[0046] The starting search range of the m-th dilated layer;
[0047] The ending search range of the m-th dilated layer;
[0048] The search step size of the m-th dilated layer;
[0049] The dilated attention range is the set of all defined layers. For a given query voxel v i , the index set of all attended voxels within its dilated attention range is obtained through the following formula:
[0050]
[0051] where, represents returning the indices of all voxels from a to b with a step size of s;
[0052] 2.3) Calculate the embeddings of queries, keys, and values: For each query voxel v i , calculate three embeddings, namely query embedding Q i , key embedding K j and value embedding V j , and the specific calculation formulas are as follows: Query embedding Q i = f i W q , where f i is the feature of the voxel, and Wq is the weight matrix of the query; the key embedding K j = f j W k + E pos ; the value embedding V j = f j W v + E pos , where f j is the feature of other voxels, W k and W v are the weight matrices of the key and the value respectively, and E pos is the positional encoding, calculated as follows: E pos = (p i - p j )W pos ;
[0053] 2.4) Calculate the attention scores and update the voxel features: The attention scores are obtained by the dot product of the query embedding and the key embedding. The new feature i of each query voxel v is calculated by the following formula:
[0054]
[0055] where Ω(i) is the set of voxels to be attended to, and exp is the exponential function used to convert the dot product result to an integer for normalization, is to mitigate the impact of the dot product result being too large at higher dimensions;
[0056] In the sparse attention coding block, the initial 3D sparse feature map is processed by sequentially applying two submanifold convolutions using the dilated attention mechanism and one ordinary sparse convolution using the dilated attention mechanism. Among them, the submanifold sparse convolution is only performed when the center of the convolution kernel coincides with non-empty voxels to maintain the sparsity of the features. The ordinary sparse convolution is responsible for performing the downsampling of each sparse attention coding block, and its operations are only performed on non-empty voxels.
[0057] The 3D sparse convolutional backbone network then uses multi-scale feature fusion to perform 16× and 32× downsamplings on the features output by the sparse attention coding block to further expand the receptive field, and uses sparse deconvolution to upsample the feature maps obtained by the 16× and 32× downsamplings to the 8× feature map size and splice the features with the 8× feature map to generate the final 3D sparse feature map;
[0058] S3. Pass the 3D sparse feature map through sparse height compression to obtain the initial 2D sparse feature map;
[0059] The specific implementation steps of sparse height compression are as follows:
[0060] 3.1) Convert the 3D sparse feature map to the bird's-eye view plane: For the voxels on the input 3D sparse feature map, only retain the x and y coordinates of each voxel in the horizontal coordinate system, and ignore its z coordinate;
[0061] 3.2) Height compression: For all voxels with the same position in the x and y coordinates, that is, those voxels that coincide in the two-dimensional plane, add their feature values;
[0062] By adding the features of the voxels located at the same two-dimensional coordinate point, the 3D sparse feature map is compressed into a 2D sparse feature map while maintaining sparsity;
[0063] S4. Input the initial 2D sparse feature map into the 2D sparse convolutional backbone network. The 2D sparse convolutional backbone network enhances and extracts the features of the initial 2D sparse feature map through dynamic feature diffusion and 2D sparse coding blocks, generating a 2D sparse feature map;
[0064] As Figure 3 shown, the 2D sparse backbone network first performs voxel classification on the input initial 2D sparse feature map, determines which voxels should undergo feature diffusion based on the classification results, then determines the diffusion region, and finally uses a 2D sparse convolutional block to extract and enhance the features of the 2D sparse feature map after feature diffusion, generating the final 2D sparse feature map;
[0065] The 2D sparse convolutional backbone network first uses dynamic feature diffusion to diffuse the features of the voxels belonging to large objects in the input initial 2D sparse feature map to a larger region, the features of the voxels belonging to small objects to a smaller region, and the points belonging to the background do not undergo feature diffusion. The specific implementation steps are as follows:
[0066] 4.1) Predict the size category of voxels: Assume there are N non-empty voxels in the initial 2D sparse feature map, and use a voxel classification method to predict the probability P that all non-empty voxels N belong to a certain size category group i. i The specific implementation steps of this voxel classification method are as follows:
[0067] 4.1.1) Group according to object category size: Assume there are O object categories in this three-dimensional object detection task, divide the object categories with similar sizes into one group, and a total of G size category groups are divided. For example, the small object group includes pedestrians and bicycles, and the large object group includes cars, etc.;
[0068] 4.1.2) Generate true labels: For the G size category groups, use the labeled true bounding boxes to generate a true label T of length N for each size category group i ∈ G(i). i For voxel j ∈ N(j), if the voxel center (x j , y jIf it is within the true bounding box of an object in a certain size category group i, then Otherwise
[0069] 4.1.3) Voxel Classification Prediction: Train a three-layer MLP network with a sigmoid function to predict a vector P of length N for size category group i ∈ G(i) i , representing the probability that all non-empty voxels N belong to size category group i, and the predicted probability value is between [0, 1];
[0070] 4.1.4) Loss Function: Use sigmoid focal loss to calculate the loss between the predicted value P i and the true label T i The loss is as follows:
[0071]
[0072] 4.2) Determine the Size Category of Voxels: For the probability P that all non-empty voxels N obtained through step 4.1) belong to size category group i i , use a binary mask M of length N i to indicate which voxels belong to size category group i. Define a probability threshold t. For voxel j ∈ N(j), if is greater than the probability threshold t, then mark the corresponding position of the binary mask M i as 1, otherwise mark it as 0;
[0073] 4.3) Define the Diffusion Region: Diffuse the voxels marked as 1 in the binary mask M i . The diffusion range is determined according to the average size of the objects in size category group i. First, define the diffusion convolution kernel size K i . The formula is as follows:
[0074] K i = α · S i
[0075] where α is a coefficient controlling the diffusion range, and S i is the average size of the objects in size category group i;
[0076] Then calculate the diffusion region R i , R i is calculated based on the binary mask M i and the diffusion convolution kernel K i . R i determines which regions are diffused. In this example, each voxel with a value of 1 in M i is diffused to the surrounding K i ×Ki area;
[0077] 4.4) Integrate the diffusion regions of all size categories: Combine the diffusion regions R of G groups of size categories i into a total diffusion region R;
[0078] The 2D sparse convolutional backbone network then uses a 2D sparse convolutional block to further extract features from the 2D sparse feature map after dynamic feature diffusion to generate the final 2D sparse feature map;
[0079] S5. Use the 2D sparse feature map to predict the target category and generate its bounding box.
[0080] For the O object categories in the 3D object detection task of this example, the specific classification and regression steps are as follows:
[0081] 5.1) Generate a heatmap: First, use the 2D sparse feature map to generate a heatmap. The number of channels of the heatmap is the number of categories O, and the value of each voxel on the heatmap represents the probability that the voxel is an object of a certain category;
[0082] 5.2) Voxel selection: Use sparse max pooling to extract the voxel with the highest probability from the heatmap;
[0083] 5.3) Bounding box regression: The extracted voxels are used as candidate centers for object detection. For each selected voxel, a 3×3 submanifold sparse convolutional layer and a fully connected layer are used to further predict the specific parameters of the bounding box, including position offsets (Δx, Δy), height h, 3D size s, and rotation angle α. These parameters are directly predicted from the voxel features through a regression task.
[0084] A 3D object detection system based on sparse attention and dynamic diffusion, for the above 3D object detection method based on sparse attention and dynamic diffusion, the system includes:
[0085] Point cloud voxelization module: This module is used to convert the input point cloud data into an initial 3D sparse feature map. First, divide the cropped point cloud into a 3D voxel grid through the defined voxel size, then divide each point in the input point cloud data into the voxel it belongs to, and finally extract features from the points contained in each voxel to obtain a 3D sparse feature map;
[0086] 3D sparse backbone network module: This module is used to extract 3D sparse features from the initial 3D sparse feature map. First, use four sparse attention encoding blocks to extract features from the initial 3D sparse feature map and gradually downsample, then use multi-scale feature fusion to downsample the feature map output by the fourth sparse attention encoding block twice and then splice it with the feature map output by the fourth sparse attention encoding block through an upsampling operation to generate a 3D sparse feature map;
[0087] 2D Sparse Backbone Network Module: This module is used to process the initial 2D sparse feature map compressed from the 3D sparse feature map and generate a 2D sparse feature map for object detection. First, classify the voxels in the initial 2D sparse feature map to find the voxels that need to be diffused, then dynamically diffuse the features of the voxels that need to be diffused to adjacent regions according to the size of the classification category, and finally use a 2D sparse convolution block to extract features from the diffused feature map;
[0088] Sparse Detection Head Module: This module uses the 2D sparse feature map for final object classification and bounding box regression. This module uses sparse max pooling to select the voxels with the highest probability from the classification heat map, and then regresses the bounding box size, position offset, height, and orientation of the features of the selected voxels with the highest probability.
Claims
1. A three-dimensional object detection method based on sparse attention and dynamic diffusion, characterized in that: The steps include: S1, converting the original point cloud data into a three-dimensional voxel grid through voxelization, and then extracting features from each non-empty voxel containing point cloud data to obtain an initial 3D sparse feature map with spatial structure; S2, input the initial 3D sparse feature map into the 3D sparse convolution backbone network, the 3D sparse convolution backbone network uses multiple sparse attention encoding blocks composed of dilated attention, submanifold sparse convolution and ordinary sparse convolution and multi-scale feature fusion to perform feature learning on the input initial 3D sparse feature map to generate a 3D sparse feature map; S3, the 3D sparse feature map is processed by sparse height compression to obtain an initial 2D sparse feature map; S4, inputting the initial 2D sparse feature map into the 2D sparse convolution backbone network, the 2D sparse convolution backbone network performs feature enhancement and extraction on the initial 2D sparse feature map through dynamic feature diffusion and 2D sparse coding blocks to generate a 2D sparse feature map; S5. Use the 2D sparse feature map to predict the target category and generate its bounding box.
2. The three-dimensional object detection method based on sparse attention and dynamic diffusion according to claim 1, characterized in that: The specific operation of the voxelization processing in step S1 is: first, the depth, height and width of the three-dimensional point cloud space of the input point cloud data are cropped along the z-axis, y-axis and x-axis respectively, and then the size of the voxel is defined, and the size of the generated three-dimensional voxel grid is calculated according to the cropped three-dimensional point cloud space and the size of the defined voxel; then, according to the position of each point in the input point cloud data in the three-dimensional point cloud space, all points are divided into the corresponding voxels; for each voxel, if the number of points contained in the voxel exceeds the threshold, the voxel is randomly sampled to ensure that the number of points in the voxel does not exceed the specified threshold, so as to facilitate the subsequent feature extraction of the backbone network.
3. The three-dimensional object detection method based on sparse attention and dynamic diffusion according to claim 1, characterized in that: In the step S1, features are extracted for each non-empty voxel containing point cloud data. Specifically, for each non-empty voxel containing point cloud data, that is, a voxel containing at least one point, the mean value of the contained points is calculated as the feature representation of the voxel.
4. The three-dimensional object detection method based on sparse attention and dynamic diffusion according to claim 1, characterized in that: In the step S2, feature learning is performed on the input initial 3D sparse feature map to generate a 3D sparse feature map, specifically: the initial 3D sparse feature map is input into the 3D sparse convolution backbone network for feature learning, the 3D sparse convolution backbone network first uses four sparse attention coding blocks in sequence to perform feature extraction and four downsampling, then uses multi-scale feature fusion to perform two downsampling on the feature map output by the sparse attention coding block, and performs feature fusion on the twice downsampled feature map and the feature map output by the sparse attention coding block; The sparse attention coding block is composed of two sub-manifold sparse convolution modules using dilated attention and an ordinary sparse convolution module using dilated attention. The sub-manifold sparse convolution module is responsible for feature enhancement and extraction, and the ordinary sparse convolution module is responsible for downsampling. The multi-scale feature fusion downsamples the feature map output by the fourth sparse attention coding block twice using two ordinary sparse convolutions, and then uses sparse deconvolution to upsample the feature maps obtained by the two downsamplings to the size of the feature map output by the fourth sparse attention coding block. Finally, the feature map obtained by the two upsamplings is spliced with the feature map output by the fourth sparse attention coding block to generate a 3D sparse feature map.
5. The three-dimensional object detection method based on sparse attention and dynamic diffusion according to claim 1, characterized in that: The specific operation of the sparse height compression processing in step S3 is: converting the input 3D sparse feature map to the bird's-eye view plane, and generating an initial 2D sparse feature map with a two-dimensional spatial representation by accumulating and summing the voxel features with the same coordinates on the bird's-eye view plane.
6. The three-dimensional object detection method based on sparse attention and dynamic diffusion according to claim 1, characterized in that: The specific operation of the 2D sparse convolution backbone network in step S4 for feature enhancement and extraction through dynamic feature diffusion and 2D sparse coding blocks is as follows: first, voxel classification is performed on the input initial 2D sparse feature map to predict whether the center point of each voxel belongs to a certain size category group or background; according to the result of voxel classification, dynamic feature diffusion is used to diffuse the voxel features belonging to large objects to a large area, and the voxel features belonging to small objects to a small area, and no feature diffusion is performed on the points belonging to the background; finally, feature extraction is performed on the 2D sparse feature map after feature diffusion through the 2D sparse convolution block to generate the final 2D sparse feature map for target detection.
7. The three-dimensional object detection method based on sparse attention and dynamic diffusion according to claim 1, characterized in that: The specific operations of using the 2D sparse feature map to predict the target category and generate its bounding box in step S5 are: predicting the probability of each voxel in the 2D sparse feature map belonging to a certain type of object, generating a heat map with the same number of channels as the number of categories; using sparse maximum pooling to extract the voxels with the highest probability from the heat map; for each extracted voxel with the highest probability, further regressing the position offset, height, 3D size and rotation angle of the bounding box.
8. A three-dimensional object detection system based on sparse attention and dynamic diffusion, characterized in that: The system is used for the three-dimensional target detection method based on sparse attention and dynamic diffusion according to any one of claims 1 to 7, and the system comprises: Point cloud voxelization module: This module is used to convert the input point cloud data into the initial 3D sparse feature map. First, the cropped point cloud is divided into a three-dimensional voxel grid by the defined voxel size. Then each point in the input point cloud data is divided into the voxel to which it belongs. Finally, the features of the points contained in each voxel are extracted to obtain the 3D sparse feature map. 3D sparse backbone network module: This module is used to extract 3D sparse features from the initial 3D sparse feature map. First, four sparse attention coding blocks are used to extract features from the initial 3D sparse feature map and gradually downsample it. Then, multi-scale feature fusion is used to downsample the feature map output by the fourth sparse attention coding block twice, and then the feature map output by the fourth sparse attention coding block is spliced with the feature map output by the fourth sparse attention coding block through an upsampling operation to generate a 3D sparse feature map. 2D sparse backbone network module: This module is used to process the initial 2D sparse feature map after compressing the 3D sparse feature map to generate a 2D sparse feature map for target detection. First, the voxels in the initial 2D sparse feature map are classified to find the voxels that need to be diffused, and then the features of the voxels that need to be diffused are dynamically diffused to the adjacent area according to the classified category size. Finally, the 2D sparse convolution block is used to extract features from the diffused feature map; Sparse detection head module: This module uses 2D sparse feature maps for final target classification and bounding box regression. The module uses sparse maximum pooling to select the voxels with the highest probability from the classification heat map, and then regresses the bounding box size, position offset, height and direction of the selected voxel features with the highest probability.
Citation Information
Patent Citations
Transform-based time sequence point cloud three-dimensional target detection
CN116740424A
Three-dimensional target detection method for capturing ground plane
CN117542041A