Point cloud detection method based on enhanced sub-manifold sparse convolution and dynamic diffusion
Through the enhancer manifold sparse convolution and dynamic feature diffusion module, the insufficient feature interaction and central area loss caused by point cloud data sparsity are solved, and the accuracy and accuracy of point cloud detection are improved.
Patent Information
- Application Number
- CN202510552595.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-29
- Publication Date
- 2025-08-01
AI Technical Summary
The sparseness of point cloud data and the execution mechanism of submanifold convolution leads to insufficient feature interaction, limited receptive field, and affects detection accuracy, especially when the features are missing in the center of the object.
The enhanced manifold sparse convolution block and dynamic feature diffusion module are used to replace the three-dimensional sub-manifold sparse convolution with the sub-manifold sparse convolution in the XOY plane and Z-axis direction, which enhances the feature interaction capability, and the dynamic feature diffusion module is used to diffuse the features from the object surface to the central area, increasing the number of features in the central area.
It significantly improves the accuracy of point cloud detection, especially in the central area of the object, and improves the detection effect of large objects or point cloud density.
Smart Images

Figure CN120411481A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of point cloud object detection, and particularly relates to a point cloud detection method based on enhancer manifold sparse convolution and dynamic diffusion. Background Art
[0002] Object detection, as a core task in computer vision, aims to automatically identify specific target objects in images or videos through algorithms and accurately label their location and category information. Traditional two-dimensional object detection realizes object localization and classification through RGB image data and has been widely applied in fields such as intelligent security and industrial quality inspection. With the growing demand for three-dimensional perception scenarios in autonomous driving, robotics, etc., three-dimensional object detection technology based on point clouds has become a research hotspot. This technology processes the discrete three-dimensional spatial point cloud data collected by lidar (LiDAR), reconstructs the geometric structure and spatial distribution of objects, and then outputs detection results with three-dimensional bounding boxes, providing high-precision distance, size, and orientation information for the environmental perception system. However, the inherent sparsity, disorder, and unstructured characteristics of point cloud data pose challenges to lidar-based object detection.
[0003] Against this backdrop, researchers have been continuously exploring new methods and technologies to overcome these challenges and drive the development of 3D object detection technology towards greater efficiency and accuracy. Early methods, such as voxelizing point cloud data and using traditional 3D convolutions for feature extraction, although demonstrating excellent detection performance, have limitations in terms of computational efficiency and resource consumption. To improve computational efficiency, columnar structure methods were proposed, which convert point clouds into columns and use 2D convolutions for processing, thus retaining spatial information while increasing processing speed, but potentially losing some information in the height dimension. Subsequently, semi-dense detectors combined the advantages of sparse and dense features, using 3D sparse convolutions to process voxelized point cloud data and 2D dense heads to complete object detection, thereby enhancing computational efficiency while maintaining three-dimensional spatial information. With the deepening of research, fully sparse detectors have gradually become mainstream due to their efficient computational characteristics. Their core lies in improving operational efficiency by eliminating redundant calculations while maintaining detection accuracy. However, such detectors still have problems: taking submanifold sparse convolution as an example, it only performs calculations when the center of the convolution kernel covers non-empty voxels. Although this mechanism significantly reduces the amount of computation and accelerates the inference process, due to the inherent sparsity of point clouds combined with the execution mechanism of submanifold convolution, it results in insufficient feature interaction and limited receptive field range, thereby compromising detection accuracy. To enhance the long-range dependence relationship between features, some studies have adopted attention mechanisms or increased the size of convolution kernels. However, using attention mechanisms leads to a quadratic increase in computational complexity, increasing the size of convolution kernels causes a cubic increase in the number of parameters, and may lead to overfitting and optimization problems. In addition, the fully sparse architecture also faces the core challenge of missing central features: limited by the observation characteristics of radar, scanned data can only capture the point cloud on the surface of objects, while the interior and central regions of objects are missing due to occlusion (as shown in Figure 1 ). To address this challenge, researchers have proposed various strategies, including introducing virtual voxels to simplify the clustering process and directly predicting from the voxels closest to the center of the object, etc.
[0004] In summary, the problems existing in the current prior art are as follows: the inherent sparsity of point clouds and the execution mechanism of submanifold convolution lead to insufficient feature interaction and limited receptive field, thus failing to effectively capture the complete features of objects and affecting detection accuracy. For some large objects or cases with low point cloud density, the point cloud in the central region of the object may be very sparse or even completely missing, resulting in the difficulty for detection heads based on the center point to work accurately, thereby affecting the final detection accuracy. Summary of the Invention
[0005] To solve the problems existing in the above prior art, the present invention proposes a point cloud detection method based on enhanced manifold sparse convolution and dynamic diffusion. The method includes: obtaining the point cloud data to be detected and performing enhancement processing on the point cloud data to be detected; inputting the enhanced point cloud data into a point cloud detection model to obtain a point cloud detection result;
[0006] Training the point cloud detection model includes: obtaining an original data set and performing enhancement processing on the data in the original data set; inputting the enhanced point cloud data into a voxel encoding module to allocate the points in the point cloud to corresponding voxels; inputting the voxelized data into a 3D backbone network for feature extraction to obtain point cloud features; inputting the point cloud features into a dynamic feature diffusion module to obtain optimized point cloud features; constructing a BEV feature map according to the optimized point cloud features; inputting the BEV feature map into a detection head network to obtain a detection result; constructing a loss function of the model according to the detection result and adjusting the parameters of the model. When the loss function converges, the training of the model is completed.
[0007] To achieve the above object, the present invention also provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it implements any one of the above point cloud detection methods based on enhanced manifold sparse convolution and dynamic diffusion.
[0008] To achieve the above object, the present invention also provides a point cloud detection device based on enhanced manifold sparse convolution and dynamic diffusion, including a processor and a memory; the memory is used to store a computer program; the processor is connected to the memory and is used to execute the computer program stored in the memory, so that the point cloud detection device based on enhanced manifold sparse convolution and dynamic diffusion executes any one of the above point cloud detection methods based on enhanced manifold sparse convolution and dynamic diffusion.
[0009] Advantages of the present invention:
[0010] The present invention introduces an enhanced manifold sparse convolution block, which effectively enhances the interaction ability between features. It replaces the three-dimensional manifold sparse convolution with a sub-manifold sparse convolution in the XOY plane and a sub-manifold sparse convolution in the Z-axis direction, significantly improving the interaction ability between local features and thus improving the detection accuracy. The present invention designs a dynamic feature diffusion module. This module can selectively diffuse features from the surface of an object to the central region according to the annotation box, increasing the number of features near the central region and alleviating the situation that the voxel distribution in the central region of an object may be relatively sparse or even missing in practical applications, thereby improving the accuracy of object detection. Description of the Drawings
[0011] Figure 1 It is a diagram of a point cloud missing example at the center of a car;
[0012] Figure 2 is the overall flowchart of the present invention;
[0013] Figure 3 is the structural diagram of the enhancer manifold sparse convolution block of the present invention;
[0014] Figure 4 is the structural diagram of the dynamic feature diffusion module of the present invention;
[0015] Figure 5 is the diagram of the hybrid diffusion process of the voxel feature and the original feature of the present invention. Specific embodiments
[0016] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0017] Aiming at the problems of insufficient feature interaction and lack of central features resulting in insufficient detection accuracy, the present invention proposes a point cloud object detection method based on enhanced submanifold sparse convolution and dynamic diffusion. First, aiming at the problem of insufficient detection accuracy caused by the limitation of feature interaction, the present invention proposes an enhanced submanifold sparse convolution block. This module replaces the three-dimensional submanifold sparse convolution with a submanifold sparse convolution on the XOY plane and a submanifold sparse convolution in the Z-axis direction. This method significantly enhances the interaction ability between features with almost no increase in the number of parameters, improving the detection accuracy. In addition, considering the problem that the voxel distribution in the central region of the object may be sparse or missing, the present invention proposes a dynamic feature diffusion module, which selectively diffuses features from the object surface to the central region according to the annotation box, increasing the number of features near the central region and further improving the detection accuracy.
[0018] In this embodiment, a point cloud detection method based on enhanced submanifold sparse convolution and dynamic diffusion includes obtaining point cloud data to be detected and performing enhancement processing on the point cloud data to be detected; inputting the enhanced point cloud data into a point cloud detection model to obtain a point cloud detection result; as Figure 2As shown in the figure, training the point cloud detection model includes: obtaining the original dataset and performing augmentation processing on the data in the original dataset; inputting the augmented point cloud data into the voxel encoding module to assign the points in the point cloud to the corresponding voxels; inputting the voxelized data into the 3D backbone network for feature extraction to obtain point cloud features; inputting the point cloud features into the dynamic feature diffusion module to obtain optimized point cloud features; constructing a BEV feature map based on the optimized point cloud features; inputting the BEV feature map into the detection head network to obtain detection results; constructing a loss function for the model based on the detection results and adjusting the parameters of the model. When the loss function converges, the training of the model is completed.
[0019] Embodiment 1
[0020] The present invention designs a point cloud detection method based on enhancer manifold sparse convolution and dynamic diffusion. Three types of targets, namely vehicles, pedestrians, and cyclists, are detected for two datasets. The entire design method includes a voxel encoding module, a 3D backbone network (including enhancer manifold sparse convolution blocks), a dynamic feature diffusion module, a BEV feature map, and a sparse detection head. The overall structure of its network is as Figure 2 shown.
[0021] In the training stage, the model is trained and evaluated on two publicly available datasets, KITTI and Waymo, which are divided into a training set and a validation set according to the default partitioning strategy in OpenPCDet. For the KITTI dataset, since only the targets in front of the acquisition vehicle are annotated in this dataset, the detection range is set to [0, 70.4] meters in the X-axis direction, [-40, 40] meters in the Y-axis direction, and [-3, 1] meters in the Z-axis direction. For the Waymo dataset, the detection range is set to [-75.2, 75.2] meters in the X-axis direction, [-72.5, 72.5] meters in the Y-axis direction, and [-2, 4] meters in the Z-axis direction.
[0022] Data augmentation (including global rotation, translation, and instance augmentation) is performed on the original dataset. Due to the limitation of video memory, the batch size is set to 32 for the KITTI dataset and 4 for the Waymo dataset.
[0023] The processed point cloud data is put into the voxel encoding module. The voxel encoding module divides the point cloud data in the three-dimensional space into regular voxel grids according to the set voxel size ([0.05, 0.05, 0.1] meters for the KITTI dataset and [0.1, 0.1, 0.15] meters for the Waymo dataset). Each voxel represents a spatial unit of a fixed size, and the points in the point cloud are assigned to the corresponding voxels.
[0024] The voxelized data is input into a 3D backbone network for feature extraction. The processing of the voxelized data in the 3D backbone network includes: performing sparse convolution processing on the data to obtain voxel features, where sparse convolution can effectively handle the sparsity of point cloud data, reduce unnecessary computational amounts, and lay a foundation for subsequent feature extraction; inputting the voxel features into an enhanced submanifold sparse convolution block to obtain point cloud features. The enhanced submanifold sparse convolution block includes a sequentially connected 1×5×5 submanifold sparse convolution layer, 3×1×1 submanifold sparse convolution layer, normalization layer, activation function, 3×3×3 submanifold sparse convolution layer, normalization layer, and activation function; it is used to extract features from the voxel features to obtain point cloud features.
[0025] In this embodiment, the structure of the enhanced submanifold sparse convolution block is as Figure 3 shown, and the data passes through two enhanced submanifold sparse convolution blocks. Each enhanced submanifold sparse convolution block replaces the 3×3×3 submanifold sparse convolution with a 1×5×5 submanifold sparse convolution on the XOY plane and a 3×1×1 submanifold sparse convolution in the Z-axis direction. The process is shown in the following formula:
[0026] f o =Conv 3×1×1 (Conv 1×5×5 (f i ))
[0027] where f o represents the output feature, f i represents the input feature, Conv 3×1×1 represents a convolution with a convolution kernel size of 3×1×1, and Conv 1×5×5 represents a convolution with a convolution kernel size of 1×5×5.
[0028] In this embodiment, the processing of the point cloud features by the dynamic feature diffusion module includes: quantitatively evaluating the importance of each voxel using a supervised learning mechanism; performing a diffusion operation on the features according to the quantitative evaluation result; and fusing the diffused features with the original features to obtain optimized point cloud features.
[0029] Specifically, as Figure 4 shown, the dynamic feature diffusion module takes the output of the 3D backbone network as input. The specific implementation includes three key steps: First, quantitatively evaluate the importance of each voxel using a supervised learning mechanism; second, based on the importance evaluation result, perform a feature diffusion operation to ensure the effective propagation of key features; finally, as Figure 5 shown, fuse the diffused features with the original features to optimize the final feature representation.
[0030] Importance Score of Predicted Voxels: The importance of voxels is determined based on their positional relationship relative to the ground truth bounding box. Specifically, voxels located within the ground truth bounding box are considered important, while those outside the bounding box are regarded as less important. This process can be formally expressed as:
[0031]
[0032] where T i,j represents the importance of the voxel located at coordinates (x i , y j ). Voxels within the bounding box are assigned an importance value of 1, indicating that they are important foreground information and crucial for predicting the object's position and size; while voxels outside the bounding box are assigned an importance value of 0, indicating that they are background information.
[0033] Considering the transformation of voxel coordinates during the downsampling process, a specific formula is needed to accurately restore these coordinates to their original positions, ensuring that even after the network performs downsampling, it can still accurately map back to the spatial positions of the original point cloud data, thereby maintaining the spatial consistency of the feature representation. The specific restoration formula is as follows:
[0034] V (x,y) = (V s(x,y) + 0.5) × s × V size + Range
[0035] where V (x,y) represents the voxel coordinates after restoration of the downsampled voxel V s(x,y) , s represents the downsampling step size, V size represents the voxel size, and Range represents the range of the point cloud.
[0036] The loss function is used to evaluate the performance of voxel importance prediction and is defined as follows:
[0037] L imp = FocalLoss(T, V)
[0038] Define the loss function for evaluating voxel importance, called importance loss. This loss function is the FocalLoss formula. Specifically, T represents the true importance label of the voxel, while V is the voxel importance score predicted by the proposed network model. By minimizing this loss function, it aims to optimize the network's prediction ability for voxel importance, thereby improving the overall performance of the model.
[0039] Example 1
[0040] As Figure 4As shown, feature diffusion is performed according to the importance score of voxels: The diffusion is divided into two steps. First, the voxel features are classified into foreground and background, and the foreground voxels are extracted. According to the positions of the foreground voxels, the importance of the voxels at the surrounding positions of the foreground voxels is predicted. When predicting the importance of the voxels at the surrounding positions of the foreground voxels, the voxels at the repeated positions are removed, and the voxels with the importance score greater than 0.8 after deduplication are restored to the original feature map. Extracting the foreground voxels includes:
[0041] V N×1 = Conv(V)
[0042] where V is the input voxel feature, and the foreground voxels are output The number of foreground voxels is N×1.
[0043]
[0044] V M×1 = Deduplicate(V N×8 )
[0045] Pos diff = POS(V M×1 > 0.8)
[0046] where V N×8 represents predicting the importance of 8 voxels around the foreground voxel according to the foreground voxel, and these importance scores are still used to calculate the loss with the ground truth through FocalLoss. Deduplicate means removing duplicate voxels according to the voxel coordinate positions. Subsequently, the positions of the voxels with the importance score greater than 0.8 are selected as the positions of the final diffused voxels, and Pos diff represents the new voxel coordinates obtained after the diffusion of the voxels with the importance score exceeding 0.8.
[0047] In this embodiment, as Figure 5 shown, the mixed diffusion of voxel features and original features: After determining the coordinates of the newly created voxels, the corresponding features are established. Assigning zero-value features to these diffused voxels may affect the effective learning ability of the network. Therefore, feature mixing is performed before outputting the new features. The formula is as follows:
[0048]
[0049] where the superscript of f represents the number of layers of feature downsampling, represents the original voxel feature map of the fourth layer, while represents the feature obtained by diffusing the fourth layer feature map. By default, is initialized as a zero feature vector. is the diffused voxel feature and the original voxel feature The result after convolution. represents the transpose of the third-layer feature with a side length of 5 centered at the diffusion position, linear represents the linear layer, · represents matrix multiplication, and softmax represents the softmax function. represents The output after feature mixing.
[0050] In this embodiment, the BEV feature map is obtained by converting the 3D feature map into a 2D bird's-eye view. For clarity of expression, s is used to represent the downsampling step size, f represents the sparse feature matrix with the shape of [N, C], where N is the number of features (i.e., the number of voxels), C is the number of channels of each feature, and I is the coordinate matrix corresponding to the feature, with the shape of [N, 3]. For a given coordinate I, the coordinates after upsampling the corresponding feature can be obtained through the formula I = I s × s. To compress the 3D voxel features into 2D sparse features, the neck module will perform an addition operation on the features corresponding to the same coordinates according to the x and y values corresponding to the coordinates:
[0051]
[0052] where The x and y coordinate values corresponding to the feature at position p are x p , y p .
[0053] The sparse detection head consists of multiple submanifold sparse convolution blocks. It is similar to the dense detection head based on the center point and outputs a heat map of the object center point, as well as the sine and cosine values of the length, width, height, and deflection angle of the object. Its corresponding loss function has four parts, and the specific loss function is as follows:
[0054] L det = L k + L size + L off + L a
[0055] where L k represents the center point loss, L size represents the loss of the target length, width, and height, L off represents the center point offset loss, and L a represents the angle loss.
[0056] In this embodiment, the detection results are evaluated, which specifically includes: for the KITTI dataset, the Average Precision (AP) is used as the performance evaluation index for each category, and the overall performance of the model is comprehensively evaluated by calculating the mean Average Precision (mAP) of all categories. The finally obtained AP value is used as the evaluation index to reflect the detection performance of the model. For the Waymo dataset, this paper uses the Average Precision (AP) and the Average Precision Weighted by Heading (APH) as the performance evaluation indexes for each category. By calculating the mean Average Precision (mAP) and the mean APH (mAPH) of all categories as the evaluation indexes to reflect the detection performance of the model. The formulas are as follows:
[0057]
[0058] Among them, p(r ′ ) represents the P-R curve at different confidence levels.
[0059] To more accurately measure the angular relationship between the Ground Truth (GT) and the predicted bounding box, the official of the Waymo dataset introduced a new evaluation index - the Average Precision Weighted by Heading (APH). Based on the calculation of the Average Precision (AP), this index takes into account the difference in the object's orientation angle. Its formula is expressed as:
[0060]
[0061] Among them, the calculation methods of h(r ′ ) and p(r ′ ) are similar, but for each true positive example, a weighted calculation is performed according to the predicted angle and the true angle, and the weighting method is Among them represents the predicted orientation angle, and θ ∈ [-π, π] represents the true orientation angle.
[0062] Embodiment 2
[0063] In this embodiment, the method of the present invention is implemented using the PyTorch and OpenPCDet frameworks through the Python programming language. The operating system used is Ubuntu 18.04, with a memory of more than 16GB, a hard disk space requirement of 2.1TB or more, and the GPU is an NVIDIA A100 GPU with a video memory of 40G.
[0064] Embodiment 3
[0065] The data augmentation techniques used in the experiment include global rotation, translation, and instance augmentation. The Adam optimizer is used to update the model parameters, with an initial learning rate of 0.0003, a weight decay coefficient of 0.01, and a momentum of 0.9. For the KITTI dataset, since only the targets in front of the acquisition vehicle are labeled in this dataset, the detection range is set to [0, 70.4] meters in the X-axis direction, [-40, 40] meters in the Y-axis direction, and [-3, 1] meters in the Z-axis direction. The voxel size is set to [0.05, 0.05, 0.1] meters for length, width, and height respectively. Each training batch contains 32 samples and is trained for 40 epochs. For the Waymo dataset, the detection range is set to [-75.2, 75.2] meters in the X-axis direction, [-72.5, 72.5] meters in the Y-axis direction, and [-2, 4] meters in the Z-axis direction. The voxel size is set to [0.1, 0.1, 0.15] meters for length, width, and height respectively. Each training batch contains 4 samples and is trained for 12 epochs.
[0066] The datasets used are the KITTI and Waymo datasets, and the division of the training set and the validation set both adopts the default division rules in the OpenPCDet open-source framework.
[0067] In an embodiment of the present invention, the present invention further includes a computer-readable storage medium, on which a computer program is stored, and when the program is executed by a processor, it implements any one of the above-mentioned point cloud detection methods based on enhancer manifold sparse convolution and dynamic diffusion.
[0068] Those of ordinary skill in the art can understand that all or part of the steps of implementing the above method embodiments can be completed by hardware related to a computer program. The aforementioned computer program can be stored in a computer-readable storage medium. When the program is executed, it executes the steps including the above method embodiments; and the aforementioned storage medium includes: various media such as ROM, RAM, magnetic disk, or optical disc that can store program codes.
[0069] A point cloud detection device based on enhancer manifold sparse convolution and dynamic diffusion, comprising a processor and a memory; the memory is used for storing a computer program; the processor is connected to the memory and is used for executing the computer program stored in the memory, so that the point cloud detection device based on enhancer manifold sparse convolution and dynamic diffusion executes any of the above-mentioned point cloud detection methods based on enhancer manifold sparse convolution and dynamic diffusion.
[0070] Specifically, the memory includes: various media such as ROM, RAM, magnetic disk, USB flash drive, memory card or optical disc that can store program codes.
[0071] Preferably, the processor may be a general-purpose processor, including a central processing unit (abbreviated as CPU), a network processor (abbreviated as NP), etc.; it may also be a digital signal processor (abbreviated as DSP), an application specific integrated circuit (abbreviated as ASIC), a field programmable gate array (abbreviated as FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components.
[0072] The above-mentioned embodiments further elaborate on the purpose, technical solution and advantages of the present invention. It should be understood that the above-mentioned embodiments are only the preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent replacements, improvements, etc. made to the present invention within the spirit and principle of the present invention shall be included in the protection scope of the present invention.
Claims
1. A point cloud detection method based on enhancer manifold sparse convolution and dynamic diffusion, characterized in that Including: Obtain the point cloud data to be detected, and perform enhancement processing on the point cloud data to be detected; Input the enhanced point cloud data into the point cloud detection model to obtain the point cloud detection result; Training the point cloud detection model includes: obtaining the original data set, and performing enhancement processing on the data in the original data set; inputting the enhanced point cloud data into the voxel encoding module, and allocating the points in the point cloud to the corresponding voxels; inputting the voxelized data into the 3D backbone network for feature extraction to obtain point cloud features; inputting the point cloud features into the dynamic feature diffusion module to obtain optimized point cloud features; constructing a BEV feature map according to the optimized point cloud features; inputting the BEV feature map into the detection head network to obtain the detection result; constructing the loss function of the model according to the detection result, adjusting the parameters of the model, and when the loss function converges, completing the training of the model.
2. The point cloud detection method based on enhancer manifold sparse convolution and dynamic diffusion according to claim 1, characterized in that, Performing enhancement processing on the data includes: performing global rotation, translation, and instance enhancement processing on the point cloud image.
3. A point cloud detection method based on enhancer manifold sparse convolution and dynamic diffusion according to claim 1, characterized in that The voxel encoding module processes the input data including: setting the voxel size, dividing the point cloud data in the three-dimensional space into regular voxel grids, each voxel representing a spatial unit of a fixed size, and allocating the points in the point cloud to the corresponding voxels.
4. A point cloud detection method based on enhancer manifold sparse convolution and dynamic diffusion according to claim 1, characterized in that, Processing the voxelized data in the 3D backbone network includes: performing sparse convolution processing on the data to obtain voxel features; inputting the voxel features into the enhanced submanifold sparse convolution block to obtain point cloud features.
5. A point cloud detection method based on enhancer manifold sparse convolution and dynamic diffusion according to claim 4, characterized in that The enhanced submanifold sparse convolution block includes a sequentially connected 1×5×5 submanifold sparse convolution layer, a 3×1×1 submanifold sparse convolution layer, a normalization layer, an activation function, a 3×3×3 submanifold sparse convolution layer, a normalization layer, and an activation function; used for feature extraction of the voxel features to obtain point cloud features.
6. A point cloud detection method based on enhancer manifold sparse convolution and dynamic diffusion according to claim 1, characterized in that, The dynamic feature diffusion module processes the point cloud features including: quantitatively evaluating the importance of each voxel using a supervised learning mechanism; performing a diffusion operation on the features according to the quantitative evaluation result; fusing the diffused features with the original features to obtain optimized point cloud features.
7. A point cloud detection method based on enhancer manifold sparse convolution and dynamic diffusion according to claim 1, characterized in that, The loss function of the model is: L det = L k + L size + L off + L a Among them, L k represents the center point loss, L size represents the loss of the target length, width, and height, L off represents the center point offset loss, L a represents the angle loss.
8. A computer-readable storage medium having a computer program stored thereon, characterized in that, The computer program is executed by a processor to implement the point cloud detection method based on enhanced submanifold sparse convolution and dynamic diffusion as described in any one of claims 1 to 7.
9. A point cloud detection device based on enhancer manifold sparse convolution and dynamic diffusion, characterized in that Including a processor and a memory; the memory is used to store a computer program; the processor is connected to the memory and is used to execute the computer program stored in the memory, so that the point cloud detection device based on enhanced submanifold sparse convolution and dynamic diffusion executes the point cloud detection method based on enhanced submanifold sparse convolution and dynamic diffusion as described in any one of claims 1 to 7.