Point cloud data enhancement method and computer readable storage medium
By adopting data perception and multi-grained enhancement strategies in the three-dimensional object detection network, the point cloud data enhancement method is dynamically adjusted, which solves the problem of low prediction accuracy caused by insufficient samples, and improves the prediction accuracy and generalization ability of the model.
Patent Information
- Application Number
- CN202311821035.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-12-27
- Publication Date
- 2025-07-04
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
The existing three-dimensional object detection network has low prediction accuracy due to insufficient samples, and the existing point cloud enhancement technology has limitations in terms of enhancement mode, enhancement strategy and enhancement mechanism.
A point cloud data augmentation method is proposed, which uses a deep neural network model to perform data perception, dynamically adjust the enhancement strategy, and adopts a hybrid enhancement strategy combined with multiple particle sizes to construct a diverse enhancement sample, including the movement of point particle size, the movement of object particle size and the rotation of object particle size, and applies it to the target object that needs to be enhanced.
It improves the prediction accuracy of the three-dimensional object detection model, improves the training effect of the model, and enhances the generalization ability of the model.
Smart Images

Figure FT_1 
Figure FT_2 
Figure FT_3
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer vision technology, and in particular, to a point cloud data enhancement method and a computer-readable storage medium. Background Art
[0002] With the rapid development of the field of artificial intelligence represented by deep learning, scenarios such as autonomous driving and intelligent transportation have gradually moved from theory to reality. In this process, 3D object detection technology plays a key role as a basic technology. However, the performance of 3D object detectors usually highly depends on the scale and quality of training data. On the one hand, in the real world, the acquisition process of point cloud images requires high-precision dedicated equipment and professional personnel; on the other hand, since traffic conditions can vary greatly due to geographical location, time lapse, or even weather changes, in order to improve the robustness of the model, it is usually necessary to collect point cloud images of the target multiple times in different locations, at different times, and under different weather conditions for a long period, which further increases the research burden of relevant institutions. Therefore, at a certain cost, how to transform limited data through techniques such as data augmentation to obtain more abundant training data, so as to help the model learn more knowledge more effectively and improve the accuracy of the model (such as detection accuracy), has gradually attracted the attention of more and more researchers.
[0003] Although the above data augmentation has been widely applied in the field of 2D images, such as image recognition, 2D object detection, instance segmentation, etc., in the field of 3D point clouds, data augmentation techniques have not been widely studied, and existing point cloud augmentation techniques still have certain limitations in terms of augmentation mode, augmentation strategy, augmentation mechanism, etc.
[0004] In order to help the deep neural network learn more useful knowledge and thus improve the prediction accuracy of the 3D object detection model, this paper proposes a point cloud data augmentation method for 3D object detection tasks. Summary of the Invention
[0005] The object of the present invention is to propose a point cloud data augmentation method for 3D object detection tasks, so as to solve the problem of low prediction accuracy of the current 3D object detection network due to insufficient samples.
[0006] To achieve the above object, the present application proposes a point cloud data augmentation method, including:
[0007] Input the collected point cloud data into the trained deep neural network model to obtain enhanced point cloud data. Among them, the deep neural network model is used to obtain a data evaluation result at the category granularity based on the input point cloud; select the target object to be enhanced from the target objects of the original sample according to the evaluation result; apply a hybrid enhancement strategy combining multiple granularities to the target to be enhanced to obtain enhanced point cloud data.
[0008] In one alternative embodiment, the deep neural network model includes a feature extractor and a detector. The feature extractor adopts a network structure based on convolutional layers, and the detector adopts a network structure based on fully connected layers. The loss function of the deep neural network model includes two parts: classification loss and regression loss.
[0009] In one alternative embodiment, the method further includes:
[0010] Construct a deep neural network model, input the original point cloud data into the model, and perform multiple forward-backward propagation iterative calculations to obtain the model after gradient descent;
[0011] Infer the deep neural network model on a small number of validation set point cloud samples to obtain the inference result of the model;
[0012] Perform data perception and calculate the data evaluation result at the category granularity;
[0013] Select the target object to be enhanced from the target objects of the original sample according to the evaluation result;
[0014] Apply a hybrid enhancement strategy combining multiple granularities to the target to be enhanced to obtain enhanced point cloud data for the model iteration process of the next cycle.
[0015] Repeat the above data enhancement steps until the training is completed.
[0016] In one alternative embodiment, the step of inferring the deep neural network model on a small number of validation set point cloud samples to obtain the inference result of the model includes: freezing the parameters of the deep neural network model and post-processing the inference result.
[0017] In one alternative embodiment, the post-processing of the inference result includes:
[0018] Calculate the average prediction accuracy corresponding to each category of validation set samples respectively.
[0019] In one alternative embodiment, performing data perception to obtain the model evaluation result includes:
[0020] Perform perception based on category granularity, invert the average accuracy of the inference results of the model on the validation set samples, and perform normalization calculation based on category granularity to obtain the enhanced probabilities of various category target objects in the original samples.
[0021] In one optional embodiment, selecting the target objects that need to be enhanced from the target objects of the original samples according to the evaluation results includes:
[0022] Sample the target objects in each sample based on the enhancement probabilities of each category in the data evaluation results, and regard the sampled target objects as the target objects that need to be enhanced.
[0023] In one optional embodiment, the hybrid enhancement strategy is composed of enhancement methods of multiple granularities, including point granularity movement, object granularity movement, and object granularity rotation; the enhancement methods of multiple granularities are applied to the target objects that need to be enhanced in a random sampling manner.
[0024] In one optional embodiment, the method further includes: determining whether to stop the data enhancement process by whether the prediction accuracy of the model on the validation data stops increasing. After the data enhancement process no longer proceeds, regard the model as having been trained.
[0025] This application also proposes a computer-readable storage medium. The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, it implements the method proposed in any embodiment of this application. Compared with the prior art, the present invention has the following advantages and technical effects:
[0026] The present invention can judge the learning situation of the model for different data at different training stages of the model, and drive the dynamic adjustment of the data enhancement strategy according to the perception results, abandoning the constant enhancement mode in the prior art during the entire training process. At the same time, due to the high learning difficulty caused by the complexity of the point cloud itself and the natural receptive field characteristics of the convolutional operation in the deep neural network, on the basis of data perception, the present invention also adopts an enhancement strategy combining multiple granularities to enhance the generalization ability of the model at different granularities, thereby more effectively improving the final model prediction accuracy. Brief Description of the Drawings
[0027] The drawings constituting a part of this application are used to provide a further understanding of this application. The schematic embodiments of this application and their descriptions are used to explain this application and do not constitute an improper limitation to this application. In the drawings:
[0028] Figure 1 It is a schematic structural diagram of the method of the present invention;
[0029] Figure 2Schematic diagram of the data perception process in the method of the present invention;
[0030] Figure 3 Schematic structural diagram of a point cloud enhancement device provided by an embodiment of the present application;
[0031] Figure 4 Schematic structural diagram of a terminal device provided by an embodiment of the present application. Detailed implementation manners
[0032] It should be noted that, without conflict, the embodiments in the present application and the features in the embodiments may be combined with each other. The present application will be described in detail below with reference to the drawings and in conjunction with the embodiments.
[0033] It should be noted that the steps shown in the flowchart of the drawings can be executed in a computer system such as a set of computer-executable instructions, and although the logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in a different order than here.
[0034] The present invention proposes a point cloud data enhancement method, which includes: inputting the collected point cloud data into a trained deep neural network model to obtain enhanced point cloud data, where the deep neural network model is used to obtain a data evaluation result at the category granularity based on the input point cloud; selecting the target object to be enhanced in the target object of the original sample according to the evaluation result; using a hybrid enhancement strategy combining multiple granularities to be applied to the target to be enhanced to obtain enhanced point cloud data.
[0035] This method first collects the inference results of the model on a small number of validation samples during the iterative training of the model; secondly, analyzes the inference results through the data perception process to obtain the learning situation of the model for each category and the enhancement probability of each category of target objects in the corresponding original samples; finally, determines the target objects to be enhanced, and uses a hybrid enhancement strategy combining multiple granularity enhancement methods to be applied to the target objects of each category to achieve point cloud data enhancement. For a three-dimensional object detection network model, this enhancement method can dynamically adjust the enhancement strategy according to the learning situation of the model, and use a multi-granularity enhancement strategy to construct more diverse enhanced samples, helping the model learn richer knowledge, improving the training effect of the model, and enhancing the final prediction accuracy.
[0036] Before using the deep neural network model to process data, network construction and model training are usually required. The deep neural network model in the present application may include a feature extractor and a detector. The feature extractor adopts a network structure based on convolutional layers, the detector adopts a network structure based on fully connected layers, and the loss function of the deep neural network model includes two parts: classification loss and regression loss.
[0037] In one embodiment, with reference to Figure 1 and Figure 2 , the present application proposes a method for training the above-mentioned deep neural network model, and the method includes:
[0038] Step 1: Construct a deep neural network model, input the original point cloud data into the model, and perform multiple forward-backward propagation iterative calculations to obtain the model after gradient descent;
[0039] Step 1.1: Construct a basic convolutional layer, fully connected layer, and attention mechanism layer, then combine multiple network layers to obtain a feature extraction backbone network part and a detection head part, and connect the backbone network with the detection head to obtain a three-dimensional object detection network model;
[0040] Step 1.2: Sample from the training samples in the dataset, sample 2 each time as a data batch, input it into the model obtained in Step 1.1 for forward inference, and compare the inference result with the true annotation. For the classification result, use the binary cross-entropy algorithm to obtain the classification loss value, and for the localization result, use the smooth L1 algorithm to obtain the regression loss value. Sum the loss values and perform backpropagation to calculate the gradients of each weight in the model, and perform gradient descent to complete one iteration. Repeat Step 1.2 to complete multiple iterations;
[0041] Step 2: Input a small number of samples for verification in the dataset into the model obtained in Step 1 for inference to obtain the inference result. Among them, inferring the deep neural network model on a small number of validation set point cloud samples to obtain the inference result of the model may include: freezing the parameters of the deep neural network model and post-processing the inference result. The post-processing method may be: calculating the average prediction accuracy corresponding to each category of validation set samples respectively.
[0042] Compare the inference result with the true annotation, and statistically obtain the inference accuracies {mAP1, mAP2,..., mAP6} of 6 categories (respectively {c1, c2,..., c6}) in the dataset, where 0 ≤ mAP i ≤ 1, i ∈ {1, 2,..., 6};
[0043] Step 3: Perform data perception to calculate the data evaluation result at the category granularity. Specifically, it can be: perform perception based on the category granularity, invert the average accuracy of the inference result of the model on the validation set samples, and perform normalization calculation at the category granularity to obtain the enhanced probabilities of various category target objects in the original samples. More specifically, use the inference result of the model on the validation data as the learning situation of the model for each category of data currently, and perform data perception at the category granularity on the inference result obtained in Step 2, such asFigure 2 As shown, the data evaluation result is obtained;
[0044] Step 3.1: Perform a subtraction operation on the inference accuracy rate obtained in Step 2 to obtain the inference error rate for each category {1 - mAP1, 1 - mAP2,..., 1 - mAP6};
[0045] Step 3.2: Normalize the error rate for each category based on the SoftMax algorithm, and use the normalized result as the enhanced probability of the target object for each category {P1, P2,..., P6};
[0046] Step 3.3: Obtain the maximum value among the enhanced probabilities for each category. When the maximum value is greater than or equal to the threshold τ = 0.9, go to Step 4; otherwise, go to Step 3.4;
[0047] Step 3.4: Calculate the multiple of the preset enhanced probability threshold τ relative to the current maximum probability value, that is, τ / max(P i ), and synchronously amplify the enhanced probabilities for each category according to the above multiple to obtain the updated {P1, P2,..., P6},
[0048] Step 4: Select the target objects to be enhanced from the target objects of the original samples according to the evaluation result. The target objects in each sample can be sampled based on the enhanced probability for each category in the data evaluation result, and the sampled target objects are regarded as the target objects to be enhanced. Specifically, traverse each point cloud image sample used for training in the dataset. For each sample, traverse each target object according to the true annotation. When the category of the target object is the i-th category c i , its probability of being enhanced is P i , where i ∈ {1, 2,..., 6}. Then, obtain a random variable X based on the uniform distribution U(0, 1). When X ≤ P i holds, add the target object to the set S aug of objects to be enhanced.
[0049] Step 5: Apply the hybrid enhancement strategy combining multiple granularities to the target objects to be enhanced. The hybrid enhancement strategy consists of enhancement methods with multiple granularities, including point - level movement, object - level movement, and object - level rotation; the multiple granularity enhancement methods are applied to the target objects to be enhanced in a randomly sampled manner
[0050] Step 5.1: Randomly select one method from 3 different granularity enhancement methods. The 3 methods include:
[0051] 1) Point movement. This enhancement method belongs to the "point" granularity. First, find all the points D that make up the target object according to the annotated bounding boxi where \(i\in\{1, 2, \ldots, M\}\), and \(M\) represents the number of points the object has. The coordinates of this point in the coordinate system established by the \(x\), \(y\), and \(z\) axes are Next, perform a random displacement on this point along the three axes respectively to obtain point \(D\). i The new coordinates of
[0052]
[0053]
[0054]
[0055] where \(R\) represents the maximum displacement distance of \(0.1\) meters, and \(Random(\cdot, \cdot)\) represents randomly selecting a floating-point value within the parameter range.
[0056] 2) Object movement. This enhancement method belongs to the "target object" granularity. Different from point movement, when traversing each target object during object movement enhancement, it will first randomly generate random displacements corresponding to the \(x\) and \(y\) axes, and then apply this displacement to all points \(D\) of the current target object. i In this way, the movement of the entire target object is achieved without changing the relative position relationship between different points within the same object, reducing the impact on the convolutional layers in the lower layers of the deep neural network with a smaller receptive field.
[0057] 3) Object rotation. This enhancement method belongs to the "target object" granularity. First, randomly select a rotation angle \(\alpha\) within the range of the maximum rotation angle \([-A, A]\), then calculate the corresponding rotation matrix \(Matrix(\alpha)\) according to \(\alpha\), and apply the rotation matrix to all points in the current target object respectively to achieve the rotation of the object. Similar to the object movement enhancement method, the object rotation enhancement method also does not have a large impact on the relative relationship between different points within the same object, but the relative positions between different objects in the same point cloud image may change to varying degrees.
[0058] Step 5.2: Select an unenhanced target object from the set \(S\) of objects to be enhanced aug and apply the enhancement method obtained in Step 5.1 to this target object.
[0059] Step 5.3: When there is no unenhanced target object in the set \(S\) aug go to Step 6; otherwise, go to Step 5.1.
[0060] Step 6: When the model training is completed, end the training; otherwise, use the point cloud data containing the enhanced target object as new training data and go back to Step 1.1. It can be judged whether to stop the data augmentation process by whether the prediction accuracy of the model on the validation data stops increasing. After the data augmentation process no longer proceeds, the model is regarded as having been trained completed.
[0061] Please refer to Figure 3 , Figure 3 FIG. 1 is a schematic structural diagram of a point cloud data augmentation device 10 provided by an embodiment of the present application. In this embodiment, each module included in the point cloud data augmentation device is used to execute each step in the corresponding embodiment of the present application. For specific details, please refer to the relevant descriptions in the corresponding embodiments. For the sake of simplicity, only the parts related to this embodiment are shown. Refer to Figure 3 , the point cloud augmentation device 10 may include: a calculation module 100 and a storage module 200, where: the storage module 200 stores a deep neural network model, and the calculation module 100 may call the neural network model on the storage module 200 to process the input point cloud to obtain enhanced point cloud data. The deep neural network model is used to obtain a data evaluation result at the category granularity based on the input point cloud; select the target object to be enhanced from the target objects of the original sample according to the evaluation result; and apply a hybrid augmentation strategy combining multiple granularities to the target object to be enhanced to obtain the enhanced point cloud data.
[0062] In one embodiment, the calculation module 100 is further configured to: construct a deep neural network model, input the original point cloud data into the model, and perform multiple forward-backward propagation iterative calculations to obtain the model after gradient descent;
[0063] infer the deep neural network model on a small number of validation set point cloud samples to obtain the inference result of the model;
[0064] perform data perception to calculate the data evaluation result at the category granularity;
[0065] select the target object to be enhanced from the target objects of the original sample according to the evaluation result;
[0066] apply a hybrid augmentation strategy combining multiple granularities to the target object to be enhanced to obtain the enhanced point cloud data for the next cycle of model iteration process.
[0067] Repeat the above data augmentation steps until the training is completed.
[0068] In one embodiment, the calculation module 100 is configured to freeze the parameters of the deep neural network model and post-process the inference result.
[0069] In one embodiment, the calculation module 100 is configured to calculate the average prediction accuracy corresponding to the verification set samples of each category respectively.
[0070] In one embodiment, the calculation module 100 is configured to perform perception based on category granularity, invert the average accuracy of the inference results of the model on the verification set samples, and perform normalization calculation based on category granularity to obtain the enhanced probabilities of the target objects of each category in the original samples.
[0071] In one embodiment, the calculation module 100 is configured to sample the target objects in each sample based on the enhanced probabilities of each category in the data evaluation results, and the sampled target objects are regarded as the target objects to be enhanced.
[0072] In one embodiment, the calculation module 100 is configured to determine whether to stop the data augmentation process by whether the prediction accuracy of the model on the verification data stops growing. After the data augmentation process no longer proceeds, the model is regarded as having been trained.
[0073] Figure 4 It is a schematic structural diagram of a terminal device provided by an embodiment of the present application. As Figure 4 shown, the terminal device 700 of this embodiment includes: a processor 710, a memory 720, and a computer program 730 stored in the memory 720 and executable on the processor 710, such as a program for a point cloud rendering method. When the processor 710 executes the computer program 730, the steps in each of the above embodiments of the point cloud enhancement method are implemented, such as Figure 1 S101 to S103 shown. Alternatively, when the processor 710 executes the computer program 730, the functions of each module in the above Figure 4 corresponding embodiments are implemented. For example, Figure 3 the functions of each module shown. For details, please refer to Figure 3 the relevant descriptions in the corresponding embodiments.
[0074] Exemplarily, the computer program 730 can be divided into one or more modules. One or more modules are stored in the memory 720 and executed by the processor 710 to implement the point cloud enhancement method provided by the embodiments of the present application. One or more modules can be a series of computer program instruction segments capable of completing specific functions, and these instruction segments are used to describe the execution process of the computer program 730 in the terminal device 700. For example, the computer program 730 can implement the point cloud rendering method provided by the embodiments of the present application.
[0075] The terminal device 700 may include, but is not limited to, a processor 710 and a memory 720. Those skilled in the art can understand that Figure 4This is only an example of the terminal device 700, and does not limit the terminal device 700. It may include more or fewer components than those shown in the figure, or combine some components, or different components. For example, the terminal device may also include input and output devices, network access devices, buses, etc.
[0076] The so-called processor 710 may be a central processing unit, or may also be other general-purpose processors, digital signal processors, application-specific integrated circuits, off-the-shelf programmable gate arrays or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc.
[0077] The memory 720 may be an internal storage unit of the terminal device 700, such as the hard disk or memory of the terminal device 700. The memory 720 may also be an external storage device of the terminal device 700, such as a plug-in hard disk, smart memory card, flash card, etc. equipped on the terminal device 700. Further, the memory 720 may also include both the internal storage unit and the external storage device of the terminal device 700.
[0078] The embodiments of the present application provide a computer-readable storage medium storing a computer program, and the computer program is executed by a processor to perform the point cloud rendering method in each of the above embodiments.
[0079] The embodiments of the present application provide a computer program product, which when running on a terminal device, causes the terminal device to execute the point cloud rendering method in each of the above embodiments.
[0080] The above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit it; although the present application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements on some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of each embodiment of the present application, and should all be included in the protection scope of the present application.
[0081] The above is only the preferred specific implementation manner of the present application, but the protection scope of the present application is not limited thereto. Any changes or replacements that can be easily thought of by those skilled in the art within the technical scope disclosed by the present application should be covered by the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.
Claims
1. A method for enhancing point cloud data, characterized in that, Including: Input the collected point cloud data into the trained deep neural network model to obtain enhanced point cloud data, where the deep neural network model is used to obtain a data evaluation result at the category granularity based on the input point cloud; select the target object to be enhanced from the target objects of the original sample according to the evaluation result; Apply a hybrid enhancement strategy combining multiple granularities to the target to be enhanced to obtain enhanced point cloud data.
2. The method according to claim 1, wherein The deep neural network model includes a feature extractor and a detector. The feature extractor adopts a network structure based on convolutional layers, and the detector adopts a network structure based on fully connected layers. The loss function of the deep neural network model includes two parts: classification loss and regression loss.
3. The method according to claim 1, characterized in that, The method further includes: Construct a deep neural network model, input the original point cloud data into the model, and perform multiple forward-backward propagation iterative calculations to obtain the model after gradient descent; Infer the deep neural network model on a small number of validation set point cloud samples to obtain the inference result of the model; Perform data perception to calculate the data evaluation result at the category granularity; Select the target object to be enhanced from the target objects of the original sample according to the evaluation result; Apply a hybrid enhancement strategy combining multiple granularities to the target to be enhanced to obtain enhanced point cloud data for the model iteration process of the next cycle; Repeat the above data enhancement steps until the training is completed.
4. The method according to claim 1, characterized in that Infer the deep neural network model on a small number of validation set point cloud samples to obtain the inference result of the model, including: freezing the parameters of the deep neural network model and post-processing the inference result.
5. The method according to claim 4, characterized in that, The post-processing of the inference result includes: Calculate the average prediction accuracy corresponding to each category of validation set samples respectively.
6. The method according to claim 3, wherein Performing data perception to obtain the model evaluation result includes: Perceiving based on the category granularity; Invert the average accuracy of the inference result of the model on the validation set samples and perform normalization calculation at the category granularity to obtain the enhanced probability of each category of target objects in the original sample.
7. The method according to claim 1, wherein Selecting the target object to be enhanced from the target objects of the original sample according to the evaluation result includes: Sampling the target objects in each sample based on the enhancement probability of each category in the data evaluation result, and regarding the sampled target objects as the target objects to be enhanced.
8. The method according to claim 1, wherein The hybrid enhancement strategy is composed of enhancement methods of multiple granularities, including point granularity movement, object granularity movement, and object granularity rotation; the enhancement methods of multiple granularities are applied to the target objects to be enhanced in a random sampling manner.
9. The method according to claim 3, wherein The method further includes: judging whether to stop the data enhancement process by whether the prediction accuracy of the model on the validation data stops increasing. When the data enhancement process no longer proceeds, the model is regarded as having been trained completed.
10. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by a processor, it implements the method according to any one of claims 1 to 9.