Model post-quantification method and device, equipment and storage medium
By preprocessing, normalizing and model training on point cloud data, and building a calibration set with target classification confidence data, the problem of insufficient accuracy of the existing medium- and low-proportion cloud quantization method in autonomous driving scenarios is solved, and a high-precision and wide-appropriate quantitative deployment model is realized.
Patent Information
- Application Number
- CN202510140711.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-08
- Publication Date
- 2025-06-13
AI Technical Summary
The existing low-profile cloud quantization method is difficult to meet the high requirements for model accuracy in autonomous driving scenarios, especially on embedded devices of different chip manufacturers, which leads to a reduction in model accuracy after quantization, which cannot meet the requirements for detection rate and detection accuracy of surrounding obstacles.
By preprocessing, normalizing and model training on point cloud data, model weight data and target classification confidence data are obtained, calibration sets are constructed based on target classification confidence data, and a quantitative deployment model is obtained, which can be adapted to embedded devices from different manufacturers.
It realizes model deployment with high quantitative accuracy and wide application range, which can meet the high requirements for model accuracy in autonomous driving scenarios and improves the model's detection accuracy for dynamic obstacles.
Smart Images

Figure CN120147774A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of model quantization, and particularly to a post-quantization method, device, equipment, and storage medium for a model. Background Art
[0002] A point cloud is a set of a series of discrete points in a three-dimensional space obtained by sensors such as lidar. These points contain the position information of objects and can be used for subsequent obstacle detection and environmental modeling, etc. With the rapid development of deep learning algorithms and sensor devices, 3D object detection technology based on lidar point clouds has been widely applied in the field of autonomous driving. And in the field of autonomous driving, 3D point cloud object detection technology is crucial. Since the real-time requirement for model inference in autonomous driving scenarios is relatively high, and the memory and computing power of in-vehicle embedded devices are very limited, usually, a convolutional neural network is compressed to accelerate the model inference process, and low-bit quantization is one of the most commonly used model compression methods.
[0003] Currently, in traditional implementation methods, low-bit quantization methods mainly include two types: PTQ (Post-training Quantization) and QAT (Quantization-aware Training). QAT mainly inserts pseudo-quantization operators in a floating-point model and simulates the error brought by the quantization process through retraining, feeds the quantization error back to the loss function, so that the model gradually adapts to the quantized parameters and calculation methods during the training process, achieving the effect of further reducing the accuracy loss of the quantized model. PTQ is mainly a quantization process directly performed after the floating-point model training is completed. It does not require retraining again. Only by analyzing the behavior of the model on typical input data and selecting appropriate quantization parameters to minimize the error introduced by quantization, the effect of keeping the accuracy of the model as much as possible is achieved.
[0004] However, when using the QAT method, it is necessary to manually perform pseudo-quantization adaptation for each operator in the model and be compatible with the multi-machine multi-card scheme used during floating-point model training. The development is relatively complex and requires a long training time; when using the PTQ method, although no retraining is required, existing post-quantization schemes mostly use GPUs or embedded devices as carriers. Since the specific processes and algorithms are encapsulated inside the inference engine and not open-sourced externally, different chip manufacturers have different implementation schemes. If directly performing post-quantization on the model based on existing methods and the quantization configuration parameters defaulted by the manufacturers, it will lead to a reduction in the accuracy of the quantized model and cannot meet the requirements of the detection rate and detection accuracy of surrounding obstacles for autonomous driving vehicles. Summary of the Invention
[0005] Based on this, the present application provides a method, apparatus, device, and storage medium for post-quantizing a model. By performing preprocessing, normalization, and model training operations on point cloud data, model weight data and target classification confidence data are obtained. Calibration set data is obtained based on the target classification confidence data, and then a quantized deployment model is obtained. This model can be adapted to embedded devices of different manufacturers and has the advantages of high quantization accuracy and wide application range.
[0006] In a first aspect, a method for post-quantizing a model is provided. The method includes:
[0007] Obtain point cloud data, perform preprocessing on the point cloud data to obtain preprocessed data;
[0008] Perform a normalization operation on the preprocessed data to obtain normalized data;
[0009] Perform model training on the normalized data to obtain model weight data and target classification confidence data;
[0010] According to the target classification confidence data and a preset calibration set threshold, obtain calibration set data;
[0011] Input the model weight data, the calibration set data, and a preset configuration file into a preset quantization tool to obtain a quantized deployment model.
[0012] According to an implementable manner in an embodiment of the present application, performing preprocessing on point cloud data to obtain preprocessed data includes:
[0013] Perform regional division on the point cloud data to obtain voxel unit data;
[0014] Perform spatial division on the voxel unit data to obtain non-empty point cloud column data;
[0015] According to the non-empty point cloud column data and preset manual feature data, obtain feature vector data;
[0016] Perform a projection operation on the feature vector data to obtain preprocessed data.
[0017] According to an implementable manner in an embodiment of the present application, performing a normalization operation on the preprocessed data to obtain normalized data includes:
[0018] Obtain preset range data and preset interval data, and according to the preset range data, the preset interval data, and a preset normalization model, obtain offset data and scaling factor data;
[0019] According to the offset data and the scaling factor data, perform a normalization operation on the preprocessed data to obtain normalized data.
[0020] According to an implementable manner in an embodiment of the present application, model training is performed on the normalized data to obtain model weight data and target classification confidence data, including:
[0021] Iteratively update according to the preset initial weight and the preset weight model to obtain the model weight data;
[0022] Perform model training inference according to the normalized data, the preset activation function, and the model weight data to obtain the target classification confidence data.
[0023] According to an implementable manner in an embodiment of the present application, the method further includes:
[0024] Perform mapping processing on the target classification confidence data according to the preset mapping function to obtain the mapped confidence data.
[0025] According to an implementable manner in an embodiment of the present application, the preset calibration set threshold includes a preset confidence threshold; according to the target classification confidence data and the preset calibration set threshold, calibration set data is obtained, including:
[0026] Use the target points corresponding to the mapped confidence in the mapped confidence data whose mapped confidence is greater than the preset confidence threshold as valid target points;
[0027] Obtain the total number of valid target points, and obtain the calibration set data according to the total number of valid target points.
[0028] According to an implementable manner in an embodiment of the present application, the preset calibration set threshold further includes a preset valid threshold and a preset cumulative threshold; according to the total number of valid target points, calibration set data is obtained, including:
[0029] When the total number of valid target points is greater than the preset valid threshold, obtain the valid target point data;
[0030] When the cumulative number of frames of the valid target point data is greater than the preset cumulative threshold, the calibration set construction is completed, and the calibration set data is obtained.
[0031] In a second aspect, a model post-quantization device is provided, and the device includes:
[0032] A preprocessing unit, configured to obtain point cloud data and perform preprocessing on the point cloud data to obtain preprocessed data;
[0033] A normalization unit, configured to perform a normalization operation on the preprocessed data to obtain normalized data;
[0034] A model training unit, configured to perform model training on the normalized data to obtain model weight data and target classification confidence data;
[0035] A calibration unit for obtaining calibration set data according to target classification confidence data and a preset calibration set threshold;
[0036] A quantization unit for inputting model weight data, calibration set data, and a preset configuration file into a preset quantization tool to obtain a quantized deployment model.
[0037] In a third aspect, a computer device is provided, including:
[0038] At least one processor; and
[0039] A memory communicatively connected to the at least one processor; wherein,
[0040] The memory stores computer instructions executable by the at least one processor, and the computer instructions are executed by the at least one processor to enable the at least one processor to execute the method involved in the first aspect above.
[0041] In a fourth aspect, a computer-readable storage medium is provided, on which computer instructions are stored, characterized in that the computer instructions are used to cause a computer to execute the method involved in the first aspect above.
[0042] According to the technical content provided by the embodiments of the present application, point cloud data is acquired, the point cloud data is preprocessed to obtain preprocessed data; the preprocessed data is normalized to obtain normalized data; the normalized data is subjected to model training to obtain model weight data and target classification confidence data; calibration set data is obtained according to the target classification confidence data and a preset calibration set threshold; the model weight data, the calibration set data, and a preset configuration file are input into a preset quantization tool to obtain a quantized deployment model. Through the above operations, by preprocessing, normalizing, and model training on the point cloud data, model weight data and target classification confidence data are obtained, calibration set data is obtained based on the target classification confidence data, and then a quantized deployment model is obtained. This model can be adapted to embedded devices of different manufacturers and has the advantages of high quantization accuracy and wide application range. Description of the Drawings
[0043] Figure 1 It is a schematic flowchart of a post-quantization method for a model in an embodiment;
[0044] Figure 2 It is a preferred schematic flowchart of a post-quantization method for a model in an embodiment;
[0045] Figure 3 It is a structural block diagram of a post-quantization device for a model in an embodiment;
[0046] Figure 4 It is a schematic structural diagram of a computer device in an embodiment. Detailed implementation manners
[0047] The present application will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.
[0048] For the convenience of understanding, the system applicable to the present application will be described first. A post-quantization method for a model provided by the present application can be applied to a computer device, and the computer device may include a terminal or a server. Specifically, the computer device acquires point cloud data, preprocesses the point cloud data to obtain preprocessed data; performs a normalization operation on the preprocessed data to obtain normalized data; performs model training on the normalized data to obtain model weight data and target classification confidence data; obtains calibration set data according to the target classification confidence data and a preset calibration set threshold; inputs the model weight data, the calibration set data, and a preset configuration file into a preset quantization tool to obtain a quantized deployment model. Among them, the terminal may be, but is not limited to, an in-vehicle terminal.
[0049] Figure 1 The flowchart of a post-quantization method for a model provided by an embodiment of the present application, and this method can be executed by a computer device. As Figure 1 shown, this method may include the following steps:
[0050] Step S101: Acquire point cloud data, preprocess the point cloud data to obtain preprocessed data.
[0051] Among them, the point cloud data refers to a set of vectors in a three-dimensional coordinate system. The scanning information is recorded in the form of points, and each point contains three-dimensional coordinates (X, Y, Z), and some points may also contain color information or reflection intensity information. The color information is usually obtained by a camera to acquire a color image, and then the color information of the corresponding pixel at the corresponding position is assigned to the corresponding point in the point cloud. The intensity information is the echo intensity collected by the receiving device of the laser scanner, and this intensity information is related to the surface material, roughness, incident angle direction of the target, as well as the emission energy and laser wavelength of the instrument.
[0052] Here, the computer device acquires point cloud data through sensors such as lidar installed on the in-vehicle terminal, and the point cloud data is stored in a matrix form with a dimension of (n, 4). Among them, n represents the number of target points included in the point cloud data, and 4 represents the dimension information, that is, the spatial coordinates x, y, z of the position of each target point and the reflection intensity r respectively. By performing a preprocessing operation on the point cloud data, preprocessed data can be obtained, and the preprocessed data is finally presented in the form of a two-dimensional BEV (Bird's Eye View) pseudo-image.
[0053] Step S103: Perform a normalization operation on the preprocessed data to obtain normalized data.
[0054] Here, since many embedded manufacturers only support symmetric quantization during quantization deployment, that is, during the quantization process, the quantization range is symmetric about zero, and no complex zero-point calculations and adjustments are required after quantization. At the same time, it can make the quantization and dequantization processes more easily utilize hardware acceleration and the speed is relatively fast. Specifically, general symmetric quantization can map the numerical range of int8 type from -128 to 127 to -2 n ~2 n , where n is an integer. The quantization process here is a normalization operation. After performing the normalization operation on the preprocessed data and then training the model, the model training process can be accelerated.
[0055] Step S105: Perform model training on the normalized data to obtain model weight data and target classification confidence data.
[0056] Among them, model training can adopt a convolutional neural network model, which generally consists of a backbone network, a neck network, and a head network; the backbone network usually uses a conventional ResNet (Residual Network) structure for feature extraction; the neck network usually uses an FPN (Feature Pyramid Network) structure to perform top-down feature fusion on the multi-scale features extracted from the backbone network; the head network usually adopts a structure similar to the SSD (Single Shot Multibox Detector) algorithm to achieve the regression and classification of bounding boxes in 3D object detection.
[0057] Here, the normalized data is used as the training set and input into the convolutional neural network model for model training. During the model training process, model weight data can be obtained. At the same time, by introducing loss functions for tasks such as regression and classification, after the model training is completed, model inference is performed on the convolutional neural network model, and the target classification confidence of each target point in the corresponding point cloud data can be obtained, that is, the target classification confidence data.
[0058] Step S107: Obtain calibration set data according to the target classification confidence data and a preset calibration set threshold.
[0059] Here, model inference is performed on the normalized data, that is, the training set, and samples for constructing the calibration set are screened according to the inference results. Specifically, calibration set data can be obtained according to the target classification confidence data and a preset calibration set threshold. It should be noted that during the process of constructing the calibration set, the actually saved calibration set data are all BEV pseudo-images after performing point cloud preprocessing, and can be saved locally in the form of.npy, that is, a numpy array.
[0060] Step S109: Input the model weight data, calibration set data, and a preset configuration file into a preset quantization tool to obtain a quantized deployment model.
[0061] Among them, both the preset configuration file and the preset quantization tool are set by the manufacturer based on their own needs.
[0062] Here, inputting the model weight data, calibration set data, and a preset configuration file into the commands and preset quantization tool provided by the manufacturer can automatically execute processes such as model parsing, graph optimization, calibration, quantization, and compilation, and then obtain the final deployed model for use after quantization, that is, the quantized deployment model.
[0063] It can be seen that in the embodiment of the present application, by acquiring point cloud data, preprocessing the point cloud data to obtain preprocessed data; performing a normalization operation on the preprocessed data to obtain normalized data; performing model training on the normalized data to obtain model weight data and target classification confidence data; obtaining calibration set data according to the target classification confidence data and a preset calibration set threshold; inputting the model weight data, calibration set data, and a preset configuration file into a preset quantization tool to obtain a quantized deployment model. The present application preprocesses, normalizes, and performs model training on point cloud data to obtain model weight data and target classification confidence data, obtains calibration set data based on the target classification confidence data, and then obtains a quantized deployment model. This model can be adapted to embedded devices of different manufacturers and has the advantages of high quantization accuracy and wide application range.
[0064] The following describes the different steps in the above method process in detail. First, in combination with an embodiment, the step of "preprocessing the point cloud data to obtain preprocessed data" in step 101 above is described in detail.
[0065] Perform regional division on the point cloud data to obtain voxel unit data; perform spatial division on the voxel unit data to obtain non-empty point cloud column data; obtain feature vector data according to the non-empty point cloud column data and preset manual feature data; perform a projection operation on the feature vector data to obtain preprocessed data.
[0066] Specifically, performing regional division on the point cloud data, that is, a voxelization operation, can obtain voxel unit data. For example, assume that the current three-dimensional space scene contains some point cloud data about a target, the side length of the preset voxel is a certain value, and the space where the point cloud data is located is divided into multiple voxels, that is, multiple cube meshes. For each point cloud in the point cloud data, it can be assigned to the corresponding voxel unit to obtain voxel unit data.
[0067] Here, for the mainstream 3D point cloud detection algorithms in the autonomous driving industry, that is, the voxel-based perception algorithm can convert irregular point cloud data into regular 3D grids and extract features, but most hardware platforms are difficult to support the deployment and quantization of its 3D convolution or sparse convolution. While the cylinder-based point cloud perception algorithm can convert point cloud data into cylinders and then use a 2D convolutional neural network to learn features, with a faster inference speed and easier deployment at the same accuracy, making it more suitable for autonomous driving scenarios. Therefore, spatial partitioning can also be performed on voxel unit data, that is, spatial partitioning in the x and y horizontal plane directions (point cloud coordinate system direction: x forward, y left, z up), and a columnar point cloud representation in three-dimensional space can be obtained, and then non-empty point cloud column data can be obtained. The columnar point cloud representation is a tensor with dimensions (P, N, 4); a tensor can be regarded as a generalization of a matrix in a high-dimensional space; where 4 represents the four dimensions of the x, y, z, and reflection intensity r of the target points in the point cloud data, P represents the number of non-empty point cloud columns, N represents the maximum number of target points that will be retained in each point cloud column, and each point cloud column retains at most N target points through random sampling, and if there are less than N target points, zero padding can be used to make up the difference. And N generally takes a value of 64 here.
[0068] After obtaining the non-empty point cloud column data, the general traditional approach is to calculate the tensor of the non-empty point cloud column data, perform voxel feature encoding on the tensor, and then extract the features of the point cloud column based on the max pooling layer, with the dimension reduced to (c, P); where c is the number of encoded feature channels, defaulting to 64; P still represents the number of non-empty point cloud columns. However, the traditional approach has certain limitations. Since the pooling layer of the convolutional neural network is introduced, and the other links in the preprocessing process implement the corresponding logic through C++ programming and run on the CPU, a large amount of data exchange processes between the neural network unit and the central processing unit are introduced, resulting in additional overhead during model inference and increasing the development cost of the model deployment framework. Therefore, from the perspective of deployment friendliness, the present application adopts the method of manually designing multiple features, that is, presetting manual feature data, and statistically calculating the mean and maximum values of the x, y, z, and r of the target points in each non-empty point cloud column, the minimum values of x, y, and z, and the point cloud density, a total of 12 features; these features are concatenated in the channel dimension to obtain feature vector data with dimensions (c, P). Among them, P still represents the number of non-empty point cloud columns, and c represents the number of manually designed features, which can be set to 12 in the present application. Specifically, the features of the target points in the non-empty point cloud column can be represented in the form of an array. For example, index [0][0] can represent the mean value of x of the first non-empty point cloud column, index [0][1] can represent the mean value of y of the first non-empty point cloud column, and so on, and index [N]
[11] can represent the point cloud density of the Nth non-empty point cloud column.
[0069] Perform a projection operation on the feature vector data to obtain preprocessed data. Here, for the feature vector data with dimensions (c, P), a projection operation is performed, that is, rasterize the feature vector data according to the preset voxel size, and sequentially back-project all the processed non-empty point cloud columns to their corresponding positions in the previous three-dimensional space, and reduce the dimension to obtain a 2D pseudo-image in the BEV view, that is, the preprocessed data, whose dimensions are (c, h, w). Among them, c represents the number of manually designed features, and h and w respectively represent the height and width of the pseudo-image. That is, after the preprocessing operation, the point cloud data can be converted into a BEV pseudo-image with a certain number of channels, height, and width.
[0070] Next, in combination with the embodiments, the above step S103, that is, "perform a normalization operation on the preprocessed data to obtain normalized data", will be described in detail.
[0071] Obtain preset range data and preset interval data, and according to the preset range data, preset interval data, and preset normalization model, obtain offset data and scaling factor data; according to the offset data and scaling factor data, perform a normalization operation on the preprocessed data to obtain normalized data.
[0072] Here, since it is considered that the output of the convolutional neural network model is usually an uncertain value, it is generally necessary to perform per-channel statistics on the generated features after output, including multiple dimensions such as maximum value, minimum value, mean value, and variance to analyze the numerical distribution of the features, discard an appropriate amount of extreme values, and at the same time, it is necessary to ensure that the numerical range of the output features cannot be too wide, otherwise quantization accuracy loss will also occur during floating-point to fixed-point conversion. Therefore, in this application, in the way of manually designed features, after obtaining the preprocessed data, first perform per-channel analysis and statistics on the manual features to obtain the interval range of the corresponding manual features, and then perform the normalization operation. Assume that the range of the point cloud data to be detected in the actual business is 80m in the forward (x direction) and 20m in each of the left and right (y direction). Then the point cloud data collected during the training set and the actual vehicle operation can be filtered according to the corresponding range, and the point cloud data within this range are all valid information with actual physical meanings. Therefore, during quantization, it can be directly mapped to the maximum and minimum values configured during symmetric quantization, without considering the problems of numerical discarding and truncation.
[0073] Specifically, obtain the preset range data and preset interval data input manually, and according to the preset range data, preset interval data, and preset normalization model, obtain offset data and scaling factor data, where the specific expression of the preset normalization model is as follows:
[0074] (e min -offset) / sca l e=-2 n
[0075] (e max-offset) / scale = 2 n
[0076] Wherein, e min represents the minimum value of the preset range corresponding to the manual feature; e max represents the maximum value of the preset range corresponding to the manual feature; offset represents the offset; scale represents the scaling factor; 2 n represents the range value after mapping. n is an arbitrary integer. In this application document, n can generally take 2.
[0077] Suppose that for this manual feature of the non-empty point cloud column in the x direction in the preprocessed data, the minimum value of the preset range corresponding to the manual feature is 0m, and the maximum value of the preset range corresponding to the manual feature is 80m, that is, e min = 0, e max = 80. By substituting the above preset normalization model, the offset and scale of the corresponding channel can be obtained, that is, offset = 40.0, scale = 10.0. Furthermore, the input values from 0 to 80m can be mapped to the interval from -4 to 4. By analogy, all the manual features in the preprocessed data can be normalized to obtain the normalized data. Thus, it can be ensured that the preprocessed data is completely covered in the symmetric interval bounded by the integer power with 2 as the base, and the problem of quantization precision loss caused by numerical truncation will no longer occur.
[0078] In the above operation, after preprocessing the point cloud data with manually designed features, through the method of per-channel statistics, combined with the preset normalization model for quantization processing, the normalized data is obtained. And because the manual features all have actual physical meanings, the point cloud data collected during the actual vehicle operation later will also be filtered according to the corresponding range, and the point clouds within this range are all valid information with actual physical meanings, without considering the problems of numerical discarding and truncation. Furthermore, the effect of improving the later quantization fineness and reducing the quantization precision loss is achieved.
[0079] Then, in combination with the embodiments, the above step S105, that is, "performing model training on the normalized data to obtain the model weight data and the target classification confidence data", is described in detail.
[0080] According to the preset initial weight and the preset weight model, iterative update is performed to obtain the model weight data; according to the normalized data, the preset activation function, and the model weight data, model inference is performed to obtain the target classification confidence data.
[0081] Here, model training is performed through a convolutional neural network model. During the training process, a loss function can be introduced. Based on the loss function, the total loss value corresponding to each iteration number can be obtained. Based on the obtained total loss value and the preset initial weights, the model weight gradient can be calculated and transmitted in the form of backpropagation. The weights are iteratively updated based on the gradient values generated by different weights to obtain the model weight data.
[0082] Specifically, the expression of the preset weight model can be represented as follows:
[0083]
[0084] Among them, w represents the preset initial weight; lr represents the learning rate; L represents the total loss value; represents the weight gradient value; w, lr, L, and can all be automatically obtained by the training framework during the model training process. It should be noted that since w’ is iteratively updated, the w’ output in the first iteration will be used as the w for the second iteration update and so on, until the convolutional neural network model gradually acquires the ability to detect target obstacles through multiple rounds of iteration, at which point the iteration stops, and then the model weight data is obtained. After obtaining the model weight data, it can be stored in the form of.pth, and the data type of the weight data is float; at the same time, the saved model weight data is loaded, and based on the model export interface provided by the model training framework, the model weight data is converted into the onnx format.
[0085] Based on the normalized data, the preset activation function, and the model weight data, model training inference is performed to obtain the target classification confidence data.
[0086] Here, the preset activation function uses the ReLU6 activation function, which can limit the output range of each layer of the convolutional neural network model to 0 to 6. In this way, it can not only ensure that during the floating-point model training process, the expression ability of the floating-point model will not be overly suppressed, but also effectively reduce the occurrence of outliers, which is beneficial to the subsequent post-training quantization process. Under the same quantization bit width (such as int8, with a numerical range of -128 to 127), the quantization granularity is finer, and the effect of reducing the quantization accuracy loss and improving the model performance can be achieved.
[0087] Specifically, when performing model training based on the normalized data, the preset activation function, and the model weight data, a specific number of point cloud data, that is, the normalized data, will be input into the convolutional neural network model in each round of training iteration to complete the forward propagation process of the model. Among them, the convolutional neural network model can be abstracted as a combination of multiple convolutional units, and each convolutional unit can be represented as follows:
[0088] y = f(g(W·X + b))
[0089] Among them, X represents the normalized data obtained after a certain amount of point cloud data undergoes preprocessing and normalization operations, or the output of the previous convolutional unit; W represents the weight matrix data of each convolutional unit; b represents the preset offset; y represents the output of each convolutional unit, g represents the batch normalization network layer, and f represents the preset activation function. It should be noted that the above convolutional unit is a process of multiple iterations. The weight matrix data corresponding to each iteration corresponds to the weight data obtained by iterating the preset weight model. After multiple rounds of iteration, the weight data output after the last iteration is input into the convolutional unit for model inference, and then the target classification confidence data of each target point in the corresponding normalized data can be obtained. It should be emphasized that since the process of using a convolutional neural network for model training is a prior art, only the key content involved in this application is briefly described here, and the specific detailed process of model training will not be elaborated.
[0090] In an implementable manner, the method further includes: mapping the target classification confidence data according to a preset mapping function to obtain mapped confidence data.
[0091] Among them, the expression of the preset mapping function is as follows:
[0092] f(x) = 1 / (1 + e^(-x))
[0093] Here, x represents the input target classification confidence data, and f(x) represents the output mapped confidence data. The reason for performing the mapping process is that for the outputs with certain physical meanings, after model training, the output range may be relatively wide. Therefore, through the preset mapping function, the target classification confidence values in the finally output target classification confidence data are all covered between 0 and 1. While meeting the actual usage requirements, it can ensure a relatively fine quantization granularity, which is beneficial to improving the detection accuracy of the quantized model.
[0094] Through the above operations, a convolutional neural network model is used to train the normalized data to obtain model weight data and target classification confidence data, which is convenient for the construction of the calibration set and the post-quantization process in the later stage. At the same time, the setting of the ReLU6 activation function can achieve the effects of improving the fineness of the quantization granularity, reducing the loss of quantization accuracy, and enhancing the model performance.
[0095] Finally, the above step S107, that is, "obtaining calibration set data according to the target classification confidence data and the preset calibration set threshold", is described in detail in combination with the embodiments.
[0096] Take the target points in the mapping confidence data whose corresponding mapping confidence is greater than the preset confidence threshold as valid target points; obtain the total number of valid target points; and obtain the calibration set data based on the total number of valid target points.
[0097] Among them, the preset calibration set thresholds include a preset confidence threshold, a preset valid threshold, and a preset cumulative threshold. Here, no specific values are limited. However, generally, the preset confidence threshold can be set to 0.9; the preset valid threshold can be set to 10; and the preset cumulative threshold can be set to 500 frames.
[0098] Here, through the model inference process, target classification confidence data can be obtained. Based on the preset mapping function, mapping confidence data can be obtained. Furthermore, the mapping confidence corresponding to each target point in the mapping confidence data can be compared with the preset confidence threshold. The target points with mapping confidence less than or equal to the preset confidence threshold are directly ignored, and the target points with mapping confidence greater than the preset confidence threshold are taken as valid target points.
[0099] Obtain the total number of valid target points, and obtain the calibration set data based on the total number of valid target points. In an achievable manner, when the total number of valid target points is greater than the preset valid threshold, valid target point data is obtained; when the cumulative number of frames of the valid target point data is greater than the preset cumulative threshold, the calibration set construction is completed, and the calibration set data is obtained.
[0100] Here, count the total number of valid target points. When the total number of valid target points is greater than the preset valid threshold, it is considered that the current sample is a valid sample and can be used as the valid target point data in the calibration set. It should be noted that the valid target point data here is for the current frame, and during the model inference process, there will be outputs of multiple frames of data, and thus multiple frames of valid target point data can be obtained. When the cumulative number of frames of the valid target point data is greater than the preset cumulative threshold, the calibration set construction is completed, and the calibration set data is obtained.
[0101] Through the above operations, through targeted screening, the post-quantization process can be achieved with only a small number of calibration sets, shortening the quantization time, improving the detection accuracy of the model for dynamic obstacles (such as pedestrians, vehicles, etc.), ensuring that the accuracy loss of each category is reduced to about 1%; at the same time, it can ensure that there will be no extreme scenarios or abnormal data in the calibration set, which is more efficient than manually selecting a large number of calibration set samples, and can quickly verify the effect of the quantized model, and maintain a low quantization accuracy loss, meeting the requirements of the autonomous driving scenario for high model accuracy and efficiency.
[0102] Here, the comparison results of the floating-point and fixed-point model accuracies obtained by using the method of this application are attached as follows (the model accuracy uses the BEV AP40 index; when evaluating, the matching threshold of IoU 0.5 is used for cars and trucks, and the matching threshold of IoU 0.25 is used for pedestrians and cyclists):
[0103]
[0104] Combined with the implementation methods in the above embodiments, the following will combine Figure 2 to give an example description of a preferred method flow provided by the embodiments of this application. As Figure 2 shown, the method may include the following steps:
[0105] Step S201: Obtain point cloud data, perform regional division on the point cloud data, and obtain voxel unit data.
[0106] Step S202: Perform spatial division on the voxel unit data to obtain non-empty point cloud column data.
[0107] Step S203: According to the non-empty point cloud column data and the preset manual feature data, obtain feature vector data.
[0108] Step S204: Perform a projection operation on the feature vector data to obtain preprocessed data.
[0109] Step S205: Obtain preset range data and preset interval data, and according to the preset range data, preset interval data, and preset normalization model, obtain offset data and scaling coefficient data.
[0110] Step S206: According to the offset data and the scaling coefficient data, perform a normalization operation on the preprocessed data to obtain normalized data.
[0111] Step S207: Perform iterative update according to the preset initial weight and the preset weight model to obtain model weight data.
[0112] Step S208: Perform model training and inference according to the normalized data, the preset activation function, and the model weight data to obtain target classification confidence data.
[0113] Step S209: According to the preset mapping function, perform a mapping process on the target classification confidence data to obtain mapped confidence data.
[0114] Step S210: Use the target points corresponding to the mapped confidence in the mapped confidence data that are greater than the preset confidence threshold as valid target points.
[0115] Step S211: Obtain the total number of valid target points. When the total number of valid target points is greater than the preset valid threshold, obtain valid target point data.
[0116] In step S212, when the cumulative number of frames of valid target point data is greater than a preset cumulative threshold, the calibration set construction is completed, and calibration set data is obtained.
[0117] It should be understood that although Figure 1 - Figure 2 each step in the flowchart of Figure 1 - Figure 2 is shown in sequence according to the indication of the arrow, these steps are not necessarily executed in the order indicated by the arrow. Unless there is a clear description in this application, the execution of these steps has no strict order limit, and these steps can be executed in other orders. Moreover,
[0118] Figure 3 FIG. Figure 1 is a schematic structural diagram of a model post-quantization device provided by an embodiment of this application. This device can be set in a computer device to execute the method flow as shown in Figure 2 . As shown in Figure 3 , this device can include: a preprocessing unit 301, a normalization unit 303, a model training unit 305, a calibration unit 307, and a quantization unit 309. The main functions of each component module are as follows:
[0119] The preprocessing unit 301 is used to obtain point cloud data, perform preprocessing on the point cloud data, and obtain preprocessed data;
[0120] The normalization unit 303 is used to perform a normalization operation on the preprocessed data to obtain normalized data;
[0121] The model training unit 305 is used to perform model training on the normalized data to obtain model weight data and target classification confidence data;
[0122] The calibration unit 307 is used to obtain calibration set data according to the target classification confidence data and a preset calibration set threshold;
[0123] The quantization unit 309 is used to input the model weight data, the calibration set data, and a preset configuration file into a preset quantization tool to obtain a quantized deployment model.
[0124] In one embodiment, the preprocessing unit 301 is further used to:
[0125] perform regional division on the point cloud data to obtain voxel unit data;
[0126] Perform spatial partitioning on the voxel unit data to obtain non-empty point cloud column data;
[0127] Based on the non-empty point cloud column data and the preset manual feature data, obtain the feature vector data;
[0128] Perform a projection operation on the feature vector data to obtain the preprocessed data.
[0129] In one embodiment, the normalization unit 303 is further configured to:
[0130] Obtain the preset range data and the preset interval data, and based on the preset range data, the preset interval data, and the preset normalization model, obtain the offset data and the scaling coefficient data;
[0131] Based on the offset data and the scaling coefficient data, perform a normalization operation on the preprocessed data to obtain the normalized data.
[0132] In one embodiment, the model training unit 305 is further configured to:
[0133] Perform iterative update according to the preset initial weight and the preset weight model to obtain the model weight data;
[0134] Based on the normalized data, the preset activation function, and the model weight data, perform model training and inference to obtain the target classification confidence data.
[0135] In one embodiment, the apparatus is further configured to:
[0136] Based on the preset mapping function, perform a mapping process on the target classification confidence data to obtain the mapped confidence data.
[0137] In one embodiment, the preset calibration set threshold includes a preset confidence threshold, and the calibration unit 307 is further configured to:
[0138] Take the target points corresponding to the mapped confidence greater than the preset confidence threshold in the mapped confidence data as the valid target points;
[0139] Obtain the total number of valid target points, and based on the total number of valid target points, obtain the calibration set data.
[0140] In one embodiment, the preset calibration set threshold further includes a preset valid threshold and a preset cumulative threshold, and the calibration unit 307 is further configured to:
[0141] When the total number of valid target points is greater than the preset valid threshold, obtain the valid target point data;
[0142] When the cumulative number of frames of the valid target point data is greater than the preset cumulative threshold, the calibration set construction is completed, and the calibration set data is obtained.
[0143] For the same or similar parts among the above embodiments, reference may be made to each other, and each embodiment focuses on the differences from other embodiments. In particular, for the apparatus embodiments, since they are basically similar to the method embodiments, the description is relatively simple, and for the relevant parts, reference may be made to the description in the method embodiments.
[0144] It should be noted that the embodiments of the present application may involve the use of user data. In actual applications, user-specific personal data may be used in the solutions described herein within the scope permitted by applicable laws and regulations in the country where it is located (for example, with the user's explicit consent, giving the user a practical notice, the user's explicit authorization, etc.).
[0145] According to the embodiments of the present application, the present application also provides a computer device and a computer-readable storage medium.
[0146] As Figure 4 shown, it is a block diagram of a computer device according to an embodiment of the present application. The computer device is intended to represent various forms of digital computers or mobile devices. Among them, the digital computer may include a desktop computer, a portable computer, a workbench, a personal digital assistant, a server, a mainframe computer, and other suitable computers. The mobile device may include a tablet computer, a smart phone, a wearable device, etc.
[0147] As Figure 4 shown, the device 400 includes a computing unit 401, a ROM 402, a RAM 403, a bus 404, and an input / output (I / O) interface 405. The computing unit 401, the ROM 402, and the RAM 403 are connected to each other through the bus 404. The input / output (I / O) interface 405 is also connected to the bus 404.
[0148] The computing unit 401 can execute various processes in the method embodiments of the present application according to the computer instructions stored in the read-only memory (ROM) 402 or the computer instructions loaded from the storage unit 408 into the random access memory (RAM) 403. The computing unit 501 can be various general-purpose and / or special-purpose processing components with processing and computing capabilities. The computing unit 401 may include, but is not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. In some embodiments, the method provided by the embodiments of the present application may be implemented as a computer software program, which is tangibly contained in a computer-readable storage medium, such as the storage unit 408.
[0149] The RAM 403 can also store various programs and data required for the operation of the device 400. Part or all of the computer programs can be loaded and / or installed onto the device 400 via the ROM 402 and / or the communication unit 409.
[0150] The input unit 406, output unit 407, storage unit 408, and communication unit 409 in the device 500 can be connected to the I / O interface 405. Among them, the input unit 506 can be, for example, a keyboard, a mouse, a touch screen, a microphone, etc.; the output unit 407 can be, for example, a display, a speaker, an indicator light, etc. The device 400 can exchange information, data, etc. with other devices through the communication unit 409.
[0151] It should be noted that the device may also include other components necessary for normal operation. It may also only include the components necessary to implement the solution of this application, and does not necessarily include all the components shown in the figure.
[0152] The various embodiments of the systems and techniques described herein can be implemented in digital electronic circuit systems, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on a chip (SOCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof.
[0153] The computer instructions for implementing the methods of this application can be written in any combination of one or more programming languages. These computer instructions can be provided to the computing unit 401, such that when the computer instructions are executed by a computing unit 401 such as a processor, the steps involved in the method embodiments of this application are executed.
[0154] The computer-readable storage medium provided by this application can be a tangible medium that can contain or store computer instructions for executing the steps involved in the method embodiments of this application. The computer-readable storage medium can include, but is not limited to, storage media in the forms of electronic, magnetic, optical, electromagnetic, etc.
[0155] The above specific embodiments do not constitute a limitation on the protection scope of this application. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this application shall be included within the protection scope of this application.
Claims
1. A model post-quantization method, characterized in that: The method comprises: Acquire point cloud data, and preprocess the point cloud data to obtain preprocessed data; Performing a normalization operation on the preprocessed data to obtain normalized data; Performing model training on the normalized data to obtain model weight data and target classification confidence data; Obtaining calibration set data according to the target classification confidence data and a preset calibration set threshold; The model weight data, the calibration set data and the preset configuration file are input into a preset quantization tool to obtain a quantization deployment model.
2. The method according to claim 1, characterized in that The preprocessing of the point cloud data to obtain preprocessed data includes: Dividing the point cloud data into regions to obtain voxel unit data; Performing spatial division on the voxel unit data to obtain non-empty point cloud column data; Obtaining feature vector data according to the non-empty point cloud column data and preset manual feature data; A projection operation is performed on the feature vector data to obtain preprocessed data.
3. The method according to claim 2, characterized in that The normalizing operation is performed on the preprocessed data to obtain normalized data, including: Acquire preset range data and preset interval data, and obtain offset data and scaling coefficient data according to the preset range data, preset interval data and a preset normalization model; The preprocessed data is normalized according to the offset data and the scaling factor data to obtain normalized data.
4. The method according to claim 1, characterized in that: The performing model training on the normalized data to obtain model weight data and target classification confidence data includes: Iteratively update according to the preset initial weight and the preset weight model to obtain model weight data; Model training and inference are performed based on the normalized data, the preset activation function and the model weight data to obtain target classification confidence data.
5. The method according to claim 4, characterized in that The method further comprises: According to a preset mapping function, the target classification confidence data is mapped to obtain mapping confidence data.
6. The method according to claim 5, characterized in that The preset calibration set threshold includes a preset confidence threshold; The obtaining of calibration set data according to the target classification confidence data and a preset calibration set threshold comprises: The target points whose mapping confidences corresponding to the target points in the mapping confidence data are greater than the preset confidence threshold are taken as valid target points; The total number of valid target points is obtained, and the calibration set data is obtained according to the total number of valid target points.
7. The method according to claim 6, characterized in that The preset calibration set threshold also includes a preset effective threshold and a preset cumulative threshold; the calibration set data is obtained according to the total number of valid target points, including: When the total number of the valid target points is greater than the preset valid threshold, obtaining valid target point data; When the accumulated number of frames of the valid target point data is greater than the preset accumulation threshold, the calibration set construction is completed and the calibration set data is obtained.
8. A model post-quantization device, characterized in that: The device comprises: A preprocessing unit, used to acquire point cloud data, and preprocess the point cloud data to obtain preprocessed data; A normalization unit, used for performing a normalization operation on the preprocessed data to obtain normalized data; A model training unit, used to perform model training on the normalized data to obtain model weight data and target classification confidence data; A calibration unit, configured to obtain calibration set data according to the target classification confidence data and a preset calibration set threshold; The quantization unit is used to input the model weight data, the calibration set data and the preset configuration file into a preset quantization tool to obtain a quantized deployment model.
9. A computer device comprising: at least one processor; as well as a memory communicatively connected to the at least one processor; wherein, The memory stores computer instructions that can be executed by the at least one processor, and the computer instructions are executed by the at least one processor to enable the at least one processor to perform the method according to any one of claims 1 to 7.
10. A computer-readable storage medium having computer instructions stored thereon, characterized in that: The computer instructions are used to cause a computer to execute the method according to any one of claims 1 to 7.