Lightweight and hardware acceleration method for point cloud image classification

By removing T-Net networks, full-in quantification and layer fusion processing, and optimizing the PointNet model structure and parameter configuration, the problem of PointNet deployment on resource-constrained hardware platforms is solved, and efficient point cloud image classification is achieved.

CN120070959APending Publication Date: 2025-05-30WUHU RES INST OF XIAN UNIV OF ELECTRONIC SCI & TECH +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510071361.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-16
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

When deploying on resource-constrained hardware platforms, existing PointNet models face high complexity and high computing resource requirements, and are difficult to efficiently implement on embedded systems such as FPGAs or ASICs.

Method used

By removing the T-Net network, performing full quantization and layer fusion processing, designing a parallel computing multiplier array, optimizing the model structure and parameter configuration, and deploying it on edge devices.

Benefits of technology

It realizes efficient deployment of PointNet models on FPGAs or ASICs, reducing model complexity and resource requirements, while maintaining the accuracy and reliability of point cloud processing and improving processing efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120070959A_ABST
    Figure CN120070959A_ABST
Patent Text Reader

Abstract

The invention discloses a lightweight and hardware acceleration method for point cloud image classification, and the method comprises the steps: removing a T-Net network in a PointNet network for point cloud image classification, carrying out the training, carrying out the full integer quantization and layer fusion processing, and deploying the obtained network on an edge device; according to the method, PointNet clipping, full integer quantization and layer fusion processing are utilized, the parameter quantity and the calculation quantity are reduced, a cache structure of edge equipment is designed to achieve hardware cache calculation, the dimensionality of a multiplier-adder array is designed to achieve parallel calculation point convolution, and the hardware logic design and the cache capacity are reduced; lightweight processing of the PointNet network covers multiple aspects of model compression, parameter simplification, calculation optimization and the like, the calculation amount and the storage capacity required by the model can be greatly reduced, the processing speed can be improved, the power consumption can be reduced, and point cloud image classification can be accurately and quickly realized on a resource-limited hardware platform.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of point cloud image classification, and particularly relates to a lightweight and hardware acceleration method for point cloud image classification. Background Art

[0002] In the current era of rapid technological development, the rapid progress of artificial intelligence and deep learning technologies has made the processing of point cloud data increasingly prominent in many key fields such as autonomous driving, drone navigation, robotics, and augmented reality. Among them, PointNet, as a neural network with great influence in the field of point cloud processing, has won wide application and recognition in the vast field of three-dimensional data processing with its unique architecture design and relatively excellent performance.

[0003] PointNet is a deep learning network for processing point cloud data. It directly operates on point cloud data, avoiding complex preprocessing steps in traditional methods, such as meshing or voxelization. PointNet extracts features independently for each point and aggregates global features using a max pooling layer, thereby achieving efficient processing of point cloud data.

[0004] Although PointNet performs well in processing point cloud data, it cannot be ignored that the PointNet model itself often shows high complexity. Its complex network structure and large number of parameters pose great challenges for implementation on resource-constrained hardware platforms. For example, the fully connected layers and high-dimensional feature spaces of PointNet require a large amount of storage and computing resources, which are difficult to bear for embedded systems such as FPGA (Field-Programmable Gate Array) or ASIC (Application-Specific Integrated Circuit).

[0005] FPGA and ASIC have become ideal choices for deploying deep learning algorithms on edge devices due to their unique low-power characteristics, excellent high-performance performance, and powerful highly parallel computing capabilities. However, due to the limited resources, how to efficiently deploy complex deep learning models such as PointNet within the framework of these limited resources has become a hot topic in the current scientific research field and is also a difficult point that needs to be overcome urgently. Summary of the Invention

[0006] In order to solve the above problems existing in the prior art, the present invention provides a lightweight and hardware acceleration method for point cloud image classification. The technical problems to be solved by the present invention are realized through the following technical solutions:

[0007] A lightweight and hardware-accelerated method for point cloud image classification, comprising:

[0008] After removing the T-Net network in the PointNet network for point cloud image classification, train it to obtain the first PointNet_Lite network;

[0009] Perform full integer quantization on the first PointNet_Lite network to obtain the second PointNet_Lite network;

[0010] Perform layer fusion processing on the second PointNet_Lite network to obtain the third PointNet_Lite network; wherein, the layer fusion processing includes fusing the target layer in the mlp structure of the second PointNet_Lite network with the bn layer and then with the Relu layer, and the target layer includes the conv layer and the dense layer;

[0011] Deploy the third PointNet_Lite network on an edge device, the cache structure of the edge device includes off-chip storage and on-chip storage of a neural network accelerator, and the on-chip storage includes unified cache, feature map cache, weight cache, multiplier-accumulator array, temporary cache, output cache; and the multiplier-accumulator array is designed to be of atomic_c*atomic_k dimension to achieve parallelized calculation of point convolution; wherein, the dimension of the network input image is n*1*c, and atomic_c and atomic_k represent adaptively selected parts of c and k.

[0012] In an embodiment of the present invention, the edge device includes an FPGA and an ASIC.

[0013] In an embodiment of the present invention, performing full integer quantization on the first PointNet_Lite network includes:

[0014] Perform int8 symmetric quantization on the weights of the first PointNet_Lite network, perform int32 symmetric quantization on the biases, and perform int8 asymmetric quantization on the input feature maps and output feature maps to achieve full integer quantization.

[0015] In an embodiment of the present invention, performing int8 symmetric quantization on the weight parameters of the first PointNet_Lite network, performing int32 symmetric quantization on the bias parameters, and performing int8 asymmetric quantization on the input feature maps and output feature maps to achieve full integer quantization, includes:

[0016] The input feature maps of each conv layer use a unified scaling factor scale and zero point zero_point, denoted as s 1 and z1 ; Each convolutional kernel weight in the conv layer uses a separate scale and zero_point, denoted as s 2 and z 2 = 0; The bias uses a unified scale and zero_point, denoted as sb = s 1 *s 2 and zb = 0. The output feature map of the conv layer uses a unified scale and zero_point, denoted as s 3 and z 3 ; The weights in each dense layer use a unified scale and zero_point, and the remaining operations are the same as those of the conv layer.

[0017] In an embodiment of the present invention, after performing full integer quantization on the first PointNet_Lite network, the method further includes:

[0018] For the network that has completed full integer quantization, connect a softmax at the output end for normalization processing to obtain a second PointNet_Lite network.

[0019] In an embodiment of the present invention, for the target layer, the process of layer fusion processing includes:

[0020] By substituting the convolutional layer calculation formula into the bn layer calculation formula and then simplifying the formula, fuse the target layer and the bn layer to obtain a fused layer;

[0021] By using the scaling factor and zero point corresponding after the output of the Relu layer as the scaling factor and zero point of the target layer, fuse the fused layer and the subsequent Relu layer.

[0022] In an embodiment of the present invention, the process of fusing the target layer and the bn layer by substituting the convolutional layer calculation formula into the bn layer calculation formula and then simplifying the formula to obtain a fused layer includes:

[0023] Obtain the convolutional layer calculation formula as follows:

[0024] Yconv = Wconv * X + Bconv;

[0025] where Yconv is the output of the target layer, Wconv is the convolutional kernel weight of the target layer, X is the input feature map, and Bconv is the convolutional kernel bias of the target layer;

[0026] Obtain the bn layer calculation formula as follows:

[0027]

[0028] Among them, γ is the scaling factor of the BN layer, β is the offset of the BN layer, μ is the mean in units of the channels of the output feature map of the target layer, and σ 2 is the variance in units of the channels of the output feature map of the target layer, ε is a very small constant used to avoid being 0, and Ybn is the output of the BN layer;

[0029] Substituting the calculation formula of the convolutional layer into the calculation formula of the BN layer, we get:

[0030]

[0031] After simplifying the formula, we get:

[0032]

[0033] Let We get:

[0034] Ybn = Wmerge * X + Bmerge;

[0035] Among them, Wmerge and Bmerge are the convolutional kernel weights and biases of the fusion layer respectively.

[0036] In an embodiment of the present invention, the method further includes:

[0037] Reducing the number of point convolutional kernels used in the MLP structure of five layers in the third PointNet_Lite network to reduce the number of features for each point dimension elevation, and performing network performance tests to determine the MLP dimension elevation strategy that meets the network performance requirements.

[0038] In an embodiment of the present invention, the method further includes:

[0039] For the MLP structure after adopting the MLP dimension elevation strategy, when implementing the conv operation in hardware, adjust the operation order, and use the calculation order of local first and then global to replace the layer-by-layer calculation to achieve MLP hardware acceleration.

[0040] In an embodiment of the present invention, the multiplier-accumulator array is designed to be 16 * 16 dimensions to achieve parallel computing with a parallelism of 16 on the input channels and a parallelism of 16 on the convolutional kernels.

[0041] Advantages of the present invention:

[0042] The lightweight and hardware acceleration method for point cloud image classification provided by the embodiment of the present invention is a lightweight and hardware accelerated implementation of the point cloud image classification convolutional neural network PointNet. The method successfully achieves efficient deployment on FPGA or ASIC by deeply optimizing the model structure and carefully adjusting the parameter configuration. The present invention specifically uses model pruning, full integer quantization, layer fusion processing and other means to significantly reduce the number of model parameters and the complexity of calculation. At the same time, through the specific design of the cache structure of the edge device and the multiplier array therein, hardware cache calculation and parallel calculation of point convolution are realized, reducing the hardware logic design and cache capacity. On the basis of maintaining the performance of the original PointNet model, while reducing the model complexity and resource requirements, the accuracy and reliability of point cloud processing are effectively guaranteed, and it is deployed on the FPGA platform to achieve the hardware acceleration effect of the point cloud neural network in a software-hardware collaborative manner, thereby ensuring the accuracy while improving the processing efficiency of point cloud image classification on edge devices. BRIEF DESCRIPTION OF THE DRAWINGS

[0043] Figure 1 A schematic flow chart of a lightweight and hardware accelerated method for point cloud image classification provided by an embodiment of the present invention;

[0044] Figure 2(a) is a schematic diagram of the structure of the PointNet network;

[0045] Figure 2(b) is a schematic diagram of the structure of the T-Net network in the PointNet network;

[0046] Figure 3 It is a schematic diagram of the structure of the first PointNet_Lite network according to an embodiment of the present invention;

[0047] Figure 4 A schematic diagram of the first 150 training accuracy of the PointNet network and the first PointNet_Lite network according to an embodiment of the present invention;

[0048] Figure 5 This is a schematic diagram of the principle that the Relu front and back scaling factors and zero points should be kept consistent in an embodiment of the present invention;

[0049] Figure 6 This is a schematic diagram of the combination of Conv and Relu in an embodiment of the present invention;

[0050] Figure 7 A schematic diagram of a cache structure according to an embodiment of the present invention;

[0051] Figure 8 A schematic diagram of parallel computation of point convolution according to an embodiment of the present invention;

[0052] Figure 9Schematic diagram of the circuit structure of the multiplier-accumulator array with dimensions atomic_c * atomic_k according to an embodiment of the present invention;

[0053] Figure 10 Schematic diagram of the implementation process of the lightweight and hardware acceleration method for point cloud image classification according to an embodiment of the present invention;

[0054] Figure 11 Schematic diagram of the calculation of continuous point convolution according to an embodiment of the present invention;

[0055] Figure 12 Schematic diagram of the structure of the PointNet_Lite1 network according to an embodiment of the present invention;

[0056] Figure 13 Schematic diagram of the structure of the PointNet_Lite2 network according to an embodiment of the present invention. Detailed implementation manners

[0057] The present invention will be further described in detail below with reference to specific embodiments, but the implementation manners of the present invention are not limited thereto.

[0058] In order to successfully and effectively deploy the PointNet network for point cloud image classification on edge devices such as FPGA and ASIC, it has become an inevitable choice to lightweight the PointNet network. The lightweight processing of the PointNet network in the embodiments of the present invention covers multiple key aspects such as model compression, parameter reduction, and calculation optimization. Through these measures, not only can the amount of calculation and storage capacity required by the model be significantly reduced, but also the processing speed can be significantly improved, and the power consumption can be effectively reduced, thus accurately meeting the stringent requirements of resource-constrained hardware platforms.

[0059] Specifically, the embodiments of the present invention provide a lightweight and hardware acceleration method for point cloud image classification, as Figure 1 shown, the method may include the following steps:

[0060] S1, remove the T-Net network in the PointNet network for point cloud image classification and then train to obtain the first PointNet_Lite network;

[0061] As shown in Figure 2(a), the input of the PointNet network is the position coordinates (x, y, z) of n point cloud data in space, with a size of n*3, where n refers to the number of input point clouds. The feature map of n*3 outputs a 3*3 matrix through the T-Net network. After multiplying this matrix with the input through matrix multiplication, the feature number of each point is increased to 64 and 64 through two consecutive point convolutions of the mlp (multi-layer perceptron) structure. After multiplying with the 64*64 matrix output by the T-Net network, the feature number of each point is increased to 64, 128, and 1024 through three consecutive point convolutions of the mlp structure. Then, the global feature is obtained through the max pooling layer (see max pool in Figure 2(a)). Finally, 40 classifications are output through three consecutive dense layers (i.e., fully connected layers) of the mlp structure. The PointNet network is built in the TensorFlow 2.4, Python 3.7, and cuda 11.6 environment, and trained using NVIDIA GeForce RTX 4090 (the following training environment is the same). The classification accuracy is 88.86%. The computational complexity of the PointNet network is 878M, and the number of parameters is 3.47M.

[0062] Among them, the T-Net network is a trainable network. The input is a feature map of n*p, and the output is a p*p matrix. In the first T-Net network of the PointNet network, p = 3, and in the second T-Net network, p = 64. The model of the T-Net network is shown in Figure 2(b). The input feature map of n*p increases the feature number of each point to 64, 128, and 1024 through three-layer perceptrons of the mlp. The global feature is extracted through the max pooling layer, and then a p*p matrix is output through three-layer perceptrons of the mlp.

[0063] For the specific structures of the PointNet network and the T-Net network therein, please refer to the related technologies for a detailed understanding.

[0064] The PointNet network contains two T-Net networks, which are used to align and transform the input point cloud and intermediate features. However, these two T-Net networks introduce a large amount of computation and parameters, increasing the complexity and resource requirements of the model. In the present invention, by removing these two T-Net networks, the impacts on the computational complexity, the number of parameters, and the accuracy are analyzed comparatively, and the feasibility and reliability of removing the T-Net are studied.

[0065] Figure 3 The first PointNet_Lite network is obtained after removing the T-Net network in PointNet and training. Its classification accuracy is 87.88%, the computational complexity is 303M, and the number of parameters is 0.8M.

[0066] Figure 4To remove the comparison graph of the last 150 training times out of the first 250 training times before and after two T-Net networks, the performance improvement effect of the T-Net on the PointNet network is not significant, but a large number of parameters and computational amounts are introduced. Therefore, removing the two T-Net networks in the present invention can save the parameter amount and computational amount, while still ensuring the training accuracy. Figure 4 In it, PointNet_Lite is the first PointNet_Lite network.

[0067] S2. Perform full integer quantization on the first PointNet_Lite network to obtain a second PointNet_Lite network;

[0068] The weights and biases in the first PointNet_Lite network are of the float32 type (a common floating-point type), each parameter needs to occupy 4 bytes, and the multiplier-accumulator structure of the float32 type is complex. In order to reduce the memory usage of edge devices, the present invention needs to perform full integer quantization on the entire network.

[0069] Among them, the floating-point to fixed-point (quantization) formula is:

[0070]

[0071] The fixed-point to floating-point (dequantization) formula is:

[0072] R = (Q - Z) * S (2);

[0073] The calculation formulas for the scaling factor and zero point are:

[0074]

[0075]

[0076] In formulas (1) to (4), R is the original floating-point data, Q is the quantized fixed-point data, S is the scaling factor scale, Z is the zero point zero_point, Rmax and Rmin respectively represent the maximum and minimum values of R, and Qmax and Qmin respectively represent the maximum and minimum values of Q.

[0077] The quantization target of int8 quantization is Q ∈ [-128, 127], and the int8 quantization formula is as shown in formula (5):

[0078]

[0079] In formula (5), the clip operation truncates and limits the data within [-128, 127].

[0080] Similarly, the int32 quantization formula is as shown in formula (6):

[0081]

[0082] In Equation (6), the clip operation truncates and limits the data to [-2 32 , 2 32 -1].

[0083] For S2, specifically: during the quantization process, the weights of the first PointNet_Lite network are quantized symmetrically to int8, the biases are quantized symmetrically to int32, and the input feature maps and output feature maps (activation maps) are quantized asymmetrically to int8 to achieve full integer quantization.

[0084] More specifically, the input feature maps of each conv layer use a unified scaling factor scale and zero point zero_point, denoted as s 1 and z 1 ; each convolutional kernel weight in the conv layer uses a separate scale and zero point, denoted as s 2 and z 2 = 0; the biases use a unified scale and zero point, denoted as sb = s 1 * s 2 and zb = 0, and the output feature maps of the conv layer use a unified scale and zero point, denoted as s 3 and z 3 ; the weights in each dense layer use a unified scale and zero point, and the remaining operations are the same as those of the conv layer.

[0085] Among them, symmetric quantization means zero_point = 0; for int8 symmetric quantization, int8 asymmetric quantization, and int32 symmetric quantization, please refer to the related technologies for understanding and will not be elaborated here.

[0086] Furthermore, after performing full integer quantization on the first PointNet_Lite network, the lightweight and hardware acceleration method for point cloud image classification further includes:

[0087] For the network after completing full integer quantization, a softmax is connected at the output end for normalization processing to obtain the second PointNet_Lite network.

[0088] The output after connecting the softmax layer can intuitively reflect the classification probability rather than the classification score. For the process of connecting the softmax for normalization processing, please refer to the related technologies for understanding and will not be elaborated here.

[0089] S3. Perform layer fusion processing on the second PointNet_Lite network to obtain a third PointNet_Lite network;

[0090] For simplicity, the second PointNet_Lite network is not shown in the figure. Please refer to Figure 3 the understanding of the first PointNet_Lite network. In the second PointNet_Lite network, the conv block in the mlp uses a three-layer conv structure, and the dense block uses a three-layer dense structure. Specifically, each layer of the conv structure contains a three-layer combination of conv + bn + Relu, and each layer of the dense structure contains a three-layer combination of dense + bn + Relu.

[0091] The layer fusion processing includes fusing the target layer in the mlp structure of the second PointNet_Lite network with the bn layer and then with the Relu layer. The target layer includes the conv layer and the dense layer;

[0092] The layer fusion processing can be denoted as conv / dense + bn + Relu fusion, that is, conv + bn + Relu fusion and dense + bn + Relu fusion. Using the conv / dense + bn + Relu fusion method, conv / dense can be fused with the subsequent bn and Relu, reducing the three-layer structure of conv / dense + bn + Relu to one layer of conv / dense, thus reducing the design of hardware logic.

[0093] Specifically, for the target layer, the process of layer fusion processing can include step A1 and step A2:

[0094] Step A1. Substitute the convolution layer calculation formula into the bn layer calculation formula and then simplify the formula to fuse the target layer and the bn layer to obtain a fusion layer;

[0095] Step A1 can specifically include:

[0096] Step A11. Obtain the convolution layer calculation formula as follows:

[0097] Yconv = Wconv * X + Bconv (7);

[0098] where Yconv is the output of the target layer, Wconv is the weight of the target layer convolution kernel, X is the input feature map, and Bconv is the bias of the target layer convolution kernel; if the target layer is the conv layer, then Yconv is the output of the convolution layer, Wconv is the weight of the convolution kernel, X is the input feature map, and Bconv is the bias of the convolution kernel. Since the dense layer is a special conv layer, it is similar to the conv layer.

[0099] Step A12, obtain the calculation formula of the BN layer as follows:

[0100]

[0101] where γ is the scaling factor of the BN layer, β is the offset of the BN layer, μ is the mean in units of the channels of the output feature map of the target layer, and σ 2 is the variance in units of the channels of the output feature map of the target layer, and ε is a very small constant used to avoid being 0, and Y bn is the output of the BN layer; similarly, if the target layer is a conv layer, μ is the mean in units of the channels of the output feature map of the conv layer, and σ 2 is the variance in units of the channels of the output feature map of the conv layer.

[0102] Step A13, substitute the calculation formula of the convolutional layer into the calculation formula of the BN layer to obtain:

[0103]

[0104] That is, substitute formula (7) into formula (8) to obtain formula (9) to achieve the fusion of the target layer and the BN layer.

[0105] Step A14, perform formula simplification to obtain:

[0106]

[0107] Step A15, let to obtain:

[0108] Ybn = Wmerge * X + Bmerge (11);

[0109] where Wmerge and Bmerge are the convolutional kernel weights and biases of the fusion layer respectively. If the target layer is a convolutional layer conv, the fusion layer is also a conv layer.

[0110] It can be seen from formulas (7) to (11) that the conv layer and the BN layer can be fused into a conv layer, and the parameters of this conv layer are the fused parameters, that is, using a fused conv layer achieves the effects of the conv layer and the BN layer before fusion. The present invention verifies through experiments that the fusion of the conv layer and the BN layer has no impact on the accuracy. The fusion of the dense layer and the BN layer is the same as that of the conv layer and the BN layer, and will not be elaborated here.

[0111] Step A2, fuse the fusion layer and the subsequent Relu layer by using the scaling factor and zero point corresponding after the output of the Relu layer as the scaling factor and zero point of the target layer.

[0112] For ease of understanding, the following takes the target layer as the conv layer as an example to illustrate how to achieve the fusion of conv and Relu. The dense layer is similar and will not be described separately.

[0113] Specifically, in quantization, a structure like Conv+Relu is generally also merged into one Conv for operation, which cannot be achieved in a full-precision model (float32).

[0114] The same scale and zero_point should be used before and after Relu. Since Relu itself does not perform any mathematical operations and is just a truncation function, using different scales and zero_points will cause inability to dequantize back to the float domain.

[0115] For example Figure 5 , assume the numerical range before Relu is r in ∈[-α,α], α>0, then the numerical range after passing through Relu is r out ∈[0,α], quantized to the int8 type, i.e., [-128,127]. Then the scale and zero_point before Relu are S in =2α / 255, Z in =0, and the scale and zero_point after Relu are S out =α / 255, Z out =-128. Assume again that the floating-point number before Relu is r in =α / 2, then the value after passing through Relu is still r out =α / 2, quantized to the int8 type, the int8 integer before Relu is 64. Since Zin = 0, the value after passing through Relu is still 64. However, Sout and Zout have changed, so the dequantized r out ≠α / 2. Therefore, if you want to ensure the consistency between the quantized Relu and the floating-point Relu, you must ensure that Sin, Sout, Zin, and Zout are consistent.

[0116] However, ensuring the consistency of the scale and zero_point before and after Relu does not stipulate that Sin and Zin must be used. The present invention can also use Sout and Zout after Relu. However, the meaning of using which scale and zero_point is completely different. If Sout and Zout after Relu are used, the function of Relu can be achieved by using the truncation function of quantization itself.

[0117] As can be seen from Equation (5), in int8 quantization, in addition to the round operation of rounding, there is also a clip operation that limits the data to [-128, 127], and the essence of Relu is to perform clip. Therefore, the quantization truncation function can be used to simulate the function of Relu.

[0118] Such as Figure 6 the structure of Conv+Relu in the upper part, where the numerical range after Conv is r in ∈[-α,α]. Before using Relu, S in and Z in need to go through quantization, Relu after quantization, and dequantization to obtain r out ∈[0,α]. For example, if r in =-α / 2, after int8 quantization, q in =-64, because Z in =0, so after Relu, qout = 0, and after dequantization, rout = 0.

[0119] If using S out and Z out after using Relu, then as Figure 6 the lower part, only data with r in ∈[0,α] will be quantized to [-128, 127], and after dequantization, r out ∈[0,α]. Data with r in ∈[-α,0] will be quantized to qout = -128, and after dequantization, rout = 0.

[0120] Therefore, for the conv layer, using the scaling factor after the output of the Relu layer as the scaling factor of the conv layer and using the zero point after the output of the Relu layer as the zero point of the conv layer can fuse the conv layer after being fused with the bn layer and the subsequent Relu layer. The same is true for the dense layer, which will not be elaborated here.

[0121] Through the above processing, the third PointNet_Lite network can be obtained.

[0122] S4. Deploy the third PointNet_Lite network on an edge device. The cache structure of the edge device includes off-chip storage and on-chip storage of a neural network accelerator. The on-chip storage includes unified cache, feature map cache, weight cache, multiplier-accumulator array, temporary cache, and output cache; and the multiplier-accumulator array is designed to be of atomic_c*atomic_k dimension to implement parallelized calculation of point convolution;

[0123] Among them, the dimension of the network input image is n*1*c, where atomic_c and atomic_k represent adaptively selected parts of c and k. k represents the number of convolution kernels.

[0124] The third PointNet_Lite network in the embodiments of the present invention needs to be deployed on edge devices such as FPGAs. Specifically, network parameters such as weights are extracted and deployed on edge devices to construct the network. The storage and exchange of activation parameters and weight parameters of the network have a significant impact on the efficiency of the system. As Figure 7 shown in the schematic diagram of the cache structure, the cache structure of edge devices such as FPGAs includes off-chip storage and on-chip storage of the neural network accelerator; the off-chip storage is located outside the neural network accelerator, and the off-chip storage directly exchanges data with the CPU; the on-chip storage includes unified cache, feature map cache, weight cache, multiplier-accumulator array, temporary cache, and output cache;

[0125] All parameters and point cloud data required by the neural network accelerator are initially placed on the off-chip storage. When the neural network accelerator starts working, the parameters and point cloud image data to be used are transported to the unified cache. When performing the first-layer conv calculation, the weight parameters and point cloud data of this layer are respectively transported from the unified cache to the feature map cache and the weight cache, and then the weight parameters and point cloud data enter the multiplier-accumulator array for calculation. The intermediate data is stored in the temporary cache. When a batch of output points are calculated, the output data enters the output cache.

[0126] Generally, the output of the previous layer of the convolutional neural network is used as the input of the next layer. If the output cache is large enough to cache the entire output feature map of one layer, it can directly enter the feature map cache as the input of the next level without entering the unified cache, which can reduce the data exchange between the output cache, the unified cache, and the off-chip storage. It should be noted that such a cache structure requires the output cache to be large enough to cache the largest output feature map in the current PointNet_Lite network.

[0127] Therefore, in the embodiments of the present invention, the cache is designed to be large enough to meet the above requirements. The subsequent layers after the first layer are processed according to the above process, which will not be specifically described here. This can reduce the data exchange between the output cache, the unified cache, and the off-chip storage, and reduce the cache capacity and computational complexity.

[0128] Moreover, in the embodiments of the present invention, the multiplier-accumulator array is specifically designed. The following specifically describes its design concept.

[0129] For ease of understanding, it is set that the dimension of the network input image in the embodiments of the present invention is n*1*c, where n corresponds to the image height, 1 corresponds to the image width, and c corresponds to the number of channels.

[0130] As Figure 8As shown, the point convolution calculation determines the multi-dimensions of the output feature map due to the multiple dimensions of the input feature map and the number of convolutional kernels (the number of convolutional kernels in any layer is denoted as k). Using a multiply-accumulate array or a systolic array composed of multiple multipliers can significantly improve the calculation efficiency. In the PointNet network, there are only point convolutions and fully connected layers. It is necessary to select a suitable multiply-accumulate array or systolic array in combination with the calculation rules of point convolutions and fully connected layers. How to make full use of the multiply-accumulate array and reduce the ineffective calculations of multipliers becomes the key to the selection.

[0131] The systolic array has obvious effects in calculating matrix multiplication, and it calculates while the data flows in the time dimension. As Figure 8 shown, the essence of the systolic array calculating matrix multiplication is the parallelism in the atomic_n and atomic_k dimensions in space and the parallelism in the atomic_c dimension in time. At the same time, the number of convolutional kernels atomic_k calculated simultaneously determines the parallelism of the output channels, and the input feature map atomic_n calculated simultaneously determines the number of points calculated simultaneously for each channel of the output feature map. Atomic means automatically selected.

[0132] Considering that the point convolution and the fully connected layer use a unified multiply-accumulate array, and the dimension n of the fully connected layer is 1, the data of the fully connected layer is placed on the c dimension. When using a systolic array for calculation, only one group of multipliers in the atomic_n dimension will work. The utilization rate of using a systolic array to calculate the fully connected layer is very low, and the utilization rate is 1 / atomic_n.

[0133] Therefore, the embodiment of the present invention uses a flexible multiply-accumulate array to calculate point convolutions and fully connected layers.

[0134] As Figure 9 shown, the present invention designs a multiply-accumulate array of atomic_c * atomic_k dimensions. In the spatial dimension, atomic_c data are calculated simultaneously on the c dimension of the input feature map as shown in Figure 8 shown, and atomic_k convolutional kernels are calculated simultaneously. On the output feature map, atomic_k output channels are calculated simultaneously, and one point is calculated for each channel. At one point, atomic_c multiplication operations are performed simultaneously. The output of the multiplication array is the intermediate result (partial sum), and the entire output feature map is calculated in the time dimension.

[0135] When calculating the fully connected layer, the fully connected data is placed on the dimension c, and the multiply-accumulate array of atomic_c * atomic_k can be fully utilized to improve the utilization rate of the multiply-accumulate array. Atomic_c and atomic_k represent the self-adaptively selected parts of c and k, and the specific values can be set according to needs.

[0136] It can be seen that in the embodiment of the present invention, by designing the multiplier-accumulator array to be of the atomic_c*atomic_k dimension, parallelized calculation of point convolution can be achieved, the utilization rate of the multiplier-accumulator array can be improved, and the calculation efficiency can be improved.

[0137] Please refer to Figure 10 the flowchart of, and understand the main process of the lightweight and hardware acceleration method for point cloud image classification described in S1 to S4. Figure 10 Among them, preparing the dataset means preparing the point cloud image dataset for network training; simplifying the network means removing the T-Net network in the PointNet network; adding the softmax layer means connecting the softmax at the network output end for normalization processing; extracting network parameters means extracting network parameters such as weights; designing the hardware, storage, and computing structure means designing the cache structure of the edge device and the multiplier-accumulator array. For the classification task with the same effect, the present invention ensures the classification accuracy of the network when simplifying the point cloud classification task, and deploys the simplified network on the hardware accelerator, which can greatly reduce the hardware pressure while achieving the same classification effect. For specific details, please refer to the relevant content above and will not be elaborated here.

[0138] In summary, the lightweight and hardware acceleration method for point cloud image classification provided by the embodiment of the present invention is the lightweight and hardware acceleration implementation of the point cloud image classification convolutional neural network PointNet. Through in-depth optimization of the PointNet model structure and careful adjustment of parameter configuration, the method has successfully achieved efficient deployment on FPGA or ASIC. Specifically, the present invention uses means such as model pruning, full integer quantization, and layer fusion processing to significantly reduce the number of model parameters and the computational complexity. At the same time, through the specific design of the cache structure of the edge device and the multiplier-accumulator array therein, hardware cache calculation and parallelized calculation of point convolution are realized, reducing the hardware logic design and cache capacity. While maintaining the performance of the original PointNet model, the model complexity and resource requirements are reduced, and the accuracy and reliability of point cloud processing are effectively guaranteed. Moreover, it is deployed on the FPGA platform to achieve the hardware acceleration effect of the point cloud neural network in a software-hardware co-operative manner.

[0139] Furthermore, the lightweight and hardware acceleration method for point cloud image classification may further include:

[0140] Reducing the number of point convolution kernels used in the point convolution in the five-layer MLP structure of the third PointNet_Lite network to reduce the number of features for each point dimension increase, and performing network performance testing to determine the MLP dimension increase strategy that meets the network performance requirements.

[0141] This method is to redesign the MLP dimension increase strategy, and the following specifically explains its design concept.

[0142] In PointNet, the dimensionality of point cloud features is increased through a multi-layer perceptron (MLP). As Figure 3 shown, a five-layer multi-layer perceptron mlp(64, 64, 64, 128, 1024) is used to increase the dimensionality of point cloud features, and 64, 64, 64, 128, and 1024 point convolution kernels are used for operations respectively. However, this way of increasing dimensionality introduces a large amount of computational complexity and the number of parameters, increasing the complexity of the model and the difficulty of hardware implementation. To further reduce the computational complexity and the number of parameters, the present invention proposes a simplification strategy for the MLP. For the third PointNet_Lite network, without changing the network structure and connection relationship, the number of point convolution kernels used in the five-layer multi-layer point convolution is gradually reduced to reduce the number of features for each point to increase dimensionality, and the impacts on computational complexity, the number of parameters, and accuracy are analyzed by comparison to study the feasibility and reliability of the simplified MLP dimensionality increase strategy.

[0143] As shown in Table 1, the present invention proposes 8 mlp dimensionality increase strategies and builds a network for training. When the number of point convolution kernels used n = 2048, the number of parameters and the computational complexity are calculated. During the process of gradually reducing the mlp dimensionality increase dimension, both the number of parameters and the computational complexity decrease significantly. For the first 5 networks in terms of serial number, there is no significant difference in accuracy. For the networks with serial numbers 6 - 8, due to the too low dimensionality increase dimension, the accuracy decreases significantly.

[0144] Table 1 Computational complexity, number of parameters, and accuracy of different dimensionality increase strategies when n = 2048

[0145]

[0146]

[0147] Note: 1) mlp(64, 64, 64, 128, 1024) is the third PointNet_Lite network.

[0148] 2) For the 5th and 8th networks, one fully connected layer is reduced due to the too low dimensionality increase.

[0149] Therefore, by reducing the number of features for each point to increase dimensionality and conducting network performance tests, at least one MLP dimensionality increase strategy that meets the network performance requirements can be selected, thereby optimizing and determining appropriate network structure parameters for the third PointNet_Lite network.

[0150] Furthermore, the lightweight and hardware acceleration method for point cloud image classification may further include:

[0151] For the MLP structure after adopting the MLP dimension elevation strategy, when implementing the conv operation in hardware, adjust the operation order, and adopt the calculation order of local first and then global to replace the layer-by-layer calculation, so as to achieve MLP hardware acceleration.

[0152] This method can be executed after redesigning the MLP dimension elevation strategy, and the processing method can be simply referred to as the hardware acceleration method of the mlp with local first and then global. The following specifically explains its design concept.

[0153] In the PointNet series of networks, mlp is used as a method for point cloud feature dimension elevation. Mlp introduces a large amount of computational complexity and number of parameters into the network. In hardware design, highly parallel computing units are used to relieve the computational pressure, but a larger output cache is required to cache intermediate variables. The essence of mlp is continuous point convolution. The calculation process of point convolution is similar to matrix multiplication. The calculation of point convolution has an additional bias calculation compared to matrix multiplication, and is the same as matrix multiplication in terms of multiplication operation. In the calculation process of neural networks, generally layer-by-layer calculation is adopted, where the previous convolutional layer calculates first, and after the calculation is completed, the activation is input to the next convolutional layer for calculation. Such calculation adopts the Figure 7 cache structure shown. To reduce the data exchange with off-chip storage, the on-chip output cache needs to adapt to the size of the largest activation layer in the PointNet_Lite network. The output activation of the conv5 layer in Table 2 is 2M.

[0154] Table 2 Computational complexity and number of parameters of the PointNet_Lite network with n = 2048

[0155]

[0156]

[0157] As Figure 11 shown, for the calculation of continuous point convolution, it is not necessary to calculate all the output activations of the previous layer to calculate the next layer of point convolution. This enables the cached data to be consumed in a timely manner without waiting for all the activation outputs to be calculated before being consumed. Therefore, the output cache in the Figure 7 cache structure can be allocated according to the design intention. From the perspective of point cloud features, the 3 coordinate features of n points respectively pass through 5 layers of point convolution for dimension elevation, and then the effect of extracting global features is the same as first performing 5 layers of point convolution dimension elevation on some points, and then sequentially performing 5 layers of point convolution dimension elevation on the remaining points until all points are completed for dimension elevation, and then extracting global features. According to the calculation property of mlp continuous point convolution, adjusting the calculation order of mlp continuous point convolution and adopting the calculation order of local first and then global to replace the layer-by-layer calculation can consume while calculating, and can reduce the hardware cache size.

[0158] In summary, the embodiments of the present invention can further redesign the MLP dimension-raising strategy and adopt a hardware acceleration method for MLP that is local-first and then global, thereby further reducing the number of parameters and the amount of computation and decreasing the output cache.

[0159] To facilitate understanding of the lightweight and hardware acceleration methods for point cloud image classification proposed by the present invention, the following presents three specific embodiments.

[0160] (1) Embodiment 1

[0161] After removing two T-Net networks based on PointNet and then training, the first PointNet_Lite network is obtained. Then, the entire network is quantized with full integers to obtain the second PointNet_Lite network; afterwards, fusion processing of the Relu / dense+bn+Relu layers is performed, and softmax is connected to the output of the network for normalization processing, obtaining the Figure 12 third PointNet_Lite network shown, also known as the PointNet Lite1 network.

[0162] The input of the third PointNet_Lite network is the position coordinates (x, y, z) of n point cloud data in space, with a size of n*3. The input is subjected to point convolution by the mlp twice consecutively to increase the feature number of each point to 64 and 64, and then subjected to point convolution by the mlp three times consecutively to increase the feature number of each point to 64, 128, and 1024. After that, the global feature is obtained through the max pooling layer, and then 40 classifications are output through the mlp three times consecutively by the dense layer. Finally, the probabilities of each classification are output after normalization through the softmax layer. The classification accuracy of the third PointNet_Lite network is 87.31% (decreasing by approximately 0.5% after int8 quantization), the number of parameters is 815KB, and the amount of computation is 303MB, where the conv / dense structure in the mlp is a fused single-layer structure.

[0163] In the cache design for hardware implementation, the cache structure as shown in Figure 7 is adopted, where the output cache is 2MB;

[0164] As an example of the atomic_c*atomic_k multiplier-accumulator array shown in Figure 9 , in an optional implementation manner, the multiplier-accumulator array is designed to be 16*16 dimensions to achieve parallel computing with a parallelism of 16 on the input channels and a parallelism of 16 on the convolution kernels, taking into account the parallel computing of point convolution and the fully connected layer.

[0165] (2) Embodiment 2

[0166] Based on Embodiment 1, the mlp dimensionality increase strategy is simplified, and Network No. 5 shown in Table 1 is adopted to obtain the Figure 13 PointNet_Lite2 network as shown.

[0167] The input of the PointNet_Lite2 network is the position coordinates (x, y, z) of n point cloud data in space, with a size of n*3. After the input passes through two consecutive point convolutions of the mlp to increase the feature number of each point to 32 and 32, and then passes through three consecutive point convolutions of the mlp to increase the feature number of each point to 32, 64, and 256, the global feature is obtained through the max pooling layer, and then 40 classifications are output through three consecutive dense layers of the mlp. Finally, the probabilities of each classification are output after normalization through the softmax layer. The classification accuracy of the PointNet_Lite2 network is 87.56% (decreasing by about 0.5% after int8 quantization), the number of parameters is 65KB, and the computational complexity is 42MB, where the conv / dense structure in the mlp is a fused single-layer structure.

[0168] In the cache design for hardware implementation, the cache structure as shown in Figure 7 is adopted, where the output cache is 0.5MB, the overall network computational complexity is reduced, and the use of a 16*16 multiplier-accumulator array also reduces the calculation time.

[0169] (3) Embodiment 3

[0170] Based on Embodiment 2, when implementing the conv operation in hardware, the operation order is adjusted, and the calculation order of first local and then global is adopted to replace the layer-by-layer calculation. According to the properties of consecutive point convolutions in Figure 11 , 16 points are first dimensionally increased, and then the other points are dimensionally increased. In the cache design for hardware implementation, the cache structure as shown in Figure 7 is adopted, where the output cache is 4KB, greatly reducing the area of the output cache.

[0171] In summary, in the solution of the present invention, an innovative design strategy specifically for the lightweight PointNet network is proposed. The lightweight processing of the PointNet network is realized to efficiently implement it on resource-constrained hardware platforms (such as FPGA or ASIC).

[0172] The present invention reduces the number of parameters and computational complexity by trimming and simplifying PointNet, adopts the fusion processing of the full integer quantization layer, significantly reduces the hardware logic design and cache capacity, reduces the cache capacity by adjusting the calculation order according to the structure of consecutive point convolutions, and uses a multiplication array such as 16*16 to calculate point convolutions in parallel, greatly improving the utilization rate of the multiplication array by the fully connected layer during calculation.

[0173] It is particularly worth mentioning that while this method effectively retains the advantages of high accuracy and strong robustness of PointNet in point cloud data processing, it greatly reduces the resource consumption and power consumption levels during the hardware implementation process, providing an extremely innovative, efficient, and practical solution for 3D data processing on edge computing devices.

[0174] With the technical power contained in the present invention, in practical application scenarios such as autonomous driving, drone control, and robot operation, it is possible to smoothly achieve real-time and accurate perception and efficient processing of the surrounding environment, strongly promoting the overall improvement of the system's intelligence level and response speed. Especially in the embedded system environment with relatively scarce resources, the method of the present invention will fully demonstrate its unique value and important role, injecting strong impetus and solid support for the implementation and popularization of point cloud processing technology in a wider range of practical application fields.

[0175] It should be noted that in the description of the present invention, the terms "first" and "second" are only used for descriptive purposes and cannot be understood as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, the features defined with "first" and "second" may explicitly or implicitly include one or more of such features. In the description of the present invention, "a plurality" means two or more, unless otherwise specifically defined.

[0176] In the description of this specification, the description referring to terms such as "one embodiment", "some embodiments", "example", "specific example", or "some examples" means that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described can be combined in a suitable manner in any one or more embodiments or examples. In addition, those skilled in the art can combine and combine the different embodiments or examples described in this specification.

[0177] The above description is only a preferred embodiment of the present invention and is not intended to limit the protection scope of the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention are included in the protection scope of the present invention.

Claims

1. A lightweight and hardware accelerated method for point cloud image classification, characterized in that: include: After removing the T-Net network from the PointNet network used for point cloud image classification, training is performed to obtain the first PointNet_Lite network; Perform full integer quantization on the first PointNet_Lite network to obtain a second PointNet_Lite network; The second PointNet_Lite network is subjected to layer fusion processing to obtain a third PointNet_Lite network; wherein the layer fusion processing includes fusing the target layer with the bn layer in the mlp structure of the second PointNet_Lite network and then fusing it with the Relu layer, and the target layer includes a conv layer and a dense layer; The third PointNet_Lite network is deployed on an edge device, and the cache structure of the edge device includes off-chip storage and on-chip storage of a neural network accelerator, and the on-chip storage includes a unified cache, a feature map cache, a weight cache, a multiplier array, a temporary cache, and an output cache; and the multiplier array is designed to be atomic_c*atomic_k in dimension to realize parallel calculation of point convolution; wherein the dimension of the network input image is n*1*c, and atomic_c and atomic_k represent adaptive selection of parts c and k.

2. The lightweight and hardware accelerated method for point cloud image classification according to claim 1, characterized in that: The edge device includes FPGA and ASIC.

3. The lightweight and hardware accelerated method for point cloud image classification according to claim 1, characterized in that: The first PointNet_Lite network is fully integer quantized, including: The weights of the first PointNet_Lite network are quantized to int8 symmetric, the biases are quantized to int32 symmetric, and the input feature maps and output feature maps are quantized to int8 asymmetric to achieve full integer quantization.

4. The lightweight and hardware accelerated method for point cloud image classification according to claim 3, characterized in that: The weight parameters of the first PointNet_Lite network are quantized by int8 symmetric quantization, the bias parameters are quantized by int32 symmetric quantization, and the input feature map and the output feature map are quantized by int8 asymmetric quantization to achieve full integer quantization, including: The input feature map of each conv layer uses a unified scaling factor scale and zero point zero_point, denoted as s1 and z1; each convolution kernel weight in the conv layer uses a separate scale and zero_point, denoted as s2 and z2=0; the bias uses a unified scale and zero_point, denoted as sb=s1*s2 and zb=0, and the output feature map of the conv layer uses a unified scale and zero_point, denoted as s3 and z3; the weights in each dense layer use a unified scale and zero_point, and the rest of the operations are the same as the conv layer operations.

5. The lightweight and hardware accelerated method for point cloud image classification according to claim 3 or 4, characterized in that: After performing full integer quantization on the first PointNet_Lite network, the method further includes: For the network after full integer quantization, softmax is connected at the output end for normalization to obtain the second PointNet_Lite network.

6. The lightweight and hardware accelerated method for point cloud image classification according to claim 1, characterized in that: For the target layer, the layer fusion processing process includes: By substituting the convolutional layer calculation formula into the BN layer calculation formula and simplifying the formula, the target layer and the BN layer are fused to obtain the fusion layer; The fusion layer is fused with the subsequent Relu layer by using the scaling factor and zero point correspondence after the Relu layer output as the scaling factor and zero point of the target layer.

7. The lightweight and hardware accelerated method for point cloud image classification according to claim 6, characterized in that: The convolutional layer calculation formula is substituted into the BN layer calculation formula to simplify the formula, and the target layer and the BN layer are fused to obtain a fusion layer, including: The calculation formula for obtaining the convolutional layer is as follows: Yconv=Wconv*X+Bconv; Among them, Yconv is the output of the target layer, Wconv is the convolution kernel weight of the target layer, X is the input feature map, and Bconv is the convolution kernel bias of the target layer; The calculation formula for obtaining the bn layer is as follows: Among them, γ is the bn layer scaling factor, β is the bn layer offset, μ is the mean value in units of the target layer output feature map channel, and σ 2 is the variance in units of target layer output feature map channels, and ε is a very small constant used to avoid is 0, Ybn is the output of the bn layer; Substituting the convolutional layer calculation formula into the BN layer calculation formula, we get: Simplifying the formula, we get: make get: Ybn=Wmerge*X+Bmerge; Among them, Wmerge and Bmerge are the convolution kernel weight and bias of the fusion layer respectively.

8. The lightweight and hardware accelerated method for point cloud image classification according to claim 1, characterized in that: The method further comprises: Reduce the number of point convolution kernels used in the point convolution within the five-layer MLP structure in the third PointNet_Lite network to reduce the number of features of each point dimension increase, and perform network performance testing to determine the MLP dimension increase strategy that meets the network performance requirements.

9. The lightweight and hardware accelerated method for point cloud image classification according to claim 8, characterized in that: The method further comprises: For the MLP structure after adopting the MLP dimensionality increase strategy, the operation order is adjusted when the conv operation is implemented in hardware, and the calculation order of local first and then overall is adopted instead of layer-by-layer calculation to achieve MLP hardware acceleration.

10. The lightweight and hardware accelerated method for point cloud image classification according to claim 1, characterized in that: The multiplier-adder array is designed to be 16*16 dimensional to achieve parallel computing with a parallelism of 16 on the input channels and a parallelism of 16 on the convolution kernels.