Point cloud data accelerated processing method, device, electronic device, and storage medium
By dividing the processing of the cylindrical feature network into the first working group and the second working group, the acceleration framework of the cylindrical feature network is optimized, the problem of long-term inference in the existing technology is solved, and more efficient point cloud data processing is achieved.
Patent Information
- Application Number
- CN202111574050.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-12-21
- Publication Date
- 2025-06-06
- Estimated Expiration
- 2041-12-21
AI Technical Summary
The existing inference acceleration framework of cylindrical feature networks has the problem of low optimization efficiency and the inability to effectively reduce the interaction time-consuming between layers, resulting in long inference time.
By dividing the cylindrical feature network processing into a first working group and a second working group, the first feature data of the to-process point cloud data is calculated based on the first working group, and comprehensively process the first feature data based on the second working group to generate a pseudo-image, thereby optimizing the acceleration framework and reducing inference time.
By dividing the working groups and decomposing the computing tasks, the inference speed of the cylindrical feature network is significantly improved and the inference time is reduced.
Smart Images

Figure CN114241219B_ABST
Abstract
Description
Background Art
[0002] In recent years, with the rapid development of artificial intelligence technologies such as autonomous driving and robotics, point cloud learning has attracted more and more attention. As the main technology in the field of artificial intelligence, deep learning has been successfully used to solve various two-dimensional visual problems. However, deep neural networks face unique challenges in processing point clouds, and deep learning on point clouds is still in its infancy. Three-dimensional object detection on point clouds is usually achieved by three-dimensional convolution and projection to the front view or bird's-eye view. Among them, the projection to the front view or bird's-eye view is limited by the sparsity of the point cloud, which makes it difficult for convolution to extract features well and inefficient. The disadvantage of three-dimensional convolution is that the amount of calculation is large, resulting in a slow reasoning speed of the network. In order to speed up, the Pillar Feature Net (PFE) converts point cloud voxels into three-dimensional grids and does not split the columnar voxels on the vertical columns of the voxels, thereby removing the three-dimensional convolution operation. The Pillar Feature Net does not require manual coding, utilizes all the information of the point cloud, and has no parameters that need to be adjusted; and its operations are all two-dimensional convolutions, which are more efficient; at the same time, the Pillar Feature Net can also be migrated to other point cloud data.
[0003] The cylindrical feature network needs to convert the three-dimensional coordinates of the lidar into a two-dimensional pseudo image. However, there are a lot of non-convolutional calculations in the cylindrical feature network, and the speed in various inference frameworks is not ideal.
[0004] The current inference acceleration framework of the columnar feature network has the following problems: for complex addition, subtraction, multiplication and division operations, it will generate very complex operation graphs, and the optimization efficiency is low; for single-layer operator acceleration optimization, it cannot effectively reduce the interaction time between layers.
[0005] Therefore, how to optimize the acceleration framework of the columnar feature network and reduce the time consumption of inference is a technical problem that needs to be solved urgently by technical personnel in this field. Summary of the invention
[0006] In order to overcome the defects of the above-mentioned prior art, the present invention provides a point cloud data accelerated processing method, device, electronic device, and storage medium to optimize the acceleration framework of the cylindrical feature network and reduce the time consumption of reasoning.
[0007] According to one aspect of the present invention, a method for accelerating processing of point cloud data is provided, wherein the point cloud data is processed based on a cylindrical feature network, and the method comprises:
[0008] Based on the first working group, first feature data of the point cloud data to be processed is calculated, the feature data including radar features and image features, the point cloud data to be processed is discretized into a grid of a first coordinate plane to form a voxel columnar set, the voxel columnar set includes P columnar voxels, the first working group includes P first processing units, each of the first processing units executes a first kernel function to process a columnar voxel of the point cloud data to be processed, and obtains the radar feature of the columnar voxel;
[0009] Based on the second working group, the first feature data of the point cloud data to be processed is comprehensively processed to obtain second feature data, wherein the comprehensive processing includes one or more of a splicing process, a linear transformation process, a batch normalization process, an activation process, and a maximum value process;
[0010] A pseudo image of the point cloud data to be processed is obtained according to the second feature data.
[0011] In some embodiments of the present application, outputs of each first processing unit of the first working group are connected to a global memory, the first working group and the second working group are executed by a graphics processor, and the global memory is a shared memory of the graphics processor and a central processing unit.
[0012] In some embodiments of the present application, the second working group has a three-dimensional work queue, the second working group includes P second computing units for parallel computing, each of the second computing units includes 2*M second processing units for parallel computing, each of the second processing units executes a second kernel function to comprehensively process a row of radar features / image features in the matrix formed by the first feature data, and M is an integer greater than 0.
[0013] In some embodiments of the present application, in the second kernel function, the splicing process and the linear transformation process are combined into one operator.
[0014] In some embodiments of the present application, the linear transformation process includes:
[0015] multiplying the radar signature by the radar weight matrix;
[0016] multiplying the image features by the image weight matrix,
[0017] Among them, the number of rows of the radar weight matrix is the number of the radar features, and the number of columns of the radar weight matrix is 2*M; the number of rows of the image weight matrix is the number of the image features, and the number of columns of the image weight matrix is 2*M. The radar weight matrix, the image weight matrix and the input data of the second computing unit are stored in the local memory of the second computing unit.
[0018] In some embodiments of the present application, the radar weight matrix and the image weight matrix include the offset of the batch normalization.
[0019] In some embodiments of the present application, the first working group and the second working group are implemented based on OpenCL, and the local memory of the second computing unit in the second working group and the global memory are implemented using an asynchronous copy function of OpenCL.
[0020] According to another aspect of the present application, there is also provided a point cloud data accelerated processing device, wherein the point cloud data is processed based on a cylindrical feature network, comprising:
[0021] A first feature data calculation module is configured to calculate first feature data of the point cloud data to be processed based on the first working group, the feature data including radar features and image features, the point cloud data to be processed is discretized into a grid of a first coordinate plane to form a voxel columnar set, the voxel columnar set includes P columnar voxels, the first working group includes P first processing units, each first processing unit executes a first kernel function to process a columnar voxel of the point cloud data to be processed to obtain the radar feature of the columnar voxel;
[0022] A second feature data calculation module is configured to perform comprehensive processing on the first feature data of the point cloud data to be processed based on the second working group to obtain second feature data, wherein the comprehensive processing includes one or more of a splicing process, a linear transformation process, a batch normalization process, an activation process, and a maximum value process;
[0023] The pseudo image acquisition module is configured to obtain a pseudo image of the point cloud data to be processed according to the second feature data.
[0024] According to another aspect of the present invention, there is further provided an electronic device, comprising: a processor; and a storage medium storing a computer program, wherein the computer program executes the above steps when executed by the processor.
[0025] According to yet another aspect of the present invention, a storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the steps described above are executed.
[0026] Compared with the prior art, the advantages of the present invention are:
[0027] The cylindrical feature network processing is divided into a first working group and a second working group for processing, wherein the first feature data of the point cloud data to be processed is calculated based on the first working group, and the first working group is split into multiple one-dimensional first processing units to execute the first kernel function, the first feature data of the point cloud data to be processed is comprehensively processed based on the second working group, and a pseudo image of the point cloud data to be processed is obtained according to the second feature data, thereby optimizing the acceleration framework of the cylindrical feature network and reducing the inference time. BRIEF DESCRIPTION OF THE DRAWINGS
[0028] The above and other features and advantages of the present invention will become more apparent by describing in detail exemplary embodiments thereof with reference to the attached drawings.
[0029] Figure 1 A flowchart of a method for accelerating processing of point cloud data according to an embodiment of the present invention is shown;
[0030] Figure 2 A flowchart of column feature network processing according to an embodiment of the present invention is shown;
[0031] Figure 3 A flowchart of a first kernel function according to an embodiment of the present invention is shown;
[0032] Figure 4 A calculation schematic diagram of a second working group according to an embodiment of the present invention is shown;
[0033] Figure 5 A module diagram of a point cloud data acceleration processing device according to an embodiment of the present invention is shown;
[0034] Figure 6 A schematic diagram of a computer-readable storage medium in an exemplary embodiment of the present disclosure is schematically shown;
[0035] Figure 7 A schematic diagram of an electronic device in an exemplary embodiment of the present disclosure is schematically shown. DETAILED DESCRIPTION
[0036] Example embodiments will now be described more fully with reference to the accompanying drawings. However, example embodiments can be implemented in a variety of forms and should not be construed as limited to the examples set forth herein; rather, these embodiments are provided so that the disclosure will be more comprehensive and complete and to fully convey the concepts of the example embodiments to those skilled in the art. The described features, structures, or characteristics may be combined in any suitable manner in one or more embodiments.
[0037] In addition, the accompanying drawings are only schematic illustrations of the present disclosure and are not necessarily drawn to scale. The same reference numerals in the figures represent the same or similar parts, and their repeated description will be omitted. Some of the block diagrams shown in the accompanying drawings are functional entities and do not necessarily correspond to physically or logically independent entities. These functional entities can be implemented in software form, or implemented in one or more hardware modules or integrated circuits, or implemented in different networks and / or processor devices and / or microcontroller devices.
[0038] In order to solve the defects of the prior art, the present invention provides a method for accelerating the processing of point cloud data, wherein the point cloud data is processed based on a cylindrical feature network. The process 210 of the cylindrical feature network processing can be seen in Figure 2 . During the process of columnar feature network processing, point cloud data will be obtained first, and each point cloud can be represented by three-dimensional coordinates and reflection intensity (x, y, z, i). The obtained point cloud data is discretized into a uniformly spaced grid (H*W) of the first coordinate plane formed by the x-axis and the y-axis, so that a voxel columnar set can be generated. The voxel columnar set includes P columnar voxels. In addition, the corresponding image can also be pre-fused to obtain the category probability of each pixel through the segmentation network. In this embodiment, taking 7 categories (such as background, pedestrians, motor vehicles, non-motor vehicles, etc.) as an example, the image features of each pixel include the category probabilities of 7 categories and RGB values, a total of 10 dimensional features. Furthermore, the pixels of the image and the point cloud data can be associated with each other.
[0039] exist Figure 2 In the illustrated embodiment, point cloud data [P][N][C], center point coordinates [P][4], and the number of points [P] can be input into the cylindrical feature network, where P is the number of cylindrical voxels (further, P is the number of non-empty cylindrical voxels), N is the number of points in the cylindrical voxels, and C is the number of channels. Thus, the point cloud data is three-dimensional data of P*N*C. The center point coordinates [P][4] are the coordinates (x, y, z, i) of the center points of P cylindrical voxels, i.e., P*4 two-dimensional data. The number of points [P] is the number of points in the P cylindrical voxels, i.e., one-dimensional data.
[0040] Based on the point cloud data [P] [N] [C] and the number of points [P], the cluster distance of each point can be calculated. The cluster distance is the distance between the point and the arithmetic mean of the points in the columnar voxel, which can be expressed as (x c ,y c , z c). Thus, the cluster distance data of each point of P columnar voxels can be obtained, thereby forming the three-dimensional cluster distance data of [P][N][3]. According to the point cloud data [P][N][C] and the center point coordinates [P][4], the center point offset of each point can be calculated. The center point offset is the offset between the point and the center point in the columnar voxel, which can be expressed as (x p ,y p , z p ). Thus, the center point offset data of each point of P columnar voxels can be obtained, thereby forming three-dimensional center point offset data of [P][N][3].
[0041] According to the cluster distance (x c ,y c , z c ), the offset of the center point (x p ,y p , z p ) and the coordinates of the point (x, y, z, i) can obtain the 10-dimensional radar features of the point.
[0042] The 10-dimensional radar feature and the 10-dimensional image feature are spliced to obtain the fused first feature data, and the first feature data is normalized. At the same time, a filling mask is calculated according to the number of points of each columnar voxel. The filling mask is used to sample columnar voxels with a point number greater than N and fill the points with a point number less than N (for example, with 0). The normalized first feature data is multiplied by the filling mask to obtain the first feature data [P][N]
[20] .
[0043] Then, the radar features [P][N][0-9] and image features [P][N][10-19] in the first feature data [P][N]
[20] are linearly transformed using the radar weight matrix
[10]
[64] and the image weight matrix
[10]
[64] , and then concatenated, activated, and maximized to obtain the second feature data [P]
[128] . After encoding, the second feature data [P]
[128] can be spread back to the original pillar positions to create a pseudo image of size (H, W, C).
[0044] from Figure 2 It can be seen that PFE calculation can be divided into two parts. The first part is data processing, including calculating the distance to the geometric center of the effective point (cluster distance), calculating the distance to the voxel center (center point offset), etc. The data processing mainly uses basic multiplication and division. The second part calls one or more neural network operators in the comprehensive processing (such as linear transformation operators (fully connected operators), splicing operators, activation operators, maximum operators, batch normalization operators, etc.).
[0045] According to the above analysis, the calculation of PFE can be divided into two distinct units. The first part has less calculation amount, more complex branch logic, poor automated model optimization effect, and cannot effectively splice operators. The second part is basically a neural network operator with large calculation amount and simple logic.
[0046] The main input data of PFE is a three-dimensional tensor, and the highest dimension is the number of non-empty voxels P, which can be regarded as the batch size. The data of each batch has no dependencies, which provides good conditions for parallel computing of graphics processors.
[0047] Therefore, the present application can use the OpenCL framework to accelerate the calculation of PFE. The OpenCL framework may include a computing device, a computing unit and a processing unit. The computing device may include multiple computing units, and the computing unit may include multiple processing units. One or more computing devices may be connected to a host, and the host processes the computing instructions. Among them, the processing unit is the most basic processing unit, and each computing instruction is ultimately processed by the processing unit. The processing unit is located in the computing unit. Each processing unit can be configured as private memory, the processing units on the same computing unit can share local memory, and the entire computing device can access global memory. The access speed of private memory is greater than the access speed of the local memory, and the access speed of local memory is greater than the access speed of global memory.
[0048] The execution model of OpenCL is that the application manages the kernel program on the OpenCL device (computing device) through the host. The model can be divided into two modules: one is the management program (Hostprogram) executed on the host, and the other is the program executed on the computing unit, also known as the kernel function. Before executing the kernel function, you can first establish an index space to identify each node in the device. Each node will execute the same kernel function program. Each processing unit can have a local ID in the computing unit, a group ID in the work group, and a global ID in the global. OpenCL can use the NDRange function to define this index space. The kernel function program can obtain the ID of each dimension through the functions get_global_id(dim), get_group_id(dim), and get_local_id(dim) to distinguish the data to be processed.
[0049] See below Figure 1 , Figure 1 A flow chart of a method for accelerating processing of point cloud data according to an embodiment of the present invention is shown. Figure 1 The steps are as follows:
[0050] Step S110: Based on the first working group, first feature data of the point cloud data to be processed is calculated, the feature data includes radar features and image features, the point cloud data to be processed is discretized into a grid of a first coordinate plane to form a voxel columnar set, the voxel columnar set includes P columnar voxels, the first working group includes P first processing units, each first processing unit executes a first kernel function to process a columnar voxel of the point cloud data to be processed to obtain the radar feature of the columnar voxel.
[0051] Step S120: Based on the second working group, the first feature data of the point cloud data to be processed is comprehensively processed to obtain second feature data, wherein the comprehensive processing includes one or more of splicing processing, linear transformation processing, batch normalization processing, activation processing and maximum value processing.
[0052] Step S130: obtaining a pseudo image of the point cloud data to be processed according to the second feature data.
[0053] In the point cloud data acceleration processing method provided by the present invention, the cylindrical feature network processing is divided into a first working group and a second working group for processing, wherein the first feature data of the point cloud data to be processed is calculated based on the first working group, and the first working group is split into a plurality of one-dimensional first processing units to execute the first kernel function, and the first feature data of the point cloud data to be processed is comprehensively processed based on the second working group, and a pseudo image of the point cloud data to be processed is obtained according to the second feature data, thereby optimizing the acceleration framework of the cylindrical feature network and reducing the inference time.
[0054] Specifically, the present application fuses the PFE operator into two OpenCL kernel functions, a first kernel function PFE_Preprocess and a second kernel function LinearCatReluMax, wherein the output of the first kernel function is the input of the second kernel function.
[0055] The first kernel function can be declared in actual use through the following code:
[0056] __kernel void PFE_Preprocess(__global const float*inVoxel,
[0057] __global const int* inCoords,
[0058] __global const int*inNumPoints,
[0059] __global float*outFeatures)
[0060] Among them, inVoxel is the input point cloud data; inCoords is the input coordinate data; inNumPoints is the number of input points, and outFeatures is the output data.
[0061] The second kernel function can be declared in actual use through the following code:
[0062] __kernel void LinearCatReluMax(__global float*inMat,
[0063] __global const float*lidarWeights,
[0064] __global const float*imageWeights,
[0065] __global const float*normBias,
[0066] __global float*outMat)
[0067] Among them, inMat is the input feature data; lidarWeights is the radar weight matrix; imageWeights is the image weight matrix, normBias is the normalization offset, and outMat is the output data.
[0068] The first kernel function PFE_Preproces integrates the operations of cluster distance calculation, center point offset calculation, feature fusion, filling mask acquisition, feature normalization and mask multiplication.
[0069] The second kernel function LinearCatReluMax integrates operators such as linear transformation processing, splicing processing, activation processing, maximum processing, and batch normalization processing.
[0070] By integrating the operations and operators of the first kernel function and the second kernel function, the data exchange of the graphics processor is greatly reduced and the speed is greatly improved.
[0071] For the first kernel function, refer to Figure 2 The flowchart of the corresponding operation in the figure is shown in Figure 1. Each operation needs to be calculated repeatedly based on the number of non-empty columnar voxels. Figure 2 The amount of calculation of each step corresponding to the first kernel function is not very large, and the calculation execution time accounts for a small proportion, so the split work items should not be too small. Therefore, in this application, the first working group is split into P processing units (P work items), and each processing unit processes all calculations of a columnar voxel.
[0072] Therefore, the host can control the workgroup as follows:
[0073] cl::NDRange global(P);
[0074] cl::NDRange local(1);
[0075] The workflow 220 of the first kernel function can be found in Figure 3 , the point cloud data [P][C], center point coordinates [4] and the number of points of a columnar voxel are taken as input data, and cluster distance calculation, center point offset calculation, splicing processing, feature normalization processing, filling mask acquisition and mask multiplication processing are performed according to the input data, so as to obtain the first feature data [N]
[20] of the columnar voxel.
[0076] Thus, each columnar voxel can be allocated to each processing unit of each graphics processor, and the intermediate result of the calculation step of the first kernel function does not need to be returned to the central processing unit, thereby reducing the data transmission time.
[0077] In some embodiments, the results of cluster distance calculation and center point offset calculation can be cached using local memory to facilitate multiple calls in subsequent steps and reduce input / output time.
[0078] In some embodiments, the data movement from the graphics processor, especially the integrated graphics card, to the central processing unit is relatively time-consuming. Therefore, the present application can connect the output of each first processing unit of the first working group to the global memory (the shared memory of the graphics processor and the central processing unit, such as DDR double data rate synchronous dynamic random access memory), so as to directly store the output of each first processing unit of the first working group in the global memory without using the local cache, thereby reducing bus conflicts. Compared with waiting for all the first processing units to complete the processing, the operation efficiency of uniformly writing from the local memory to the global memory is greatly improved.
[0079] See also Figure 4 , Figure 4 The calculation diagram of the second working group according to the embodiment of the present invention is shown. First, obtain Figure 3Output data: the first feature data of the columnar voxel
[32]
[20] , where
[32] [0-9] is the radar feature (as shown in the dashed box in 301), and
[32] [10-19] is the image feature (as shown in the dotted box in 301). Then the linear multiplication of the linear transformation operator is performed: the radar feature is multiplied by the radar weight matrix 302 to obtain the radar linear transformation processing matrix 304; the image feature is multiplied by the image weight matrix 303 to obtain the image linear transformation processing matrix 305. The radar linear transformation processing matrix 304 and the image linear transformation processing matrix 305 are merged together to obtain the fusion matrix 306, which is then output after batch normalization processing, activation processing (such as Relu calculation) and maximum value processing to obtain the second feature data 307 of the columnar voxel. The essence of the linear transformation processing operator is matrix multiplication, that is, a row of the first matrix is multiplied by a column of the second matrix to obtain a point in the output matrix. Figure 4 The dashed box in the middle is the computation flow of radar features, and the dashed box is the computation flow of image features. The kernel function of the second working group is mainly split by the granularity of feature calculation.
[0080] Specifically, the second working group can have a three-dimensional work queue, the second working group includes P second computing units for parallel computing, each of the second computing units includes 2*M second processing units for parallel computing, each of the second processing units executes a second kernel function to comprehensively process a row of radar features / image features in the matrix formed by the first feature data, and M is an integer greater than 0.
[0081] In this embodiment, since the shape of the first feature data is (batch, 32, 20), the shape of the two-way weight matrix is (10, 64). Each time a column of the weight matrix is calculated and multiplied by half a row of the first feature data, the last dimension of NDRange determines whether to process radar data or image data. Since the weight matrix has 64 columns, the second dimension of NDRange can be set to 64, and the global and local dimensions can be set to the same. The highest dimension of NDRange is still P (batch, that is, the number of columnar voxels).
[0082] Thus, NDRange can be set as follows:
[0083] cl::NDRange global(P,64,2);
[0084] cl::NDRange local(1,64,2).
[0085] Specifically, in the second kernel function, the splicing process and the linear transformation process are combined into one operator. The splicing process is to directly splice the image linear transformation process matrix 305 after the radar linear transformation process matrix 304. Therefore, the output address of the linear transformation process can be offset so that the linear transformation process output obtains the image linear transformation process matrix 305 and is stored adjacent to the radar linear transformation process matrix 304. In this way, the image linear transformation process matrix 305 and the radar linear transformation process matrix 304 can be stored in different locations, and the splicing process needs to be spliced by moving the memory.
[0086] Specifically, the number of rows of the radar weight matrix is the number of the radar features, and the number of columns of the radar weight matrix is 2*M; the number of rows of the image weight matrix is the number of the image features, and the number of columns of the image weight matrix is 2*M. The radar weight matrix, the image weight matrix, and the input data of the second working group can be stored in the local memory of the second computing unit. Thus, multiple accesses to the global memory are reduced. Furthermore, in the present application, the first working group and the second working group are implemented based on OpenCL, and the local memory of the second computing unit in the second working group is implemented with the OpenCL asynchronous copy function async_work_group_copy. Thus, data copying can be accelerated, and memory barriers are used for synchronization. After the data is moved between the local memory and the global memory, the kernel function operation can be started.
[0087] Specifically, the radar weight matrix and the image weight matrix may include the offset of the batch normalization, thereby the size scaling and translation in the batch normalization can be integrated into the weight matrix, thereby reducing one matrix multiplication.
[0088] In a test scenario, Intel OpenVINO is an acceleration library for neural network reasoning, which has outstanding performance on Intel processors. The following is a comparative test result of the method of this application and OpenVINO:
[0089] Testbed details:
[0090] - Processor: 11th Gen Core TM i7-1185GRE@2.8GHz x 8
[0091] Graphics card: Mesa Xe Graphics (TGL GT2)
[0092] Memory: 14.9 GiB
[0093] Operating system name: Ubuntu 20.04.2LTS
[0094] Intel OpenVINO version: openvino_2021.4.582
[0095] The following table lists the comparison of inference time for different numbers of columnar voxels.
[0096] Number of columnar voxels OpenVINO CPU time (milliseconds) OpenVINO GPU time consumption (milliseconds) This application takes time (milliseconds) 100 23.9ms 2.2ms 0.18ms 500 30.6ms 5.8ms 0.34ms 1000 40.5ms 10.7ms 0.59ms 2000 57.7ms 21.8ms 1ms 3000 73.8ms 40.8ms 1.6ms
[0097] It can be seen that the method of the present application can greatly improve the reasoning time of PFE.
[0098] The above are only a number of specific implementations of the point cloud data acceleration processing method of the present invention. Each implementation can be implemented independently or in combination, and the present invention is not limited to this. Furthermore, the flowchart of the present invention is only illustrative, and the execution order between the steps is not limited to this. The splitting, merging, order exchange, and other synchronous or asynchronous execution methods of the steps are all within the scope of protection of the present invention.
[0099] The present invention also provides a point cloud data acceleration processing device, Figure 5 The module diagram of the point cloud data acceleration processing device according to an embodiment of the present invention is shown. The point cloud data is processed based on a columnar feature network, and the point cloud data acceleration processing device 400 includes a first feature data calculation module 410, a second feature data calculation module 420 and a pseudo image acquisition module 430.
[0100] The first feature data calculation module 410 is configured to calculate first feature data of the point cloud data to be processed based on the first working group, the feature data including radar features and image features, the point cloud data to be processed is discretized into a grid of a first coordinate plane to form a voxel columnar set, the voxel columnar set includes P columnar voxels, the first working group includes P first processing units, each first processing unit executes a first kernel function to process a columnar voxel of the point cloud data to be processed to obtain the radar feature of the columnar voxel;
[0101] The second feature data calculation module 420 is configured to perform comprehensive processing on the first feature data of the to-be-processed point cloud data based on the second working group to obtain second feature data, wherein the comprehensive processing includes one or more of a splicing process, a linear transformation process, a batch normalization process, an activation process, and a maximum value process;
[0102] The pseudo image acquisition module 430 is configured to obtain a pseudo image of the point cloud data to be processed according to the second feature data.
[0103] In the point cloud data acceleration processing device provided by the present invention, the cylindrical feature network processing is divided into a first working group and a second working group for processing, wherein the first feature data of the point cloud data to be processed is calculated based on the first working group, and the first working group is split into a plurality of one-dimensional first processing units to execute the first kernel function, and the first feature data of the point cloud data to be processed is comprehensively processed based on the second working group, and a pseudo image of the point cloud data to be processed is obtained according to the second feature data, thereby optimizing the acceleration framework of the cylindrical feature network and reducing the inference time.
[0104] Figure 5 The point cloud data acceleration processing device 400 provided by the present invention is only schematically shown. Without violating the concept of the present invention, the splitting, merging and adding of modules are all within the protection scope of the present invention. The point cloud data acceleration processing device 400 provided by the present invention can be implemented by software, hardware, firmware, plug-in and any combination thereof, and the present invention is not limited thereto.
[0105] In an exemplary embodiment of the present disclosure, a computer-readable storage medium is also provided, on which a computer program is stored, and when the program is executed by, for example, a processor, the steps of the point cloud data acceleration processing method described in any of the above embodiments can be implemented. In some possible implementations, various aspects of the present invention can also be implemented in the form of a program product, which includes a program code. When the program product is run on a terminal device, the program code is used to enable the terminal device to execute the steps of various exemplary embodiments of the present invention described in the above point cloud data acceleration processing method section of this specification.
[0106] refer to Figure 6 As shown, a program product 800 for implementing the above method according to an embodiment of the present invention is described, which can adopt a portable compact disk read-only memory (CD-ROM) and include program code, and can be run on a terminal device, such as a personal computer. However, the program product of the present invention is not limited thereto. In this document, a readable storage medium can be any tangible medium containing or storing a program, which can be used by or in combination with an instruction execution system, an apparatus or a device.
[0107] The program product may use any combination of one or more readable media. The readable medium may be a readable signal medium or a readable storage medium. The readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or device, or any combination of the above. More specific examples (non-exhaustive list) of readable storage media include: an electrical connection with one or more wires, a portable disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above.
[0108] The computer readable storage medium may include a data signal propagated in a baseband or as part of a carrier wave, wherein a readable program code is carried. This propagated data signal may take a variety of forms, including but not limited to an electromagnetic signal, an optical signal, or any suitable combination of the above. The readable storage medium may also be any readable medium other than a readable storage medium, which may send, propagate, or transmit a program for use by an instruction execution system, an apparatus, or a device or used in combination with it. The program code contained on the readable storage medium may be transmitted with any appropriate medium, including but not limited to wireless, wired, optical cable, RF, etc., or any suitable combination of the above.
[0109] Program code for performing the operations of the present invention may be written in any combination of one or more programming languages, including object-oriented programming languages such as Java, C++, etc., and conventional procedural programming languages such as "C" or similar programming languages. The program code may be executed entirely on the tenant computing device, partially on the tenant device, as a separate software package, partially on the tenant computing device and partially on a remote computing device, or entirely on a remote computing device or server. In cases involving remote computing devices, the remote computing device may be connected to the tenant computing device via any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computing device (e.g., via the Internet using an Internet service provider).
[0110] In an exemplary embodiment of the present disclosure, an electronic device is further provided, which may include a processor and a memory for storing executable instructions of the processor, wherein the processor is configured to execute the steps of the point cloud data acceleration processing method in any of the above embodiments by executing the executable instructions.
[0111] It will be appreciated by those skilled in the art that various aspects of the present invention may be implemented as a system, method or program product. Therefore, various aspects of the present invention may be specifically implemented in the following forms, namely: a complete hardware implementation, a complete software implementation (including firmware, microcode, etc.), or a combination of hardware and software, which may be collectively referred to herein as a "circuit", "module" or "system".
[0112] Refer to the following Figure 7 The electronic device 600 according to this embodiment of the present invention is described. Figure 7 The electronic device 600 shown is only an example and should not bring any limitation to the functions and scope of use of the embodiments of the present invention.
[0113] like Figure 7 As shown, the electronic device 600 is in the form of a general computing device. The components of the electronic device 600 may include, but are not limited to: at least one processing unit 610, at least one storage unit 620, a bus 630 connecting different system components (including the storage unit 620 and the processing unit 610), a display unit 640, etc.
[0114] The storage unit stores program codes, which can be executed by the processing unit 610, so that the processing unit 610 performs the steps of various exemplary embodiments of the present invention described in the above-mentioned point cloud data acceleration processing method section of this specification. For example, the processing unit 610 can perform the following steps: Figure 1 Follow the steps shown in .
[0115] The storage unit 620 may include a readable medium in the form of a volatile storage unit, such as a random access memory unit (RAM) 6201 and / or a cache memory unit 6202 , and may further include a read-only memory unit (ROM) 6203 .
[0116] The storage unit 620 may also include a program / utility 6204 having a set (at least one) of program modules 6205, such program modules 6205 including but not limited to: an operating system, one or more application programs, other program modules and program data, each of which or some combination may include the implementation of a network environment.
[0117] Bus 630 may represent one or more of several types of bus structures, including a memory unit bus or memory unit controller, a peripheral bus, an accelerated graphics port, a processing unit, or a local bus using any of a variety of bus architectures.
[0118] The electronic device 600 may also communicate with one or more external devices 700 (e.g., keyboards, pointing devices, Bluetooth devices, etc.), may communicate with one or more devices that enable tenants to interact with the electronic device 600, and / or may communicate with any device that enables the electronic device 600 to communicate with one or more other computing devices (e.g., routers, modems, etc.). Such communication may be performed through an input / output (I / O) interface 650. In addition, the electronic device 600 may also communicate with one or more networks (e.g., a local area network (LAN), a wide area network (WAN), and / or a public network, such as the Internet) through a network adapter 660. The network adapter 660 may communicate with other modules of the electronic device 600 through a bus 630. It should be understood that, although not shown in the figure, other hardware and / or software modules may be used in conjunction with the electronic device 600, including but not limited to: microcode, device drivers, redundant processing units, external disk drive arrays, RAID systems, tape drives, and data backup storage systems.
[0119] Through the description of the above implementation, it is easy for those skilled in the art to understand that the example implementation described here can be implemented by software, or by software combined with necessary hardware. Therefore, the technical solution according to the implementation of the present disclosure can be embodied in the form of a software product, which can be stored in a non-volatile storage medium (which can be a CD-ROM, a USB flash drive, a mobile hard disk, etc.) or on a network, including a number of instructions to enable a computing device (which can be a personal computer, a server, or a network device, etc.) to execute the above-mentioned point cloud data acceleration processing method according to the implementation of the present disclosure.
[0120] Compared with the prior art, the advantages of the present invention are:
[0121] The cylindrical feature network processing is divided into a first working group and a second working group for processing, wherein the first feature data of the point cloud data to be processed is calculated based on the first working group, and the first working group is split into multiple one-dimensional first processing units to execute the first kernel function, the first feature data of the point cloud data to be processed is comprehensively processed based on the second working group, and a pseudo image of the point cloud data to be processed is obtained according to the second feature data, thereby optimizing the acceleration framework of the cylindrical feature network and reducing the inference time.
[0122] Those skilled in the art will readily appreciate other embodiments of the present disclosure after considering the specification and practicing the invention disclosed herein. This application is intended to cover any modification, use or adaptation of the present disclosure, which follows the general principles of the present disclosure and includes common knowledge or customary techniques in the art that are not disclosed in the present disclosure. The specification and examples are intended to be exemplary only, and the true scope and spirit of the present disclosure are indicated by the appended claims.
Claims
1. A method for accelerating processing of point cloud data, It is characterized in that The point cloud data is processed based on a cylindrical feature network, and the method includes: Based on the first working group, first feature data of the point cloud data to be processed is calculated, the feature data including radar features and image features, the point cloud data to be processed is discretized into a grid of a first coordinate plane to form a voxel columnar set, the voxel columnar set includes P columnar voxels, the first working group includes P first processing units, each of the first processing units executes a first kernel function to process a columnar voxel of the point cloud data to be processed, and obtains the radar feature of the columnar voxel; Based on the second working group, the first feature data of the point cloud data to be processed are comprehensively processed to obtain second feature data, wherein the comprehensive processing includes one or more of a splicing process, a linear transformation process, a batch normalization process, an activation process, and a maximum value process, the output of each first processing unit of the first working group is connected to a global memory, the first working group and the second working group are executed by a graphics processor, the global memory is a shared memory of the graphics processor and the central processing unit, the second working group has a three-dimensional work queue, the second working group includes P second computing units for parallel computing, each of the second computing units includes 2*M second processing units for parallel computing, each of the second processing units executes a second kernel function to comprehensively process a row of radar features / image features in a matrix formed by the first feature data, and M is an integer greater than 0; A pseudo image of the point cloud data to be processed is obtained according to the second feature data.
2. The point cloud data accelerated processing method according to claim 1, It is characterized in that In the second kernel function, the concatenation process and the linear transformation process are combined into one operator.
3. The point cloud data accelerated processing method according to claim 1, It is characterized in that The linear transformation process includes: multiplying the radar signature by a radar weight matrix; Multiplying the image features with the image weight matrix, Among them, the number of rows of the radar weight matrix is the number of the radar features, and the number of columns of the radar weight matrix is 2*M; the number of rows of the image weight matrix is the number of the image features, and the number of columns of the image weight matrix is 2*M. The radar weight matrix, the image weight matrix and the input data of the second computing unit are stored in the local memory of the second computing unit.
4. The point cloud data accelerated processing method according to claim 3, It is characterized in that The radar weight matrix and the image weight matrix contain the offset of the batch normalization.
5. The point cloud data accelerated processing method according to claim 3, It is characterized in that The first working group and the second working group are implemented based on OpenCL, and the local memory of the second computing unit in the second working group and the global memory are implemented using an asynchronous copy function of OpenCL.
6. A point cloud data acceleration processing device, It is characterized in that The point cloud data is processed based on a cylindrical feature network, including: A first feature data calculation module is configured to calculate first feature data of the point cloud data to be processed based on the first working group, the feature data including radar features and image features, the point cloud data to be processed is discretized into a grid of a first coordinate plane to form a voxel columnar set, the voxel columnar set includes P columnar voxels, the first working group includes P first processing units, each first processing unit executes a first kernel function to process a columnar voxel of the point cloud data to be processed to obtain the radar feature of the columnar voxel; A second feature data calculation module is configured to perform comprehensive processing on the first feature data of the point cloud data to be processed based on a second working group to obtain second feature data, wherein the comprehensive processing includes one or more of a splicing process, a linear transformation process, a batch normalization process, an activation process, and a maximum value process, wherein the output of each first processing unit of the first working group is connected to a global memory, the first working group and the second working group are executed by a graphics processor, the global memory is a shared memory of the graphics processor and the central processing unit, the second working group has a three-dimensional work queue, the second working group includes P second computing units for parallel computing, each of the second computing units includes 2*M second processing units for parallel computing, and each of the second processing units executes a second kernel function to perform comprehensive processing on a row of radar features / image features in a matrix formed by the first feature data, where M is an integer greater than 0; The pseudo image acquisition module is configured to obtain a pseudo image of the point cloud data to be processed according to the second feature data.
7. An electronic device, It is characterized in that The electronic device comprises: processor; A storage medium having a computer program stored thereon, wherein the computer program, when executed by the processor, performs the method according to any one of claims 1 to 5.
8. A storage medium, It is characterized in that The storage medium stores a computer program, and when the computer program is executed by a processor, the method according to any one of claims 1 to 5 is executed.
Citation Information
Patent Citations
Three-dimensional reconstruction algorithm parallelization method based on GPU cluster
CN111968218A
Target detection method and device, and computer storage medium
CN113205515A