Heterogeneous hardware acceleration system and recognition system for point cloud neural network
By designing a heterogeneous hardware acceleration system for point cloud neural networks, the problem of real-time processing of 3D point cloud data in resource-constrained environments is solved, efficient and low-power point cloud data processing is achieved, hardware development is simplified and design cycles are shortened.
Patent Information
- Application Number
- CN202510156234.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-12
- Publication Date
- 2025-05-30
AI Technical Summary
When processing 3D point cloud data, the prior art is difficult to meet the real-time processing requirements in resource-constrained environments, especially in deep learning algorithm scenarios with low power consumption and fast iteration.
A heterogeneous hardware acceleration system for point cloud neural networks is designed, including PS terminal, PL terminal and DDR. The system performs point cloud data processing through the accelerator, uses PE computing module, storage module, DMA and control module for efficient calculations, and optimizes the data transmission process through double buffering and ping-pong Buffer mechanisms.
It realizes efficient processing of 3D point cloud data in resource-constrained environments, reduces the cost of hardware resource usage and platform power consumption, simplifies the difficulty of hardware development, and greatly shortens the design cycle.
Smart Images

Figure CN120068963A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of hardware circuit acceleration design, and particularly relates to a heterogeneous hardware acceleration system and an identification system for point cloud neural networks. Background Art
[0002] In recent years, with the development of 3D sensor technology and its wide application in multiple fields, the processing and application of 3D data have become increasingly important. Compared with traditional 2D images, 3D data can provide more accurate distance, geometric shape, and surface information, showing significant advantages in fields such as autonomous driving, robot navigation, and virtual / augmented reality. However, these application scenarios pose higher requirements for 3D data processing algorithms, especially in cases where real-time response and high-precision spatial perception are needed. Existing two-dimensional image deep learning algorithms are difficult to directly apply to the processing of such 3D data due to their inability to effectively handle the characteristics of disordered and sparse point cloud data, which limits their performance and technical potential in practical applications.
[0003] To address these issues, researchers have explored various hardware acceleration solutions to improve the speed and efficiency of 3D point cloud data processing. Among them, GPU (Graphics Processing Unit) and ASIC (Application-Specific Integrated Circuit) are regarded as two of the closest technical solutions. With its powerful parallel computing ability, GPU can significantly enhance the execution speed of algorithms such as object detection and is suitable for scenarios requiring high-speed computing; while ASIC, through optimized design for specific algorithms, can greatly improve the computing efficiency while ensuring accuracy, and is particularly suitable for devices that need to operate with low power consumption. Both of these methods meet the needs of 3D point cloud data analysis to a certain extent and demonstrate adaptability and flexibility in different application scenarios.
[0004] Although the above technologies provide effective solutions, they each have obvious limitations. Although GPU has excellent parallel processing ability, due to its high energy consumption, it is not suitable for edge mobile devices with low power consumption requirements. On the other hand, although ASIC performs well in terms of speed and energy efficiency, its long design cycle and high manufacturing cost make it difficult to keep up with the rapid iteration of deep learning algorithms. In addition, for resource-constrained environments such as in-vehicle devices or intelligent Internet of Things terminals, the high power consumption and large hardware overhead of existing general machine learning hardware platforms become a major obstacle to the realization of efficient real-time applications. The existence of these problems urgently requires a new technical solution to overcome the deficiencies of the existing technologies to better serve the real-time processing needs of 3D point cloud data in resource-constrained environments. Summary of the Invention
[0005] To solve the above problems existing in the prior art, the present invention provides a heterogeneous hardware acceleration system and an identification system for a point cloud neural network.
[0006] The technical problems to be solved by the present invention are realized through the following technical solutions:
[0007] In a first aspect, the present invention provides a heterogeneous hardware acceleration system for a point cloud neural network, including: a PS side, a PL side, and a DDR;
[0008] The PS side, as the CPU, runs the data path and control path of the accelerator through the ARM core;
[0009] An accelerator is provided in the PL side, and the accelerator is used to process the point cloud processing data to obtain the point cloud calculation result;
[0010] The data path includes: transmitting the point cloud processing data to the DDR, and transferring part of the data in the point cloud processing data from the DDR to the memory storage of the accelerator; the point cloud processing data includes: point cloud feature data and the training result corresponding to the point cloud feature data;
[0011] The control path is used to control the accelerator to complete the corresponding point cloud operation for the point cloud processing data;
[0012] The accelerator includes: a PE calculation module, a storage module, a DMA, and a control module. The PE calculation module includes multiple PE sub-units, and each PE sub-unit is composed of multiple multipliers and adders; a double buffer and a ping-pong Buffer mechanism are provided in the storage module.
[0013] Optionally, in the data path, the point cloud processing data is transferred from the DDR to the storage module through the DMA;
[0014] In the control path, the accelerator is controlled according to the AXI-lite protocol to complete the corresponding point cloud operation for the point cloud processing data under the action of the accelerator.
[0015] Optionally, the storage module includes: an input cache module and an output cache module;
[0016] The input cache module is provided with a feature cache module and a weight cache module;
[0017] The output ends of the DMA are respectively connected in parallel to the input ends of the feature cache module and the weight cache module; the output ends of the feature cache module and the weight cache module are both connected to the input end of the PE calculation module; the output end of the PE calculation module is connected to the input end of the output cache module, and the output end of the output cache module is connected to the input end of the DMA;
[0018] The control module is respectively connected to the input ends of the weight cache module, the PE calculation module, and the output cache module.
[0019] Optionally, the feature cache module is provided with a first feature cache module and a second feature cache module;
[0020] The first feature cache module and the second feature cache module store data using the ping-pong Buffer mechanism.
[0021] Optionally, the output cache module is provided with a first output cache and a second output cache;
[0022] The first output cache is used to store the intermediate data during matrix multiplication by the PE calculation module;
[0023] The second output cache is used to store the final point cloud calculation result and transmits the point cloud calculation result to the CPU through DMA.
[0024] Optionally, multiple PE sub-units form a calculation array within the PE calculation module, and the PE calculation module further includes: an addition module and a comparator unit;
[0025] The calculation array, the addition module, and the comparator unit are connected in series in sequence; the output end of the first output cache is also connected to the input end of the addition module;
[0026] The addition module is used to perform an addition operation on the intermediate data and the array data corresponding to the calculation array according to the preset calculation rules in the control module;
[0027] The comparator unit is used to perform a max pooling operation and a Relu activation operation on the output result of the addition module.
[0028] Optionally, the bit width of the first output cache is greater than the bit width of the second output cache.
[0029] Optionally, the control module is used to run a control file, and the control file includes: network structure parameters, batch, and preset calculation rules;
[0030] The network structure parameters include: the number of layers of the neural network, the number of neurons, the size of the convolutional kernel, and the stride.
[0031] Optionally, the PE calculation module is provided with N PE sub-units; each PE sub-unit is provided with 8 multipliers and 7 adders;
[0032] Among them, N takes the value of 2 (5+n) , where n is an integer greater than or equal to 0 and less than or equal to 2.
[0033] In a second aspect, the present invention provides an identification system, including: a lidar, a data preprocessing module, a display, and the heterogeneous hardware acceleration system in the first aspect above;
[0034] The lidar is used to collect the original point cloud data;
[0035] The data preprocessing module is used to preprocess the original point cloud data to obtain point cloud feature data;
[0036] The heterogeneous hardware acceleration system is used to perform hardware acceleration operations on the point cloud feature data and the training results corresponding to the point cloud feature data, obtain the point cloud calculation results, and send the point cloud calculation results to the display;
[0037] The display is used to display the point cloud calculation results.
[0038] The present invention provides a heterogeneous hardware acceleration system and an identification system for a point cloud neural network. Among them, the heterogeneous hardware acceleration system for a point cloud neural network includes: a PS side, a PL side, and a DDR; the PS side, as the CPU, runs the data path and control path of the accelerator through the ARM core; an accelerator is provided in the PL side, and the accelerator is used to process the point cloud processing data to obtain the point cloud calculation results; the data path includes: transmitting the point cloud processing data to the DDR, and moving a part of the data in the point cloud processing data from the DDR to the memory storage of the accelerator; the point cloud processing data includes: point cloud feature data and the training results corresponding to the point cloud feature data; the control path is used to control the accelerator to complete the corresponding point cloud operations for the point cloud processing data; the accelerator includes: a PE calculation module, a storage module, a DMA, and a control module, the PE calculation module includes a plurality of PE sub-units, and each PE sub-unit is composed of a plurality of multipliers and adders; a double buffer and ping-pong Buffer mechanism is provided in the storage module. In the present invention, the double buffer and ping-pong Buffer mechanism optimizes the data transmission process, reduces the demand for high-bandwidth and high-performance storage, thereby reducing the usage cost of hardware resources and the platform power consumption. And the DDR, as a large-capacity storage medium, further shares the temporary storage pressure of the point cloud data, enabling the accelerator memory to focus on high-performance computing tasks and realizing the efficient utilization of resource allocation. In addition, since the PL side of the present invention can be based on the existing FPGA platform to realize the rapid deployment and iteration of the algorithm logic, the design cycle is greatly shortened, and the PS side, based on the general-purpose processor of the ARM core, is responsible for the data path and control logic, simplifying the design of complex control circuits and reducing the hardware development difficulty and related design costs.
[0039] The following will further describe the present invention in detail with reference to the accompanying drawings and embodiments. Description of the Drawings
[0040] Figure 1 It is a schematic structural diagram of a heterogeneous hardware acceleration system for a point cloud neural network provided by an embodiment of the present invention;
[0041] Figure 2 The structural schematic diagram of the accelerator is exemplarily shown;
[0042] Figure 3 The structural schematic diagram of the parallel computing unit of the computing array is exemplarily shown;
[0043] Figure 4 The structural schematic diagram of an identification system provided by an embodiment of the present invention. Detailed implementation manners
[0044] The present invention will be further described in detail below in conjunction with specific embodiments, but the implementation manners of the present invention are not limited thereto.
[0045] In order to reduce the usage cost of hardware resources and the difficulty of hardware development, and at the same time achieve the efficient utilization of resource allocation, an embodiment of the present invention provides a heterogeneous hardware acceleration system for point cloud neural networks. Figure 1 The structural schematic diagram of a heterogeneous hardware acceleration system for point cloud neural networks provided by an embodiment of the present invention. As Figure 1 shown, it includes: a PS side, a PL side, and a DDR;
[0046] The PS side, as the CPU, runs the data path and control path of the accelerator through the ARM core;
[0047] An accelerator is provided in the PL side, and the accelerator is used to process the point cloud processing data to obtain the point cloud calculation result;
[0048] The data path includes: transmitting the point cloud processing data to the DDR, and moving a part of the data in the point cloud processing data from the DDR to the memory storage of the accelerator; the point cloud processing data includes: point cloud feature data and the training result corresponding to the point cloud feature data;
[0049] The control path is used to control the accelerator to complete the corresponding point cloud operations for the point cloud processing data;
[0050] The accelerator includes: a PE calculation module, a storage module, a DMA, and a control module. The PE calculation module includes a plurality of PE sub-units, and each PE sub-unit is composed of a plurality of multipliers and adders; a double buffer and a ping-pong Buffer mechanism are provided in the storage module.
[0051] The training result in this embodiment is specifically the weights and biases obtained by the point cloud feature data during the neural network training process.
[0052] It should be noted that the point cloud processing data in the embodiment of the present invention may specifically be 3D point cloud processing data.
[0053] In the present invention, the PS (Processing System) side is mainly responsible for running control logic and data management tasks, usually implemented based on an ARM core. The PL (Programmable Logic) side is used to implement functions such as hardware accelerators and specifically processes high-performance computing tasks, such as the acceleration processing of point cloud data.
[0054] DDR (Double Data Rate SDRAM) is a high-speed storage technology that can transfer data on both the rising and falling edges of each clock cycle, thus achieving a double data transfer rate. Compared with dedicated caches or on-chip memories, DDR has a lower cost per unit capacity and is suitable as a large-capacity storage solution at the system level. This helps control the overall hardware cost while providing sufficient storage space. In the present invention, by temporarily storing point cloud processing in DDR, more flexible data scheduling and management can be achieved. For example, in the point cloud processing of the present invention, data can be transferred from DDR to the accelerator memory in batches according to requirements, avoiding resource waste caused by loading all data at once. In addition, the on-chip memory of the accelerator usually has limited capacity but is faster. By using DDR as an external memory, the pressure on on-chip storage can be effectively relieved, enabling the accelerator to focus on high-performance computing tasks and thus improving the overall efficiency.
[0055] Optionally, in the data path, the point cloud processing data is transferred from DDR to the storage module through DMA;
[0056] In the control path, the accelerator is controlled according to the AXI-lite protocol to complete corresponding point cloud operations on the point cloud processing data under the action of the accelerator.
[0057] DMA (Direct Memory Access) allows data to be directly transferred between DDR and the storage module without the CPU on the PS side participating in the specific transfer operation.
[0058] Correspondingly, Figure 2 Exemplarily shows a schematic structural diagram of the accelerator. As Figure 2 shown, the accelerator includes: a PE computing module, a storage module, DMA, and a control module. The storage module includes: an input buffer module and an output buffer module;
[0059] The input buffer module is provided with a feature buffer module and a weight buffer module;
[0060] The output ends of the DMA are respectively connected in parallel to the input ends of the feature cache module and the weight cache module; the output ends of the feature cache module and the weight cache module are both connected to the input end of the PE calculation module; the output end of the PE calculation module is connected to the input end of the output cache module, and the output end of the output cache module is connected to the input end of the DMA;
[0061] The control module is respectively connected to the input ends of the weight cache module, the PE calculation module, and the output cache module.
[0062] Optionally, the feature cache module is provided with a first feature cache module and a second feature cache module;
[0063] The first feature cache module and the second feature cache module store data using the ping-pong Buffer mechanism.
[0064] Optionally, the output cache module is provided with a first output cache and a second output cache;
[0065] The first output cache is used to store the intermediate data during the matrix multiplication of the PE calculation module;
[0066] The second output cache is used to store the final point cloud calculation result, and the point cloud calculation result is transmitted to the CPU through the DMA.
[0067] It should be noted that there are two important reasons for designing this two-level output buffer structure: one is to avoid the possible problem of reduced accuracy during the matrix partitioning process of the computing array, and to ensure the accuracy and reliability of the point cloud calculation result through reasonable buffer design and data storage methods; the other is to reduce the frequent reading pressure on the second output cache. By storing some intermediate data in the first output cache and realizing the reuse of some data, the number of data transmissions and readings is reduced, and the overall performance of the heterogeneous hardware acceleration system is improved.
[0068] Optionally, multiple PE sub-units form a computing array within the PE calculation module, and the PE calculation module further includes: an addition module and a comparator unit;
[0069] The computing array, the addition module, and the comparator unit are connected in series in sequence; the output end of the first output cache is also connected to the input end of the addition module;
[0070] The addition module is used to perform an addition operation on the intermediate data and the array data corresponding to the computing array according to the preset calculation rules in the control module;
[0071] The comparator unit is used to perform a max pooling operation and a Relu activation operation on the output result of the addition module.
[0072] Figure 3The structural schematic diagram of the parallel computing units of the computing array is exemplarily shown. In this embodiment, 32 PE sub-units are taken as an example for illustration. As Figure 3 shown, 32 PE sub-units are connected in series and perform corresponding arithmetic processing on the point cloud feature data according to the corresponding weights.
[0073] It should be noted that the max-pooling operation in the embodiments of the present invention can effectively reduce the data volume while retaining the key information of the output result (input feature) of the summation module. For example, when processing a point cloud object with a large number of detailed features, the most representative feature information can be selected through the max-pooling operation to avoid the influence of data redundancy on subsequent calculations. The Relu activation operation can enhance the non-linear expression ability of the data, enabling the neural network to better process complex point cloud data relationships and improving the expression ability and generalization ability of the preset model in the control module.
[0074] Optionally, the bit width of the first output buffer is greater than the bit width of the second output buffer.
[0075] Optionally, the control module is used to run a control file, and the control file includes: network structure parameters, batch, and preset calculation rules;
[0076] The network structure parameters include: the number of layers of the neural network, the number of neurons, the size and stride of the convolutional kernel.
[0077] Optionally, the PE computing module is provided with N PE sub-units; each PE sub-unit is provided with 8 multipliers and 7 adders;
[0078] wherein, N takes a value of 2 (5+n) , and n is an integer greater than or equal to 0 and less than or equal to 2.
[0079] In addition, in the embodiments of the present invention, N PE sub-units are arranged in a matrix form.
[0080] It should be noted that the above structural design of the PE computing module enables it to efficiently process the mathematical operations in the MLP (Multi-Layer Perceptron). For example, when processing the feature extraction and transformation of point cloud data, it can quickly perform multiplication and addition operations on a large number of data points, realizing the efficient extraction and integration of data features.
[0081] The embodiment of the present invention provides a heterogeneous hardware acceleration system for point cloud neural networks. By adopting a double-buffer and ping-pong buffer mechanism, the data transmission process is optimized, the demand for high-bandwidth and high-performance storage is reduced, thereby reducing the usage cost of hardware resources and the platform power consumption. Using DDR as a large-capacity storage medium further shares the temporary storage pressure of point cloud data, enabling the accelerator memory to focus on high-performance computing tasks and achieving efficient utilization of resource allocation. In addition, since the PL side of the present invention can be based on an existing FPGA platform to achieve rapid deployment and iteration of algorithm logic, the design cycle is significantly shortened, and the PS side, a general-purpose processor based on an ARM core, is responsible for the data path and control logic, simplifying the design of complex control circuits, reducing the hardware development difficulty and related design costs.
[0082] Corresponding to the heterogeneous hardware acceleration system in the above embodiment, the embodiment of the present invention also provides an identification system. Figure 4 It is a schematic structural diagram of an identification system provided by an embodiment of the present invention. As Figure 4 shown, it includes: a lidar, a data preprocessing module, a display, and the above heterogeneous hardware acceleration system;
[0083] The lidar is used to collect raw point cloud data;
[0084] The data preprocessing module is used to preprocess the raw point cloud data to obtain point cloud feature data;
[0085] The heterogeneous hardware acceleration system is used to perform hardware acceleration operations on the point cloud feature data and the training results corresponding to the point cloud feature data to obtain point cloud calculation results, and send the point cloud calculation results to the display;
[0086] The display is used to display the point cloud calculation results.
[0087] The preprocessing in this embodiment mainly filters and denoises the raw point cloud data, removes redundant point information, etc., and finally obtains point cloud feature data.
[0088] After obtaining the point cloud processing data, taking the application of point cloud data in autonomous driving as an example, the processing is described. The detailed processing process can refer to the following process: The CPU initializes and configures the control file according to preset network structure parameters, such as the resolution of the point cloud feature data (which can be dynamically adjusted according to the vehicle driving speed and road conditions, for example, a lower resolution is adopted at high speed to reduce the calculation amount, and a higher resolution is adopted in complex urban road conditions to ensure accurate perception), the categories of target detection (including multiple categories such as pedestrians, vehicles, traffic signs and signals, road boundaries, etc.).
[0089] In the initialization stage, the pre-trained point cloud neural network weight data is quickly transferred to the on-chip parameter cache module through DMA. This weight data is trained with a large amount of actual driving scenario data and can accurately identify various target objects and road features. Then, when the computing is enabled, the point cloud data is input from the memory to the hardware acceleration circuit through DMA. The PE computing module in the hardware acceleration circuit first performs calculations on the MLP, and the multipliers and adder trees in the PE sub-units perform fast operations. Each PE sub-unit can process a certain number of data points within one clock cycle, and multiple PE sub-units work in parallel, greatly improving the computing efficiency. The PE sub-units initially extract and transform the features of the input point cloud data according to the calculation rules of the MLP, and extract the initial point cloud feature information. Then, it is further processed through the max pooling and fully connected modules. In this process, the key road features are retained through the max pooling operation of the comparator unit. For example, when detecting multiple small potholes on the road, the most representative pothole feature information is selected through max pooling to reduce the data volume for subsequent fast processing. In addition, the Relu activation processing can enhance the non-linear expression of the data, enabling the model to better process complex road scenario data. At the end of the entire process, the classification of the point cloud data (such as objects like trees, people, cars, etc.) is finally achieved, and the point cloud computing result is sent to the display for display.
[0090] In the description of this specification, the description with reference to terms such as "one embodiment", "some embodiments", "example", "specific example", or "some examples" means that the specific features or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic expressions of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features or characteristics described can be combined in a suitable manner in any one or more embodiments or examples. In addition, those skilled in the art can combine and combine the different embodiments or examples described in this specification.
[0091] Although the present invention has been described in conjunction with various embodiments herein, however, in the process of implementing the claimed invention, those skilled in the art can understand and achieve other variations of the above-described disclosed embodiments by viewing the drawings and the disclosure. In the description of the present invention, the term "including" does not exclude other components or steps, the term "a" or "one" does not exclude a plurality of cases, and the meaning of "a plurality" is two or more, unless otherwise specifically defined. In addition, certain measures are described in different embodiments, but this does not mean that these measures cannot be combined to produce good results.
[0092] The above content is a further detailed description of the present invention in combination with specific preferred embodiments. It cannot be determined that the specific implementation of the present invention is only limited to these descriptions. For those of ordinary skill in the technical field to which the present invention pertains, without departing from the concept of the present invention, several simple deductions or substitutions can still be made, which should all be regarded as falling within the protection scope of the present invention.
Claims
1. A heterogeneous hardware acceleration system for point cloud neural networks, characterized in that: include: PS side, PL side and DDR; The PS side acts as a CPU to run the data path and control path of the accelerator through the ARM core; The PL end is provided with the accelerator, and the accelerator is used to process the point cloud processing data to obtain the point cloud computing result; The data path includes: transmitting the point cloud processing data to the DDR, and moving part of the point cloud processing data from the DDR to the memory storage of the accelerator; the point cloud processing data includes: point cloud feature data and training results corresponding to the point cloud feature data; The control path is used to control the accelerator to perform corresponding point cloud operations on the point cloud processing data; The accelerator comprises: a PE calculation module, a storage module, a DMA and a control module. The PE calculation module comprises a plurality of PE subunits, each of which is composed of a plurality of multipliers and adders. The storage module is provided with a double buffer and a ping-pong buffer mechanism.
2. The heterogeneous hardware acceleration system for point cloud neural networks according to claim 1, characterized in that: In the data path, the point cloud processing data is transferred from the DDR to the storage module via DMA; In the control path, the accelerator is controlled according to the AXI-lite protocol, so that the corresponding point cloud operation is completed for the point cloud processing data under the action of the accelerator.
3. The heterogeneous hardware acceleration system for point cloud neural networks according to claim 1, characterized in that: The storage module includes: an input cache module and an output cache module; The input cache module is provided with a feature cache module and a weight cache module; The output end of the DMA is connected in parallel with the input end of the feature cache module and the weight cache module respectively; the output end of the feature cache module and the weight cache module are both connected to the input end of the PE calculation module; the output end of the PE calculation module is connected to the input end of the output cache module, and the output end of the output cache module is connected to the input end of the DMA; The control module is connected to the input ends of the weight cache module, the PE calculation module and the output cache module respectively.
4. The heterogeneous hardware acceleration system for point cloud neural networks according to claim 3, characterized in that: The feature cache module is provided with a first feature cache module and a second feature cache module; The first feature cache module and the second feature cache module use a ping-pong buffer mechanism to store data.
5. The heterogeneous hardware acceleration system for point cloud neural networks according to claim 3, characterized in that: The output cache module is provided with a first output cache and a second output cache; The first output buffer is used to store intermediate data of the PE calculation module during matrix multiplication; The second output buffer is used to store the final point cloud computing result, and transmit the point cloud computing result to the CPU through the DMA.
6. The heterogeneous hardware acceleration system for point cloud neural networks according to claim 5, characterized in that: The plurality of PE sub-units form a calculation array within the PE calculation module, and the PE calculation module further comprises: a summing module and a comparator unit; The calculation array, the summing module and the comparator unit are connected in series in sequence; the output end of the first output buffer is also connected to the input end of the summing module; The summing module is used to perform a summing operation on the intermediate data and the array data corresponding to the calculation array according to the calculation rules preset in the control module; The comparator unit is used to perform a maximum pooling operation and a ReLU activation operation on the output result of the summing module.
7. The heterogeneous hardware acceleration system for point cloud neural networks according to claim 5, characterized in that: The bit width of the first output buffer is greater than the bit width of the second output buffer.
8. The heterogeneous hardware acceleration system for point cloud neural networks according to claim 1, characterized in that: The control module is used to run a control file, which includes: network structure parameters, batch and preset calculation rules; The network structure parameters include: the number of layers of the neural network, the number of neurons, the size of the convolution kernel and the step size.
9. The heterogeneous hardware acceleration system for point cloud neural networks according to claim 1, characterized in that: The PE calculation module is provided with N PE subunits; each of the PE subunits is provided with 8 multipliers and 7 adders; Among them, N is 2 (5+n) , n is an integer greater than or equal to 0 and less than or equal to 2.
10. A recognition system, characterized in that: include: A laser radar, a data preprocessing module, a display, and a heterogeneous hardware acceleration system as described in any one of claims 1 to 9; The laser radar is used to collect original point cloud data; The data preprocessing module is used to preprocess the original point cloud data to obtain point cloud feature data; The heterogeneous hardware acceleration system is used to perform hardware acceleration operations on the point cloud feature data and the training results corresponding to the point cloud feature data to obtain a point cloud computing result, and send the point cloud computing result to the display; The display is used to display the point cloud computing result.