Real-time point cloud target detection system based on lightweight model and collaborative acceleration device
Through the point cloud-projection dual-stream architecture and collaborative acceleration device, the resource limitation problem of real-time point cloud target detection on edge devices is solved, achieving efficient and low-power real-time processing, while maintaining high detection accuracy, and is suitable for scenarios such as autonomous driving, intelligent robots and augmented reality.
Patent Information
- Application Number
- CN202510496648.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-21
- Publication Date
- 2025-07-25
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
The prior art is difficult to realize real-time point cloud object detection on resource-constrained edge devices, and the existing accelerators have failed to fully utilize the synergistic advantages of point cloud data characteristics and multimodal representation, resulting in reduced detection accuracy and high power consumption.
Using the point cloud-projection dual-stream architecture, combined with the graph convolution edge attenuation module, multi-resolution projection strategy and adaptive feature fusion unit, memory access is optimized through adaptive weight allocation and heterogeneous memory hierarchy scheduling, and power consumption configuration is dynamically adjusted to achieve lightweight and efficient processing.
Real-time processing performance above 25FPS is achieved on edge devices, with a reduction of about 65% in model parameters, a reduction of about 45% in memory footprint, and an average power consumption is reduced by about 40%. It maintains high detection accuracy in complex scenarios.
Smart Images

Figure CN120375081A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical fields of computer vision and artificial intelligence, and particularly to a real-time point cloud object detection system and a collaborative acceleration device based on a lightweight model. Background Art
[0002] As a key component of 3D scene understanding, point cloud object detection technology has shown extensive application value in fields such as autonomous driving, robot navigation, and augmented reality. Traditional point cloud object detection is mainly based on the design of handcrafted features, such as segmentation and clustering methods based on geometric features, but its robustness and accuracy in complex scenes are relatively low. With the development of deep learning, point cloud object detection methods based on deep neural networks have gradually become the mainstream. Models such as PointNet, PointNet++, and VoxelNet have significantly improved the detection accuracy and robustness. However, these methods usually require a large amount of computing resources and are difficult to achieve real-time processing on resource-constrained edge devices.
[0003] Existing lightweight point cloud object detection technologies mainly adopt methods such as model compression, knowledge distillation, and low-precision quantization to reduce computational overhead. For example, redundant connections are removed through pruning techniques, or knowledge distillation is used to transfer the knowledge of a large teacher model to a small student model. However, these methods often lead to a significant decrease in detection accuracy during the lightweight process, especially for long-distance, small targets, or partially occluded scenes. In addition, single-modal point cloud processing methods are difficult to fully utilize the complementary advantages of different representation forms, which limits the overall performance of the system.
[0004] In terms of hardware acceleration, existing point cloud processing accelerators are mainly designed for general deep learning operations, such as NVIDIA's GPUs or Google's TPUs, and lack targeted optimization for the characteristics of point cloud data. These general architectures are less efficient in processing the irregular and sparse structures unique to point clouds and often consume high power, making them unsuitable for edge computing scenarios. At the same time, traditional memory architectures do not consider the spatial locality and access pattern characteristics of point cloud data, resulting in frequent memory access conflicts and cache misses, further reducing the system efficiency.
[0005] In summary, the existing technologies have not effectively solved the contradiction between real-time performance and accuracy, lack dedicated hardware acceleration designs for the characteristics of point cloud data, and have not fully explored the collaborative advantages of multi-modal representations and heterogeneous computing. The present invention aims to solve the above key technical problems through an innovative dual-stream heterogeneous collaborative architecture and a dedicated acceleration device, and achieve efficient real-time processing of point cloud object detection on edge devices. Summary of the Invention
[0006] Based on the above objectives, the present invention provides a real-time point cloud object detection system and a collaborative acceleration device based on a lightweight model.
[0007] A real-time point cloud object detection system based on a lightweight model, comprising: A point cloud-projection dual-stream architecture model, which includes a point cloud stream processing part and a projection stream processing part; The point cloud stream processing part includes a point cloud feature extraction unit and a graph convolutional edge attenuation module, which are used for feature extraction and lightweight processing of point cloud data; The projection stream processing part includes a multi-view projection unit and a lightweight CNN processing unit, which are used for projecting the point cloud into multiple two-dimensional views and extracting complementary features; An adaptive feature fusion unit, which is used for automatically adjusting the weights of the point cloud stream and projection stream features according to the scene complexity and generating fused features.
[0008] Furthermore, the graph convolutional edge attenuation module realizes adaptive weight assignment of features by calculating the edge importance score, and the edge importance score is calculated by the following formula: EIS(p)=α·S(p)+β·D(p)+γ·V(p) Wherein, S(p) represents the shape feature score of point p, D(p) represents the density change score of point p, V(p) represents the visibility score of point p, and α, β, γ are adaptive weight coefficients.
[0009] Furthermore, the multi-view projection unit adopts a multi-resolution projection strategy, including: Adopting high-resolution projection for the close-range area (0 - 20m); Adopting medium-resolution projection for the medium-range area (20 - 50m); Adopting low-resolution projection for the long-range area (>50m).
[0010] Furthermore, the adaptive feature fusion unit includes: A dynamic weight calculation module, which is used for calculating the fusion weights of the point cloud stream and projection stream according to the scene complexity; A feature alignment module, which is used for mapping features of different dimensions to a unified representation space; A complementary feature enhancement module, which is used for enhancing complementary information through feature interaction.
[0011] A collaborative acceleration device for point cloud object detection, comprising: An intelligent task dispatching controller, which is used for dynamically allocating computing resources according to the point cloud characteristics; A heterogeneous memory hierarchy scheduling unit, which is used for optimizing memory access according to the point cloud data access pattern; A heterogeneous parallel acceleration unit, including a point cloud stream processing unit, a projection stream processing unit, and a feature fusion unit; A dynamic power management unit for automatically adjusting the power consumption configuration according to task requirements and system status.
[0012] Furthermore, the heterogeneous memory hierarchy scheduling unit implements hierarchical cache management by calculating the access priority score of data blocks, and the access priority score is calculated by the following formula: APS(b)=w1·I(b)+w2·F(b)+w3·T(b) Where, I(b) represents the information entropy of data block b, F(b) represents the access frequency of data block b, T(b) represents the temporal correlation of data block b, and w1, w2, w3 are weight coefficients.
[0013] Furthermore, the point cloud stream processing unit is implemented by a dedicated point cloud graph convolution accelerator, including: A k-nearest neighbor search accelerator using spatial hashing and parallel search techniques; An edge feature calculation unit for parallelly calculating edge features; A graph convolution core for performing graph convolution operations; A residual connection unit for implementing feature residual connection.
[0014] Furthermore, the projection stream processing unit is implemented by a reconfigurable logic array, including: A projection transformation module for implementing the transformation from point cloud to multi-view; A depthwise separable convolution array for accelerating convolution operations; A channel attention calculation unit for implementing the SE attention mechanism; A feature aggregation unit for implementing multi-scale feature fusion.
[0015] Furthermore, the dynamic power management unit includes: A dynamic voltage and frequency adjustment module for adjusting the working frequency and voltage of the processing unit according to the processing load; A selective unit sleep module for putting some processing units into a low-power state under light load conditions; A model hierarchical quantization module for applying quantization strategies with different bit widths to different layers of the model.
[0016] Preferably, the system and device are combined and applied to at least one application scenario of an autonomous driving perception system, intelligent robot navigation, drone scene understanding, or augmented reality spatial mapping.
[0017] Advantages of the present invention: Achieve real-time processing performance above 25FPS on edge devices; The number of model parameters is reduced by about 65%, and the memory occupancy is reduced by about 45%; The average power consumption is reduced by about 40%; Maintain a high detection accuracy in complex scenarios. BRIEF DESCRIPTION OF THE DRAWINGS
[0018] In order to more clearly illustrate the technical solutions in the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are only those of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0019] Figure 1 Schematic diagram of the point cloud-projection dual-stream architecture of the present invention; Figure 2 Schematic diagram of the hardware co-acceleration device architecture of the present invention; Figure 3 Schematic diagram of the structure of the graph convolutional edge decay module (GCBD) of the present invention; Figure 4 Schematic diagram of the multi-resolution projection strategy of the present invention; Figure 5 Schematic diagram of the heterogeneous memory scheduling process of the present invention; Figure 6 Schematic diagram of the performance comparison test results of the embodiments of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0020] The present invention will be described in detail below in conjunction with the drawings and specific embodiments. At the same time, it should be noted here that in order to make the embodiments more detailed, the following embodiments are the best and preferred embodiments. For some well-known technologies, those skilled in the art can also adopt other alternative methods for implementation; and the drawings are only for more specific description of the embodiments, and are not intended to specifically limit the present invention.
[0021] It should be noted that in the specification, it is mentioned that "an embodiment", "embodiments", "exemplary embodiments", "some embodiments", etc. indicate that the described embodiments may include specific features, structures or characteristics, but not necessarily every embodiment includes the specific features, structures or characteristics. In addition, when combining embodiments to describe specific features, structures or characteristics, implementing such features, structures or characteristics in combination with other embodiments (whether explicitly described or not) should be within the knowledge of those skilled in the relevant art.
[0022] In general, terms can be understood at least in part from their use in context. For example, depending at least in part on the context, the term "one or more" as used herein can be used to describe any feature, structure, or property in a singular sense, or can be used to describe a combination of features, structures, or properties in a plural sense. Additionally, the term "based on" can be understood to not necessarily be intended to convey a set of exclusive factors, but rather can alternatively, depending at least in part on the context, allow for the existence of other factors that are not necessarily explicitly described.
[0023] System overall architecture As Figure 1 shown, the real-time point cloud object detection system of the present invention mainly includes two main parts: a point cloud processing stream and a projection processing stream. The system input is the original point cloud data, and after dual-stream processing, the final object detection result is output through adaptive feature fusion.
[0024] Point cloud processing stream The point cloud processing stream mainly includes the following processing units: Point cloud feature extraction unit: An improved PointNet structure is used for preliminary feature extraction; Graph convolutional edge decay module (GCBD): Performs lightweight feature processing and optimization.
[0025] Projection processing stream The projection processing stream mainly includes the following processing units: Multi-view projection unit: Projects the three-dimensional point cloud into multiple two-dimensional views; Lightweight CNN processing unit: Processes the projected views and extracts complementary features.
[0026] Adaptive feature fusion The adaptive feature fusion unit automatically adjusts the weights of the point cloud stream and the projection stream according to the scene complexity, realizing complementary enhancement of features between the two streams.
[0027] Specific implementation of the graph convolutional edge decay module (GCBD) As Figure 3 shown, the specific implementation of the GCBD module includes the following steps: Construct a k-nearest neighbor point cloud graph structure, where k is set to 20 in this embodiment; Calculate the edge importance score (EIS) for each point; Apply an adaptive decay function to adjust the feature weights; Perform lightweight graph convolutional operations to extract features; Retain key feature information through residual connections.
[0028] The calculation method of the edge importance score (EIS) is as follows: EIS(p) = α·S(p) + β·D(p) + γ·V(p) Wherein: S(p) represents the shape feature score of point p, and the calculation formula is: S(p) = 1 - exp(-λ1·||∇N(p)||²) Where ∇N(p) is the normal vector gradient at point p, and λ1 is a tuning parameter, which is set to 0.5 in this embodiment; D(p) represents the density change score of point p, and the calculation formula is: D(p) = |ρ(p) - ρ̄(N(p))| / max(ρ̄(N(p)), ε) Where ρ(p) is the local density at point p, ρ̄(N(p)) is the neighborhood average density of p, and ε is a small constant to prevent division by zero, which is set to 1e-6 in this embodiment; V(p) represents the visibility score of point p, and the calculation formula is: V(p) = (1 + cos(θ(p))) / 2 Where θ(p) is the angle between the normal vector of point p and the line-of-sight direction; α, β, γ are adaptive weight coefficients, and the initial values are 0.4, 0.3, 0.3 respectively, and are automatically adjusted through network training.
[0029] The adaptive attenuation function based on EIS is defined as follows: W(p) = σ(EIS(p))·W0(p) Wherein: W(p) is the attenuated feature weight; W0(p) is the original feature weight; σ(·) is a variant of the sigmoid activation function, and is defined as: σ(x) = 1 / (1 + exp(-τ·(x - θ))) Where τ is the temperature parameter, which is set to 2.0 in this embodiment; θ is the threshold parameter, which is set to 0.5 in this embodiment.
[0030] Specific implementation of the multi-resolution projection strategy As Figure 4 shown, the specific implementation of the multi-resolution projection strategy includes: The point cloud is divided into three regions according to the distance: Near-distance region (0 - 20m): Project with a resolution of 128×128; Mid-distance region (20 - 50m): Project with a resolution of 64×64; Far-distance region (>50m): Project with a resolution of 32×32.
[0031] The projection methods include three types: Bird's Eye View (BEV), Side View, and Front View.
[0032] The processing of the projected image uses a lightweight CNN network, and the specific structure is as follows: Input layer: Projected images with different resolutions; Feature extraction layer: 3 depthwise separable convolutional layers, with the convolutional kernel sizes being 7×7, 5×5, and 3×3 respectively; Attention layer: Channel attention mechanism, adopting the squeeze-and-excitation structure; Feature aggregation layer: A combination of global average pooling and global max pooling.
[0033] Specific implementation of the collaborative acceleration device As Figure 2 shown, the specific implementation of the collaborative acceleration device includes the following main components: 5.4.1 Intelligent task dispatching controller The intelligent task dispatching controller dynamically determines the task allocation strategy according to the characteristics of the input point cloud (such as the number of points, distribution density, etc.), and its main functions include: Preprocessing and partitioning of point cloud data; Load balancing of processing units; Dynamic allocation of computing resources.
[0034] The core algorithm of the controller is as follows: Algorithm 1: Intelligent task dispatching algorithm Input: Point cloud data P, system resource status S Output: Task allocation strategy T 1. Calculate the point cloud complexity index CI = f(P), where f() is the complexity evaluation function 2. Estimate the expected load of each processing unit: L_point = g_point(CI, S) / / Expected load of point cloud stream L_proj = g_proj(CI, S) / / Expected load of projection stream L_fusion = g_fusion(CI, S) / / Expected load of fusion unit 3. Allocate computing resources according to the load: R_point = h_alloc(L_point, S) R_proj = h_alloc(L_proj, S) R_fusion = h_alloc(L_fusion, S) 4. Generate the task allocation strategy T = {R_point, R_proj, R_fusion} 5. Return T.
[0035] Among them, the calculation method of the complexity evaluation function f(P) is as follows: f(P) = w1·N + w2·Var(d) + w3·H(P) Among them: N is the number of points in the point cloud; Var(d) is the variance of the distance between points; H(P) is the information entropy of the point cloud; w1, w2, w3 are weight coefficients, which are set to 0.5, 0.3, and 0.2 respectively in this embodiment.
[0036] Heterogeneous memory hierarchy scheduling unit As Figure 5 shown, the heterogeneous memory hierarchy scheduling unit optimizes memory access according to the access pattern of the point cloud data, and the main functions include: Data block partitioning and marking; Access priority calculation; Cache hierarchy management; Predictive data loading.
[0037] The core implementation of the heterogeneous memory scheduling algorithm is as follows: Algorithm 2: Heterogeneous Memory Scheduling Algorithm Input: Set of point cloud data blocks B Output: Optimized memory mapping M 1. Calculate the access priority score APS(b) for each data block b ∈ B 2. Sort the data blocks according to APS to obtain the ordered set B' 3. Initialize the memory mapping M = {} 4. For each data block b in B': 4.1 If APS(b) > θ1: Map b to the L1 cache and update M 4.2 Otherwise, if APS(b) > θ2: Map b to the L2 cache and update M 4.3 Otherwise: Keep b in the main memory and update M 5. Return the optimized memory mapping M The calculation method of the access priority score (APS) is as follows: APS(b) = w1·I(b) + w2·F(b) + w3·T(b) Among them: I(b) represents the information entropy of data block b, and its calculation method is as follows: I(b)=-∑p(x)·log(p(x)) where p(x) is the probability distribution of feature x in the data block; F(b) represents the access frequency of data block b, which is statistically obtained from historical access records; T(b) represents the temporal correlation of data block b, and its calculation method is as follows: T(b)=exp(-λ·Δt) where Δt is the difference between the current time and the last access time, and λ is the decay factor, which is set to 0.1 in this embodiment; w1, w2, w3 are weight coefficients, which are set to 0.2, 0.5, 0.3 respectively in this embodiment.
[0038] Point cloud stream processing unit The point cloud stream processing unit adopts a dedicated point cloud graph convolution accelerator (PGA), and its main features include: Parallel graph structure processing: Simultaneously process graph convolution operations of multiple nodes; Sparse matrix optimization: Optimize storage and calculation for the sparse characteristics of the point cloud graph; Pipeline design: Implement pipeline processing of multi-layer graph convolution networks.
[0039] The core circuit modules of PGA include: k-nearest neighbor search accelerator: Adopt spatial hashing and parallel search technologies; Edge feature calculation unit: Parallelly calculate edge features; Graph convolution core: Execute graph convolution operations; Residual connection unit: Implement feature residual connection.
[0040] Projection stream processing unit The projection stream processing unit is implemented by a reconfigurable logic array (FPGA), and its main features include: Multi-view projection parallel processing: Simultaneously calculate projection images of multiple viewpoints; Depthwise separable convolution acceleration: Hardware acceleration for the depthwise separable convolution of lightweight CNN; Reconfigurable computing architecture: Dynamically adjust the allocation of computing resources according to projection images of different resolutions.
[0041] The logic design of FPGA includes: Projection transformation module: Implement the transformation from point cloud to multi-view; Depthwise separable convolution array: Implement convolution operation acceleration; Channel attention calculation unit: Implement the SE attention mechanism; Feature aggregation unit: Achieves multi-scale feature fusion.
[0042] Feature fusion unit The feature fusion unit realizes the adaptive fusion of point cloud flow and projection flow features. Its main functions include: Dynamic weight calculation: Calculates the fusion weights of the two flows according to the scene complexity; Feature alignment: Maps features of different dimensions to a unified representation space; Complementary feature enhancement: Enhances complementary information through feature interaction.
[0043] The core implementation of the fusion unit is as follows: Algorithm 3: Adaptive Feature Fusion Algorithm Input: Point cloud flow feature F_point, projection flow feature F_proj, scene complexity SC Output: Fusion feature F_fusion 1. Calculate the dynamic fusion weights: w_point = sigmoid(a·SC + b) w_proj = 1 - w_point 2. Feature alignment: F_point_aligned = align(F_point) F_proj_aligned = align(F_proj) 3. Initial fusion: F_initial = w_point·F_point_aligned + w_proj·F_proj_aligned 4. Complementary feature enhancement: Attention_matrix = softmax(F_point_aligned·F_proj_aligned^T) F_enhanced = F_initial + δ·Attention_matrix·F_initial 5. Return the fusion feature F_fusion = F_enhanced Where: SC is the scene complexity index, calculated according to the distribution characteristics of the point cloud; a, b are learnable parameters, and the initial values are set to 1.0 and 0.5 respectively; δ is the enhancement coefficient, set to 0.3 in this embodiment.
[0044] 5.4.6 Dynamic Power Management Unit The dynamic power management unit automatically adjusts the power consumption configuration according to the task requirements and system status. Its main functions include: Dynamic Voltage and Frequency Scaling (DVFS): Adjust the operating frequency and voltage of the processing unit according to the processing load; Selective unit sleep: Put some processing units into a low-power state under light load conditions; Model hierarchical quantization: Apply quantization strategies with different bit widths to different layers of the model.
[0045] The core algorithm of dynamic power management is as follows: Algorithm 4: Dynamic Power Management Algorithm Input: System load status L, real-time requirement RT, current power consumption configuration PC Output: Optimized power consumption configuration PC_new 1. Calculate the power-performance ratio metric: PPI = performance(PC) / power(PC) 2. Evaluate the current real-time performance: current_RT = estimate_RT(L, PC) 3. If current_RT < RT (meeting the real-time requirement): 3.1 Calculate the potential configuration set PC_down for reducing power consumption 3.2 Select the configuration PC_candidate with the highest PPI from PC_down 3.3 If performance(PC_candidate) ≥ RT: PC_new = PC_candidate 3.4 Otherwise: PC_new = PC 4. Otherwise (not meeting the real-time requirement): 4.1 Calculate the potential configuration set PC_up for improving performance 4.2 Select the configuration PC_candidate that meets RT and has the lowest power consumption from PC_up 4.3 PC_new = PC_candidate 5. Apply the hierarchical quantization strategy: 5.1 Adjust the quantization bit width of each layer according to PC_new 5.2 Update the quantization parameters 6. Return the optimized power consumption configuration PC_new The model hierarchical quantization strategy is as follows: Key feature extraction layer: Maintain 8-bit quantization; Intermediate processing layer: 4-8 bit adaptive quantization according to the load; Feature fusion layer: Maintain 8-bit quantization to ensure accuracy.
[0046] System embodiments Embodiment 1: Application of the lightweight object detection system in the autonomous driving scenario This embodiment is for the autonomous driving perception system and is implemented on the Jetson AGX Xavier platform. The system configuration is as follows: Hardware configuration: Main processor: Jetson AGX Xavier (8-core ARM CPU + 512-core NVIDIA Volta GPU) Coprocessor: Xilinx Kintex-7 FPGA Memory: 32GB LPDDR4x Power consumption limit: 30W Software configuration: Operating system: Ubuntu 18.04 Deep learning framework: TensorRT 7.1 FPGA development tool: Vivado 2019.2 Data processing parameters: Point cloud input: 64-line LiDAR, 10Hz scanning frequency Detection range: 0 - 100 meters Target categories: Vehicles, pedestrians, cyclists System performance metrics: Processing speed: 28 FPS Average detection accuracy (mAP): 85.3% Average power consumption: 22.5W The test results on the KITTI dataset are as Figure 6 shown. This system achieves significantly lower power consumption and higher processing speed while maintaining high detection accuracy.
[0047] Embodiment 2: Application of the lightweight object detection system in intelligent robot navigation This embodiment is for indoor navigation robots and is implemented on the Nvidia Jetson Nano platform. The system configuration is as follows: Hardware configuration: Main processor: Jetson Nano (4-core ARM CPU + 128-core NVIDIA Maxwell GPU) Coprocessor: Intel Cyclone 10 FPGA Memory: 4GB LPDDR4 Power consumption limit: 10W Software configuration: Operating system: Ubuntu18.04 Deep learning framework: TensorRT7.0 FPGA development tool: QuartusPrime19.1 Data processing parameters: Point cloud input: 16-line LiDAR, 20Hz scanning frequency Detection range: 0 - 20 meters Target categories: people, furniture, obstacles System performance metrics: Processing speed: 25FPS Average detection accuracy (mAP): 83.1% Average power consumption: 7.8W Tests on the ScanNet dataset show that this system reduces energy consumption by about 42% compared to traditional methods while maintaining high detection accuracy.
[0048] This invention covers any alternatives, modifications, equivalent methods, and solutions made within the spirit and scope of this invention. For the public to have a thorough understanding of this invention, specific details are described in detail in the following preferred embodiments of this invention, but those skilled in the art can also fully understand this invention without these detailed descriptions. Additionally, well-known methods, processes, procedures, components, and circuits, etc. are not described in detail to avoid unnecessary confusion to the essence of this invention.
[0049] The above are only the preferred embodiments of this invention. It should be noted that for those of ordinary skill in the art, without departing from the principle of this invention, several improvements and refinements can be made, and these improvements and refinements should also be regarded as the protection scope of this invention.
Claims
1. A real-time point cloud object detection system based on a lightweight model, characterized in that, Including: A point cloud-projection dual-stream architecture model, which includes a point cloud stream processing part and a projection stream processing part; The point cloud stream processing part includes a point cloud feature extraction unit and a graph convolutional edge attenuation module, which are used for feature extraction and lightweight processing of point cloud data; The projection stream processing part includes a multi-view projection unit and a lightweight CNN processing unit, which are used to project the point cloud into multiple two-dimensional views and extract complementary features; An adaptive feature fusion unit, which is used to automatically adjust the weights of the point cloud stream and projection stream features according to the scene complexity and generate fused features.
2. The system according to claim 1, wherein The graph convolutional edge attenuation module realizes adaptive weight assignment of features by calculating the edge importance score, and the edge importance score is calculated by the following formula: EIS(p)=α·S(p)+β·D(p)+γ·V(p) Where, S(p) represents the shape feature score of point p, D(p) represents the density change score of point p, V(p) represents the visibility score of point p, and α, β, γ are adaptive weight coefficients.
3. The system according to claim 1, characterized in that, The multi-view projection unit adopts a multi-resolution projection strategy, including: Using high-resolution projection for the close-range area (0-20m); Using medium-resolution projection for the medium-range area (20-50m); Using low-resolution projection for the long-range area (>50m).
4. The system according to claim 1, wherein The adaptive feature fusion unit includes: A dynamic weight calculation module, which is used to calculate the fusion weights of the point cloud stream and the projection stream according to the scene complexity; A feature alignment module, which is used to map features of different dimensions to a unified representation space; A complementary feature enhancement module, which is used to enhance complementary information through feature interaction.
5. A collaborative acceleration device for point cloud object detection, characterized in that, Including: An intelligent task dispatching controller, which is used to dynamically allocate computing resources according to the point cloud characteristics; A heterogeneous memory hierarchy scheduling unit, which is used to optimize memory access according to the point cloud data access pattern; A heterogeneous parallel acceleration unit, including a point cloud stream processing unit, a projection stream processing unit, and a feature fusion unit; A dynamic power management unit, which is used to automatically adjust the power consumption configuration according to the task requirements and system status.
6. The device according to claim 5, characterized in that The heterogeneous memory hierarchy scheduling unit realizes hierarchical cache management by calculating the access priority score of the data block, and the access priority score is calculated by the following formula: APS(b)=w1·I(b)+w2·F(b)+w3·T(b) Where, I(b) represents the information entropy of data block b, F(b) represents the access frequency of data block b, T(b) represents the temporal correlation of data block b, and w1, w2, w3 are weight coefficients.
7. The device according to claim 5, characterized in that The point cloud stream processing unit is implemented by a dedicated point cloud graph convolutional accelerator, including: A k-nearest neighbor search accelerator, which adopts spatial hashing and parallel search techniques; An edge feature calculation unit, which is used to calculate edge features in parallel; A graph convolutional core, which is used to execute graph convolutional operations; A residual connection unit, which is used to realize feature residual connection.
8. The device according to claim 5, characterized in that, The projection stream processing unit is implemented by a reconfigurable logic array, including: A projection transformation module, which is used to realize the transformation from point cloud to multi-view; A depthwise separable convolution array, which is used to accelerate convolution operations; A channel attention calculation unit, which is used to realize the SE attention mechanism; A feature aggregation unit for implementing multi-scale feature fusion.
9. The device according to claim 5, characterized in that The dynamic power management unit includes: A dynamic voltage and frequency adjustment module for adjusting the operating frequency and voltage of the processing unit according to the processing load; A selective unit sleep module for putting some processing units into a low-power state under light load conditions; A model hierarchical quantization module for applying quantization strategies with different bit widths to different layers of the model.
10. The system and device according to claims 1 and 5, characterized in that, The system and the device are combined and applied to at least one application scenario among an autonomous driving perception system, intelligent robot navigation, drone scene understanding, or augmented reality spatial mapping.