Machine Learning Inference Optimization for Lidar Point Cloud Processing
Overview of Technical Issues:
The inference computing unit processes lidar point cloud data with insufficient speed, unable to meet real-time requirements for downstream applications such as autonomous navigation or obstacle detection. This functional insufficiency is compounded by inefficient memory access patterns caused by the irregular spatial structure of point cloud data, creating transmission bottlenecks between memory and computing units. The goal is to optimize the inference pipeline to achieve real-time processing throughput while operating within the computational and power constraints of the target hardware platform.
Solution directions generated for this problem
Problem Direction 1 :
ImproveProcessing throughput rate
VSConstraintPower consumption
Inspiration 1 : Cross-domain reference
Application Principle: #19 Periodic action
Cross-domain applicability
Time division duplex (TDD) uplink downlink (UL-DL) reconfiguration
Innovative Solution Refine solution
Burst-mode inference engine with thermal-aware duty cycling for lidar processing
Burst-mode inference with thermal recovery
How to solve :
- Operate inference engine in burst cycles: process 4 frames at peak frequency (1.8 GHz) in 50ms active phase, then idle 83ms for thermal dissipation, achieving 30 fps average throughput
- Implement thermal-aware scheduler monitoring junction temperature via on-chip sensor (polling every 10ms): trigger burst when T<75°C, extend idle when T>80°C, maintaining 18-22W average power within thermal envelope
- Deploy frame buffering with priority queuing: 6-frame FIFO buffer absorbs input during idle phase, prioritize obstacle-rich frames (point density >5000/m²) for immediate burst processing, defer sparse background frames to next cycle
Expected Effect : 30 fps sustained, 19W average power, no thermal throttling
Risk Control :
- burst timing synchronization with lidar frame rate
- temperature sensor calibration drift over lifetime
- frame buffer overflow under sustained high-density scenes
Problem Direction 2 :
ImproveProcessing throughput rate
VSConstraintAlgorithm implementation complexity
Inspiration 1 : Cross-domain reference
Application Principle: #1 Segmentation
Cross-domain applicability
System for recursive recombination of streaming interactive video
Innovative Solution Refine solution
Spatial-tile point cloud partitioning for parallel inference acceleration
Divide point cloud into fixed spatial tiles for independent parallel processing
How to solve :
- Partition incoming lidar point cloud into fixed 10m×10m×5m spatial grid tiles using simple modulo indexing (x_tile = floor(x/10), y_tile = floor(y/10))
- each tile processed independently by separate inference threads with local 2MB cache, eliminating global data reorganization overhead
- Assign one compute thread per tile with dedicated memory buffer
- tiles processed in parallel across 4-8 CPU cores or GPU streaming multiprocessors, achieving 30+ fps aggregate throughput through spatial parallelism rather than complex sequential optimization
- Implement lightweight tile boundary handling via 0.5m overlap zones between adjacent tiles
- objects spanning boundaries detected in both tiles then merged via simple centroid-distance clustering (threshold 0.3m), adding only 2-3ms overhead versus 80-100ms for global restructuring
Expected Effect : Throughput 30-35 fps; complexity overhead <15%; latency 28-32ms per frame; bandwidth utilization 55-65%
Risk Control :
- tile size mismatch with object distribution causing load imbalance
- boundary object duplication rate exceeding 10%
- thread synchronization overhead above 5ms
Problem Direction 3 :
ImproveMemory bandwidth utilization efficiency
VSConstraintAlgorithm implementation complexity
Inspiration 1 : Cross-domain reference
Application Principle: #35 Parameter changes
Cross-domain applicability
Route switching device and data cashing method thereof
Innovative Solution Refine solution
Adaptive voxel resolution point cloud encoding for bandwidth-efficient inference
Dynamically adjust voxel grid resolution based on spatial density
How to solve :
- Implement multi-resolution voxel encoding: dense regions (>500 pts/m³) use 0.1m voxels, sparse regions (<100 pts/m³) use 0.5m voxels, automatically selected per 10m×10m tile during sensor readout
- Encode each voxel as fixed 16-byte structure (occupancy flag + mean XYZ + intensity), enabling sequential memory access with 128-byte cache-line alignment for 8 voxels per fetch
- Configure lidar driver firmware to output pre-voxelized data stream with resolution metadata in tile header, eliminating runtime conversion overhead and complex prefetching logic.
Expected Effect : Bandwidth utilization 68-75%; latency <30ms; code complexity +12%
Risk Control :
- voxel resolution threshold calibration per scene type
- cache-line alignment enforcement across hardware platforms
- firmware modification compatibility with existing lidar models
Problem Direction 4 :
ImproveInference latency duration
VSConstraintPower consumption
Inspiration 1 : Cross-domain reference
Application Principle: #19 Periodic action
Cross-domain applicability
Methods and apparatus for improved low energy data communications
Innovative Solution Refine solution
Burst-mode inference engine with thermal-aware duty cycling for lidar processing
Burst-mode inference with thermal recovery
How to solve :
- Execute inference in high-speed burst windows at peak frequency (1.8 GHz) for 18ms per frame, achieving <33ms latency target, then enter thermal recovery idle state for 82ms
- Implement thermal-aware duty cycling controller monitoring junction temperature via on-chip sensor (sample rate 10 Hz), dynamically adjusting burst duration (15-22ms) and idle period (78-85ms) to maintain Tj <85°C within 15-25W envelope
- Preload 8MB network weights into on-chip SRAM during sensor readout phase (overlapped, zero added latency), enabling inference to run entirely from low-latency SRAM during burst, eliminating DRAM access bottlenecks and power spikes
Expected Effect : Latency <20ms per burst, avg power 18W, 30fps sustained
Risk Control :
- thermal sensor calibration drift
- SRAM preload timing synchronization failure
- burst scheduling jitter under varying point cloud density
Problem Direction 5 :
ImproveInference latency duration
VSConstraintAlgorithm implementation complexity
Inspiration 1 : Cross-domain reference
Application Principle: #10 Preliminary action
Cross-domain applicability
Stream-based accelerator processing of computational graphs
Innovative Solution Refine solution
Lidar-firmware spatial pre-sorting for zero-latency point cloud restructuring
Embed spatial sorting into lidar firmware
How to solve :
- Configure lidar firmware to output points pre-sorted by Morton Z-order curve (space-filling curve mapping 3D→1D) during sensor acquisition, eliminating 80-100ms runtime restructuring overhead
- Partition output into cache-line-aligned 64-byte blocks grouped by 10m×10m spatial grid cells, enabling inference engine to read sequentially at 65-75% bandwidth utilization without runtime reorganization
- Implement double-buffered DMA transfer (ping-pong buffers) where firmware writes sorted data to buffer A while inference reads buffer B, achieving 25ms end-to-end latency with simple sequential memory access patterns
Expected Effect : Latency reduced to 25ms (83% improvement); bandwidth utilization 65-75% vs 30-40% baseline; zero algorithm complexity increase
Risk Control :
- firmware modification compatibility across lidar models
- Morton code computation overhead in sensor MCU
- sorting stability under high point density scenes
