Autonomous Driving Data Pipeline Bottlenecks for Training at Scale

Overview of Technical Issues:

The data preprocessing module provides insufficient transformation and conversion capacity for raw autonomous driving sensor data, causing accumulation backlogs that starve the training computation cluster and create GPU idle time; the goal is to achieve continuous high-throughput data flow that fully utilizes training resources and enables rapid model iteration at scale.

Solution directions generated for this problem

Problem Direction 1 :

ImproveData transformation throughput rate
VS
ConstraintSystem computational complexity

Inspiration 1 : Cross-domain reference

Application Principle: #1 Segmentation
Cross-domain applicability Assess applicability
Uplink coordinated multi-point
Innovative Solution Refine solution

Frame-level autonomous preprocessing workers with embedded sensor fusion

Divide preprocessing into independent frame-level workers handling complete sensor fusion end-to-end
How to solve :
  • Deploy containerized preprocessing workers where each instance processes one complete frame (lidar + camera + radar) independently from ingestion to output, eliminating centralized orchestration
  • each worker executes a fixed pipeline: raw data decode → sensor calibration → coordinate transformation → format conversion → output serialization, with processing time target <30ms per frame
  • provision 80 stateless worker instances with simple round-robin frame assignment at ingestion layer, achieving 2TB/hour aggregate throughput (25GB/hour per worker) without load balancing logic
Expected Effect : Throughput 4× to 2TB/hour; GPU idle <5%; worker count scales linearly
Risk Control :
  • frame assignment skew under variable data rates
  • worker failure detection latency
  • output ordering consistency

Problem Direction 2 :

ImprovePreprocessing processing speed
VS
ConstraintResource scheduling difficulty

Inspiration 1 : Cross-domain reference

Application Principle: #28 Mechanics substitution
Cross-domain applicability Assess applicability
Hardware for performing a database operation
Innovative Solution Refine solution

FPGA-accelerated fixed-pipeline preprocessing for autonomous driving data

Replace CPU-based dynamic scheduling with FPGA hardware pipeline
How to solve :
  • Deploy FPGA preprocessing accelerators with fixed-function pipelines—each FPGA handles dedicated sensor type (lidar/camera/radar) with hardwired transformation logic, eliminating runtime scheduling decisions
  • Implement static resource partitioning: allocate 4 FPGA cards with fixed 500GB/hour capacity each, totaling 2TB/hour throughput without load balancing—each card processes sequential frame batches via deterministic routing based on sensor ID modulo 4
  • Embed hardware transformation kernels for point cloud filtering (voxel downsampling at 0.1m resolution), image normalization (fixed 8-bit to float32 conversion), and radar clustering (DBSCAN with ε=2m, minPts=5) directly in FPGA fabric, achieving <25ms per frame with zero software intervention
  • Quality control: FPGA output validated via hardware CRC-32 checksums (error rate <10⁻⁹), frame completeness verified by embedded counters, transformation accuracy spot-checked every 1000th frame against golden reference (tolerance ±0.01% for numerical precision)
  • Implementation: integrate Xilinx Alveo U280 cards via PCIe Gen4 x16, deploy pre-compiled bitstreams for each sensor pipeline, configure input DMA buffers (512MB per card) and output queues (1TB NVMe cache), monitor via FPGA health registers (temperature <85°C, utilization 70-90%)
Expected Effect : Frame latency reduced to 22ms; throughput 2.1TB/hour; GPU idle time <4%; zero dynamic scheduling overhead
Risk Control :
  • FPGA bitstream development cycle 8-12 weeks
  • hardware failure requires cold spare swap
  • sensor format changes need bitstream recompilation

Problem Direction 3 :

ImprovePipeline continuous flow capacity
VS
ConstraintResource scheduling difficulty

Inspiration 1 : Cross-domain reference

Application Principle: #11 Beforehand cushioning
Cross-domain applicability Assess applicability
Aviation crew scheduling method and device, computer equipment and storage medium
Innovative Solution Refine solution

Fatigue-aware static preprocessing worker allocation with pre-provisioned capacity buffer

Pre-provision worker capacity buffer to absorb throughput fluctuations
How to solve :
  • Deploy 20% over-capacity preprocessing workers (2.4TB/hour vs 2TB/hour demand) with static allocation—each worker handles fixed sensor type and frame range without dynamic rebalancing
  • Implement fatigue threshold monitoring per worker: track cumulative processing time and frame backlog depth
  • when worker fatigue index exceeds 0.75 (processing time >90% of target), buffer workers automatically absorb overflow without coordination
  • Establish constant-rate output mode: each worker targets fixed 150GB/hour throughput with ±5% tolerance, verified via per-worker output counters sampled every 60 seconds
  • acceptance criteria requires 95% of samples within tolerance over 24-hour window
Expected Effect : GPU idle time ≤5%, zero dynamic scheduling
Risk Control :
  • buffer capacity under-utilization during normal operation
  • fatigue threshold calibration inaccuracy
  • worker output rate drift over time
Patsnap Eureka Solution