How to Select Autonomous Driving Processors for Multi-Task Inference
Overview of Technical Issues:
The computing processor's parallel processing capability is insufficient to simultaneously execute multiple inference tasks (object detection, path planning, sensor fusion) within the required real-time window, causing decision latency that compromises autonomous driving safety; the goal is to select a processor that balances multi-task computing throughput, power efficiency, and thermal constraints while meeting strict real-time performance requirements.
Solution directions generated for this problem
Problem Direction 1 :
ImproveParallel computational throughput
VSConstraintThermal dissipation requirement
Inspiration 1 : Cross-domain reference
Application Principle: #35 Parameter changes
Cross-domain applicability
Technologies for predictively managing heat generation in a datacenter
Innovative Solution Refine solution
Adaptive workload-thermal state machine for parallel inference scaling
Workload state machine with thermal feedback
How to solve :
- Implement a three-state thermal governor: State A (≤90W) runs all 5+ tasks at nominal frequency
- State B (90-110W) reduces prediction task precision to INT8 and drops localization update rate to 5Hz
- State C (>110W) suspends non-critical prediction tasks entirely, maintaining detection and planning at full throughput
- Embed real-time thermal sensors (±1°C accuracy, 10ms sampling) at processor hotspots with hardware interrupt triggering state transitions within 20ms, preventing thermal runaway before throttling occurs
- Deploy workload pre-emption scheduler with task priority queue: safety-critical detection (priority 1, always active), path planning (priority 2, 50ms budget), sensor fusion (priority 3, 30ms budget), prediction (priority 4, droppable), ensuring sub-100ms latency for top-3 tasks under all thermal states
Expected Effect : 5+ tasks at 115W avg, latency <100ms, zero throttling events
Risk Control :
- state transition hysteresis causing oscillation
- sensor calibration drift over temperature cycles
- task priority conflicts during edge cases
Problem Direction 2 :
ImprovePower efficiency
VSConstraintSystem integration complexity
Inspiration 1 : Cross-domain reference
Application Principle: #6 Universality
Cross-domain applicability
Systems and methods for improving efficiencies of a memory system
Innovative Solution Refine solution
Unified heterogeneous SoC with adaptive task routing for multi-domain inference
Unified SoC handles all inference tasks
How to solve :
- Deploy a single-die heterogeneous SoC (e.g., NVIDIA Orin or Tesla FSD architecture) integrating CPU clusters, GPU cores, and dedicated neural accelerators on one substrate with unified memory architecture, eliminating inter-chip PCIe transfers and separate power domains
- Implement hardware task router that dynamically assigns detection to GPU (150+ GOPS/W), planning to CPU (80 GOPS/W), and fusion to DLA accelerators based on real-time workload profiling, maintaining aggregate efficiency ≥150 GOPS/W within 150W budget
- Use shared LPDDR5 memory pool (64GB, 204GB/s bandwidth) accessible by all compute domains via coherent interconnect, reducing data movement energy by 60% compared to discrete multi-chip solutions and cutting BOM cost by 35%
Expected Effect : Power efficiency 150+ GOPS/W; integration complexity -40% vs multi-chip; latency <100ms for 5+ tasks; BOM cost -35%
Risk Control :
- SoC vendor lock-in and supply chain risk
- thermal hotspot management within single die requiring advanced packaging
- software stack complexity for heterogeneous scheduling
Problem Direction 3 :
ImproveReal-time response latency
VSConstraintThermal dissipation requirement
Inspiration 1 : Cross-domain reference
Application Principle: #19 Periodic action
Cross-domain applicability
Light curing dental system
Innovative Solution Refine solution
Pulsed computation with thermal recovery intervals for latency control
Execute tasks in time-sliced bursts with cooling gaps
How to solve :
- Partition the 100ms cycle into pulsed computation windows: 30ms object detection at peak frequency (2.5GHz), 10ms passive cooling, 25ms path planning at nominal frequency (1.8GHz), 10ms cooling, 25ms sensor fusion at nominal frequency
- total latency 100ms with peak thermal load ≤115W per burst, average ≤95W across cycle
- Implement thermal-aware scheduler that monitors junction temperature in real-time (sampling at 1kHz via on-die sensors) and triggers computation bursts only when temperature drops below 85°C threshold, ensuring no sustained overheating
- Pre-load inference models into on-chip SRAM (≥32MB capacity) during idle intervals to eliminate memory access latency during computation bursts, enabling full GPU utilization (≥90%) within each 25-30ms window without thermal penalty from prolonged DRAM transfers
Expected Effect : Latency reduced to 95-100ms; peak thermal load 115W vs 200W continuous; no throttling events; throughput 5+ tasks sustained
Risk Control :
- temperature sensor calibration drift ±3°C
- task scheduling jitter ±5ms under interrupt load
- SRAM capacity insufficient for model size growth
Problem Direction 4 :
ImproveReal-time response latency
VSConstraintSystem integration complexity
Inspiration 1 : Cross-domain reference
Application Principle: #24 Intermediary
Cross-domain applicability
Integrated insulin delivery system with continuous glucose sensor
Innovative Solution Refine solution
Lightweight FPGA pre-processor gateway for sensor data aggregation and latency buffering
FPGA gateway aggregates multi-sensor streams
How to solve :
- Deploy a dedicated FPGA pre-processor module between sensors (camera, lidar, radar) and main processor to perform hardware-accelerated time-stamping, data alignment, and format conversion — reducing main processor interface count from 8+ to 1 unified stream
- Configure FPGA with deterministic 10-15ms fixed latency pipeline: sensor data arrives asynchronously, FPGA synchronizes timestamps to ±1μs precision, packages into standardized message frames, and forwards via single PCIe Gen4 x4 link at 64 Gbps to main processor
- Implement plug-and-play sensor abstraction layer in FPGA firmware: each sensor type (camera/lidar/radar) maps to predefined data slots, allowing sensor upgrades without main processor software changes — validation reduced to FPGA firmware update (2-4 weeks) vs full system re-integration (6+ months)
Expected Effect : End-to-end latency 85-95ms; integration complexity -60%; validation cycle 8-10 weeks
Risk Control :
- FPGA firmware timing closure failure under multi-sensor peak load
- PCIe bandwidth saturation during simultaneous 8-camera 30fps streaming
- sensor hot-swap causing transient data corruption
Problem Direction 5 :
ImproveParallel computational throughput
VSConstraintSystem integration complexity
Inspiration 1 : Cross-domain reference
Application Principle: #6 Universality
Cross-domain applicability
Technologies for big data analytics accelerator
Innovative Solution Refine solution
Unified heterogeneous SoC with domain-specific accelerator clusters for multi-task inference
Single-die SoC consolidates all functions
How to solve :
- Deploy a unified heterogeneous SoC architecture integrating CPU (4-core ARM A78, 2.2GHz), GPU (1024 CUDA cores), and dedicated DLA accelerator clusters (2×INT8 engines, 40 TOPS each) on a single 7nm die with shared coherent memory fabric (128GB/s bandwidth) to eliminate inter-chip communication overhead
- Assign task-specific domains: DLA cluster-A handles object detection (YOLOv5, 25ms), DLA cluster-B executes sensor fusion (Kalman filter, 15ms), GPU processes path planning (A-star, 30ms), CPU manages localization and prediction (20ms), all accessing unified 16GB LPDDR5 memory to avoid PCIe data transfers
- Implement hardware task scheduler with priority queuing (safety-critical detection priority-1, prediction priority-3) and automatic load balancing across accelerator clusters, maintaining total latency under 100ms for all 5+ concurrent tasks while keeping integration to standard JEDEC interfaces (no custom interconnects)
Expected Effect : Throughput 5+ tasks within 95ms; integration complexity -60% vs multi-chip; BOM cost +18% vs baseline; validation cycle 3 months
Risk Control :
- DLA compiler optimization for model compatibility
- memory bandwidth contention under peak load
- thermal density in 7nm process node
