Autonomous Driving Compute Scheduling for Mixed-Criticality Workloads
Overview of Technical Issues:
The compute scheduling unit insufficiently isolates and prioritizes mixed-criticality workloads, allowing lower-priority tasks to interfere with safety-critical computations through resource contention, resulting in unpredictable execution timing that violates the deterministic performance guarantees required for autonomous driving safety functions; the goal is to achieve guaranteed worst-case execution time for critical tasks while efficiently utilizing compute resources for all workload types.
Solution directions generated for this problem
Problem Direction 1 :
ImproveResource isolation strength
VSConstraintOverall compute resource utilization
Inspiration 1 : Cross-domain reference
Application Principle: #15 Dynamics
Cross-domain applicability
Resource allocation method, identification method, base station, mobile station, and program
Innovative Solution Refine solution
Adaptive criticality-aware resource elasticity with real-time boundary reconfiguration
Elastic isolation via workload-driven reconfiguration
How to solve :
- Implement hardware-assisted dynamic partition resizing triggered by criticality state transitions—when ASIL-D perception tasks enter idle state (vehicle stationary, sensor fusion paused), release reserved CPU cores and cache ways to QM workloads within 50μs via memory management unit remapping, reclaim with guaranteed 100μs latency upon critical task resumption detected by interrupt controller
- Deploy criticality-aware resource broker with three-tier elasticity: Tier-1 (ASIL-D) maintains 30% hard reservation with microsecond-level reclaim priority, Tier-2 (ASIL-B) holds 25% soft reservation borrowable to Tier-3 (QM) when idle >5ms, all transitions enforced by hardware timer and cache coloring mechanism ensuring zero interference during handoff
- Establish runtime workload profiling monitoring critical task duty cycle every 20ms frame—if ASIL-D utilization <40% for three consecutive frames, expand QM allocation from 45% to 60%, shrink back when critical workload exceeds 35% threshold, maintaining isolation via temporal guards (10μs buffer zones) between partition switches
Expected Effect : Utilization maintained at 72-75%, WCET compliance 100%, reconfiguration latency <100μs
Risk Control :
- partition switch timing jitter exceeding 100μs bound
- cache way remapping causing transient interference
- workload prediction error triggering false expansions
Problem Direction 2 :
ImproveScheduling determinism
VSConstraintSystem implementation complexity
Inspiration 1 : Cross-domain reference
Application Principle: #26 Copying
Cross-domain applicability
Method and apparatus for transmitting/receiving data on multiple carriers in mobile communication system
Innovative Solution Refine solution
Offline pre-computed schedule table with runtime table-driven dispatch for deterministic mixed-criticality scheduling
Replace runtime scheduler with offline schedule tables to eliminate runtime complexity
How to solve :
- Generate exhaustive schedule tables offline covering all safety-critical task combinations (perception, planning, control variants) with pre-computed WCET-optimal slot assignments, store in ROM lookup tables indexed by vehicle state and active task set
- Implement lightweight table-driven dispatcher at runtime that reads current vehicle state (driving/parked/emergency), indexes corresponding schedule table, and executes pre-determined task dispatch sequence without runtime arbitration logic—dispatcher complexity <500 SLOC vs 15000+ SLOC for dynamic scheduler
- Deploy dual-layer execution model: ASIL-D tasks follow rigid table entries with deterministic 10ms major cycle and <5μs jitter, QM tasks opportunistically fill table-defined idle slots using simple FIFO queue, maintaining 70-75% utilization while guaranteeing critical task isolation
Expected Effect : WCET guarantee 100% with <5μs jitter; certification effort reduced 60%; implementation complexity from 15000 to <2000 SLOC; utilization maintained 70-75%
Risk Control :
- offline table generation completeness verification
- ROM storage capacity for all scenario tables
- table update latency during mode transitions
Problem Direction 3 :
ImproveExecution time predictability
VSConstraintOverall compute resource utilization
Problem Direction 4 :
ImproveScheduling determinism
VSConstraintMust not deteriorate
